You want to turn a meeting recording, an interview or a voice memo into text. But most free transcription services upload your audio to their servers — not ideal when the recording is confidential or personal.
This guide shows you how to transcribe audio with AI right in your browser, without sending it anywhere, using the free Audio Transcription tool. No sign-up needed, and you can save the result as text or as subtitle files (SRT/VTT).
I couldn’t find a transcription tool that was both free and kept the audio on my own device, so I built one. For confidential meetings and personal notes, keeping everything local just feels safer.
How AI transcription works in your browser
Most transcription services send your audio file to a server, where an AI converts it to text and sends the result back. It’s convenient, but it means handing your audio to the service.
This tool runs Whisper, the speech recognition AI released by OpenAI, directly inside your browser. The AI model is downloaded the first time you use it, but your audio never leaves your computer or phone. Because it uses your device’s processing power, it’s slower than a server-based service — the trade-off is that nobody else ever receives your recording.
The tool doesn’t connect to OpenAI’s paid service (the API). It only uses the Whisper model files that OpenAI has made freely available, so neither OpenAI nor FlowPocket ever receives your audio or the transcript.
What this tool does
- Transcribes audio files (MP3, WAV, M4A and more)
- Records from your microphone and transcribes the recording
- Recognizes English or Japanese, or detects the language automatically
- Saves results as plain text, timestamped text, subtitles (SRT/VTT) or Markdown
How to use it
Step 1. Choose your audio
Open Audio Transcription and choose an audio file. To record on the spot, click Record with microphone. If your file won’t load, convert it to MP3 with the Audio Format Converter and try again.
Step 2. Choose accuracy and language, then start
For meetings and interviews, choose More accurate. If you only need a rough idea of what’s in a short memo, Faster is fine. Pick the language and click Start transcription.
The first time, downloading the AI model (tens of megabytes) takes a moment. After that, the model is stored in your browser and starts right away.
For meetings and interviews, choose “More accurate”
Step 3. Review the result
When processing finishes, the transcript appears. On a computer, expect it to take roughly 0.5–1.5 times the length of the audio. Names and numbers are easy to mishear, so give the text a quick read.
Transcript without speaker labels. Reviewing the entire text for typos
Step 4. Pick a format and save
Choose an Output format to suit what you’ll do next. Switching formats doesn’t require transcribing again.
- Plain text: for notes and meeting minute drafts
- Timestamped text: to find the spot in the audio again later
- Subtitle file (SRT/VTT): to load as captions for a video
- Markdown: to paste into Notion or other document tools
Use Copy result to copy it to your clipboard, or Save as file to download it.
Tips for more accurate results
AI transcription depends heavily on recording quality. A few simple habits will save you a lot of correcting:
- Keep the mic close to the speaker: even placing your phone in the middle of the table helps pick up voices from across the room
- Reduce background noise: air conditioning, fans and keyboard clatter all interfere with recognition
- Avoid talking over each other: overlapping speech is rarely recognized correctly
- Trim long silences: recordings with long pauses process faster after you shorten them with the Silence Trimmer
For recordings over an hour, I recommend using a computer. It works on a phone, but the phone may heat up or stop partway through.
When this comes in handy
- Meeting records: turn recordings of internal meetings into a draft of the minutes
- Interviews and research: get text you can quote from for articles and reports
- Video subtitles: export SRT and load it into your video editor or YouTube
- Voice memos: make spoken ideas searchable as text
To turn a transcript into finished meeting minutes, try the Meeting Minutes Builder.
Choosing a transcription tool for work
Meeting recordings often contain confidential information or personal details about clients and customers. When you transcribe for work, it’s important to check not just accuracy and convenience, but also where your audio and text are sent and how they’re handled.
Watch how free AI tools handle your data
Some free transcription services and AI tools state in their terms that audio or text you submit may be used to improve the service or train AI models. Because recordings could be stored by the service or used for training, many companies ban or restrict free AI tools for work to prevent data leaks.
That’s why businesses usually rely on paid services they have a contract with, or on transcription features built into tools they already use.
What paid transcription tools offer
Business-grade paid tools — and AI features included in business plans — often guarantee the following in their contracts or specifications, depending on the service:
- Your audio and text are not used to train AI models
- Encryption of data in transit and at rest
- Third-party security certifications such as SOC 2 or ISO/IEC 27001
Many also offer features that make minutes easier, such as speaker identification, AI summaries and action items, and custom vocabulary for industry or company terms.
For online meetings, use your meeting app’s transcription
For online meetings, the easiest option is usually the transcription feature built into your meeting app. There’s no recording to prepare, and speaker names are captured too.
- Microsoft Teams: meeting transcripts, plus meeting summaries and action items with a paid Microsoft 365 Copilot license
- Zoom and Google Meet: both offer transcription features
Availability depends on your plan and on settings chosen by your organization’s administrator.
Built-in features on your phone and computer
Transcription is also built into devices you already use:
- iPhone: the built-in Voice Memos app can transcribe recordings (on recent versions of iOS)
- Android: some phones, such as Google Pixel with its Recorder app, can transcribe as you record (varies by device)
- Windows: Voice typing (Windows key + H) turns what you say into the microphone into text in real time. It can’t transcribe existing audio files
Even built-in features may process some data in the cloud. For example, Google says Pixel’s Recorder transcribes on the device while recording, but audio may be processed on Google’s servers when you re-run a transcription later. For confidential material, check how each feature works.
Which one should you use?
If you’re transcribing for work, first check that the tool is allowed under your company’s security rules — for example, that your data won’t be used for training or sent outside. Then choose based on the job:
- Online meetings: your company’s meeting app transcription
- Automatic speaker labels and summaries: a paid transcription tool your company subscribes to
- Quick notes and short in-person recordings: your phone’s or computer’s built-in features
- Transcribing an existing recording without sending it anywhere: this tool
Because this tool processes audio only on your device, you don’t have to worry about where your data goes. In exchange, it doesn’t identify speakers or write summaries. Think of it as a trade-off between privacy and simplicity on one side and features on the other, and choose what fits the task.
If you plan to use it on a work computer, the surest route is to check with your IT department first. The verification steps in the FAQ below can help with that conversation.
FAQ
What is Whisper?
Whisper is a speech recognition AI developed by OpenAI and released for free as open source. It supports many languages with high accuracy and powers many transcription apps.
Could my audio be used to train AI?
No. The tool downloads the Whisper model files to your browser once, then converts speech to text entirely on your computer or phone. Nothing sends your audio or the transcript anywhere, so it never reaches OpenAI or FlowPocket and can’t be used for training.
The downloaded model is fixed: it only converts speech to text and doesn’t learn or change as you use it.
Can I check for myself that nothing is sent?
Yes. After transcribing once so the model is loaded, keep the page open (don’t reload it), disconnect your computer from the internet, and transcribe again with the same settings. If it still finishes, it’s running entirely on your device. You can also open your browser’s developer tools (F12) and watch the Network tab to confirm that no audio is sent.
Can it tell speakers apart?
No. The tool converts speech to text but doesn’t detect who said what. If you need speaker labels, add names by hand while reviewing the result.
How is this different from paid transcription services?
Paid services run on powerful servers, so they’re faster and offer more features, such as speaker identification and summaries. This tool keeps things simple in exchange for being free and keeping your audio on your device. See “Choosing a transcription tool for work” above for how to decide.
Things to watch out for
- Always review the result: names, technical terms and numbers are especially easy to mishear
- The first use downloads data: the AI model is tens of megabytes, so use Wi-Fi on a phone
- Get consent before recording: let meeting and interview participants know you’re recording and transcribing
Summary
With FlowPocket’s Audio Transcription, you can transcribe audio with AI for free, without sending it anywhere. Save the result as text or subtitles and use it for minutes or video captions.
For work, check your company’s rules first, and combine this tool with your meeting app’s and your devices’ built-in features.
Start with a short voice memo to get a feel for the accuracy and processing time.

