How to transcribe audio to text
Five methods compared, with speed and accuracy measured for two Whisper models on 1 October 2026.
Free toolAudio to TextTranscribe a recording or a video into text with timestamps. Whisper runs on your device, and the file is never uploaded.Open the transcriberThe short answer
Open the recording in a transcriber that accepts files, check the result against the audio, and save it with timestamps. Which transcriber depends on two questions: may the recording leave your computer, and do you need to know who is speaking?
| If you | Use |
|---|---|
| Already pay for Microsoft 365 and transcribe a few hours a month | Word for the web |
| Need speaker labels and can upload the file | A cloud service such as TurboScribe or Otter |
| Cannot upload the recording | Whisper on your own computer, in a desktop app or a browser |
| Want to speak and have it typed as you go | Dictation, not file transcription |
Method 1: Word for the web
Word’s Transcribe button takes an MP3, WAV, M4A or MP4 and returns a transcript with timestamps and speaker labels. It is in Word for the web only, and it needs a Microsoft 365 subscription.
- Open a document at office.com and choose Home, then Dictate, then Transcribe.
- Choose Upload audio and pick the file. Microsoft processes it on its servers.
- Edit the sections in the side panel, then add them to the document.
The limits are 300 minutes of uploaded audio a month and 300 MB per file, according to Microsoft’s support page. There is no SRT export.
Method 2: Google Docs voice typing
Google Docs cannot transcribe a file. Voice typing listens to the microphone in Chrome, and there is nowhere to open a recording.
The common workaround is to play the file through a speaker next to the microphone. It loses words to room echo and volume, needs the whole recording played in real time, and produces no timestamps. It is a reasonable way to dictate, and a poor way to transcribe.
Method 3: a cloud transcription service
Upload the file, wait a few minutes, download the transcript. Cloud services run large models on dedicated hardware, so they are fast, and most of them label speakers.
| Service | Free tier, as stated on 1 October 2026 |
|---|---|
| TurboScribe | 3 files a day, 30 minutes each. Unlimited from $10 a month billed yearly |
| Wave | 5 files a day, 200 MB each, no account |
| Otter | Account required. English, Spanish and French |
The cost is that the recording leaves your computer. Read the retention terms before you upload anything you were not given permission to share.
Method 4: a desktop Whisper app
Whisper is the open speech model OpenAI released in 2022. Desktop apps run it on your own machine, so the file never leaves it. MacWhisper is the best known on a Mac, and Buzz is a free open source option for Windows, Mac and Linux.
They can run the large Whisper models, which are more accurate than anything a browser can comfortably download. They need installing, which a locked-down work laptop may not allow.
Method 5: Whisper in the browser
The audio to text tool runs Whisper inside the browser tab. Nothing is installed and nothing is uploaded. The model downloads once and is cached.
- Open the recording or video in the tool.
- Choose Standard or Accurate, and name the language if you know it.
- Press Transcribe, correct the lines, then save TXT, SRT or VTT.
It does not label speakers, and it is slower than a cloud service. It suits private recordings and anyone who would otherwise hit a daily limit.
How fast and how accurate Whisper is in a browser
We transcribed the same 8 minute 36 second reading with both models the tool offers, and counted every word that came out wrong. Measured on 1 October 2026.
| Model | Download | Time for 8:36 of speech | Word error rate |
|---|---|---|---|
| Standard (Whisper base) | 142 MB | 38 s | 6.95% (118 of 1,699 words) |
| Accurate (Whisper small) | 299 MB | 64 s | 4.00% (68 of 1,699 words) |
Method. The input was chapters I and II of Pride and Prejudice from Project Gutenberg, 1,691 words read by the macOS voice Samantha at 175 words a minute and saved as a 64 kb/s mono MP3. The machine was a MacBook Pro with an M1 Max and 32 GB, in Chrome 154 with WebGPU. Time excludes the one-off model download.
Errors are word-level edits against the source text, after lower-casing and removing punctuation. Several of the errors were names, such as “Bennett” for Bennet, which is the mistake to check for first.
Without WebGPU the model runs on the processor through WebAssembly. Standard then read the first two minutes of the same file in 33.2 seconds on the same machine, about four times slower than WebGPU.
A synthetic voice is the easiest speech there is. Real interviews, with two people, room echo and interruptions, will have a higher error rate on every method. Use the numbers to compare the two models, not to predict your own result.
Getting a better transcript
- Clean the audio first. Hum and room noise cost more words than anything else. The noise remover runs before transcription.
- Name the language. Automatic detection listens to the first spoken part. A quiet or music-led opening can fool it.
- Use the larger model for anything you will publish. In our test it made 42% fewer errors.
- Check names, numbers and technical terms. Those are where every model guesses.
- Keep the timestamps. They let you find a quote in the audio in seconds.
From transcript to subtitles
A subtitle file is a transcript cut into short timed lines. SRT is the most widely read format, and VTT is the one web players prefer. Both can be opened and retimed in the subtitle editor, which can also burn them into a video.