Google Analytics helps us understand which tools you use and which actions work. Your files and entered text stay out of analytics.

How to transcribe audio to text

Five methods compared, with speed and accuracy measured for two Whisper models on 1 October 2026.

Free toolAudio to TextTranscribe a recording or a video into text with timestamps. Whisper runs on your device, and the file is never uploaded.Open the transcriber

The short answer

Open the recording in a transcriber that accepts files, check the result against the audio, and save it with timestamps. Which transcriber depends on two questions: may the recording leave your computer, and do you need to know who is speaking?

If youUse
Already pay for Microsoft 365 and transcribe a few hours a monthWord for the web
Need speaker labels and can upload the fileA cloud service such as TurboScribe or Otter
Cannot upload the recordingWhisper on your own computer, in a desktop app or a browser
Want to speak and have it typed as you goDictation, not file transcription

Method 1: Word for the web

Word’s Transcribe button takes an MP3, WAV, M4A or MP4 and returns a transcript with timestamps and speaker labels. It is in Word for the web only, and it needs a Microsoft 365 subscription.

  1. Open a document at office.com and choose Home, then Dictate, then Transcribe.
  2. Choose Upload audio and pick the file. Microsoft processes it on its servers.
  3. Edit the sections in the side panel, then add them to the document.

The limits are 300 minutes of uploaded audio a month and 300 MB per file, according to Microsoft’s support page. There is no SRT export.

Method 2: Google Docs voice typing

Google Docs cannot transcribe a file. Voice typing listens to the microphone in Chrome, and there is nowhere to open a recording.

The common workaround is to play the file through a speaker next to the microphone. It loses words to room echo and volume, needs the whole recording played in real time, and produces no timestamps. It is a reasonable way to dictate, and a poor way to transcribe.

Method 3: a cloud transcription service

Upload the file, wait a few minutes, download the transcript. Cloud services run large models on dedicated hardware, so they are fast, and most of them label speakers.

ServiceFree tier, as stated on 1 October 2026
TurboScribe3 files a day, 30 minutes each. Unlimited from $10 a month billed yearly
Wave5 files a day, 200 MB each, no account
OtterAccount required. English, Spanish and French

The cost is that the recording leaves your computer. Read the retention terms before you upload anything you were not given permission to share.

Method 4: a desktop Whisper app

Whisper is the open speech model OpenAI released in 2022. Desktop apps run it on your own machine, so the file never leaves it. MacWhisper is the best known on a Mac, and Buzz is a free open source option for Windows, Mac and Linux.

They can run the large Whisper models, which are more accurate than anything a browser can comfortably download. They need installing, which a locked-down work laptop may not allow.

Method 5: Whisper in the browser

The audio to text tool runs Whisper inside the browser tab. Nothing is installed and nothing is uploaded. The model downloads once and is cached.

  1. Open the recording or video in the tool.
  2. Choose Standard or Accurate, and name the language if you know it.
  3. Press Transcribe, correct the lines, then save TXT, SRT or VTT.

It does not label speakers, and it is slower than a cloud service. It suits private recordings and anyone who would otherwise hit a daily limit.

How fast and how accurate Whisper is in a browser

We transcribed the same 8 minute 36 second reading with both models the tool offers, and counted every word that came out wrong. Measured on 1 October 2026.

ModelDownloadTime for 8:36 of speechWord error rate
Standard (Whisper base)142 MB38 s6.95% (118 of 1,699 words)
Accurate (Whisper small)299 MB64 s4.00% (68 of 1,699 words)

Method. The input was chapters I and II of Pride and Prejudice from Project Gutenberg, 1,691 words read by the macOS voice Samantha at 175 words a minute and saved as a 64 kb/s mono MP3. The machine was a MacBook Pro with an M1 Max and 32 GB, in Chrome 154 with WebGPU. Time excludes the one-off model download.

Errors are word-level edits against the source text, after lower-casing and removing punctuation. Several of the errors were names, such as “Bennett” for Bennet, which is the mistake to check for first.

Without WebGPU the model runs on the processor through WebAssembly. Standard then read the first two minutes of the same file in 33.2 seconds on the same machine, about four times slower than WebGPU.

A synthetic voice is the easiest speech there is. Real interviews, with two people, room echo and interruptions, will have a higher error rate on every method. Use the numbers to compare the two models, not to predict your own result.

Getting a better transcript

  • Clean the audio first. Hum and room noise cost more words than anything else. The noise remover runs before transcription.
  • Name the language. Automatic detection listens to the first spoken part. A quiet or music-led opening can fool it.
  • Use the larger model for anything you will publish. In our test it made 42% fewer errors.
  • Check names, numbers and technical terms. Those are where every model guesses.
  • Keep the timestamps. They let you find a quote in the audio in seconds.

From transcript to subtitles

A subtitle file is a transcript cut into short timed lines. SRT is the most widely read format, and VTT is the one web players prefer. Both can be opened and retimed in the subtitle editor, which can also burn them into a video.

Frequently asked questions

What is the easiest free way to transcribe an audio file?

Open it in a transcriber that takes files directly. Word for the web does it if you already pay for Microsoft 365. Otherwise a browser tool running Whisper, or a cloud service’s free tier, will give you a transcript in minutes. Google Docs cannot open an audio file at all.

Can Google Docs transcribe an audio file?

No. Voice typing in Google Docs listens to a microphone in Chrome. It has no way to open a recording. Playing the file out of a speaker into the microphone works badly, and the transcript has no timestamps.

How accurate is AI transcription?

On clear, single-voice speech, very. In our test Whisper small got 96 words in 100 right and Whisper base 93. Real recordings with crosstalk, accents and room noise will do worse, and names are the most common mistake. Always read the transcript against the audio before you quote it.

How long does it take to transcribe an hour of audio?

A cloud service takes a few minutes. On a 2021 MacBook Pro with WebGPU, Whisper base in the browser read 8 minutes 36 seconds of speech in 38 seconds, so an hour would take about four and a half minutes. A laptop without WebGPU takes several times longer.

Is it safe to upload a recording to a transcription site?

It depends on the recording and the site’s terms. Most free sites keep the file for hours or days. If you are not allowed to share the recording, for example a research interview or a client call, use a method that keeps it on your own computer.

How do I get timestamps or subtitles from a transcript?

Use a transcriber that writes SRT or VTT. Both are plain text files with a start and end time on every line, and every video editor and player reads them. Word’s transcript has timestamps but no subtitle export.

↑↓ move↵ openesc close