Transcribe audio to text
Turn an interview, lecture, voice memo or video into text with a timestamp on every line. Whisper runs on your own device, so a recording you could not send to a cloud service never has to leave it.
Open the recording, transcribe it, correct and save.
Drop a recording or a video here
Speech becomes text with timestamps, on this device. The file is never uploaded.
MP3, WAV, M4A, FLAC, OGG, MP4, MOV or WebM, up to two hours. .
Whisper base. Quick enough for long recordings.
Naming the language avoids a wrong guess on a short or noisy opening.
Writes down what was said, in the language it was said in.
Click a time to hear that line. Click the words to correct them before you save.
The recording is decoded and transcribed in this tab. It is never uploaded, and the transcript is not saved to your account.
Signed in, your model, language and format choices save to your private account. The recording and the transcript stay in this browser. Account saving
How to transcribe an audio file
Open a recording or a video
Drop in an MP3, WAV, M4A, FLAC, OGG, MP4, MOV or WebM file up to two hours long. Your browser decodes it and nothing is uploaded.
Choose the model and press Transcribe
Standard is quick. Accurate makes fewer mistakes and takes about 1.7 times as long. Name the spoken language, or let Whisper detect it.
Correct the lines and save
Click a time to hear that line, fix any word in place, then copy the text or download TXT, SRT or VTT.
Using the audio to text tool
The tool reads the file, runs Whisper over it in 28 second windows and writes each line as soon as it is read. You can stop at any point and keep what is done.
Opening a recording or a video
Drop a file on the panel or press Open a file. Try a short example plays a 20 second synthetic reading if you want to see the result before using your own file.
| What you open | What happens |
|---|---|
| MP3, WAV, M4A, AAC, FLAC, OGG or Opus | Your browser decodes it to a 16 kHz mono copy, which is what Whisper reads |
| MP4, MOV or WebM video | Only the soundtrack is read. The picture is ignored and the file stays where it is |
| Longer than two hours, or over 2 GB | The file is refused before the model runs, because it is held in memory |
| A file your browser cannot play | It will not decode. Convert it to MP3 or WAV first |
Choosing the settings
Set these before you press Transcribe. They are remembered on this device.
| Control | Effect |
|---|---|
| Model: Standard | Whisper base, a 142 MB download. The quick choice for long recordings |
| Model: Accurate | Whisper small, a 299 MB download. In our test it made 42% fewer mistakes and took 1.7 times as long |
| Spoken language | Detect automatically reads the first spoken part and picks one of 99 languages. Naming it avoids a wrong guess |
| Write it in: The language spoken | A transcript in the original language |
| Write it in: English, translated | An English translation of any language Whisper understands |
Watching it run
The status line shows how far through the recording Whisper is and how long the rest should take at the current speed. Lines appear as each window finishes.
Stop ends the run and keeps every line already written. Pressing Stop while the model is still downloading discards that run, and the recording stays open.
Correcting the transcript
Every line has its start time beside it. Click the time to play the recording from there, then click the words to fix them. The line being played is highlighted.
The four figures count the words, the lines, the length of the recording and the time the run took. The word count updates as you edit.
Saving and copying
| Save as | What you get |
|---|---|
| Plain text (.txt) | Running text. A pause of two seconds or more starts a new paragraph |
| Text with timestamps (.txt) | One line per segment, each starting with its time, such as [12:04] |
| Subtitles (.srt) | Numbered cues of up to 84 characters, for video players and editors |
| Web subtitles (.vtt) | The same cues in WebVTT, for the HTML video element and most web players |
Copy puts the chosen format on the clipboard. Edit as subtitles opens the lines in the subtitle editor, where you can retime them and burn them into a video.
Keyboard controls
There are no tool-specific shortcuts. Tab moves between the times and the lines,Enter on a time plays from it, and Enter in a line finishes the edit.
What this tool will not do
- It does not label speakers. A conversation comes out as one stream of lines.
- It is slower than a cloud service. Your device does the work, not a rack of graphics cards.
- It does not accept links. It cannot fetch a YouTube video or a cloud file. Download the file first.
- It does not keep the transcript in your account. Save or copy it before you close the tab.
- It translates only into English. Whisper cannot translate in any other direction.
How it compares with the free transcription sites
Most free transcribers upload your file and limit how much you can do in a day. These terms were read from each service’s own page on 1 October 2026.
| Service | Where the file goes | Free limit | Speaker labels |
|---|---|---|---|
| floi | Stays on your device | None. Two hours per file | No |
| TurboScribe | Uploaded | 3 files a day, 30 minutes each | Yes |
| Wave | Uploaded, deleted afterwards | 5 files a day, 200 MB each | Yes |
| Audio Transcriber | Uploaded, deleted after 24 hours | 100 MB per file | Yes |
| Otter | Uploaded | Account required | Yes |
The cloud services win on speed and on speaker labels. floi wins when the recording is private, long, or one of many. The transcription guide compares every method, including Word and Google Docs, with measured accuracy.
Frequently asked questions
Is my recording uploaded?
No. Your browser decodes the file and Whisper runs in the same tab, on your processor or graphics chip. The only download is the speech model, once, from floi’s own model domain. The transcript is not saved to a floi account either, even when you are signed in.
Is it really free, with no limit?
Yes. There is no daily file count, no minute allowance and no sign-up. The work happens on your device, so a tenth recording costs floi nothing more than the first. The two hour ceiling is about browser memory, not a quota.
Can I transcribe a video file?
Yes, when your browser can play its audio. MP4 and MOV with AAC audio work in current Chrome, Edge, Firefox and Safari. WebM works in Chrome, Edge and Firefox. Some MKV files, and video with an unusual audio codec, will not decode here.
How accurate is it?
It is good on clear speech and worse on crosstalk, heavy accents, music beds and far-away microphones. On an 8 minute reading, Standard got 6.95% of words wrong and Accurate 4.00%. Both are smaller than the large models cloud services run, so expect to correct some words, names above all.
Does it label who is speaking?
No. Every line is one stream of text with a start time. Speaker labels need a second model, and the browser versions of that are not reliable enough yet to ship. Cloud services such as TurboScribe and Otter do label speakers.
Which languages does it understand?
Whisper was trained on 99 languages. The list names 20 common ones, and Detect automatically lets Whisper pick from all 99 using the first spoken part of the file. English, Spanish, French, German and Portuguese are the strongest. Smaller languages produce more errors.
Can it translate into English?
Yes. Set Write it in to English, translated. Whisper then writes English for any language it hears. It only translates into English. That is a limit of the model, not a setting.
Can I turn the transcript into subtitles?
Yes. Save as SRT or VTT for a subtitle file, or press Edit as subtitles to open the lines in the subtitle editor. Long lines are split into cues of up to two 42 character lines.
Something missing? Request a feature