Google Analytics helps us understand which tools you use and which actions work. Your files and entered text stay out of analytics.

Free · no sign-upAudio and videoNothing uploaded

Transcribe audio to text

Turn an interview, lecture, voice memo or video into text with a timestamp on every line. Whisper runs on your own device, so a recording you could not send to a cloud service never has to leave it.

No daily file limit or minute allowanceRecordings up to two hoursTXT, SRT and VTT with timestampsWorks offline after the first run

Open the recording, transcribe it, correct and save.

No file open

Drop a recording or a video here

Speech becomes text with timestamps, on this device. The file is never uploaded.

MP3, WAV, M4A, FLAC, OGG, MP4, MOV or WebM, up to two hours. .

Whisper base. Quick enough for long recordings.

Naming the language avoids a wrong guess on a short or noisy opening.

Writes down what was said, in the language it was said in.

Open a recording to startThe first run downloads the speech model once. After that it works offline.

The recording is decoded and transcribed in this tab. It is never uploaded, and the transcript is not saved to your account.

Signed in, your model, language and format choices save to your private account. The recording and the transcript stay in this browser. Account saving

How to transcribe an audio file

  1. Open a recording or a video

    Drop in an MP3, WAV, M4A, FLAC, OGG, MP4, MOV or WebM file up to two hours long. Your browser decodes it and nothing is uploaded.

  2. Choose the model and press Transcribe

    Standard is quick. Accurate makes fewer mistakes and takes about 1.7 times as long. Name the spoken language, or let Whisper detect it.

  3. Correct the lines and save

    Click a time to hear that line, fix any word in place, then copy the text or download TXT, SRT or VTT.

Using the audio to text tool

The tool reads the file, runs Whisper over it in 28 second windows and writes each line as soon as it is read. You can stop at any point and keep what is done.

Opening a recording or a video

Drop a file on the panel or press Open a file. Try a short example plays a 20 second synthetic reading if you want to see the result before using your own file.

What you openWhat happens
MP3, WAV, M4A, AAC, FLAC, OGG or OpusYour browser decodes it to a 16 kHz mono copy, which is what Whisper reads
MP4, MOV or WebM videoOnly the soundtrack is read. The picture is ignored and the file stays where it is
Longer than two hours, or over 2 GBThe file is refused before the model runs, because it is held in memory
A file your browser cannot playIt will not decode. Convert it to MP3 or WAV first

Choosing the settings

Set these before you press Transcribe. They are remembered on this device.

ControlEffect
Model: StandardWhisper base, a 142 MB download. The quick choice for long recordings
Model: AccurateWhisper small, a 299 MB download. In our test it made 42% fewer mistakes and took 1.7 times as long
Spoken languageDetect automatically reads the first spoken part and picks one of 99 languages. Naming it avoids a wrong guess
Write it in: The language spokenA transcript in the original language
Write it in: English, translatedAn English translation of any language Whisper understands

Watching it run

The status line shows how far through the recording Whisper is and how long the rest should take at the current speed. Lines appear as each window finishes.

Stop ends the run and keeps every line already written. Pressing Stop while the model is still downloading discards that run, and the recording stays open.

Correcting the transcript

Every line has its start time beside it. Click the time to play the recording from there, then click the words to fix them. The line being played is highlighted.

The four figures count the words, the lines, the length of the recording and the time the run took. The word count updates as you edit.

Saving and copying

Save asWhat you get
Plain text (.txt)Running text. A pause of two seconds or more starts a new paragraph
Text with timestamps (.txt)One line per segment, each starting with its time, such as [12:04]
Subtitles (.srt)Numbered cues of up to 84 characters, for video players and editors
Web subtitles (.vtt)The same cues in WebVTT, for the HTML video element and most web players

Copy puts the chosen format on the clipboard. Edit as subtitles opens the lines in the subtitle editor, where you can retime them and burn them into a video.

Keyboard controls

There are no tool-specific shortcuts. Tab moves between the times and the lines,Enter on a time plays from it, and Enter in a line finishes the edit.

What this tool will not do

  • It does not label speakers. A conversation comes out as one stream of lines.
  • It is slower than a cloud service. Your device does the work, not a rack of graphics cards.
  • It does not accept links. It cannot fetch a YouTube video or a cloud file. Download the file first.
  • It does not keep the transcript in your account. Save or copy it before you close the tab.
  • It translates only into English. Whisper cannot translate in any other direction.

How it compares with the free transcription sites

Most free transcribers upload your file and limit how much you can do in a day. These terms were read from each service’s own page on 1 October 2026.

ServiceWhere the file goesFree limitSpeaker labels
floiStays on your deviceNone. Two hours per fileNo
TurboScribeUploaded3 files a day, 30 minutes eachYes
WaveUploaded, deleted afterwards5 files a day, 200 MB eachYes
Audio TranscriberUploaded, deleted after 24 hours100 MB per fileYes
OtterUploadedAccount requiredYes

The cloud services win on speed and on speaker labels. floi wins when the recording is private, long, or one of many. The transcription guide compares every method, including Word and Google Docs, with measured accuracy.

Frequently asked questions

Is my recording uploaded?

No. Your browser decodes the file and Whisper runs in the same tab, on your processor or graphics chip. The only download is the speech model, once, from floi’s own model domain. The transcript is not saved to a floi account either, even when you are signed in.

Is it really free, with no limit?

Yes. There is no daily file count, no minute allowance and no sign-up. The work happens on your device, so a tenth recording costs floi nothing more than the first. The two hour ceiling is about browser memory, not a quota.

Can I transcribe a video file?

Yes, when your browser can play its audio. MP4 and MOV with AAC audio work in current Chrome, Edge, Firefox and Safari. WebM works in Chrome, Edge and Firefox. Some MKV files, and video with an unusual audio codec, will not decode here.

How accurate is it?

It is good on clear speech and worse on crosstalk, heavy accents, music beds and far-away microphones. On an 8 minute reading, Standard got 6.95% of words wrong and Accurate 4.00%. Both are smaller than the large models cloud services run, so expect to correct some words, names above all.

Does it label who is speaking?

No. Every line is one stream of text with a start time. Speaker labels need a second model, and the browser versions of that are not reliable enough yet to ship. Cloud services such as TurboScribe and Otter do label speakers.

Which languages does it understand?

Whisper was trained on 99 languages. The list names 20 common ones, and Detect automatically lets Whisper pick from all 99 using the first spoken part of the file. English, Spanish, French, German and Portuguese are the strongest. Smaller languages produce more errors.

Can it translate into English?

Yes. Set Write it in to English, translated. Whisper then writes English for any language it hears. It only translates into English. That is a limit of the model, not a setting.

Can I turn the transcript into subtitles?

Yes. Save as SRT or VTT for a subtitle file, or press Edit as subtitles to open the lines in the subtitle editor. Long lines are split into cues of up to two 42 character lines.

Something missing? Request a feature

↑↓ move↵ openesc close