Speech to text
Speak into an editable transcript without sending your voice to a transcription service. Correct the text as you go, then copy it or save a plain TXT file.
Choose a language, speak, then copy or download the text.
English works best. Other languages can be much less accurate because this small local model has uneven multilingual training.
Audio is processed in this tab and is never saved. Signed in, only your transcript and language choice save privately.
Signed in, typed text and settings save automatically to your private account. Text opened from files stays in this browser’s local storage. Files stay in your browser. Account saving
How to turn speech into text
Choose the spoken language
Select a language for steadier recognition, or leave Auto-detect on for mixed or uncertain input.
Start dictating
Allow microphone access and speak naturally. The first run downloads the local model while your speech waits in this tab.
Edit and keep the text
Stop when you are done. Correct the editable transcript, then copy it or download a plain TXT file.
Using the speech to text tool
Pick the language you plan to speak, then press Start dictating. Your browser asks for microphone access the first time. A moving waveform confirms that the tab can hear you, and the status line tells you whether a section is waiting or finished.
You do not have to wait for the model before speaking. The first short sections wait in memory while the download finishes. They are then processed in order without leaving the tab. Keep the page open until the queue says Caught up.
Spoken language
Choose a named language when you know it. That gives the model one less decision to make. Auto-detect is useful for a quick test or when the language is uncertain, but it can guess badly from a very short opening.
Listening and model status
The top row separates microphone state from model state. Listening means audio is being captured. Local model ready means Whisper can process it. A queue count above zero is normal when your device transcribes more slowly than you speak.
WebGPU is tried first on supported devices. If it cannot start, the tool says so and uses WebAssembly on the processor. That fallback is broader but can be slower.
Transcript and exports
New text appears in sections of about eight seconds. The transcript remains an ordinary editable text area, so you can fix names and punctuation while the next section runs. Copy records a completed task. Download creates a UTF-8 plain text file in this browser.
Keyboard controls
| Shortcut | Action |
|---|---|
| Ctrl or Cmd + Enter | Start or stop dictation while a tool control has focus. |
| Esc | Stop an active microphone. |
| Tab | Move through language, recording, editor and export controls. |
What happens inside the browser
This page runs the quantised Whisper tiny ONNX model through Transformers.js in a Web Worker. The microphone is sampled in the tab, converted to 16 kHz mono audio, and divided into short sections. Only model files travel over the network.
The q4 encoder and merged decoder weights total 95,734,369 bytes. Small configuration and tokenizer files are also fetched. The size was checked against the model repository on 24 September 2026 rather than copied from a marketing page.
How it compares with other speech to text options
| Option | Useful advantage | Tradeoff |
|---|---|---|
| floi local dictation | Live, editable text with no audio upload, account or quota. | About 96 MB on first use. The small model can lag or miss difficult speech. |
| Built-in browser recognition | Very small interface and quick live results. | The default recognition service can process audio on a server. Local mode is still experimental. |
| Local file transcriber | Can handle an existing recording and may export subtitles. | Often needs a larger model and is not a live dictation surface. |
| Hosted transcription service | Larger models, speaker labels and team workflows. | Uploads audio, usually requires an account, and often limits free minutes. |
| Operating-system dictation | Works across apps and may already be installed. | Privacy, languages and local processing vary by system and setting. |
Read the private speech to text guide for a closer look at those methods and the privacy wording that matters.
What the tool will not do
- It will not accept an audio file. This page is for live microphone dictation. To transcribe a recording, use audio to text.
- It will not identify speakers. Everyone becomes one continuous transcript.
- It will not create SRT or VTT subtitles. There are no timestamps in the TXT result.
- It will not recognise every language equally well. English is the model's strongest language. Other languages can produce substantial errors.
- It will not match a large cloud model on every recording. Whisper tiny favours a manageable browser download.
- It will not save audio. Stop or Clear discards any captured sections that have not produced text.
Frequently asked questions
Is this speech to text tool free?
Yes. There is no account requirement, minute allowance, daily cap or paid export. Recognition runs on your device, so another minute does not create a transcription bill that needs a quota.
Is my voice uploaded anywhere?
No. Your browser records short sections from the microphone and gives them to a local copy of Whisper. The audio is not uploaded or saved. The model files are downloaded from Hugging Face on first use. If you sign in to floi, the transcript and language choice save privately to your account, but the audio never does.
Why does the first use download about 96 MB?
Speech recognition needs learned model weights. This tool uses the q4 Whisper tiny encoder and decoder, which total about 95.7 MB. Your browser normally caches them for later visits. Private browsing, cleared site data or storage pressure can remove that cache.
How accurate is the transcript?
Clear speech, a close microphone and the correct language give the best result. Whisper tiny is deliberately small enough to run in a browser. A larger cloud model can be more accurate with names, heavy accents, several speakers or background noise. Always review important text.
Can it transcribe an existing audio file?
Not on this page. This tool is for live microphone dictation. It does not accept MP3, WAV or video files, and it does not make subtitles or identify speakers.
Which languages are supported?
The menu includes 31 languages and automatic detection, but English works best. Other languages can be much less accurate because Whisper tiny has uneven multilingual training. Choosing a language helps the model start correctly, but it cannot remove that underlying limitation.
Does speech to text work offline?
The recognition step can run without a speech service after the page and model are available. The first use needs a connection for the model download. Browser caches are not permanent, so offline reuse cannot be guaranteed after site data is cleared or storage is reclaimed.
Which browser should I use?
Use a current version of Chrome, Edge, Firefox or Safari on a device with enough free memory. WebGPU is used when available. The tool falls back to the processor through WebAssembly when it is not. A slower phone may take longer than the speech itself to finish a section.
Something missing? Request a feature