How to use speech to text privately
Use a named local model when the recording must stay on your device. Ordinary browser speech recognition can send audio to a service, even when the page itself stores nothing.
Free toolSpeech to TextDictate into an editable transcript with a local Whisper model. Copy or download the text without uploading your voice.Open the dictation toolThe useful distinction: capture is not recognition
A browser always captures microphone samples on your device. The important question is where those samples are recognised. A local tool passes them to a model in the tab. A cloud tool uploads them to a model on somebody else's computer.
The Web Speech API does not settle that question by itself. MDN explains that recognition on a web page is server-based by default and sends audio to a service. Its newer processLocally option requests on-device recognition, but the feature is experimental and needs installed language packs.
Four practical methods
| Method | Voice goes | Choose it when |
|---|---|---|
| Local browser model | Into a model running in the tab. | You want a private, inspectable tool and accept a model download. |
| On-device system dictation | Stays local only when the system documents that mode. | You need dictation across several apps. |
| Web Speech recognition | Often to a browser or operating-system service. | You want a small live interface and the audio is not sensitive. |
| Hosted transcription | To the provider's servers. | You need large models, collaboration, speaker labels or an API. |
Method 1: run Whisper in the browser
Whisper is a multilingual speech recognition model. Browser versions use ONNX weights and a JavaScript runtime to execute it through WebGPU or WebAssembly. The model download is large because the page is bringing the recogniser to the recording instead of taking the recording to a recogniser.
- Open the speech to text tool.
- Choose the spoken language when you know it.
- Press Start, allow microphone access and begin speaking.
- Watch the separate microphone, model and queue states.
- Stop, correct names and punctuation, then copy or download the result.
Floi uses the q4 Whisper tiny release. On 24 September 2026, its encoder was 9,020,667 bytes and its merged decoder was 86,713,702 bytes. That is 95,734,369 bytes of weights before small configuration and tokenizer files. The figures came from the public model repository API.
The size is a deliberate compromise. A larger Whisper model can improve difficult recognition, but would make the first visit and phone memory cost much higher. The tiny model is suitable for clear personal dictation. It is not a court transcription service.
Method 2: use system dictation
Windows, macOS, iOS, Android and ChromeOS all provide voice typing. This is the easiest option when the destination is another app. It can place text directly into an email, document or message without a copy step.
Do not assume that every setting means the same thing. Some languages support on-device recognition while others need a connection. Product releases also change which speech is retained or reviewed. Check the privacy note for your system, language and current version.
Method 3: use browser speech recognition carefully
The browser API can produce fast interim words with almost no page download. That makes it attractive for simple voice typing. The privacy claim must still name the recognition mode. A statement that the page has no database does not prove that the browser recognition service kept the audio local.
If a page uses processLocally = true, it should also handle unavailable language packs and installation failure. A silent fallback to remote recognition defeats the point. The MDN Web Speech guide describes both on-device setup and the default server-based path.
Method 4: choose a cloud service for cloud-sized work
A hosted product earns its place for meetings, long files, shared projects, speaker labels, searchable archives and APIs. It can run a larger model than a browser tab and process a backlog without keeping your laptop awake.
Read the provider's retention, training and deletion terms before uploading sensitive speech. Check limits in minutes, not just the word “free”. Several products let you record without payment but reserve export, longer sessions or their accurate model for an account.
What seven current tools actually offer
The following pages were checked on 24 September 2026. It is a product audit, not a claim about every plan or later release. Cloudflare blocked two interactive pages, so their rows use the public search description and are marked accordingly.
| Tool | Observed approach | Useful limit or tradeoff |
|---|---|---|
| SoundTools | Local Whisper for uploaded or recorded audio. | 180 MB or 240 MB models, desktop only, with SRT export. |
| OpenBroca | Polished local Whisper editor. | 12 language choices and no download button observed. |
| 712Tools | Live Web Speech interface. | Copy and clear only. Its local-only wording did not explain the default remote path. |
| utils.com | Web Speech with language choice. | Records audio and offers a download, but the privacy path was not made clear. |
| OmniTools | Live recognition with output translation. | Broad language choice. Its browser-only claim did not identify the recognition engine. |
| OpenL | Hosted AI transcription. | Minimal editor with a ten-run daily allowance observed. |
| FreeTTS | Live preview plus a hosted accurate pass. | Public result listed 90 minutes a month and 15 minutes a session. Interactive page blocked. |
A five-point privacy check
- Name the engine. “AI” says nothing about where it runs.
- Look for the first download. A local model has weights to fetch or install.
- Check the failure mode. Local recognition should stop instead of falling back silently.
- Separate audio from text. A tool may keep voice local but sync the transcript to an account.
- Test the network if the work matters. Browser developer tools show requests made after recording starts.
Improve recognition before changing tools
- Put the microphone 15 to 30 cm from your mouth.
- Choose the language instead of relying on a one-word detection sample.
- Speak in complete phrases, with a short pause at sentence boundaries.
- Turn away from fans and hard walls that add steady noise or echo.
- Review names, numbers and specialist terms before using the transcript.