Text to speech
Turn English text into a natural voice inside this browser. Choose the voice and pace, listen to the result, then save a WAV or MP3 without an account or usage allowance.
Write the text, choose a voice, then play or save the result.
Your clips 0
Newest firstClips stay until you leave or reload. Download any you want to keep.
Your first clip will appear here.
Speech is generated in your browser. Signed in, your text and settings save privately. Audio never enters your account.
Signed in, your text and settings save automatically to your private account. Files stay in your browser. Account saving
How to turn text into speech
Paste the text you want to hear
Use up to 3,000 characters. A short script, pronunciation check, or accessibility copy all fit.
Choose a voice and speed
Pick an American or British English voice. Set the pace from 0.70× to 1.30×.
Generate, listen, and save
The model builds speech in this browser. Scrub through the result, then download a WAV or MP3.
Using the text to speech tool
The editor is the whole job in one view. Paste your script, pick a voice, and press Generate speech. You can also press Ctrlor Cmd + Enter while a tool control has focus.
The first run is different from the rest. Your browser has to fetch the 92.4 MB q8 model before it can speak. The status row shows that transfer and then shows synthesis progress. Leave the tab open until the player appears.
Choose a voice
The list is deliberately short. Seven American and four British English voices cover a useful range without asking you to audition a hundred names. Heart is the balanced default. Bella and Michael are composed, while Puck and Fable carry more energy.
Voice labels are descriptions, not promises about age, identity or suitability. Generate a sentence before committing to a long script. Names, abbreviations and unusual punctuation can change how any speech model reads a line.
Set the pace
Speed runs from 0.70× to 1.30×. Start at 1.00×. Slow it for instructions or language practice. Raise it for a draft review where you care more about finding awkward wording than producing finished narration.
Listen before you download
Every generated clip stays in a list, newest first. Play, pause, or drag its position slider before saving. Try the female and male voice filters to compare different readings. Download any clips you want to keep before reloading or leaving the page.
What happens inside the browser
This page uses Kokoro-82M, an Apache 2.0 speech model with 82 million parameters. The q8 weights come from the Kokoro ONNX release. They run through WASM in a separate Web Worker, so loading and inference do not freeze the controls.
Long input is divided at sentence and word boundaries. The clips are joined with 120 ms of silence, then written as 16-bit mono PCM at 24 kHz. Nothing in that process needs your text to leave the tab.
MP3 downloads use the lamejs encoder, based on LAME and distributed under the LGPL. Encoding runs in your browser.
How this compares with other options
| Option | Best at | Tradeoff |
|---|---|---|
| floi local AI | Private English speech with a downloadable file | 92.4 MB first load, English only |
| Built-in browser speech | Reading a page aloud with no model download | Voices vary by browser and saving a file is unreliable |
| Hosted AI service | Many languages, acting controls and studio workflows | Your script is processed remotely and free use is usually capped |
| Operating-system voice | Accessibility and offline reading across apps | Export controls vary by system |
The full text to speech guide helps you choose among those four paths.
What the tool will not do
- It will not speak every language. The current voice set is English only.
- MP3 is a compressed copy. Choose WAV if you want the original audio for editing.
- It will not clone a voice. You choose from the model's included synthetic voices.
- It will not hide the model cost. The first 92.4 MB download is shown before you press Generate.
- It will not promise the same speed on every device. Synthesis uses your processor, so older phones take longer.
Frequently asked questions
Is this text to speech tool really free?
Yes. There is no account requirement, credit balance, daily allowance, watermark or paid download. The speech model runs on your device, so generating another clip does not create a server bill that needs a usage limit.
Does my text get sent to an AI service?
No. Your browser downloads the speech model, then gives your text to that local copy. It is not submitted to Hugging Face, Kokoro or another speech service. If you are signed in to floi, your text and settings save automatically to your private account. The generated audio never enters that account.
Why is there a 92.4 MB download?
That is the quantised Kokoro speech model. A natural voice needs learned weights that ordinary browser speech does not include. The download begins only after you press Generate. The browser normally caches it, so later runs do not need to fetch the same file again. Private browsing, cleared site data or storage pressure can remove that cache.
Which languages does it support?
English only in this version, with American and British voices. Some hosted tools support dozens of languages and are the better choice when that matters. This page keeps the model small enough to run locally instead of claiming broad language support it does not have.
Can I download an MP3?
Yes. Each clip has WAV and MP3 download buttons. MP3 creates a smaller, compressed copy at 128 kbps inside your browser. Its encoder loads only when you request an MP3. Choose WAV for the original 24 kHz mono audio.
Does it work offline?
After the page and model are cached, it can generate without contacting a speech service. A first visit needs a connection for the 92.4 MB model. Browsers control their own caches, so offline reuse cannot be guaranteed after you clear site data or the browser reclaims storage.
Why is the text limited to 3,000 characters?
Long speech synthesis uses memory and processor time inside one tab. The limit keeps a slow phone from spending several minutes on a job it may not finish. Split a longer script at a paragraph boundary and generate several WAV files.
Can I use the generated voice commercially?
The Kokoro model is released under Apache 2.0, but a licence does not settle every question about a particular script, character or use. You remain responsible for the text you enter and for laws that apply where you publish the audio. Do not use a generated voice to impersonate a real person.