Remove silence from audio
Shorten the pauses in a podcast, interview or voice note without cutting every quiet word. A speech model marks the waveform first, so you can hear the result before one edit becomes a download.
Open the audio, inspect the speech map, preview and save.
Drop a spoken recording here
The speech model finds words rather than guessing at a noise level. Your audio stays in this tab.
MP3, WAV, M4A, AAC, FLAC, OGG and Opus, when your browser can decode them. .
Longer natural pauses stay at this length. Short ones stay untouched.
Protect quiet speech keeps more doubtful audio. Tight makes more cuts.
Extra audio kept before and after each detected spoken region.
Keyboard controls
| Space | Play or pause |
| B | Switch between Original and Result |
| ←→ | Skip five seconds |
| Home | Return to the start |
The recording is decoded and edited in this tab. It is never uploaded, and it never enters your account.
Signed in, typed text and settings save automatically to your private account. Text opened from files stays in this browser’s local storage. Files stay in your browser. Account saving
How to remove silence from an audio file
Open a spoken recording
Drop in an MP3, WAV, M4A, FLAC, OGG or Opus file. The browser decodes it locally and never uploads the recording.
Find speech and inspect the map
Press Find speech. Purple marks detected speech and amber marks audio that the current settings will leave out.
Listen, adjust and download
Switch between Original and Result, change the pause treatment if needed, then download the edited recording as WAV.
Using the audio silence remover
The model makes the first pass. You decide what happens to its gaps. Nothing is changed in the original file, and every setting rebuilds a fresh result in memory.
Opening a spoken recording
Drop a file onto the waveform or press Open a file. Try the room-tone example if you want to see the speech map before using your own recording.
| What you open | What the tool does |
|---|---|
| MP3, WAV, M4A, AAC, FLAC, OGG or Opus | Your browser decodes it. Format support can differ between browsers |
| Stereo audio | The channels stay separate in the WAV. A mono mix is used only for speech detection |
| More than 45 minutes | The file is refused before the model runs, because the decoded copies may exhaust tab memory |
| Music or singing | The file may open, but the speech model can remove material you meant to keep |
Reading the speech map
Press Find speech after the waveform appears. Purple regions are audio the model classifies as speech. Amber hatching is audio the current plan will leave out.
Click anywhere on the waveform to listen from that point. Original and Result keep the nearest matching place when you switch between them, even after several gaps have gone.
Choosing how pauses change
Shorten pauses is the safer default for conversation. Remove pauses is useful for isolated prompts or datasets where any gap is unwanted.
| Control | Effect |
|---|---|
| Shorten pauses | Keeps short natural gaps and caps longer gaps at the Pause length |
| Remove pauses | Joins the padded speech regions with an 8 millisecond equal-power crossfade |
| Pause length | Sets the longest internal gap from 0.2 to 1.2 seconds |
| Speech sensitivity | Protect quiet speech keeps more doubtful audio. Tight makes more cuts |
| Speech padding | Keeps 80, 140 or 240 milliseconds around each detected region |
Checking and downloading the result
The four figures report detected regions, changed pauses, removed time and result length. Use Original and Result while the file plays to judge the same part both ways.
Download WAV writes a 16-bit PCM file. The result is larger than an MP3 because the browser has no built-in MP3 encoder, and this tool does not send the audio away to get one.
Keyboard controls
The waveform takes these keys after you click or tab into it.
| Space | Play or pause |
| B | Switch between Original and Result |
| ← → | Move five seconds |
| Home | Return to the start |
What this tool will not do
It will not clean hiss under speech, recognise individual words or make a useful edit of music. It also cannot promise that every breath is kept. Protect quiet speech and generous padding are the safer settings when a recording contains whispers or distant voices.
The default settings turned the 8.950563 second room-tone example into 4.956 seconds. They kept both the normal and quiet spoken phrases. The method and every competing result are in the silence removal guide.
Frequently asked questions
Is my audio uploaded?
No. Your browser decodes the recording, runs the speech model and writes the WAV in the same tab. The recording never enters a floi account. The first use downloads a 2.3 MB model and the shared browser engine from floi’s own model domain.
How does it know what is silence?
It does not define silence as “below -40 dB”. Silero VAD 6.2 gives the tool one speech probability every 32 milliseconds. The tool joins those decisions into spoken regions, adds your chosen padding, then edits the gaps between them. This is why a quiet sentence can survive above steady room tone.
What is the difference between Shorten pauses and Remove pauses?
Shorten pauses keeps up to the chosen pause length between spoken regions. A 2.4 second gap becomes 0.5 seconds at the default setting, while a natural 0.3 second pause stays untouched. Remove pauses joins the padded speech regions directly with an 8 millisecond crossfade.
Can it remove silence from music?
It should not. The model is trained to recognise speech, so singing, a quiet intro or an instrumental break can be classified as non-speech and removed. Use the audio trimmer for music when you know the section you want to keep.
Why does it download as WAV instead of MP3?
Browsers decode MP3 but do not include an MP3 encoder. WAV keeps the decoded result without another lossy encode. It is larger. Converting that WAV to MP3 later is reasonable when file size matters more than avoiding another lossy generation.
Will it remove background noise too?
No. It shortens the gaps, but any room tone kept under a word stays under that word. Clean a difficult recording in the audio noise remover first, then bring the cleaned file here. Keeping these as two visible steps makes each change easier to judge.
How long can the recording be?
The tool stops at 45 minutes. A decoded stereo recording and its edited copy can use more than a gigabyte before the WAV is written. The limit prevents a browser tab from failing halfway through a long file. Split a longer recording in the audio trimmer first.
Which model does it use?
Silero VAD 6.2, using the publisher’s official 2,327,524 byte ONNX file under the MIT licence. It runs with ONNX Runtime Web at 16 kHz and keeps the original decoded sample rate for the final WAV.
Something missing? Request a feature