How to remove silence from audio
Speech detection, Audacity, FFmpeg and a manual cut, with six ranking tools inspected on the same file. Measured 26 September 2026.
Free toolSilence RemoverFind speech instead of guessing at a noise level, then shorten or remove the pauses you can see.Open the silence removerThe short answer
Use voice activity detection for spoken audio with room tone, distant voices or whispers. Use a decibel threshold for music or a clean studio recording. The difference is what each method believes silence means.
| Your recording | Best starting method |
|---|---|
| Podcast, interview or voice note | Speech detection, with 0.4 to 0.7 seconds left between phrases |
| Quiet speech over steady room tone | Speech detection with a lower sensitivity threshold and generous padding |
| Music with silent starts or ends | A threshold or a manual trim. A speech model is the wrong detector |
| A few obvious mistakes | Manual cuts on a waveform, because automation adds risk without saving much time |
| Hundreds of consistent files | FFmpeg, after tuning one representative file and listening to the result |
Why a noise threshold cuts the wrong thing
A threshold tool calls everything below one level silence. That works when the speaker is always loud and the background is always quiet. A real room often reverses that assumption.
A distant word can sit below a fan, while the fan runs through every pause. Raising the threshold risks the word. Lowering it keeps the room tone and finds no gap at all.
| Detector | Question it asks | Main failure |
|---|---|---|
| Peak or RMS threshold | Is this frame loud enough? | Quiet words and loud background share one line |
| Voice activity detection | Does this frame contain speech? | Music, singing and unusual voices can confuse it |
| Manual editing | Does this passage belong? | Accurate, but slow on a long recording |
Silero's official ONNX wrapper reads 512 samples at 16 kHz for each decision. That is one probability every 32 milliseconds. Its published timestamp code uses a lower exit threshold than entry threshold, which stops speech flickering on and off at one borderline frame. See the Silero VAD implementation.
Six ranking tools and floi on the same recording
I tested the first six relevant organic tool results on 26 September 2026. The input was an 8.950563 second mono recording made for this test, with one normal phrase, one quiet phrase, three long gaps and low pink room tone. The exact sample is the Try the room-tone example file in the paired tool.
| Tool and default | Result | What you can control |
|---|---|---|
| floi, speech detection | 4.956 s | Pause length, sensitivity and speech padding. Both spoken phrases were detected |
| AudioUtils, about -50 dB | 4.02375 s | No settings on the free path. Only a ten-second preview for free |
| CanDoYa, -40 dB | 4.0 s | Minimum pause and padding. Three pauses reported |
| TunePocket, Normal preset | 4.2 s | Threshold, minimum silence, padding and crossfade |
| Calvio, -40 dB | 5.1 s | Minimum silence, padding, retained silence and optional noise reduction |
| SuperUtils, Balanced preset | 4.2 s | Threshold, minimum silence and padding |
| FlowPocket | Not measured | The browser automation could not complete its native file picker |
This is not a universal quality ranking. Each threshold can be tuned, and a shorter file is not automatically a better file. The finding is narrower: every measured rival exposed a loudness threshold, while floi exposed the speech decision and its planned cuts.
Method 1: use speech detection in the browser
This is the useful default for a spoken recording. It gives you an editable map before it produces a file, so the model's judgement is visible rather than final.
Open the recording
Use the paired silence remover. The file is decoded in the tab and is not uploaded.
Find speech
Start with Balanced and 140 milliseconds of padding. Purple regions should cover every spoken phrase.
Shorten before you remove
Set Pause length to 0.5 seconds, then compare Original and Result. Remove pauses only when the material should have no breathing room.
The tool uses Silero VAD 6.2. The official model file is 2,327,524 bytes and MIT licensed. It is fetched from floi's own model domain only when you hover, focus or use the editor.
Method 2: truncate silence in Audacity
Audacity is the better choice when the recording is already in a desktop edit. Its Truncate Silence effect can shorten repeated gaps across a track without deleting them completely.
Select the track
Use the whole track, or test one difficult minute before committing to a long recording.
Open Effect, Special, Truncate Silence
Set the detection threshold below the quietest word but above the room tone. Set a minimum duration so normal word gaps are ignored.
Preview the quiet passage
Listen for clipped first consonants and missing breaths. Undo and lower the threshold if speech is touched.
Audacity notes that Truncate Silence only removes audio. It does not clean the noise inside the silence it keeps. The Audacity manual documents the current controls.
Method 3: automate it with FFmpeg
FFmpeg is the practical choice for a repeatable batch. This command shortens every internal passage under -40 dB that lasts at least 0.5 seconds, while retaining 0.5 seconds:
ffmpeg -i input.wav -af "silenceremove=stop_periods=-1:stop_duration=0.5:stop_threshold=-40dB:stop_silence=0.5:detection=rms" output.wav| Option | What it changes |
|---|---|
stop_periods=-1 | Restarts detection after each kept passage, so internal gaps are processed |
stop_duration=0.5 | Ignores shorter quiet passages |
stop_threshold=-40dB | Defines which RMS level counts as silence |
stop_silence=0.5 | Retains half a second from a longer detected gap |
The names and behaviour come from the official FFmpeg silenceremove documentation. Test one output before starting a batch. A threshold copied from another recording is only a guess.
Method 4: cut a few gaps by hand
Manual editing wins when there are only two or three obvious pauses. Put each boundary in a low-energy point, keep a breath where the delivery needs it, and use a short crossfade if the join clicks.
Use the audio trimmer when you know the exact section to remove. It is also the right answer for music, because it does not assume that non-speech is waste.
How much pause should remain?
Start at 0.5 seconds for ordinary conversation. The value is a ceiling on long gaps, not a command to make every pause identical.
| Result you want | Starting point |
|---|---|
| Natural interview or lesson | 0.6 to 0.9 seconds, with generous padding |
| Podcast with obvious dead air | 0.4 to 0.7 seconds |
| Fast social clip | 0.2 to 0.4 seconds, checked sentence by sentence |
| Speech dataset or prompt list | Remove pauses, if each item is already a separate utterance |
What to check before you export
Listen to the quietest speaker, the first consonant after a pause and the end of a sentence. Those are the three places an aggressive edit reveals itself first.
Breaths
A breath can be rhythm, not waste. Increase padding if every phrase starts too abruptly.
Clicks
A join outside a zero crossing may click. Use a short crossfade rather than a long fade that softens words.
Format
Keep a WAV master. Re-encoding an MP3 is a separate quality decision from removing pauses.