Google Analytics helps us understand which tools you use and which actions work. Your files and entered text stay out of analytics.

How to remove silence from audio

Speech detection, Audacity, FFmpeg and a manual cut, with six ranking tools inspected on the same file. Measured 26 September 2026.

Free toolSilence RemoverFind speech instead of guessing at a noise level, then shorten or remove the pauses you can see.Open the silence remover

The short answer

Use voice activity detection for spoken audio with room tone, distant voices or whispers. Use a decibel threshold for music or a clean studio recording. The difference is what each method believes silence means.

Your recordingBest starting method
Podcast, interview or voice noteSpeech detection, with 0.4 to 0.7 seconds left between phrases
Quiet speech over steady room toneSpeech detection with a lower sensitivity threshold and generous padding
Music with silent starts or endsA threshold or a manual trim. A speech model is the wrong detector
A few obvious mistakesManual cuts on a waveform, because automation adds risk without saving much time
Hundreds of consistent filesFFmpeg, after tuning one representative file and listening to the result

Why a noise threshold cuts the wrong thing

A threshold tool calls everything below one level silence. That works when the speaker is always loud and the background is always quiet. A real room often reverses that assumption.

A distant word can sit below a fan, while the fan runs through every pause. Raising the threshold risks the word. Lowering it keeps the room tone and finds no gap at all.

DetectorQuestion it asksMain failure
Peak or RMS thresholdIs this frame loud enough?Quiet words and loud background share one line
Voice activity detectionDoes this frame contain speech?Music, singing and unusual voices can confuse it
Manual editingDoes this passage belong?Accurate, but slow on a long recording

Silero's official ONNX wrapper reads 512 samples at 16 kHz for each decision. That is one probability every 32 milliseconds. Its published timestamp code uses a lower exit threshold than entry threshold, which stops speech flickering on and off at one borderline frame. See the Silero VAD implementation.

Six ranking tools and floi on the same recording

I tested the first six relevant organic tool results on 26 September 2026. The input was an 8.950563 second mono recording made for this test, with one normal phrase, one quiet phrase, three long gaps and low pink room tone. The exact sample is the Try the room-tone example file in the paired tool.

Tool and defaultResultWhat you can control
floi, speech detection4.956 sPause length, sensitivity and speech padding. Both spoken phrases were detected
AudioUtils, about -50 dB4.02375 sNo settings on the free path. Only a ten-second preview for free
CanDoYa, -40 dB4.0 sMinimum pause and padding. Three pauses reported
TunePocket, Normal preset4.2 sThreshold, minimum silence, padding and crossfade
Calvio, -40 dB5.1 sMinimum silence, padding, retained silence and optional noise reduction
SuperUtils, Balanced preset4.2 sThreshold, minimum silence and padding
FlowPocketNot measuredThe browser automation could not complete its native file picker

This is not a universal quality ranking. Each threshold can be tuned, and a shorter file is not automatically a better file. The finding is narrower: every measured rival exposed a loudness threshold, while floi exposed the speech decision and its planned cuts.

Method 1: use speech detection in the browser

This is the useful default for a spoken recording. It gives you an editable map before it produces a file, so the model's judgement is visible rather than final.

  1. Open the recording

    Use the paired silence remover. The file is decoded in the tab and is not uploaded.

  2. Find speech

    Start with Balanced and 140 milliseconds of padding. Purple regions should cover every spoken phrase.

  3. Shorten before you remove

    Set Pause length to 0.5 seconds, then compare Original and Result. Remove pauses only when the material should have no breathing room.

The tool uses Silero VAD 6.2. The official model file is 2,327,524 bytes and MIT licensed. It is fetched from floi's own model domain only when you hover, focus or use the editor.

Method 2: truncate silence in Audacity

Audacity is the better choice when the recording is already in a desktop edit. Its Truncate Silence effect can shorten repeated gaps across a track without deleting them completely.

  1. Select the track

    Use the whole track, or test one difficult minute before committing to a long recording.

  2. Open Effect, Special, Truncate Silence

    Set the detection threshold below the quietest word but above the room tone. Set a minimum duration so normal word gaps are ignored.

  3. Preview the quiet passage

    Listen for clipped first consonants and missing breaths. Undo and lower the threshold if speech is touched.

Audacity notes that Truncate Silence only removes audio. It does not clean the noise inside the silence it keeps. The Audacity manual documents the current controls.

Method 3: automate it with FFmpeg

FFmpeg is the practical choice for a repeatable batch. This command shortens every internal passage under -40 dB that lasts at least 0.5 seconds, while retaining 0.5 seconds:

ffmpeg -i input.wav -af "silenceremove=stop_periods=-1:stop_duration=0.5:stop_threshold=-40dB:stop_silence=0.5:detection=rms" output.wav
OptionWhat it changes
stop_periods=-1Restarts detection after each kept passage, so internal gaps are processed
stop_duration=0.5Ignores shorter quiet passages
stop_threshold=-40dBDefines which RMS level counts as silence
stop_silence=0.5Retains half a second from a longer detected gap

The names and behaviour come from the official FFmpeg silenceremove documentation. Test one output before starting a batch. A threshold copied from another recording is only a guess.

Method 4: cut a few gaps by hand

Manual editing wins when there are only two or three obvious pauses. Put each boundary in a low-energy point, keep a breath where the delivery needs it, and use a short crossfade if the join clicks.

Use the audio trimmer when you know the exact section to remove. It is also the right answer for music, because it does not assume that non-speech is waste.

How much pause should remain?

Start at 0.5 seconds for ordinary conversation. The value is a ceiling on long gaps, not a command to make every pause identical.

Result you wantStarting point
Natural interview or lesson0.6 to 0.9 seconds, with generous padding
Podcast with obvious dead air0.4 to 0.7 seconds
Fast social clip0.2 to 0.4 seconds, checked sentence by sentence
Speech dataset or prompt listRemove pauses, if each item is already a separate utterance

What to check before you export

Listen to the quietest speaker, the first consonant after a pause and the end of a sentence. Those are the three places an aggressive edit reveals itself first.

Breaths

A breath can be rhythm, not waste. Increase padding if every phrase starts too abruptly.

Clicks

A join outside a zero crossing may click. Use a short crossfade rather than a long fade that softens words.

Format

Keep a WAV master. Re-encoding an MP3 is a separate quality decision from removing pauses.

Frequently asked questions

What is the best way to remove silence from audio?

Use speech detection for a podcast, interview or voice note. It can keep a quiet word above steady room tone because it is looking for speech, not merely sound above a decibel threshold. Use a threshold tool for music or a controlled studio recording where quiet really does mean empty.

How much silence should I leave between words?

Do not change the gaps between words inside a sentence. For longer pauses between phrases, 0.4 to 0.7 seconds is a useful starting range. The right value depends on delivery. Instructional speech usually needs more room than a fast social clip.

Can Audacity automatically remove pauses?

Yes. Select the track, then use Effect, Special, Truncate Silence. Set the detection threshold and minimum duration, then choose how much of each long silence remains. Preview a passage with quiet speech before applying it to the whole track.

How do I remove silence with FFmpeg?

Use the silenceremove filter. The command on this page shortens every passage below -40 dB that lasts at least 0.5 seconds, while retaining 0.5 seconds. It is fast and scriptable, but its threshold still needs tuning for each recording.

Does removing silence reduce audio quality?

The cuts do not have to reduce sample quality, but the export format can. A WAV export is not lossy. Re-encoding an MP3 creates another lossy generation. A bad boundary can also click or clip a breath, which is why padding and a short crossfade matter.

Should I remove every pause from a podcast?

No. Pauses carry emphasis, separate ideas and give a listener time to follow. Shorten unusually long dead air first. Removing every gap suits datasets and rapid clips more than natural conversation.

↑↓ move↵ openesc close