How to remove background noise from audio
Four methods on one recording at four noise levels, with the numbers each one earned. Measured 20 September 2026.
Free toolAudio Noise RemoverTake the hiss, hum, traffic and room echo off a voice recording, with a speech model that runs on your own machine.Open the noise removerThe short answer
Use a speech model. On the measurements below, one improved the speech in a badly noisy recording by 13 dB, while a noise gate improved it by nothing and the usual tutorial chain made it 18 dB worse. The older methods are not simply weaker versions of the new ones: they clean the silence between phrases and leave the noise under the voice exactly where it was.
| Your situation | What to use |
|---|---|
| A voice recording, and you want it done | A speech model. The noise remover here runs one in your browser |
| You are already editing in Audacity | Noise Reduction with a proper noise profile, at about 9 dB |
| You are in Premiere Pro or DaVinci Resolve | The built-in speech enhancer, which is a model of the same kind |
| Hundreds of files, on a command line | ffmpeg with arnndn, or the DeepFilterNet binary |
| Hum or rumble only, and nothing else | A high-pass filter and a notch. Do not reach for a denoiser at all |
The four methods, and what each one is doing
Every technique in this article is one of four ideas. Knowing which one a button is running tells you what it will do to your recording.
A noise gate
A gate is a switch. When the signal is below a threshold it turns the volume down, and when it is above it lets the sound through untouched. It is the effect behind "remove silence" buttons and behind most of what a podcast plugin chain does first.
So a gate cannot reduce noise that happens while somebody is speaking, because during speech the gate is open and the signal is passing through unchanged. What it does is make the gaps silent. That is worth something: hiss that stops between sentences is much less tiring than hiss that never stops. It is not noise removal.
Spectral subtraction
This is what Audacity's Noise Reduction does, and what afftdn in ffmpeg does. You give it a passage of noise on its own, it measures the average level in every frequency band, and then it subtracts that much from every frame of the whole recording.
It works on noise that is steady and predictable, which is why it is good at mains hum and fan whirr and poor at traffic. Push it too far and it produces musical noise: isolated surviving fragments that ring on their own and make a voice sound like it is under water. The method has no idea what speech is, so what it removes from a vowel is the same thing it removes from silence.
A speech model
A neural network trained on thousands of hours of speech mixed with noise. It works out which parts of each frame are voice and keeps those, rather than subtracting an average. Because it knows what speech looks like, it can take noise out from under a word instead of only out of the gaps.
Two are free and widely available. RNNoise is tiny, old and fast, and is what powers the noise suppression in several conferencing apps. DeepFilterNet 3 is larger, newer and works at the full 48 kHz, and it does dereverberation as well as denoising. Both are in ffmpeg or one command away from it.
The stacked chain
Most tutorials tell you to run all three: high-pass the rumble out, then denoise, then gate the gaps. It is worth measuring because it is the most commonly recommended answer on the internet, and because each stage damages the speech a little before the next one starts.
What each one did, measured
One clean speech recording and one noise recording, mixed at four signal to noise ratios, so the only thing changing down each column is how much noise there is. Every method ran on all four mixes. Outputs were delay-aligned before measurement, because a filter with latency is not a worse filter. The full record, including the settings swept for each ffmpeg filter, is in the research notes kept with this site's source, at docs/research/audio-noise-remover.md.
The number is SI-SDR in dB, measured only over the parts where somebody is speaking. Higher is better, and the row to compare everything against is the first one, which is the recording left alone.
| Method | SNR 0 dB | SNR 5 dB | SNR 10 dB | SNR 20 dB |
|---|---|---|---|---|
| Nothing, the recording as it is | 6.87 | 11.87 | 16.87 | 26.87 |
| A noise gate | 6.68 | 11.39 | 15.57 | 20.20 |
| Spectral subtraction | 6.81 | 11.84 | 16.66 | 21.65 |
| High-pass, then denoise, then gate | 4.72 | 7.18 | 8.33 | 8.79 |
| RNNoise | 12.36 | 14.66 | 16.43 | 19.35 |
| DeepFilterNet 3 | 19.84 | 21.63 | 23.28 | 26.91 |
Three things in that table are worth saying out loud.
- The gate and the spectral filter do nothing for the speech. At SNR 0 they score 6.68 and 6.81 against 6.87 for leaving the recording alone. Both are slightly worse than nothing.
- The recommended chain is the worst option in the table. At SNR 20, a mild case, it takes the speech from 26.87 down to 8.79. That is a recording made dramatically worse by following the standard advice.
- The speech models are a different category. DeepFilterNet gains 13 dB at SNR 0 and is the only method that is still improving things at SNR 20.
So what were the gate and the filter doing?
Cleaning the silence. This is the same four mixes, measured on the quietest 100 ms window in each file, which is a gap between phrases. Lower is quieter.
| Method | SNR 0 dB | SNR 5 dB | SNR 10 dB | SNR 20 dB |
|---|---|---|---|---|
| Nothing, the recording as it is | -55.60 | -60.50 | -65.50 | -75.20 |
| A noise gate | -79.80 | -84.70 | -89.50 | -98.60 |
| Spectral subtraction | -62.60 | -73.90 | -109.40 | silent |
| High-pass, then denoise, then gate | silent | silent | silent | silent |
| RNNoise | -66.90 | -69.90 | -72.40 | -78.70 |
| DeepFilterNet 3 | -78.40 | -78.90 | -78.80 | -133.10 |
The gate takes 24 dB off the noise floor between phrases and 0.19 dB off the quality of the speech, in the wrong direction. The stacked chain makes the gaps completely silent, which is exactly why it is recommended so often: play the first two seconds of the result and it sounds transformed. Play a sentence and the voice is worse than when you started.
This is the trap in judging a noise removal tool by ear on its first second. Silence is easy and audible. Noise under a word is hard and is the thing you actually wanted removed.
How to do it, whichever tool you have
In the browser
The audio noise remover on this site runs DeepFilterNet 3 in the page. Drop a recording in and press the button. Nothing is uploaded, and you get the whole file rather than a preview, which is not true of the other browser tools that rank for this.
In Audacity
Find a patch of noise on its own
Half a second is enough. Select it. If there is no such patch anywhere in the recording, this method has nothing to work from and you should use a model instead.
Effect, Noise Removal and Repair, Noise Reduction, Get Noise Profile
The dialog closes. That is expected: you have only taught it what the noise sounds like.
Select the whole track and open the same dialog again
Set Noise reduction to about 9 dB rather than the default 12, sensitivity around 6, and press OK. Listen to a sentence, not to a gap.
Audacity's Noise Reduction is spectral subtraction, the same family as the afftdn row above. Those numbers are from ffmpeg's implementation rather than from Audacity, so treat them as the method's character rather than as a score for Audacity specifically.
In Premiere Pro or DaVinci Resolve
Both now ship a speech model, and it is the one to use rather than the older noise reduction sliders beside it. In Premiere it is Enhance Speech in the Essential Sound panel, with a mix amount you should pull back from 100% if the result sounds boxy. In Resolve it is Voice Isolation on the Fairlight page. Both are doing what the last row of the table does.
On the command line
RNNoise is built into ffmpeg and needs a model file, which the GregorR/rnnoise-models repository publishes:
ffmpeg -i noisy.wav -af arnndn=m=sh.rnnn clean.wavDeepFilterNet ships a standalone binary for macOS, Linux and Windows, and it is the same model this site runs:
deep-filter -D noisy.wav -o out/For hum and nothing else, do not use either. A high-pass at 80 Hz removes rumble without touching a voice, and it is the one piece of the standard chain worth keeping:
ffmpeg -i noisy.wav -af highpass=f=80 clean.wavWhat none of them can fix
| The problem | Why it survives |
|---|---|
| A second person talking | It is speech, so a speech model protects it. You want speaker separation, which is a different model |
| Clipping and distortion | The information was destroyed when it was recorded. Nothing reconstructs it |
| Music under the voice | Not speech, so a speech model removes or mangles it. Clean the voice before the music goes on |
| Heavy reverb in a big hard room | DeepFilterNet reduces it. A cathedral is beyond it, and the early reflections are as loud as the voice |
| Noise louder than the voice | Technically it works, and what comes back is thin, because there was very little voice to keep |
A note on how the browser tool compares to the reference
The noise remover on this site is an independent implementation of the signal path around the published model, so it was checked against the DeepFilterNet project's own command line build. On the project's test pair the two agree at 51.11 dB SI-SDR, which is a waveform RMSE of 1.5e-4.
On the four mixes above, the two agree within 0.1 dB over the speaking parts and the browser version trails by up to 3.8 dB over the whole file. The difference is entirely in the first 100 ms and the last 700 ms, passages where the reference is silent and the input is loud noise, and where the offline export handles the edges differently. It is in the table above as the row that ships, not the better one.