How to Remove Sibilance Without Dulling the Vocal
Split-band de-essing settings by voice type, where the de-esser belongs in the chain, and how to prove you removed esses and not the vocal.
Use a split-band de-esser, not a wideband one: it reduces only the band above the crossover, so the body of the voice never ducks. Set the crossover between 5 and 8 kHz, take 3 to 6 dB off the worst esses, and audition the removed signal to confirm you are catching esses and not words. Dullness comes from wideband gain reduction, not from de-essing itself.
Disclosure: we build audio plugins at Ghostnote, including a free de-esser. Its measurements appear below where they answer a question; the tool itself sits in one labeled section near the end. No affiliate links on this page.
Work in this order. Each step makes the next one smaller.
- Move the mic. Ten to fifteen degrees off-axis. Costs nothing and prevents the problem instead of masking it.
- Clip-gain the worst esses by hand. Usually a dozen or so edits per verse, 3 to 6 dB each.
- Split-band de-esser. The main tool. Set the crossover by voice type (table below) and keep gain reduction at zero through vowels.
- Dynamic EQ. For sibilance parked in one narrow band, where a whole-band split takes too much with it.
Get VEIL free
Our de-esser for rap vocals. Free, keyless, no account. Enter an email and the download appears right here.
A copy goes to your inbox. Unsubscribe anytime.
VEIL is yours. Pick your build.
What settings remove sibilance without dulling the vocal?
These are the starting points I dial in, then adjust by ear.
| Voice | Crossover start | Reduction on worst esses | Release | Watch for |
|---|---|---|---|---|
| Deep male rap vocal, close mic | 5–6 kHz | 3–5 dB | 50–80 ms | "sh" and "ch" sit lower than "s" and can slip under the crossover |
| Mid-range male | 6–7 kHz | 3–6 dB | 40–80 ms | Soft consonants mean the crossover dropped too low |
| Bright or female vocal | 7–8 kHz | 2–5 dB | 30–60 ms | Short, fast esses need a quicker release or the tail pumps |
| Ad-libs and doubles | Same as the lead | 1–2 dB less than the lead | 40–70 ms | Stacked layers add sibilance, so treat each layer, not the bus |
Set the crossover before the amount. Then set the threshold by the meter, not by the number: lower it until gain reduction moves on esses only and rests at zero through vowels and breaths. On a bright condenser into a top-heavy chain, use two passes of 2 to 4 dB rather than one pass of 8. A single heavy pass lisps where two gentle ones do not.
Why does de-essing make a vocal sound dull?
Because most of the time the tool is turning down the whole voice, not the ess.
A wideband de-esser is a compressor with a high-frequency filter on its detector. It listens up top, but the gain reduction lands on the full-range signal, so the 200 Hz body of the voice ducks with every ess. On a verse with an ess every half second, the vocal pumps and reads small. Push the fader to get the presence back and the esses come back with it.
A split-band de-esser divides the signal into a low and a high band, reduces gain on the high band only, then sums them again. The body cannot move because nothing is happening to it.
The second cause is a crossover set too low. Drop to 4 kHz and you are inside the presence region that carries intelligibility, along with the "t" and "k" transients that make consonants land. Sibilance is a band, not a point. In most adult voices the energy of an "s" and the other sibilant consonants sits somewhere in the 5 to 10 kHz span, while "sh" and "ch" land lower, closer to 2 to 4 kHz. Set the crossover for the "s" and fix a hard "sh" with clip gain.
The third cause is quantity. Take more than about 6 dB off every ess in a phrase and the "s" starts to read as "th".
How to remove sibilance: the four fixes in order of preference
1. Fix it at the mic
Free, and part of why professional vocals need less de-essing. Rotate the vocalist 10 to 15 degrees off-axis so the mouth is not firing straight into the capsule, or set the diaphragm above mouth height and let them rap past it. For a consistently sibilant rapper, reach for a dynamic mic instead of the brightest condenser in the locker. Sibilance spikes when someone pushes air on a punch-in, so a retake at the same energy as the rest of the verse beats an hour of plugins.
2. Clip-gain the worst esses by hand
A rap verse has maybe eight to fifteen esses that offend. Split the clip around each one and pull it down 3 to 6 dB with 20 to 30 ms ramps either side so the edit does not click. Ten minutes of that and every plugin downstream has an easier job. You are making a loud syllable quieter, the way engineers did it by riding a fader.
3. A split-band de-esser
What separates a good one is not knob count. Three things do: the band split recombines flat, so the plugin is inaudible when nothing is being reduced; the detector is level-independent, so a whispered line and a shouted line get the same treatment; and there is enough lookahead that the gain is down before the transient arrives.
This is why the crossover filter matters. A fourth-order Linkwitz-Riley crossover sums back to a flat magnitude response, so with the de-esser idle the vocal comes out as it went in. A sloppier split leaves a dip or bump at the crossover point, which is dullness you did not ask for.
4. Dynamic EQ
Reach for this when sibilance sits in a narrow band. A dynamic bell at 7 kHz with a fast attack leaves the rest of the top end alone, whereas a de-esser moves everything above the crossover together. The trade-off runs both ways: a bell can miss sibilance that drifts between syllables, which a band split catches by design. If your DAW has no dynamic EQ, TDR Nova is free and does the job. What fails is a static shelf cut: a fixed 4 dB dip at 7 kHz dulls every syllable, not just the harsh ones.
How do you know you did not dull the vocal?
Listen to what you took out. That is the whole test.
Some de-essers include a mode that plays only the removed signal. Solo it across a full verse. It should sound like a track of hiss and spit, "sss", "ts", "sh", with silence between. If you hear vowels, words, or the rhythm of the performance, you are turning the vocal down rather than the esses. If your plugin has no such mode, duplicate the track, invert the polarity of one copy, and process only the other. The difference is what you removed.
Then run three checks:
- Level-matched bypass. Match output to input before you A/B, or the louder version wins on loudness alone and you learn nothing. Mastering engineer Ian Shepherd covers this at Production Advice.
- Body and presence. Listen across 200 to 400 Hz and 2.5 to 5 kHz with the de-esser in and out. Nothing there should change. If it does, your tool is wideband.
- Phone speaker. Small drivers push the top end forward, so a vocal that sits fine on monitors can spit on a phone. An over-de-essed vocal sounds hollow there first.
Hold any de-esser to that standard, ours included. We built VEIL's probe suite around those questions, and it reports the answers as numbers: 300 Hz body at −0.00 dB while 4.6 dB of de-essing was active, and 0.00 dB of gain reduction on a bright vowel. An "ss" at −9 dBFS and at −29 dBFS both drew exactly 10.0 dB of gain reduction, so a whisper and a shout are treated identically.
Where does the de-esser go in the vocal chain?
Before the compressor, and often a second time before saturation.
A compressor with a 1 to 3 ms attack sees an untreated ess as the loudest thing in the phrase, so it triggers on sibilance instead of words. You get heavy gain reduction on the esses, barely any on the vowels, and a vocal that ducks on every "s". De-ess first and the gain reduction follows the performance again, so you need less of it. Numbers for that stage are in our guide to vocal compression settings for rap.
The second pass earns its place because saturation generates new harmonics from whatever you feed it, so a little sibilance going in becomes more coming out, higher up the spectrum. An air shelf causes the same trouble by a different route: boosting 10 kHz puts back part of what you took away, so de-ess after the boost, or boost less.
At the master, sibilant transients often set the true-peak reading, which is measured on an oversampled signal per Annex 2 of ITU-R BS.1770. An over-bright vocal costs you limiter headroom across the whole record. High-frequency detail is also expensive for a lossy codec, part of why Apple's Digital Masters program cares how a master survives encoding.
For numbers by delivery style, see de-esser settings for rap. For how this stage sits with the others, read our rap vocal chain guide.
When is a de-esser the wrong tool?
Plenty of harshness is not sibilance, and pointing a de-esser at it is how vocals get dull for no benefit. Diagnose first.
| What you hear | What it is | The right tool |
|---|---|---|
| "S" and "ts" jump out, vowels are fine | Sibilance | Split-band de-esser |
| Certain vowels honk around 2 to 5 kHz | Resonance from the mic, room or voice | Dynamic EQ notch, or a resonance suppressor |
| Harsh across every word of the take | Chain or mic choice, or a bright master | Fix it upstream, not on the vocal |
| Crackle and edge on loud syllables only | Clipping at the preamp or interface | Re-record with more headroom |
| Thumps on "p" and "b" | Plosives | High-pass filter or clip gain |
Row two is the one worth spending money on. Broadband resonance suppression is a different problem from sibilance, and oeksound's soothe2 is the best-known tool for that job: it tracks the spectrum continuously and dips whatever sticks out, wherever it is, which a fixed-band de-esser cannot do. It costs more than a dedicated de-esser, and a dynamic EQ gets part of the way if you can find the frequency by ear. Among dedicated de-essers, FabFilter Pro-DS has been a reference point for years, with metering clear enough to show which syllables it is catching.
From Our Own Line
We make a de-esser called VEIL and it is free, so this section is short.
De-esser · Free forever
Ghostnote VEIL
Split-band by design, so the body of the voice cannot be touched. Our probe measured 300 Hz at −0.00 dB while 4.6 dB of de-essing was working, and an "ss" at −9 and −29 dBFS both got exactly 10.0 dB. The full plugin, free, keyless.
Get VEIL free →An LR4 Linkwitz-Riley split with an allpass-flat sum puts gain reduction on the high band only, so the vocal body cannot be touched by construction.
| Stage or control | What VEIL does |
|---|---|
| Crossover | 3–10 kHz, default 5.5 kHz |
| Detection | High-frequency envelope against full-band envelope in dB, so it is level-independent |
| SENSE | Relative threshold of −2 to −14 dB, 4 dB soft knee, 1.3 slope |
| RANGE | Caps reduction, 1–24 dB |
| Lookahead | 3 ms (144 samples at 48 kHz), so the gain is in place about 2 ms before a "ts" reaches the output |
| Attack and release | 1 ms gain attack; program-dependent release from a 20–200 ms base, up to 3× on a sustained "shhh" |
| Transient probe | A 30 ms "ts" burst: 8.8 dB caught, click-free |
| LISTEN | Output, Sibilants, Removed. Removed plays only what the plugin took away |
The limit: VEIL is a dedicated de-esser, not a broadband resonance suppressor. If your problem is row two of the table above, buy soothe2 and use VEIL for the esses. VEIL is VST3 on macOS and Windows, AU on macOS, and keyless: no license server, no machine limit. The rest of the line is on our plugins page.
Frequently Asked Questions
What frequency is sibilance?
For most adult voices the energy in an "s" runs roughly 5 to 10 kHz, while "sh" and "ch" sit lower, nearer 2 to 4 kHz. Deeper voices center lower in that span, brighter voices higher. Sweep a narrow boost through the top end on a sibilant word to find yours.
Why does my vocal sound lispy after de-essing?
You removed too much, or your crossover is too low. Past about 6 dB of reduction on every ess, an "s" starts sounding like a "th". Back the amount off until the lisp goes, raise the crossover a few hundred hertz, and split the work across two gentle stages.
Should the de-esser go before or after the compressor?
Before. A compressor with a fast attack triggers on untreated sibilance instead of on the words, so it ducks the whole vocal on every "s", and de-essing first lets the compressor follow the performance. A second light de-esser before saturation helps too, since saturation generates more high-frequency energy.
How much gain reduction should a de-esser show?
Three to six decibels on the worst esses, and nothing at all on vowels and breaths. Continuous gain reduction through a whole phrase means your threshold is too low and you are compressing the top end of the entire vocal rather than catching sibilance.
Can I fix sibilance with EQ instead of a de-esser?
A dynamic EQ can, because it only moves while the sibilance is present. A static shelf or bell cut cannot, since it dulls the whole vocal to fix a handful of syllables. With no dynamic tools at all, clip gain on the offending esses beats a fixed cut.
Take VEIL with you
The de-esser this site was built around. Free, keyless, yours on every machine you own.
A copy goes to your inbox. Unsubscribe anytime.
VEIL is yours. Pick your build.




