A de-esser removes the loud, harsh part of an s or a t and leaves the rest of the vocal alone. A dull vocal is almost always the opposite mistake: a fixed high cut that pulls down every bright syllable, not only the sharp ones. The fix is to treat sibilance as a level problem in a narrow band, then aim a tool that only acts when the ess is actually there. Below is the band to work in, the two tools that do it, how to aim them, and how much to take before the singer starts to lisp. I do a lot of vocal production, and this is the move that comes up on nearly every lead.
What is a de-esser actually removing?
Sibilance is the burst of high-frequency energy in s, sh, z, t and ch sounds, and the technique for taming it has a name of its own (de-essing). It sits high: roughly 5 to 10 kHz, with male voices peaking around 5 to 6 kHz and brighter female voices closer to 7 to 8 kHz (Sound On Sound on vocal de-essing). A close condenser with a presence lift above 5 kHz, a bright preamp, or a singer who leans into the capsule all push those bursts past the rest of the words. The job is to pull the peak of each ess down by a few decibels for the 30 to 80 milliseconds it lasts, then let go.
Why does a static high shelf make the vocal dull?
A shelf or a fixed cut at 7 kHz is working all the time. It lowers the esses, and it also lowers the breath, the consonant transients, and the air that makes a vocal sound close and present. You solve one harsh syllable and lose the top of the whole take. Sibilance is intermittent, a few frames per line, so the tool has to be intermittent too: reduction only in the moment the ess crosses a threshold, and unity gain the rest of the time.
De-esser or dynamic EQ?
Both run on the same principle, a band that turns down only when it passes a threshold. A dedicated de-esser is faster to set: one band, a frequency, a threshold, and a listen button that solos the detected sibilance so you tune by ear. A dynamic EQ is the better choice when the sibilance moves, because a bright line can peak at 6 kHz on one word and 8 kHz on the next, and one fixed band will miss half of it (iZotope on the dos and don'ts of de-essing). On a plain pop vocal I reach for a single-band de-esser. On a busy top line with a lot of movement I use a dynamic EQ with a wider band centered on the worst frequency.
How do you find the sibilant band?
Solo the detection. Every serious de-esser and dynamic EQ has a way to hear only the band it is acting on, and the iZotope, FabFilter and stock plugins all label it. I sweep a narrow boost first to find where the ess is loudest, then flip that same band to cut and to dynamic, so the boost becomes the trigger. On most lead vocals the worst energy is a 1 to 2 kHz zone somewhere between 6 and 8 kHz. Set the band too low, near 4 kHz, and you dull the body of the consonant so the word starts to sound like a lisp. That mistake is the single most common cause of a de-essed vocal sounding wrong.
How much reduction is enough?
Aim for 3 to 6 dB on the loudest esses and less on the rest. The gain-reduction meter should twitch on the s and sit still on the vowels. If the de-esser is pumping the whole word, the threshold is too low or the band too wide. On a very bright or badly tracked vocal I run two gentle de-essers in series, each taking 2 to 3 dB, instead of one taking 8, because two small cuts sound more natural than one deep one. That is the same logic behind serial compression on a vocal: a stack of small moves reads cleaner than a single large one.
Where does de-essing sit in the chain?
After the main compression, not before it. A compressor with a fast attack raises the level of the esses against the rest of the word, so de-essing first gets partly undone downstream. I put a broad compressor for level, then the de-esser, then any tone EQ and the limiter. If a vocal has been tuned hard I check sibilance again afterward, because pitch correction can sharpen the s (how much pitch correction a lead vocal needs).
When is the problem the mic or the singer, not the mix?
Sometimes the ess is baked into the recording. A cardioid condenser 6 inches from the mouth with no pop filter and a hard 8 kHz presence lift will print sibilance no plugin fully repairs later. The cheap fixes live at the source: move the singer a few degrees off axis, back off two inches, or change to a microphone with a flatter top octave. Fifteen seconds of angle at the stand saves twenty minutes of automation in the mix.
Frequently asked
What frequency is vocal sibilance? Usually 5 to 10 kHz. Male voices tend to peak around 5 to 6 kHz, brighter female voices around 7 to 8 kHz. Solo the detection band and sweep to find the exact spot rather than guessing at a fixed number.
Why does my vocal sound dull after de-essing? The band is set too low or the threshold too aggressive, so it cuts the body of the consonant and the air, not only the harsh peak. Raise the frequency, narrow the band, and reduce less.
Should I de-ess before or after compression? After. A fast compressor raises the level of the esses relative to the rest of the word, so de-essing before it gets partly undone. Compress for level first, then de-ess.
If your vocal is fighting its own top end
Send the lead and one line on what bothers you, the harshness or the dullness. What comes back is a plain answer on whether it is a de-essing move, a microphone choice, or a compression-order problem, and what to change first. Background and credits at /about.