A compressor ducking straight off a dialogue track usually gets the timing backward: the loudest chord or drum hit tends to land exactly when the vocal is loudest too, so the music ducks hardest at the moment it was already covered and barely moves the rest of the time. A gate set up on a send, not an insert, gives control a plain compressor does not: threshold, release and depth set independently of how loud the dialogue happens to be. What follows is the mechanics of that setup, the narrow band worth cutting instead of the whole track, and the loudness number a streaming deliverable actually gets measured against.
Why does a plain compressor duck the wrong amount at the wrong moment?
Because gain reduction on an insert compressor tracks the level of whatever triggers it, and dialogue level is not constant. Feed a compressor on the music bus from a dialogue sidechain and a shouted line ducks the music hard, a mumbled one barely ducks it at all, and the two effects trade off in a way that draws attention to itself instead of hiding. Sound On Sound's Mike Senior, writing about the same mechanism in a different context, sets the compressor version up with the threshold pulled down toward -60dB, a ratio as gentle as 1.04:1, and a release around 10ms, aiming for no more than 2dB of gain reduction on the meter. Even tuned that carefully, a compressor-based duck stays tied to how loud the dialogue happens to be take by take.
What does a gated send do differently?
It stops caring how loud the dialogue is and starts caring only whether dialogue is present. Route the music to an effects channel, trigger a gate on that channel from the dialogue track, and set the gate to open, not close, when dialogue starts. Senior's method inverts the gated channel's polarity so it cancels part of the direct signal instead of muting it outright, which avoids the on/off snap a gate produces on its own. A threshold around -40dB and a release near 40ms are reasonable starting points. From there, the depth of the duck is set by the send level rather than by a compressor ratio: a send around -20dB yields roughly 1dB of gain reduction, and one around -11dB yields roughly 3dB, so the two ends of a usable range sit a short fader move apart.
Why cut a narrow band instead of the whole track?
Because most of what makes dialogue unintelligible under music lives in a specific, well-documented range. Consonants, the part of speech that actually carries word recognition, sit mostly between 500Hz and 4kHz, and the octave centered on 2kHz alone accounts for close to a third of perceived intelligibility. DPA's technical reference on speech puts a number on the fix: cutting background music 5 to 10dB in the 1 to 4kHz range, rather than turning the whole track down, is what actually protects intelligibility when music sits under narration or dialogue. A full-band duck takes out low end and air along with the part that was actually competing with the voice. A band-limited one leaves the rest of the mix alone and removes the fight only where it was happening.
What loudness number does the finished mix actually have to hit?
Netflix's own sound mix specification sets dialogue-gated loudness at -27 LKFS, plus or minus 2 LU, measured under ITU-R BS.1770-1 across the entire program, with true peak held near -2dBFS. Dialogue-gated means the meter only counts the sections where dialogue is present when it calculates the average, so a mix can run louder in its music-only passages and still pass, as long as the gated dialogue measurement lands in range. There is a second wrinkle worth knowing before it causes a failed QC pass: when dialogue makes up less than 15 percent of a program, the spec drops dialogue-gating and measures the whole thing at a flat -24 LKFS instead. An action cue or a mostly-instrumental passage is being judged against a different number than a dialogue-heavy scene, on the same deliverable.
What actually goes wrong when this step gets skipped?
The note that comes back is almost never about tone. It is some version of "can't hear the line," which is a level and frequency problem wearing a performance complaint as a disguise. Riding the whole music fader down by ear on every pass is the usual workaround, and it does not survive a revision: the next picture cut changes where the dialogue sits, the manual rides no longer line up, and someone redoes the same work from scratch. A gate tied to the dialogue stem keeps working after a picture change because it responds to whatever dialogue is actually there, not to where it used to be.
Frequently asked
Does ducking mean the music always gets quieter whenever someone talks? Not by a fixed amount. A gate or compressor tuned as above reduces gain only enough to clear the 1 to 4kHz band dialogue needs, typically 2 to 4dB, not a blanket cut.
Is dialogue-gated loudness the same measurement as ordinary program loudness? No. Ordinary loudness measurement averages the whole program. Dialogue-gated measurement only counts the sections where dialogue is present.
Does this need special hardware? No. A gate and a compressor with a side-chain input cover it.
If a cue needs this before delivery
Send the dialogue stem separately from a rough combine, and the duck gets set against the dialogue that is actually there rather than guessed from a temp mix. Full credit list at /credits, background and contact at /about.