17 July 2026 | 6 min read
Audio ducking: how to keep music under dialogue without riding a fader
Ducking lowers the music whenever someone speaks. Done well it is invisible. Done badly it pumps like a car alarm.
Ducking is automatic level reduction of one track triggered by another. Music drops when dialogue starts, and comes back when it stops. Technically it is a compressor whose detector listens to a different signal, which is why it is also called sidechain compression.
The four controls that matter
- Threshold: how loud the dialogue must be before the music moves.
- Range or ratio: how far the music drops, typically 4 to 9 dB for narration.
- Attack: how fast the drop happens. Too fast clips the first syllable, too slow buries it.
- Release or hold: how quickly the music returns. This is what makes ducking audible.
Sensible starting points
For voiceover over music: 6 dB of reduction, attack around 20 ms, hold 200 ms, release 400 to 800 ms. For drama, less reduction and longer release, so the music breathes rather than pumps. Set the release long enough that short gaps between words do not let the music rise back up.
When to automate by hand instead
Ducking reacts, it does not anticipate. It cannot know that a music cue should swell before a line ends, or that a specific word needs air. In drama and documentary, hand written automation on the music track almost always sounds better. Ducking earns its place in podcasts, live streams, radio and anything long form where consistency beats nuance.
Prepare the trigger track
The detector is only as good as what it listens to. Feed it a clean dialogue stem, not a mix. If the dialogue lives on several channels of a poly WAV, split it out, level the tracks so the detector sees a consistent input, and only then set the threshold.
Split, rename and level those stems here in the browser before they reach your session.