How to Mix Background Music Under a Voiceover: Levels, Ducking, and Timing

How to Mix Background Music Under a Voiceover: Levels, Ducking, and Timing

I delivered a three-minute product video and got one line back from the client that day: “Great music — but when I played it at the office, my coworkers turned around.”

Then I metered: the music sat 4 dB under the voiceover. Not a bed — a duet. By hour three my ears had decided that level was normal, so pulling it down felt wrong. I was still 8 dB away from right.

Mixing does not care how you feel; it cares what the meter says. Here are the numbers I use: bed levels, ducking, tempo against speech rate, royalty-free sources, and the license traps I keep seeing.

The Level Chart I Keep Pinned

Decibels are relative, so I only think in “how far below the voice.” I park voice peaks between -6 and -3 dB, which puts the music near -21 dB — rarely far off.

Situation Music vs. voiceover The mistake I made
General narration bed 15 to 20 dB lower 4 to 5 dB lower
Data line, key claim 20 to 25 dB lower No change at all
No-dialogue transition 3 to 6 dB lower 15 dB lower, music vanished
Opening intro, before speech 6 to 10 dB lower Full level
Outro 8 to 12 dB lower, then fade Hard cut

Fifteen and twenty decibels are different animals. At -15 the mood survives; at -20 the music is an outline and attention goes to the voice. I mix dense tutorials at 20 and vlogs at 15.

Ducking: Two Ways to Make the Music Move

Ducking, plainly: the music gets quieter when the person talks. I use both.

Manual keyframes — in Premiere, DaVinci, CapCut. Drop the music 6 to 10 dB starting 0.3 seconds before the voice, bring it back 0.8 to 1 second after. Tedious but controllable — key lines get manual moves.

Sidechain compression — a compressor on the music bus, keyed from the voice track:

  • Threshold: low enough that the voice triggers it every time
  • Ratio: 3:1 to 4:1
  • Range: -6 to -10 dB, how deep the music dips
  • Attack 0.2 to 0.3 seconds, release 0.8 to 1.5 seconds

Attack too fast and the music gets pinched; too slow and it eats the first half of the sentence. Release too short and the bed springs back with a thud; too long and everything feels muffled. My first year I ran attack at 0.05 seconds and wondered why the music sounded breathless.

Waveform of a ducking keyframe in my editor, the music track dipping where the voiceover starts

Tempo, Speech Rate, and the First Three Seconds

The music needs more than low volume; it needs to agree with the rhythm of speech.

BPM is the fastest handle: one beat every half second is 120 BPM. Comfortable English narration runs 140 to 160 words a minute with stressed beats about every half second, so 90 to 110 BPM sits underneath. Drop 128 BPM dance music under a long sentence and you get a tug-of-war: music hurries, voice strolls.

Three habits I never skip:

  1. Three seconds of music intro before the first word. The ear needs to accept the key and mood before someone speaks, or the opening feels blunt.
  2. Cut music sections on script sections. When the narration moves to a new chapter, move the music to a new section. Nobody can say why it feels good. It just does.
  3. Bring the music back in the gaps. Over b-roll and title cards with no narration, lift the bed from 18 dB down to 3 or 6 dB down. More than 1.5 seconds of dead silence loses scrollers.

Royalty-Free Music: Sources and License Traps

Royalty-free does not mean free. You pay once, or nothing, and never pay per use again — it says nothing about quality. Sources I actually use:

  • Artlist and Epidemic Sound: subscriptions with full libraries, ideal if you ship weekly. Read the cancellation terms — whether old downloads survive into new videos differs by plan.
  • YouTube Audio Library: free and clearly labeled, but the license protects you on YouTube. Move the file to a client’s website or an ad and you are on your own.
  • Pixabay Music and Free Music Archive: free for commercial use, uneven quality. Check for CC0 or an explicit commercial clause.
  • Single-track purchases (AudioJungle and friends): for campaigns that need a paper trail.

Three traps: songs from Spotify or Apple Music are never cleared for video; free tracks that require attribution go in the description; and Content ID claims burn days to appeal. Source only from libraries with a paper trail.

Six Mistakes I Still See Every Week

  1. The music is too full. A pop track with vocals and drums fights you at any level. Beds want atmosphere without a lead: pads, minimal electronica, instrumental lo-fi.
  2. The vocal band gets buried. Voice intelligibility lives between 200 Hz and 4 kHz, and so does most music energy. High-pass at 80 Hz, cut 2 to 3 dB at 200 Hz, cut 3 dB across 2 to 4 kHz. It steps backward, fader untouched.
  3. Obvious loop seams. An eight-second loop played thirty times clicks audibly around take twenty. Loop in 8- or 16-bar phrases, cut on the downbeat.
  4. One volume for the whole video. If the intro and the payoff are equally loud, you have no payoff. Level is typography for the ear.
  5. Lyrics in the bed. Listeners process two language streams at once and retain neither.
  6. A hard-cut ending. Half a second of fade, minimum.

TTS Voiceovers Mix Differently

A human take carries a noise floor, breaths, and room reverb — all that dirt gives the music something to blend with, so a bed at 14 dB down still sounds glued. A TTS take is surgically clean: no floor, no breaths, dead-consistent level. Two consequences:

  • Push the bed lower. Swapping a human take for TTS, I move the bed from 15 dB down to 18 or 20 dB down. Nothing about the music changed; with no noise floor to hide behind, anything louder climbs over the voice.
  • Add the breaths yourself. TTS does not inhale. I leave 0.3 to 0.5 seconds between sentences so the groove fills the gap, and 0.5 to 0.8 seconds before key lines. Listeners cannot name what improved; they just say the pacing is smooth.

The upside: a TTS track is level-consistent, so your keyframes line up on every sentence instead of drifting like a human take.

My pass takes six steps: peak the voice between -6 and -3 dB; drop the music in at -15 dB; EQ the low end and the vocal band; set ducking from the table above; three seconds of intro before line one, one second of outro after the last; then listen on headphones, speakers, and a phone. The phone is honest — low end gone, and if the music still reads clearly, it is too loud. Steps one through five take five minutes; step six has saved three deliveries.

When the Numbers Are Right, Your Ear Agrees

Back to that video: I moved the bed from 4 dB under the voice to 17 dB under, same song, same narration. The second reply was “this one I can play at work.” The song had not changed; it had gone back to being furniture.

Bed at 15 to 20 dB down, ducking at 6 to 10, three seconds of intro, gaps lifted to 3 or 6 dB down. Start there, then bend it into your own habit. If your voiceover comes from text-to-speech, this gets easier still: a clean vocal track means the numbers either work or they do not. I generate mine on Soundwaver — plans are on /pricing/, or grab a clean take at /tts/ and run tonight’s numbers through it.

Share this post
XFacebookLINE