You band-limit it, you collapse it to mono, you put it in the room the camera is standing in, and you deliver it as a separate pass instead of baking it into the song. The band-limiting is the part everybody does: a steep pair of filters that leaves roughly 250 Hz to 5 kHz, which is close to what a dashboard driver can actually reproduce. The part that carries the illusion is the room. A radio playing in a shot is a small loudspeaker in a hard box with glass on five sides, and no equaliser curve contains that.
What is source music, and why is it a different job from score?
Score sits outside the story and gets mixed for clarity against dialogue. Source music comes from something the audience can see or infer in the frame, a car radio, a diner jukebox, a phone on a table, a PA in a corridor, and it gets mixed for geometry. The question stops being how loud the music should be and becomes where the speaker is, how far the microphone is from it, and what is between them.
That is also why source music is the harder of the two to place under dialogue. The futz band, roughly 300 Hz to 4 kHz, is the same band the voice lives in. Score can be moved out of the way with a shelf. A car radio cannot, because the moment you take out the band that collides with the dialogue you have taken out the only band the effect had.
What does futzing actually change?
Three stages, in this order: bandwidth, dynamics, distortion.
Bandwidth first. A high-pass at 300 Hz and a low-pass at 4 kHz, both at 24 dB per octave, gets a generic small speaker. A mobile phone wants the corners tighter and the slopes steeper, a landline wants the low corner nearer 250 Hz and a gentler curve with more low mid in it, and a period wireless set wants a narrow window around 350 Hz to 4 kHz with the filters resonant enough to ring slightly (344 Audio).
Dynamics second. A cheap amplifier driving a cheap driver runs out of headroom early, so a 4:1 ratio with a 1 ms attack and around 6 dB of gain reduction on the loud sections does more for realism than another filter. Transients that survive intact are the giveaway that a mix is a filtered studio recording.
Distortion third, and it is the stage people skip. Two or three percent of harmonic distortion on peaks, from a saturation stage placed after the compressor, is what stops the cue reading as a band-limited file. Without it the ear hears an equaliser. With it the ear hears a loudspeaker being pushed.
Why does an equaliser alone not sell a car radio?
Because filtering removes information and adds none, and the thing the shot needs added is a space.
The reference case is fifty years old and still the clearest. On *American Graffiti*, Walter Murch and George Lucas took the master of the two-hour Wolfman Jack radio programme, played it back on a Nagra through a loudspeaker in a suburban backyard, and re-recorded it with a second Nagra and a microphone, Lucas walking the speaker across the yard while Murch walked the microphone to match. That gave the final mix three versions of the same programme, one clean and two full of yard, so the music could be pushed forward or dissolved into the background depending on what the scene needed (Lucasfilm). Murch called the technique worldizing, and the term is still used with that meaning (filmsound.org).
The modern version costs an afternoon. Print the futzed stem, play it through a single small speaker in a hard-walled room, record it with one microphone at the distance the camera implies, and blend that re-recorded pass under the direct futz at 10 to 30 percent. If there is production room tone from the location, record the pass in a room with a similar decay time so the two do not argue.
What happens to the low end and the stereo image?
Fold to mono before the filters, not after. One speaker in the frame means one source, and a stereo image that survives the futz reads as two speakers in a place that has one.
Order matters through the whole chain. A saturation stage on a stereo pair generates different harmonics on each side, and folding that down afterwards produces cancellation in the exact band the effect depends on. Mono first, then band-limit, then compress, then saturate, then place in the room. Anything below the high-pass corner is gone by design, so the pressure modes of a real car cabin are not being simulated at all, which is fine: the audience is hearing the radio, not sitting on the parcel shelf.
How do you handle the cut from inside the car to outside?
With two printed passes and an automated send rather than a plugin bypass.
The interior pass is the futz plus a short, tight room. The exterior pass is the same futz, filtered harder at both ends, dropped 8 to 12 dB, and fed into a longer, more open space. Cross them across the cut with two to four frames of overlap rather than switching on the frame, and keep the reverb on the futz bus so the tail of the interior sound survives into the first frames of the exterior. Bypassing a plugin at a cut kills the tail and the edit announces itself.
What do you deliver so a dub stage can undo it?
Never bake the futz into the music print. Four files, all at 24-bit, 48 kHz, all starting at the same timecode and all the same length:
1. The clean cue, full bandwidth, unprocessed. 2. The futzed pass as its own file. 3. The re-recorded or worldized pass on its own, unmixed with the other two. 4. A one-page note with the filter corners, the compressor settings and the device the futz was built to imitate.
The reason is that the final relationship between music, dialogue and effects belongs to the re-recording mixer, not to the person who prepared the cue (re-recording mixer). If the futz is inside the music print, that mixer cannot dip the music under a line without dipping the artefacts with it, and cannot widen the cue when the camera pulls out of the car. The same logic governs everything else in a sync package, which is set out in what a sync delivery actually contains, and the dialogue side of the problem is in how music gets ducked under dialogue.
What goes wrong most often?
Loudness, first. Removing everything below 300 Hz takes a large amount of integrated loudness out of a cue, and the reflex is to push the fader on the way into the futz chain, which clips the saturation stage and produces a fizz that sounds like a codec rather than a speaker. Set the level after the chain, never before it.
Sibilance, second. A low-pass at 4 or 5 kHz leaves the band a vocal was already pushed into by a de-esser, and the compressor in the futz chain then rides it. A narrow cut of 2 or 3 dB around 4.5 kHz on the futz bus alone fixes it without touching the clean cue.
Mastered material, third. A cue that arrives as a finished master has had its peak-to-loudness ratio squeezed, so the compression stage of the futz has nothing to grip and the effect stays flat. Ask for the mix.
I have built alternate versions of finished mixes under exactly these constraints, including contest versions for Eurovision entries where the same master had to exist in more than one form. Those credits with their sources are at /credits.
If you have a scene with a radio in it
Send the cue, the picture with timecode, and one sentence about what the speaker is and where the camera sits relative to it. What comes back is the clean cue, a futzed pass, a worldized pass and the settings sheet, in the four-file layout above. Background and contact at /about.