Count in bars, cut on downbeats, and spend the last bar on an ending. At 120 BPM a bar of 4/4 lasts exactly two seconds, so a 30-second spot is fifteen bars, and a usable edit is fourteen bars of song plus one bar of composed ending with its reverb tail dying inside the runtime. I composed, produced, mixed and mastered "Stefania" for Kalush Orchestra, with Sony's release metadata on Qobuz listing all four roles, and when I prepare alternate lengths for sync use the arithmetic above is where every edit starts. The craft is in which fourteen bars, and in hiding the seams.
Why count bars instead of seconds?
Because the song only has legal places to cut, and they fall on the grid. Divide 240 by the tempo and you have the length of a 4/4 bar in seconds: two seconds at 120 BPM, 2.4 at 100, 2.67 at 90. From there the standard ad lengths turn into bar budgets. A :30, still the most common television spot length, gives you fifteen bars at 120 BPM but only eleven at 90, and that difference decides whether the edit can afford a four-bar intro or has to open cold on the chorus. A :15 at 90 BPM is five and a half bars, which is why slow songs get half-time edits built from a single phrase.
Odd tempos refuse to divide evenly, and the spot still has to end on time. The fix is to let the music end early rather than late: a 96 BPM song fills twelve bars in exactly 30 seconds, but at 93 BPM I take eleven bars, land the final hit around 28.4 seconds, and give the tail the remaining 1.6. Nobody has ever complained that the music resolved a second before the end card. Music that is still sustaining when the file ends sounds like a mistake every time.
Which fourteen bars survive?
The hook, an entrance, and an ending, in that order of priority. My working order on an edit session: mute everything, find the strongest four-bar statement of the hook, and build outward from it. The verse usually dies first, since a :30 has no time to earn a payoff. What listeners need in the first two bars is the sonic identity of the song, so if the signature of the record is a folk flute line or a specific drum sound, that goes in the opening bar even when the full song makes you wait for it.
Cuts land on downbeats at phrase boundaries, bar 4 to bar 5, bar 8 to bar 9, and never in the middle of a sung word. When two sections refuse to join cleanly, I steal a connective event from elsewhere in the song: a drum fill, a riser, a crash, flown in over the seam. A one-beat fill covers a key-adjacent join better than any crossfade, because it gives the ear a reason for the discontinuity. The crossfades themselves stay short at musical seams, 10 to 30 ms equal-power, placed at zero crossings so low frequencies do not thump at the join. Longer crossfades smear transients and announce the edit they were supposed to hide.
What is a button, and why does every spot need one?
A button is a composed final event: one last downbeat hit, usually the tonic chord with a cymbal, that decays naturally inside the remaining runtime. Broadcast spots do not fade, because a fade under a voice-over tag sounds like the music gave up. I print the button from the record's own material: take the biggest chorus downbeat, clip it with its full tail from the stems, and fly it onto the final bar. If the natural decay runs long, a gentle fade on the reverb tail over the last 1.5 seconds reads as a room dying away. A fade on sustained music reads as an edit.
The same move scales. A :60 cutdown can afford a short outro before the button. A :15 is often just the hook twice and the button. In every length, the button is the one bar I refuse to compromise, for the same reason an edit delivered for film or TV placement always includes alternate endings: the person cutting the picture needs the music to stop on purpose.
Why do edits on stems beat edits on the master?
Because different elements want to be cut in different places. A vocal phrase ends a beat after the chord changes; a bass note sustains across the barline; the drums are the only stem that actually wants the edit exactly on the downbeat. Editing the stereo master forces one compromise cut on all of them at once. On stems I cut the drums on the grid, slip the vocal edit back to the end of the word, and let the bass crossfade over a longer 40 ms span on its own, and the seam that was audible on the master simply is not there. Eight to ten stems is enough: lead vocal, backing vocals, drums, bass, main harmonic instruments, hooks, FX.
Stems also survive contact with the editor. Whoever cuts the spot to picture will move your downbeat to land on a product shot, duck the music under the voice-over, and sometimes rebuild the back half, the same negotiation I described in ducking music under dialogue. A stereo-only delivery means every one of those changes degrades the mix. Stems mean the spot mixer works with the record instead of against it.
How loud should the spot master be?
Programme loudness, measured, and no louder. US television normalizes commercials to the loudness of the surrounding show under the FCC's CALM Act rules, which adopted the ATSC A/85 practice built around -24 LKFS. Broadcast normalization turns a crushed spot master down to the same measured loudness as everything around it, and only the distortion survives the ride down. So my cutdown masters keep the dynamics of the record, peaks controlled to -2 dBTP for the broadcast chain, and the loudness war stays home. Deliver 48 kHz, 24-bit WAV, full mix plus instrumental for every length, labeled with the length in the filename, and the stems beside them.
What to do with this
Editing your own song for a pitch or a placement? Work out your bar length from the tempo, build the edit around the hook, print a real button, and cut on stems. The songs this practice comes from, including all four roles on "Stefania," are documented on my credits page, and the longer background is on the about page.