Comp in two passes and keep them separate. The first pass happens with the singer in the room and produces nothing but marks: which line, which take, in one word. The second pass happens after they leave and is pure assembly. Mixing the two is what produces a comp that is technically correct and emotionally flat, because every edit becomes a negotiation while the person who sang it is watching. Cut on breaths rather than inside words, keep joints between 5 and 15 ms, and stop when the performance is coherent rather than when every syllable is the best available syllable.
Why does a comp go flat when every part of it was the best take?
Because the thing you are choosing between is not syllables, and choosing syllable by syllable destroys it.
A sung phrase has an internal shape. The singer pushes into the fourth word because they were heading somewhere, and the small strain on the sixth word exists because of the push. Take the fourth word from take 2 and the sixth from take 5 and both of them are individually better, while the line no longer goes anywhere. This is the failure mode of every comp built by scrolling through lanes looking for the cleanest version of each moment.
The practical guard against it is a rule about size. The smallest unit you are allowed to take from a different playlist is a phrase, and you break that rule only for a specific fault, a flat note or a cracked consonant, that you can name out loud. If you cannot name what is wrong with the word, you do not replace the word.
How many takes are actually worth recording?
Four full passes, plus punches on the two or three lines that were still not working.
The first pass is usually a warm-up that nobody admits is a warm-up. Passes two and three carry most of the usable material. Pass four is where the singer stops thinking about the microphone. Somewhere around pass six the takes start converging: the singer is copying their own phrasing from twenty minutes earlier, so the versions differ in tuning and not in intent, and comping between them becomes a coin toss dressed up as a decision.
Set the session up so that recording those passes costs nothing in bookkeeping. In Pro Tools, the preference that does this lives in Setup, Operation tab, Recording section, and is called Automatically Create New Playlists When Loop Recording. Avid document the behaviour on their own playlists and comping page: each looped pass lands on its own playlist instead of overwriting the previous one. Logic, Cubase and Reaper all have an equivalent, and the point is the same in every one of them. If saving a take requires a decision, the singer will feel the pause.
How should the takes be named?
By what happened in them, and the naming has to be done during the session, not after.
Left alone, Pro Tools auto-names duplicated playlists with the track name, a period and a number, so a track called Kick produces Kick.01, Kick.02 and Kick.03. That is a fine default for drums and useless for a lead vocal, because by the following morning nothing distinguishes Lead.04 from Lead.07 except opening both.
Rename each lane the moment the pass ends, in two or three words, describing the take rather than the clock: quiet one, pushed chorus, no vibrato, the one where they laughed. Fifteen seconds per pass. A week later that lane list is the only reason you can find the take the artist is asking for over the phone.
Where in the audio should the joint go?
In the breath, and if there is no breath, on the front of a consonant.
Breaths are the natural seams of a sung line. They are quieter than the material either side, they have no pitch to mismatch, and the ear is already expecting a discontinuity there. A joint placed in the middle of a breath is inaudible even between takes with fairly different tone.
When two lines run together with no breath, put the joint on the attack of a hard consonant, a T, K, P or B. The transient masks the splice. What you avoid is the middle of a sustained vowel, where the two takes will differ in pitch by a few cents and in vibrato phase by more, and no crossfade length rescues that. If the only possible joint is mid vowel, the honest move is to take the whole phrase from one playlist and fix the fault another way.
One more thing about breaths: keep them. Comping out every inhale produces a vocal that sounds like it was assembled by a machine, because it was. If a breath is too loud, pull it down 4 to 6 dB with clip gain rather than deleting it.
How long should the crossfade be, and which shape?
Short, and equal power when the two sides sit at different levels.
For a joint inside a phrase, 5 to 15 ms is the working range. Avid's own guidance on fades puts short fades of 5 to 10 ms as the fix for pops and clicks after trimming, which is exactly the artefact a comp joint creates, and sets 10 ms as a sensible default for the Quick Punch crossfade during recording. Anything longer than about 20 ms inside a phrase starts to audibly blend two performances, and blended vibrato has a distinctive wobble that listeners notice without being able to name.
On shape, the same document draws the line clearly: equal power crossfades suit transitions between clips of different amplitude, equal gain suits material where the level difference is small. Two takes of the same singer at the same distance are close enough for equal gain. Two takes recorded twenty minutes apart, one of them noticeably louder, need equal power or the joint will dip in the middle.
In a breath, you can go longer, 30 to 50 ms, because there is nothing tonal to smear.
Why assemble the comp after the singer has left?
Because the two activities need different states of mind, and one of them is expensive in studio time.
The marking pass is fast and the singer is essential to it. Play the song down in Playlist View with the lanes visible, and for each phrase say a take number out loud while they say yes or no. Avid describe the mechanism for hearing an alternate lane: soloing the playlist lane auditions it without muting the rest of the session, so the singer hears the take in context rather than naked. Ten to fifteen minutes for a three-minute song. Write the numbers on paper.
Then the assembly. Copy Selection to Target Playlist moves each marked region into the main playlist at the same timecode, and the whole job is twenty minutes of clicking with no judgement involved, because the judgement already happened. Doing this with the artist present converts every one of those clicks into a conversation, and around the fourth conversation somebody suggests one more take.
What do you do when the best performance is out of tune?
Take the performance. Pitch is cheap and intent is not.
A line that is 25 cents flat but goes somewhere can be corrected in a few minutes with graphical pitch editing, and if the correction is applied to the sustain while the attack is left alone, nobody hears the tool. A line that is perfectly in tune and means nothing cannot be fixed by any process that exists.
The exception worth naming is a note that is flat because the singer ran out of air. That one usually comes with a thinning of the tone in the last 300 ms, and pulling the pitch up leaves the thin tone exposed. Take a different pass for that phrase.
What do you save before you flatten the comp?
The lanes, the map, and one bounce of the raw main playlist.
Duplicate the finished main playlist before you consolidate anything, and name the duplicate with the date. Keep the alternate lanes in the session rather than deleting them to save space, because the artist changing their mind about one line three weeks later is a normal event and not an unusual one.
Then write the comp map somewhere that is not the session file: phrase, playlist number, one word about why. Six lines of text. It costs two minutes and it is the difference between answering a revision request in ten minutes and rebuilding the reasoning from scratch. Sound On Sound's walkthrough of playlist workflow makes the same point about keeping alternates live rather than treating the comp as final.
Frequently asked
How many takes should I record before I stop? Four full passes plus targeted punches. Past six the singer is copying their own phrasing and the takes stop differing in intent.
Should the singer be there while I comp? For the marking pass yes, for the assembly no. Marking needs their memory of what each line meant. Assembly is clicking, and an audience turns each click into a discussion.
What crossfade length should I use? 5 to 15 ms inside a phrase, 30 to 50 ms in a breath, equal power when the two sides differ in level.
If your comp sounds correct and dead
Send me the session with the alternate playlists still in it, not a flattened bounce, and tell me which line you rebuilt the most times. That line is almost always where the shape got lost, and it is usually fixable by putting one whole phrase back rather than by editing further. Full credit list at /credits, background and contact at /about.