Industry · 2026-10-04 · 5 min read

What Can AI Stem Separation Actually Fix?

Machine de-mixing went from a free research tool to a Grammy-winning Beatles single in four years. Where it is release-grade and where it still breaks.

The honest answer: AI stem separation is release-grade for a surprisingly narrow list of jobs and misleading everywhere else. If you need an isolated lead vocal from a clean stereo master, an instrumental for a sync request when the session files are gone, or dialogue pulled off a location recording, current models deliver results that survive a commercial release. If you need a full multitrack back, they do not. The gap between those two sentences is where catalog owners waste money, so this piece walks the line between them with the specific cases on each side.

How did de-mixing reach commercial releases?

The timeline is short. Deezer open-sourced Spleeter in November 2019, and its pretrained two-stem, four-stem and five-stem models turned separation from a research topic into a free command-line tool, with one practical catch: the standard models reconstructed nothing above 11 kHz, so everything that came out sounded like it had a blanket over the cymbals unless you used the 16 kHz variants. Four years later the same family of techniques, trained and refined by Peter Jackson's WingNut Films team for the Get Back documentary, pulled John Lennon's voice off a late-1970s mono cassette cleanly enough that the Beatles released Now and Then as a single in November 2023, and that recording went on to win the Grammy for best rock performance in 2025. A de-mixed vocal from a cassette demo now sits on an officially released, award-winning Beatles record. That is the ceiling of the technology when a dedicated team trains a model for one specific voice on one specific tape. The floor, which is what you get from a free model on a dense modern master, is lower, and the craft is knowing where your job sits between the two.

What does separation handle well today?

Four jobs come out clean enough to ship.

Instrumentals when the multitrack is gone. A catalog that reaches back a decade always has releases whose session drives did not survive, and a sync request for one of them used to be a dead end. A two-stem split of a lossless master produces an instrumental that holds up under dialogue, which is exactly how sync uses it. The paperwork side of that delivery is in what a sync package actually contains.

Isolated vocals for new versions. A clean lead pulled from a stereo master is usable for an official remix, a duet version, or a translation release where the original vocal returns in the bridge. Vocals are the best case for every current model because the models were trained hardest on them and because a voice sits in a register the ear forgives least, so the training effort went where the scrutiny is.

Dialogue and location audio. Separating speech from room tone, traffic and music bleed is the most mature branch of the family, and tools like iZotope RX's Music Rebalance and its dialogue modules run inside every post house I know of. Podcast producers get the same benefit for free.

Rebalance instead of full separation. The least glamorous use is the most reliable: do not solo the stem, just move it. Lifting a buried vocal 1.5 dB inside a finished stereo file, or pulling a snare down 1 dB in a live recording where the drums ran hot, keeps every artifact masked by the rest of the mix. The separation only has to be good enough to nudge, not to stand naked.

Where does it still break?

The failures cluster in predictable places. Cymbals and sibilance smear first: hi-hats come back sounding underwater, an "s" on the vocal drags part of the hi-hat with it, and the 8 to 12 kHz band is where you audition any engine before trusting it. Shared reverb breaks second: the models cannot decide whether the tail of a vocal plate belongs to the voice or the room, so the instrumental keeps a ghost of the singer and the a cappella arrives half dry. Dense midrange breaks third: piano under voice, stacked rhythm guitars, string pads and synth pads all occupy the same 200 Hz to 4 kHz region, and what comes out is phasey, hollow and unusable in solo. And mono or lossy sources lower every ceiling at once, because the models lean on stereo position as one of their separation cues and on content above 16 kHz that an MP3 already threw away. The Now and Then case looks like a counterexample, a mono cassette, until you remember it took a bespoke model trained for that one task, which is not what a subscription plugin does.

What protocol keeps separated stems out of trouble?

The same discipline as any other repair job, written down. Start from the best source that exists, the original lossless master, never a streaming rip, because every generation of loss becomes an artifact generator. Run two engines and pick per element, not per song: one model wins on the vocal, another on the drums, and nothing forces you to use one output for both. Null-test the recombined stems against the original file; whatever does not cancel is what the process invented, and hearing that residue solo tells you exactly what to hide. Then hide it: keep separated material under new arrangement elements, automate the solo moments down to the few bars that need them, and check the result in mono, where separation artifacts phase-cancel into audibility. The delivery logic for versions built this way, instrumental, clean, TV mix, follows the same sheet as any other release, which I documented in what leaves a mastering session.

What does this mean for anyone who owns a catalog?

It means old releases quietly regained commercial surface area. An instrumental for sync, a clean version for broadcast, a vocal-up version for a film trailer: five years ago each of those required the multitrack, and today a lossless master plus an afternoon of careful de-mixing covers most requests. Contest and television work shows the same shift, because broadcast formats run on backing tracks and alternate versions, a world I described from the inside in how Eurovision stems get delivered. One boundary stays fixed: separation creates audio, not rights. A de-mixed a cappella of someone else's record is still their record, and using it in public needs the same clearance as sampling the original file. The technology moved; the paperwork did not.

What to do with this

Pick one release you own whose multitrack is lost, take the best master you have, and run it through two different engines this week. Compare the instrumentals in the 8 to 12 kHz band, null-test against the original, and file whichever version survives next to the master with its loudness and true-peak numbers logged. The next sync request for that song gets answered in an hour instead of with an apology. The releases this practice comes from are documented with sources on my credits page, and the longer background is on the about page.

I(Questions)

Is AI stem separation good enough to replace lost multitracks?

For an isolated lead vocal or a usable instrumental, often yes, if the source is a clean lossless master. For a full remix that solos every instrument, no. Drums and bass come out usable; dense midrange content like pianos, pads and stacked guitars still smears.

Does separating stems from a release change who owns it?

No. A de-mixed a cappella or instrumental is still the same recording, and the master owner and publishers keep every right they had. Separation is a technical act, so clear the recording the same way you would clear a sample of the original file.

Which separation tool should you try first?

Run the same song through two different engines and compare. The open-source Demucs models and iZotope RX's Music Rebalance fail in different places, and the comparison tells you in one listen which artifacts your material triggers.

Need this on your record?

Start a session →