Skip to content
All articles

Production

Stem splitting explained: what AI separation can (and can't) do

Pulling a finished song apart used to be impossible. AI makes it routine — with a few limits worth knowing.

SoundBooster Team·· 5 min read

A mixed song is a single waveform: every instrument is added together. Separating it back into stems — vocals, drums, bass and the rest — was long considered impossible to do cleanly. Modern AI models, trained on thousands of songs with their original stems, have learnt what each instrument sounds like and can pull them apart remarkably well.

What people use stems for

  • Instrumentals for karaoke, or a cappellas for remixes.
  • Practising an instrument by muting it and playing along.
  • Sampling a drum break or bassline.
  • DJ edits and mashups.

Where it struggles

  • Bleed — faint traces of one instrument left in another stem, most noticeable in dense arrangements.
  • Similar instruments — piano and guitar, or two vocals, share a range and are harder to split.
  • Heavy effects — long reverbs and distortion blur the boundaries the model relies on.
  • Low-quality sources — a heavily compressed MP3 gives the model less to work with.

Getting cleaner stems

  • Start from the best file you have: WAV or a high-bitrate file beats a low-quality stream rip.
  • Pick the right mode — four stems for most jobs, six when you need guitar and piano apart.
  • Use the stems in context: small artefacts that stand out in solo usually disappear in a mix.

Try it in the stem splitter, then rebalance the stems or send them to the studio. Only separate music you have the right to use.