Separate the vocals from a song with AI

Open your song and choose Effects ▸ Separate Stems…. Auditius runs Demucs, an AI source separation model, on your own computer and gives you the vocals and the accompaniment as new files, ready to play, edit or save. You download the model once, when you ask for it, and the song never leaves your computer.

  • Free
  • No install, no account
  • Files never leave your computer
  • English and Spanish

How to separate the vocals from a song

The first time, the Separate Stems window offers to download the model. After that it is stored in your browser and loads from there.

  1. Click Open a song to separate, or use File ▸ Open… (Ctrl+O), and choose your song.
  2. To separate only part of it, select that part first; otherwise the whole file is used.
  3. Choose Effects ▸ Separate Stems… from the menu bar.
  4. If the window shows a Download button, click it once. The download is about 198 MB.
  5. Tick what you want: Vocals and Accompaniment (drums, bass and other) are ticked by default; Drums, Bass and Other are there too.
  6. Click Separate and wait. A progress bar shows how far it has got, and Cancel stops it.
  7. Each part appears as a new file, such as “My song (Vocals)”. Play it, edit it, or save it with File ▸ Save As… as WAV, FLAC, MP3 or Ogg.

What you get

The model splits a song into four parts: vocals, drums, bass and everything else. Auditius turns them into the files you ticked.

  • Vocals: the voice on its own, for an a cappella, a remix or to study a melody.
  • Accompaniment: drums, bass and other mixed together, which is the song without vocals. Use it as a karaoke track or to sing over.
  • Drums, Bass and Other: each part on its own, to practise an instrument or rebuild the mix in the Multitrack view.
  • The new files keep the sample rate and length of the original, so they line up exactly if you put them back together.

How long it takes

Separation is heavy work. With a graphics card that the browser can use through WebGPU it is much faster; without one it runs on the processor, and a song takes roughly three times its own length, around ten minutes for a three and a half minute song. The window tells you which of the two it will use.

It needs a few gigabytes of free memory while it works. To try it quickly, select thirty seconds of the song first and separate only that.

Tips for a cleaner result

Separation is never perfect: a little of the music can stay in the vocals and the other way round.

  • Start from the best file you have. A WAV or a high bitrate MP3 separates better than a low quality one.
  • Clean up what is left with the usual tools: Noise Gate or Envelope on the vocals, or a little EQ on the accompaniment.
  • Effects ▸ Amplitude ▸ Channel Mixer… has a Vocal Cut (center) preset that removes what sits in the middle of a stereo mix. It is instant but rough; Separate Stems is the one to use when quality matters.

Frequently asked questions

Is my song uploaded to separate the vocals?

No. The model runs in your browser, on your own computer. Only the model itself is downloaded, once, and your audio never leaves the device.

Can I make a karaoke version of a song?

Yes. Tick only Accompaniment (drums, bass and other) and you get the song without the vocals, ready to save as MP3.

Why is vocal separation slow on my computer?

Without a graphics card available to the browser, the model runs on the processor. It still works, but a whole song takes several minutes. Separating a shorter selection is quicker.

Can I use the separated vocals in my own releases?

That depends on the song, not on Auditius: you need the rights to it. Also note that the Demucs model weights are offered for research and personal use, as the models window explains.

Which AI model separates the vocals?

Demucs (htdemucs), a source separation model by Meta that splits music into vocals, drums, bass and other. Auditius runs it with ONNX Runtime in the browser.

Audio tools