Separate Voice and Music
Split a song into the vocal and the backing track.
Split a song into the vocal and the backing track. Free, with no sign-up, no watermark and no upload: it all runs in your browser.
What Separate Voice and Music is for
Every other way of doing this asks you to upload the track. That is a strange thing to agree to for a song you are working on, and it is why the paid services can charge for it and the free ones queue you behind everyone else. Here the separation runs on your own machine: the models come down once, about 76 MB, and after that the tool works with the network off. What comes out is two files — the voice on its own, and everything else on its own — cut from the same recording, the same length, still lined up with each other.
How to use it
- Add your files. Drop them on the page or choose them. They stay on your device.
- Set the output format and quality. The defaults suit most files.
- Press convert and save the results, individually or as one ZIP.
Questions
How good is it?
Good enough to sing over and good enough to sample, and not good enough to fool anyone listening closely. A separated vocal usually carries a faint ghost of the backing, and the instrumental keeps a trace of the loudest notes of the vocal. Dense, loud mixes are harder than sparse ones, and a track that was quietly mixed comes apart more cleanly than one that was mastered to be as loud as possible.
Is this how you make a karaoke track?
Yes, and it is the most common reason people use it. Keep the instrumental and throw the vocal away. It will not be the official backing track, because that was mixed from the original separate recordings, but for singing along it is usually indistinguishable.
Can I get just the vocal, for a remix?
That is the other output. You get both every time, so take whichever one you came for and ignore the other. The isolated vocal is what people sample, put over a new beat, or feed into a pitch editor.
What format do the stems come out as?
MP3 by default, because a separated stem left as raw WAV is roughly eight times the size of the track that went in, and a 24 MB backing track is awkward to carry around. FLAC is there if you want lossless at about half the size of WAV, and WAV itself is there for anyone taking the stems into a session, where re-compressing before the real work starts is the wrong order to do things in.
How long does it take?
Several times faster than the song is long, on a laptop. A three minute track is well under a minute of work once the models are cached. A phone is slower, and the first run on any device includes the one-off download.
Does it work on a podcast, not a song?
Yes, and it is good at it. Lifting a voice off a music bed is the same problem, and speech is easier for it than singing. It is a reasonable way to rescue an interview that was mixed too hot against its own intro music.