Vocal Remover
Take the singer out and keep the band.
or paste one from your clipboard
This tool needs a 64 MB download, once
It is the separation network itself, and it is served from this domain rather than a third party. It is kept on your device afterwards, so this happens once and then the tool works offline forever. Separation then runs at roughly real time, so a four-minute song takes a few minutes.
One time only. Nothing is uploaded — this is the model coming down to you.
Model ready on this device. Works offline, and nothing you load here leaves the browser.
Solo a track to hear it alone. What you can hear is what gets mixed down.
How to use it
- Drop the song onto the page. The first time you use this it fetches a 64 MB model, once, and keeps it — after that the tool works with the wifi off.
- Press Separate. A neural network goes through the track in chunks and the progress bar counts them, so you can see how long is left rather than guessing.
- Listen before you commit. You get both halves as a small mixer: play them together, mute the vocal to check it is properly gone rather than merely quieter, or solo it to hear exactly what was taken out.
- Pick a format and download. Choose WAV if this is going into an editor; MP3 is fine if it is going straight into a video or a karaoke night.
What it is actually doing
A mix is a sum. Once a voice and a guitar have been added together into one waveform, no amount of arithmetic recovers the originals, because an infinite number of pairs add up to the same total. So separation is not undoing the mix — it is a guess, made by a network that has been shown several thousand songs where both the mix and the separate parts were known, and has learned what a voice tends to look like when it is drawn as a picture of frequency against time.
That picture is the key to the whole thing. The track is chopped into short overlapping windows, each window is turned into a spectrum, and the result is an image: time along one axis, pitch up the other, brightness for energy. In that image a sung note is a stack of horizontal lines that wobble together, a drum is a vertical smear, a bass note is a thick band near the bottom. The network's whole job is to look at that image and decide, for every point in it, what fraction belongs to the voice. Then the picture is turned back into sound.
This is why separation quality depends so much on the song rather than on the tool. A voice alone over a piano is two things that barely overlap in the picture, and any decent model will pull them apart cleanly. A shouted chorus over distorted guitars is two things occupying the same pixels, and the model has to apportion energy that genuinely could have come from either. Where it guesses wrong you hear it as a faint watery quality, because the error is spread thinly across many frequencies rather than concentrated anywhere.
Why the old phase-cancellation trick is not this
There is a much older method that people still recommend: invert one channel, add it to the other, and anything sitting exactly in the middle of the stereo image cancels out. Vocals are usually centred, so vocals mostly vanish. It costs nothing and takes no time.
It also takes the bass, the kick drum and the snare with it, because those are centred too. What survives is the reverb and the guitars, in mono, sounding like a rehearsal heard through a wall. The trick is not separating anything; it is deleting the middle of the stereo field and hoping the vocal was the only thing there. A model that has learned what a voice looks like can remove a centred vocal while leaving the centred kick drum exactly where it was, and that difference is the entire point.
The honest limitations
Reverb and delay are the hardest problem. When a voice is recorded in a room, or has a reverb added, its energy is smeared across time and mixed into the whole track. The model can identify the direct voice cleanly and still leave its tail behind, which is why a heavily-produced pop vocal can disappear while a faint halo of itself remains.
Backing vocals are a genuine ambiguity rather than a failure. A model asked to remove the voice will usually remove the harmonies too, since they are voices. If you wanted the harmonies kept, there is no setting for that, because the question you are asking is one the model was never trained to distinguish.
Anything already heavily compressed will separate worse. A low-bitrate MP3 has had its quiet detail thrown away by the encoder, and some of that discarded detail is exactly what the model uses to tell a voice from an instrument. Start from the best copy you have; separating a 96 kbps file gives a noticeably rougher result than separating the same song at 320.