Acapella Extractor
Just the voice, with the band taken away.
or paste one from your clipboard
This tool needs a 64 MB download, once
It is the separation network, shared with the vocal remover and served from this domain rather than a third party. It is kept on your device afterwards, so this happens once and then the tool works offline forever. Extraction then runs at roughly real time, so a four-minute song takes a few minutes.
One time only. Nothing is uploaded — this is the model coming down to you.
Model ready on this device. Works offline, and nothing you load here leaves the browser.
How to use it
- Drop the song in. If you have already used the vocal remover on this device there is nothing to download — the two tools share one model, and it is already here.
- Press Extract. The network reads the whole track in chunks, working out which part of every moment is voice, and the counter tells you how far along it is.
- Play it back before downloading. Solo the vocal to hear it alone — an isolated vocal is the hardest thing to get right, so listen for cymbal bleed and reverb tails rather than assuming.
- Download as WAV if this is going into a DAW. It is the format that will not add a second layer of artefacts on top of the separation.
Why an isolated vocal is harder than an instrumental
Both halves come from the same decision, so it seems as though they ought to be equally good. They are not, and the asymmetry is worth understanding because it tells you what to expect.
An instrumental is mostly a subtraction problem. There is a great deal of music and comparatively little voice, so removing the voice leaves something whose errors are buried under everything else still playing. A missed fragment of vocal sits inside a full arrangement where you may not notice it.
An isolated vocal has nowhere to hide. Once the arrangement is gone, every error is exposed against near-silence: a fragment of snare that was left behind is obvious, a moment where the model removed too much is an audible hole. The same separation, judged on the quieter half, sounds worse simply because there is nothing left to mask it.
What to do with the result
For learning a part, this is close to ideal. Hearing exactly what a singer does with timing and phrasing, with nothing else in the way, is worth more than any transcription, and artefacts do not matter for that at all.
For a remix, treat it as a starting point that will need help. An isolated vocal usually benefits from a high-pass filter around 80 Hz to remove low-frequency mud the separation left behind, and from a new reverb, since the original room is smeared and partially missing. Adding a deliberate space to a dry-ish vocal sounds far better than trying to rescue a half-removed one.
For transcription, run it through the audio transcriber afterwards. Speech recognition is markedly more accurate on a vocal with the band removed, because the model doing the transcribing is no longer competing with instruments occupying the same frequencies as the voice.