ihateaudio

Acapella Extractor

Just the voice, with the band taken away.

Drop an audio file here

or paste one from your clipboard

MP3 · WAV · M4A · AAC · OGG · FLAC and moreYour file stays on your device. Always.

How to use it

  1. Drop the song in. If you have already used the vocal remover on this device there is nothing to download — the two tools share one model, and it is already here.
  2. Press Extract. The network reads the whole track in chunks, working out which part of every moment is voice, and the counter tells you how far along it is.
  3. Play it back before downloading. Solo the vocal to hear it alone — an isolated vocal is the hardest thing to get right, so listen for cymbal bleed and reverb tails rather than assuming.
  4. Download as WAV if this is going into a DAW. It is the format that will not add a second layer of artefacts on top of the separation.

Why an isolated vocal is harder than an instrumental

Both halves come from the same decision, so it seems as though they ought to be equally good. They are not, and the asymmetry is worth understanding because it tells you what to expect.

An instrumental is mostly a subtraction problem. There is a great deal of music and comparatively little voice, so removing the voice leaves something whose errors are buried under everything else still playing. A missed fragment of vocal sits inside a full arrangement where you may not notice it.

An isolated vocal has nowhere to hide. Once the arrangement is gone, every error is exposed against near-silence: a fragment of snare that was left behind is obvious, a moment where the model removed too much is an audible hole. The same separation, judged on the quieter half, sounds worse simply because there is nothing left to mask it.

What to do with the result

For learning a part, this is close to ideal. Hearing exactly what a singer does with timing and phrasing, with nothing else in the way, is worth more than any transcription, and artefacts do not matter for that at all.

For a remix, treat it as a starting point that will need help. An isolated vocal usually benefits from a high-pass filter around 80 Hz to remove low-frequency mud the separation left behind, and from a new reverb, since the original room is smeared and partially missing. Adding a deliberate space to a dry-ish vocal sounds far better than trying to rescue a half-removed one.

For transcription, run it through the audio transcriber afterwards. Speech recognition is markedly more accurate on a vocal with the band removed, because the model doing the transcribing is no longer competing with instruments occupying the same frequencies as the voice.

Questions

Is this as clean as the instrumental?
Usually slightly less clean, and it is worth knowing why. The model predicts the instrumental directly, so the vocal comes out as what is left when the instrumental is subtracted from the original. That means every small error in the instrumental estimate lands in the vocal. In practice it is fine for a reference track, a remix, a transcription or a practice guide, and it is not what you would master a release from.
Why does it need no download when the vocal remover did?
They are the same 64 MB model, stored once. Separation produces both halves in a single pass, so whichever of the two tools you use first pays the download and the other is free from then on. It also means the model is stored under one name, so clearing site data clears it once rather than twice.
There is cymbal bleed all over the vocal.
Cymbals are the classic hard case. A crash is broadband noise spread across exactly the frequencies a sibilant "s" occupies, and in the spectrogram the two genuinely look alike. Models trained on real music learn to favour keeping consonants intelligible, which means erring toward letting some cymbal through rather than lisping the vocal. If the bleed matters more than the sibilance, a gentle high shelf cut around 8 kHz on the result costs you very little of the voice.
Can I get each backing vocal separately?
No, and no tool can. The model separates voice from not-voice; it has no concept of which voice. Two singers, or a lead and its own doubled harmony, are a single object as far as it is concerned. What comes out is every voice in the track mixed together.
The vocal sounds thin or watery.
That is separation artefact and it has a specific cause. The model decides how much of each frequency at each moment belongs to the voice, and where it is uncertain it splits the difference. Those partial decisions, scattered across the spectrum and changing every few milliseconds, are what the ear hears as a watery, phasey quality. Denser mixes produce more of it. Starting from a higher-quality source file reduces it, because the model has more detail to work from.
Will this work on speech rather than singing?
Less well than you would hope. The model was trained on music, so it has learned what a sung voice over instruments looks like. A podcast with background music does get separated, but a purpose-built denoiser is better at that job — try the noise remover, which is built for pulling a voice out of a noisy room rather than out of an arrangement.