Song Splitter
Take a song apart and hear the pieces on their own.
or paste one from your clipboard
This tool runs a neural network on your machine
Each stem is its own 28 MB network. Only the ones you tick are fetched, and they are kept afterwards. Expect roughly the length of the track per stem.
One time only. Nothing is uploaded — this is the model coming down to you.
Model ready on this device. Works offline, and nothing you load here leaves the browser.
Solo a track to hear it alone. What you can hear is what gets mixed down.
How to use it
- Drop the song in, then tick the stems you actually want. Each one is a separate 28 MB network and a separate pass over the audio, so asking for two instead of four halves both the download and the wait.
- Trim first if you only need part of the track. The tool works on whatever is loaded, and separating one chorus takes a fraction of the time of a whole song.
- Press Split. The stems are produced one after another and the progress counter tells you which one is running.
- When it finishes you get a mixer: play everything together, mute a part, or solo one to hear it alone. Judge the split by ear before you download anything, then take individual stems, or a mix of whatever balance you set.
What a stem is, and what these are not
In a studio, stems are the real thing: the actual separate recordings, bounced out of the session before anyone mixed them together. They are perfect because they were never combined in the first place.
What this produces is an estimate of those recordings, reconstructed from the finished mix by a network that has learned what drums and bass and voices tend to look like. It is genuinely useful and it is not the same object. If you are remixing something officially, ask for the real stems. If you are learning a bassline, building a mashup, practising along, or working out what a producer did, an estimate is entirely sufficient.
Why the frequency limits matter more here than elsewhere
Each of these networks reads a fixed number of frequency bands, and because the transform size differs per stem, so does the ceiling. Bass is analysed with a very long window, which buys excellent frequency resolution down low — exactly what you need to tell a bass note from a kick drum — at the cost of only covering the bottom quarter of the spectrum. Nothing above roughly 5.5 kHz appears in the bass stem, which is fine, because nothing you want from a bass lives up there.
It becomes visible when you add the stems back together and find the top end missing. That is not a bug to work around; it is what these models are. If you need a complete complement to one stem, subtract that stem from the original yourself rather than summing the other three.
Getting a better result
Start from the best source you have. A 320 kbps MP3 or a lossless file gives the models detail that a 128 kbps file has already thrown away, and some of that detail is what distinguishes a hi-hat from a sibilant.
Separate a section rather than a whole track when you can. It is not only faster — the result is identical, because the models work in short chunks and have no memory of the wider song, so there is nothing to lose by cutting first.
Expect the drums to be the most convincing thing you get, and judge the tool on the stem you actually need rather than on the weakest one. "Other" being mediocre says very little about whether the drums are good.