JamScore AI

Home / Guides / Remove Vocals

How to Remove Vocals from a Song

AI stem separation can produce a clean instrumental or a cappella from almost any track, straight in your browser, no plugins, no installs. Here is what the process actually involves, what to expect from the results, and how to get the most out of it.

How to remove vocals from a song using AI stem separation

Last updated August 2026 ยท 6 min read

How AI vocal removal works

Traditional phase-cancellation vocal removers worked by inverting one channel of a stereo file and adding it back, which only removed anything panned dead-centre. The results were famously unreliable: any instrument sharing that centre space (kick drum, bass, lead synth) would vanish too.

Modern AI separation is fundamentally different. Dedicated deep-learning models, like the ones JamScore AI uses, are trained on thousands of songs where the stems are known. Rather than exploiting stereo phase tricks, the model learns the spectral and temporal signatures of vocals, drums, bass, guitar, and piano. It then masks the spectrogram of your song accordingly, producing independent stem tracks. The quality ceiling is much higher, though no model is infallible.

Step-by-step: removing vocals with JamScore AI

  1. Upload your audio file

    Drag an MP3, WAV, M4A, or OGG file onto the upload area, or click to browse. Files up to 10 minutes long are supported. For the cleanest results, use the highest-quality version of the track you have, a 320 kbps MP3 will give the model more to work with than a heavily compressed 128 kbps stream rip.

  2. Let the AI separate the stems

    The AI separation model processes your audio with high fidelity. Separation typically takes between 30 seconds and two minutes depending on track length and your device. There is a progress indicator; you do not need to keep the tab in focus.

  3. Mute the vocal stem in the mixer

    Once separation finishes, the studio mixer appears with a channel for each stem. Press Mute on the Vocals channel. The remaining stems, drums, bass, guitar, piano, and other, play together as the backing track. If you are on the free plan, you will have a Vocals stem and an Instrumental stem; muting vocals gives you the instrumental directly.

  4. Adjust levels and EQ

    Even a good separation sometimes leaves one stem slightly too prominent. Use the per-stem volume faders and the three-band EQ (low, mid, high) to rebalance the mix. If the guitar is masking the bass, pull down its mid slightly. If the backing feels thin without the vocal, nudging the "other" channel up a touch can add warmth.

  5. Set an A/B loop if needed

    If you want to practise singing over the backing, set loop in and out points to drill a specific section, a chorus, a bridge, a tricky key change, without having to scrub back manually each time.

  6. Export as WAV

    Click the download icon on the channels you want. Each stem exports as a lossless WAV file. To get the full mixed instrumental (all stems except vocals), mute only the vocal channel and export each remaining stem, then combine them in a DAW, or use the Solo function on each non-vocal stem one at a time and mix down in your recording software.

Getting an a cappella instead

The same process works in reverse. Instead of muting vocals, Solo the Vocals channel in the mixer. Everything else drops out and you hear the isolated vocal track. This is useful for analysing a singer's phrasing and timing, for creating remix stems, or for pitch-checking your own vocal against a reference. Export the soloed stem as WAV in the same way.

If you want to go further, for example converting the vocal melody to MIDI or sheet music, see our guide on converting a song to MIDI and sheet music, which walks through the transcription step after separation.

Honest caveat on quality: Heavily brick-walled or loudness-maximised masters (many pop and rock records from the 2000s onwards) can produce a faint "watery" or phasey vocal ghost on the instrumental. This is not a bug, it reflects how little headroom exists between the vocal and the surrounding mix on those masters. Quieter, more dynamically open recordings (jazz, acoustic, lo-fi) tend to separate more cleanly. Older mono recordings fare the worst: without any stereo-field information, the model has less evidence for spatial placement and artefacts are more pronounced.

What the free plan covers

The free plan processes three songs per month and separates them into two stems: Vocals and Instrumental. That is enough to get a karaoke backing or an a cappella. For six-stem separation, individual drums, bass, guitar, piano, and other channels, you need a Pro subscription ($9/month) or Studio ($24/month). Both paid plans include unlimited processing.

The instrument-level control is worth having if you want to, say, remove just the guitar from a jazz trio while keeping the piano and bass, or practise a drum pattern by isolating the kit. See the guide on how to isolate instrument stems for that workflow.

Try it now, no account required

Upload any MP3, WAV, or M4A and get your stems in seconds. Free plan: 3 songs per month.

Open JamScore AI

Tips for better results

Frequently asked questions

Will the instrumental be completely clean?

Not always. AI separation has improved enormously, but heavily limited or loud masters can leave a faint, "watery" vocal residue on the instrumental. Quieter, well-produced tracks tend to separate more cleanly.

Does the free plan let me remove vocals?

Yes. The free plan separates vocals and instrumental (two stems) and allows three songs per month. You need a Pro or Studio subscription to access the full six-stem separation.

Can I export the instrumental as MP3?

Exports are WAV only. If you need MP3, convert the WAV file afterwards with any free tool such as Audacity or FFmpeg.

Why does an old or mono recording separate badly?

Older recordings were often mixed to mono, or the vocal sits on the same frequencies as the instruments with very little separation in the original mix. The AI has less information to work with, so leakage between stems is more noticeable.

Is my audio file uploaded to a server?

Processing happens in your browser. Your audio is not sent to a remote server, and nothing is stored after you close the tab.