Skip to content

Vocal Remover guide

How to Set Up and Use Vocal Remover

Setting up a vocal remover is simpler than you might think — here's a step-by-step walkthrough from download to your first clean stem separation.

How to Set Up and Use Vocal Remover

Getting a vocal remover up and running for the first time is straightforward once you know where the decisions lie. This guide focuses on Ultimate Vocal Remover (UVR) — the free, open-source desktop application that is the most capable option available without spending a penny — but the general principles apply to most tools, including web-based services like LALAL.AI and Moises.

What You’ll Need Before You Start

For a web-based service, you need only a browser and an audio file. For UVR, make sure your machine meets these basics:

  • Windows 10/11 or macOS 11 (Big Sur) or later
  • At least 8 GB of RAM (16 GB is more comfortable for larger models)
  • A modern CPU; an NVIDIA GPU (for CUDA) or Apple Silicon chip speeds processing significantly
  • Your source audio file in FLAC, WAV, or high-bitrate MP3 (320kbps preferred)

Step-by-Step: Setting Up and Using UVR

  1. Download and install UVR. Visit the official Ultimate Vocal Remover project on GitHub (github.com/Anjok07/ultimatevocalremovergui) and download the installer for your operating system. Run the installer and follow the prompts. The application will install a Python environment and its dependencies automatically.
  2. Launch UVR and open the settings. On first launch, check the Settings panel. Under Processing, select your processing device: choose your NVIDIA GPU if you have one (shown as a CUDA device), or choose MPS if you’re on an Apple Silicon Mac. If neither is available, CPU is the fallback — slower, but it works.
  3. Download an AI model. Click the Download Centre button. For most vocal-removal tasks, start with a Demucs model — the htdemucs_ft (fine-tuned) variant is a reliable starting point. Click Download and wait for the model file to arrive; it is several hundred megabytes.
  4. Select your model and configure the output. In the main window, choose your downloaded model from the AI Model dropdown. Set the Output Format to WAV or FLAC for lossless quality. Choose a Save To directory where the separated stems will be stored.
  5. Load your audio file. Drag and drop your track into the input area, or click Select Input and browse to the file. UVR shows the file name and its duration. You can add multiple files for batch processing.
  6. Run the separation. Click Start Processing. A progress bar shows the status. Processing time depends on your hardware and model size — typically between fifteen seconds and a few minutes for a single track. When finished, UVR saves the stems to your chosen output folder.
  7. Check the output. Open your output folder and listen to both the instrumental and the vocal stems. Look for ghost vocals — faint traces of the voice remaining in the instrumental. If the bleed is noticeable, return to UVR and try a different model. MDX-Net models often handle certain tracks better than Demucs, and vice versa.
  8. Fine-tune if needed. For stubborn vocal bleed, you can run the instrumental stem through UVR again with a different model — this “stacking” technique can clean up residual artefacts, though it may also reduce some low-level detail in the instrumental.

Setting Up a Web Service (LALAL.AI as an Example)

If you prefer not to install software, web-based services are quicker to get started with:

  1. Go to the service’s website and create a free account if required.
  2. Upload your audio file using the drag-and-drop interface. Most services accept MP3, WAV, FLAC, and M4A.
  3. Choose your output type — typically “Vocal” and “Instrumental,” or a four-stem split if available.
  4. Start processing and wait. Web services queue files server-side; times vary.
  5. Use the preview player to listen to a short segment before downloading your full files.
  6. Download the output stems to your computer.

Tips for Consistently Better Results

  • Use lossless source files wherever you can. Compressed MP3s at 128kbps or below confuse the AI model and produce noticeably worse separations.
  • Normalise your input to a consistent level before processing, if your tool allows it. Very quiet or very loud files can occasionally produce less stable separations.
  • Keep the original file safe. Always work from a copy of your source file, never the original. You may want to re-run with a different model.
  • Try multiple models. No single model is best for every genre or production style. Dance and electronic music, acoustic folk, and heavily layered pop each tend to suit different model strengths.

Troubleshooting Common Issues

If UVR crashes or fails to process, the most common culprits are:

  • Insufficient RAM — try a lighter model (smaller ONNX models use less memory than large Demucs variants)
  • Corrupted source file — re-export or re-download the audio from a fresh source
  • Outdated UVR version — check for updates in the app; model support improves with each release

If you’re new to the concept of vocal removal and would like a broader understanding of the technology before diving into setup, the vocal remover explainer is a good place to start. And if you’re comparing which tools to use, take a look at the best vocal removers by budget before installing anything.

Frequently Asked Questions

How do I set up Ultimate Vocal Remover?

Download the installer from the UVR GitHub page, run it, then on first launch download an AI model from the built-in Download Centre. Select Demucs htdemucs_ft as a starting point, choose WAV output, load your file, and click Start Processing.

Do I need a GPU to use a vocal remover?

No — a CPU works fine, but processing will be significantly slower. An NVIDIA GPU (CUDA) or Apple Silicon chip accelerates processing considerably, often cutting a four-minute track to under thirty seconds.

What file formats can I use as input?

Most vocal removers accept WAV, FLAC, and MP3. High-quality source files (lossless or 320kbps MP3) produce noticeably better separations than compressed or low-bitrate audio.

Why is there still a ghost vocal in my instrumental?

Residual vocal bleed is a known limitation of current AI models. Try a different model within UVR — MDX-Net and Demucs variants handle different tracks differently. You can also 'stack' two models by running the output through UVR a second time.

How long does vocal removal take?

With GPU acceleration, a typical four-minute track takes fifteen to forty-five seconds. On CPU alone, the same track may take two to six minutes depending on the model and your hardware.

Can I process multiple files at once?

Yes — Ultimate Vocal Remover supports batch processing. Add multiple files to the input list before clicking Start Processing, and UVR will work through them in sequence.

About Lucy Edwards

Lucy grew up singing in Cardiff, where choirs and communal singing are woven into daily life, and trained her ear long before she trained her pen. She's endlessly interested in the voice as an instrument and how singers look after it. She writes about vocals, singing and vocal gear.

Comments

Leave a comment

Comments are read before they appear. Links cannot be published.

Your rating (optional)