Every audio person has a folder of recordings they wish they could fix: the podcast interview taped in a cafe with an espresso machine screaming in the background, the archival tape with a faint 60Hz hum baked into every second, the video where a music bed is fighting with the narrator for attention. You can't re-record the past. What you can do is separate the voice from everything around it, and clean up the result. That's exactly what PlayHT's Voice Isolator is built for.
This tutorial walks through when to reach for it, how the workflow actually works, what quality to expect, and how to plug the output into other tools like transcription and voice cloning.
When a Voice Isolator Earns Its Keep
Not every noisy file needs AI intervention, but a surprising number of everyday situations do. A few common ones:
- The noisy-cafe interview. You recorded a great conversation on a phone or lavalier mic, and it's underscored by espresso machines, clinking cups, and other tables' chatter. The dialogue is there, but it's buried.
- The old recording with hum or hiss. Archival interviews, old voicemails, or recordings made on aging equipment often carry a steady electrical hum, tape hiss, or room rumble sitting underneath the speech the whole time.
- The video with a competing music bed. A promo video, a course recording, or a livestream clip has background music mixed in at a level that muddies the dialogue, and you need clean narration, the music alone, or both, separated into their own tracks.
In all three cases, the underlying task is the same: take one mixed file and split it into a voice layer and a non-voice layer.
How the Workflow Actually Works
PlayHT's Voice Isolator is deliberately simple to use. There's no manual EQ, no notch filters to tune by ear, no waveform surgery. The AI listens to the whole file and learns to tell foreground speech apart from everything else in it.
- Upload your noisy audio file. This can be a podcast recording, a phone-recorded interview, an old tape transfer, or the audio track pulled from a video. Most common audio formats work fine.
- The AI separates vocal and non-vocal elements. Under the hood, the model has learned what human speech tends to look like in a spectrogram versus what ambient noise, hum, hiss, and music tend to look like, and it uses that to pull the two apart rather than simply filtering by frequency.
- Download the isolated voice track. This is your clean dialogue, with the background noise or music suppressed.
- Optionally, download the separated background track too. If you need the noise bed or music bed on its own, for a remix, for evidence, or just out of curiosity, it's available separately rather than thrown away.
That's the entire loop. No audio engineering background required, though understanding the concepts underneath helps you get better results and set the right expectations, which brings us to the next part.
What to Expect: The Honest Version
AI separation has gotten remarkably good, but it isn't magic, and being upfront about its limits will save you time. Here's the honest breakdown.
Where it shines: a single clearly dominant foreground voice against a distinct background. Think one host talking over cafe ambience, one narrator over a music bed, one interview subject with hum underneath. The bigger the gap between the voice and everything else in the mix, the cleaner the separation.
Where it struggles:
- Heavily overlapping simultaneous speakers. If two people are talking over each other at similar volume, the model has to guess which voice is the foreground one, and it can bleed between the two rather than cleanly isolate either.
- Extremely loud or dominant noise sources. If the noise is actually louder than the speech, or shares a lot of acoustic characteristics with a voice (like a droning crowd or a vocal-heavy music track), separation quality drops. The model can only recover what's still detectably there in the mix.
- Very low-quality source audio. If the original recording is heavily clipped, extremely quiet, or already through several rounds of lossy compression, there's less signal for the model to work with.
Practical tip: before you commit to a big batch job, run one representative clip through first. It takes a minute and tells you immediately whether your specific noise situation is a good candidate or a hard case.
Two Places Isolated Audio Pays Off Downstream
Cleaning up a file isn't usually the end goal, it's a step toward something else. Two workflows benefit especially well from starting with isolated audio.
Better Transcripts via Speech to Text
Background noise is one of the biggest silent killers of transcription accuracy. Running a noisy file straight through Speech to Text often means more misheard words, dropped short phrases, and garbled names, especially in noisy environments like cafes or events. Isolate the voice first, then transcribe the clean track, and you'll typically see meaningfully fewer errors, particularly in the noisiest stretches of a recording.
Cleaner Source Material for Voice Cloning
Voice Cloning quality depends heavily on the quality of the sample you feed it. A clone trained on audio with hum, hiss, or a music bed running underneath tends to inherit some of that noise character, or at minimum gives the model a harder job telling the voice apart from everything else during training. Running your source clip through the Voice Isolator first, and feeding the model a clean, isolated voice track, is one of the simplest ways to improve a clone's starting point.
Getting Started
The Voice Isolator is available directly from your PlayHT dashboard alongside PlayHT's other audio tools, so there's no separate tool to learn or account to manage. Upload a file, download the isolated voice, and see for yourself. If you're evaluating it as part of a broader workflow, check the pricing page for how it fits into your plan.
Noisy audio doesn't have to be a dead end. With a realistic sense of what the tool does well, and a quick test on your specific source material, you can turn a lot of unusable recordings into something you can actually publish, transcribe, or build on.