Forum Discussion
Is AI tool a good choice to remove vocal from audio?
To be completely clear, the search results do not provide a concrete guide on how to remove vocal from audio with WeSep. If you are determined to explore how to remove vocal from audio using this framework.
WeSep is a modular framework that reformulates target speaker extraction as a "heterogeneous cue-conditioned learning problem". This means it isolates a target speaker based on auxiliary cues, such as:
- A short sample of the target speaker's voice (enrollment speech).
- Spatial information (like the direction the speaker is in).
- Visual signals (like lip movement from a video).
- Textual descriptions of the speaker.
Since WeSep's core design is for speaker isolation, using it to remove vocal from audio would require adapting its underlying technology. A promising approach could be to use a speaker embedding model (like Wespeaker, which is noted for integration with WeSep) to represent a generic "singer" voice. You would then potentially use that embedding as a cue to extract the vocal stem.
Given that you are seeking a straightforward method how to remove vocal from audio, WeSep is likely not the right tool. It is a complex research toolkit built for a different task. I would recommend exploring other specialized music source separation tools that are designed for vocal removal from the outset.