Retrieval-based-Voice-Conversion-WebUI
View on GitHubEasily train a good VC model with voice data <= 10 mins!
RVC WebUI is a PyTorch voice-conversion framework: train a custom timbre model from as little as 10 minutes of audio, then convert speech or singing vocals via CLI, Gradio WebUI, or a low-latency realtime voice changer. Includes RMVPE pitch extraction, HuBERT features, and pymss vocal separation.
Use Cases
Convert one voice into another from short audio samplesTrain a custom voice-conversion model with about 10 minutes of clean speechReal-time voice changer with ~90-170ms end-to-end latencyCreate AI singing voice covers from vocalsSeparate vocals from instrumentals before conversionExtract F0 pitch to avoid muted/unvoiced artifacts on hard audioMerge checkpoints to blend timbresRun inference/training through a local Gradio WebUIRun on AMD/Intel via CPU or Windows DirectMLCLI inference for batch audio conversion
Built With
- Language
- Python
- Frameworks
- PyTorch · Gradio · torchaudio · FFmpeg · ContentVec · VITS · HiFi-GAN · RMVPE · pymss/MSST · FAISS · ONNX/DirectML · soundfile
Tags
voice-conversion · voice-cloning · rvc · audio-generation · realtime-voice-changer · singing-voice · model-training · pitch-extraction · vocal-separation · gradio-webui · pytorch · hubert · rmvpe · vits · retrieval-index · checkpoint-merge