VoxCPM
View on GitHubVoxCPM2: Tokenizer-Free TTS for Multilingual Speech Generation, Creative Voice Design, and True-to-Life Cloning
VoxCPM2 is a tokenizer-free TTS system with a 2B diffusion-autoregressive model supporting 30 languages, voice design from text descriptions, controllable voice cloning, and 48kHz output. Ships Python API, CLI, Gradio demo, and vLLM-based serving.
Use Cases
Multilingual text-to-speech in 30 languagesZero-shot voice cloning from a short reference clipCreating new voices from natural-language descriptionsStyle-controlled cloning (emotion, pace, tone)Real-time streaming speech synthesisBatch audio generation from text filesAudiobook and narration productionDubbing and localization workflowsFine-tuning custom voices with SFT/LoRAProduction TTS serving via vLLM with OpenAI-compatible APIPost-generation word/character timestamp alignmentOn-device inference via llama.cpp-omni
Built With
- Language
- Python
- Frameworks
- PyTorch · Transformers · Gradio · vLLM · Nano-vLLM · ModelScope · Hugging Face Hub · Docker · FastAPI
Tags
tts · text-to-speech · voice-cloning · voice-design · multilingual · speech-synthesis · audio-generation · diffusion · streaming · tokenizer-free · lora · fine-tuning · 48khz · pytorch · minicpm