Vibe Coding Discover

AI Tools

Multi-lingual large voice generation model, providing inference, training and deployment full-stack ability.

★ 24K2,701 forksPythonApache-2.0QwenAudio

CosyVoice is an LLM-based text-to-speech system (Fun-CosyVoice3 / CosyVoice2 / CosyVoice1) supporting multilingual zero-shot voice cloning, streaming output down to ~150ms latency, instruct control and voice conversion. Ships inference, fine-tuning and deployment paths (vLLM, TensorRT-LLM, gRPC/FastAPI, Gradio demo).

Use Cases

Zero-shot voice cloning from a short reference clipMultilingual and cross-lingual text-to-speechLow-latency streaming speech synthesis for real-time voice appsInstruction-controlled TTS (emotion, speed, volume, dialect)Fine-tuning custom voices on your own audio dataVoice conversionDeploying a TTS inference server via gRPC/FastAPI/DockerAccelerating TTS inference with vLLM or TensorRT-LLMBuilding voice chatbots and interactive assistantsAudiobook, dubbing and content narration pipelines

Built With

Language
Python
Frameworks
PyTorch · vLLM · TensorRT-LLM · Triton Inference Server · Gradio · FastAPI · gRPC · ONNX Runtime · PyTorch Lightning · DeepSpeed · Hydra · Transformers · Diffusers · ModelScope · WeNet

Tags

text-to-speech · tts · voice-cloning · zero-shot · multilingual · speech-synthesis · streaming · audio-generation · voice-conversion · llm · fine-tuning · cross-lingual · low-latency · gradio · instruct-control · deployment