Amphion
View on GitHubAmphion (/æmˈfaɪən/) is a toolkit for Audio, Music, and Speech Generation. Its purpose is to support reproducible research and help junior researchers and engineers get started in the field of audio, music, and speech generation research and development.
OpenMMLab toolkit for audio, music and speech generation: TTS, singing voice synthesis/conversion, voice conversion and text-to-audio, with neural codecs, vocoders and evaluation metrics. Ships reproducible PyTorch recipes, pretrained checkpoints and the Emilia dataset.
Use Cases
Built With
- Language
- Python
- Frameworks
- PyTorch · Hugging Face Transformers · Hugging Face Hub · Hugging Face Datasets · PyTorch Lightning · ModelScope · ONNX
Tags
text-to-speech · speech-synthesis · voice-conversion · singing-voice-conversion · singing-voice-synthesis · text-to-audio · music-generation · audio-generation · neural-audio-codec · vocoder · zero-shot-tts · multilingual-speech · reproducible-research · speech-dataset · evaluation-metrics · pytorch