pyannote-audio
View on GitHubNeural building blocks for speaker diarization: speech activity detection, speaker change detection, overlapped speech detection, speaker embedding
Python toolkit for training and running pretrained neural speech models and pipelines. It supports speaker diarization, speech activity and speaker-change detection, overlapped-speech detection, and speaker embeddings.
Use Cases
Identify speakers and their speaking turns in recordingsDetect speech and silenceDetect speaker changesDetect overlapping speechExtract speaker embeddingsFine-tune speech models on custom audio dataEvaluate speaker diarization pipelines
Built With
- Language
- Jupyter Notebook
- Frameworks
- PyTorch · PyTorch Lightning · Hugging Face Hub · pyannote.pipeline
Tags
speaker diarization · speech processing · speaker recognition · voice activity detection · speaker embeddings · pretrained models · audio ML · model fine-tuning