whisperX
View on GitHubWhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Python ASR tool that transcribes audio with batched Whisper inference, aligns words to audio for precise timestamps, and optionally labels segments by speaker using diarization.
Use Cases
Transcribe audio to textGenerate word-level timestampsLabel transcript segments by speakerCreate synchronized subtitlesProcess multilingual audioBatch transcribe recordings
Built With
- Language
- Python
- Frameworks
- faster-whisper · PyTorch · torchaudio · pyannote-audio · Hugging Face Transformers
Tags
speech recognition · transcription · word timestamps · speaker diarization · forced alignment · voice activity detection · batched inference · audio