Vibe Coding Discover

AI Tools

WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)

★ 24K2,460 forksPythonBSD-2-Clausem-bain

Python ASR tool that transcribes audio with batched Whisper inference, aligns words to audio for precise timestamps, and optionally labels segments by speaker using diarization.

Use Cases

Transcribe audio to textGenerate word-level timestampsLabel transcript segments by speakerCreate synchronized subtitlesProcess multilingual audioBatch transcribe recordings

Built With

Language
Python
Frameworks
faster-whisper · PyTorch · torchaudio · pyannote-audio · Hugging Face Transformers

Tags

speech recognition · transcription · word timestamps · speaker diarization · forced alignment · voice activity detection · batched inference · audio