Speech Recognition
lhotse-speech/lhotse๐ฅ active
Tools for handling multimodal data in machine learning projects.
PythonApache-2.0updated Sep 22, 2026
Voice activity detection (VAD) toolkit including DNN, bDNN, LSTM and ACAM based VAD. We also provide our directly recorded dataset.
$ git clone https://github.com/jtkim-kaist/VAD.gitTools for handling multimodal data in machine learning projects.
A list of publically available audio data that anyone can download for ASR or other speech activities
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
kaldi-asr/kaldi is the official location of the Kaldi project.
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Data from GitHub ยท snapshot Sep 24, 2026