Speech Recognition
lhotse-speech/lhotse๐ฅ active
Tools for handling multimodal data in machine learning projects.
PythonApache-2.0updated Sep 22, 2026
Python API & command-line tool to easily transcribe speech-based video files into clean text
$ git clone https://github.com/pszemraj/vid2cleantxt.gitTools for handling multimodal data in machine learning projects.
Speech recognition module for Python, supporting several engines and APIs, online and offline.
Open-Source Large Vocabulary Continuous Speech Recognition Engine
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
A Implementation of SpecAugment with Tensorflow & Pytorch, introduced by Google Brain
Data from GitHub ยท snapshot Sep 24, 2026