speechbrain/speechbrain
A PyTorch-based Speech Toolkit
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
$ git clone https://github.com/sgl-project/sglang-omni.gitA PyTorch-based Speech Toolkit
Production First and Production Ready End-to-End Speech Recognition Toolkit
Fine-tune the Whisper speech recognition model to support training without timestamp data, training with timestamp data, and training without speech data. Accelerate inference and support Web deployment, Windows desktop deployment, and Android deployment
[Unofficial] PyTorch implementation of "Conformer: Convolution-augmented Transformer for Speech Recognition" (INTERSPEECH 2020)
Speech-to-text server framework with next-gen Kaldi
The hub for audio AI research: papers, open models, benchmarks & datasets across audio LLMs, speech recognition, TTS, music & audio generation.
Data from GitHub ยท snapshot Sep 24, 2026