YuanGongND/ltu
Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
Code and Pretrained Models for Interspeech 2023 Paper "Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong Audio Event Taggers"
$ git clone https://github.com/YuanGongND/whisper-at.gitCode, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
Python API & command-line tool to easily transcribe speech-based video files into clean text
A PyTorch-based Speech Toolkit
Speech recognition module for Python, supporting several engines and APIs, online and offline.
A text-to-speech (TTS), speech-to-text (STT) and speech-to-speech (STS) library built on Apple's MLX framework, providing efficient speech analysis on Apple Silicon.
Distilled variant of Whisper for speech recognition. 6x faster, 50% smaller, within 1% word error rate.
Data from GitHub ยท snapshot Sep 24, 2026