Speech Recognition
YuanGongND/ltu
Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
๐ค Language Models
Pythonno licenseupdated Apr 24, 2024
SALMONN family: A suite of advanced multi-modal LLMs
$ git clone https://github.com/bytedance/SALMONN.gitCode, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".
A PyTorch-based Speech Toolkit
SincNet is a neural architecture for efficiently processing raw audio samples.
Code and Pretrained Models for Interspeech 2023 Paper "Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong Audio Event Taggers"
Python API & command-line tool to easily transcribe speech-based video files into clean text
State-of-the-Art Deep Learning scripts organized by models - easy to train and deploy with reproducible accuracy and performance on enterprise-grade infrastructure.
Data from GitHub ยท snapshot Sep 24, 2026