๐Ÿ† #1,606 overall#188 of 301 in Speech Recognition

YuanGongND /ltu

Code, Dataset, and Pretrained Models for Audio and Speech Large Language Model "Listen, Think, and Understand".

$ git clone https://github.com/YuanGongND/ltu.git
GitHub social preview for YuanGongND/ltu
Stars
479
479
Forks
41
41
Language
Python
License
None
Created
May 19, 2023
3.4 years old
Last push
Apr 24, 2024

Categories

GitHub topics

More in Speech Recognition

Speech Recognition

speechbrain/speechbrain

A PyTorch-based Speech Toolkit

๐Ÿค– Language Models
PythonApache-2.0updated Aug 27, 2026
GitHub โ†—โ˜… 11.8Kโ‘‚ 1.7K
Speech Recognition

bytedance/SALMONN

SALMONN family: A suite of advanced multi-modal LLMs

๐Ÿค– Language Models
OtherApache-2.0updated Aug 24, 2026
GitHub โ†—โ˜… 1.5Kโ‘‚ 125
Speech Recognition

mravanelli/SincNet

SincNet is a neural architecture for efficiently processing raw audio samples.

PythonMITupdated Apr 28, 2021
GitHub โ†—โ˜… 1.2Kโ‘‚ 273
Speech Recognition

lhotse-speech/lhotse๐Ÿ”ฅ active

Tools for handling multimodal data in machine learning projects.

PythonApache-2.0updated Sep 22, 2026
GitHub โ†—โ˜… 1.2Kโ‘‚ 281
Speech Recognition

YuanGongND/whisper-at

Code and Pretrained Models for Interspeech 2023 Paper "Whisper-AT: Noise-Robust Automatic Speech Recognizers are Also Strong Audio Event Taggers"

PythonBSD-2-Clauseupdated Feb 21, 2024
GitHub โ†—โ˜… 426โ‘‚ 36

Data from GitHub ยท snapshot Sep 24, 2026