QuentinFuxa/WhisperLiveKitπ₯ active
Real-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.
Wav2Vec for speech recognition, classification, and audio classification
$ git clone https://github.com/m3hrdadfi/soxan.gitReal-time, local speech-to-text with streaming ASR, speaker diarization, translation, and OpenAI/Deepgram-compatible APIs.
Production First and Production Ready End-to-End Speech Recognition Toolkit
πΈSTT - The deep learning toolkit for Speech-to-Text. Training and deploying STT models has never been so easy.
A 0.9B model for long-form transcription in 50+ languages with speaker diarization, timestamps, and acoustic event awareness
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
Data from GitHub Β· snapshot Sep 24, 2026