FireRedTeam/FireRedASR
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
$ git clone https://github.com/Soul-AILab/SoulX-Transcriber.gitOpen-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
WhisperX: Automatic Speech Recognition with Word-level Timestamps (& Diarization)
Easy-to-use Speech Toolkit including Self-Supervised Learning model, SOTA/Streaming ASR with punctuation, Streaming TTS with text frontend, Speaker Verification System, End-to-End Speech Translation and Keyword Spotting. Won NAACL2022 Best Demo Award.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Automatic Speech Recognition with Speaker Diarization based on OpenAI Whisper
Data from GitHub ยท snapshot Sep 24, 2026