modelscope/FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Set of ๐ with ๐ to help those building Voice AI agents ๐๏ธ๐ค
$ git clone https://github.com/mahimairaja/voiceai.gitFunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Open-source industrial-grade ASR models supporting Mandarin, Chinese dialects and English, achieving a new SOTA on public Mandarin ASR benchmarks, while also offering outstanding singing lyrics recognition capability.
A SOTA Industrial-Grade All-in-One ASR system with ASR, VAD, LID, and Punc modules. FireRedASR2 supports Chinese (Mandarin, 20+ dialects/accents), English, code-switching, and both speech and singing ASR. FireRedVAD supports speech/singing/music in 100+ langs. FireRedLID supports 100+ langs and 20+ zh dialects. FireRedPunc supports zh and en.
Stream-Omni is a GPT-4o-like language-vision-speech chatbot that simultaneously supports interaction across various modality combinations.
An end-to-end framework for multi-speaker transcription that jointly models who spoke, when, and what.
Data from GitHub ยท snapshot Sep 24, 2026