modelscope/FunClip
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
$ git clone https://github.com/modelscope/FunASR.gitFunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
A PyTorch-based Speech Toolkit
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Fun-ASR speech recognition models, with native Hugging Face Transformers support for Fun-ASR-Nano and separate FunASR, vLLM and llama.cpp deployment paths.
Controllable Transcription. Verbatim ( every, filler, pause, stutter, vocal sound) , or intended ( what the speaker meant to say, optimized for readability) with word-level timestamps.
SincNet is a neural architecture for efficiently processing raw audio samples.
Data from GitHub ยท snapshot Sep 24, 2026