QwenAudio/SenseVoice๐ฅ active
Open-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Fun-ASR speech recognition models, with native Hugging Face Transformers support for Fun-ASR-Nano and separate FunASR, vLLM and llama.cpp deployment paths.
$ git clone https://github.com/QwenAudio/Fun-ASR.gitOpen-source SenseVoiceSmall model for Mandarin, Cantonese, English, Japanese, and Korean ASR, language ID, emotion recognition, and audio event detection.
Arabic-first generative speech recognition โ Audar-ASR-V1 (Flash + Turbo). #1 on the Open Universal Arabic ASR Leaderboard. Model cards, benchmarks & inference.
Open-source speech recognition toolkit for training, inference, streaming ASR, VAD, punctuation, speaker diarization pipelines, and OpenAI-compatible/MCP serving.
FunASR-powered video transcription, subtitle generation, and LLM-assisted clipping tool with a local Gradio UI.
Real-time audio translation, captures system audio + mic, runs ASR (Whisper/SenseVoice), translates via LLM API with streaming display. Perfect for VTubers, livestreamers, and watching foreign content. Windows ๅฎๆถ้ณ้ข็ฟป่ฏ๏ผASR ่ฏญ้ณ่ฏๅซๅ LLM ๆตๅผ็ฟป่ฏๆพ็คบ๏ผ้ๅ VTuberใไธปๆญๅๅค่ฏญ่ง้ข่ง็ใ
Open-source voice typing for Windows โ a Wispr Flow / Superwhisper alternative. Press a shortcut, speak, and AI-polished text lands at your cursor. Local models, your own API keys, or a self-hosted backend.
Data from GitHub ยท snapshot Sep 24, 2026