ictnlp/LLaMA-Omni
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
[ICML 2024] Mastering Text-to-Image Diffusion: Recaptioning, Planning, and Generating with Multimodal LLMs (RPG)
$ git clone https://github.com/YangLing0818/RPG-DiffusionMaster.gitLLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
Cambrian-1 is a family of multimodal LLMs with a vision-centric design.
NEO Series: Native Vision-Language Models from First Principles
β¨β¨[CVPR 2025] Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
[NeurIPS 2025 Spotlight] A Unified Tokenizer for Visual Generation and Understanding
β¨β¨Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
Data from GitHub Β· snapshot Sep 24, 2026