Language Models
QwenLM/Qwen-VL
The official repo of Qwen-VL (้ไนๅ้ฎ-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
PythonOtherupdated Aug 7, 2024
Official repo for "Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models"
$ git clone https://github.com/JIA-Lab-research/MGM.gitThe official repo of Qwen-VL (้ไนๅ้ฎ-VL) chat & pretrained large vision language model proposed by Alibaba Cloud.
Align Anything: Training All-modality Model with Feedback
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
The official repo of MiniMax-Text-01 and MiniMax-VL-01, large-language-model & vision-language-model based on Linear Attention
[CVPR 2024 Highlight๐ฅ] Chat-UniVi: Unified Visual Representation Empowers Large Language Models with Image and Video Understanding
VoxPoser: Composable 3D Value Maps for Robotic Manipulation with Language Models
Data from GitHub ยท snapshot Sep 24, 2026