OpenBMB/VisCPM
[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | εΊδΊCPMεΊη‘樑εηδΈθ±εθ―ε€ζ¨‘ζ倧樑εη³»ε
[ICLR & NeurIPS 2025] Repository for Show-o series, One Single Transformer to Unify Multimodal Understanding and Generation.
$ git clone https://github.com/showlab/Show-o.git[ICLR'24 spotlight] Chinese and English Multimodal Large Model Series (Chat and Paint) | εΊδΊCPMεΊη‘樑εηδΈθ±εθ―ε€ζ¨‘ζ倧樑εη³»ε
[NeurIPS 2024 Best Paper Award][GPT beats diffusionπ₯] [scaling laws in visual generationπ] Official impl. of "Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction". An *ultra-simple, user-friendly yet state-of-the-art* codebase for autoregressive image generation!
Align Anything: Training All-modality Model with Feedback
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
Code and models for ICML 2024 paper, NExT-GPT: Any-to-Any Multimodal Large Language Model
Get clean data from tricky documents, powered by vision-language models β‘
Data from GitHub Β· snapshot Sep 24, 2026