MMMU-Benchmark/MMMU
This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!
$ git clone https://github.com/huggingface/evaluation-guidebook.gitThis repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"
Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.
๐ชข Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Machine Learning Engineering Open Book
Anomaly detection related books, papers, videos, and toolboxes. Last update late 2025 for LLM and VLM works!
๐ง Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 ๐
Data from GitHub ยท snapshot Sep 24, 2026