๐Ÿ† #598 overall#178 of 601 in Language Models

huggingface /evaluation-guidebook

Sharing both practical insights and theoretical knowledge about LLM evaluation that we gathered while managing the Open LLM Leaderboard and designing lighteval!

$ git clone https://github.com/huggingface/evaluation-guidebook.git
GitHub social preview for huggingface/evaluation-guidebook
Stars
2.1K
2,147
Forks
125
125
Language
Jupyter Notebook
License
Other
Created
Oct 9, 2024
2.0 years old
Last push
Dec 3, 2025

Categories

GitHub topics

More in Language Models

Language Models

MMMU-Benchmark/MMMU

This repo contains evaluation code for the paper "MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI"

โ“ Question Answering๐Ÿง  Natural Language Processing
PythonApache-2.0updated Jul 28, 2026
GitHub โ†—โ˜… 598โ‘‚ 56
Language Models

mlabonne/llm-course

Course to get into Large Language Models (LLMs) with roadmaps and Colab notebooks.

OtherApache-2.0updated Feb 5, 2026
GitHub โ†—โ˜… 83.1Kโ‘‚ 9.7K
Language Models

langfuse/langfuse๐Ÿ”ฅ active

๐Ÿชข Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.

TypeScriptOtherupdated Sep 24, 2026
GitHub โ†—โ˜… 35Kโ‘‚ 3.8K
Language Models

stas00/ml-engineering๐Ÿ”ฅ active

Machine Learning Engineering Open Book

PythonCC-BY-SA-4.0updated Sep 23, 2026
GitHub โ†—โ˜… 19Kโ‘‚ 1.2K
Language Models

Helicone/helicone

๐ŸงŠ Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 ๐Ÿ“

TypeScriptApache-2.0updated Sep 16, 2026
GitHub โ†—โ˜… 6.2Kโ‘‚ 674

Data from GitHub ยท snapshot Sep 24, 2026