Language Models
EvolvingLMMs-Lab/lmms-evalπ₯ active
One-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
PythonOtherupdated Sep 24, 2026
A Dynamic Environment to Evaluate Attacks and Defenses for LLM Agents.
$ git clone https://github.com/ethz-spylab/agentdojo.gitOne-for-All Multimodal Evaluation Toolkit Across Text, Image, Video, and Audio Tasks
A series of large language models developed by Baichuan Intelligent Technology
β‘FlashRAG: A Python Toolkit for Efficient RAG Research (WWW2025 Resource)
A 13B large language model developed by Baichuan Intelligent Technology
Rigourous evaluation of LLM-synthesized code - NeurIPS 2023 & COLM 2024
The official GitHub page for the survey paper "A Survey on Evaluation of Large Language Models".
Data from GitHub Β· snapshot Sep 24, 2026