SWE-bench/SWE-benchπ₯ active
SWE-bench: Can Language Models Resolve Real-world Github Issues?
CL-bench: A Benchmark for Context Learning
$ git clone https://github.com/Tencent-Hunyuan/CL-bench.gitSWE-bench: Can Language Models Resolve Real-world Github Issues?
δΈζθ―θ¨ηθ§£ζ΅θ―εΊε Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Z-Bench 1.0 by ηζ ΌεΊιοΌδΈδΈͺιΊ»ηηε€§θ―θ¨ζ¨‘εδΈζζ΅θ―ιγZ-Bench is a LLM prompt dataset for non-technical users, developed by an enthusiastic AI-focused team in Zhenfund.
OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks
CVE-Bench: A Benchmark for AI Agentsβ Ability to Exploit Real-World Web Application Vulnerabilities
Data from GitHub Β· snapshot Sep 24, 2026