πŸ† #1,453 overall#396 of 601 in Language Models

Tencent-Hunyuan /CL-bench

CL-bench: A Benchmark for Context Learning

$ git clone https://github.com/Tencent-Hunyuan/CL-bench.git
GitHub social preview for Tencent-Hunyuan/CL-bench
Stars
583
583
Forks
27
27
Language
Python
License
Other
Created
Jan 23, 2026
0.7 years old
Last push
May 12, 2026

Categories

GitHub topics

More in Language Models

Language Models

SWE-bench/SWE-benchπŸ”₯ active

SWE-bench: Can Language Models Resolve Real-world Github Issues?

PythonMITupdated Sep 18, 2026
GitHub β†—β˜… 5.9Kβ‘‚ 979
Language Models

CLUEbenchmark/CLUE

δΈ­ζ–‡θ―­θ¨€η†θ§£ζ΅‹θ―„εŸΊε‡† Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard

Pythonno licenseupdated Feb 6, 2026
GitHub β†—β˜… 4.3Kβ‘‚ 543
Language Models

xlang-ai/OSWorld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

🧠 Natural Language Processing
PythonApache-2.0updated Sep 14, 2026
GitHub β†—β˜… 3.2Kβ‘‚ 534
Language Models

zhenbench/z-bench

Z-Bench 1.0 by ηœŸζ ΌεŸΊι‡‘οΌšδΈ€δΈͺιΊ»η“œηš„ε€§θ―­θ¨€ζ¨‘εž‹δΈ­ζ–‡ζ΅‹θ―•ι›†γ€‚Z-Bench is a LLM prompt dataset for non-technical users, developed by an enthusiastic AI-focused team in Zhenfund.

OtherCC-BY-4.0updated Jun 28, 2023
GitHub β†—β˜… 504β‘‚ 41
Language Models

xlang-ai/OSWorld-V2

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

🧠 Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub β†—β˜… 330β‘‚ 50
Language Models

uiuc-kang-lab/cve-bench

CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities

PythonApache-2.0updated Sep 15, 2026
GitHub β†—β˜… 291β‘‚ 56

Data from GitHub Β· snapshot Sep 24, 2026