πŸ† #2,058 overall#492 of 601 in Language Models

uiuc-kang-lab /cve-bench

CVE-Bench: A Benchmark for AI Agents’ Ability to Exploit Real-World Web Application Vulnerabilities

$ git clone https://github.com/uiuc-kang-lab/cve-bench.git
GitHub social preview for uiuc-kang-lab/cve-bench
Stars
291
291
Forks
56
56
Language
Python
License
Apache-2.0
Created
Feb 25, 2025
1.6 years old
Last push
Sep 15, 2026

Categories

GitHub topics

More in Language Models

Language Models

SWE-bench/SWE-benchπŸ”₯ active

SWE-bench: Can Language Models Resolve Real-world Github Issues?

PythonMITupdated Sep 18, 2026
GitHub β†—β˜… 5.9Kβ‘‚ 979
Language Models

CLUEbenchmark/CLUE

δΈ­ζ–‡θ―­θ¨€η†θ§£ζ΅‹θ―„εŸΊε‡† Chinese Language Understanding Evaluation Benchmark: datasets, baselines, pre-trained models, corpus and leaderboard

Pythonno licenseupdated Feb 6, 2026
GitHub β†—β˜… 4.3Kβ‘‚ 543
Language Models

xlang-ai/OSWorld

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

🧠 Natural Language Processing
PythonApache-2.0updated Sep 14, 2026
GitHub β†—β˜… 3.2Kβ‘‚ 534
Language Models

zhenbench/z-bench

Z-Bench 1.0 by ηœŸζ ΌεŸΊι‡‘οΌšδΈ€δΈͺιΊ»η“œηš„ε€§θ―­θ¨€ζ¨‘εž‹δΈ­ζ–‡ζ΅‹θ―•ι›†γ€‚Z-Bench is a LLM prompt dataset for non-technical users, developed by an enthusiastic AI-focused team in Zhenfund.

OtherCC-BY-4.0updated Jun 28, 2023
GitHub β†—β˜… 504β‘‚ 41
Language Models

xlang-ai/OSWorld-V2

OSWorld 2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks

🧠 Natural Language Processing
PythonApache-2.0updated Sep 16, 2026
GitHub β†—β˜… 330β‘‚ 50

Data from GitHub Β· snapshot Sep 24, 2026