Benchmarks / CyberGym
CyberGym
Built by UC Berkeley · released 3 Jun 2025
1,507 real vulnerabilities from 188 open-source projects, drawn from Google's OSS-Fuzz. Given a text description of a bug and the unpatched codebase, an agent must write a proof-of-concept input that triggers the vulnerability. UC Berkeley maintains a public leaderboard with team submissions.
Frontier
98.5%
Success rate (Level 1)
Creation (天工), multi-model · agent: Creation
7 Sep 2026 · Source: UC Berkeley (third party)
Team submission listed first on the official leaderboard. Uses a runnable vulnerable Docker image ("dynamic" label). Sangfor AI (GLM-5.3) follows at 97.21% and Alipay AI4SDL at 96.75%.
The best result is at 90% or more of the ceiling.
What it measures
Whether an agent can reproduce a known vulnerability in a large real codebase. A task counts as solved when the generated proof of concept crashes the pre-patch build but not the post-patch build. The headline (Level 1) gives the agent the vulnerability description and the source. Other levels give less or more information. The leaderboard notes when a team used a runnable vulnerable image or test-time memory.
Success rate (Level 1): Share of the 1,507 instances where the agent produces a working proof of concept for the target vulnerability. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- repo, cli
- Grading
- automated-tests
- Tasks
- 1,507
- Contamination
- Vulnerabilities and their fixes are public in OSS-Fuzz and project histories, so training data may include them.
- Reuse
- Apache-2.0 (GitHub repository license). Dataset hosted on Hugging Face (about 240 GB). (open-apache)
Limits to keep in mind
- Leaderboard results are run and submitted by the teams themselves. Agent runs are stochastic, and the maintainers say small score differences may not reflect real capability gaps now that leading systems score high. Source
- Vulnerability descriptions can be ambiguous, which adds noise to grading. Source
- Top entries use multiple models, orchestration, and test-time memory across instances. They measure agent systems, not single models. Source
- The task is reproduction of known bugs, not discovery. The paper separately reports 34 zero-days found by agents in open-ended runs. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Success rate (Level 1) | Date | Source |
|---|---|---|---|
| Creation (天工), multi-model · frontier | 98.5% | 7 Sep 2026 | UC Berkeley · third-party |
| MDASH (multi-model) | 91% | 17 Jun 2026 | Microsoft Research · lab-reported |
| Claude Mythos Preview | 83.1% | 7 Apr 2026 | Anthropic · lab-reported |
| OpenHands (GPT-5) | 39.4% | 5 Dec 2025 | UC Berkeley · primary |
| OpenHands (Claude Sonnet 3.7) | 11.9% | 15 May 2025 | UC Berkeley · primary |
Timeline
Where it sits in the atlas
Go to the source
- Website www.cybergym.io
- Paper arxiv.org
- Full leaderboard www.cybergym.io
- Code github.com
- Dataset huggingface.co
Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: CyberGym by UC Berkeley (https://arxiv.org/abs/2506.02548).
Then, if you used this page:
Can Agents Work. "CyberGym: frontier results and sources." https://canagentswork.com/benchmarks/cybergym/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-cybergym,
title = {{CyberGym: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/cybergym/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}