Benchmarks / Cybench
Cybench
Built by Stanford University · released 15 Aug 2024
40 professional-level Capture the Flag tasks from four recent CTF competitions. An agent works in a Kali Linux container, runs commands against local files and task servers, and submits a flag. Each task also has subtasks for partial credit. Used by the US and UK AI Safety Institutes and in many model system cards.
Frontier
100%
Unguided % solved
Claude Mythos Preview
7 Apr 2026 · Source: Anthropic (reported by the system's developer)
Pass@1 of 100% on a 35-task subset, 10 trials per task, no extended thinking. The Cybench site lists this entry on its leaderboard.
The best result is at 90% or more of the ceiling.
What it measures
Whether an agent can solve a CTF task on its own (unguided) or with subtask hints (guided), across cryptography, web, reverse engineering, forensics, binary exploitation, and misc categories. The headline is the unguided share of tasks solved. The site also reports the hardest task solved by first-solve time of human teams (up to 24 hours 54 minutes).
Unguided % solved: Share of tasks solved without subtask guidance. Lab-reported entries use pass@1 on subsets of 35 to 39 tasks. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- automated-tests
- Tasks
- 40
- Human reference
- First solve time by competition teams, from 2 minutes to 24 hours 54 minutes per task.
- Contamination
- Tasks come from public CTF competitions with published writeups, so training data may include solutions. One HAL run used a framework fork that leaked an answer; the site adjusted those scores down.
- Reuse
- Tasks are public in the GitHub repository (Apache-2.0); the site asks users to cite the ICLR 2025 paper. (open-apache)
Limits to keep in mind
- Recent leaderboard entries are copied from lab system cards on 35 to 39 task subsets with different trial counts, not from a common harness. Source
- A fork of the Inspect framework leaked an answer to one task. Scores for o3-mini and o1-mini were adjusted down by 2.5 points. Source
- Anthropic found grading errors in its earlier Cybench runs, so its newer numbers may not match previously reported ones. Source
- CTF tasks are gamified. Anthropic says real-world vulnerability work (for example CyberGym) now reflects capability better. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Unguided % solved | Date | Source |
|---|---|---|---|
| Claude Opus 4.7 | 96% | 16 Apr 2026 | Anthropic · lab-reported |
| Claude Mythos Preview · frontier | 100% | 7 Apr 2026 | Anthropic · lab-reported |
| Claude Opus 4.5 | 82% | Nov 2025 | Anthropic · lab-reported |
| Claude 3.5 Sonnet harness: Cybench agent (structured bash) | 17.5% | 15 Aug 2024 | Stanford University · primary |
Timeline
Where it sits in the atlas
Go to the source
- Website cybench.github.io
- Paper arxiv.org
- Full leaderboard cybench.github.io
- Code github.com
Last checked 23 Sep 2026 against 6 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: Cybench by Stanford University (https://arxiv.org/abs/2408.08926).
Then, if you used this page:
Can Agents Work. "Cybench: frontier results and sources." https://canagentswork.com/benchmarks/cybench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-cybench,
title = {{Cybench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/cybench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}