Benchmarks / Cybench

Cybench

Built by Stanford University · released 15 Aug 2024

40 professional-level Capture the Flag tasks from four recent CTF competitions. An agent works in a Kali Linux container, runs commands against local files and task servers, and submits a flag. Each task also has subtasks for partial credit. Used by the US and UK AI Safety Institutes and in many model system cards.

Frontier

100%

Unguided % solved

Claude Mythos Preview

7 Apr 2026 · Source: Anthropic (reported by the system's developer)

Pass@1 of 100% on a 35-task subset, 10 trials per task, no extended thinking. The Cybench site lists this entry on its leaderboard.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at Stanford University

Cybench: Unguided % solved over time, 4 recorded results. 0%20%40%60%80%100%Oct 2024Jan 2025Apr 2025Jul 2025Oct 2025Jan 2026Apr 2026 Claude 3.5 Sonnet: 17.5% (15 Aug 2024) Claude Opus 4.5: 82% (Nov 2025) Claude Mythos Preview: 100% (7 Apr 2026) Claude Opus 4.7: 96% (16 Apr 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/cybench/"><img src="https://canagentswork.com/og/benchmarks-cybench.png" width="600" height="315" alt="Cybench: the best result is 100% (Claude Mythos Preview, 7 Apr 2026)." loading="lazy"></a>

Markdown:

[![Cybench: the best result is 100% (Claude Mythos Preview, 7 Apr 2026).](https://canagentswork.com/og/benchmarks-cybench.png)](https://canagentswork.com/benchmarks/cybench/)

What it measures

Whether an agent can solve a CTF task on its own (unguided) or with subtask hints (guided), across cryptography, web, reverse engineering, forensics, binary exploitation, and misc categories. The headline is the unguided share of tasks solved. The site also reports the hardest task solved by first-solve time of human teams (up to 24 hours 54 minutes).

Unguided % solved: Share of tasks solved without subtask guidance. Lab-reported entries use pass@1 on subsets of 35 to 39 tasks. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
automated-tests
Tasks
40
Human reference
First solve time by competition teams, from 2 minutes to 24 hours 54 minutes per task.
Contamination
Tasks come from public CTF competitions with published writeups, so training data may include solutions. One HAL run used a framework fork that leaked an answer; the site adjusted those scores down.
Reuse
Tasks are public in the GitHub repository (Apache-2.0); the site asks users to cite the ICLR 2025 paper. (open-apache)

Limits to keep in mind

  • Recent leaderboard entries are copied from lab system cards on 35 to 39 task subsets with different trial counts, not from a common harness. Source
  • A fork of the Inspect framework leaked an answer to one task. Scores for o3-mini and o1-mini were adjusted down by 2.5 points. Source
  • Anthropic found grading errors in its earlier Cybench runs, so its newer numbers may not match previously reported ones. Source
  • CTF tasks are gamified. Anthropic says real-world vulnerability work (for example CyberGym) now reflects capability better. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemUnguided % solvedDateSource
Claude Opus 4.796%16 Apr 2026Anthropic · lab-reported
Claude Mythos Preview · frontier100%7 Apr 2026Anthropic · lab-reported
Claude Opus 4.582%Nov 2025Anthropic · lab-reported
Claude 3.5 Sonnet
harness: Cybench agent (structured bash)
17.5%15 Aug 2024Stanford University · primary

Timeline

  • 7 Apr 2026 — Claude Mythos Preview reaches 100% on Cybench subset. Source
  • 15 Aug 2024 — Stanford releases Cybench with 40 professional CTF tasks. Source

Where it sits in the atlas

Security

Vulnerability research

Last checked 23 Sep 2026 against 6 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: Cybench by Stanford University (https://arxiv.org/abs/2408.08926).

Then, if you used this page:

Can Agents Work. "Cybench: frontier results and sources." https://canagentswork.com/benchmarks/cybench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-cybench,
  title        = {{Cybench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/cybench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}