Benchmarks / CyberBench v1.1
CyberBench v1.1 (CyberBench)
Built by Vals AI · released 30 Jun 2026
Vals AI's private test of whether an agent can reproduce and fix real crash bugs from Google's OSS-Fuzz. It follows CyberGym's method but uses vulnerabilities found after CyberGym's task window, so it is a fresh successor to that benchmark. The PoC track (60 tasks) asks for an input that crashes the vulnerable build; the Patch track (56 tasks) asks for a source fix. Vals runs every model itself with the mini-swe-agent harness.
Frontier
75.4%
Overall accuracy (mean of PoC and Patch pass rates)
MiMo V2.6 Flash · harness: mini-swe-agent (Vals run)
23 Sep 2026 · Source: Vals AI (benchmark maintainers)
Page data: overall 75.357 (stderr 5.438, $0.046 per test); PoC 65 (stderr 6.158); Patch 85.714 (stderr 4.718). Date is the page's "updated" field. Vals' model page ranks it 1 of 9 on CyberBench v1.1. Patch track leaders: Gemini 3.8 Flash and Claude Fable 5.1, 87.5 each.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Overall accuracy, which Vals ranks by default: the mean of the PoC and Patch track pass rates. A PoC passes when it crashes the vulnerable build and not the fixed build. A patch passes when the edited tree compiles, the original PoC no longer triggers the sanitizer error, and hidden holdout inputs match the maintainer's fix. In the PoC track the agent gets the source tree, fuzz target, and task metadata, but no crash input or bug description; in the Patch track it gets the source, the PoC, and the sanitizer report. Provider refusals count as failures.
Overall accuracy (mean of PoC and Patch pass rates): Unweighted mean of the PoC track pass rate (60 tasks) and the Patch track pass rate (56 tasks), as shown in the page's default "Overall" view. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- repo, cli
- Grading
- automated-tests
- Tasks
- 116
- Human reference
- No human score. The Patch grader compares behavior with the maintainer's own fix on hidden inputs.
- Contamination
- Tasks are OSS-Fuzz regressions harvested after CyberGym's original task window, chosen to be fresh. The bugs and fixes are public in the projects' histories, so later models may have seen them.
- Reuse
- Private task set ("dataset_type: private" in the page data). No license stated. (cite-only)
Limits to keep in mind
- Narrow scope: memory-safety, sanitizer, and undefined-behavior crashes in native code from OSS-Fuzz, not web, identity, malware, or network intrusion work. Source
- Small tracks (60 and 56 tasks): standard errors are 5 to 6 points, so many differences between models are within noise. Only nine models were on the board on 2026-09-23. Source
- Provider refusals count as failures, and some requests fall back to another model. GPT-6 Astra and Gemini 3.8 Flash score 0% on the PoC track, which pulls their Overall score down. Source
- v1.1 reran everything with offline sandboxes and a hardened grader; scores are not comparable with v1, and Vals posts about v1 results (for example GPT-5.6 Sol at 88.14%) refer to the old version. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Overall accuracy (mean of PoC and Patch pass rates) | Date | Source |
|---|---|---|---|
| MiMo V2.6 Flash · frontier harness: mini-swe-agent (Vals run) | 75.4% | 23 Sep 2026 | Vals AI · primary |
| DeepSeek V4.1 Flash harness: mini-swe-agent (Vals run) | 73.7% | 23 Sep 2026 | Vals AI · primary |
| Claude Fable 5.1 harness: mini-swe-agent (Vals run) | 70.4% | 23 Sep 2026 | Vals AI · primary |
Where sources disagree
The CyberBench v1.1 page's Key Takeaways say the PoC track "is led by Muse Spark 1.3 Max (63.3%)". The page's own leaderboard data (read 2026-09-24, "updated 2026-09-23") lists MiMo V2.6 Flash at 65 on the PoC track, ahead of Muse Spark 1.3 Max at 63.333. The text likely predates the addition of the Xiaomi models. We show the table values; the PoC track is not our headline metric.
- PoC track led by Muse Spark 1.3 Max, 63.3% — www.vals.ai (primary) · Key Takeaways text on the page, read 2026-09-24.
- PoC track led by MiMo V2.6 Flash, 65; Muse Spark 1.3 Max 63.333 — www.vals.ai (primary) · Leaderboard data embedded in the page HTML ("poc" section), read 2026-09-24.
We show MiMo V2.6 Flash 65 (PoC track, in the results notes). Status: open.
Timeline
- 30 Jun 2026 — Vals AI releases CyberBench, a fresh OSS-Fuzz successor to CyberGym with PoC and Patch tracks. Source
Where it sits in the atlas
Work ladder: Direct evidence for Security. On the work ladder it counts as 75.4% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Test performance of computer or information systems; Implement security measures for computer or information systems.
Go to the source
- Website www.vals.ai
- Full leaderboard www.vals.ai
Last checked 24 Sep 2026 against 4 primary sources. See an error? Tell us.
How to cite
Credit the original work first: CyberBench v1.1 by Vals AI.
Then, if you used this page:
Can Agents Work. "CyberBench v1.1: frontier results and sources." https://canagentswork.com/benchmarks/vals-cyberbench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-vals-cyberbench,
title = {{CyberBench v1.1: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/vals-cyberbench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}