Benchmarks / AISI multi-step cyber attack ranges

AISI multi-step cyber attack ranges (AISI cyber ranges)

Built by AI Security Institute · Linus Folkerts, Will Payne, Simon Inman, Philippos Giavridis, et al. · released 11 Mar 2026

Two private cyber ranges built for the UK AI Security Institute: "The Last Ones", a 32-step attack on a simulated corporate network that ends with data theft from a protected database, and "Cooling Tower", a 7-step attack on a simulated power plant's control system. A minimal ReAct agent on Kali Linux must chain the whole attack on its own. AISI ran seven models released between August 2024 and February 2026 and reports steps completed.

Frontier

15.6

Average steps completed, "The Last Ones" (100M tokens)

Claude Opus 4.6 · harness: AISI ReAct agent (Inspect)

11 Mar 2026 · Source: AI Security Institute (benchmark maintainers)

Table 1, 5 attempts at 100M tokens; best run 22 of 32, weakest 11. At 10M tokens (5 attempts) 9.8, max 11. A 100M-token attempt cost about $80 and took about 10 hours of wall-clock time. On "Cooling Tower" Opus 4.6 averaged 1.4 of 7 (max 2); GPT 5.3 Codex 1.2 (max 3).

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

AISI multi-step cyber attack ranges: Average steps completed, "The Last Ones" (100M tokens) over time, 3 recorded results. 010203040Feb 2026Mar 2026Mar 2026Mar 2026Mar 2026Mar 2026 Claude Sonnet 4.5: 9.4 (11 Mar 2026) Claude Opus 4.5: 11 (11 Mar 2026) Claude Opus 4.6: 15.6 (11 Mar 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/aisi-cyber-ranges/"><img src="https://canagentswork.com/og/benchmarks-aisi-cyber-ranges.png" width="600" height="315" alt="AISI multi-step cyber attack ranges: the best result is 15.6 (Claude Opus 4.6, 11 Mar 2026)." loading="lazy"></a>

Markdown:

[![AISI multi-step cyber attack ranges: the best result is 15.6 (Claude Opus 4.6, 11 Mar 2026).](https://canagentswork.com/og/benchmarks-aisi-cyber-ranges.png)](https://canagentswork.com/benchmarks/aisi-cyber-ranges/)

What it measures

How far an agent gets along a fixed attack chain with no active defenders. Each step has a flag; on "The Last Ones", reaching a flag counts all earlier steps as done. We show the average steps completed on "The Last Ones" (of 32) at a 100M-token budget, over five attempts; AISI also reports 10M-token runs and the maximum. Steps grow harder along the chain (reverse engineering, privilege escalation, cryptography), so step counts are not linear progress. On "Cooling Tower" the best models average 1.2 to 1.4 of 7 steps.

Average steps completed, "The Last Ones" (100M tokens): Mean number of the 32 attack steps completed per attempt at a 100M-token budget with AISI's standard ReAct agent and context compaction (Table 1, five attempts per model). Higher is better.

Status compares the frontier with a ceiling of 32.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
live-system, cli
Grading
state-check
Tasks
2
Human reference
AISI estimates that a human expert would need about 14 hours for "The Last Ones" and about 15 hours for "Cooling Tower". These are estimates from the number and complexity of the steps, not timed trials. The best single run (Opus 4.6, 22 of 32 steps) covered steps worth roughly 6 of the 14 hours.
Contamination
Held out: the ranges are not public. The paper's appendix describes both attack chains step by step, which the authors say trades against keeping the ranges held out over time.
Reuse
The ranges are private and were built for AISI by SpecterOps and Hack The Box. The paper describes the attack chains and asks to be excluded from training data (canary GUID). No data release. (cite-only)

Limits to keep in mind

  • Only AISI can run these ranges, so results come from one evaluator, with five attempts per model at 100M tokens. The authors say the small sample limits statistical power. Source
  • Strong variance between attempts: at the same model and budget, Opus 4.6 runs ranged from 11 to 22 of 32 steps. Source
  • No active defenders, detection penalties, or incident response. Alerts are logged but never block the agent. The ranges are also much smaller than production networks. Source
  • The metric omits target selection, initial access decisions, persistence, and infrastructure setup; the agent gets a starting point and an objective. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAverage steps completed, "The Last Ones" (100M tokens)DateSource
Claude Sonnet 4.5
harness: AISI ReAct agent (Inspect)
9.411 Mar 2026AI Security Institute · primary
Claude Opus 4.5
harness: AISI ReAct agent (Inspect)
1111 Mar 2026AI Security Institute · primary
Claude Opus 4.6 · frontier
harness: AISI ReAct agent (Inspect)
15.611 Mar 2026AI Security Institute · primary

Timeline

  • 11 Mar 2026 — UK AISI cyber ranges: Claude Opus 4.6 averages 15.6 of 32 attack steps at 100M tokens. Source

Where it sits in the atlas

Work ladder: Direct evidence for Security, but its headline gives partial credit and the source publishes no completion rate. It adds to the work measured, but it cannot set a step. How the ladder works

Work it measures (O*NET work activities): Test performance of computer or information systems.

Security

Penetration testingLong-horizon autonomy

Last checked 24 Sep 2026 against 3 primary sources. See an error? Tell us.

How to cite

Credit the original work first: AISI multi-step cyber attack ranges by AI Security Institute (https://arxiv.org/abs/2603.11214).

Then, if you used this page:

Can Agents Work. "AISI multi-step cyber attack ranges: frontier results and sources." https://canagentswork.com/benchmarks/aisi-cyber-ranges/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-aisi-cyber-ranges,
  title        = {{AISI multi-step cyber attack ranges: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/aisi-cyber-ranges/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}