Benchmarks / ARTEMIS live penetration test (agents vs. professionals)

ARTEMIS live penetration test (agents vs. professionals) (ARTEMIS pentest)

Built by Stanford University, Carnegie Mellon University and Gray Swan AI · Justin W. Lin, Eliot Krzysztof Jones, Donovan Julian Jasper, Ethan Jun-shen Ho, et al. · released 10 Dec 2025

A live penetration test on a large university's computer science network (about 8,000 hosts on 12 subnets), run once with ten professional penetration testers, six existing agent scaffolds, and ARTEMIS, the authors' multi-agent scaffold. Everyone worked from the same Kali Linux VM with student-level credentials for a 10-hour engagement. The study ranks people and agents on one board.

Frontier

95.2

Total score (complexity plus weighted severity of valid findings)

ARTEMIS A2 (ensemble supervisor, Claude Sonnet 4 sub-agents) · agent: ARTEMIS

10 Dec 2025 · Source: Stanford University (benchmark maintainers)

Second of 15 ranked entries (Table 1). 11 submissions, 82% valid (9 valid findings), severity score 54, complexity score 41.2. Ran for 16 hours; only the first 10 hours were scored, to match the humans. Run cost $944.07 ($59 per hour). The top human (P1) scored 111.4.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

ARTEMIS live penetration test (agents vs. professionals): Total score (complexity plus weighted severity of valid findings) over time, 3 recorded results. 020406080100120Nov 2025Nov 2025Dec 2025Dec 2025Dec 2025Dec 2025human parity ARTEMIS A2 (ensemble supervisor, Claude Sonnet 4 sub-agents): 95.2 (10 Dec 2025) ARTEMIS A1 (GPT-5): 53.2 (10 Dec 2025) Codex (GPT-5): 38.6 (10 Dec 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/artemis-live-pentest/"><img src="https://canagentswork.com/og/benchmarks-artemis-live-pentest.png" width="600" height="315" alt="ARTEMIS live penetration test (agents vs. professionals): the best result is 95.2 (ARTEMIS A2 (ensemble supervisor, Claude Sonnet 4 sub-agents), 10 Dec 2025)." loading="lazy"></a>

Markdown:

[![ARTEMIS live penetration test (agents vs. professionals): the best result is 95.2 (ARTEMIS A2 (ensemble supervisor, Claude Sonnet 4 sub-agents), 10 Dec 2025).](https://canagentswork.com/og/benchmarks-artemis-live-pentest.png)](https://canagentswork.com/benchmarks/artemis-live-pentest/)

What it measures

The value of the vulnerabilities each participant found and reported in 10 hours on a real production network. Each valid finding earns a technical complexity score (detection plus exploit complexity, with a penalty for unexploited findings) plus a severity weight (8 for critical down to 1 for informational); the total is the participant's score. The best ARTEMIS configuration (A2, an ensemble of supervisor models with Claude Sonnet 4 sub-agents) placed second of 15 ranked entries with 11 submissions, 9 of them valid (82%), behind one professional and ahead of nine. Codex, CyAgent, Claude Code, Incalmo, and MAPTA did far worse.

Total score (complexity plus weighted severity of valid findings): Sum over valid findings of technical complexity and severity weight, as in Table 1 of the paper. The scale has no fixed maximum; the best human scored 111.4. Higher is better.

Status compares the frontier with human parity at 59. Median total score of the ten professional participants in Table 1 (65.0 and 53.0 are the fifth and sixth values). We computed the median; the paper ranks but does not state it.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
live-system, cli
Grading
human-expert, rubric
Tasks
1
Human reference
Ten professional penetration testers, recruited for their experience and paid a flat $2,000, each committed at least 10 working hours. Their total scores ranged from 25.4 to 111.4 (median 59.0, computed by us from Table 1). The paper prices human pentesters at $60 per hour; ARTEMIS A1 cost $291.47 for the run ($18.21 per hour) and A2 $944.07 ($59 per hour).
Contamination
Not applicable in the usual sense: the target was a live network, not a dataset. Agents and people attacked the same environment in the same period.
Reuse
Study artifacts are described as released with the code, but we found no data license. The target is a live network and cannot be re-run. The ARTEMIS repository is Apache-2.0. (cite-only)

Limits to keep in mind

  • One engagement on one network, with ten people and seven agent configurations. The authors say the sample sizes preclude hypothesis tests with statistical power, and the network cannot be re-tested. Source
  • Ten hours of active time (and four days of access), where real penetration tests usually run one to two weeks. The IT team knew about the test and approved flagged actions, so there were no real defenders. Source
  • Agents submitted more false positives than people (A1 had a 55% valid-submission rate, A2 82%) and struggled with GUI-only targets. The scoring rewards technically complex exploits over easy, high-impact findings, by design. Source
  • The authors built and evaluated their own scaffold, and the Cybench check found no significant scaffold uplift on single-host CTF tasks. The comparison with other scaffolds is not a controlled ablation. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemTotal score (complexity plus weighted severity of valid findings)DateSource
ARTEMIS A2 (ensemble supervisor, Claude Sonnet 4 sub-agents) · frontier95.210 Dec 2025Stanford University · primary
ARTEMIS A1 (GPT-5)53.210 Dec 2025Stanford University · primary
Codex (GPT-5)38.610 Dec 2025Stanford University · primary

Timeline

  • 10 Dec 2025 — ARTEMIS agent places second of 15 against ten professional pentesters on a live university network. Source

Where it sits in the atlas

Work ladder: Direct evidence for Security, but it is a field study of real deployments. It adds to the work measured, but it cannot set a step. How the ladder works

Work it measures (O*NET work activities): Test performance of computer or information systems.

Security

Penetration testingVulnerability researchLong-horizon autonomy

Last checked 24 Sep 2026 against 4 primary sources. See an error? Tell us.

How to cite

Credit the original work first: ARTEMIS live penetration test (agents vs. professionals) by Stanford University, Carnegie Mellon University, and Gray Swan AI (https://arxiv.org/abs/2512.09882).

Then, if you used this page:

Can Agents Work. "ARTEMIS live penetration test (agents vs. professionals): frontier results and sources." https://canagentswork.com/benchmarks/artemis-live-pentest/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-artemis-live-pentest,
  title        = {{ARTEMIS live penetration test (agents vs. professionals): frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/artemis-live-pentest/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}