Benchmarks / DARPA AI Cyber Challenge, Final Competition
DARPA AI Cyber Challenge, Final Competition (AIxCC final)
Built by DARPA and ARPA-H · released 8 Aug 2025
A two-year DARPA competition, run with ARPA-H, in which seven teams built autonomous cyber reasoning systems (CRSs) that find and patch vulnerabilities in open-source software without human help. In the August 2025 final, the systems worked unattended over 63 synthetic vulnerabilities injected into real projects (54 million lines of code). Team Atlanta won. The competition is over, so no new results will come.
Frontier
43
Synthetic vulnerabilities found (scored round)
Team Atlanta
8 Aug 2025 · Source: DARPA (benchmark maintainers)
First place. Scored round: 43 vulnerabilities found, 31 successful patches, total score 393. Team Atlanta is Georgia Tech, Samsung Research, KAIST, and POSTECH (DARPA). The AIxCC site lists results by team name, not by system name.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether an autonomous system can find a vulnerability in a large real codebase, prove it with a crashing input, write a patch that fixes it, and report it (bug report and pairing of proof and patch). The official scoring rewards fast, accurate, and high-quality submissions; Team Atlanta scored 393 points. We show the count of synthetic vulnerabilities each team found, out of 63. All seven systems together found 54 of 63 (86%) and patched 43; they also found 18 real, previously unknown vulnerabilities.
Synthetic vulnerabilities found (scored round): Number of the 63 synthetic vulnerabilities that one team's cyber reasoning system found in the scored round of the Final Competition, as listed on aicyberchallenge.com. Higher is better.
Status compares the frontier with a ceiling of 63.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- repo, cli
- Grading
- automated-tests, state-check
- Tasks
- 63
- Human reference
- No human baseline. DARPA reports an average cost of about $152 per competition task and an average of 45 minutes to submit a patch.
- Contamination
- The challenges were real open-source projects with injected synthetic vulnerabilities. The projects are public; the injected bugs were new for the competition.
- Reuse
- DARPA says all seven CRSs are released under an OSI-approved open-source license and that the competition challenges, framework, and telemetry will also be open sourced (archive.aicyberchallenge.com). We did not read the licenses of the individual releases. (cite-only)
Limits to keep in mind
- A one-off competition. Seven systems ran once in August 2025; there is no leaderboard to which new systems can be added. Source
- DARPA's result page says the systems "patched 68% of the vulnerabilities identified". 43 of 63 is 68%; 43 of the 54 found would be 79.6%. The editor's note ties 68% to the 63 total. We show counts, not this percent (see the conflict file). Source
- The systems were competition-scale CRSs that teams built over two years, "combining AI with other cyber defense techniques" (DARPA). Results measure the teams' systems, not one model. Source
- The official ranking uses a composite score (speed, accuracy, patch quality, bug reports), not the count of vulnerabilities found. 42-b3yond-6ug found 41 but patched 3 and placed sixth. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Where sources disagree
DARPA's result page says the systems "patched 68% of the vulnerabilities identified". The systems found 54 of 63 synthetic vulnerabilities and patched 43. 43 of 63 is 68%; 43 of the 54 identified would be 79.6%. The sentence and the number do not match. The same page's editor's note ties 68% to the 63 total. We show the counts (54 found, 43 patched, of 63) and not the percent.
- 68% of the vulnerabilities identified were patched (43 of 54 would be 79.6%) — www.darpa.mil (primary) · Body text of the DARPA results page, dated Aug. 8, 2025, read 2026-09-24.
- 68% patched, of the 63 synthetic vulnerabilities (43 of 63) — www.darpa.mil (primary) · Editor's note on the same page: the total was corrected from 70 to 63, "which means competitors discovered 86% of the synthetic vulnerabilities and patched 68%". The same note is on aicyberchallenge.com.
We show 43 patched of 63 (counts only). Status: resolved.
Timeline
- 8 Aug 2025 — AIxCC final: seven autonomous systems find 54 of 63 planted vulnerabilities and patch 43. Source
Where it sits in the atlas
Work ladder: Direct evidence for Security, but its headline gives partial credit and the source publishes no completion rate. It adds to the work measured, but it cannot set a step. How the ladder works
Work it measures (O*NET work activities): Test performance of computer or information systems; Implement security measures for computer or information systems.
Go to the source
- Website www.aicyberchallenge.com
- Full leaderboard www.aicyberchallenge.com
- Code archive.aicyberchallenge.com
- Announcement www.darpa.mil
Last checked 24 Sep 2026 against 3 primary sources. See an error? Tell us.
How to cite
Credit the original work first: DARPA AI Cyber Challenge, Final Competition by DARPA and ARPA-H.
Then, if you used this page:
Can Agents Work. "DARPA AI Cyber Challenge, Final Competition: frontier results and sources." https://canagentswork.com/benchmarks/aixcc-final/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-aixcc-final,
title = {{DARPA AI Cyber Challenge, Final Competition: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/aixcc-final/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}