Benchmarks / ITBench-AA
ITBench-AA
Built by Artificial Analysis and IBM Research · released 27 May 2026
Kubernetes incident diagnosis tasks built from IBM's ITBench and run by Artificial Analysis in a fixed agent harness. The model reads an offline incident snapshot with alerts, logs, traces, metrics, and topology, then names the root-cause Kubernetes entities.
Frontier
47%
ITBench-AA score
Claude Opus 4.7 (Adaptive Reasoning, Max Effort) · harness: Stirrup
27 May 2026 · Source: Artificial Analysis (benchmark maintainers)
Led the launch leaderboard. Most expensive at $5.38 per task.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Whether the model finds the exact set of root-cause entities for a Kubernetes incident. Scoring is average precision at full recall: a repeat scores 0 if any true root cause is missed, and otherwise scores the precision of the submitted list. The headline is the mean over 59 tasks and 3 repeats. Models run in the open-source Stirrup harness with shell access and a 100-turn cap.
ITBench-AA score: Mean average precision at full recall over 59 SRE tasks x 3 repeats. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- state-check
- Tasks
- 59
- Contamination
- 19 of the 59 tasks are new and held out. The other 40 are public ITBench scenarios.
- Reuse
- 40 public tasks and 19 held-out tasks; the held-out tasks are private to Artificial Analysis. (cite-only)
Limits to keep in mind
- Diagnosis only. Models name root-cause entities from a snapshot; they do not repair a live system. Repair is what the original ITBench SRE track scores. Source
- Strict scoring: naming one extra entity lowers the score, and missing one true root cause gives zero for that repeat. Models that investigate longer tend to add false positives. Source
- FinOps and CISO tasks were announced but not yet included at launch. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | ITBench-AA score | Date | Source |
|---|---|---|---|
| Gemini 3.1 Pro Preview harness: Stirrup | 30% | 27 May 2026 | Artificial Analysis · primary |
| GLM-5.1 (Reasoning) harness: Stirrup | 40% | 27 May 2026 | Artificial Analysis · primary |
| GPT-5.5 (xhigh) harness: Stirrup | 46% | 27 May 2026 | Artificial Analysis · primary |
| Claude Opus 4.7 (Adaptive Reasoning, Max Effort) · frontier harness: Stirrup | 47% | 27 May 2026 | Artificial Analysis · primary |
Timeline
- 27 May 2026 — Artificial Analysis and IBM launch ITBench-AA. Source
Where it sits in the atlas
Go to the source
- Website artificialanalysis.ai
- Full leaderboard artificialanalysis.ai
- Announcement artificialanalysis.ai
Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.
How to cite
Credit the original work first: ITBench-AA by Artificial Analysis and IBM Research.
Then, if you used this page:
Can Agents Work. "ITBench-AA: frontier results and sources." https://canagentswork.com/benchmarks/itbench-aa/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-itbench-aa,
title = {{ITBench-AA: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/itbench-aa/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}