Benchmarks / ITBench
ITBench
Built by IBM Research and UIUC · released 7 Feb 2025
IT automation scenarios that run in live Kubernetes environments. Agents must diagnose and repair incidents (SRE), assess and enforce compliance (CISO), and find cost problems (FinOps). IBM Research hosts the environments and a leaderboard.
Frontier
25%
SRE incidents resolved
ITBench-SRE-Agent-GPT-4o · harness: ITBench-SRE-Agent
2 May 2025 · Source: IBM Research (benchmark maintainers)
Single-trial table, 16 trials across incidents. The multi-trial table (162 trials) shows 24.79% for the same agent. Leaderboard date read as 2 May 2025.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
For SRE scenarios, whether the agent localizes the fault and repairs the incident so that the alert clears. The headline number here is the share of SRE incidents resolved. The CISO track scores compliance assessments and policy generation (1.0 is perfect). The FinOps track scores cost optimization and anomaly detection. The ICML 2025 paper covers 102 scenarios; the public repo ships a smaller open-source subset.
SRE incidents resolved: Share of SRE incidents that the agent repaired, as reported on the ITBench SRE leaderboard. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- live-system, cli
- Grading
- state-check, automated-tests
- Tasks
- 102
- Reuse
- Apache-2.0 (repository LICENSE file). Scenario tooling and sample scenarios are open source; hosted environments require registration. (open-apache)
Limits to keep in mind
- The SRE leaderboard was last updated in May 2025 and lists only IBM's own reference agent with three models. Newer models have not been posted there. Source
- The paper's headline (11.4% of SRE scenarios resolved) and the leaderboard's best entry (25.0%) use different scenario sets and agents, so they are not directly comparable. Source
- SREGym, a later benchmark from some of the same authors, reports that problems ported from ITBench and AIOpsLab are now easy for strong agents (mitigation above 80%). Source
- The arXiv v1 abstract (94 scenarios, 13.8% SRE) and the ICML version (102 scenarios, 11.4% SRE) report different counts and rates. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | SRE incidents resolved | Date | Source |
|---|---|---|---|
| Agents powered by state-of-the-art models (paper, ICML 2025) harness: ITBench reference agents | 11.4% | Jul 2025 | IBM Research · primary |
| ITBench-SRE-Agent-LLama-3-3-70B harness: ITBench-SRE-Agent | 12.5% | 2 May 2025 | IBM Research · primary |
| ITBench-SRE-Agent-GPT-4o · frontier harness: ITBench-SRE-Agent | 25% | 2 May 2025 | IBM Research · primary |
Where sources disagree
The ICML 2025 paper says agents resolve 11.4% of SRE scenarios. The arXiv v1 abstract says 13.8%. The official SRE leaderboard shows 25.0% for IBM's GPT-4o reference agent. The three numbers come from different scenario sets and dates. We show the leaderboard value because it is the maintainers' most recent published result.
- 11.4 — proceedings.mlr.press (primary) · ICML 2025 camera-ready abstract, 102 scenarios.
- 13.8 — arxiv.org (primary) · arXiv v1 abstract (2025-02-07), 94 scenarios.
- 25 — github.com (primary) · SRE leaderboard, ITBench-SRE-Agent-GPT-4o, 16 trials, updated 2 May 2025. Multi-trial table shows 24.79%.
We show 25. Status: open.
Timeline
- 7 Feb 2025 — IBM Research releases ITBench. Source
Where it sits in the atlas
Go to the source
- Website github.com
- Paper proceedings.mlr.press
- Full leaderboard github.com
- Code github.com
Last checked 23 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: ITBench by IBM Research and UIUC (https://proceedings.mlr.press/v267/jha25a.html).
Then, if you used this page:
Can Agents Work. "ITBench: frontier results and sources." https://canagentswork.com/benchmarks/itbench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-itbench,
title = {{ITBench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/itbench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}