Benchmarks / Terminal-Bench 2.0
Terminal-Bench 2.0 (TB 2.0)
Built by Laude Institute and Stanford University · Mike Merrill, Alex Shaw · released 7 Nov 2025
89 hand-verified tasks that an agent must finish inside a Docker container from the command line. Tasks include compiling old software, training small models, debugging code, cracking archives, and setting up servers. It replaced Terminal-Bench 1.0 in November 2025 and runs on the Harbor framework.
Frontier
84.7%
Resolution rate
NexAU-AHE + GPT-5.5 · harness: NexAU-AHE
23 Apr 2026 · Source: Laude Institute (benchmark maintainers)
Rank 1 on the frozen 2.0 leaderboard (84.7 ± 2.1). Self-submitted by china-qijizhifeng and not verified by the maintainers. Statistically tied with LemonHarness (84.5 ± 2.6), Capy + GPT-5.5 (83.1 ± 2.1), and the verified Codex CLI + GPT-5.5 (82.2 ± 2.2).
Maintainers or a major user stopped using it as a frontier measure.
Retired on 6 May 2026: The maintainers replaced 2.0 with Terminal-Bench 2.1 on 2026-05-06 after finding problems in 28 of the 89 tasks (changed external dependencies, resource limits, and instructions that did not match tests). They then moved to 3.0 (2026-07-30) and 4.0 (2026-08-28), saying many tasks had saturated. The 2.0 leaderboard is frozen but still hosted. Source.
What it measures
Share of tasks whose tests pass after the agent finishes, averaged over several trials with a 95% confidence interval. Each leaderboard row is a harness plus a model, so scores depend on both. Some rows are verified by the maintainers; most are self-submitted.
Resolution rate: Percentage of tasks resolved across trials, shown with a 95% confidence interval. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- automated-tests
- Tasks
- 89
- Contamination
- Tasks are public on GitHub. The maintainers add a canary string and ask that the data never appear in training corpora.
- Reuse
- Apache-2.0 (terminal-bench-2 repository). (open-apache)
Limits to keep in mind
- 28 of 89 tasks had defects that 2.1 fixed: nine depended on external resources that changed, eight had resource budgets too small for valid solutions, and some instructions did not match their tests. Scores on 2.0 understate some systems. Source
- The top entries overlap within their confidence intervals (84.7 ± 2.1, 84.5 ± 2.6, 83.1 ± 2.1, 82.2 ± 2.2), so the leader is not statistically distinct. Source
- Most top rows are self-submitted harnesses that the maintainers did not verify. The best maintainer-verified row is Codex CLI with GPT-5.5 at 82.2%. Source
- The maintainers say many Terminal-Bench tasks have saturated, which is why they built 3.0. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Resolution rate | Date | Source |
|---|---|---|---|
| LemonHarness (Gemini 3.1 Pro Preview + GPT-5.3-Codex) harness: LemonHarness | 84.5% | 14 May 2026 | Laude Institute · primary |
| Codex CLI + GPT-5.5 harness: Codex CLI | 82.2% | 23 Apr 2026 | Laude Institute · primary |
| NexAU-AHE + GPT-5.5 · frontier harness: NexAU-AHE | 84.7% | 23 Apr 2026 | Laude Institute · primary |
| Ante + Gemini 3 Pro harness: Ante | 69.4% | 18 Nov 2025 | Laude Institute · primary |
| Codex CLI + GPT-5 harness: Codex CLI | 49.6% | 7 Aug 2025 | Laude Institute · primary |
Timeline
Where it sits in the atlas
DevOps, SRE, and IT operationsSoftware engineering (partial)ML and research engineering (partial)
Go to the source
- Website www.tbench.ai
- Full leaderboard www.tbench.ai
- Code github.com
- Announcement www.tbench.ai
- Dataset hub.harborframework.com
Last checked 23 Sep 2026 against 8 primary sources. See an error? Tell us.
How to cite
Credit the original work first: Terminal-Bench 2.0 by Laude Institute and Stanford University.
Then, if you used this page:
Can Agents Work. "Terminal-Bench 2.0: frontier results and sources." https://canagentswork.com/benchmarks/terminal-bench-2/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-terminal-bench-2,
title = {{Terminal-Bench 2.0: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/terminal-bench-2/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}