Benchmarks / Terminal-Bench 2.0

Terminal-Bench 2.0 (TB 2.0)

Built by Laude Institute and Stanford University · Mike Merrill, Alex Shaw · released 7 Nov 2025

89 hand-verified tasks that an agent must finish inside a Docker container from the command line. Tasks include compiling old software, training small models, debugging code, cracking archives, and setting up servers. It replaced Terminal-Bench 1.0 in November 2025 and runs on the Harbor framework.

Frontier

84.7%

Resolution rate

NexAU-AHE + GPT-5.5 · harness: NexAU-AHE

23 Apr 2026 · Source: Laude Institute (benchmark maintainers)

Rank 1 on the frozen 2.0 leaderboard (84.7 ± 2.1). Self-submitted by china-qijizhifeng and not verified by the maintainers. Statistically tied with LemonHarness (84.5 ± 2.6), Capy + GPT-5.5 (83.1 ± 2.1), and the verified Codex CLI + GPT-5.5 (82.2 ± 2.2).

Retired

Maintainers or a major user stopped using it as a frontier measure.

Retired on 6 May 2026: The maintainers replaced 2.0 with Terminal-Bench 2.1 on 2026-05-06 after finding problems in 28 of the 89 tasks (changed external dependencies, resource limits, and instructions that did not match tests). They then moved to 3.0 (2026-07-30) and 4.0 (2026-08-28), saying many tasks had saturated. The 2.0 leaderboard is frozen but still hosted. Source.

See the full leaderboard at Laude Institute

Terminal-Bench 2.0: Resolution rate over time, 5 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026 Codex CLI + GPT-5: 49.6% (7 Aug 2025) Ante + Gemini 3 Pro: 69.4% (18 Nov 2025) Codex CLI + GPT-5.5: 82.2% (23 Apr 2026) NexAU-AHE + GPT-5.5: 84.7% (23 Apr 2026) LemonHarness (Gemini 3.1 Pro Preview + GPT-5.3-Codex): 84.5% (14 May 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/terminal-bench-2/"><img src="https://canagentswork.com/og/benchmarks-terminal-bench-2.png" width="600" height="315" alt="Terminal-Bench 2.0: the best result is 84.7% (NexAU-AHE + GPT-5.5, 23 Apr 2026)." loading="lazy"></a>

Markdown:

[![Terminal-Bench 2.0: the best result is 84.7% (NexAU-AHE + GPT-5.5, 23 Apr 2026).](https://canagentswork.com/og/benchmarks-terminal-bench-2.png)](https://canagentswork.com/benchmarks/terminal-bench-2/)

What it measures

Share of tasks whose tests pass after the agent finishes, averaged over several trials with a 95% confidence interval. Each leaderboard row is a harness plus a model, so scores depend on both. Some rows are verified by the maintainers; most are self-submitted.

Resolution rate: Percentage of tasks resolved across trials, shown with a 95% confidence interval. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
automated-tests
Tasks
89
Contamination
Tasks are public on GitHub. The maintainers add a canary string and ask that the data never appear in training corpora.
Reuse
Apache-2.0 (terminal-bench-2 repository). (open-apache)

Limits to keep in mind

  • 28 of 89 tasks had defects that 2.1 fixed: nine depended on external resources that changed, eight had resource budgets too small for valid solutions, and some instructions did not match their tests. Scores on 2.0 understate some systems. Source
  • The top entries overlap within their confidence intervals (84.7 ± 2.1, 84.5 ± 2.6, 83.1 ± 2.1, 82.2 ± 2.2), so the leader is not statistically distinct. Source
  • Most top rows are self-submitted harnesses that the maintainers did not verify. The best maintainer-verified row is Codex CLI with GPT-5.5 at 82.2%. Source
  • The maintainers say many Terminal-Bench tasks have saturated, which is why they built 3.0. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemResolution rateDateSource
LemonHarness (Gemini 3.1 Pro Preview + GPT-5.3-Codex)
harness: LemonHarness
84.5%14 May 2026Laude Institute · primary
Codex CLI + GPT-5.5
harness: Codex CLI
82.2%23 Apr 2026Laude Institute · primary
NexAU-AHE + GPT-5.5 · frontier
harness: NexAU-AHE
84.7%23 Apr 2026Laude Institute · primary
Ante + Gemini 3 Pro
harness: Ante
69.4%18 Nov 2025Laude Institute · primary
Codex CLI + GPT-5
harness: Codex CLI
49.6%7 Aug 2025Laude Institute · primary

Timeline

  • 6 May 2026 — Terminal-Bench 2.1 fixes 28 tasks and replaces 2.0. Source
  • 7 Nov 2025 — Terminal-Bench 2.0 and Harbor released. Source

Where it sits in the atlas

DevOps, SRE, and IT operationsSoftware engineering (partial)ML and research engineering (partial)

Terminal operationsFeature development

Last checked 23 Sep 2026 against 8 primary sources. See an error? Tell us.

How to cite

Credit the original work first: Terminal-Bench 2.0 by Laude Institute and Stanford University.

Then, if you used this page:

Can Agents Work. "Terminal-Bench 2.0: frontier results and sources." https://canagentswork.com/benchmarks/terminal-bench-2/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-terminal-bench-2,
  title        = {{Terminal-Bench 2.0: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/terminal-bench-2/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}