Benchmarks / Terminal-Bench 4.0

Terminal-Bench 4.0 (TB 4.0)

Built by Laude Institute and Stanford University · Ryan Marten, Alex Shaw, Andy Konwinski · released 30 Jul 2026

66 hard, expert-written tasks that an agent must finish from the command line inside a container, across software, science, ML, business operations, hardware, security, and media. The maintainers run each agent and model pair five times on every task. It is a continuous benchmark: 3.0 (July 2026, 74 tasks) became 4.0 in August 2026 after task removals and fixes.

Frontier

58.2%

Resolution rate

Codex + GPT-6 Astra (max) · harness: Codex

3 Sep 2026 · Source: Laude Institute (benchmark maintainers)

Rank 1 on the 4.0.0 board: 192 of 330 trials resolved (58.2% ± 2.8), pass@2 0.6485, pass@5 0.7121, total cost $3,267.18, mean trial 2,796 s. Run by the maintainers with reasoning effort max. Statistically tied with four rows at 57.88. Anthropic reports Opus 5.5 at 66.4% in its own setup; that model is not on the board.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Laude Institute

Terminal-Bench 4.0: Resolution rate over time, 4 recorded results. 0%20%40%60%80%100%Aug 2026Aug 2026Aug 2026Aug 2026Sep 2026Sep 2026Sep 2026 Codex + GPT-5.6 Sol (max): 37.3% (27 Aug 2026) Claude Code + Opus 5 (max): 51.8% (27 Aug 2026) Claude Code + Fable 5.1 (max): 57.9% (3 Sep 2026) Codex + GPT-6 Astra (max): 58.2% (3 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/terminal-bench-4/"><img src="https://canagentswork.com/og/benchmarks-terminal-bench-4.png" width="600" height="315" alt="Terminal-Bench 4.0: the best result is 58.2% (Codex + GPT-6 Astra (max), 3 Sep 2026)." loading="lazy"></a>

Markdown:

[![Terminal-Bench 4.0: the best result is 58.2% (Codex + GPT-6 Astra (max), 3 Sep 2026).](https://canagentswork.com/og/benchmarks-terminal-bench-4.png)](https://canagentswork.com/benchmarks/terminal-bench-4/)

What it measures

Share of trials in which the agent's output passes the task's verifier, over 66 tasks with 5 trials each (330 trials), shown with a 95% confidence interval. The agent works in its own container with an 8-hour limit and open internet; artifacts are copied to a separate verifier container. Task metadata gives expert time estimates from 45 minutes to 60 hours (median 4 hours). Categories at 4.0.0: software 18, science 14, ML 11, operations 9, hardware 5, security 5, media 4.

Resolution rate: Percentage of the 330 trials (66 tasks, 5 trials each) that the verifier marks as resolved, with a 95% confidence interval. The leaderboard ranks by this value. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli
Grading
automated-tests
Tasks
66
Contamination
Tasks are public on GitHub with a canary GUID that the maintainers ask never to appear in training data. Agents have open internet access and are told not to look for task-specific solutions online. 4.0 removed two tasks because public solutions existed.
Reuse
Apache-2.0 LICENSE at the root of the terminal-bench repository, which holds the tasks. No separate data license is stated. (open-apache)

Limits to keep in mind

  • The top five rows overlap within their confidence intervals (58.18 ± 2.79, then four rows at 57.88), so the leader is not statistically distinct. Each row is an agent plus a model at one reasoning effort, run by the maintainers. Source
  • The benchmark is "continuous" with semantic versioning. Breaking changes (4.0 removed 8 tasks, fixed 19, and reset all agent timeouts to 8 hours) re-run every trial, and 4.1 (verifier changes) and 5.0 (new tasks) are planned, so numbers do not transfer between major versions. Source
  • Lab-reported scores are higher than the board: Anthropic reports Claude Opus 5.5 at 66.4% (xhigh) in its own setup and says its setup reproduces the board's Opus 5 at 52.3% versus 51.8%. Opus 5.5 is not on the public board. Source
  • The maintainers observed that Sonnet 5 hit timeouts and output-token limits, and that remaining errors in 4.0 are largely model refusals, so some low scores reflect refusals and limits rather than capability. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemResolution rateDateSource
Claude Code + Fable 5.1 (max)
harness: Claude Code
57.9%3 Sep 2026Laude Institute · primary
Codex + GPT-6 Astra (max) · frontier
harness: Codex
58.2%3 Sep 2026Laude Institute · primary
Codex + GPT-5.6 Sol (max)
harness: Codex
37.3%27 Aug 2026Laude Institute · primary
Claude Code + Opus 5 (max)
harness: Claude Code
51.8%27 Aug 2026Laude Institute · primary

Timeline

  • 3 Sep 2026 — GPT-6 Astra in Codex reaches 58.2% on Terminal-Bench 4.0. Source
  • 28 Aug 2026 — Terminal-Bench 4.0 removes 8 tasks, fixes 19, and re-runs every trial on 66 tasks. Source
  • 30 Jul 2026 — Terminal-Bench 3.0 launches a continuous benchmark with 74 tasks; the best agent scores 34.4%. Source

Where it sits in the atlas

Work ladder: Direct evidence for Software engineering. On the work ladder it counts as 58.2% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Program computer systems or production equipment; Design computer or information systems or applications; Set up computer systems, networks, or other information systems; Resolve computer problems; Analyze scientific or applied data using mathematical principles.

Software engineeringScience and research (partial)ML and research engineering (partial)DevOps, SRE, and IT operations (partial)Management and business operations (partial)Security (partial)Architecture and engineering (partial)

Terminal operationsFeature developmentML engineering

Last checked 24 Sep 2026 against 8 primary sources. See an error? Tell us.

How to cite

Credit the original work first: Terminal-Bench 4.0 by Laude Institute and Stanford University.

Then, if you used this page:

Can Agents Work. "Terminal-Bench 4.0: frontier results and sources." https://canagentswork.com/benchmarks/terminal-bench-4/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-terminal-bench-4,
  title        = {{Terminal-Bench 4.0: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/terminal-bench-4/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}