Benchmarks / Terminal-Bench 4.0
Terminal-Bench 4.0 (TB 4.0)
Built by Laude Institute and Stanford University · Ryan Marten, Alex Shaw, Andy Konwinski · released 30 Jul 2026
66 hard, expert-written tasks that an agent must finish from the command line inside a container, across software, science, ML, business operations, hardware, security, and media. The maintainers run each agent and model pair five times on every task. It is a continuous benchmark: 3.0 (July 2026, 74 tasks) became 4.0 in August 2026 after task removals and fixes.
Frontier
58.2%
Resolution rate
Codex + GPT-6 Astra (max) · harness: Codex
3 Sep 2026 · Source: Laude Institute (benchmark maintainers)
Rank 1 on the 4.0.0 board: 192 of 330 trials resolved (58.2% ± 2.8), pass@2 0.6485, pass@5 0.7121, total cost $3,267.18, mean trial 2,796 s. Run by the maintainers with reasoning effort max. Statistically tied with four rows at 57.88. Anthropic reports Opus 5.5 at 66.4% in its own setup; that model is not on the board.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Share of trials in which the agent's output passes the task's verifier, over 66 tasks with 5 trials each (330 trials), shown with a 95% confidence interval. The agent works in its own container with an 8-hour limit and open internet; artifacts are copied to a separate verifier container. Task metadata gives expert time estimates from 45 minutes to 60 hours (median 4 hours). Categories at 4.0.0: software 18, science 14, ML 11, operations 9, hardware 5, security 5, media 4.
Resolution rate: Percentage of the 330 trials (66 tasks, 5 trials each) that the verifier marks as resolved, with a 95% confidence interval. The leaderboard ranks by this value. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli
- Grading
- automated-tests
- Tasks
- 66
- Contamination
- Tasks are public on GitHub with a canary GUID that the maintainers ask never to appear in training data. Agents have open internet access and are told not to look for task-specific solutions online. 4.0 removed two tasks because public solutions existed.
- Reuse
- Apache-2.0 LICENSE at the root of the terminal-bench repository, which holds the tasks. No separate data license is stated. (open-apache)
Limits to keep in mind
- The top five rows overlap within their confidence intervals (58.18 ± 2.79, then four rows at 57.88), so the leader is not statistically distinct. Each row is an agent plus a model at one reasoning effort, run by the maintainers. Source
- The benchmark is "continuous" with semantic versioning. Breaking changes (4.0 removed 8 tasks, fixed 19, and reset all agent timeouts to 8 hours) re-run every trial, and 4.1 (verifier changes) and 5.0 (new tasks) are planned, so numbers do not transfer between major versions. Source
- Lab-reported scores are higher than the board: Anthropic reports Claude Opus 5.5 at 66.4% (xhigh) in its own setup and says its setup reproduces the board's Opus 5 at 52.3% versus 51.8%. Opus 5.5 is not on the public board. Source
- The maintainers observed that Sonnet 5 hit timeouts and output-token limits, and that remaining errors in 4.0 are largely model refusals, so some low scores reflect refusals and limits rather than capability. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Resolution rate | Date | Source |
|---|---|---|---|
| Claude Code + Fable 5.1 (max) harness: Claude Code | 57.9% | 3 Sep 2026 | Laude Institute · primary |
| Codex + GPT-6 Astra (max) · frontier harness: Codex | 58.2% | 3 Sep 2026 | Laude Institute · primary |
| Codex + GPT-5.6 Sol (max) harness: Codex | 37.3% | 27 Aug 2026 | Laude Institute · primary |
| Claude Code + Opus 5 (max) harness: Claude Code | 51.8% | 27 Aug 2026 | Laude Institute · primary |
Timeline
Where it sits in the atlas
Work ladder: Direct evidence for Software engineering. On the work ladder it counts as 58.2% of tasks completed. How the ladder works
Work it measures (O*NET work activities): Program computer systems or production equipment; Design computer or information systems or applications; Set up computer systems, networks, or other information systems; Resolve computer problems; Analyze scientific or applied data using mathematical principles.
Software engineeringScience and research (partial)ML and research engineering (partial)DevOps, SRE, and IT operations (partial)Management and business operations (partial)Security (partial)Architecture and engineering (partial)
Go to the source
- Website www.tbench.ai
- Full leaderboard www.tbench.ai
- Code github.com
- Announcement www.tbench.ai
- Dataset hub.harborframework.com
Last checked 24 Sep 2026 against 8 primary sources. See an error? Tell us.
How to cite
Credit the original work first: Terminal-Bench 4.0 by Laude Institute and Stanford University.
Then, if you used this page:
Can Agents Work. "Terminal-Bench 4.0: frontier results and sources." https://canagentswork.com/benchmarks/terminal-bench-4/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-terminal-bench-4,
title = {{Terminal-Bench 4.0: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/terminal-bench-4/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}