Benchmarks / TheAgentCompany
TheAgentCompany
Built by Carnegie Mellon University · Frank F. Xu, Yufan Song, Boxuan Li, Yuxuan Tang, et al. · released 18 Dec 2024
A simulated software company with self-hosted GitLab, a project tracker (Plane), ownCloud file storage, and RocketChat. An agent gets work tasks like a digital employee and must browse, code, run programs, and talk to simulated coworkers to finish them.
Frontier
42.9%
Tasks resolved
TTE-MatrixAgent + DeepSeek-V3.2 · harness: TTE-MatrixAgent
10 Nov 2025 · Source: Carnegie Mellon University (benchmark maintainers)
Top entry on the leaderboard, marked checked by the maintainers. Partial-credit score 52.4, 29.91 steps and $0.40 per task on average. Environment model Qwen Plus. The agent code is not open source. Newest entry on the board as of 2026-09-23.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Whether an agent completes 175 work tasks in a sealed company environment by browsing the web, writing code, running programs, and messaging coworkers. Checkers inspect the final state of the systems (for example a merged change, a filed document, or a message sent) and award full credit only when every checkpoint passes. Simulated coworkers are language-model agents that the agent must ask for information.
Tasks resolved: Share of the 175 tasks fully completed. The leaderboard also shows a partial-credit score that rewards passed checkpoints. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- simulated-workplace, browser, cli
- Grading
- state-check
- Tasks
- 175
- Contamination
- All tasks and evaluators are public, so models trained after December 2024 may have seen them. The maintainers have not published a contamination analysis.
- Reuse
- MIT license on the GitHub repository. Task images, evaluators, and company data are public. (open-mit)
Limits to keep in mind
- The leaderboard's newest entry is from November 2025, so it does not reflect models released in 2026. Source
- Coworkers are simulated by a language model, and results depend on which model plays the environment (the board lists it per entry). Source
- The company is a small software firm, so tasks lean toward engineering and internal tooling rather than the full range of office work. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Tasks resolved | Date | Source |
|---|---|---|---|
| TTE-MatrixAgent + DeepSeek-V3.2 · frontier harness: TTE-MatrixAgent | 42.9% | 10 Nov 2025 | Carnegie Mellon University · primary |
| OpenHands-Versa + Claude Sonnet 4 harness: OpenHands-Versa | 33.1% | 14 Jun 2025 | Carnegie Mellon University · primary |
| OpenHands + Gemini 2.5 Pro harness: OpenHands | 30.3% | 10 May 2025 | Carnegie Mellon University · primary |
| OpenHands + Claude 3.5 Sonnet harness: OpenHands | 24% | 17 Dec 2024 | Carnegie Mellon University · primary |
Timeline
- 18 Dec 2024 — CMU releases TheAgentCompany. Source
Where it sits in the atlas
Software engineeringManagement and business operations (partial)Office and administrative support (partial)
Enterprise workflowsFeature developmentLong-horizon autonomy
Go to the source
- Website the-agent-company.com
- Paper arxiv.org
- Full leaderboard the-agent-company.com
- Code github.com
Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.
How to cite
Credit the original work first: TheAgentCompany by Carnegie Mellon University (https://arxiv.org/abs/2412.14161).
Then, if you used this page:
Can Agents Work. "TheAgentCompany: frontier results and sources." https://canagentswork.com/benchmarks/theagentcompany/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-theagentcompany,
title = {{TheAgentCompany: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/theagentcompany/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}