Benchmarks / OSWorld-Verified
OSWorld-Verified (OSWorld)
Built by HKU, Salesforce AI Research, Carnegie Mellon University and University of Waterloo · Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, et al. · released 28 Jul 2025
A real Ubuntu desktop in a virtual machine with 369 computer tasks across Chrome, LibreOffice, GIMP, VLC, VS Code, Thunderbird, the file system, and multi-app workflows. The agent sees the screen and acts with mouse and keyboard. A script checks the final state of the machine.
Frontier
90.2%
Success rate
Intelligence-Indeed Agent
25 Jul 2026 · Source: HKU (benchmark maintainers)
Rank 1 on the verified board as of 2026-09-23. Agentic framework, 100 steps, 325.59 of 361 tasks. The board does not name the underlying model. Verified runs are executed by the OSWorld team.
The best result is at 90% or more of the ceiling.
What it measures
Share of desktop tasks an agent finishes correctly. Each task starts from a set initial state and ends with an execution-based check of files, application state, or settings. OSWorld-Verified (July 2025) fixed about 300 reported problems with tasks and checkers and moved evaluation to a parallel AWS setup. Eight Google Drive tasks may be skipped, giving a 361-task set. The official board runs agents at a 100-step budget.
Success rate: Share of tasks whose final-state check passes, on the 361-task set (or 369 with Google Drive tasks). Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- computer-use
- Grading
- state-check
- Tasks
- 369
- Human reference
- Humans completed 72.36% of tasks in the original 2024 study. The OSWorld-Verified report cites human performance as about 72%.
- Contamination
- All tasks and checkers are public. The maintainers have not published a contamination analysis. Some tasks depend on live websites that change over time.
- Reuse
- Apache 2.0 on the GitHub repository. Windows tasks need a licensed image. (open-apache)
Limits to keep in mind
- Agents now score above the 72% human figure, and the maintainers released OSWorld 2.0 in June 2026 as the next version with 108 longer tasks. OSWorld-Verified is still maintained but is no longer the hardest test. Source
- The original task set avoided long tasks, deep professional software, video, and real-time work, so it under-covers those parts of computer use. Source
- Scores depend on the step budget. The same model can score much higher at 100 steps than at 15, so compare entries at the same budget. Source
- Verified entries need the maintainers to run the agent on their own AWS platform, so the board lags model releases and lab-reported numbers (for example GPT-5.5 at 78.7%) are not on it. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
Where it sits in the atlas
Office and administrative supportDevOps, SRE, and IT operations (partial)Design, media, and writing (partial)
Go to the source
- Website os-world.github.io
- Paper arxiv.org
- Full leaderboard os-world.github.io
- Code github.com
- Announcement xlang.ai
Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: OSWorld-Verified by HKU, Salesforce AI Research, Carnegie Mellon University, and University of Waterloo (https://arxiv.org/abs/2404.07972).
Then, if you used this page:
Can Agents Work. "OSWorld-Verified: frontier results and sources." https://canagentswork.com/benchmarks/osworld-verified/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-osworld-verified,
title = {{OSWorld-Verified: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/osworld-verified/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}