Benchmarks / OSWorld-Verified

OSWorld-Verified (OSWorld)

Built by HKU, Salesforce AI Research, Carnegie Mellon University and University of Waterloo · Tianbao Xie, Danyang Zhang, Jixuan Chen, Xiaochuan Li, et al. · released 28 Jul 2025

A real Ubuntu desktop in a virtual machine with 369 computer tasks across Chrome, LibreOffice, GIMP, VLC, VS Code, Thunderbird, the file system, and multi-app workflows. The agent sees the screen and acts with mouse and keyboard. A script checks the final state of the machine.

Frontier

90.2%

Success rate

Intelligence-Indeed Agent

25 Jul 2026 · Source: HKU (benchmark maintainers)

Rank 1 on the verified board as of 2026-09-23. Agentic framework, 100 steps, 325.59 of 361 tasks. The board does not name the underlying model. Verified runs are executed by the OSWorld team.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at HKU

OSWorld-Verified: Success rate over time, 5 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026Jul 2026 CoACT-1: 60.8% (4 Aug 2025) agent s3 w/ GPT-5 bBoN (N=10): 69.9% (4 Oct 2025) Holo3-35B-A3B: 82.6% (20 Apr 2026) Intelligence-Indeed Agent: 90.2% (25 Jul 2026) claude-fable-5[1m]: 86% (1 Aug 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/osworld-verified/"><img src="https://canagentswork.com/og/benchmarks-osworld-verified.png" width="600" height="315" alt="OSWorld-Verified: the best result is 90.2% (Intelligence-Indeed Agent, 25 Jul 2026)." loading="lazy"></a>

Markdown:

[![OSWorld-Verified: the best result is 90.2% (Intelligence-Indeed Agent, 25 Jul 2026).](https://canagentswork.com/og/benchmarks-osworld-verified.png)](https://canagentswork.com/benchmarks/osworld-verified/)

What it measures

Share of desktop tasks an agent finishes correctly. Each task starts from a set initial state and ends with an execution-based check of files, application state, or settings. OSWorld-Verified (July 2025) fixed about 300 reported problems with tasks and checkers and moved evaluation to a parallel AWS setup. Eight Google Drive tasks may be skipped, giving a 361-task set. The official board runs agents at a 100-step budget.

Success rate: Share of tasks whose final-state check passes, on the 361-task set (or 369 with Google Drive tasks). Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
computer-use
Grading
state-check
Tasks
369
Human reference
Humans completed 72.36% of tasks in the original 2024 study. The OSWorld-Verified report cites human performance as about 72%.
Contamination
All tasks and checkers are public. The maintainers have not published a contamination analysis. Some tasks depend on live websites that change over time.
Reuse
Apache 2.0 on the GitHub repository. Windows tasks need a licensed image. (open-apache)

Limits to keep in mind

  • Agents now score above the 72% human figure, and the maintainers released OSWorld 2.0 in June 2026 as the next version with 108 longer tasks. OSWorld-Verified is still maintained but is no longer the hardest test. Source
  • The original task set avoided long tasks, deep professional software, video, and real-time work, so it under-covers those parts of computer use. Source
  • Scores depend on the step budget. The same model can score much higher at 100 steps than at 15, so compare entries at the same budget. Source
  • Verified entries need the maintainers to run the agent on their own AWS platform, so the board lags model releases and lab-reported numbers (for example GPT-5.5 at 78.7%) are not on it. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemSuccess rateDateSource
claude-fable-5[1m]86%1 Aug 2026HKU · primary
Intelligence-Indeed Agent · frontier90.2%25 Jul 2026HKU · primary
Holo3-35B-A3B82.6%20 Apr 2026HKU · primary
agent s3 w/ GPT-5 bBoN (N=10)
harness: Agent S3
69.9%4 Oct 2025HKU · primary
CoACT-160.8%4 Aug 2025HKU · primary

Timeline

  • 25 Jul 2026 — First verified OSWorld score above 90%. Source
  • 26 Jun 2026 — OSWorld 2.0 succeeds OSWorld-Verified. Source
  • 28 Jul 2025 — OSWorld becomes OSWorld-Verified. Source

Where it sits in the atlas

Office and administrative supportDevOps, SRE, and IT operations (partial)Design, media, and writing (partial)

Computer useSpreadsheet work

Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: OSWorld-Verified by HKU, Salesforce AI Research, Carnegie Mellon University, and University of Waterloo (https://arxiv.org/abs/2404.07972).

Then, if you used this page:

Can Agents Work. "OSWorld-Verified: frontier results and sources." https://canagentswork.com/benchmarks/osworld-verified/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-osworld-verified,
  title        = {{OSWorld-Verified: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/osworld-verified/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}