Benchmarks / WorkArena

WorkArena

Built by ServiceNow · Alexandre Drouin, Maxime Gasse, Massimo Caccia, Issam H. Laradji, et al. · released 12 Mar 2024

Browser tasks on a live ServiceNow enterprise instance, the software many companies use for IT and HR requests. WorkArena-L1 has 33 task types (19,912 instances) that cover the main parts of the ServiceNow interface: forms, lists, knowledge base search, service catalog orders, and dashboards.

Frontier

90.3%

WorkArena-L1 success rate

IpaziaHPA-Gemini-3-flash-preview · harness: IpaziaHPA · agent: IpaziaHPA (Ipazia S.p.a.)

7 Mar 2026 · Source: Ipazia S.p.a. (submitted to the ServiceNow BrowserGym leaderboard) (third party)

Top WorkArena-L1 entry on the leaderboard as of 2026-09-23. Standard error 1.6. Uses the accessibility tree, no screenshots. Max steps raised from 15 to 30, so the file marks it as not following the standard protocol. Self-submitted; not re-run by ServiceNow.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at ServiceNow

WorkArena: WorkArena-L1 success rate over time, 4 recorded results. 0%20%40%60%80%100%Oct 2024Jan 2025Apr 2025Jul 2025Oct 2025Jan 2026Apr 2026 GenericAgent-GPT-4o: 45.5% (23 Oct 2024) GenericAgent-Claude-4-Sonnet: 63.3% (7 Aug 2025) GenericAgent-GPT-5: 79.1% (7 Aug 2025) IpaziaHPA-Gemini-3-flash-preview: 90.3% (7 Mar 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/workarena/"><img src="https://canagentswork.com/og/benchmarks-workarena.png" width="600" height="315" alt="WorkArena: the best result is 90.3% (IpaziaHPA-Gemini-3-flash-preview, 7 Mar 2026)." loading="lazy"></a>

Markdown:

[![WorkArena: the best result is 90.3% (IpaziaHPA-Gemini-3-flash-preview, 7 Mar 2026).](https://canagentswork.com/og/benchmarks-workarena.png)](https://canagentswork.com/benchmarks/workarena/)

What it measures

Whether a web agent completes routine enterprise-software actions from a text instruction. Each task samples a fresh configuration, and a checker validates the result in the ServiceNow instance. WorkArena++ (L2 and L3, 682 tasks) composes these atomic actions into longer workflows that need planning and reasoning, but the public leaderboard currently lists results only for L1.

WorkArena-L1 success rate: Share of sampled WorkArena-L1 task instances completed, as reported on the BrowserGym leaderboard with standard error. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
browser, live-system
Grading
state-check
Tasks
33
Contamination
Task templates are public, but each run samples new instances. The maintainers have not published a contamination analysis.
Reuse
Apache 2.0 for the benchmark code. Running tasks needs access to a ServiceNow developer instance. (open-apache)

Limits to keep in mind

  • Leaderboard entries are self-submitted result files. The top entry raised the step budget from 15 to 30 and is marked as not following the standard evaluation protocol. Source
  • L1 tasks are short atomic actions. The harder WorkArena++ (L2, L3) tasks have no entries on the public leaderboard, so the board understates how far agents are from full workflows. Source
  • Few agents have been submitted (about 20 folders), and several have no WorkArena result at all. Source
  • All tasks run on one vendor's platform, so results say little about other enterprise software. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemWorkArena-L1 success rateDateSource
IpaziaHPA-Gemini-3-flash-preview · frontier
harness: IpaziaHPA
90.3%7 Mar 2026Ipazia S.p.a. (submitted to the ServiceNow BrowserGym leaderboard) · third-party
GenericAgent-Claude-4-Sonnet
harness: BrowserGym GenericAgent
63.3%7 Aug 2025ServiceNow · primary
GenericAgent-GPT-5
harness: BrowserGym GenericAgent
79.1%7 Aug 2025ServiceNow · primary
GenericAgent-GPT-4o
harness: BrowserGym GenericAgent
45.5%23 Oct 2024ServiceNow · primary

Timeline

  • 12 Mar 2024 — ServiceNow Research releases WorkArena and BrowserGym. Source

Where it sits in the atlas

Office and administrative supportDevOps, SRE, and IT operations (partial)

Enterprise workflowsComputer use

Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: WorkArena by ServiceNow (https://arxiv.org/abs/2403.07718).

Then, if you used this page:

Can Agents Work. "WorkArena: frontier results and sources." https://canagentswork.com/benchmarks/workarena/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-workarena,
  title        = {{WorkArena: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/workarena/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}