Benchmarks / EnterpriseOps-Gym

EnterpriseOps-Gym

Built by ServiceNow, Mila and Université de Montréal · Shiva Krishna Reddy Malay, Shravan Nayak, Jishnu Sethumadhavan Nair, Aman Tiwari, et al. · released 13 Mar 2026

A containerized enterprise simulation from ServiceNow AI Research. An agent works through MCP tools against live databases in eight domains: Calendar, Customer Service Management, Drive, Email, HR, IT Service Management, Teams, and Hybrid tasks that span several systems. Expert-written SQL checks grade the final state, not the actions.

Frontier

45.9%

Task success rate

Claude Opus 4.6 · harness: ReAct (EnterpriseOps-Gym)

16 Mar 2026 · Source: ServiceNow (benchmark maintainers)

README leaderboard, full benchmark, oracle mode. By domain: Teams 52.0, CSM 45.1, Email 57.7, ITSM 33.3, Calendar 43.3, HR 45.1, Drive 57.1, Hybrid 34.0. The table is undated; it appears in the repository's first commit (2026-03-16). Artificial Analysis's separate run of the public split scores Claude Fable 5 at 51.1%, which is not comparable.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at ServiceNow

EnterpriseOps-Gym: Task success rate over time, 2 recorded results. 0%20%40%60%80%100%Mar 2026Apr 2026 Claude Opus 4.5: 39.4% (16 Mar 2026) Claude Opus 4.6: 45.9% (16 Mar 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/enterpriseops-gym/"><img src="https://canagentswork.com/og/benchmarks-enterpriseops-gym.png" width="600" height="315" alt="EnterpriseOps-Gym: the best result is 45.9% (Claude Opus 4.6, 16 Mar 2026)." loading="lazy"></a>

Markdown:

[![EnterpriseOps-Gym: the best result is 45.9% (Claude Opus 4.6, 16 Mar 2026).](https://canagentswork.com/og/benchmarks-enterpriseops-gym.png)](https://canagentswork.com/benchmarks/enterpriseops-gym/)

What it measures

Task success rate: a task passes only when every verification condition on the final database state holds (goal reached, data intact, policy followed, no side effects). 1,150 expert-curated tasks, 512 tools, and 164 tables; expert trajectories average 9.15 steps (up to 34). The headline setting is oracle tool mode, where the agent gets the tools the task needs. 30 tasks are infeasible and test whether the agent refuses cleanly.

Task success rate: Share of tasks where all SQL verifiers pass, in oracle tool mode on the full benchmark. The README leaderboard reports the mean of the eight domain rates. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, simulated-workplace
Grading
state-check
Tasks
1,150
Contamination
60% of the tasks are public. The README leaderboard reports the full benchmark and, separately, the public split.
Reuse
Apache 2.0 on the Hugging Face dataset card and the repository. The public split holds 60% of the tasks (649 rows in the oracle config); the rest is private. (open-apache)

Limits to keep in mind

  • Three numbers for the best model appear at the source: the README intro says 34.1%, the paper says 37.4% (Claude Opus 4.5), and the README leaderboard says 45.9% (Claude Opus 4.6). The paper averages over tasks; the README board averages the eight domain rates. See the conflict file. Source
  • The maintainers' site now names the Artificial Analysis board the official leaderboard. That board runs the public dataset in AA's own Stirrup harness with 3 repeats, and AA says its numbers are not directly comparable with the paper. Its top score is 51.1% (Claude Fable 5, Opus 4.8 fallback), undated. Source
  • Oracle tool mode gives the agent the right tools. The paper reports that adding 5 to 15 distractor tools changed Claude Sonnet 4.5's score by about one point, so the effect of tool retrieval is small but the headline is still the easiest setting. Source
  • Agents refuse infeasible tasks cleanly only 53.9% of the time at best, so side effects on policy-violating requests are common. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemTask success rateDateSource
Claude Opus 4.5
harness: ReAct (EnterpriseOps-Gym)
39.4%16 Mar 2026ServiceNow · primary
Claude Opus 4.6 · frontier
harness: ReAct (EnterpriseOps-Gym)
45.9%16 Mar 2026ServiceNow · primary

Where sources disagree

Three "best model" numbers appear at the source. The paper says Claude Opus 4.5 reaches 37.4%. The README leaderboard shows 39.4% for the same model, with the same eight domain scores. The README intro says "best model achieves only 34.1%", which equals the board's row for Claude Sonnet 4.5, while the board's top row is Claude Opus 4.6 at 45.9%.

  • 37.4arxiv.org (primary) · Paper abstract and Table 2 (2026-03-13), Claude Opus 4.5, oracle tool mode.
  • 39.4github.com (primary) · README leaderboard "Avg" for Claude Opus 4.5, full benchmark, oracle mode.
  • 34.1github.com (primary) · README intro line "Best model achieves only 34.1% success rate". Matches the board row for Claude Sonnet 4.5.

We show 39.4. Status: resolved.

Timeline

  • 13 Mar 2026 — ServiceNow releases EnterpriseOps-Gym; best model passes 37.4% of enterprise tasks. Source

Where it sits in the atlas

Work ladder: Direct evidence for Office and administrative support. On the work ladder it counts as 45.9% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Perform administrative or clerical activities; Schedule appointments; Maintain operational records; Communicate with others about operational plans or activities; Respond to customer problems or inquiries.

Office and administrative supportCustomer support (partial)DevOps, SRE, and IT operations (partial)Management and business operations (partial)

Enterprise workflowsCustomer serviceIncident response

Last checked 24 Sep 2026 against 8 primary sources. See an error? Tell us.

How to cite

Credit the original work first: EnterpriseOps-Gym by ServiceNow, Mila, and Université de Montréal (https://arxiv.org/abs/2603.13594).

Then, if you used this page:

Can Agents Work. "EnterpriseOps-Gym: frontier results and sources." https://canagentswork.com/benchmarks/enterpriseops-gym/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-enterpriseops-gym,
  title        = {{EnterpriseOps-Gym: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/enterpriseops-gym/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}