Benchmarks / APEX-Agents

APEX-Agents

Built by Mercor, Box and Harvey · Bertie Vidgen, Austin Mann, Abby Fennelly, John Wright Stanly, et al. · released 21 Jan 2026

Long, multi-application work tasks written by investment banking analysts, management consultants, and corporate lawyers. Each task sits inside a simulated project "world" with files, spreadsheets, and chat threads. An agent must find the right information and produce a client-ready output.

Frontier

73.5%

Pass@1

Claude Opus 5.5 (max) · harness: Mercor Loop agent

23 Sep 2026 · Source: Mercor (benchmark maintainers)

Rank 1 on the APEX-Agents 1.1 board, seen 2026-09-23. Pass@1 73.5% plus or minus 4.9; Mean Score 81.3% plus or minus 4.0; 950 samples. Model release date shown as 2026-09-22. Graded by an LLM judge against expert rubrics.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Mercor

APEX-Agents: Pass@1 over time, 4 recorded results. 0%20%40%60%80%100%Jan 2026Apr 2026Jul 2026Oct 2026 Gemini 3 Flash (Thinking=High): 24% (Jan 2026) Claude Fable 5.1 (max): 68.6% (8 Sep 2026) GPT-5.5 (xhigh): 55.1% (23 Sep 2026) Claude Opus 5.5 (max): 73.5% (23 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/apex-agents/"><img src="https://canagentswork.com/og/benchmarks-apex-agents.png" width="600" height="315" alt="APEX-Agents: the best result is 73.5% (Claude Opus 5.5 (max), 23 Sep 2026)." loading="lazy"></a>

Markdown:

[![APEX-Agents: the best result is 73.5% (Claude Opus 5.5 (max), 23 Sep 2026).](https://canagentswork.com/og/benchmarks-apex-agents.png)](https://canagentswork.com/benchmarks/apex-agents/)

What it measures

Whether an agent completes a professional task end to end inside a realistic workspace. Experts from top firms built 31 worlds (for example a week-long consulting project for a fictional oil and gas company) and wrote 240 tasks (80 per job) with 1 to 10 pass or fail criteria each. Version 1.1 (September 2026) tightened task specifications, added a judge that gives zero credit for hedged multiple answers, and fixed tool reliability.

Pass@1: Share of tasks where the agent passes every rubric criterion on a single attempt. Mercor also reports Mean Score, the average share of criteria passed per task. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
simulated-workplace, documents, api-tools
Grading
rubric, llm-judge
Tasks
240
Contamination
Mercor says the full task set stays private so that models cannot be trained on it, while an open subset is public. The exact split between public and private tasks is not stated on the leaderboard page.
Reuse
CC BY 4.0 on the Hugging Face dataset. Mercor says the full task set used for the leaderboard stays private. (open-cc-by)

Limits to keep in mind

  • Grading uses an LLM judge (DeepSeek-V4-Flash-0731 in v1.1) against expert rubrics, not human review of each output. Mercor reports the judge's false negative rate rose from 5.3% to 8.0% in v1.1. Source
  • Version 1.1 changed tasks (480 to 240), grading, and prompts, so scores before September 2026 are not comparable with the current board. Source
  • Confidence intervals are about plus or minus 5 percentage points, so the top five models overlap. Source
  • Only three jobs are covered, and all worlds run in a Google Workspace style environment with Mercor's own agent loop. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemPass@1DateSource
GPT-5.5 (xhigh)
harness: Mercor Loop agent
55.1%23 Sep 2026Mercor · primary
Claude Opus 5.5 (max) · frontier
harness: Mercor Loop agent
73.5%23 Sep 2026Mercor · primary
Claude Fable 5.1 (max)
harness: Mercor Loop agent
68.6%8 Sep 2026Mercor · primary
Gemini 3 Flash (Thinking=High)24%Jan 2026Mercor · primary

Timeline

  • 8 Sep 2026 — APEX-Agents 1.1 stops rewarding hedged answers. Source
  • 21 Jan 2026 — Mercor releases APEX-Agents. Source

Where it sits in the atlas

Finance and accountingManagement and business operations (partial)Legal (partial)

Financial analysisConsulting workLegal workProfessional deliverables

Last checked 23 Sep 2026 against 4 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: APEX-Agents by Mercor, Box, and Harvey (https://arxiv.org/abs/2601.14242).

Then, if you used this page:

Can Agents Work. "APEX-Agents: frontier results and sources." https://canagentswork.com/benchmarks/apex-agents/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-apex-agents,
  title        = {{APEX-Agents: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/apex-agents/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}