Benchmarks / APEX-Accounting

APEX-Accounting

Built by Mercor and Ramp · Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, et al. · released 31 Jul 2026

Month-end close work for staff accountants and bookkeepers: reconcile accounts, post entries, accrue expenses, explain variances, and build close schedules. Each task runs inside a synthetic company frozen at month-end, with an accounting system, spreadsheets, PDFs, and other files.

Frontier

61.8%

Mean Criteria@3

Claude Opus 5.5 (Max) · harness: Mercor Loop harness

24 Sep 2026 · Source: Mercor (benchmark maintainers)

Rank 1 on the board seen 2026-09-24; model release date shown as 2026-09-22. Mean Score 61.8% plus or minus 4.0; Pass@1 14.8% plus or minus 4.9; 637 samples. Fable 5.1 (Max, 61.0%) and Opus 5.5 (Medium, 59.9%) are within the confidence interval.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Mercor

APEX-Accounting: Mean Criteria@3 over time, 3 recorded results. 0%20%40%60%80%100%Aug 2026Sep 2026Oct 2026 Claude Fable 5 (Max): 56.4% (30 Jul 2026) Claude Fable 5.1 (Max): 61% (24 Sep 2026) Claude Opus 5.5 (Max): 61.8% (24 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/apex-accounting/"><img src="https://canagentswork.com/og/benchmarks-apex-accounting.png" width="600" height="315" alt="APEX-Accounting: the best result is 61.8% (Claude Opus 5.5 (Max), 24 Sep 2026)." loading="lazy"></a>

Markdown:

[![APEX-Accounting: the best result is 61.8% (Claude Opus 5.5 (Max), 24 Sep 2026).](https://canagentswork.com/og/benchmarks-apex-accounting.png)](https://canagentswork.com/benchmarks/apex-accounting/)

What it measures

Whether an agent can do close-cycle bookkeeping the way an accountant would. 42 accounting experts (median 11 years; more than half with Big Four experience) built 10 private company worlds and 160 tasks in four categories: Reconciliation (61), Data Entry (26), Variance Analysis (28), and Schedules and Accruals (45). The agent answers in a console message, and an LLM judge scores it against a rubric of binary, outcome-based criteria (13.7 on average). Mercor runs every model eight times per task with 91 tools, including a QuickBooks-style accounting interface, and a cap of 500 steps and 5 million tokens.

Mean Criteria@3: Share of rubric criteria met, averaged over 3 of the 8 runs per task. Mercor calls it Mean Score on the leaderboard. Partial credit: a task counts even when only some criteria pass. Mercor also publishes Pass@1 (all criteria met on one attempt), Pass@8, and Pass^8. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
simulated-workplace, documents, api-tools
Grading
rubric, llm-judge
Tasks
160
Human reference
No human score. Task authors estimate real-world completion time for each task; the 10 public dev tasks carry estimates of 0.75 to 4.0 hours.
Contamination
The scored set is held out. Mercor says every world document is novel and screened against public sources. The dev world is the easiest of the 11 qualifying worlds, so dev-set scores run higher.
Reuse
The 160 scored tasks are private and never released. A public dev set (1 world, 10 tasks) is on Hugging Face under CC BY 4.0. The tool layer (accounting software interface) is not released. (cite-only)

Limits to keep in mind

  • Grading uses an LLM judge (DeepSeek-v4-Flash) against expert rubrics. Mercor reports 97.1% judge accuracy against majority-vote human labels on 1,687 criteria. Source
  • The headline metric gives partial credit. Whole-task completion is much lower: Pass@1 is 14.8% for the top model, no model exceeds 2.6% Pass^8, and 58% of tasks were never fully solved by any model in the paper's runs. Source
  • Covers close-cycle bookkeeping only. Tax, audit, consolidation, multi-entity and multi-currency work, external reporting, and tasks that need a human in the loop were left out. Source
  • Confidence intervals are about plus or minus 4 points, so the top four entries overlap. The page's stat bar ("Highest score 61.0%") lags the table (61.8%). Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemMean Criteria@3Pass@1DateSource
Claude Fable 5.1 (Max)
harness: Mercor Loop harness
61%11.7%24 Sep 2026Mercor · primary
Claude Opus 5.5 (Max) · frontier
harness: Mercor Loop harness
61.8%14.8%24 Sep 2026Mercor · primary
Claude Fable 5 (Max)
harness: Mercor Loop harness
56.4%9.5%30 Jul 2026Mercor · primary

Timeline

  • Sep 2026 — Claude Opus 5.5 reaches 61.8% on APEX-Accounting. Source
  • 31 Jul 2026 — Mercor and Ramp release APEX-Accounting; Claude Fable 5 meets 56.4% of criteria. Source

Where it sits in the atlas

Work ladder: Direct evidence for Finance and accounting. On the work ladder it counts as 14.8% of tasks fully completed (Pass@1). How the ladder works

Work it measures (O*NET work activities): Examine financial activities, operations, or systems; Evaluate the quality or accuracy of data; Execute financial transactions; Analyze business or financial data; Calculate financial data; Prepare financial documents, reports, or budgets.

Finance and accountingOffice and administrative support (partial)

Bookkeeping and month-end closeSpreadsheet work

Last checked 24 Sep 2026 against 5 primary sources. See an error? Tell us.

How to cite

Credit the original work first: APEX-Accounting by Mercor and Ramp (https://arxiv.org/abs/2607.27189).

Then, if you used this page:

Can Agents Work. "APEX-Accounting: frontier results and sources." https://canagentswork.com/benchmarks/apex-accounting/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-apex-accounting,
  title        = {{APEX-Accounting: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/apex-accounting/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}