Benchmarks / APEX-Accounting
APEX-Accounting
Built by Mercor and Ramp · Julien Benchek, Austin Bennett, Jasmin Kern, Ryan Stevens, et al. · released 31 Jul 2026
Month-end close work for staff accountants and bookkeepers: reconcile accounts, post entries, accrue expenses, explain variances, and build close schedules. Each task runs inside a synthetic company frozen at month-end, with an accounting system, spreadsheets, PDFs, and other files.
Frontier
61.8%
Mean Criteria@3
Claude Opus 5.5 (Max) · harness: Mercor Loop harness
24 Sep 2026 · Source: Mercor (benchmark maintainers)
Rank 1 on the board seen 2026-09-24; model release date shown as 2026-09-22. Mean Score 61.8% plus or minus 4.0; Pass@1 14.8% plus or minus 4.9; 637 samples. Fable 5.1 (Max, 61.0%) and Opus 5.5 (Medium, 59.9%) are within the confidence interval.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether an agent can do close-cycle bookkeeping the way an accountant would. 42 accounting experts (median 11 years; more than half with Big Four experience) built 10 private company worlds and 160 tasks in four categories: Reconciliation (61), Data Entry (26), Variance Analysis (28), and Schedules and Accruals (45). The agent answers in a console message, and an LLM judge scores it against a rubric of binary, outcome-based criteria (13.7 on average). Mercor runs every model eight times per task with 91 tools, including a QuickBooks-style accounting interface, and a cap of 500 steps and 5 million tokens.
Mean Criteria@3: Share of rubric criteria met, averaged over 3 of the 8 runs per task. Mercor calls it Mean Score on the leaderboard. Partial credit: a task counts even when only some criteria pass. Mercor also publishes Pass@1 (all criteria met on one attempt), Pass@8, and Pass^8. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- simulated-workplace, documents, api-tools
- Grading
- rubric, llm-judge
- Tasks
- 160
- Human reference
- No human score. Task authors estimate real-world completion time for each task; the 10 public dev tasks carry estimates of 0.75 to 4.0 hours.
- Contamination
- The scored set is held out. Mercor says every world document is novel and screened against public sources. The dev world is the easiest of the 11 qualifying worlds, so dev-set scores run higher.
- Reuse
- The 160 scored tasks are private and never released. A public dev set (1 world, 10 tasks) is on Hugging Face under CC BY 4.0. The tool layer (accounting software interface) is not released. (cite-only)
Limits to keep in mind
- Grading uses an LLM judge (DeepSeek-v4-Flash) against expert rubrics. Mercor reports 97.1% judge accuracy against majority-vote human labels on 1,687 criteria. Source
- The headline metric gives partial credit. Whole-task completion is much lower: Pass@1 is 14.8% for the top model, no model exceeds 2.6% Pass^8, and 58% of tasks were never fully solved by any model in the paper's runs. Source
- Covers close-cycle bookkeeping only. Tax, audit, consolidation, multi-entity and multi-currency work, external reporting, and tasks that need a human in the loop were left out. Source
- Confidence intervals are about plus or minus 4 points, so the top four entries overlap. The page's stat bar ("Highest score 61.0%") lags the table (61.8%). Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Mean Criteria@3 | Pass@1 | Date | Source |
|---|---|---|---|---|
| Claude Fable 5.1 (Max) harness: Mercor Loop harness | 61% | 11.7% | 24 Sep 2026 | Mercor · primary |
| Claude Opus 5.5 (Max) · frontier harness: Mercor Loop harness | 61.8% | 14.8% | 24 Sep 2026 | Mercor · primary |
| Claude Fable 5 (Max) harness: Mercor Loop harness | 56.4% | 9.5% | 30 Jul 2026 | Mercor · primary |
Timeline
Where it sits in the atlas
Work ladder: Direct evidence for Finance and accounting. On the work ladder it counts as 14.8% of tasks fully completed (Pass@1). How the ladder works
Work it measures (O*NET work activities): Examine financial activities, operations, or systems; Evaluate the quality or accuracy of data; Execute financial transactions; Analyze business or financial data; Calculate financial data; Prepare financial documents, reports, or budgets.
Finance and accountingOffice and administrative support (partial)
Go to the source
- Website www.mercor.com
- Paper arxiv.org
- Full leaderboard www.mercor.com
- Code github.com
- Announcement www.mercor.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: APEX-Accounting by Mercor and Ramp (https://arxiv.org/abs/2607.27189).
Then, if you used this page:
Can Agents Work. "APEX-Accounting: frontier results and sources." https://canagentswork.com/benchmarks/apex-accounting/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-apex-accounting,
title = {{APEX-Accounting: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/apex-accounting/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}