Benchmarks / SpreadsheetBench 2

SpreadsheetBench 2

Built by Renmin University of China, Aptura, AfterQuery and Shortcut · Jian Zhu, Yuzheng Zhang, Zeyao Ma, Bohan Zhang, et al. · released 29 Jun 2026

321 business spreadsheet tasks built from real financial reports and corporate filings and checked by domain experts: complete a template, build out a financial model, find and fix errors in a workbook, or draw charts. Workbooks average 11.8 sheets, and a task needs 593.5 cell changes on average. The successor to SpreadsheetBench from the same Renmin University group.

Frontier

50.9%

Overall accuracy

WPS AI

29 Aug 2026 · Source: Renmin University of China (benchmark maintainers)

Rank 1 on the V2 full leaderboard. Commercial product by Kingsoft Office; model and scaffold not disclosed. Marked verified by the maintainers. Subscores: template 76.29, financial model 56.0, debugging 16.0, visualization 71.94. The site's static top-score box still shows 34.89%.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Renmin University of China

SpreadsheetBench 2: Overall accuracy over time, 3 recorded results. 0%20%40%60%80%100%Jul 2026Aug 2026Sep 2026 Claude Opus 4.6 (SWE-agent): 34.9% (29 Jun 2026) arito: 45.5% (4 Aug 2026) WPS AI: 50.9% (29 Aug 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/spreadsheetbench-2/"><img src="https://canagentswork.com/og/benchmarks-spreadsheetbench-2.png" width="600" height="315" alt="SpreadsheetBench 2: the best result is 50.9% (WPS AI, 29 Aug 2026)." loading="lazy"></a>

Markdown:

[![SpreadsheetBench 2: the best result is 50.9% (WPS AI, 29 Aug 2026).](https://canagentswork.com/og/benchmarks-spreadsheetbench-2.png)](https://canagentswork.com/benchmarks/spreadsheetbench-2/)

What it measures

Share of tasks solved. For template, financial-model, and debugging tasks the output workbook must match the golden file in every target cell with no other cell changed. For visualization tasks a vision-language model checks the chart against a checklist of expert assertions and reports the share passed. The overall score aggregates the four categories (template, financial model, debugging, visualization). Rows marked verified were run or checked by the maintainers; scaffolded models use a SWE-agent-based scaffold.

Overall accuracy: Share of the 321 tasks solved, as listed in the site's V2 full leaderboard, with subscores for template, financial model, debugging, and visualization. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
documents, cli
Grading
state-check, llm-judge
Tasks
321
Contamination
Tasks and golden files are public on Hugging Face. The maintainers have not published a contamination analysis. Results are submitted by email with logs and outputs and checked by the maintainers, who process submissions weekly.
Reuse
MIT on the Hugging Face dataset card (KAKA22/SpreadsheetBench-v2). The code repository has no license file. (open-mit)

Limits to keep in mind

  • The site's static "Top Score (OVERALL)" box for V2 showed 34.89% on 2026-09-24 while the leaderboard data file listed WPS AI at 50.86%. See data/conflicts/spreadsheetbench-2-top-score.yaml. Source
  • The two leading rows are commercial spreadsheet products (WPS AI, arito) that do not disclose the model or scaffold. The best scaffolded model is Claude Opus 4.6 with SWE-agent at 34.89%. Source
  • Visualization tasks are scored by a vision-language model against assertion checklists, and the site gives two readings of that subscore (average assertion pass rate, or a task counted correct above 70), so the overall score mixes exact-match completion with rubric credit. Source
  • Debugging is the weak spot: the best debugging subscore is 28.0% (arito) and the leader WPS AI scores 16.0% there. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemOverall accuracyDateSource
WPS AI · frontier50.9%29 Aug 2026Renmin University of China · primary
arito45.5%4 Aug 2026Renmin University of China · primary
Claude Opus 4.6 (SWE-agent)
harness: SWE-agent
34.9%29 Jun 2026Renmin University of China · primary

Where sources disagree

The SpreadsheetBench site shows two different top scores for V2. The static "Top Score (OVERALL)" box on the V2 overview reads 34.89%, the best scaffolded model in the paper. The leaderboard data file that the page loads lists WPS AI at 50.86% (verified, 2026-08-29). The box looks stale. We show the leaderboard value.

  • 50.86spreadsheetbench.github.io (primary) · Leaderboard data file, seen 2026-09-24. Entry is marked verified.
  • 34.89spreadsheetbench.github.io (primary) · Static statistics box on the V2 overview page, seen 2026-09-24. Matches the Claude Opus 4.6 (SWE-agent) row and the paper abstract.

We show 50.86. Status: open.

Timeline

  • 29 Aug 2026 — WPS AI reaches 50.86% on SpreadsheetBench 2. Source
  • 29 Jun 2026 — SpreadsheetBench 2 launches with 321 business spreadsheet workflows; best model 34.89%. Source

Where it sits in the atlas

Work ladder: Direct evidence for Finance and accounting. On the work ladder it counts as 50.9% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Analyze business or financial data; Develop financial or business plans; Create visual designs or displays.

Finance and accountingOffice and administrative support (partial)Data and analytics (partial)

Financial analysisSpreadsheet work

Last checked 24 Sep 2026 against 5 primary sources. See an error? Tell us.

How to cite

Credit the original work first: SpreadsheetBench 2 by Renmin University of China, Aptura, AfterQuery, and Shortcut (https://arxiv.org/abs/2606.29955).

Then, if you used this page:

Can Agents Work. "SpreadsheetBench 2: frontier results and sources." https://canagentswork.com/benchmarks/spreadsheetbench-2/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-spreadsheetbench-2,
  title        = {{SpreadsheetBench 2: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/spreadsheetbench-2/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}