Benchmarks / SpreadsheetBench
SpreadsheetBench
Built by Renmin University of China · Zeyao Ma, Bohan Zhang, Jing Zhang, Jifan Yu, et al. · released 21 Jun 2024
912 spreadsheet manipulation questions taken from real Excel forum posts, each paired with the user's actual workbook. Workbooks have multiple tables, odd layouts, and non-text elements. A solution is checked like an online judge: it must work on several test-case spreadsheets with different values.
Frontier
83.1%
Overall pass@1 (V1, 912 questions)
Qingqiu Agent
23 Jun 2026 · Source: Renmin University of China (benchmark maintainers)
Verified entry at the top of the V1 Full (912) table in the leaderboard data file on 2026-09-23. JT AlphaData (CMCC JIUTIAN) is second at 77.85, dated 2026-08-15.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether a system can produce the exact cell values a user asked for in a real spreadsheet, across cell-level and sheet-level edits. The headline is overall pass@1 on the full 912-question V1 set. The site also keeps a 400-question expert-verified V1 subset (released December 2025, best 99.25%) and SpreadsheetBench 2 (321 workflow tasks on financial modeling, debugging, and charts; best 50.86%).
Overall pass@1 (V1, 912 questions): Share of the 912 questions answered correctly across all 2,729 test-case spreadsheets, as listed on the official leaderboard. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- documents, cli
- Grading
- state-check
- Tasks
- 912
- Human reference
- Four Excel experts on a 50-instruction subset (3 test cases each) scored 71.33% under the soft restriction and 62.00% under the hard restriction. GPT-4o scored 18.35% and 15.02% on the same measures in the paper.
- Contamination
- Questions were rewritten from forum posts by GPT-4 and annotators, and spreadsheet values were changed to build test cases, so exact forum solutions do not apply directly. The full data set and answers are public.
- Reuse
- CC BY-SA 4.0, stated in the paper's maintenance plan and the repo README. (open-other)
Limits to keep in mind
- Some leaderboard rows are marked unverified. They come from outside evaluations by OpenAI and Microsoft, not from the maintainers' own runs. Source
- The site's hard-coded "Top Score" box showed 70.48% on 2026-09-23 while its leaderboard data file listed 83.11%. See data/conflicts/spreadsheetbench-v1-top-score.yaml. Source
- Most top entries are commercial spreadsheet products (WPS, Google Sheets, Univer) that do not disclose the model or scaffold. Source
- The human baseline covers only 50 of 912 instructions, so it is not directly comparable with leaderboard scores. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Overall pass@1 (V1, 912 questions) | Date | Source |
|---|---|---|---|
| Qingqiu Agent · frontier | 83.1% | 23 Jun 2026 | Renmin University of China · primary |
| WPS AI (Seed 2.0) | 73.5% | 16 Jun 2026 | Renmin University of China · primary |
| Gemini in Google Sheets | 70.5% | 10 Mar 2026 | Renmin University of China · primary |
| Shortcut.ai | 59.3% | 16 Oct 2025 | Renmin University of China · primary |
| ChatGPT Agent w/ .xlsx | 45.5% | 17 Jul 2025 | Renmin University of China · primary |
Where sources disagree
The SpreadsheetBench site shows two different top scores for the 912-question V1 set. The static "Top Score (OVERALL)" box reads 70.48%. The leaderboard data file that the page loads lists Qingqiu Agent at 83.11% (verified, 2026-06-23). The box looks stale. We show the leaderboard value.
- 83.11 — spreadsheetbench.github.io (primary) · Leaderboard data file, seen 2026-09-23. Entry is marked verified.
- 70.48 — spreadsheetbench.github.io (primary) · Static statistics box on the V1 overview page, seen 2026-09-23. Matches the Gemini in Google Sheets row dated 2026-03-10.
We show 83.11. Status: open.
Timeline
Where it sits in the atlas
Office and administrative supportData and analytics (partial)Finance and accounting (partial)
Go to the source
- Website spreadsheetbench.github.io
- Paper arxiv.org
- Full leaderboard spreadsheetbench.github.io
- Code github.com
- Dataset huggingface.co
Last checked 23 Sep 2026 against 7 primary sources. See an error? Tell us.
How to cite
Credit the original work first: SpreadsheetBench by Renmin University of China (https://arxiv.org/abs/2406.14991).
Then, if you used this page:
Can Agents Work. "SpreadsheetBench: frontier results and sources." https://canagentswork.com/benchmarks/spreadsheetbench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-spreadsheetbench,
title = {{SpreadsheetBench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/spreadsheetbench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}