Benchmarks / Tax Agent Bench
Tax Agent Bench
Built by Vals AI · released Sep 2026
Vals AI's test of whether an agent can do the research work of a corporate tax associate. Each task is one written US tax question, mostly federal corporate income tax, with the rest in partnership tax, state and local tax, transfer pricing, and employment tax. The agent researches with five tools (web search, authority lookup by citation, document fetch, retrieval, calculator) and submits one written answer with citations.
Frontier
77.6%
Accuracy
Claude Fable 5.1 · harness: Vals Tax Agent Bench harness
23 Sep 2026 · Source: Vals AI (benchmark maintainers)
Page data gives 77.642 (stderr 2.835) at about $13.18 per task and about 29 minutes per question, compute effort max. All-Pass (every rubric check passed) 49.22%. Claude Opus 5.5 scores 70.50%, sixth. Date is the page's "updated" field.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Accuracy on a private test set of 193 expert-written questions, graded by an LLM judge (GPT 5.4) against an expert rubric of 3 to 89 weighted checks per question (mean 23.6). About 40% of checks are must-pass; failing one zeroes the question. The rubric score is scaled by the share of cited sources that resolve to real authoritative documents (down to 70% of the score). Each task has a three-hour limit. A stricter All-Pass rate (every check passed) is reported as a secondary metric.
Accuracy: Weighted partial-credit rubric score, zeroed on any failed must-pass check and scaled by citation quality, averaged over the 193 private test questions. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, documents
- Grading
- llm-judge, rubric
- Tasks
- 391
- Human reference
- Tax professionals wrote and reviewed every question, gold answer, and rubric. Reference answers average 601 words. No human score is published.
- Contamination
- The test set is private. Vals reports scores only on the test set.
- Reuse
- 391 questions in total: 5 public samples, 193 validation questions available for license, and a 193-question test set that Vals says will stay private. No license stated for the public samples. (cite-only)
Limits to keep in mind
- Grading is by an LLM judge (GPT 5.4) against expert rubrics. The headline Accuracy gives partial credit; under the strict All-Pass metric no model reaches 50% (Claude Fable 5.1 49.22%). Source
- Standard errors are about 2.8 to 3.3 points, so the top three models (73% to 78%) overlap. Source
- Tools are fixed by the Vals harness, and the private set can only be run by Vals. Some document fetches fail on blocked or missing pages (3.8% of tool calls fail). Source
- The questions are research questions with a written answer. They do not cover preparing a return end to end or client work, so the score speaks to tax research, not all tax work. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Accuracy | All-Pass | Date | Source |
|---|---|---|---|---|
| GLM 5.3 harness: Vals Tax Agent Bench harness | 73.1% | 39.9% | 23 Sep 2026 | Vals AI · primary |
| Claude Opus 5 harness: Vals Tax Agent Bench harness | 75.1% | 45.6% | 23 Sep 2026 | Vals AI · primary |
| Claude Fable 5.1 · frontier harness: Vals Tax Agent Bench harness | 77.6% | 49.22% | 23 Sep 2026 | Vals AI · primary |
Timeline
- Sep 2026 — Vals AI launches Tax Agent Bench; Claude Fable 5.1 leads at 77.64% on corporate tax research. Source
Where it sits in the atlas
Work ladder: Direct evidence for Finance and accounting. On the work ladder it counts as 49.2% of tasks fully completed (All-Pass). How the ladder works
Work it measures (O*NET work activities): Advise others on legal or regulatory matters; Calculate financial data.
Go to the source
- Website www.vals.ai
- Full leaderboard www.vals.ai
Last checked 24 Sep 2026 against 1 primary source. See an error? Tell us.
How to cite
Credit the original work first: Tax Agent Bench by Vals AI.
Then, if you used this page:
Can Agents Work. "Tax Agent Bench: frontier results and sources." https://canagentswork.com/benchmarks/vals-tax-agent-bench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-vals-tax-agent-bench,
title = {{Tax Agent Bench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/vals-tax-agent-bench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}