Benchmarks / Tax Agent Bench

Tax Agent Bench

Built by Vals AI · released Sep 2026

Vals AI's test of whether an agent can do the research work of a corporate tax associate. Each task is one written US tax question, mostly federal corporate income tax, with the rest in partnership tax, state and local tax, transfer pricing, and employment tax. The agent researches with five tools (web search, authority lookup by citation, document fetch, retrieval, calculator) and submits one written answer with citations.

Frontier

77.6%

Accuracy

Claude Fable 5.1 · harness: Vals Tax Agent Bench harness

23 Sep 2026 · Source: Vals AI (benchmark maintainers)

Page data gives 77.642 (stderr 2.835) at about $13.18 per task and about 29 minutes per question, compute effort max. All-Pass (every rubric check passed) 49.22%. Claude Opus 5.5 scores 70.50%, sixth. Date is the page's "updated" field.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Vals AI

Tax Agent Bench: Accuracy over time, 3 recorded results. 0%20%40%60%80%100%Sep 2026Sep 2026Sep 2026Sep 2026Oct 2026Oct 2026 GLM 5.3: 73.1% (23 Sep 2026) Claude Opus 5: 75.1% (23 Sep 2026) Claude Fable 5.1: 77.6% (23 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/vals-tax-agent-bench/"><img src="https://canagentswork.com/og/benchmarks-vals-tax-agent-bench.png" width="600" height="315" alt="Tax Agent Bench: the best result is 77.6% (Claude Fable 5.1, 23 Sep 2026)." loading="lazy"></a>

Markdown:

[![Tax Agent Bench: the best result is 77.6% (Claude Fable 5.1, 23 Sep 2026).](https://canagentswork.com/og/benchmarks-vals-tax-agent-bench.png)](https://canagentswork.com/benchmarks/vals-tax-agent-bench/)

What it measures

Accuracy on a private test set of 193 expert-written questions, graded by an LLM judge (GPT 5.4) against an expert rubric of 3 to 89 weighted checks per question (mean 23.6). About 40% of checks are must-pass; failing one zeroes the question. The rubric score is scaled by the share of cited sources that resolve to real authoritative documents (down to 70% of the score). Each task has a three-hour limit. A stricter All-Pass rate (every check passed) is reported as a secondary metric.

Accuracy: Weighted partial-credit rubric score, zeroed on any failed must-pass check and scaled by citation quality, averaged over the 193 private test questions. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, documents
Grading
llm-judge, rubric
Tasks
391
Human reference
Tax professionals wrote and reviewed every question, gold answer, and rubric. Reference answers average 601 words. No human score is published.
Contamination
The test set is private. Vals reports scores only on the test set.
Reuse
391 questions in total: 5 public samples, 193 validation questions available for license, and a 193-question test set that Vals says will stay private. No license stated for the public samples. (cite-only)

Limits to keep in mind

  • Grading is by an LLM judge (GPT 5.4) against expert rubrics. The headline Accuracy gives partial credit; under the strict All-Pass metric no model reaches 50% (Claude Fable 5.1 49.22%). Source
  • Standard errors are about 2.8 to 3.3 points, so the top three models (73% to 78%) overlap. Source
  • Tools are fixed by the Vals harness, and the private set can only be run by Vals. Some document fetches fail on blocked or missing pages (3.8% of tool calls fail). Source
  • The questions are research questions with a written answer. They do not cover preparing a return end to end or client work, so the score speaks to tax research, not all tax work. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAccuracyAll-PassDateSource
GLM 5.3
harness: Vals Tax Agent Bench harness
73.1%39.9%23 Sep 2026Vals AI · primary
Claude Opus 5
harness: Vals Tax Agent Bench harness
75.1%45.6%23 Sep 2026Vals AI · primary
Claude Fable 5.1 · frontier
harness: Vals Tax Agent Bench harness
77.6%49.22%23 Sep 2026Vals AI · primary

Timeline

  • Sep 2026 — Vals AI launches Tax Agent Bench; Claude Fable 5.1 leads at 77.64% on corporate tax research. Source

Where it sits in the atlas

Work ladder: Direct evidence for Finance and accounting. On the work ladder it counts as 49.2% of tasks fully completed (All-Pass). How the ladder works

Work it measures (O*NET work activities): Advise others on legal or regulatory matters; Calculate financial data.

Finance and accountingLegal (partial)

Tax researchWeb research

Last checked 24 Sep 2026 against 1 primary source. See an error? Tell us.

How to cite

Credit the original work first: Tax Agent Bench by Vals AI.

Then, if you used this page:

Can Agents Work. "Tax Agent Bench: frontier results and sources." https://canagentswork.com/benchmarks/vals-tax-agent-bench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-vals-tax-agent-bench,
  title        = {{Tax Agent Bench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/vals-tax-agent-bench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}