Benchmarks / Finance Agent Benchmark
Finance Agent Benchmark (Finance Agent)
Built by Vals AI · Antoine Bigeard, Langston Nashold, Rayan Krishnan, Shirley Wu · released 20 May 2025
Vals AI's test of whether an agent can do the research work of an entry-level financial analyst. Each question asks about public companies and their SEC filings, and the agent must find the answer with EDGAR search, web search, a page parser, and a retrieval tool.
Frontier
64.4%
Accuracy
Claude Opus 4.7 · harness: Vals finance-agent v1.1
4 Jun 2026 · Source: Vals AI (benchmark maintainers)
Top of the v1.1 leaderboard (page data 64.373, stderr 2.79, cost about $0.80 per question). Muse Spark 60.59, DeepSeek V4 Pro 60.39, and Claude Opus 4.6 (Thinking) 60.05 follow. Date is the page's "updated" field, not the run date.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Final-answer accuracy on a private test set of 337 expert-written questions (537 in total, with 50 public and 150 licensable validation questions). Questions span nine categories from simple retrieval to financial modeling and market analysis. An LLM judge (GPT-5.2, mode of three runs) compares each answer with the expert answer. Vals also records cost, latency, and tool calls.
Accuracy: Share of private test questions where the LLM judge accepts the agent's final answer. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- api-tools, documents
- Grading
- llm-judge
- Tasks
- 537
- Human reference
- Finance experts from banks, private equity firms, and hedge funds wrote and answered the questions. No human accuracy score is published.
- Contamination
- The test set is private and Vals says it will stay private.
- Reuse
- 50 public validation questions in the repository (MIT). 150 validation questions are available for license. The 337-question test set is private. (cite-only)
Limits to keep in mind
- Version 1.1 (early 2026) changed the data, tools, prompts, and judge, and re-ran every model. Scores from version 1.0, including the paper's 46.8% for o3, are not comparable. Source
- Each score has a standard error of about 2.8 points, so the top five models (60% to 64%) overlap. Source
- Grading uses an LLM judge. Vals takes the mode of three GPT-5.2 judgments to reduce variance. Source
- The agent only sees the fixed tool set in the Vals harness. Access to the Vals platform to run the private set is gated. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Accuracy | Date | Source |
|---|---|---|---|
| GPT-5.2 harness: Vals finance-agent v1.1 | 58.5% | 4 Jun 2026 | Vals AI · primary |
| Claude Sonnet 4.6 harness: Vals finance-agent v1.1 | 63.3% | 4 Jun 2026 | Vals AI · primary |
| Claude Opus 4.7 · frontier harness: Vals finance-agent v1.1 | 64.4% | 4 Jun 2026 | Vals AI · primary |
| o3 harness: Vals finance-agent v1.0 | 46.8% | 20 May 2025 | Vals AI · primary |
Timeline
- 20 May 2025 — Vals AI releases the Finance Agent Benchmark. Source
Where it sits in the atlas
Go to the source
- Website www.vals.ai
- Paper arxiv.org
- Full leaderboard www.vals.ai
- Code github.com
- Dataset github.com
Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.
How to cite
Credit the original work first: Finance Agent Benchmark by Vals AI (https://arxiv.org/abs/2508.00828).
Then, if you used this page:
Can Agents Work. "Finance Agent Benchmark: frontier results and sources." https://canagentswork.com/benchmarks/vals-finance-agent/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-vals-finance-agent,
title = {{Finance Agent Benchmark: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/vals-finance-agent/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}