Benchmarks / Harvey Legal Agent Benchmark

Harvey Legal Agent Benchmark (LAB)

Built by Harvey and Vals AI · Niko Grupen, Gabe Pereyra, Julio Pereyra · released 6 May 2026

Harvey's open-source benchmark of legal work assigned the way a partner delegates to an associate. Each task gives an agent a short instruction and a client matter of files, and asks for a reviewable work product such as a memo, a redline, or a disclosure schedule. Vals AI runs the held-out set and publishes the leaderboard that the atlas follows.

Frontier

25.4%

All-pass rate (held-out set)

Muse Spark 1.2 · harness: Harvey LAB harness (Vals run)

22 Sep 2026 · Source: Vals AI (third party)

Top of the Vals board on Harvey's final score (page data 25.417, stderr 3.777). About $2.09 and 24 minutes per task. Vals runs Harvey's harness and grading protocol on the held-out set with Harvey as evaluation partner. Date is the page's "updated" field. Open-weights status not confirmed.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Harvey

Harvey Legal Agent Benchmark: All-pass rate (held-out set) over time, 3 recorded results. 0%20%40%60%80%100%Jul 2026Oct 2026 Claude Opus 4.7: 7.1% (26 May 2026) Muse Spark 1.3 Max: 23.8% (22 Sep 2026) Muse Spark 1.2: 25.4% (22 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/harvey-lab/"><img src="https://canagentswork.com/og/benchmarks-harvey-lab.png" width="600" height="315" alt="Harvey Legal Agent Benchmark: the best result is 25.4% (Muse Spark 1.2, 22 Sep 2026)." loading="lazy"></a>

Markdown:

[![Harvey Legal Agent Benchmark: the best result is 25.4% (Muse Spark 1.2, 22 Sep 2026).](https://canagentswork.com/og/benchmarks-harvey-lab.png)](https://canagentswork.com/benchmarks/harvey-lab/)

What it measures

All-pass rate on Harvey's held-out task set: a task counts only when every rubric criterion passes, as judged by two LLM judges (GPT 5.5 and Claude Sonnet 4.6) whose task pass rates are averaged. The agent works in a sandbox with six file-system tools (Read, Edit, Write, Glob, Bash, Grep) and docx, pptx, and xlsx skills, with internet access off. The public set launched with 1,250 tasks in 24 practice areas and has grown with in-house contracting (500 tasks), M&A diligence, and law-firm knowledge (250 tasks) extensions. Vals also reports the criteria pass rate, which sits near 90% for top models.

All-pass rate (held-out set): Share of held-out tasks where every rubric criterion passes, averaged over two LLM judges (Harvey's "final score"). Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
cli, documents
Grading
llm-judge, rubric
Tasks
1,671
Human reference
Tasks were built from real client matters handled by practicing lawyers and broken into the assignments an associate would receive. Rubrics emulate partner or client review. No human score is published.
Contamination
Scores come from a private held-out set that mirrors the practice areas and task distribution of the public set. Harvey does not publish its size.
Reuse
Public tasks, rubrics, and harness are in the MIT-licensed repository. The held-out set used for scores is private. (open-mit)

Limits to keep in mind

  • Standard errors on the Vals board are about 3.8 points, so the top three Muse Spark models (25.42%, 23.75%, 22.08%) overlap. Source
  • Artificial Analysis runs its own implementation (Harvey LAB-AA) with a different harness, a single judge, and 120 private tasks, and leads with criterion pass rate (95.5% for Muse Spark 1.3 xhigh). Those numbers are a different measurement and are not shown here. Source
  • Vals fixed a harness bug so judges could see DOCX tracked changes (harveyai/harvey-labs#76) before its runs. Harvey's May 2026 numbers came from Harvey's own runs and judges, and they differ slightly from Vals's numbers for the same models. Source
  • The public task set keeps growing (contracting, diligence, and firm-knowledge extensions), so the README task count changes while the held-out scores stay on one set. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemAll-pass rate (held-out set)DateSource
Muse Spark 1.3 Max
harness: Harvey LAB harness (Vals run)
23.8%22 Sep 2026Vals AI · third-party
Muse Spark 1.2 · frontier
harness: Harvey LAB harness (Vals run)
25.4%22 Sep 2026Vals AI · third-party
Claude Opus 4.7
harness: Harvey LAB harness
7.1%26 May 2026Harvey · primary

Where sources disagree

Harvey's initial results (May 2026) put Claude Opus 4.7 at 7.1% all-pass on the held-out set. Vals AI's board, which uses Harvey's harness and grading protocol on the same set, shows 6.67% for the same model. The likely cause is separate agent runs and judge sets: Harvey averaged several gradings across model families, while Vals used GPT 5.5 and Claude Sonnet 4.6 as judges after fixing a DOCX tracked-changes bug in the harness. Neither source explains the gap.

  • 7.1 — www.harvey.ai (primary) · Harvey post dated 2026-05-26, Figure 1.
  • 6.67 — www.vals.ai (third-party) · Page data 6.667 (stderr 1.998), page updated 2026-09-22, seen 2026-09-24.

We show 7.1. Status: open.

Timeline

  • 22 Sep 2026 — Muse Spark 1.2 completes 25.42% of Harvey LAB held-out tasks on the Vals AI board. Source
  • 6 May 2026 — Harvey open-sources its Legal Agent Benchmark with 1,250 tasks in 24 practice areas. Source

Where it sits in the atlas

Work ladder: Direct evidence for Legal. On the work ladder it counts as 25.4% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Prepare legal or regulatory documents; Research laws, precedents, or other legal data; Consult legal materials or public records.

Legal

Legal workProfessional deliverables

Last checked 24 Sep 2026 against 8 primary sources. See an error? Tell us.

How to cite

Credit the original work first: Harvey Legal Agent Benchmark by Harvey and Vals AI.

Then, if you used this page:

Can Agents Work. "Harvey Legal Agent Benchmark: frontier results and sources." https://canagentswork.com/benchmarks/harvey-lab/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-harvey-lab,
  title        = {{Harvey Legal Agent Benchmark: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/harvey-lab/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}