Benchmarks / RedlineBench
RedlineBench
Built by Crosby Legal and micro1 · Sharan Ramjee, Juhi Pandit, Ryan Tanenholz, John Sarihan, et al. · released 16 Jun 2026
Contract redlining as one side of a multi-turn negotiation. An agent acts as in-house counsel for a vendor or a customer at one turn of a SaaS agreement negotiation and must return a Word file with tracked changes and comments. Attorneys wrote the reference redlines and the rubrics.
Frontier
50.5%
Turn-weighted rubric score
GPT-5.5
16 Jun 2026 · Source: Crosby Legal (benchmark maintainers)
First on the turn-weighted, cross-scenario score. Average of three rollouts. The maintainers say the spread between models is narrow. Side A (vendor) 48.2%, side B (customer) 52.9%.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
Whether the agent's redline makes the moves that attorneys made at the same point of the negotiation. 140 tasks come from three simulated deals between AgentCo, an AI hiring-platform vendor, and a large customer, each played over four turns. Turn 1 marks up a clean template; turns 2 to 4 respond to the counterparty's tracked changes. A panel of three LLM judges checks the output against attorney-written rubric items in five dimensions: legal correctness, commercial context, negotiation quality, counterparty acceptance, and deal-closing orientation. Scores are averaged per input group and per scenario-turn cell, so later turns do not dominate.
Turn-weighted rubric score: Share of rubric weight earned, averaged within each input group, then averaged equally over the 12 scenario-by-turn cells. Reported across three rollouts per model (one for Claude Fable 5). Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli, documents
- Grading
- llm-judge, rubric
- Tasks
- 140
- Human reference
- Attorney-authored golden redlines for 138 of 140 tasks. No human score is published. The report compares editing behavior: attorneys make 48.6% inline edits and 3.10 edits per touched paragraph at 101 characters each; models make fewer, longer block edits.
- Contamination
- All 140 tasks, rubrics, and golden redlines are public on Hugging Face. The maintainers do not keep a private set.
- Reuse
- CC-BY-4.0 (Hugging Face dataset card and repository README). (open-cc-by)
Limits to keep in mind
- Only four models have published scores, all run by the maintainers. Crosby Legal is a law firm that sells AI contract work, and it is a party to the benchmark. Source
- Scores average three rollouts per model, except Claude Fable 5 with one rollout. The maintainers say a full re-run is non-deterministic and the spread between models is narrow. Source
- Grading uses a panel of three small LLM judges (gpt-5.4-mini, claude-haiku-4-5, gemini-3.1-flash-lite) with a majority vote per rubric item. Source
- One deal type: three SaaS or services agreements between the same fictional parties. Scores may not carry over to other contract types. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Turn-weighted rubric score | Date | Source |
|---|---|---|---|
| GPT-5.5 · frontier | 50.5% | 16 Jun 2026 | Crosby Legal · primary |
| Claude Fable 5 | 47.3% | 16 Jun 2026 | Crosby Legal · primary |
| Gemini 3.5 Flash | 45.1% | 16 Jun 2026 | Crosby Legal · primary |
| Claude Opus 4.8 | 44.4% | 16 Jun 2026 | Crosby Legal · primary |
Timeline
- 16 Jun 2026 — RedlineBench launches; GPT-5.5 earns 50.5% of attorney rubric weight on contract redlines. Source
Where it sits in the atlas
Work ladder: Direct evidence for Legal, but its headline gives partial credit and the source publishes no completion rate. It adds to the work measured, but it cannot set a step. How the ladder works
Work it measures (O*NET work activities): Prepare legal or regulatory documents; Negotiate contracts or agreements.
Go to the source
Last checked 24 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: RedlineBench by Crosby Legal and micro1.
Then, if you used this page:
Can Agents Work. "RedlineBench: frontier results and sources." https://canagentswork.com/benchmarks/redlinebench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-redlinebench,
title = {{RedlineBench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/redlinebench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}