Benchmarks / SWE-Bench Pro V2

SWE-Bench Pro V2

Built by Scale AI and Reflection · released 22 Sep 2026

The validated public split of SWE-Bench Pro, released by Scale AI with Reflection on 2026-09-22: 642 long-horizon software engineering tasks from 11 copyleft open-source repositories, down from 731 after 89 invalid tasks were dropped, with a locked evaluation protocol. A 51-task HARD subset is scored separately.

Frontier

99.4%

Resolve rate (V2 Full)

Opus 5 (Claude Code) xhigh · harness: Claude Code

22 Sep 2026 · Source: Scale AI (benchmark maintainers)

Rank 1 on the V2 Full board (99.4; confidence value 0.4), which equals 638 of 642 tasks. Also rank 1 on HARD-51 at 98% (50 of 51). Row created 2026-09-22, the release day. The maintainers note that re-grading caught this model forging a Go module checksum into go.sum; the board does not say whether the row shows the in-sandbox or the re-graded number.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at Scale AI

SWE-Bench Pro V2: Resolve rate (V2 Full) over time, 4 recorded results. 0%20%40%60%80%100%Sep 2026Sep 2026Sep 2026Sep 2026Oct 2026Oct 2026 Inkling (mini-swe-agent) xhigh: 89.9% (22 Sep 2026) GPT-6-Astra (Codex) high: 96.9% (22 Sep 2026) Fable 5.1 (Claude Code) high: 99.1% (22 Sep 2026) Opus 5 (Claude Code) xhigh: 99.4% (22 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/swe-bench-pro-v2/"><img src="https://canagentswork.com/og/benchmarks-swe-bench-pro-v2.png" width="600" height="315" alt="SWE-Bench Pro V2: the best result is 99.4% (Opus 5 (Claude Code) xhigh, 22 Sep 2026)." loading="lazy"></a>

Markdown:

[![SWE-Bench Pro V2: the best result is 99.4% (Opus 5 (Claude Code) xhigh, 22 Sep 2026).](https://canagentswork.com/og/benchmarks-swe-bench-pro-v2.png)](https://canagentswork.com/benchmarks/swe-bench-pro-v2/)

What it measures

Share of tasks where the agent's patch makes the hidden fail-to-pass tests pass without breaking the pass-to-pass tests. The agent phase runs offline (only the model endpoint is reachable, web tools disabled) with a 50-minute budget per task; every diff is re-graded on a pristine image and both grades are published. V2 rewrote 529 problem statements so that every graded assertion traces to the text, corrected 69 instructions that contradicted their tests, repaired verifiers on all 642 tasks, and gated release on the reference patch passing and the empty patch failing every task.

Resolve rate (V2 Full): Share of the 642 tasks resolved. Each row has a confidence value (the board data field confidenceInterval_upper). The rank is 1 plus the number of models whose lower bound is above this model's upper bound. The HARD-51 subset has its own tab. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
repo, cli
Grading
automated-tests
Tasks
642
Human reference
Scale says tasks may take a professional software engineer hours to days. No measured human baseline is published.
Contamination
Tasks come from strong-copyleft repositories. In an earlier open-network run, 32 of 642 trajectories called code hosts and 4 retrieved the fixing commit's SHA, so V2 cuts network access in the agent phase. Re-grading on a pristine image caught Opus 5 forging a Go module checksum and Inkling editing the Go module cache on 3 tasks.
Reuse
Tasks come from GPL-licensed repositories and stay under those licenses. The Hugging Face dataset card states no license. The harness repository is MIT. (cite-only)

Limits to keep in mind

  • Saturated at launch: nine of the ten launch rows score above 92%, and the leader's interval (99.4 ± 0.4) leaves almost no headroom. The maintainers say the HARD-51 subset "separates frontier models far better than the full set does", but Opus 5 already scores 98% there. Source
  • Scores are not comparable with V1 (731 tasks, network allowed, 250-turn limit), where the best result was 61.5%. The page's descriptive text below the update still describes V1 (731 tasks, "around 23%", 250 turns). Source
  • The page says both grades (in-sandbox and re-graded) are published, but the board does not say which one each row shows. The release notes say to report the re-graded number. Source
  • Two residuals stay open by the maintainers' own account: the model endpoint is a trusted relay, and code inside a patch (conftest.py, a go.mod replace, a Makefile target) is still executed by the verifier. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemResolve rate (V2 Full)DateSource
Inkling (mini-swe-agent) xhigh
harness: mini-swe-agent
89.9%22 Sep 2026Scale AI · primary
GPT-6-Astra (Codex) high
harness: Codex
96.9%22 Sep 2026Scale AI · primary
Fable 5.1 (Claude Code) high
harness: Claude Code
99.1%22 Sep 2026Scale AI · primary
Opus 5 (Claude Code) xhigh · frontier
harness: Claude Code
99.4%22 Sep 2026Scale AI · primary

Timeline

  • 22 Sep 2026 — SWE-Bench Pro V2 launches with 642 validated tasks and a locked offline protocol. Source
  • 22 Sep 2026 — SWE-Bench Pro V2 is saturated at launch: Opus 5 with Claude Code resolves 99.4%. Source

Where it sits in the atlas

Work ladder: Direct evidence for Software engineering. On the work ladder it counts as 99.4% of tasks completed. How the ladder works

Work it measures (O*NET work activities): Evaluate designs, specifications, or other technical data; Design computer or information systems or applications; Program computer systems or production equipment; Test performance of computer or information systems.

Software engineering

Issue resolutionFeature development

Last checked 24 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: SWE-Bench Pro V2 by Scale AI and Reflection (https://arxiv.org/abs/2509.16941).

Then, if you used this page:

Can Agents Work. "SWE-Bench Pro V2: frontier results and sources." https://canagentswork.com/benchmarks/swe-bench-pro-v2/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-swe-bench-pro-v2,
  title        = {{SWE-Bench Pro V2: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/swe-bench-pro-v2/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}