Sources and licenses
Every benchmark, who made it, how you may reuse its data, and when we last checked it. Our own contributions are CC BY 4.0; source data keeps its original terms. Institutions · Downloads.
Benchmarks
| Benchmark | Built by | Reuse | Last checked |
|---|---|---|---|
| Aider Polyglot leaderboard | Aider | Cite; link to the source Exercises are copyright Exercism and used under the Exercism tracks' open-source licenses. The benchmark repository states no license of its own. | 23 Sep 2026 |
| AIDev | Queen's University | Unknown; cite only Not checked in this session. Dataset is on Hugging Face and Zenodo (DOI 10.5281/zenodo.16919272). | 23 Sep 2026 |
| AIOpsLab | Microsoft Research, UIUC, UC Berkeley, and IISc | Open (MIT) MIT (GitHub repository license). | 23 Sep 2026 |
| Ambig-SWE | Carnegie Mellon University | Open (MIT) Paper CC BY 4.0; repo MIT. Issues derive from SWE-Bench Verified. | 23 Sep 2026 |
| APEX-Agents | Mercor, Box, and Harvey | Open (CC BY) CC BY 4.0 on the Hugging Face dataset. Mercor says the full task set used for the leaderboard stays private. | 23 Sep 2026 |
| BrowseComp | OpenAI | Open (MIT) Released in OpenAI's simple-evals repository, which lists BrowseComp under the MIT License. | 23 Sep 2026 |
| Code Review Bench | Martian | Open (MIT) MIT (repo license covers PRs metadata, golden comments, judge prompts, pipeline, and results) | 23 Sep 2026 |
| CodeClash | Stanford University and Princeton University | Open (MIT) MIT (CodeClash repository). Arenas are third-party games with their own licenses. | 23 Sep 2026 |
| CORE-Bench | Princeton University | Open (MIT) Harness and code are MIT licensed. Papers and capsules come from CodeOcean and keep their own licenses. | 23 Sep 2026 |
| CRMArena-Pro | Salesforce AI Research | Cite; link to the source CC BY-NC 4.0 (repository LICENSE.txt). Research use only. | 23 Sep 2026 |
| CVE-Bench | UIUC | Open (Apache-2.0) Apache-2.0 (GitHub repository license). Reference exploits are available on request. | 23 Sep 2026 |
| Cybench | Stanford University | Open (Apache-2.0) Tasks are public in the GitHub repository (Apache-2.0); the site asks users to cite the ICLR 2025 paper. | 23 Sep 2026 |
| CyberGym | UC Berkeley | Open (Apache-2.0) Apache-2.0 (GitHub repository license). Dataset hosted on Hugging Face (about 240 GB). | 23 Sep 2026 |
| Finance Agent Benchmark | Vals AI | Cite; link to the source 50 public validation questions in the repository (MIT). 150 validation questions are available for license. The 337-question test set is private. | 23 Sep 2026 |
| GAIA | Meta and Hugging Face | Cite; link to the source Gated dataset on Hugging Face. Test answers are private. The maintainers ask that the public set not be reposted or used for training. | 23 Sep 2026 |
| GDPval | OpenAI | Cite; link to the source The 220-task gold subset (prompts and reference files) is public on Hugging Face. The other 1,100 tasks and the expert grades are private. Rubrics and gold deliverables were later released. | 23 Sep 2026 |
| GDPval-AA | Artificial Analysis | Cite; link to the source Tasks are OpenAI's public GDPval gold set. The Stirrup harness is open source on GitHub. Artificial Analysis publishes Elo scores and example submissions on its site. | 23 Sep 2026 |
| IaC-Eval | University of Michigan and Cisco Research | Open (CC BY) CC-BY-4.0 (Hugging Face dataset card) | 23 Sep 2026 |
| ImpossibleBench | Carnegie Mellon University and Anthropic | Open (MIT) Paper CC BY 4.0; repo MIT. Task data derives from SWE-bench Verified and LiveCodeBench. | 23 Sep 2026 |
| ITBench | IBM Research and UIUC | Open (Apache-2.0) Apache-2.0 (repository LICENSE file). Scenario tooling and sample scenarios are open source; hosted environments require registration. | 23 Sep 2026 |
| ITBench-AA | Artificial Analysis and IBM Research | Cite; link to the source 40 public tasks and 19 held-out tasks; the held-out tasks are private to Artificial Analysis. | 23 Sep 2026 |
| MedAgentBench | Stanford University | Cite; link to the source Code is MIT (repository license). The arXiv paper is CC BY 4.0. Patient data derives from de-identified Stanford STARR records and ships as a Docker image; the reference solutions are distributed separately from a Stanford Medicine Box link. | 23 Sep 2026 |
| METR Task-Completion Time Horizons | METR | Cite; link to the source Horizon estimates and run data are public in eval-analysis-public (no license file seen). Many tasks are private. | 23 Sep 2026 |
| MLE-bench | OpenAI | Cite; link to the source Code is MIT licensed. Competition data comes from Kaggle and users must accept each competition's rules to download it. | 23 Sep 2026 |
| OSWorld-Verified | HKU, Salesforce AI Research, Carnegie Mellon University, and University of Waterloo | Open (Apache-2.0) Apache 2.0 on the GitHub repository. Windows tasks need a licensed image. | 23 Sep 2026 |
| PaperBench | OpenAI | Open (MIT) Code and rubrics are in the MIT-licensed openai/frontier-evals repo. Papers belong to their authors. | 23 Sep 2026 |
| PR Arena | aavetis | Unknown; cite only No license file in the repo. Data is derived from public GitHub search counts. | 23 Sep 2026 |
| RE-Bench | METR | Open (MIT) Repo is MIT licensed. Reference solutions are password-protected. METR asks users to keep the tasks out of training data and not to publish solutions. | 23 Sep 2026 |
| Remote Labor Index | CAIS and Scale AI | Cite; link to the source Scores come from 230 private projects. 10 public projects and the open-source evaluation platform are released for qualitative analysis. | 23 Sep 2026 |
| Spider 2.0 | HKU and Salesforce AI Research | Open (MIT) Repository is MIT licensed. All examples and gold answers were released for self-evaluation in December 2024. Maintainers ask users not to fine-tune on the gold SQL. | 23 Sep 2026 |
| SpreadsheetBench | Renmin University of China | Open (other license) CC BY-SA 4.0, stated in the paper's maintenance plan and the repo README. | 23 Sep 2026 |
| SREGym | UIUC and University of Toronto | Open (MIT) MIT (GitHub repository license). Paper is CC BY 4.0. | 23 Sep 2026 |
| SWE-Bench Pro | Scale AI | Cite; link to the source Public-set tasks come from GPL-licensed repositories and stay under those licenses. The harness is MIT. The held-out and commercial sets are not released. | 23 Sep 2026 |
| SWE-bench Verified | Princeton University and OpenAI | Cite; link to the source Task content comes from 12 open-source Python repositories under their own licenses. The SWE-bench harness is MIT. | 23 Sep 2026 |
| SWE-Lancer | OpenAI | Open (MIT) Diamond split is public in the openai/preparedness repository (MIT). The remaining tasks are a private holdout. | 23 Sep 2026 |
| SWE-rebench | Nebius | Cite; link to the source Tasks come from repositories under permissive licenses (MIT, Apache-2.0, BSD, ISC, and others that Nebius checked by hand). The leaderboard dataset and Docker images are public. | 23 Sep 2026 |
| Terminal-Bench 2.0 | Laude Institute and Stanford University | Open (Apache-2.0) Apache-2.0 (terminal-bench-2 repository). | 23 Sep 2026 |
| TheAgentCompany | Carnegie Mellon University | Open (MIT) MIT license on the GitHub repository. Task images, evaluators, and company data are public. | 23 Sep 2026 |
| Vending-Bench 2 | Andon Labs | Cite; link to the source No public code or data release is linked from the benchmark page. Andon Labs runs the simulation. | 23 Sep 2026 |
| Will It Survive? Agent code survival study | Concordia University | Cite; link to the source Paper CC BY 4.0. Underlying data comes from the AIDev dataset. | 23 Sep 2026 |
| WorkArena | ServiceNow | Open (Apache-2.0) Apache 2.0 for the benchmark code. Running tasks needs access to a ServiceNow developer instance. | 23 Sep 2026 |
| τ²-bench | Sierra | Open (MIT) MIT (repository license covers tasks, policies, and code). | 23 Sep 2026 |
Where sources disagree
CodeClash · open · opened 23 Sep 2026
The CodeClash site leaderboard shows Claude Sonnet 4.5 at 1385 ± 18 and GPT-5 at 1366 ± 17. Version 2 of the paper (May 2026) shows 1389 ± 18 and 1360 ± 17 for the same models. Both come from the maintainers. The difference is small and the ranking is the same. A refit of the Elo model is one possible cause. We show the leaderboard value.
- 1385 — codeclash.ai (primary) · Site leaderboard, "Updated Nov. 3, 2025", seen 2026-09-23.
- 1389 — arxiv.org (primary) · Paper v2 (2026-05-12), Table 1 and Table 4.
We show 1385.
ITBench · open · opened 23 Sep 2026
The ICML 2025 paper says agents resolve 11.4% of SRE scenarios. The arXiv v1 abstract says 13.8%. The official SRE leaderboard shows 25.0% for IBM's GPT-4o reference agent. The three numbers come from different scenario sets and dates. We show the leaderboard value because it is the maintainers' most recent published result.
- 11.4 — proceedings.mlr.press (primary) · ICML 2025 camera-ready abstract, 102 scenarios.
- 13.8 — arxiv.org (primary) · arXiv v1 abstract (2025-02-07), 94 scenarios.
- 25 — github.com (primary) · SRE leaderboard, ITBench-SRE-Agent-GPT-4o, 16 trials, updated 2 May 2025. Multi-trial table shows 24.79%.
We show 25.
METR Task-Completion Time Horizons · resolved · opened 23 Sep 2026
METR's Time Horizon 1.1 blog post (2026-01-29) gives Claude Opus 4.5 a 50% horizon of 320 minutes. METR's live data file gives 293.0 minutes. Both are METR sources. The results page lists a 2026-03-03 update that "corrected a regularization mistake that affected our measurements", which explains the change. GPT-5 moved the same way (214 to 203 minutes).
- 293 — metr.org (primary) · Live data file linked from metr.org/time-horizons, retrieved 2026-09-23. CI 161.7 to 623.7.
- 320 — metr.org (primary) · TH1.1 launch post, table "Changes to Model Horizon Estimates". CI 170 to 729.
We show 293. Resolution: We show the live data file, which reflects the 2026-03-03 correction. The blog post predates it.
Remote Labor Index · open · opened 23 Sep 2026
CAIS, a co-author of the benchmark, reports 15.8% for Claude Fable 5. Two aggregator pages report 16.1%. We did not find the cause. A later re-grade is one possible cause.
- 15.8 — safe.ai (primary) · CAIS post dated 2026-07-01.
- 16.1 — www.benchleader.com (third-party) · BenchLeader page, seen 2026-09-23.
- 16.1 — ai-intensify.com (third-party) · News article, seen 2026-09-23.
We show 15.8.
SpreadsheetBench · open · opened 23 Sep 2026
The SpreadsheetBench site shows two different top scores for the 912-question V1 set. The static "Top Score (OVERALL)" box reads 70.48%. The leaderboard data file that the page loads lists Qingqiu Agent at 83.11% (verified, 2026-06-23). The box looks stale. We show the leaderboard value.
- 83.11 — spreadsheetbench.github.io (primary) · Leaderboard data file, seen 2026-09-23. Entry is marked verified.
- 70.48 — spreadsheetbench.github.io (primary) · Static statistics box on the V1 overview page, seen 2026-09-23. Matches the Gemini in Google Sheets row dated 2026-03-10.
We show 83.11.
SWE-Bench Pro · open · opened 23 Sep 2026
Scale's public leaderboard shows a best resolve rate of 61.5% (Muse Spark 1.1, mini-swe-agent, added 2026-07-09). OpenAI's 2026-07-08 audit post says frontier models reached 80.3% on the same 731-task public split, without naming the model or harness. The gap is likely a different agent scaffold and OpenAI's own internal runs. We show the leaderboard value.
- 61.5 — labs.scale.com (primary) · Scale AI public leaderboard, Muse Spark 1.1 with mini-swe-agent, 61.50 ± 3.10, seen 2026-09-23.
- 80.3 — openai.com (lab-reported) · OpenAI post dated 2026-07-08. "Frontier models improved from a pass rate of 23.3% to 80.3% in eight months." No model or harness named.
We show 61.5.
Reference data
- O*NET 31.0 Database (U.S. Department of Labor, Employment and Training Administration): occupations, job titles, and tasks. CC BY 4.0. onetcenter.org
- BLS OEWS, May 2025 national estimates (U.S. Bureau of Labor Statistics): employment and wages. Public domain. bls.gov/oes