Sources and licenses

Every benchmark, who made it, how you may reuse its data, and when we last checked it. Our own contributions are CC BY 4.0; source data keeps its original terms. Institutions · Downloads.

Benchmarks

BenchmarkBuilt byReuseLast checked
Aider Polyglot leaderboardAiderCite; link to the source
Exercises are copyright Exercism and used under the Exercism tracks' open-source licenses. The benchmark repository states no license of its own.
23 Sep 2026
AIDevQueen's UniversityUnknown; cite only
Not checked in this session. Dataset is on Hugging Face and Zenodo (DOI 10.5281/zenodo.16919272).
23 Sep 2026
AIOpsLabMicrosoft Research, UIUC, UC Berkeley, and IIScOpen (MIT)
MIT (GitHub repository license).
23 Sep 2026
Ambig-SWECarnegie Mellon UniversityOpen (MIT)
Paper CC BY 4.0; repo MIT. Issues derive from SWE-Bench Verified.
23 Sep 2026
APEX-AgentsMercor, Box, and HarveyOpen (CC BY)
CC BY 4.0 on the Hugging Face dataset. Mercor says the full task set used for the leaderboard stays private.
23 Sep 2026
BrowseCompOpenAIOpen (MIT)
Released in OpenAI's simple-evals repository, which lists BrowseComp under the MIT License.
23 Sep 2026
Code Review BenchMartianOpen (MIT)
MIT (repo license covers PRs metadata, golden comments, judge prompts, pipeline, and results)
23 Sep 2026
CodeClashStanford University and Princeton UniversityOpen (MIT)
MIT (CodeClash repository). Arenas are third-party games with their own licenses.
23 Sep 2026
CORE-BenchPrinceton UniversityOpen (MIT)
Harness and code are MIT licensed. Papers and capsules come from CodeOcean and keep their own licenses.
23 Sep 2026
CRMArena-ProSalesforce AI ResearchCite; link to the source
CC BY-NC 4.0 (repository LICENSE.txt). Research use only.
23 Sep 2026
CVE-BenchUIUCOpen (Apache-2.0)
Apache-2.0 (GitHub repository license). Reference exploits are available on request.
23 Sep 2026
CybenchStanford UniversityOpen (Apache-2.0)
Tasks are public in the GitHub repository (Apache-2.0); the site asks users to cite the ICLR 2025 paper.
23 Sep 2026
CyberGymUC BerkeleyOpen (Apache-2.0)
Apache-2.0 (GitHub repository license). Dataset hosted on Hugging Face (about 240 GB).
23 Sep 2026
Finance Agent BenchmarkVals AICite; link to the source
50 public validation questions in the repository (MIT). 150 validation questions are available for license. The 337-question test set is private.
23 Sep 2026
GAIAMeta and Hugging FaceCite; link to the source
Gated dataset on Hugging Face. Test answers are private. The maintainers ask that the public set not be reposted or used for training.
23 Sep 2026
GDPvalOpenAICite; link to the source
The 220-task gold subset (prompts and reference files) is public on Hugging Face. The other 1,100 tasks and the expert grades are private. Rubrics and gold deliverables were later released.
23 Sep 2026
GDPval-AAArtificial AnalysisCite; link to the source
Tasks are OpenAI's public GDPval gold set. The Stirrup harness is open source on GitHub. Artificial Analysis publishes Elo scores and example submissions on its site.
23 Sep 2026
IaC-EvalUniversity of Michigan and Cisco ResearchOpen (CC BY)
CC-BY-4.0 (Hugging Face dataset card)
23 Sep 2026
ImpossibleBenchCarnegie Mellon University and AnthropicOpen (MIT)
Paper CC BY 4.0; repo MIT. Task data derives from SWE-bench Verified and LiveCodeBench.
23 Sep 2026
ITBenchIBM Research and UIUCOpen (Apache-2.0)
Apache-2.0 (repository LICENSE file). Scenario tooling and sample scenarios are open source; hosted environments require registration.
23 Sep 2026
ITBench-AAArtificial Analysis and IBM ResearchCite; link to the source
40 public tasks and 19 held-out tasks; the held-out tasks are private to Artificial Analysis.
23 Sep 2026
MedAgentBenchStanford UniversityCite; link to the source
Code is MIT (repository license). The arXiv paper is CC BY 4.0. Patient data derives from de-identified Stanford STARR records and ships as a Docker image; the reference solutions are distributed separately from a Stanford Medicine Box link.
23 Sep 2026
METR Task-Completion Time HorizonsMETRCite; link to the source
Horizon estimates and run data are public in eval-analysis-public (no license file seen). Many tasks are private.
23 Sep 2026
MLE-benchOpenAICite; link to the source
Code is MIT licensed. Competition data comes from Kaggle and users must accept each competition's rules to download it.
23 Sep 2026
OSWorld-VerifiedHKU, Salesforce AI Research, Carnegie Mellon University, and University of WaterlooOpen (Apache-2.0)
Apache 2.0 on the GitHub repository. Windows tasks need a licensed image.
23 Sep 2026
PaperBenchOpenAIOpen (MIT)
Code and rubrics are in the MIT-licensed openai/frontier-evals repo. Papers belong to their authors.
23 Sep 2026
PR ArenaaavetisUnknown; cite only
No license file in the repo. Data is derived from public GitHub search counts.
23 Sep 2026
RE-BenchMETROpen (MIT)
Repo is MIT licensed. Reference solutions are password-protected. METR asks users to keep the tasks out of training data and not to publish solutions.
23 Sep 2026
Remote Labor IndexCAIS and Scale AICite; link to the source
Scores come from 230 private projects. 10 public projects and the open-source evaluation platform are released for qualitative analysis.
23 Sep 2026
Spider 2.0HKU and Salesforce AI ResearchOpen (MIT)
Repository is MIT licensed. All examples and gold answers were released for self-evaluation in December 2024. Maintainers ask users not to fine-tune on the gold SQL.
23 Sep 2026
SpreadsheetBenchRenmin University of ChinaOpen (other license)
CC BY-SA 4.0, stated in the paper's maintenance plan and the repo README.
23 Sep 2026
SREGymUIUC and University of TorontoOpen (MIT)
MIT (GitHub repository license). Paper is CC BY 4.0.
23 Sep 2026
SWE-Bench ProScale AICite; link to the source
Public-set tasks come from GPL-licensed repositories and stay under those licenses. The harness is MIT. The held-out and commercial sets are not released.
23 Sep 2026
SWE-bench VerifiedPrinceton University and OpenAICite; link to the source
Task content comes from 12 open-source Python repositories under their own licenses. The SWE-bench harness is MIT.
23 Sep 2026
SWE-LancerOpenAIOpen (MIT)
Diamond split is public in the openai/preparedness repository (MIT). The remaining tasks are a private holdout.
23 Sep 2026
SWE-rebenchNebiusCite; link to the source
Tasks come from repositories under permissive licenses (MIT, Apache-2.0, BSD, ISC, and others that Nebius checked by hand). The leaderboard dataset and Docker images are public.
23 Sep 2026
Terminal-Bench 2.0Laude Institute and Stanford UniversityOpen (Apache-2.0)
Apache-2.0 (terminal-bench-2 repository).
23 Sep 2026
TheAgentCompanyCarnegie Mellon UniversityOpen (MIT)
MIT license on the GitHub repository. Task images, evaluators, and company data are public.
23 Sep 2026
Vending-Bench 2Andon LabsCite; link to the source
No public code or data release is linked from the benchmark page. Andon Labs runs the simulation.
23 Sep 2026
Will It Survive? Agent code survival studyConcordia UniversityCite; link to the source
Paper CC BY 4.0. Underlying data comes from the AIDev dataset.
23 Sep 2026
WorkArenaServiceNowOpen (Apache-2.0)
Apache 2.0 for the benchmark code. Running tasks needs access to a ServiceNow developer instance.
23 Sep 2026
τ²-benchSierraOpen (MIT)
MIT (repository license covers tasks, policies, and code).
23 Sep 2026

Where sources disagree

CodeClash · open · opened 23 Sep 2026

The CodeClash site leaderboard shows Claude Sonnet 4.5 at 1385 ± 18 and GPT-5 at 1366 ± 17. Version 2 of the paper (May 2026) shows 1389 ± 18 and 1360 ± 17 for the same models. Both come from the maintainers. The difference is small and the ranking is the same. A refit of the Elo model is one possible cause. We show the leaderboard value.

We show 1385.

ITBench · open · opened 23 Sep 2026

The ICML 2025 paper says agents resolve 11.4% of SRE scenarios. The arXiv v1 abstract says 13.8%. The official SRE leaderboard shows 25.0% for IBM's GPT-4o reference agent. The three numbers come from different scenario sets and dates. We show the leaderboard value because it is the maintainers' most recent published result.

We show 25.

METR Task-Completion Time Horizons · resolved · opened 23 Sep 2026

METR's Time Horizon 1.1 blog post (2026-01-29) gives Claude Opus 4.5 a 50% horizon of 320 minutes. METR's live data file gives 293.0 minutes. Both are METR sources. The results page lists a 2026-03-03 update that "corrected a regularization mistake that affected our measurements", which explains the change. GPT-5 moved the same way (214 to 203 minutes).

We show 293. Resolution: We show the live data file, which reflects the 2026-03-03 correction. The blog post predates it.

Remote Labor Index · open · opened 23 Sep 2026

CAIS, a co-author of the benchmark, reports 15.8% for Claude Fable 5. Two aggregator pages report 16.1%. We did not find the cause. A later re-grade is one possible cause.

We show 15.8.

SpreadsheetBench · open · opened 23 Sep 2026

The SpreadsheetBench site shows two different top scores for the 912-question V1 set. The static "Top Score (OVERALL)" box reads 70.48%. The leaderboard data file that the page loads lists Qingqiu Agent at 83.11% (verified, 2026-06-23). The box looks stale. We show the leaderboard value.

We show 83.11.

SWE-Bench Pro · open · opened 23 Sep 2026

Scale's public leaderboard shows a best resolve rate of 61.5% (Muse Spark 1.1, mini-swe-agent, added 2026-07-09). OpenAI's 2026-07-08 audit post says frontier models reached 80.3% on the same 731-task public split, without naming the model or harness. The gap is likely a different agent scaffold and OpenAI's own internal runs. We show the leaderboard value.

We show 61.5.

Reference data