Benchmarks / CORE-Bench
CORE-Bench
Built by Princeton University · Zachary S. Siegel, Sayash Kapoor, Nitya Nadgir, Benedikt Stroebl, et al. · released 17 Sep 2024
Computational reproducibility tasks from 90 published papers in computer science, social science, and medicine, all sourced from CodeOcean capsules known to run. An agent gets the paper's code and data, must install dependencies, run the code, and answer questions about the results, including values read from figures.
Frontier
95.5%
CORE-Bench-Hard accuracy (public test set)
Claude Code + Claude Opus 4.5 (after manual regrading) · harness: Claude Code
3 Dec 2025 · Source: Princeton University (benchmark maintainers)
HAL regraded 8 tasks by hand and removed 1 task with a dead data URL. The agent failed 2 remaining tasks. HAL declared the benchmark solved on this basis.
Its maintainers or a major evaluator declared it solved.
Solved on 3 Dec 2025: HAL (Princeton) declared CORE-Bench solved after Claude Opus 4.5 in a Claude Code scaffold scored 77.78% on Hard, rising to 95.5% after HAL fixed grading errors in 8 tasks by manual scoring and removed 1 task whose data URL had died. HAL said it would open a private test set of new papers. Source.
What it measures
Whether an agent can reproduce a paper's reported results from its own repository. Three levels: Easy gives the outputs, Medium gives a Docker command, Hard gives only the codebase. The headline is accuracy on CORE-Bench-Hard over the 45-paper public test set, where a task counts only if every question is answered within the tolerance set from three manual runs.
CORE-Bench-Hard accuracy (public test set): Share of the 45 Hard test-set tasks where the agent answers every task question correctly. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- cli, repo
- Grading
- automated-tests
- Tasks
- 45
- Human reference
- Every task was reproduced by hand three times to set answer tolerances, so a correct answer means matching what a careful human reproduction produced.
- Contamination
- The 45 test papers are public and named. HAL keeps a second private set of papers, disclosed only after the 80% threshold was crossed, for future evaluation.
- Reuse
- Harness and code are MIT licensed. Papers and capsules come from CodeOcean and keep their own licenses. (open-mit)
Limits to keep in mind
- HAL found grading errors in 9 tasks that only surfaced with strong agents. Deterministic outputs were penalized for tiny floating-point differences, and some tasks were underspecified. Source
- Papers were filtered to those that run in under 45 minutes, and the agent reproduces only selected results, so real reproduction is harder than the benchmark. Source
- Scores depend heavily on scaffold. Opus 4.5 scored 42.22% with HAL's CORE-Agent and 77.78% with Claude Code before any regrading. Source
- HAL has paused adding new models to its leaderboards while it focuses on reliability work. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | CORE-Bench-Hard accuracy (public test set) | Date | Source |
|---|---|---|---|
| Claude Code + Claude Opus 4.5 (automated grading) harness: Claude Code | 77.8% | 3 Dec 2025 | Princeton University · primary |
| Claude Code + Claude Opus 4.5 (after manual regrading) · frontier harness: Claude Code | 95.5% | 3 Dec 2025 | Princeton University · primary |
| CORE-Agent + Claude Opus 4.1 harness: CORE-Agent | 51.1% | 2025 | Princeton University · primary |
| Claude Code + Claude Sonnet 4.5 harness: Claude Code | 62.2% | 2025 | Princeton University · primary |
| CORE-Agent + GPT-4o harness: CORE-Agent | 21.5% | 17 Sep 2024 | Princeton University · primary |
Timeline
Where it sits in the atlas
Go to the source
- Website github.com
- Paper arxiv.org
- Full leaderboard hal.cs.princeton.edu
- Code github.com
Last checked 23 Sep 2026 against 5 primary sources. See an error? Tell us.
How to cite
Credit the original work first: CORE-Bench by Princeton University (https://arxiv.org/abs/2409.11363).
Then, if you used this page:
Can Agents Work. "CORE-Bench: frontier results and sources." https://canagentswork.com/benchmarks/core-bench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-core-bench,
title = {{CORE-Bench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/core-bench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}