Benchmarks / MLE-bench
MLE-bench
Built by OpenAI · Jun Shern Chan, Neil Chowdhury, Oliver Jaffe, James Aung, et al. · released 9 Oct 2024
75 Kaggle competitions rebuilt offline. An agent gets the data and description, trains models, and submits predictions. Its score is compared with the real human leaderboard to see whether it would have won a bronze, silver, or gold medal.
Frontier
64.4%
Any medal (%)
Famou-Agent 2.0 (Gemini-3-Pro-Preview) · agent: Famou-Agent 2.0 (Baidu)
23 Feb 2026 · Source: OpenAI (benchmark maintainers)
64.44 ± 1.18 over 24 hours per competition. Top of the main leaderboard table. Disarray reports 77.78 but used test-set feedback and is listed separately as not comparable.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether an agent can do end-to-end machine learning engineering: prepare data, train and tune models, and produce a valid submission. The headline is the share of competitions where the agent's best submission reaches at least a bronze medal, averaged over at least 3 seeds on all 75 competitions. The repo also reports Low (22 competitions, the "Lite" set), Medium, and High splits.
Any medal (%): Share of the 75 competitions where the agent earns any medal, reported as the mean over seeds. Higher is better.
Status compares the frontier with a ceiling of 100%.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- cli
- Grading
- outcome-metric
- Tasks
- 75
- Human reference
- Medal thresholds come from the public Kaggle leaderboards, so a medal means the agent beat most human participants in that competition.
- Contamination
- The competitions are public and top solutions are online. The paper ran familiarity experiments and found no clear link between a model's familiarity with a competition and its score.
- Reuse
- Code is MIT licensed. Competition data comes from Kaggle and users must accept each competition's rules to download it. (cite-only)
Limits to keep in mind
- Since 2026-04-24 the maintainers accept no new leaderboard submissions while they design a fairer submission process. Source
- Two submissions (Disarray 77.78%, LoongFlow 62.66%) are listed separately because they used test-set feedback and are not comparable with the main table. Source
- Known grading issues in several competitions are left unfixed to keep the v1 leaderboard comparable. Fixes are planned for a v2 release. Source
- Agents run for 24 hours per competition with modern models, while Kaggle participants worked under different conditions, so a medal is not a like-for-like human comparison. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Any medal (%) | Date | Source |
|---|---|---|---|
| AIBuildAI (Claude Opus 4.6) | 63.1% | 6 Mar 2026 | OpenAI · primary |
| Famou-Agent 2.0 (Gemini-3-Pro-Preview) · frontier | 64.4% | 23 Feb 2026 | OpenAI · primary |
| Famou-Agent 2.0 (Gemini-2.5-Pro) | 59.6% | 27 Dec 2025 | OpenAI · primary |
| R&D-Agent (o3 + GPT-4.1) harness: R&D-Agent | 30.2% | 15 Aug 2025 | OpenAI · primary |
| AIDE + o1-preview harness: AIDE | 17.1% | 8 Oct 2024 | OpenAI · primary |
Timeline
Where it sits in the atlas
Go to the source
- Website github.com
- Paper arxiv.org
- Full leaderboard github.com
- Code github.com
Last checked 23 Sep 2026 against 3 primary sources. See an error? Tell us.
How to cite
Credit the original work first: MLE-bench by OpenAI (https://arxiv.org/abs/2410.07095).
Then, if you used this page:
Can Agents Work. "MLE-bench: frontier results and sources." https://canagentswork.com/benchmarks/mle-bench/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-mle-bench,
title = {{MLE-bench: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/mle-bench/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}