Benchmarks / CodeClash
CodeClash
Built by Stanford University and Princeton University · John Yang, Kilian Lieret, Joyce Yang, Carlos E. Jimenez, et al. · released 2 Nov 2025
Models compete in multi-round tournaments to build the best codebase for a goal, not a task. Each round has an edit phase, where the agent improves its code however it likes, and a competition phase, where the codebases fight in an arena such as a poker bot, a robot battle, or a territory game. The model that wins the most rounds wins the tournament.
Frontier
1385 Elo
Elo (all arenas)
Claude Sonnet 4.5 · harness: mini-SWE-agent
3 Nov 2025 · Source: Stanford University (benchmark maintainers)
Rank 1 on the site leaderboard (1385 ± 18, updated 2025-11-03). Paper v2 reports 1389 ± 18 and a 69.9% average tournament win rate. The leaderboard has eight models and no entries after launch.
No fixed reference point (for example Elo scores or field signals).
What it measures
Strength as an Elo rating fitted to tournament win rates across six arenas (BattleSnake, Core War, Halite, Poker, RoboCode, RobotRumble). The paper ran 1,680 tournaments of 15 rounds with 8 models in the mini-SWE-agent scaffold. Elo has a base of 1200 and is fitted by maximum likelihood with bootstrapped uncertainties.
Elo (all arenas): Overall Elo rating across the six arenas, base 1200, with a bootstrapped uncertainty. Higher is better.
No fixed reference point, so we do not rate its status. Elo is relative to the other models in the pool, so there is no fixed ceiling.
Facts
- Grain
- Project level: whole projects judged by an acceptance standard
- Environment
- repo, cli
- Grading
- outcome-metric
- Human reference
- A top open-source human RobotRumble bot (gigachad) beat Claude Sonnet 4.5 in 10 of 10 tournaments and 150 of 150 rounds. The paper says top models lose every round against expert human programmers.
- Contamination
- Low relevance. There is no fixed answer to memorize; the arenas are public games, and models may have seen strategies or bots for them in training.
- Reuse
- MIT (CodeClash repository). Arenas are third-party games with their own licenses. (open-mit)
Limits to keep in mind
- Elo is relative to the eight models in the pool. Adding or removing models shifts every rating, and the leaderboard has not added models since November 2025. Source
- Game arenas stand in for business goals. Winning at poker or robot battles is not the same as improving retention or revenue. Source
- All leaderboard results use one scaffold (mini-SWE-agent) with fixed settings, so scores reflect the model inside that harness. Source
- The maintainers report that model codebases grow messy and redundant over rounds and that models analyze competition logs only shallowly. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Elo (all arenas) | Date | Source |
|---|---|---|---|
| o3 harness: mini-SWE-agent | 1343 Elo | 3 Nov 2025 | Stanford University · primary |
| GPT-5 harness: mini-SWE-agent | 1366 Elo | 3 Nov 2025 | Stanford University · primary |
| Claude Sonnet 4.5 · frontier harness: mini-SWE-agent | 1385 Elo | 3 Nov 2025 | Stanford University · primary |
Where sources disagree
The CodeClash site leaderboard shows Claude Sonnet 4.5 at 1385 ± 18 and GPT-5 at 1366 ± 17. Version 2 of the paper (May 2026) shows 1389 ± 18 and 1360 ± 17 for the same models. Both come from the maintainers. The difference is small and the ranking is the same. A refit of the Elo model is one possible cause. We show the leaderboard value.
- 1385 — codeclash.ai (primary) · Site leaderboard, "Updated Nov. 3, 2025", seen 2026-09-23.
- 1389 — arxiv.org (primary) · Paper v2 (2026-05-12), Table 1 and Table 4.
We show 1385. Status: open.
Timeline
- 2 Nov 2025 — CodeClash launches goal-oriented coding tournaments. Source
Where it sits in the atlas
Go to the source
- Website codeclash.ai
- Paper arxiv.org
- Full leaderboard codeclash.ai
- Code github.com
Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.
How to cite
Credit the original work first: CodeClash by Stanford University and Princeton University (https://arxiv.org/abs/2511.00839).
Then, if you used this page:
Can Agents Work. "CodeClash: frontier results and sources." https://canagentswork.com/benchmarks/codeclash/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-codeclash,
title = {{CodeClash: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/codeclash/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}