Benchmarks / CodeClash

CodeClash

Built by Stanford University and Princeton University · John Yang, Kilian Lieret, Joyce Yang, Carlos E. Jimenez, et al. · released 2 Nov 2025

Models compete in multi-round tournaments to build the best codebase for a goal, not a task. Each round has an edit phase, where the agent improves its code however it likes, and a competition phase, where the codebases fight in an arena such as a poker bot, a robot battle, or a territory game. The model that wins the most rounds wins the tournament.

Frontier

1385 Elo

Elo (all arenas)

Claude Sonnet 4.5 · harness: mini-SWE-agent

3 Nov 2025 · Source: Stanford University (benchmark maintainers)

Rank 1 on the site leaderboard (1385 ± 18, updated 2025-11-03). Paper v2 reports 1389 ± 18 and a 69.9% average tournament win rate. The leaderboard has eight models and no entries after launch.

Unrated

No fixed reference point (for example Elo scores or field signals).

See the full leaderboard at Stanford University

CodeClash: Elo (all arenas) over time, 3 recorded results. 1200 Elo1300 Elo1400 Elo1500 Elo1600 EloOct 2025Oct 2025Nov 2025Nov 2025Nov 2025 o3: 1343 Elo (3 Nov 2025) GPT-5: 1366 Elo (3 Nov 2025) Claude Sonnet 4.5: 1385 Elo (3 Nov 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/codeclash/"><img src="https://canagentswork.com/og/benchmarks-codeclash.png" width="600" height="315" alt="CodeClash: the best result is 1385 Elo (Claude Sonnet 4.5, 3 Nov 2025)." loading="lazy"></a>

Markdown:

[![CodeClash: the best result is 1385 Elo (Claude Sonnet 4.5, 3 Nov 2025).](https://canagentswork.com/og/benchmarks-codeclash.png)](https://canagentswork.com/benchmarks/codeclash/)

What it measures

Strength as an Elo rating fitted to tournament win rates across six arenas (BattleSnake, Core War, Halite, Poker, RoboCode, RobotRumble). The paper ran 1,680 tournaments of 15 rounds with 8 models in the mini-SWE-agent scaffold. Elo has a base of 1200 and is fitted by maximum likelihood with bootstrapped uncertainties.

Elo (all arenas): Overall Elo rating across the six arenas, base 1200, with a bootstrapped uncertainty. Higher is better.

No fixed reference point, so we do not rate its status. Elo is relative to the other models in the pool, so there is no fixed ceiling.

Facts

Grain
Project level: whole projects judged by an acceptance standard
Environment
repo, cli
Grading
outcome-metric
Human reference
A top open-source human RobotRumble bot (gigachad) beat Claude Sonnet 4.5 in 10 of 10 tournaments and 150 of 150 rounds. The paper says top models lose every round against expert human programmers.
Contamination
Low relevance. There is no fixed answer to memorize; the arenas are public games, and models may have seen strategies or bots for them in training.
Reuse
MIT (CodeClash repository). Arenas are third-party games with their own licenses. (open-mit)

Limits to keep in mind

  • Elo is relative to the eight models in the pool. Adding or removing models shifts every rating, and the leaderboard has not added models since November 2025. Source
  • Game arenas stand in for business goals. Winning at poker or robot battles is not the same as improving retention or revenue. Source
  • All leaderboard results use one scaffold (mini-SWE-agent) with fixed settings, so scores reflect the model inside that harness. Source
  • The maintainers report that model codebases grow messy and redundant over rounds and that models analyze competition logs only shallowly. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemElo (all arenas)DateSource
o3
harness: mini-SWE-agent
1343 Elo3 Nov 2025Stanford University · primary
GPT-5
harness: mini-SWE-agent
1366 Elo3 Nov 2025Stanford University · primary
Claude Sonnet 4.5 · frontier
harness: mini-SWE-agent
1385 Elo3 Nov 2025Stanford University · primary

Where sources disagree

The CodeClash site leaderboard shows Claude Sonnet 4.5 at 1385 ± 18 and GPT-5 at 1366 ± 17. Version 2 of the paper (May 2026) shows 1389 ± 18 and 1360 ± 17 for the same models. Both come from the maintainers. The difference is small and the ranking is the same. A refit of the Elo model is one possible cause. We show the leaderboard value.

  • 1385codeclash.ai (primary) · Site leaderboard, "Updated Nov. 3, 2025", seen 2026-09-23.
  • 1389arxiv.org (primary) · Paper v2 (2026-05-12), Table 1 and Table 4.

We show 1385. Status: open.

Timeline

  • 2 Nov 2025 — CodeClash launches goal-oriented coding tournaments. Source

Where it sits in the atlas

Software engineering

Codebase evolutionLong-horizon autonomy

Last checked 23 Sep 2026 against 6 primary sources. See an error? Tell us.

How to cite

Credit the original work first: CodeClash by Stanford University and Princeton University (https://arxiv.org/abs/2511.00839).

Then, if you used this page:

Can Agents Work. "CodeClash: frontier results and sources." https://canagentswork.com/benchmarks/codeclash/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-codeclash,
  title        = {{CodeClash: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/codeclash/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}