Benchmarks / RoboArena

RoboArena

Built by UC Berkeley and Stanford University · Pranav Atreya, Karl Pertsch, Tony Lee · released 22 Jun 2025

A distributed real-robot evaluation of generalist manipulation policies (CoRL 2025). Evaluators at universities and companies run two policies back to back on the same Franka arm (DROID setup), same scene, and same instruction, without knowing which policy is which, and record which one did better. Pairwise preferences from many sites feed one ranking.

Frontier

1735 Elo

Leaderboard score (Elo-style)

DreamZero (dreaming_zebra)

24 Sep 2026 · Source: UC Berkeley (benchmark maintainers)

First on the official board (100 or more A/B evaluations): score 1735, standard deviation 42.6, 190 evaluations; the policy server was offline (0% uptime over 30 days) when read. The site links the DreamZero paper (arXiv 2602.15922). In the "All policies" view, j2-vla (Spirit v1.6) shows 1788 with a standard deviation of 106 from 25 evaluations; it is below the official threshold and within noise of DreamZero, so we do not show it as the frontier.

Unrated

No fixed reference point (for example Elo scores or field signals).

See the full leaderboard at UC Berkeley

RoboArena: Leaderboard score (Elo-style) over time, 2 recorded results. 1500 Elo1600 Elo1700 Elo1800 Elo1900 Elo2000 EloOct 2026 π0.5 DROID (pi05_droid): 1608 Elo (24 Sep 2026) DreamZero (dreaming_zebra): 1735 Elo (24 Sep 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/roboarena/"><img src="https://canagentswork.com/og/benchmarks-roboarena.png" width="600" height="315" alt="RoboArena: the best result is 1735 Elo (DreamZero (dreaming_zebra), 24 Sep 2026)." loading="lazy"></a>

Markdown:

[![RoboArena: the best result is 1735 Elo (DreamZero (dreaming_zebra), 24 Sep 2026).](https://canagentswork.com/og/benchmarks-roboarena.png)](https://canagentswork.com/benchmarks/roboarena/)

What it measures

How often a policy beats other policies on tabletop tasks that evaluators choose, expressed as an Elo-style score. Instructions in the public log are short manipulation tasks such as "put the apple into the bowl", "stack up the wooden blocks", "cover the controller with the towel", "rotate the pen by 90 degrees", and "press the yellow button". Evaluators also give a 0 to 100 progress score and a written explanation. As of 2026-09-24 the public log holds 3,994 A/B evaluations from 28 evaluator organizations.

Leaderboard score (Elo-style): Rating fitted from double-blind pairwise preferences across evaluator-chosen tasks. The official board lists policies with 100 or more A/B evaluations. Higher is better.

No fixed reference point, so we do not rate its status. Relative rating with no ceiling or human line. The scale is recomputed and has shifted (top score 2022 on 2026-08-24, 1788 a week later).

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
live-system
Grading
pairwise-human, human-expert
Reuse
MIT (Hugging Face card for RoboArena/DataDump_07-17-2026). Paper CC BY 4.0 on arXiv. (open-mit)

Limits to keep in mind

  • Tasks are open-ended lab tabletop tasks chosen by evaluators (pick up a doll, put an apple in a bowl). They are manipulation primitives, not the work of a specific occupation, so we map them as partial evidence only. Source
  • The site's default "Official leaderboard" shows only policies with 100 or more A/B evaluations. The "All policies" view includes early results with higher uncertainty; its top entry, j2-vla (shown as Spirit v1.6), has 1788 with a standard deviation of 106 from 25 evaluations and is not in the official set. Source
  • One evaluator organization (FrodoBots) contributed 2,301 of the 3,994 public A/B evaluations. The transparency page reports leave-one-organization-out reruns for this reason. Source
  • Scores are recomputed as data arrives, and the scale moved between weekly snapshots (DreamZero 1969 on 2026-08-24, 1736 on 2026-08-31). Compare ranks, not score values, across time. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemLeaderboard score (Elo-style)DateSource
π0.5 DROID (pi05_droid)1608 Elo24 Sep 2026UC Berkeley · primary
DreamZero (dreaming_zebra) · frontier1735 Elo24 Sep 2026UC Berkeley · primary

Timeline

  • 22 Jun 2025 — RoboArena launches distributed double-blind real-robot evaluation across seven academic institutions. Source

Where it sits in the atlas

Work ladder: Mapped in part only. It gives context to the families below but is not direct evidence for any of them. How the ladder works

Work it measures (O*NET work activities): Move materials, equipment, or supplies; Sort materials or products.

Production and manufacturing (partial)Food service, cleaning, and personal care (partial)

Robot manipulation

Last checked 24 Sep 2026 against 10 primary sources. See an error? Tell us.

How to cite

Credit the original work first: RoboArena by UC Berkeley and Stanford University (https://proceedings.mlr.press/v305/atreya25a.html).

Then, if you used this page:

Can Agents Work. "RoboArena: frontier results and sources." https://canagentswork.com/benchmarks/roboarena/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-roboarena,
  title        = {{RoboArena: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/roboarena/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}