Benchmarks / RoboArena
RoboArena
Built by UC Berkeley and Stanford University · Pranav Atreya, Karl Pertsch, Tony Lee · released 22 Jun 2025
A distributed real-robot evaluation of generalist manipulation policies (CoRL 2025). Evaluators at universities and companies run two policies back to back on the same Franka arm (DROID setup), same scene, and same instruction, without knowing which policy is which, and record which one did better. Pairwise preferences from many sites feed one ranking.
Frontier
1735 Elo
Leaderboard score (Elo-style)
DreamZero (dreaming_zebra)
24 Sep 2026 · Source: UC Berkeley (benchmark maintainers)
First on the official board (100 or more A/B evaluations): score 1735, standard deviation 42.6, 190 evaluations; the policy server was offline (0% uptime over 30 days) when read. The site links the DreamZero paper (arXiv 2602.15922). In the "All policies" view, j2-vla (Spirit v1.6) shows 1788 with a standard deviation of 106 from 25 evaluations; it is below the official threshold and within noise of DreamZero, so we do not show it as the frontier.
No fixed reference point (for example Elo scores or field signals).
What it measures
How often a policy beats other policies on tabletop tasks that evaluators choose, expressed as an Elo-style score. Instructions in the public log are short manipulation tasks such as "put the apple into the bowl", "stack up the wooden blocks", "cover the controller with the towel", "rotate the pen by 90 degrees", and "press the yellow button". Evaluators also give a 0 to 100 progress score and a written explanation. As of 2026-09-24 the public log holds 3,994 A/B evaluations from 28 evaluator organizations.
Leaderboard score (Elo-style): Rating fitted from double-blind pairwise preferences across evaluator-chosen tasks. The official board lists policies with 100 or more A/B evaluations. Higher is better.
No fixed reference point, so we do not rate its status. Relative rating with no ceiling or human line. The scale is recomputed and has shifted (top score 2022 on 2026-08-24, 1788 a week later).
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- live-system
- Grading
- pairwise-human, human-expert
- Reuse
- MIT (Hugging Face card for RoboArena/DataDump_07-17-2026). Paper CC BY 4.0 on arXiv. (open-mit)
Limits to keep in mind
- Tasks are open-ended lab tabletop tasks chosen by evaluators (pick up a doll, put an apple in a bowl). They are manipulation primitives, not the work of a specific occupation, so we map them as partial evidence only. Source
- The site's default "Official leaderboard" shows only policies with 100 or more A/B evaluations. The "All policies" view includes early results with higher uncertainty; its top entry, j2-vla (shown as Spirit v1.6), has 1788 with a standard deviation of 106 from 25 evaluations and is not in the official set. Source
- One evaluator organization (FrodoBots) contributed 2,301 of the 3,994 public A/B evaluations. The transparency page reports leave-one-organization-out reruns for this reason. Source
- Scores are recomputed as data arrives, and the scale moved between weekly snapshots (DreamZero 1969 on 2026-08-24, 1736 on 2026-08-31). Compare ranks, not score values, across time. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Leaderboard score (Elo-style) | Date | Source |
|---|---|---|---|
| π0.5 DROID (pi05_droid) | 1608 Elo | 24 Sep 2026 | UC Berkeley · primary |
| DreamZero (dreaming_zebra) · frontier | 1735 Elo | 24 Sep 2026 | UC Berkeley · primary |
Timeline
- 22 Jun 2025 — RoboArena launches distributed double-blind real-robot evaluation across seven academic institutions. Source
Where it sits in the atlas
Work ladder: Mapped in part only. It gives context to the families below but is not direct evidence for any of them. How the ladder works
Work it measures (O*NET work activities): Move materials, equipment, or supplies; Sort materials or products.
Production and manufacturing (partial)Food service, cleaning, and personal care (partial)
Go to the source
- Website robo-arena.github.io
- Paper proceedings.mlr.press
- Full leaderboard roboarena-api-domain-name.online
- Code github.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 10 primary sources. See an error? Tell us.
How to cite
Credit the original work first: RoboArena by UC Berkeley and Stanford University (https://proceedings.mlr.press/v305/atreya25a.html).
Then, if you used this page:
Can Agents Work. "RoboArena: frontier results and sources." https://canagentswork.com/benchmarks/roboarena/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-roboarena,
title = {{RoboArena: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/roboarena/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}