Benchmarks / RoboChallenge Table30

RoboChallenge Table30 (Table30)

Built by Dexmal and Hugging Face · released 20 Oct 2025

A real-robot evaluation service. Teams connect their robot policy over the internet to a fleet of robot arms that the maintainers host, and human testers reset the scene from a reference image and grade each rollout. Table30 is the first task set: 30 tabletop tasks on four robot types (UR5, Franka Panda, Cobot Magic Aloha, ARX-5).

Frontier

64.3%

Success rate

Era0

May 2026 · Source: Dexmal (benchmark maintainers)

First on the v1 board: success_ratio 0.6433 (193 of 300 rollouts), progress score 76.34. Task-specific protocol, submitted by Robotera (user type organization). Per-task entries carry update times from 2026-03-24 to 2026-05-28, so the month is the best date we can give.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Dexmal

RoboChallenge Table30: Success rate over time, 3 recorded results. 0%20%40%60%80%100%Oct 2025Jan 2026Apr 2026 π0.5 (task-specific, RoboChallenge baseline): 43.7% (20 Oct 2025) DM0: 62% (May 2026) Era0: 64.3% (May 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/robochallenge-table30/"><img src="https://canagentswork.com/og/benchmarks-robochallenge-table30.png" width="600" height="315" alt="RoboChallenge Table30: the best result is 64.3% (Era0, May 2026)." loading="lazy"></a>

Markdown:

[![RoboChallenge Table30: the best result is 64.3% (Era0, May 2026).](https://canagentswork.com/og/benchmarks-robochallenge-table30.png)](https://canagentswork.com/benchmarks/robochallenge-table30/)

What it measures

The share of rollouts in which the robot completes a household or light-workplace tabletop task end to end. Tasks include wipe the table, fold a dishcloth, sweep rubbish into a dustpan, make a vegetarian sandwich, set plates on a rack, stack bowls, sort books, plug in a network cable, press three buttons in order, and scan a QR code. Each task is run 10 times, so the overall rate covers 300 rollouts. A progress score with partial credit is shown alongside but does not decide the rank.

Success rate: Share of successful rollouts across the 30 tasks (10 rollouts per task), as ranked on the Table30 v1 leaderboard. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
live-system
Grading
rubric
Tasks
30
Human reference
None published. The maintainers say the tasks should look trivial enough for a person to do without training.
Contamination
Teams fine-tune on up to 1,000 demonstrations per task that the maintainers provide. Reference episodes used to reset the scene are held out of training.
Reuse
No license on the Table30 dataset card (Hugging Face reports empty metadata). The inference client repo is CC BY-NC-SA 4.0. The report is under the arXiv non-exclusive license. (cite-only)

Limits to keep in mind

  • The board mixes two protocols. Most top entries are task-specific (one fine-tuned model per task); entries marked generalist use one model for all 30 tasks and score lower. Entries are submitted by outside teams, including individuals, and no model weights are public. Source
  • The maintainer, Dexmal, also submits its own models (DM0 is second on the v1 board). The platform is now run by a committee of partner organizations, but the report's authors are almost all Dexmal staff. Source
  • Real-robot results vary from run to run. The report shows success rates on the same task and model that changed with the human tester, which is why testers now reset scenes from a reference image. Source
  • Table30 V2 (April 2026) is a separate series with a new task list and a generalist-only protocol. Its top entry is far lower (40.67% success). The v1 board shows no new entries after May 2026. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemSuccess rateDateSource
DM062%May 2026Dexmal · primary
Era0 · frontier64.3%May 2026Dexmal · primary
π0.5 (task-specific, RoboChallenge baseline)
harness: RoboChallenge rc_baseline fine-tune
43.7%20 Oct 2025Dexmal · primary

Where sources disagree

The RoboChallenge report (October 2025) gives the task-specific π0.5 baseline an average success rate of 43.7% on Table30. The live v1 leaderboard shows the same baseline (pi0.5, rc_baseline) at 42.67%. The per-task entries on the board all carry a 2026-05-28 update time, so a re-run is the likely cause, but the maintainers have not said so.

  • 43.7 — arxiv.org (primary) · Figure 9 of the report, "average" row, Pi05 SR column.
  • 42.67 — robochallenge.ai (primary) · success_ratio 0.4267 for display_name pi0.5, user rc_baseline, read 2026-09-24.

We show 43.7. Status: open.

Timeline

  • May 2026 — Era0 (Robotera) reaches 64.33% success on RoboChallenge Table30. Source
  • Apr 2026 — RoboChallenge opens Table30 V2, a new 30-task series with a generalist-only protocol. Source
  • 20 Oct 2025 — RoboChallenge opens Table30: 30 real-robot tabletop tasks, best baseline 43.7% success. Source

Where it sits in the atlas

Work ladder: Mapped in part only. It gives context to the families below but is not direct evidence for any of them. How the ladder works

Work it measures (O*NET work activities): Clean tools, equipment, facilities, or work areas; Prepare foods or beverages; Move materials, equipment, or supplies; Sort materials or products.

Food service, cleaning, and personal care (partial)Production and manufacturing (partial)

Robot manipulationHousehold chores

Last checked 24 Sep 2026 against 13 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: RoboChallenge Table30 by Dexmal and Hugging Face (https://arxiv.org/abs/2510.17950).

Then, if you used this page:

Can Agents Work. "RoboChallenge Table30: frontier results and sources." https://canagentswork.com/benchmarks/robochallenge-table30/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-robochallenge-table30,
  title        = {{RoboChallenge Table30: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/robochallenge-table30/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}