Benchmarks / BEHAVIOR Challenge 2025

BEHAVIOR Challenge 2025 (BEHAVIOR 2025)

Built by Stanford University · released 2 Sep 2025

A NeurIPS 2025 robotics challenge from the Stanford Vision and Learning Lab. A simulated mobile manipulator (R1 Pro, in OmniGibson) must complete 50 full-length household activities from the BEHAVIOR-1K collection in house-scale scenes, using 10,000 teleoperated demonstrations (1,200+ hours) for training.

Frontier

0.26

Q-score (held-out test)

Robot Learning Collective

14 Nov 2025 · Source: Stanford University (benchmark maintainers)

First place, Standard track, affiliation "Independent". Held-out full task success rate 0.1240 (12.4% of episodes finished the whole chore). Public validation Q-score 0.2605. Code and report (arXiv 2512.06951) are linked from the board.

Emerging

The best result is at 20–60% of the ceiling, or at 50–100% of human parity.

See the full leaderboard at Stanford University

BEHAVIOR Challenge 2025: Q-score (held-out test) over time, 2 recorded results. 00.20.40.60.81Nov 2025Dec 2025 Robot Learning Collective: 0.26 (14 Nov 2025) Comet (NVIDIA Research): 0.25 (17 Nov 2025)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/behavior-challenge-2025/"><img src="https://canagentswork.com/og/benchmarks-behavior-challenge-2025.png" width="600" height="315" alt="BEHAVIOR Challenge 2025: the best result is 0.26 (Robot Learning Collective, 14 Nov 2025)." loading="lazy"></a>

Markdown:

[![BEHAVIOR Challenge 2025: the best result is 0.26 (Robot Learning Collective, 14 Nov 2025).](https://canagentswork.com/og/benchmarks-behavior-challenge-2025.png)](https://canagentswork.com/benchmarks/behavior-challenge-2025/)

What it measures

How much of each household chore a policy finishes. Tasks include cooking (cook bacon, make pizza, chop an onion), cleaning (clean a patio, wash dog toys, clean up plates and food), tidying and storage (put shoes on a rack, box books, store food), and a few installation and outdoor tasks (hang pictures, chop wood, spray fruit trees). A task ends when its BDDL goal predicates are all true or a timeout of twice the average human time is reached. The ranking metric, Q-score, is the fraction of goal predicates satisfied at the end of the episode, averaged over the 50 tasks on a hidden test set that the organizers run. Full task success rate is shown alongside.

Q-score (held-out test): Fraction of BDDL goal predicates satisfied at episode end, averaged over 50 tasks on the hidden test set (0 to 1). Higher is better.

Status compares the frontier with a ceiling of 1.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
simulated-workplace
Grading
state-check
Tasks
50
Human reference
Human demonstrations average 6.6 minutes per task (397 seconds). No human score is on the board; a human teleoperator satisfies all goal predicates by construction of the demonstrations.
Contamination
Training instances and 200 demonstrations per task are public. The test instances are hidden; teams may not collect data on evaluation instances.
Reuse
MIT (Hugging Face dataset card for behavior-1k/2025-challenge-demos). (open-mit)

Limits to keep in mind

  • Simulation only (OmniGibson on Isaac Sim), with a single robot model. Results do not show transfer to a physical robot or a real kitchen. Source
  • The board is labeled "Provisional 2025 Challenge Leaderboard" and the organizers plan to move it to Hugging Face with task-level statistics. Source
  • The Standard track (robot sensors only) and the Privileged track (may query simulator state) appear on one board. The frontier is a Standard-track entry. Source
  • The 2026 challenge is a new version: 100 tasks, 7 scenes, one track, 20,000 demonstrations, submissions due 2026-10-16 and winners on 2026-11-04. Its results will need a separate entry. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemQ-score (held-out test)Full task success rateDateSource
Comet (NVIDIA Research)0.2511.4%17 Nov 2025Stanford University · primary
Robot Learning Collective · frontier0.2612.4%14 Nov 2025Stanford University · primary

Timeline

  • 2 Sep 2025 — BEHAVIOR Challenge 2025 launches: 50 full-length household chores in simulation. Source

Where it sits in the atlas

Work ladder: Direct evidence for Food service, cleaning, and personal care. On the work ladder it counts as 12.4% of tasks fully completed (Full task success rate). How the ladder works

Work it measures (O*NET work activities): Prepare foods or beverages; Clean tools, equipment, facilities, or work areas; Clean workpieces, finished products, or other objects; Move materials, equipment, or supplies; Stock supplies or products.

Food service, cleaning, and personal careTransportation and material moving (partial)

Household choresRobot manipulationLong-horizon autonomy

Last checked 24 Sep 2026 against 7 primary sources. See an error? Tell us.

How to cite

Credit the original work first: BEHAVIOR Challenge 2025 by Stanford University (https://arxiv.org/abs/2403.09227).

Then, if you used this page:

Can Agents Work. "BEHAVIOR Challenge 2025: frontier results and sources." https://canagentswork.com/benchmarks/behavior-challenge-2025/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-behavior-challenge-2025,
  title        = {{BEHAVIOR Challenge 2025: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/behavior-challenge-2025/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}