Benchmarks / BEHAVIOR Challenge 2025
BEHAVIOR Challenge 2025 (BEHAVIOR 2025)
Built by Stanford University · released 2 Sep 2025
A NeurIPS 2025 robotics challenge from the Stanford Vision and Learning Lab. A simulated mobile manipulator (R1 Pro, in OmniGibson) must complete 50 full-length household activities from the BEHAVIOR-1K collection in house-scale scenes, using 10,000 teleoperated demonstrations (1,200+ hours) for training.
Frontier
0.26
Q-score (held-out test)
Robot Learning Collective
14 Nov 2025 · Source: Stanford University (benchmark maintainers)
First place, Standard track, affiliation "Independent". Held-out full task success rate 0.1240 (12.4% of episodes finished the whole chore). Public validation Q-score 0.2605. Code and report (arXiv 2512.06951) are linked from the board.
The best result is at 20–60% of the ceiling, or at 50–100% of human parity.
What it measures
How much of each household chore a policy finishes. Tasks include cooking (cook bacon, make pizza, chop an onion), cleaning (clean a patio, wash dog toys, clean up plates and food), tidying and storage (put shoes on a rack, box books, store food), and a few installation and outdoor tasks (hang pictures, chop wood, spray fruit trees). A task ends when its BDDL goal predicates are all true or a timeout of twice the average human time is reached. The ranking metric, Q-score, is the fraction of goal predicates satisfied at the end of the episode, averaged over the 50 tasks on a hidden test set that the organizers run. Full task success rate is shown alongside.
Q-score (held-out test): Fraction of BDDL goal predicates satisfied at episode end, averaged over 50 tasks on the hidden test set (0 to 1). Higher is better.
Status compares the frontier with a ceiling of 1.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- simulated-workplace
- Grading
- state-check
- Tasks
- 50
- Human reference
- Human demonstrations average 6.6 minutes per task (397 seconds). No human score is on the board; a human teleoperator satisfies all goal predicates by construction of the demonstrations.
- Contamination
- Training instances and 200 demonstrations per task are public. The test instances are hidden; teams may not collect data on evaluation instances.
- Reuse
- MIT (Hugging Face dataset card for behavior-1k/2025-challenge-demos). (open-mit)
Limits to keep in mind
- Simulation only (OmniGibson on Isaac Sim), with a single robot model. Results do not show transfer to a physical robot or a real kitchen. Source
- The board is labeled "Provisional 2025 Challenge Leaderboard" and the organizers plan to move it to Hugging Face with task-level statistics. Source
- The Standard track (robot sensors only) and the Privileged track (may query simulator state) appear on one board. The frontier is a Standard-track entry. Source
- The 2026 challenge is a new version: 100 tasks, 7 scenes, one track, 20,000 demonstrations, submissions due 2026-10-16 and winners on 2026-11-04. Its results will need a separate entry. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
| System | Q-score (held-out test) | Full task success rate | Date | Source |
|---|---|---|---|---|
| Comet (NVIDIA Research) | 0.25 | 11.4% | 17 Nov 2025 | Stanford University · primary |
| Robot Learning Collective · frontier | 0.26 | 12.4% | 14 Nov 2025 | Stanford University · primary |
Timeline
- 2 Sep 2025 — BEHAVIOR Challenge 2025 launches: 50 full-length household chores in simulation. Source
Where it sits in the atlas
Work ladder: Direct evidence for Food service, cleaning, and personal care. On the work ladder it counts as 12.4% of tasks fully completed (Full task success rate). How the ladder works
Work it measures (O*NET work activities): Prepare foods or beverages; Clean tools, equipment, facilities, or work areas; Clean workpieces, finished products, or other objects; Move materials, equipment, or supplies; Stock supplies or products.
Food service, cleaning, and personal careTransportation and material moving (partial)
Go to the source
- Website behavior.stanford.edu
- Paper arxiv.org
- Full leaderboard behavior.stanford.edu
- Code github.com
- Dataset huggingface.co
Last checked 24 Sep 2026 against 7 primary sources. See an error? Tell us.
How to cite
Credit the original work first: BEHAVIOR Challenge 2025 by Stanford University (https://arxiv.org/abs/2403.09227).
Then, if you used this page:
Can Agents Work. "BEHAVIOR Challenge 2025: frontier results and sources." https://canagentswork.com/benchmarks/behavior-challenge-2025/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-behavior-challenge-2025,
title = {{BEHAVIOR Challenge 2025: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/behavior-challenge-2025/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}