Benchmarks / TutorMoments-Preview
TutorMoments-Preview (TutorMoments)
Built by Ai2 · Albert Zhang, Alexis Ross, Kajal Patel, Julian Bernado, et al. · released 7 Aug 2026
A replay test of whether a language model, acting as a math tutor, knows when to help and when to hold back. Experienced teachers marked decision points in real one-on-one tutoring transcripts with US students in grades 2 to 7. The model takes over as tutor at that point and talks with a simulated student for five turns. No tools: the model only writes tutoring turns.
Frontier
0.84
Mean of appropriate scaffolding and appropriate rigor (evaluation-aware prompt)
Claude Opus 4.8
7 Aug 2026 · Source: Ai2 (benchmark maintainers)
Mean of appropriate scaffolding 0.858 and appropriate rigor 0.831 (evaluation-aware prompt), computed by the atlas from Table 8. Leads the rigor column; avoids over-scaffolding 0.896. Plain prompt: 0.615, 0.208, 0.462. Results date is the Ai2 blog post; the working paper is dated 2026-06-30.
The best result is at 60–90% of the ceiling, or at or above human parity.
What it measures
Whether the model's move fits what teachers said the moment called for. 520 key moments are split evenly between moments that call for scaffolding (making the problem more accessible) and moments that call for a push for rigor (asking the student to do harder thinking). An LLM classifier, validated against teacher annotations, labels each replay. The paper reports three shares of moments: appropriate scaffolding, appropriate rigor, and avoiding over-scaffolding, under a plain prompt and an evaluation-aware prompt that spells out the trade-off. The atlas headline is the maintainers' overall measure: the mean of appropriate scaffolding and appropriate rigor with the evaluation-aware prompt.
Mean of appropriate scaffolding and appropriate rigor (evaluation-aware prompt): Average of two shares of the 520 moments: moments that called for scaffolding where the model scaffolded, and moments that called for rigor where the model pushed for rigor. The site plots this mean as "overall performance"; the atlas computes it from the two published columns. Higher is better.
Status compares the frontier with a ceiling of 1.
Facts
- Grain
- Task level: bounded tasks with a clear spec
- Environment
- chat
- Grading
- llm-judge
- Tasks
- 520
- Human reference
- The human tutors in the same transcripts, scored the same way at the same points, get 0.458 (appropriate scaffolding), 0.182 (appropriate rigor), and 0.496 (avoids over-scaffolding). The maintainers call this a naturalistic reference, not a ceiling: teachers deliberately flagged moments where tutoring could have gone better, so the set concentrates on missed opportunities. For that reason the atlas does not use the human line as a parity reference.
- Contamination
- The transcripts, key moments, and ground-truth labels are public. The maintainers do not keep a private set.
- Reuse
- CC-BY-4.0 on the Hugging Face dataset card (462 de-identified transcripts, annotations, and the 7,280 model replays). (open-cc-by)
Limits to keep in mind
- No tools and no deliverable. Each replay is five turns: three by the model tutor and two by a simulated student, so the benchmark scores tutoring judgment, not multi-step work. Source
- Scores measure tutor behavior, not learning. The student is a simulated "oracle" student, and the maintainers say automated evaluation cannot stand in for studies with real students. Source
- Rigor is noisier than scaffolding: the scoring pipeline detects rigor pushes less reliably, and the annotations hold fewer rigor moments (260) than scaffolding moments (738). Source
- Narrow scope: US elementary and middle-school math from one tutoring program, annotated by one pool of 27 teachers. Scores depend heavily on the prompt. Source
Recorded results
A selection that shows the frontier over time. The full leaderboard is at the source.
Timeline
- 7 Aug 2026 — Ai2 releases TutorMoments-Preview; Claude Opus 4.8 scores 0.831 on appropriate rigor. Source
Where it sits in the atlas
Work ladder: Context only. It tests how agents behave, not whether they complete work, so it does not set a step. How the ladder works
Go to the source
- Website tutormoments.allen.ai
- Paper tutormoments.allen.ai
- Full leaderboard tutormoments.allen.ai
- Code github.com
- Announcement allenai.org
- Dataset huggingface.co
Last checked 24 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.
How to cite
Credit the original work first: TutorMoments-Preview by Ai2 (https://tutormoments.allen.ai/static/paper/tutormoments-preview.pdf).
Then, if you used this page:
Can Agents Work. "TutorMoments-Preview: frontier results and sources." https://canagentswork.com/benchmarks/tutormoments-preview/ (accessed 2026-09-24). CC BY 4.0.@misc{caw-tutormoments-preview,
title = {{TutorMoments-Preview: frontier results and sources}},
author = {{Can Agents Work}},
year = {2026},
howpublished = {\url{https://canagentswork.com/benchmarks/tutormoments-preview/}},
note = {Accessed 2026-09-24. CC BY 4.0}
}