Benchmarks / TutorMoments-Preview

TutorMoments-Preview (TutorMoments)

Built by Ai2 · Albert Zhang, Alexis Ross, Kajal Patel, Julian Bernado, et al. · released 7 Aug 2026

A replay test of whether a language model, acting as a math tutor, knows when to help and when to hold back. Experienced teachers marked decision points in real one-on-one tutoring transcripts with US students in grades 2 to 7. The model takes over as tutor at that point and talks with a simulated student for five turns. No tools: the model only writes tutoring turns.

Frontier

0.84

Mean of appropriate scaffolding and appropriate rigor (evaluation-aware prompt)

Claude Opus 4.8

7 Aug 2026 · Source: Ai2 (benchmark maintainers)

Mean of appropriate scaffolding 0.858 and appropriate rigor 0.831 (evaluation-aware prompt), computed by the atlas from Table 8. Leads the rigor column; avoids over-scaffolding 0.896. Plain prompt: 0.615, 0.208, 0.462. Results date is the Ai2 blog post; the working paper is dated 2026-06-30.

Strong

The best result is at 60–90% of the ceiling, or at or above human parity.

See the full leaderboard at Ai2

TutorMoments-Preview: Mean of appropriate scaffolding and appropriate rigor (evaluation-aware prompt) over time, 3 recorded results. 00.20.40.60.81Jul 2026Jul 2026Aug 2026Aug 2026Aug 2026Aug 2026 Claude Opus 4.8: 0.84 (7 Aug 2026) Claude Sonnet 4.6: 0.81 (7 Aug 2026) Gemini 2.5 Pro: 0.8 (7 Aug 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/tutormoments-preview/"><img src="https://canagentswork.com/og/benchmarks-tutormoments-preview.png" width="600" height="315" alt="TutorMoments-Preview: the best result is 0.84 (Claude Opus 4.8, 7 Aug 2026)." loading="lazy"></a>

Markdown:

[![TutorMoments-Preview: the best result is 0.84 (Claude Opus 4.8, 7 Aug 2026).](https://canagentswork.com/og/benchmarks-tutormoments-preview.png)](https://canagentswork.com/benchmarks/tutormoments-preview/)

What it measures

Whether the model's move fits what teachers said the moment called for. 520 key moments are split evenly between moments that call for scaffolding (making the problem more accessible) and moments that call for a push for rigor (asking the student to do harder thinking). An LLM classifier, validated against teacher annotations, labels each replay. The paper reports three shares of moments: appropriate scaffolding, appropriate rigor, and avoiding over-scaffolding, under a plain prompt and an evaluation-aware prompt that spells out the trade-off. The atlas headline is the maintainers' overall measure: the mean of appropriate scaffolding and appropriate rigor with the evaluation-aware prompt.

Mean of appropriate scaffolding and appropriate rigor (evaluation-aware prompt): Average of two shares of the 520 moments: moments that called for scaffolding where the model scaffolded, and moments that called for rigor where the model pushed for rigor. The site plots this mean as "overall performance"; the atlas computes it from the two published columns. Higher is better.

Status compares the frontier with a ceiling of 1.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
chat
Grading
llm-judge
Tasks
520
Human reference
The human tutors in the same transcripts, scored the same way at the same points, get 0.458 (appropriate scaffolding), 0.182 (appropriate rigor), and 0.496 (avoids over-scaffolding). The maintainers call this a naturalistic reference, not a ceiling: teachers deliberately flagged moments where tutoring could have gone better, so the set concentrates on missed opportunities. For that reason the atlas does not use the human line as a parity reference.
Contamination
The transcripts, key moments, and ground-truth labels are public. The maintainers do not keep a private set.
Reuse
CC-BY-4.0 on the Hugging Face dataset card (462 de-identified transcripts, annotations, and the 7,280 model replays). (open-cc-by)

Limits to keep in mind

  • No tools and no deliverable. Each replay is five turns: three by the model tutor and two by a simulated student, so the benchmark scores tutoring judgment, not multi-step work. Source
  • Scores measure tutor behavior, not learning. The student is a simulated "oracle" student, and the maintainers say automated evaluation cannot stand in for studies with real students. Source
  • Rigor is noisier than scaffolding: the scoring pipeline detects rigor pushes less reliably, and the annotations hold fewer rigor moments (260) than scaffolding moments (738). Source
  • Narrow scope: US elementary and middle-school math from one tutoring program, annotated by one pool of 27 teachers. Scores depend heavily on the prompt. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemMean of appropriate scaffolding and appropriate rigor (evaluation-aware prompt)DateSource
Claude Opus 4.8 · frontier0.847 Aug 2026Ai2 · primary
Claude Sonnet 4.60.817 Aug 2026Ai2 · primary
Gemini 2.5 Pro0.87 Aug 2026Ai2 · primary

Timeline

  • 7 Aug 2026 — Ai2 releases TutorMoments-Preview; Claude Opus 4.8 scores 0.831 on appropriate rigor. Source

Where it sits in the atlas

Work ladder: Context only. It tests how agents behave, not whether they complete work, so it does not set a step. How the ladder works

Education and social services

Tutoring

Last checked 24 Sep 2026 against 5 primary sources, with a second independent check. See an error? Tell us.

How to cite

Credit the original work first: TutorMoments-Preview by Ai2 (https://tutormoments.allen.ai/static/paper/tutormoments-preview.pdf).

Then, if you used this page:

Can Agents Work. "TutorMoments-Preview: frontier results and sources." https://canagentswork.com/benchmarks/tutormoments-preview/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-tutormoments-preview,
  title        = {{TutorMoments-Preview: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/tutormoments-preview/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}