Benchmarks / τ²-bench

τ²-bench

Built by Sierra · Victor Barres, Honghua Dong, Soham Ray, Xujie Si, et al. · released 9 Jun 2025

Simulated customer-service conversations from Sierra. An agent chats with an LLM-played user, calls tools, and must follow a written domain policy. The telecom domain adds dual control: the user also has tools, so the agent must guide the user through steps that only the user can perform.

Frontier

97.8%

Telecom pass^1

Qwen3.5-397B-A17B (thinking) · harness: tau2-bench default agent

27 Feb 2026 · Source: Sierra (benchmark maintainers)

Best telecom pass^1 among runs that Sierra ran and verified with trajectories. User simulator gpt-5.2 (low reasoning), 4 trials, self-hosted with vLLM. pass^2 95.8, pass^3 93.9, pass^4 92.1. Same submission: airline 81.5, retail 84.4, banking knowledge 9.8 pass^1.

Saturated

The best result is at 90% or more of the ceiling.

See the full leaderboard at Sierra

τ²-bench: Telecom pass^1 over time, 5 recorded results. 0%20%40%60%80%100%Jul 2025Oct 2025Jan 2026 GPT-4.1: 34% (9 Jun 2025) o4-mini: 50.2% (9 Jun 2025) GPT-5: 95.8% (9 Aug 2025) Qwen3-Max-Thinking: 98.2% (23 Jan 2026) Qwen3.5-397B-A17B (thinking): 97.8% (27 Feb 2026)
Dots are recorded results; the line is the best result so far; the red dot is the current frontier. Sources for every point are in the table below.
Embed this card

The image names the institutions behind the numbers and links back to this page.

HTML:

<a href="https://canagentswork.com/benchmarks/tau2-bench/"><img src="https://canagentswork.com/og/benchmarks-tau2-bench.png" width="600" height="315" alt="τ²-bench: the best result is 97.8% (Qwen3.5-397B-A17B (thinking), 27 Feb 2026)." loading="lazy"></a>

Markdown:

[![τ²-bench: the best result is 97.8% (Qwen3.5-397B-A17B (thinking), 27 Feb 2026).](https://canagentswork.com/og/benchmarks-tau2-bench.png)](https://canagentswork.com/benchmarks/tau2-bench/)

What it measures

Whether the final database state and required outputs match the expected result for each task. Each task runs several times. pass^k is the chance that all k tries of a task succeed, so higher k rewards reliability. Domains are airline (50 tasks), retail (114), and telecom (114). We use telecom pass^1 as the headline because telecom is the domain that τ²-bench introduced. The repository now ships as τ³-bench with a banking knowledge domain and a voice mode; those are not scored here.

Telecom pass^1: Share of telecom tasks that the agent completes correctly on a single try, averaged over trials. The leaderboard reports pass^1 to pass^4 per domain. Higher is better.

Status compares the frontier with a ceiling of 100%.

Facts

Grain
Task level: bounded tasks with a clear spec
Environment
api-tools, chat
Grading
state-check
Tasks
278
Contamination
Tasks, policies, and tools are public in the repository. The leaderboard labels models trained on τ-bench domains or tasks as custom submissions, but it cannot detect training on the public data.
Reuse
MIT (repository license covers tasks, policies, and code). (open-mit)

Limits to keep in mind

  • The user is an LLM simulator. The choice of user model changes scores, and submissions use different user models (gpt-4.1, gpt-5.2, or the agent's own model). Sierra recommends gpt-5.2. Source
  • Telecom pass^1 is close to the ceiling: several models score 97-98%. pass^4 and the newer banking knowledge domain (best 55.2% pass^1) still separate models. Source
  • In February 2026 Sierra fixed 50+ airline and retail tasks. Airline pass^1 rose by 14 to 20 points after the fixes, so airline and retail scores from before and after are not comparable. Source
  • Some leaderboard entries were submitted by model developers with modified prompts and no trajectories. Sierra marks these as unverified in the submission files. Source

Recorded results

A selection that shows the frontier over time. The full leaderboard is at the source.

SystemTelecom pass^1DateSource
Qwen3.5-397B-A17B (thinking) · frontier
harness: tau2-bench default agent
97.8%27 Feb 2026Sierra · primary
Qwen3-Max-Thinking
harness: tau2-bench default agent
98.2%23 Jan 2026Qwen team (leaderboard submission) · lab-reported
GPT-5
harness: tau2-bench default agent
95.8%9 Aug 2025Sierra · primary
GPT-4.1
harness: tau2-bench default agent
34%9 Jun 2025Sierra · primary
o4-mini
harness: tau2-bench default agent
50.2%9 Jun 2025Sierra · primary

Timeline

  • Feb 2026 — Sierra fixes 50+ airline and retail tasks for τ³-bench. Source
  • 2 Oct 2025 — GPT-5 reaches 95.8% telecom pass^1 on τ²-bench. Source
  • 9 Jun 2025 — Sierra releases τ²-bench with a dual-control telecom domain. Source

Where it sits in the atlas

Customer support

Customer service

Last checked 23 Sep 2026 against 8 primary sources. See an error? Tell us.

How to cite

Credit the original work first: τ²-bench by Sierra (https://arxiv.org/abs/2506.07982).

Then, if you used this page:

Can Agents Work. "τ²-bench: frontier results and sources." https://canagentswork.com/benchmarks/tau2-bench/ (accessed 2026-09-24). CC BY 4.0.
@misc{caw-tau2-bench,
  title        = {{τ²-bench: frontier results and sources}},
  author       = {{Can Agents Work}},
  year         = {2026},
  howpublished = {\url{https://canagentswork.com/benchmarks/tau2-bench/}},
  note         = {Accessed 2026-09-24. CC BY 4.0}
}