tau2-bench

tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment

↑ tau-bench

Successor to tau-bench that adds a telecom troubleshooting domain (114 tasks) in which both the agent and the simulated user have tools, so the agent must coordinate and guide the user rather than act alone; the verified retail (115) and airline (50) domains are retained. Tasks are generated compositionally from atomic sub-tasks and scored on final database state; pass^k measures reliability across repeated trials. Now maintained as tau^3-bench with a banking knowledge domain and a voice modality.

87.9%
Qwen3.5-397B-A17B
official
2025-06-09 → 2026-09-04: 34% → 87.9%
Released
2025-06
Maintainer
Sierra Research (Barres et al.)
Status
active
Contamination
medium
Metric
pass^1 (percent, ↑)
Tasks
279
Domains
tool-use instruction-following general-assistant
human
no measured baseline

Notes. Task fixes in February 2026 (tau^3-bench v1.0) and a July 2026 grading update mean pre- and post-fix scores are not strictly comparable; taubench.com re-graded affected submissions. Domains, tools and policies are public, but tasks are compositional and the user simulator adds variance.

Full ledger

SystemDeveloperScoreDateSourceConditions
Qwen3.5-397B-A17BAlibaba Cloud87.9%officialsplit: core
Top of the taubench.com tau2-bench text leaderboard as of access date. No per-row date; date is the day observed.
Gemini 3.0 ProGoogle DeepMind85.4%officialsplit: core
taubench.com tau2-bench text leaderboard. Site shows no per-row date; date is the day observed.
Claude Opus 4.5Anthropic85.3%officialsplit: core
taubench.com tau2-bench text leaderboard (retail, airline, telecom). Site shows no per-row date; date is the day observed.
claude-3.7-sonnetAnthropic49%papersplit: telecom scaffold: function calling
Paper abstract (pass^1 on the new telecom tasks).
gpt-4.1OpenAI34%papersplit: telecom scaffold: function calling
Paper abstract/Section 4.2 (pass^1 on the new telecom tasks).