tau2-bench

tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment

↑ tau-bench

tau-bench 的继任者,新增电信故障排查领域(114 个任务),其中智能体和模拟用户都拥有工具,因此智能体必须协调并引导用户而非独自行动;保留了经核验的零售(115)和航空(50)领域。任务由原子子任务组合生成,并按最终数据库状态评分;pass^k 衡量重复试验间的可靠性。现以 tau^3-bench 形式维护,增加了银行知识领域和语音模态。

87.9%
Qwen3.5-397B-A17B
官方榜单
2025-06-09 → 2026-09-04: 34% → 87.9%
发布
2025-06
维护者
Sierra Research (Barres et al.)
状态
active
污染风险
medium
指标
pass^1 (percent, ↑)
题量
279
领域
tool-use instruction-following general-assistant
人类
无实测基线

备注. Task fixes in February 2026 (tau^3-bench v1.0) and a July 2026 grading update mean pre- and post-fix scores are not strictly comparable; taubench.com re-graded affected submissions. Domains, tools and policies are public, but tasks are compositional and the user simulator adds variance.

完整账本

系统开发者分数日期来源条件
Qwen3.5-397B-A17BAlibaba Cloud87.9%官方榜单split: core
Top of the taubench.com tau2-bench text leaderboard as of access date. No per-row date; date is the day observed.
Gemini 3.0 ProGoogle DeepMind85.4%官方榜单split: core
taubench.com tau2-bench text leaderboard. Site shows no per-row date; date is the day observed.
Claude Opus 4.5Anthropic85.3%官方榜单split: core
taubench.com tau2-bench text leaderboard (retail, airline, telecom). Site shows no per-row date; date is the day observed.
claude-3.7-sonnetAnthropic49%论文split: telecom scaffold: function calling
Paper abstract (pass^1 on the new telecom tasks).
gpt-4.1OpenAI34%论文split: telecom scaffold: function calling
Paper abstract/Section 4.2 (pass^1 on the new telecom tasks).