tau2-bench
tau^2-Bench: Evaluating Conversational Agents in a Dual-Control Environment
tau-bench 的继任者,新增电信故障排查领域(114 个任务),其中智能体和模拟用户都拥有工具,因此智能体必须协调并引导用户而非独自行动;保留了经核验的零售(115)和航空(50)领域。任务由原子子任务组合生成,并按最终数据库状态评分;pass^k 衡量重复试验间的可靠性。现以 tau^3-bench 形式维护,增加了银行知识领域和语音模态。
- 发布
- 2025-06
- 维护者
- Sierra Research (Barres et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- pass^1 (percent, ↑)
- 题量
- 279
- 领域
- tool-use instruction-following general-assistant
- 人类
- 无实测基线
备注. Task fixes in February 2026 (tau^3-bench v1.0) and a July 2026 grading update mean pre- and post-fix scores are not strictly comparable; taubench.com re-graded affected submissions. Domains, tools and policies are public, but tasks are compositional and the user simulator adds variance.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Qwen3.5-397B-A17B | Alibaba Cloud | 87.9% | 官方榜单 | split: core Top of the taubench.com tau2-bench text leaderboard as of access date. No per-row date; date is the day observed. | |
| Gemini 3.0 Pro | Google DeepMind | 85.4% | 官方榜单 | split: core taubench.com tau2-bench text leaderboard. Site shows no per-row date; date is the day observed. | |
| Claude Opus 4.5 | Anthropic | 85.3% | 官方榜单 | split: core taubench.com tau2-bench text leaderboard (retail, airline, telecom). Site shows no per-row date; date is the day observed. | |
| claude-3.7-sonnet | Anthropic | 49% | 论文 | split: telecom scaffold: function calling Paper abstract (pass^1 on the new telecom tasks). | |
| gpt-4.1 | OpenAI | 34% | 论文 | split: telecom scaffold: function calling Paper abstract/Section 4.2 (pass^1 on the new telecom tasks). |