Tool Decathlon (Toolathlon)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
108 个人工采集或构造的任务,要求智能体通过 604 个工具(主要为 MCP server)操作 32 个真实软件应用,如 Google Calendar、Notion、Canvas、WooCommerce、Kubernetes 和 BigQuery,从真实的初始环境状态出发,平均约 20 轮交互。每个任务由专用的基于执行的评测脚本对最终状态评分;榜单报告三次运行的 Pass@1 以及 Pass@3 和 Pass^3。是目前覆盖最广的通用 MCP 风格跨应用工具使用公开测试。
- 发布
- 2025-10
- 维护者
- HKUST NLP (Junlong Li, Junxian He et al.) with CMU / OpenHands collaborators
- 状态
- active
- 污染风险
- medium
- 指标
- Pass@1 (mean of 3 runs) (percent, ↑)
- 题量
- 108
- 领域
- tool-use general-assistant software-engineering
- 人类
- 无实测基线
备注. Rows marked with a check on the leaderboard are run by the Toolathlon team in the Default agent configuration, three runs each, with +/- standard deviation shown. Tasks touch live services (Canvas, Notion, WooCommerce, Google APIs), so results can drift with the services themselves; the Verified release added bounded retries and per-instance isolation to reduce this. Task files and evaluators are public on GitHub.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GLM 5.3 Flash (max) | Z.ai | 78.4% | 官方榜单 | split: verified scaffold: Default reasoning_effort: max Rank 1 Pass@1 as of access date; +/-1.9 over 3 runs; Pass@3 / Pass^3 not yet published. | |
| Kimi K3 (max) | Moonshot AI | 76.5% | 官方榜单 | split: verified scaffold: Default reasoning_effort: max Evaluated by the Toolathlon team; +/-1.9 over 3 runs; Pass@3 83.3, Pass^3 68.5 (highest Pass^3 on the board). | |
| Claude Opus 4.8 (max) | Anthropic | 76.2% | 官方榜单 | split: verified scaffold: Default reasoning_effort: max Evaluated by the Toolathlon team; +/-3.4 over 3 runs; Pass@3 84.3, Pass^3 66.7; 19.9 turns, 36.3 tool calls per task. | |
| GPT-5.5 (xhigh) | OpenAI | 73.5% | 官方榜单 | split: verified scaffold: Default reasoning_effort: xhigh Evaluated by the Toolathlon team; +/-1.2 over 3 runs; Pass@3 82.4, Pass^3 62.0. | |
| Claude-4.5-Sonnet | Anthropic | 38.6% | 论文 | split: original Best model in the paper abstract (20.2 tool-calling turns on average); original October 2025 task release. | |
| DeepSeek-V3.2-Exp | DeepSeek | 20.1% | 论文 | split: original Top open-weights model in the paper abstract; original October 2025 task release. |