Tool Decathlon (Toolathlon)

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

108 个人工采集或构造的任务,要求智能体通过 604 个工具(主要为 MCP server)操作 32 个真实软件应用,如 Google Calendar、Notion、Canvas、WooCommerce、Kubernetes 和 BigQuery,从真实的初始环境状态出发,平均约 20 轮交互。每个任务由专用的基于执行的评测脚本对最终状态评分;榜单报告三次运行的 Pass@1 以及 Pass@3 和 Pass^3。是目前覆盖最广的通用 MCP 风格跨应用工具使用公开测试。

78.4%
GLM 5.3 Flash (max)
官方榜单
2025-10-29 → 2026-08-30: 20.1% → 78.4%
发布
2025-10
维护者
HKUST NLP (Junlong Li, Junxian He et al.) with CMU / OpenHands collaborators
状态
active
污染风险
medium
指标
Pass@1 (mean of 3 runs) (percent, ↑)
题量
108
领域
tool-use general-assistant software-engineering
人类
无实测基线

备注. Rows marked with a check on the leaderboard are run by the Toolathlon team in the Default agent configuration, three runs each, with +/- standard deviation shown. Tasks touch live services (Canvas, Notion, WooCommerce, Google APIs), so results can drift with the services themselves; the Verified release added bounded retries and per-instance isolation to reduce this. Task files and evaluators are public on GitHub.

完整账本

系统开发者分数日期来源条件
GLM 5.3 Flash (max)Z.ai78.4%官方榜单split: verified scaffold: Default reasoning_effort: max
Rank 1 Pass@1 as of access date; +/-1.9 over 3 runs; Pass@3 / Pass^3 not yet published.
Kimi K3 (max)Moonshot AI76.5%官方榜单split: verified scaffold: Default reasoning_effort: max
Evaluated by the Toolathlon team; +/-1.9 over 3 runs; Pass@3 83.3, Pass^3 68.5 (highest Pass^3 on the board).
Claude Opus 4.8 (max)Anthropic76.2%官方榜单split: verified scaffold: Default reasoning_effort: max
Evaluated by the Toolathlon team; +/-3.4 over 3 runs; Pass@3 84.3, Pass^3 66.7; 19.9 turns, 36.3 tool calls per task.
GPT-5.5 (xhigh)OpenAI73.5%官方榜单split: verified scaffold: Default reasoning_effort: xhigh
Evaluated by the Toolathlon team; +/-1.2 over 3 runs; Pass@3 82.4, Pass^3 62.0.
Claude-4.5-SonnetAnthropic38.6%论文split: original
Best model in the paper abstract (20.2 tool-calling turns on average); original October 2025 task release.
DeepSeek-V3.2-ExpDeepSeek20.1%论文split: original
Top open-weights model in the paper abstract; original October 2025 task release.