Tool Decathlon (Toolathlon)

The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution

108 manually sourced tasks that require an agent to drive 32 real software applications through 604 tools (mostly MCP servers) such as Google Calendar, Notion, Canvas, WooCommerce, Kubernetes and BigQuery, starting from realistic seeded environment states and taking roughly 20 turns each. Every task is graded by a dedicated execution-based evaluation script on the resulting state; the leaderboard reports Pass@1 over three runs plus Pass@3 and Pass^3. It is the broadest public test of general MCP-style tool use across heterogeneous apps.

78.4%
GLM 5.3 Flash (max)
official
2025-10-29 → 2026-08-30: 20.1% → 78.4%
Released
2025-10
Maintainer
HKUST NLP (Junlong Li, Junxian He et al.) with CMU / OpenHands collaborators
Status
active
Contamination
medium
Metric
Pass@1 (mean of 3 runs) (percent, ↑)
Tasks
108
Domains
tool-use general-assistant software-engineering
human
no measured baseline

Notes. Rows marked with a check on the leaderboard are run by the Toolathlon team in the Default agent configuration, three runs each, with +/- standard deviation shown. Tasks touch live services (Canvas, Notion, WooCommerce, Google APIs), so results can drift with the services themselves; the Verified release added bounded retries and per-instance isolation to reduce this. Task files and evaluators are public on GitHub.

Full ledger

SystemDeveloperScoreDateSourceConditions
GLM 5.3 Flash (max)Z.ai78.4%officialsplit: verified scaffold: Default reasoning_effort: max
Rank 1 Pass@1 as of access date; +/-1.9 over 3 runs; Pass@3 / Pass^3 not yet published.
Kimi K3 (max)Moonshot AI76.5%officialsplit: verified scaffold: Default reasoning_effort: max
Evaluated by the Toolathlon team; +/-1.9 over 3 runs; Pass@3 83.3, Pass^3 68.5 (highest Pass^3 on the board).
Claude Opus 4.8 (max)Anthropic76.2%officialsplit: verified scaffold: Default reasoning_effort: max
Evaluated by the Toolathlon team; +/-3.4 over 3 runs; Pass@3 84.3, Pass^3 66.7; 19.9 turns, 36.3 tool calls per task.
GPT-5.5 (xhigh)OpenAI73.5%officialsplit: verified scaffold: Default reasoning_effort: xhigh
Evaluated by the Toolathlon team; +/-1.2 over 3 runs; Pass@3 82.4, Pass^3 62.0.
Claude-4.5-SonnetAnthropic38.6%papersplit: original
Best model in the paper abstract (20.2 tool-calling turns on average); original October 2025 task release.
DeepSeek-V3.2-ExpDeepSeek20.1%papersplit: original
Top open-weights model in the paper abstract; original October 2025 task release.