Tool Decathlon (Toolathlon)
The Tool Decathlon: Benchmarking Language Agents for Diverse, Realistic, and Long-Horizon Task Execution
108 manually sourced tasks that require an agent to drive 32 real software applications through 604 tools (mostly MCP servers) such as Google Calendar, Notion, Canvas, WooCommerce, Kubernetes and BigQuery, starting from realistic seeded environment states and taking roughly 20 turns each. Every task is graded by a dedicated execution-based evaluation script on the resulting state; the leaderboard reports Pass@1 over three runs plus Pass@3 and Pass^3. It is the broadest public test of general MCP-style tool use across heterogeneous apps.
paper · website · leaderboard · dataset · code
- Released
- 2025-10
- Maintainer
- HKUST NLP (Junlong Li, Junxian He et al.) with CMU / OpenHands collaborators
- Status
- active
- Contamination
- medium
- Metric
- Pass@1 (mean of 3 runs) (percent, ↑)
- Tasks
- 108
- Domains
- tool-use general-assistant software-engineering
- human
- no measured baseline
Notes. Rows marked with a check on the leaderboard are run by the Toolathlon team in the Default agent configuration, three runs each, with +/- standard deviation shown. Tasks touch live services (Canvas, Notion, WooCommerce, Google APIs), so results can drift with the services themselves; the Verified release added bounded retries and per-instance isolation to reduce this. Task files and evaluators are public on GitHub.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GLM 5.3 Flash (max) | Z.ai | 78.4% | official | split: verified scaffold: Default reasoning_effort: max Rank 1 Pass@1 as of access date; +/-1.9 over 3 runs; Pass@3 / Pass^3 not yet published. | |
| Kimi K3 (max) | Moonshot AI | 76.5% | official | split: verified scaffold: Default reasoning_effort: max Evaluated by the Toolathlon team; +/-1.9 over 3 runs; Pass@3 83.3, Pass^3 68.5 (highest Pass^3 on the board). | |
| Claude Opus 4.8 (max) | Anthropic | 76.2% | official | split: verified scaffold: Default reasoning_effort: max Evaluated by the Toolathlon team; +/-3.4 over 3 runs; Pass@3 84.3, Pass^3 66.7; 19.9 turns, 36.3 tool calls per task. | |
| GPT-5.5 (xhigh) | OpenAI | 73.5% | official | split: verified scaffold: Default reasoning_effort: xhigh Evaluated by the Toolathlon team; +/-1.2 over 3 runs; Pass@3 82.4, Pass^3 62.0. | |
| Claude-4.5-Sonnet | Anthropic | 38.6% | paper | split: original Best model in the paper abstract (20.2 tool-calling turns on average); original October 2025 task release. | |
| DeepSeek-V3.2-Exp | DeepSeek | 20.1% | paper | split: original Top open-weights model in the paper abstract; original October 2025 task release. |