AutomationBench
AutomationBench: cross-application business workflow orchestration via REST APIs
Business-workflow tasks across Sales, Marketing, Operations, Support, Finance and HR in which an agent, given one trigger message, must discover the right REST endpoints (BM25 search over ~500 endpoint schemas from 47 simulated SaaS apps), follow policy documents hidden in the environment, avoid decoy records, and mutate the state of a simulated company. Grading is deterministic end-state assertions, including negative ones; the official score is the strict fraction of tasks with every assertion passing on a private held-out set (600+ tasks); a 600-task public set is released for research.
paper · website · leaderboard · dataset · code
- Released
- 2026-04
- Maintainer
- Zapier (Daniel Shepard, Robin Salimans)
- Status
- active
- Contamination
- low
- Metric
- task pass rate (all assertions) (percent, ↑)
- Tasks
- 600
- Domains
- tool-use general-assistant instruction-following
- human
- no measured baseline
Notes. Tasks were synthetically generated from the shape of real Zapier customer workflows and hardened; partial_credit (fraction of assertions) is reported as a diagnostic and RL reward but is not the headline score. Artificial Analysis runs an independent variant, AutomationBench-AA, whose headline is the share of objectives completed without guardrail violations, so AA numbers are on a different scale from Zapier's strict pass rate. Zapier's own leaderboard is a single run per model at its highest reasoning effort; the rank-1 Fable 5.1 entry uses an Opus 5 fallback for ~40% of tasks.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Claude Opus 5 (max) | Anthropic | 50.3% | official | split: public reasoning_effort: max README table of pass rates on the 600-task public set; README undated, access date used. Public set is easier than the private leaderboard set. | |
| GPT-6 Astra | OpenAI | 41.4% | self-reported | OpenAI launch table (Professional section), maximum at any effort; set (public/private) and toolset not stated. Same table quotes Fable 5.1 at 31.4%, Opus 5 26.9%, GPT-5.6 Sol 18.1%. Not yet on Zapier's leaderboard as of access date. | |
| Claude Fable 5.1 with Opus 5 fallback (max) | Anthropic | 31.4% | official | split: private reasoning_effort: max cost_usd_per_task: 2.45 Rank 1 on leaderboard v1.0.6 as of access date; Opus 5 completed ~40% of tasks (260/657) after Fable safety refusals and cost excludes fallback tokens. Leaderboard undated, access date used. | |
| Gemini 3.7 Flash (High) | 30.44% | official | split: private reasoning_effort: high cost_usd_per_task: 0.61 Rank 2 on leaderboard v1.0.6 as of access date; best standalone model. Leaderboard undated, access date used. | ||
| Opus 4.7 (max) | Anthropic | 9.9% | paper | split: private reasoning_effort: max cost_usd_per_task: 1.8 Top of the leaderboard table in the white paper (API toolset, single run). | |
| Gemini 3.1 Pro (high) | 9.6% | paper | split: private reasoning_effort: high cost_usd_per_task: 0.54 White paper leaderboard table (API toolset, single run). |