AutomationBench

AutomationBench: cross-application business workflow orchestration via REST APIs

Business-workflow tasks across Sales, Marketing, Operations, Support, Finance and HR in which an agent, given one trigger message, must discover the right REST endpoints (BM25 search over ~500 endpoint schemas from 47 simulated SaaS apps), follow policy documents hidden in the environment, avoid decoy records, and mutate the state of a simulated company. Grading is deterministic end-state assertions, including negative ones; the official score is the strict fraction of tasks with every assertion passing on a private held-out set (600+ tasks); a 600-task public set is released for research.

50.3%
Claude Opus 5 (max)
official
2026-04-21 → 2026-09-04: 9.6% → 50.3%
Released
2026-04
Maintainer
Zapier (Daniel Shepard, Robin Salimans)
Status
active
Contamination
low
Metric
task pass rate (all assertions) (percent, ↑)
Tasks
600
Domains
tool-use general-assistant instruction-following
human
no measured baseline

Notes. Tasks were synthetically generated from the shape of real Zapier customer workflows and hardened; partial_credit (fraction of assertions) is reported as a diagnostic and RL reward but is not the headline score. Artificial Analysis runs an independent variant, AutomationBench-AA, whose headline is the share of objectives completed without guardrail violations, so AA numbers are on a different scale from Zapier's strict pass rate. Zapier's own leaderboard is a single run per model at its highest reasoning effort; the rank-1 Fable 5.1 entry uses an Opus 5 fallback for ~40% of tasks.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude Opus 5 (max)Anthropic50.3%officialsplit: public reasoning_effort: max
README table of pass rates on the 600-task public set; README undated, access date used. Public set is easier than the private leaderboard set.
GPT-6 AstraOpenAI41.4%self-reported
OpenAI launch table (Professional section), maximum at any effort; set (public/private) and toolset not stated. Same table quotes Fable 5.1 at 31.4%, Opus 5 26.9%, GPT-5.6 Sol 18.1%. Not yet on Zapier's leaderboard as of access date.
Claude Fable 5.1 with Opus 5 fallback (max)Anthropic31.4%officialsplit: private reasoning_effort: max cost_usd_per_task: 2.45
Rank 1 on leaderboard v1.0.6 as of access date; Opus 5 completed ~40% of tasks (260/657) after Fable safety refusals and cost excludes fallback tokens. Leaderboard undated, access date used.
Gemini 3.7 Flash (High)Google30.44%officialsplit: private reasoning_effort: high cost_usd_per_task: 0.61
Rank 2 on leaderboard v1.0.6 as of access date; best standalone model. Leaderboard undated, access date used.
Opus 4.7 (max)Anthropic9.9%papersplit: private reasoning_effort: max cost_usd_per_task: 1.8
Top of the leaderboard table in the white paper (API toolset, single run).
Gemini 3.1 Pro (high)Google9.6%papersplit: private reasoning_effort: high cost_usd_per_task: 0.54
White paper leaderboard table (API toolset, single run).