AutomationBench

AutomationBench: cross-application business workflow orchestration via REST APIs

覆盖销售、市场、运营、客服、财务和 HR 的业务流程任务:智能体接收单条触发消息后,需在 47 个模拟 SaaS 应用的约 500 个 REST 端点中(BM25 检索)自行发现接口,遵循隐藏在环境中的政策文档,避开干扰记录,并修改模拟公司的状态。评分为确定性的最终状态断言(含负向断言);官方得分为所有断言全部通过的任务比例,在私有保留集(600+ 任务)上运行,另有 600 个公开任务供研究。

50.3%
Claude Opus 5 (max)
官方榜单
2026-04-21 → 2026-09-04: 9.6% → 50.3%
发布
2026-04
维护者
Zapier (Daniel Shepard, Robin Salimans)
状态
active
污染风险
low
指标
task pass rate (all assertions) (percent, ↑)
题量
600
领域
tool-use general-assistant instruction-following
人类
无实测基线

备注. Tasks were synthetically generated from the shape of real Zapier customer workflows and hardened; partial_credit (fraction of assertions) is reported as a diagnostic and RL reward but is not the headline score. Artificial Analysis runs an independent variant, AutomationBench-AA, whose headline is the share of objectives completed without guardrail violations, so AA numbers are on a different scale from Zapier's strict pass rate. Zapier's own leaderboard is a single run per model at its highest reasoning effort; the rank-1 Fable 5.1 entry uses an Opus 5 fallback for ~40% of tasks.

完整账本

系统开发者分数日期来源条件
Claude Opus 5 (max)Anthropic50.3%官方榜单split: public reasoning_effort: max
README table of pass rates on the 600-task public set; README undated, access date used. Public set is easier than the private leaderboard set.
GPT-6 AstraOpenAI41.4%厂商自报
OpenAI launch table (Professional section), maximum at any effort; set (public/private) and toolset not stated. Same table quotes Fable 5.1 at 31.4%, Opus 5 26.9%, GPT-5.6 Sol 18.1%. Not yet on Zapier's leaderboard as of access date.
Claude Fable 5.1 with Opus 5 fallback (max)Anthropic31.4%官方榜单split: private reasoning_effort: max cost_usd_per_task: 2.45
Rank 1 on leaderboard v1.0.6 as of access date; Opus 5 completed ~40% of tasks (260/657) after Fable safety refusals and cost excludes fallback tokens. Leaderboard undated, access date used.
Gemini 3.7 Flash (High)Google30.44%官方榜单split: private reasoning_effort: high cost_usd_per_task: 0.61
Rank 2 on leaderboard v1.0.6 as of access date; best standalone model. Leaderboard undated, access date used.
Opus 4.7 (max)Anthropic9.9%论文split: private reasoning_effort: max cost_usd_per_task: 1.8
Top of the leaderboard table in the white paper (API toolset, single run).
Gemini 3.1 Pro (high)Google9.6%论文split: private reasoning_effort: high cost_usd_per_task: 0.54
White paper leaderboard table (API toolset, single run).