AutomationBench
AutomationBench: cross-application business workflow orchestration via REST APIs
覆盖销售、市场、运营、客服、财务和 HR 的业务流程任务:智能体接收单条触发消息后,需在 47 个模拟 SaaS 应用的约 500 个 REST 端点中(BM25 检索)自行发现接口,遵循隐藏在环境中的政策文档,避开干扰记录,并修改模拟公司的状态。评分为确定性的最终状态断言(含负向断言);官方得分为所有断言全部通过的任务比例,在私有保留集(600+ 任务)上运行,另有 600 个公开任务供研究。
- 发布
- 2026-04
- 维护者
- Zapier (Daniel Shepard, Robin Salimans)
- 状态
- active
- 污染风险
- low
- 指标
- task pass rate (all assertions) (percent, ↑)
- 题量
- 600
- 领域
- tool-use general-assistant instruction-following
- 人类
- 无实测基线
备注. Tasks were synthetically generated from the shape of real Zapier customer workflows and hardened; partial_credit (fraction of assertions) is reported as a diagnostic and RL reward but is not the headline score. Artificial Analysis runs an independent variant, AutomationBench-AA, whose headline is the share of objectives completed without guardrail violations, so AA numbers are on a different scale from Zapier's strict pass rate. Zapier's own leaderboard is a single run per model at its highest reasoning effort; the rank-1 Fable 5.1 entry uses an Opus 5 fallback for ~40% of tasks.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Opus 5 (max) | Anthropic | 50.3% | 官方榜单 | split: public reasoning_effort: max README table of pass rates on the 600-task public set; README undated, access date used. Public set is easier than the private leaderboard set. | |
| GPT-6 Astra | OpenAI | 41.4% | 厂商自报 | OpenAI launch table (Professional section), maximum at any effort; set (public/private) and toolset not stated. Same table quotes Fable 5.1 at 31.4%, Opus 5 26.9%, GPT-5.6 Sol 18.1%. Not yet on Zapier's leaderboard as of access date. | |
| Claude Fable 5.1 with Opus 5 fallback (max) | Anthropic | 31.4% | 官方榜单 | split: private reasoning_effort: max cost_usd_per_task: 2.45 Rank 1 on leaderboard v1.0.6 as of access date; Opus 5 completed ~40% of tasks (260/657) after Fable safety refusals and cost excludes fallback tokens. Leaderboard undated, access date used. | |
| Gemini 3.7 Flash (High) | 30.44% | 官方榜单 | split: private reasoning_effort: high cost_usd_per_task: 0.61 Rank 2 on leaderboard v1.0.6 as of access date; best standalone model. Leaderboard undated, access date used. | ||
| Opus 4.7 (max) | Anthropic | 9.9% | 论文 | split: private reasoning_effort: max cost_usd_per_task: 1.8 Top of the leaderboard table in the white paper (API toolset, single run). | |
| Gemini 3.1 Pro (high) | 9.6% | 论文 | split: private reasoning_effort: high cost_usd_per_task: 0.54 White paper leaderboard table (API toolset, single run). |