OSWorld 2.0
OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks
108 个长程计算机使用工作流(人类完成中位时间约 1.6 小时,智能体约 300 次工具调用),覆盖科研、创意、工程、财务、行政和医疗场景,涉及桌面应用与 31 个自托管网站。每个任务平均设 27 个基于执行的检查点,给出二元完成率(主指标,500 步预算)和部分得分。任务包含任务中途的动态事件、隐藏状态和模拟用户澄清,用于区分 OSWorld 1.0 已无法分辨的失败模式。
- 发布
- 2026-06
- 维护者
- XLANG Lab, University of Hong Kong (Yuan, Zhou, Xiong, Xie, Yu et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- binary completion rate (500 steps) (percent, ↑)
- 题量
- 108
- 领域
- computer-use multimodal tool-use
- 人类
- 无实测基线
备注. Official rows are run by the XLANG team (static/data/leaderboard/official-results.json) at 150/300/500 step budgets with either the standard or a 'batch tool' action setting; the 500-step binary metric is primary. Vendor launch tables often quote the partial score or the offline subset, and Anthropic's Fable 5.1 system card used modified tasks and grading, so treat developer-reported numbers as non-comparable to the leaderboard unless the release version, set and metric match.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 41.7% | 厂商自报 | split: full Anthropic 'strict' score on the August 2026 task release with production safeguards on (partial 77.9%); post dated only 'September 2026', 1st used. OpenAI notes Anthropic used modified tasks and grading, so not leaderboard-comparable. | |
| Claude Opus 5 (max, batch tool) | Anthropic | 31.43% | 官方榜单 | split: full reasoning_effort: max Top of the official leaderboard as of access date; 500 steps, release v2026.08.08, partial 68.31%. Leaderboard updatedAt 2026-09-03 used as date. | |
| GPT-5.6 Sol (max, batch tool) | OpenAI | 27.34% | 官方榜单 | split: full reasoning_effort: max Official XLANG run, 500 steps, release v2026.08.08, partial 62.72%. Leaderboard updatedAt 2026-09-03 used as date. | |
| Claude Opus 4.8 (max, batched tool calls) | Anthropic | 20.6% | 论文 | split: full reasoning_effort: max Best model in the paper abstract; 500-step budget, task release v2026.06.24, partial score 54.8%. | |
| GPT-5.5 (xhigh, batch tool) | OpenAI | 13% | 论文 | split: full reasoning_effort: xhigh Paper abstract ('plateaus near 13%'); official leaderboard lists 13.0 binary / 49.5 partial at 500 steps, release v2026.06.24. |