OSWorld 2.0

OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks

↑ OSWorld

108 个长程计算机使用工作流(人类完成中位时间约 1.6 小时,智能体约 300 次工具调用),覆盖科研、创意、工程、财务、行政和医疗场景,涉及桌面应用与 31 个自托管网站。每个任务平均设 27 个基于执行的检查点,给出二元完成率(主指标,500 步预算)和部分得分。任务包含任务中途的动态事件、隐藏状态和模拟用户澄清,用于区分 OSWorld 1.0 已无法分辨的失败模式。

41.7%
Claude Fable 5.1
厂商自报
2026-06-28 → 2026-09-03: 13% → 31.43%
发布
2026-06
维护者
XLANG Lab, University of Hong Kong (Yuan, Zhou, Xiong, Xie, Yu et al.)
状态
active
污染风险
medium
指标
binary completion rate (500 steps) (percent, ↑)
题量
108
领域
computer-use multimodal tool-use
人类
无实测基线

备注. Official rows are run by the XLANG team (static/data/leaderboard/official-results.json) at 150/300/500 step budgets with either the standard or a 'batch tool' action setting; the 500-step binary metric is primary. Vendor launch tables often quote the partial score or the offline subset, and Anthropic's Fable 5.1 system card used modified tasks and grading, so treat developer-reported numbers as non-comparable to the leaderboard unless the release version, set and metric match.

完整账本

系统开发者分数日期来源条件
Claude Fable 5.1Anthropic41.7%厂商自报split: full
Anthropic 'strict' score on the August 2026 task release with production safeguards on (partial 77.9%); post dated only 'September 2026', 1st used. OpenAI notes Anthropic used modified tasks and grading, so not leaderboard-comparable.
Claude Opus 5 (max, batch tool)Anthropic31.43%官方榜单split: full reasoning_effort: max
Top of the official leaderboard as of access date; 500 steps, release v2026.08.08, partial 68.31%. Leaderboard updatedAt 2026-09-03 used as date.
GPT-5.6 Sol (max, batch tool)OpenAI27.34%官方榜单split: full reasoning_effort: max
Official XLANG run, 500 steps, release v2026.08.08, partial 62.72%. Leaderboard updatedAt 2026-09-03 used as date.
Claude Opus 4.8 (max, batched tool calls)Anthropic20.6%论文split: full reasoning_effort: max
Best model in the paper abstract; 500-step budget, task release v2026.06.24, partial score 54.8%.
GPT-5.5 (xhigh, batch tool)OpenAI13%论文split: full reasoning_effort: xhigh
Paper abstract ('plateaus near 13%'); official leaderboard lists 13.0 binary / 49.5 partial at 500 steps, release v2026.06.24.