OSWorld 2.0

OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks

↑ OSWorld

108 long-horizon computer-use workflows (median human time about 1.6 hours, roughly 300 agent tool calls) spanning research, creative, engineering, finance, administrative and healthcare work across desktop apps and 31 self-hosted websites. Tasks are graded on an average of 27 execution-based checkpoints, yielding a binary completion rate (primary metric, reported at a 500-step budget) and a partial score. Tasks include dynamic mid-task events, hidden state and simulated-user clarification, so it probes the failure modes OSWorld 1.0 no longer separates.

41.7%
Claude Fable 5.1
self-reported
2026-06-28 → 2026-09-03: 13% → 31.43%
Released
2026-06
Maintainer
XLANG Lab, University of Hong Kong (Yuan, Zhou, Xiong, Xie, Yu et al.)
Status
active
Contamination
medium
Metric
binary completion rate (500 steps) (percent, ↑)
Tasks
108
Domains
computer-use multimodal tool-use
human
no measured baseline

Notes. Official rows are run by the XLANG team (static/data/leaderboard/official-results.json) at 150/300/500 step budgets with either the standard or a 'batch tool' action setting; the 500-step binary metric is primary. Vendor launch tables often quote the partial score or the offline subset, and Anthropic's Fable 5.1 system card used modified tasks and grading, so treat developer-reported numbers as non-comparable to the leaderboard unless the release version, set and metric match.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude Fable 5.1Anthropic41.7%self-reportedsplit: full
Anthropic 'strict' score on the August 2026 task release with production safeguards on (partial 77.9%); post dated only 'September 2026', 1st used. OpenAI notes Anthropic used modified tasks and grading, so not leaderboard-comparable.
Claude Opus 5 (max, batch tool)Anthropic31.43%officialsplit: full reasoning_effort: max
Top of the official leaderboard as of access date; 500 steps, release v2026.08.08, partial 68.31%. Leaderboard updatedAt 2026-09-03 used as date.
GPT-5.6 Sol (max, batch tool)OpenAI27.34%officialsplit: full reasoning_effort: max
Official XLANG run, 500 steps, release v2026.08.08, partial 62.72%. Leaderboard updatedAt 2026-09-03 used as date.
Claude Opus 4.8 (max, batched tool calls)Anthropic20.6%papersplit: full reasoning_effort: max
Best model in the paper abstract; 500-step budget, task release v2026.06.24, partial score 54.8%.
GPT-5.5 (xhigh, batch tool)OpenAI13%papersplit: full reasoning_effort: xhigh
Paper abstract ('plateaus near 13%'); official leaderboard lists 13.0 binary / 49.5 partial at 500 steps, release v2026.06.24.