OSWorld 2.0
OSWorld 2.0: Benchmarking Computer-Use Agents on Long-Horizon Real-World Tasks
108 long-horizon computer-use workflows (median human time about 1.6 hours, roughly 300 agent tool calls) spanning research, creative, engineering, finance, administrative and healthcare work across desktop apps and 31 self-hosted websites. Tasks are graded on an average of 27 execution-based checkpoints, yielding a binary completion rate (primary metric, reported at a 500-step budget) and a partial score. Tasks include dynamic mid-task events, hidden state and simulated-user clarification, so it probes the failure modes OSWorld 1.0 no longer separates.
paper · website · leaderboard · dataset · code
- Released
- 2026-06
- Maintainer
- XLANG Lab, University of Hong Kong (Yuan, Zhou, Xiong, Xie, Yu et al.)
- Status
- active
- Contamination
- medium
- Metric
- binary completion rate (500 steps) (percent, ↑)
- Tasks
- 108
- Domains
- computer-use multimodal tool-use
- human
- no measured baseline
Notes. Official rows are run by the XLANG team (static/data/leaderboard/official-results.json) at 150/300/500 step budgets with either the standard or a 'batch tool' action setting; the 500-step binary metric is primary. Vendor launch tables often quote the partial score or the offline subset, and Anthropic's Fable 5.1 system card used modified tasks and grading, so treat developer-reported numbers as non-comparable to the leaderboard unless the release version, set and metric match.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Claude Fable 5.1 | Anthropic | 41.7% | self-reported | split: full Anthropic 'strict' score on the August 2026 task release with production safeguards on (partial 77.9%); post dated only 'September 2026', 1st used. OpenAI notes Anthropic used modified tasks and grading, so not leaderboard-comparable. | |
| Claude Opus 5 (max, batch tool) | Anthropic | 31.43% | official | split: full reasoning_effort: max Top of the official leaderboard as of access date; 500 steps, release v2026.08.08, partial 68.31%. Leaderboard updatedAt 2026-09-03 used as date. | |
| GPT-5.6 Sol (max, batch tool) | OpenAI | 27.34% | official | split: full reasoning_effort: max Official XLANG run, 500 steps, release v2026.08.08, partial 62.72%. Leaderboard updatedAt 2026-09-03 used as date. | |
| Claude Opus 4.8 (max, batched tool calls) | Anthropic | 20.6% | paper | split: full reasoning_effort: max Best model in the paper abstract; 500-step budget, task release v2026.06.24, partial score 54.8%. | |
| GPT-5.5 (xhigh, batch tool) | OpenAI | 13% | paper | split: full reasoning_effort: xhigh Paper abstract ('plateaus near 13%'); official leaderboard lists 13.0 binary / 49.5 partial at 500 steps, release v2026.06.24. |