OSWorld
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
369 real computer-use tasks in a live Ubuntu VM spanning Chrome, LibreOffice, GIMP, VLC, VS Code, Thunderbird, OS file operations and multi-app workflows. Agents see screenshots (and optionally the accessibility tree), emit mouse/keyboard actions, and are scored by execution-based checkers on the final machine state. The standard benchmark for GUI computer-use agents; OSWorld-Verified (July 2025) fixed 300+ task issues and is the basis of the current official leaderboard.
paper · website · leaderboard · dataset · code
- Released
- 2024-04
- Maintainer
- XLANG Lab, University of Hong Kong (Xie et al.)
- Status
- active
- Contamination
- medium
- Metric
- success rate (percent, ↑)
- Tasks
- 369
- Domains
- computer-use multimodal tool-use
- human
- 72.36% human computer users (paper)
Notes. Official 'Verified' leaderboard rows are run by the XLANG team under unified settings (source: static/data/osworld_verified_results.xlsx on the site); self-reported rows are listed separately. Scores depend strongly on the max-step budget, so compare within the same budget. Task configs and evaluators are public, but live-environment execution limits memorisation. Superseded by OSWorld 2.0 (June 2026, 108 long-horizon tasks, tracked as `osworld-2`); the site now lives at osworld-v1.xlang.ai and the Verified leaderboard is still updated.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Intelligence-Indeed Agent | Intelligence Indeed | 90.19% | official | split: verified scaffold: Intelligence-Indeed Agent Top of the verified leaderboard as of access date (agentic framework, 100 max steps). | |
| claude-fable-5[1m] | Anthropic | 85.96% | official | split: verified Official verified run, 100 max steps; general model without a bespoke framework. | |
| agent s3 w/ Opus 4.5 + GPT-5 bBoN (N=10) | Simular | 72.58% | official | split: verified scaffold: Agent S3 with behavior best-of-N (N=10) Official verified run, 100 max steps. | |
| claude-sonnet-4-5-20250929 | Anthropic | 62.88% | official | split: verified Official verified run, 100 max steps. | |
| claude-4-sonnet-20250514 | Anthropic | 41.4% | official | split: verified XLANG re-evaluation at launch of OSWorld-Verified, 100 max steps. | |
| GPT-4V (screenshot + a11y tree) | OpenAI | 12.24% | paper | split: original Best model in the paper abstract. |