OSWorld

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

↓ OSWorld 2.0

369 real computer-use tasks in a live Ubuntu VM spanning Chrome, LibreOffice, GIMP, VLC, VS Code, Thunderbird, OS file operations and multi-app workflows. Agents see screenshots (and optionally the accessibility tree), emit mouse/keyboard actions, and are scored by execution-based checkers on the final machine state. The standard benchmark for GUI computer-use agents; OSWorld-Verified (July 2025) fixed 300+ task issues and is the basis of the current official leaderboard.

90.19% · human 72.36%
Intelligence-Indeed Agent
official
2024-04-11 → 2026-08-01: 12.24% → 85.96%
Released
2024-04
Maintainer
XLANG Lab, University of Hong Kong (Xie et al.)
Status
active
Contamination
medium
Metric
success rate (percent, ↑)
Tasks
369
Domains
computer-use multimodal tool-use
human
72.36% human computer users (paper)

Notes. Official 'Verified' leaderboard rows are run by the XLANG team under unified settings (source: static/data/osworld_verified_results.xlsx on the site); self-reported rows are listed separately. Scores depend strongly on the max-step budget, so compare within the same budget. Task configs and evaluators are public, but live-environment execution limits memorisation. Superseded by OSWorld 2.0 (June 2026, 108 long-horizon tasks, tracked as `osworld-2`); the site now lives at osworld-v1.xlang.ai and the Verified leaderboard is still updated.

Full ledger

SystemDeveloperScoreDateSourceConditions
Intelligence-Indeed AgentIntelligence Indeed90.19%officialsplit: verified scaffold: Intelligence-Indeed Agent
Top of the verified leaderboard as of access date (agentic framework, 100 max steps).
claude-fable-5[1m]Anthropic85.96%officialsplit: verified
Official verified run, 100 max steps; general model without a bespoke framework.
agent s3 w/ Opus 4.5 + GPT-5 bBoN (N=10)Simular72.58%officialsplit: verified scaffold: Agent S3 with behavior best-of-N (N=10)
Official verified run, 100 max steps.
claude-sonnet-4-5-20250929Anthropic62.88%officialsplit: verified
Official verified run, 100 max steps.
claude-4-sonnet-20250514Anthropic41.4%officialsplit: verified
XLANG re-evaluation at launch of OSWorld-Verified, 100 max steps.
GPT-4V (screenshot + a11y tree)OpenAI12.24%papersplit: original
Best model in the paper abstract.