OSWorld
OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
369 个在实时 Ubuntu 虚拟机中的真实计算机使用任务,涵盖 Chrome、LibreOffice、GIMP、VLC、VS Code、Thunderbird、操作系统文件操作和多应用工作流。智能体查看截图(可选无障碍树),输出鼠标/键盘动作,由基于执行的检查器对最终机器状态评分。是 GUI 计算机使用智能体的标准基准;OSWorld-Verified(2025 年 7 月)修复了 300 多个任务问题,是当前官方排行榜的基础。
- 发布
- 2024-04
- 维护者
- XLANG Lab, University of Hong Kong (Xie et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- success rate (percent, ↑)
- 题量
- 369
- 领域
- computer-use multimodal tool-use
- 人类
- 72.36% human computer users (paper)
备注. Official 'Verified' leaderboard rows are run by the XLANG team under unified settings (source: static/data/osworld_verified_results.xlsx on the site); self-reported rows are listed separately. Scores depend strongly on the max-step budget, so compare within the same budget. Task configs and evaluators are public, but live-environment execution limits memorisation. Superseded by OSWorld 2.0 (June 2026, 108 long-horizon tasks, tracked as `osworld-2`); the site now lives at osworld-v1.xlang.ai and the Verified leaderboard is still updated.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Intelligence-Indeed Agent | Intelligence Indeed | 90.19% | 官方榜单 | split: verified scaffold: Intelligence-Indeed Agent Top of the verified leaderboard as of access date (agentic framework, 100 max steps). | |
| claude-fable-5[1m] | Anthropic | 85.96% | 官方榜单 | split: verified Official verified run, 100 max steps; general model without a bespoke framework. | |
| agent s3 w/ Opus 4.5 + GPT-5 bBoN (N=10) | Simular | 72.58% | 官方榜单 | split: verified scaffold: Agent S3 with behavior best-of-N (N=10) Official verified run, 100 max steps. | |
| claude-sonnet-4-5-20250929 | Anthropic | 62.88% | 官方榜单 | split: verified Official verified run, 100 max steps. | |
| claude-4-sonnet-20250514 | Anthropic | 41.4% | 官方榜单 | split: verified XLANG re-evaluation at launch of OSWorld-Verified, 100 max steps. | |
| GPT-4V (screenshot + a11y tree) | OpenAI | 12.24% | 论文 | split: original Best model in the paper abstract. |