OSWorld

OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

↓ OSWorld 2.0

369 个在实时 Ubuntu 虚拟机中的真实计算机使用任务,涵盖 Chrome、LibreOffice、GIMP、VLC、VS Code、Thunderbird、操作系统文件操作和多应用工作流。智能体查看截图(可选无障碍树),输出鼠标/键盘动作,由基于执行的检查器对最终机器状态评分。是 GUI 计算机使用智能体的标准基准;OSWorld-Verified(2025 年 7 月)修复了 300 多个任务问题,是当前官方排行榜的基础。

90.19% · 人类 72.36%
Intelligence-Indeed Agent
官方榜单
2024-04-11 → 2026-08-01: 12.24% → 85.96%
发布
2024-04
维护者
XLANG Lab, University of Hong Kong (Xie et al.)
状态
active
污染风险
medium
指标
success rate (percent, ↑)
题量
369
领域
computer-use multimodal tool-use
人类
72.36% human computer users (paper)

备注. Official 'Verified' leaderboard rows are run by the XLANG team under unified settings (source: static/data/osworld_verified_results.xlsx on the site); self-reported rows are listed separately. Scores depend strongly on the max-step budget, so compare within the same budget. Task configs and evaluators are public, but live-environment execution limits memorisation. Superseded by OSWorld 2.0 (June 2026, 108 long-horizon tasks, tracked as `osworld-2`); the site now lives at osworld-v1.xlang.ai and the Verified leaderboard is still updated.

完整账本

系统开发者分数日期来源条件
Intelligence-Indeed AgentIntelligence Indeed90.19%官方榜单split: verified scaffold: Intelligence-Indeed Agent
Top of the verified leaderboard as of access date (agentic framework, 100 max steps).
claude-fable-5[1m]Anthropic85.96%官方榜单split: verified
Official verified run, 100 max steps; general model without a bespoke framework.
agent s3 w/ Opus 4.5 + GPT-5 bBoN (N=10)Simular72.58%官方榜单split: verified scaffold: Agent S3 with behavior best-of-N (N=10)
Official verified run, 100 max steps.
claude-sonnet-4-5-20250929Anthropic62.88%官方榜单split: verified
Official verified run, 100 max steps.
claude-4-sonnet-20250514Anthropic41.4%官方榜单split: verified
XLANG re-evaluation at launch of OSWorld-Verified, 100 max steps.
GPT-4V (screenshot + a11y tree)OpenAI12.24%论文split: original
Best model in the paper abstract.