AndroidWorld
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
116 个人工设计的任务,覆盖 20 个真实 Android 应用,在实时 Android 模拟器中执行。任务带参数并随机实例化(姓名、日期、金额等),可生成数百万个变体;专用的初始化、成功检查和清理逻辑读取设备状态给出可靠奖励。智能体观察截图和/或无障碍树,输出点击、输入和导航动作,得分为任务成功率。是移动端 GUI 智能体的参考基准,并兼容 MiniWoB++ 网页任务。
- 发布
- 2024-05
- 维护者
- Google Research / DeepMind (Rawles, Clinckemaillie, Riva et al.)
- 状态
- saturated
- 污染风险
- medium
- 指标
- task success rate (percent, ↑)
- 题量
- 116
- 领域
- computer-use multimodal tool-use
- 人类
- 80% human annotators, 3 trials (paper / leaderboard sheet)
备注. The leaderboard is a Google Sheet of community-submitted, self-reported pass@1 numbers with no independent verification; several entries have been disputed or revised. Since 2024-11-18 the per-task step budget is about 2x the human completion time. Task templates and success checkers are public and the top agents now exceed the 80% human baseline, with a 100% entry in August 2026, so the benchmark no longer separates frontier systems.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| FluizAI agent (gpt-4o / gpt-5.6-sol) | FluizAI | 100% | 厂商自报 | pass_k: 1 scaffold: FluizAI Rank 1 on the community sheet as of access date; screenshot + a11y tree; code not open-sourced. Sheet gives only 08/2026, 1st used. No independent verification. | |
| Qwen3.8 Max | Alibaba | 85.3% | 聚合站 | Vendor-published number republished by aggregator; harness and observation space unspecified. | |
| AutoGLM-Mobile (9B) | Zhipu AI | 80.2% | 厂商自报 | pass_k: 1 Self-reported entry on the official community leaderboard sheet (screenshot + a11y tree); updated 2025-09-26 from 75.8. No independent verification. | |
| Gemini 2.5 Computer Use | 69.7% | 厂商自报 | pass_k: 1 Self-reported entry on the community leaderboard sheet (screenshot only, single model); sheet gives only 10/2025, 1st used. | ||
| M3A (GPT-4 Turbo, a11y tree) | Google Research | 30.6% | 论文 | scaffold: M3A Best agent in the paper abstract. |