AndroidWorld

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

116 个人工设计的任务,覆盖 20 个真实 Android 应用,在实时 Android 模拟器中执行。任务带参数并随机实例化(姓名、日期、金额等),可生成数百万个变体;专用的初始化、成功检查和清理逻辑读取设备状态给出可靠奖励。智能体观察截图和/或无障碍树,输出点击、输入和导航动作,得分为任务成功率。是移动端 GUI 智能体的参考基准,并兼容 MiniWoB++ 网页任务。

100% · 人类 80%
FluizAI agent (gpt-4o / gpt-5.6-sol)
厂商自报
2024-05-23 → 2026-09-03: 30.6% → 85.3%
发布
2024-05
维护者
Google Research / DeepMind (Rawles, Clinckemaillie, Riva et al.)
状态
saturated
污染风险
medium
指标
task success rate (percent, ↑)
题量
116
领域
computer-use multimodal tool-use
人类
80% human annotators, 3 trials (paper / leaderboard sheet)

备注. The leaderboard is a Google Sheet of community-submitted, self-reported pass@1 numbers with no independent verification; several entries have been disputed or revised. Since 2024-11-18 the per-task step budget is about 2x the human completion time. Task templates and success checkers are public and the top agents now exceed the 80% human baseline, with a 100% entry in August 2026, so the benchmark no longer separates frontier systems.

完整账本

系统开发者分数日期来源条件
FluizAI agent (gpt-4o / gpt-5.6-sol)FluizAI100%厂商自报pass_k: 1 scaffold: FluizAI
Rank 1 on the community sheet as of access date; screenshot + a11y tree; code not open-sourced. Sheet gives only 08/2026, 1st used. No independent verification.
Qwen3.8 MaxAlibaba85.3%聚合站
Vendor-published number republished by aggregator; harness and observation space unspecified.
AutoGLM-Mobile (9B)Zhipu AI80.2%厂商自报pass_k: 1
Self-reported entry on the official community leaderboard sheet (screenshot + a11y tree); updated 2025-09-26 from 75.8. No independent verification.
Gemini 2.5 Computer UseGoogle69.7%厂商自报pass_k: 1
Self-reported entry on the community leaderboard sheet (screenshot only, single model); sheet gives only 10/2025, 1st used.
M3A (GPT-4 Turbo, a11y tree)Google Research30.6%论文scaffold: M3A
Best agent in the paper abstract.