AndroidWorld
AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents
116 hand-crafted tasks across 20 real Android apps, executed in a live Android emulator. Each task is parameterised and instantiated with random values (names, dates, amounts), so the suite yields millions of unique variants; dedicated set-up, success-check and tear-down logic inspect device state to produce a durable reward. Agents observe screenshots and/or the accessibility tree and emit touch, type and navigation actions; the score is the task success rate. It is the reference benchmark for mobile GUI agents and also hosts MiniWoB++ web tasks.
paper · website · leaderboard · dataset · code
- Released
- 2024-05
- Maintainer
- Google Research / DeepMind (Rawles, Clinckemaillie, Riva et al.)
- Status
- saturated
- Contamination
- medium
- Metric
- task success rate (percent, ↑)
- Tasks
- 116
- Domains
- computer-use multimodal tool-use
- human
- 80% human annotators, 3 trials (paper / leaderboard sheet)
Notes. The leaderboard is a Google Sheet of community-submitted, self-reported pass@1 numbers with no independent verification; several entries have been disputed or revised. Since 2024-11-18 the per-task step budget is about 2x the human completion time. Task templates and success checkers are public and the top agents now exceed the 80% human baseline, with a 100% entry in August 2026, so the benchmark no longer separates frontier systems.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| FluizAI agent (gpt-4o / gpt-5.6-sol) | FluizAI | 100% | self-reported | pass_k: 1 scaffold: FluizAI Rank 1 on the community sheet as of access date; screenshot + a11y tree; code not open-sourced. Sheet gives only 08/2026, 1st used. No independent verification. | |
| Qwen3.8 Max | Alibaba | 85.3% | aggregator | Vendor-published number republished by aggregator; harness and observation space unspecified. | |
| AutoGLM-Mobile (9B) | Zhipu AI | 80.2% | self-reported | pass_k: 1 Self-reported entry on the official community leaderboard sheet (screenshot + a11y tree); updated 2025-09-26 from 75.8. No independent verification. | |
| Gemini 2.5 Computer Use | 69.7% | self-reported | pass_k: 1 Self-reported entry on the community leaderboard sheet (screenshot only, single model); sheet gives only 10/2025, 1st used. | ||
| M3A (GPT-4 Turbo, a11y tree) | Google Research | 30.6% | paper | scaffold: M3A Best agent in the paper abstract. |