AndroidWorld

AndroidWorld: A Dynamic Benchmarking Environment for Autonomous Agents

116 hand-crafted tasks across 20 real Android apps, executed in a live Android emulator. Each task is parameterised and instantiated with random values (names, dates, amounts), so the suite yields millions of unique variants; dedicated set-up, success-check and tear-down logic inspect device state to produce a durable reward. Agents observe screenshots and/or the accessibility tree and emit touch, type and navigation actions; the score is the task success rate. It is the reference benchmark for mobile GUI agents and also hosts MiniWoB++ web tasks.

100% · human 80%
FluizAI agent (gpt-4o / gpt-5.6-sol)
self-reported
2024-05-23 → 2026-09-03: 30.6% → 85.3%
Released
2024-05
Maintainer
Google Research / DeepMind (Rawles, Clinckemaillie, Riva et al.)
Status
saturated
Contamination
medium
Metric
task success rate (percent, ↑)
Tasks
116
Domains
computer-use multimodal tool-use
human
80% human annotators, 3 trials (paper / leaderboard sheet)

Notes. The leaderboard is a Google Sheet of community-submitted, self-reported pass@1 numbers with no independent verification; several entries have been disputed or revised. Since 2024-11-18 the per-task step budget is about 2x the human completion time. Task templates and success checkers are public and the top agents now exceed the 80% human baseline, with a 100% entry in August 2026, so the benchmark no longer separates frontier systems.

Full ledger

SystemDeveloperScoreDateSourceConditions
FluizAI agent (gpt-4o / gpt-5.6-sol)FluizAI100%self-reportedpass_k: 1 scaffold: FluizAI
Rank 1 on the community sheet as of access date; screenshot + a11y tree; code not open-sourced. Sheet gives only 08/2026, 1st used. No independent verification.
Qwen3.8 MaxAlibaba85.3%aggregator
Vendor-published number republished by aggregator; harness and observation space unspecified.
AutoGLM-Mobile (9B)Zhipu AI80.2%self-reportedpass_k: 1
Self-reported entry on the official community leaderboard sheet (screenshot + a11y tree); updated 2025-09-26 from 75.8. No independent verification.
Gemini 2.5 Computer UseGoogle69.7%self-reportedpass_k: 1
Self-reported entry on the community leaderboard sheet (screenshot only, single model); sheet gives only 10/2025, 1st used.
M3A (GPT-4 Turbo, a11y tree)Google Research30.6%paperscaffold: M3A
Best agent in the paper abstract.