WebArena

WebArena: A Realistic Web Environment for Building Autonomous Agents

812 long-horizon natural-language tasks on self-hosted, fully functional websites (e-commerce, forum, GitLab, CMS) plus a map, calculator and wiki. Agents act through a browser and are scored by programmatic functional checks on the final site state or answer rather than by matching action sequences. The first reproducible end-to-end web-agent benchmark and still the reference for text/DOM web agents.

74.3% · human 78.24%
WebTactix + Deepseek v3.2
official
2023-06-01 → 2026-02-01: 14.9% → 74.3%
Released
2023-07
Maintainer
Carnegie Mellon University (Zhou, Xu et al.)
Status
active
Contamination
high
Metric
success rate (percent, ↑)
Tasks
812
Domains
web tool-use computer-use
human
78.24% human annotators (paper, end-to-end task success)

Notes. Leaderboard is a maintainer-curated Google Sheet of self-reported results (trajectories required since September 2024). Environments are public Docker images, so task configs are widely available. Top systems now approach the 78% human figure; the maintainers point to WebArena-Infinity and TheAgentCompany as successors, neither of which is tracked here yet.

Full ledger

SystemDeveloperScoreDateSourceConditions
WebTactix + Deepseek v3.2WebTactix74.3%officialscaffold: WebTactix
Top of the maintainer sheet as of access date. Month only; 1st used.
ColorBrowserAgentColorBrowserAgent71.2%officialscaffold: ColorBrowserAgent
Open-source entry with trajectories. Month only; 1st used.
OpenAI OperatorOpenAI58.1%officialscaffold: OpenAI CUA
Self-reported by OpenAI (system card), listed on the maintainer sheet. Month only; 1st used.
StePSteP (leaderboard result source)33.5%officialscaffold: SteP
Sheet notes high-level plans are human-derived. Month only; 1st used.
gpt-4-0613OpenAI14.9%official
WebArena team baseline without the 'not achievable' hint (paper reports 14.41% with it). Sheet gives month only; 1st used.