VisualWebArena

VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

910 web tasks on Classifieds, Shopping and Reddit sites that require reading images (product photos, listings, posted pictures) to complete, extending WebArena's functional evaluation with visually grounded checks. Agents receive screenshots (optionally Set-of-Marks annotated) plus the accessibility tree and must act in the browser. It isolates the multimodal-perception gap that text-only web agents cannot close.

54% · human 88.7%
Gemini 2.5 Flash (SGV)
official
2024-01-24 → 2025-10-01: 16.37% → 52.9%
Released
2024-01
Maintainer
Carnegie Mellon University (Koh et al.)
Status
active
Contamination
high
Metric
success rate (percent, ↑)
Tasks
910
Domains
web multimodal computer-use
human
88.7% human annotators (paper)

Notes. Shares the WebArena maintainer Google Sheet (separate tab) and the same self-reporting caveats. Best systems are still more than 30 points below the human baseline.

Full ledger

SystemDeveloperScoreDateSourceConditions
Gemini 2.5 Flash (SGV)Google DeepMind54%officialscaffold: SGV
Top of the sheet as of access date. Month only; 1st used.
GPT-5 (WALT)OpenAI52.9%officialscaffold: WALT
Month only; 1st used.
GPT-4o + R-MCTS (ExACT)ExACT (leaderboard result source)33.7%officialscaffold: ExACT R-MCTS
Month only; 1st used.
GPT-4o (SoM)OpenAI19.78%officialscaffold: VisualWebArena SoM agent
Maintainer sheet, month only; 1st used.
GPT-4V (SoM)OpenAI16.37%paperscaffold: VisualWebArena SoM agent
Paper baseline, Set-of-Marks + captions + image inputs.