VisualWebArena
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
910 web tasks on Classifieds, Shopping and Reddit sites that require reading images (product photos, listings, posted pictures) to complete, extending WebArena's functional evaluation with visually grounded checks. Agents receive screenshots (optionally Set-of-Marks annotated) plus the accessibility tree and must act in the browser. It isolates the multimodal-perception gap that text-only web agents cannot close.
paper · website · leaderboard · dataset · code
- Released
- 2024-01
- Maintainer
- Carnegie Mellon University (Koh et al.)
- Status
- active
- Contamination
- high
- Metric
- success rate (percent, ↑)
- Tasks
- 910
- Domains
- web multimodal computer-use
- human
- 88.7% human annotators (paper)
Notes. Shares the WebArena maintainer Google Sheet (separate tab) and the same self-reporting caveats. Best systems are still more than 30 points below the human baseline.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Gemini 2.5 Flash (SGV) | Google DeepMind | 54% | official | scaffold: SGV Top of the sheet as of access date. Month only; 1st used. | |
| GPT-5 (WALT) | OpenAI | 52.9% | official | scaffold: WALT Month only; 1st used. | |
| GPT-4o + R-MCTS (ExACT) | ExACT (leaderboard result source) | 33.7% | official | scaffold: ExACT R-MCTS Month only; 1st used. | |
| GPT-4o (SoM) | OpenAI | 19.78% | official | scaffold: VisualWebArena SoM agent Maintainer sheet, month only; 1st used. | |
| GPT-4V (SoM) | OpenAI | 16.37% | paper | scaffold: VisualWebArena SoM agent Paper baseline, Set-of-Marks + captions + image inputs. |