VisualWebArena

VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks

910 个在 Classifieds、Shopping 和 Reddit 站点上的网页任务,需要读取图像(商品照片、列表、发布的图片)才能完成,以视觉落地的检查扩展了 WebArena 的功能性评估。智能体接收截图(可选 Set-of-Marks 标注)和无障碍树,并需在浏览器中行动。它隔离考察纯文本网页智能体无法弥合的多模态感知差距。

54% · 人类 88.7%
Gemini 2.5 Flash (SGV)
官方榜单
2024-01-24 → 2025-10-01: 16.37% → 52.9%
发布
2024-01
维护者
Carnegie Mellon University (Koh et al.)
状态
active
污染风险
high
指标
success rate (percent, ↑)
题量
910
领域
web multimodal computer-use
人类
88.7% human annotators (paper)

备注. Shares the WebArena maintainer Google Sheet (separate tab) and the same self-reporting caveats. Best systems are still more than 30 points below the human baseline.

完整账本

系统开发者分数日期来源条件
Gemini 2.5 Flash (SGV)Google DeepMind54%官方榜单scaffold: SGV
Top of the sheet as of access date. Month only; 1st used.
GPT-5 (WALT)OpenAI52.9%官方榜单scaffold: WALT
Month only; 1st used.
GPT-4o + R-MCTS (ExACT)ExACT (leaderboard result source)33.7%官方榜单scaffold: ExACT R-MCTS
Month only; 1st used.
GPT-4o (SoM)OpenAI19.78%官方榜单scaffold: VisualWebArena SoM agent
Maintainer sheet, month only; 1st used.
GPT-4V (SoM)OpenAI16.37%论文scaffold: VisualWebArena SoM agent
Paper baseline, Set-of-Marks + captions + image inputs.