VisualWebArena
VisualWebArena: Evaluating Multimodal Agents on Realistic Visual Web Tasks
910 个在 Classifieds、Shopping 和 Reddit 站点上的网页任务,需要读取图像(商品照片、列表、发布的图片)才能完成,以视觉落地的检查扩展了 WebArena 的功能性评估。智能体接收截图(可选 Set-of-Marks 标注)和无障碍树,并需在浏览器中行动。它隔离考察纯文本网页智能体无法弥合的多模态感知差距。
- 发布
- 2024-01
- 维护者
- Carnegie Mellon University (Koh et al.)
- 状态
- active
- 污染风险
- high
- 指标
- success rate (percent, ↑)
- 题量
- 910
- 领域
- web multimodal computer-use
- 人类
- 88.7% human annotators (paper)
备注. Shares the WebArena maintainer Google Sheet (separate tab) and the same self-reporting caveats. Best systems are still more than 30 points below the human baseline.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Gemini 2.5 Flash (SGV) | Google DeepMind | 54% | 官方榜单 | scaffold: SGV Top of the sheet as of access date. Month only; 1st used. | |
| GPT-5 (WALT) | OpenAI | 52.9% | 官方榜单 | scaffold: WALT Month only; 1st used. | |
| GPT-4o + R-MCTS (ExACT) | ExACT (leaderboard result source) | 33.7% | 官方榜单 | scaffold: ExACT R-MCTS Month only; 1st used. | |
| GPT-4o (SoM) | OpenAI | 19.78% | 官方榜单 | scaffold: VisualWebArena SoM agent Maintainer sheet, month only; 1st used. | |
| GPT-4V (SoM) | OpenAI | 16.37% | 论文 | scaffold: VisualWebArena SoM agent Paper baseline, Set-of-Marks + captions + image inputs. |