BrowseComp
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
1,266 个人工编写的问题,其简短可验证的答案只能通过持续浏览大量网页并组合相互交织的约束才能找到;人类训练者验证了他人无法在十分钟内作答,且当时的模型均告失败。通过评分器对照参考答案计算准确率。它隔离考察深度网络研究的持久性而非潜在知识,因此能显著区分各浏览智能体。
- 发布
- 2025-04
- 维护者
- OpenAI (Wei et al.)
- 状态
- 污染风险
- medium
- 指标
- accuracy (percent, ↑)
- 题量
- 1,266
- 领域
- web tool-use factuality reasoning
- 人类
- 无实测基线
备注. No official leaderboard; numbers are vendor self-reports collected by aggregators, and browsing budget, search backend and parallel-sampling settings are rarely disclosed. Questions and encrypted answers are public in simple-evals. Top frontier scores are now clustered above 90%.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 92.2% | 聚合站 | Top of the aggregator table as of access date; developer-reported, scaffold and effort unspecified. | |
| Claude Opus 5 | Anthropic | 90.8% | 聚合站 | Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified. | |
| Claude Opus 4.6 | Anthropic | 83.7% | 聚合站 | Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified. | |
| Deep Research | OpenAI | 51.5% | 论文 | tools Table 3 of the paper; model trained specifically for browsing tasks. | |
| OpenAI o1 | OpenAI | 9.9% | 论文 | tools: no Table 3 of the paper (no browsing). | |
| GPT-4o w/ browsing | OpenAI | 1.9% | 论文 | tools Table 3 of the paper. |