BrowseComp

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

1,266 个人工编写的问题,其简短可验证的答案只能通过持续浏览大量网页并组合相互交织的约束才能找到;人类训练者验证了他人无法在十分钟内作答,且当时的模型均告失败。通过评分器对照参考答案计算准确率。它隔离考察深度网络研究的持久性而非潜在知识,因此能显著区分各浏览智能体。

92.2%
GPT-5.6 Sol
聚合站
2025-04-16 → 2026-09-03: 1.9% → 92.2%
发布
2025-04
维护者
OpenAI (Wei et al.)
状态
saturating
污染风险
medium
指标
accuracy (percent, ↑)
题量
1,266
领域
web tool-use factuality reasoning
人类
无实测基线

备注. No official leaderboard; numbers are vendor self-reports collected by aggregators, and browsing budget, search backend and parallel-sampling settings are rarely disclosed. Questions and encrypted answers are public in simple-evals. Top frontier scores are now clustered above 90%.

完整账本

系统开发者分数日期来源条件
GPT-5.6 SolOpenAI92.2%聚合站
Top of the aggregator table as of access date; developer-reported, scaffold and effort unspecified.
Claude Opus 5Anthropic90.8%聚合站
Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified.
Claude Opus 4.6Anthropic83.7%聚合站
Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified.
Deep ResearchOpenAI51.5%论文tools
Table 3 of the paper; model trained specifically for browsing tasks.
OpenAI o1OpenAI9.9%论文tools: no
Table 3 of the paper (no browsing).
GPT-4o w/ browsingOpenAI1.9%论文tools
Table 3 of the paper.