BrowseComp

BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents

1,266 human-written questions whose short, verifiable answers can only be found by persistently browsing many web pages and combining entangled constraints; human trainers verified that other people could not answer them within ten minutes and that then-current models failed. Scored as accuracy against the reference answer via a grader. It isolates deep web-research persistence rather than latent knowledge, so it separates browsing agents sharply.

92.2%
GPT-5.6 Sol
aggregator
2025-04-16 → 2026-09-03: 1.9% → 92.2%
Released
2025-04
Maintainer
OpenAI (Wei et al.)
Status
saturating
Contamination
medium
Metric
accuracy (percent, ↑)
Tasks
1,266
Domains
web tool-use factuality reasoning
human
no measured baseline

Notes. No official leaderboard; numbers are vendor self-reports collected by aggregators, and browsing budget, search backend and parallel-sampling settings are rarely disclosed. Questions and encrypted answers are public in simple-evals. Top frontier scores are now clustered above 90%.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-5.6 SolOpenAI92.2%aggregator
Top of the aggregator table as of access date; developer-reported, scaffold and effort unspecified.
Claude Opus 5Anthropic90.8%aggregator
Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified.
Claude Opus 4.6Anthropic83.7%aggregator
Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified.
Deep ResearchOpenAI51.5%papertools
Table 3 of the paper; model trained specifically for browsing tasks.
OpenAI o1OpenAI9.9%papertools: no
Table 3 of the paper (no browsing).
GPT-4o w/ browsingOpenAI1.9%papertools
Table 3 of the paper.