BrowseComp
BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
1,266 human-written questions whose short, verifiable answers can only be found by persistently browsing many web pages and combining entangled constraints; human trainers verified that other people could not answer them within ten minutes and that then-current models failed. Scored as accuracy against the reference answer via a grader. It isolates deep web-research persistence rather than latent knowledge, so it separates browsing agents sharply.
- Released
- 2025-04
- Maintainer
- OpenAI (Wei et al.)
- Status
- Contamination
- medium
- Metric
- accuracy (percent, ↑)
- Tasks
- 1,266
- Domains
- web tool-use factuality reasoning
- human
- no measured baseline
Notes. No official leaderboard; numbers are vendor self-reports collected by aggregators, and browsing budget, search backend and parallel-sampling settings are rarely disclosed. Questions and encrypted answers are public in simple-evals. Top frontier scores are now clustered above 90%.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-5.6 Sol | OpenAI | 92.2% | aggregator | Top of the aggregator table as of access date; developer-reported, scaffold and effort unspecified. | |
| Claude Opus 5 | Anthropic | 90.8% | aggregator | Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified. | |
| Claude Opus 4.6 | Anthropic | 83.7% | aggregator | Developer-reported number republished by aggregator (page 'data verified' date used); browsing scaffold and effort unspecified. | |
| Deep Research | OpenAI | 51.5% | paper | tools Table 3 of the paper; model trained specifically for browsing tasks. | |
| OpenAI o1 | OpenAI | 9.9% | paper | tools: no Table 3 of the paper (no browsing). | |
| GPT-4o w/ browsing | OpenAI | 1.9% | paper | tools Table 3 of the paper. |