DeepSearchQA
DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
900 个人工构造、时间锚定的网页研究问题,覆盖 17 个领域,答案为完整集合而非单一事实(65% 为集合型答案)。每题是一条依赖链式检索,智能体需规划多步搜索、汇总多来源碎片、消歧去重并判断何时停止。由固定的 Gemini 2.5 Flash 自动评审判定语义集合匹配;主指标为逐题平均 F1(精确率对召回率),并报告完全正确、完全错误和多答比例作为诊断。针对 BrowseComp 等单答案基准忽视的全面性缺口。
- 发布
- 2025-12
- 维护者
- Google DeepMind (Gupta, Chatterjee, Haas, Tao et al.); leaderboard run by Kaggle
- 状态
- active
- 污染风险
- medium
- 指标
- F1 (mean over prompts) (percent, ↑)
- 题量
- 900
- 领域
- web research factuality tool-use
- 人类
- 无实测基线
备注. Announced 2025-12-11 with the Gemini Deep Research agent; arXiv paper followed 2026-01-28. The official Kaggle leaderboard (last updated 2025-12-11, 13 systems) is independently run by Kaggle and reports F1 with 95% CIs; the dataset card warns that a different autorater or grading prompt gives statistically different scores, so vendor-reported numbers (Meta, Anthropic, Moonshot via aggregators) may not be comparable. Prompts and gold answers are public on Hugging Face/Kaggle, and ground truth can drift as web sources change.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 95% | 聚合站 | tools Vendor-published number republished by aggregator (tied with Kimi K3 at 95.0); harness and autorater unspecified, so not comparable to the Kaggle-run board. | |
| Gemini Deep Research Agent | 81.9% | 官方榜单 | tools Kaggle-run leaderboard, rank 1; 66.1% fully correct, 10.0% fully incorrect, 10.3% correct with excessive answers. Leaderboard last updated 2025-12-11. | ||
| GPT-5 Pro | OpenAI | 79% | 官方榜单 | tools Kaggle-run leaderboard, rank 2; 65.2% fully correct, 14.1% fully incorrect. | |
| GPT-5.4 | OpenAI | 78.6% | 官方榜单 | tools Kaggle-run leaderboard, rank 3; 63.7% fully correct. Date is the leaderboard's 'last updated' stamp. | |
| Claude Sonnet 4.6 | Anthropic | 56.5% | 官方榜单 | tools Kaggle-run leaderboard, rank 6; 50.8% fully correct, 41.1% fully incorrect. Date is the leaderboard's 'last updated' stamp. |