DeepSearchQA

DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents

900 个人工构造、时间锚定的网页研究问题,覆盖 17 个领域,答案为完整集合而非单一事实(65% 为集合型答案)。每题是一条依赖链式检索,智能体需规划多步搜索、汇总多来源碎片、消歧去重并判断何时停止。由固定的 Gemini 2.5 Flash 自动评审判定语义集合匹配;主指标为逐题平均 F1(精确率对召回率),并报告完全正确、完全错误和多答比例作为诊断。针对 BrowseComp 等单答案基准忽视的全面性缺口。

95%
Claude Opus 5
聚合站
2025-12-11 → 2026-09-03: 56.5% → 95%
发布
2025-12
维护者
Google DeepMind (Gupta, Chatterjee, Haas, Tao et al.); leaderboard run by Kaggle
状态
active
污染风险
medium
指标
F1 (mean over prompts) (percent, ↑)
题量
900
领域
web research factuality tool-use
人类
无实测基线

备注. Announced 2025-12-11 with the Gemini Deep Research agent; arXiv paper followed 2026-01-28. The official Kaggle leaderboard (last updated 2025-12-11, 13 systems) is independently run by Kaggle and reports F1 with 95% CIs; the dataset card warns that a different autorater or grading prompt gives statistically different scores, so vendor-reported numbers (Meta, Anthropic, Moonshot via aggregators) may not be comparable. Prompts and gold answers are public on Hugging Face/Kaggle, and ground truth can drift as web sources change.

完整账本

系统开发者分数日期来源条件
Claude Opus 5Anthropic95%聚合站tools
Vendor-published number republished by aggregator (tied with Kimi K3 at 95.0); harness and autorater unspecified, so not comparable to the Kaggle-run board.
Gemini Deep Research AgentGoogle81.9%官方榜单tools
Kaggle-run leaderboard, rank 1; 66.1% fully correct, 10.0% fully incorrect, 10.3% correct with excessive answers. Leaderboard last updated 2025-12-11.
GPT-5 ProOpenAI79%官方榜单tools
Kaggle-run leaderboard, rank 2; 65.2% fully correct, 14.1% fully incorrect.
GPT-5.4OpenAI78.6%官方榜单tools
Kaggle-run leaderboard, rank 3; 63.7% fully correct. Date is the leaderboard's 'last updated' stamp.
Claude Sonnet 4.6Anthropic56.5%官方榜单tools
Kaggle-run leaderboard, rank 6; 50.8% fully correct, 41.1% fully incorrect. Date is the leaderboard's 'last updated' stamp.