DROP
DROP: Discrete Reasoning Over the content of Paragraphs
96,000 个众包、对抗式构造的阅读理解问题,基于维基百科段落,需要加法、计数、排序和日期运算等离散操作,答案为文本片段、数字或日期。以精确匹配和数值感知的 F1(主要指标)评分,通常在 9,535 题的开发集上采用 3-shot。曾是有依据的数值推理的严苛测试;前沿模型如今已接近专家人类的 F1。
- 发布
- 2019-03
- 维护者
- Allen Institute for AI / UC Irvine (Dua et al.)
- 状态
- saturated
- 污染风险
- high
- 指标
- F1 (percent, ↑)
- 题量
- 9,535
- 领域
- reasoning knowledge
- 人类
- 96.4% expert annotators (paper Table 4, test F1)
备注. Model-vendor reports use the dev split (test labels are hidden) and vary in shots and answer normalization; DROP was dropped from most vendor cards after 2025.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| o1 | OpenAI | 90.2% | 厂商自报 | shots: 3 simple-evals README, DROP F1 3-shot; o3-high 89.8 (April 2025). | |
| DeepSeek-V4-Pro-Base | DeepSeek | 88.7% | 厂商自报 | shots: 1 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, F1, 1-shot. | |
| NAQANet (paper baseline) | AI2 / UC Irvine (Dua et al.) | 47% | 论文 | Complete model in Table 4, test F1 47.01; best prior baseline 32.7; expert humans 96.4. |