DROP

DROP: Discrete Reasoning Over the content of Paragraphs

96,000 个众包、对抗式构造的阅读理解问题,基于维基百科段落,需要加法、计数、排序和日期运算等离散操作,答案为文本片段、数字或日期。以精确匹配和数值感知的 F1(主要指标)评分,通常在 9,535 题的开发集上采用 3-shot。曾是有依据的数值推理的严苛测试;前沿模型如今已接近专家人类的 F1。

90.2% · 人类 96.4%
o1
厂商自报
2019-03-01 → 2026-04-22: 47% → 88.7%
发布
2019-03
维护者
Allen Institute for AI / UC Irvine (Dua et al.)
状态
saturated
污染风险
high
指标
F1 (percent, ↑)
题量
9,535
领域
reasoning knowledge
人类
96.4% expert annotators (paper Table 4, test F1)

备注. Model-vendor reports use the dev split (test labels are hidden) and vary in shots and answer normalization; DROP was dropped from most vendor cards after 2025.

完整账本

系统开发者分数日期来源条件
o1OpenAI90.2%厂商自报shots: 3
simple-evals README, DROP F1 3-shot; o3-high 89.8 (April 2025).
DeepSeek-V4-Pro-BaseDeepSeek88.7%厂商自报shots: 1
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, F1, 1-shot.
NAQANet (paper baseline)AI2 / UC Irvine (Dua et al.)47%论文
Complete model in Table 4, test F1 47.01; best prior baseline 32.7; expert humans 96.4.