DROP

DROP: Discrete Reasoning Over the content of Paragraphs

96,000 crowd-sourced, adversarially created reading-comprehension questions over Wikipedia paragraphs that require discrete operations such as addition, counting, sorting and date arithmetic, with span, number or date answers. Scored by exact match and a numerically-aware F1 (the headline metric), usually 3-shot on the 9,535-question dev set. Once a hard test of grounded numerical reasoning; frontier models now approach expert human F1.

90.2% · human 96.4%
o1
self-reported
2019-03-01 → 2026-04-22: 47% → 88.7%
Released
2019-03
Maintainer
Allen Institute for AI / UC Irvine (Dua et al.)
Status
saturated
Contamination
high
Metric
F1 (percent, ↑)
Tasks
9,535
Domains
reasoning knowledge
human
96.4% expert annotators (paper Table 4, test F1)

Notes. Model-vendor reports use the dev split (test labels are hidden) and vary in shots and answer normalization; DROP was dropped from most vendor cards after 2025.

Full ledger

SystemDeveloperScoreDateSourceConditions
o1OpenAI90.2%self-reportedshots: 3
simple-evals README, DROP F1 3-shot; o3-high 89.8 (April 2025).
DeepSeek-V4-Pro-BaseDeepSeek88.7%self-reportedshots: 1
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, F1, 1-shot.
NAQANet (paper baseline)AI2 / UC Irvine (Dua et al.)47%paper
Complete model in Table 4, test F1 47.01; best prior baseline 32.7; expert humans 96.4.