DROP
DROP: Discrete Reasoning Over the content of Paragraphs
96,000 crowd-sourced, adversarially created reading-comprehension questions over Wikipedia paragraphs that require discrete operations such as addition, counting, sorting and date arithmetic, with span, number or date answers. Scored by exact match and a numerically-aware F1 (the headline metric), usually 3-shot on the 9,535-question dev set. Once a hard test of grounded numerical reasoning; frontier models now approach expert human F1.
- Released
- 2019-03
- Maintainer
- Allen Institute for AI / UC Irvine (Dua et al.)
- Status
- saturated
- Contamination
- high
- Metric
- F1 (percent, ↑)
- Tasks
- 9,535
- Domains
- reasoning knowledge
- human
- 96.4% expert annotators (paper Table 4, test F1)
Notes. Model-vendor reports use the dev split (test labels are hidden) and vary in shots and answer normalization; DROP was dropped from most vendor cards after 2025.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| o1 | OpenAI | 90.2% | self-reported | shots: 3 simple-evals README, DROP F1 3-shot; o3-high 89.8 (April 2025). | |
| DeepSeek-V4-Pro-Base | DeepSeek | 88.7% | self-reported | shots: 1 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, F1, 1-shot. | |
| NAQANet (paper baseline) | AI2 / UC Irvine (Dua et al.) | 47% | paper | Complete model in Table 4, test F1 47.01; best prior baseline 32.7; expert humans 96.4. |