WinoGrande
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
44,000 binary fill-in-the-blank pronoun-resolution problems inspired by the Winograd Schema Challenge, crowd-sourced and then debiased with the AfLite adversarial filter to remove word-association shortcuts. Scored as accuracy on the 1,267-item validation set (test labels are hidden). A long-running commonsense probe in pretraining evaluations; frontier models are within a few points of the 94% human figure.
- Released
- 2019-07
- Maintainer
- Allen Institute for AI (Sakaguchi et al.)
- Status
- saturated
- Contamination
- high
- Metric
- accuracy (percent, ↑)
- Tasks
- 1,267
- Domains
- commonsense reasoning
- human
- 94% crowd workers (paper Table 3, test set)
Notes. Most harnesses report 5-shot accuracy on the validation split of the 'xl' (40k-train) configuration.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro-Base | DeepSeek | 81.5% | self-reported | shots: 0 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot. | |
| RoBERTa (fine-tuned) | Facebook AI (evaluated by Sakaguchi et al.) | 79.1% | paper | Table 3 of the WinoGrande paper, test accuracy on the debiased set; human performance 94.0. |