WinoGrande

WinoGrande: An Adversarial Winograd Schema Challenge at Scale

44,000 binary fill-in-the-blank pronoun-resolution problems inspired by the Winograd Schema Challenge, crowd-sourced and then debiased with the AfLite adversarial filter to remove word-association shortcuts. Scored as accuracy on the 1,267-item validation set (test labels are hidden). A long-running commonsense probe in pretraining evaluations; frontier models are within a few points of the 94% human figure.

81.5% · human 94%
DeepSeek-V4-Pro-Base
self-reported
2019-07-24 → 2026-04-22: 79.1% → 81.5%
Released
2019-07
Maintainer
Allen Institute for AI (Sakaguchi et al.)
Status
saturated
Contamination
high
Metric
accuracy (percent, ↑)
Tasks
1,267
Domains
commonsense reasoning
human
94% crowd workers (paper Table 3, test set)

Notes. Most harnesses report 5-shot accuracy on the validation split of the 'xl' (40k-train) configuration.

Full ledger

SystemDeveloperScoreDateSourceConditions
DeepSeek-V4-Pro-BaseDeepSeek81.5%self-reportedshots: 0
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot.
RoBERTa (fine-tuned)Facebook AI (evaluated by Sakaguchi et al.)79.1%paper
Table 3 of the WinoGrande paper, test accuracy on the debiased set; human performance 94.0.