WinoGrande

WinoGrande: An Adversarial Winograd Schema Challenge at Scale

44,000 个受 Winograd Schema Challenge 启发的二选一填空代词消解问题,众包构建后经 AfLite 对抗式过滤去偏,以消除词语关联捷径。以 1,267 项验证集上的准确率评分(测试集标签隐藏)。是预训练评估中长期使用的常识探针;前沿模型与 94% 的人类水平仅差几分。

81.5% · 人类 94%
DeepSeek-V4-Pro-Base
厂商自报
2019-07-24 → 2026-04-22: 79.1% → 81.5%
发布
2019-07
维护者
Allen Institute for AI (Sakaguchi et al.)
状态
saturated
污染风险
high
指标
accuracy (percent, ↑)
题量
1,267
领域
commonsense reasoning
人类
94% crowd workers (paper Table 3, test set)

备注. Most harnesses report 5-shot accuracy on the validation split of the 'xl' (40k-train) configuration.

完整账本

系统开发者分数日期来源条件
DeepSeek-V4-Pro-BaseDeepSeek81.5%厂商自报shots: 0
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot.
RoBERTa (fine-tuned)Facebook AI (evaluated by Sakaguchi et al.)79.1%论文
Table 3 of the WinoGrande paper, test accuracy on the debiased set; human performance 94.0.