WinoGrande
WinoGrande: An Adversarial Winograd Schema Challenge at Scale
44,000 个受 Winograd Schema Challenge 启发的二选一填空代词消解问题,众包构建后经 AfLite 对抗式过滤去偏,以消除词语关联捷径。以 1,267 项验证集上的准确率评分(测试集标签隐藏)。是预训练评估中长期使用的常识探针;前沿模型与 94% 的人类水平仅差几分。
- 发布
- 2019-07
- 维护者
- Allen Institute for AI (Sakaguchi et al.)
- 状态
- saturated
- 污染风险
- high
- 指标
- accuracy (percent, ↑)
- 题量
- 1,267
- 领域
- commonsense reasoning
- 人类
- 94% crowd workers (paper Table 3, test set)
备注. Most harnesses report 5-shot accuracy on the validation split of the 'xl' (40k-train) configuration.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro-Base | DeepSeek | 81.5% | 厂商自报 | shots: 0 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot. | |
| RoBERTa (fine-tuned) | Facebook AI (evaluated by Sakaguchi et al.) | 79.1% | 论文 | Table 3 of the WinoGrande paper, test accuracy on the debiased set; human performance 94.0. |