HellaSwag
HellaSwag: Can a Machine Really Finish Your Sentence?
常识句子补全:给定来自 ActivityNet 字幕或 WikiHow 的上下文,从四个选项中选出最合理的续写,其中错误结尾经过对抗式筛选以迷惑 2019 年前后的模型。包含 10,042 个验证项和 10,003 个测试项,以准确率评分(通常 10-shot)。对人类而言极其简单,曾是预训练的头条基准;现代模型已超过 95%,不再能区分各系统。
- 发布
- 2019-05
- 维护者
- Allen Institute for AI / University of Washington (Zellers et al.)
- 状态
- saturated
- 污染风险
- high
- 指标
- accuracy (percent, ↑)
- 题量
- 10,042
- 领域
- commonsense
- 人类
- 95.6% crowd workers (paper Table 1, overall)
备注. Reported scores use the validation split because test labels are hidden. Roughly a third of items contain grammatical or labeling errors, which caps meaningful accuracy.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro-Base | DeepSeek | 88% | 厂商自报 | shots: 0 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot; not comparable with 10-shot harness numbers. | |
| BERT-Large | Google (evaluated by Zellers et al.) | 47.3% | 论文 | Table 1 of the HellaSwag paper, overall test accuracy; humans 95.6. |