HellaSwag

HellaSwag: Can a Machine Really Finish Your Sentence?

常识句子补全:给定来自 ActivityNet 字幕或 WikiHow 的上下文,从四个选项中选出最合理的续写,其中错误结尾经过对抗式筛选以迷惑 2019 年前后的模型。包含 10,042 个验证项和 10,003 个测试项,以准确率评分(通常 10-shot)。对人类而言极其简单,曾是预训练的头条基准;现代模型已超过 95%,不再能区分各系统。

88% · 人类 95.6%
DeepSeek-V4-Pro-Base
厂商自报
2019-05-19 → 2026-04-22: 47.3% → 88%
发布
2019-05
维护者
Allen Institute for AI / University of Washington (Zellers et al.)
状态
saturated
污染风险
high
指标
accuracy (percent, ↑)
题量
10,042
领域
commonsense
人类
95.6% crowd workers (paper Table 1, overall)

备注. Reported scores use the validation split because test labels are hidden. Roughly a third of items contain grammatical or labeling errors, which caps meaningful accuracy.

完整账本

系统开发者分数日期来源条件
DeepSeek-V4-Pro-BaseDeepSeek88%厂商自报shots: 0
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot; not comparable with 10-shot harness numbers.
BERT-LargeGoogle (evaluated by Zellers et al.)47.3%论文
Table 1 of the HellaSwag paper, overall test accuracy; humans 95.6.