HellaSwag

HellaSwag: Can a Machine Really Finish Your Sentence?

Commonsense sentence-completion: given a context from ActivityNet captions or WikiHow, pick the most plausible continuation among four options, where the wrong endings were adversarially filtered to fool 2019-era models. 10,042 validation and 10,003 test items, scored as accuracy (typically 10-shot). Trivial for humans and once a headline pretraining benchmark; modern models exceed 95%, so it no longer separates systems.

88% · human 95.6%
DeepSeek-V4-Pro-Base
self-reported
2019-05-19 → 2026-04-22: 47.3% → 88%
Released
2019-05
Maintainer
Allen Institute for AI / University of Washington (Zellers et al.)
Status
saturated
Contamination
high
Metric
accuracy (percent, ↑)
Tasks
10,042
Domains
commonsense
human
95.6% crowd workers (paper Table 1, overall)

Notes. Reported scores use the validation split because test labels are hidden. Roughly a third of items contain grammatical or labeling errors, which caps meaningful accuracy.

Full ledger

SystemDeveloperScoreDateSourceConditions
DeepSeek-V4-Pro-BaseDeepSeek88%self-reportedshots: 0
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot; not comparable with 10-shot harness numbers.
BERT-LargeGoogle (evaluated by Zellers et al.)47.3%paper
Table 1 of the HellaSwag paper, overall test accuracy; humans 95.6.