HellaSwag
HellaSwag: Can a Machine Really Finish Your Sentence?
Commonsense sentence-completion: given a context from ActivityNet captions or WikiHow, pick the most plausible continuation among four options, where the wrong endings were adversarially filtered to fool 2019-era models. 10,042 validation and 10,003 test items, scored as accuracy (typically 10-shot). Trivial for humans and once a headline pretraining benchmark; modern models exceed 95%, so it no longer separates systems.
- Released
- 2019-05
- Maintainer
- Allen Institute for AI / University of Washington (Zellers et al.)
- Status
- saturated
- Contamination
- high
- Metric
- accuracy (percent, ↑)
- Tasks
- 10,042
- Domains
- commonsense
- human
- 95.6% crowd workers (paper Table 1, overall)
Notes. Reported scores use the validation split because test labels are hidden. Roughly a third of items contain grammatical or labeling errors, which caps meaningful accuracy.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro-Base | DeepSeek | 88% | self-reported | shots: 0 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match, zero-shot; not comparable with 10-shot harness numbers. | |
| BERT-Large | Google (evaluated by Zellers et al.) | 47.3% | paper | Table 1 of the HellaSwag paper, overall test accuracy; humans 95.6. |