GSM8K

Grade School Math 8K

8,500 linguistically diverse grade-school math word problems (7,473 train, 1,319 test) that take two to eight elementary arithmetic steps to solve, with natural-language step-by-step solutions. Scored as exact match of the final numeric answer on the test split. It drove chain-of-thought research and was the standard math-reasoning check from 2021 to 2024; frontier models now exceed 95%, and errors are dominated by label noise.

92.6%
DeepSeek-V4-Pro-Base
self-reported
2021-10-27 → 2026-04-22: 20.6% → 92.6%
Released
2021-10
Maintainer
OpenAI (Cobbe et al.)
Status
saturated
Contamination
high
Metric
accuracy (percent, ↑)
Tasks
1,319
Domains
math reasoning
human
no measured baseline

Notes. Roughly 1-2% of test items are believed mislabeled, so scores above ~97% are within noise. Few-shot count (0, 5 or 8) and use of a calculator or code affect comparability.

Full ledger

SystemDeveloperScoreDateSourceConditions
DeepSeek-V4-Pro-BaseDeepSeek92.6%self-reportedshots: 8
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match; chat models are no longer reported on GSM8K by most vendors.
GPT-3 6B (finetuned)OpenAI20.6%paper
Section 4.1 of the GSM8K paper: a 6B model finetuned with full natural-language solutions reaches 20.6% test accuracy (5.2% when outputting the answer directly). Larger 175B results are given only in figures.