GSM8K
Grade School Math 8K
8,500 linguistically diverse grade-school math word problems (7,473 train, 1,319 test) that take two to eight elementary arithmetic steps to solve, with natural-language step-by-step solutions. Scored as exact match of the final numeric answer on the test split. It drove chain-of-thought research and was the standard math-reasoning check from 2021 to 2024; frontier models now exceed 95%, and errors are dominated by label noise.
- Released
- 2021-10
- Maintainer
- OpenAI (Cobbe et al.)
- Status
- saturated
- Contamination
- high
- Metric
- accuracy (percent, ↑)
- Tasks
- 1,319
- Domains
- math reasoning
- human
- no measured baseline
Notes. Roughly 1-2% of test items are believed mislabeled, so scores above ~97% are within noise. Few-shot count (0, 5 or 8) and use of a calculator or code affect comparability.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro-Base | DeepSeek | 92.6% | self-reported | shots: 8 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match; chat models are no longer reported on GSM8K by most vendors. | |
| GPT-3 6B (finetuned) | OpenAI | 20.6% | paper | Section 4.1 of the GSM8K paper: a 6B model finetuned with full natural-language solutions reaches 20.6% test accuracy (5.2% when outputting the answer directly). Larger 175B results are given only in figures. |