GSM8K

Grade School Math 8K

8,500 道语言多样的小学数学应用题(7,473 训练、1,319 测试),需要两到八步基础算术求解,并附自然语言分步解答。以测试集上最终数值答案的精确匹配评分。它推动了思维链研究,是 2021 至 2024 年的标准数学推理检验;前沿模型如今已超过 95%,错误主要来自标注噪声。

92.6%
DeepSeek-V4-Pro-Base
厂商自报
2021-10-27 → 2026-04-22: 20.6% → 92.6%
发布
2021-10
维护者
OpenAI (Cobbe et al.)
状态
saturated
污染风险
high
指标
accuracy (percent, ↑)
题量
1,319
领域
math reasoning
人类
无实测基线

备注. Roughly 1-2% of test items are believed mislabeled, so scores above ~97% are within noise. Few-shot count (0, 5 or 8) and use of a calculator or code affect comparability.

完整账本

系统开发者分数日期来源条件
DeepSeek-V4-Pro-BaseDeepSeek92.6%厂商自报shots: 8
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base model, exact match; chat models are no longer reported on GSM8K by most vendors.
GPT-3 6B (finetuned)OpenAI20.6%论文
Section 4.1 of the GSM8K paper: a 6B model finetuned with full natural-language solutions reaches 20.6% test accuracy (5.2% when outputting the answer directly). Larger 175B results are given only in figures.