MATH

MATH: Measuring Mathematical Problem Solving

12,500 道竞赛数学题(7,500 训练、5,000 测试),来自 AMC 10/12、AIME 及类似竞赛,涵盖七个学科和五个难度等级,每题附完整的 LaTeX 分步解答。归一化后以加框最终答案的精确匹配评分。在推理模型将其推过 95% 之前,它是主要的高难数学基准;最难的 Level 5 子集和 500 题的 MATH-500 划分仍在使用。

98.1% · 人类 90%
o3-high
厂商自报
2021-03-05 → 2025-04-16: 5.2% → 98.1%
发布
2021-03
维护者
Dan Hendrycks et al. (UC Berkeley)
状态
saturated
污染风险
high
指标
accuracy (percent, ↑)
题量
5,000
领域
math reasoning
人类
90% three-time IMO gold medalist (single participant; a CS PhD student scored about 40%)

备注. Answer-equivalence checking differs across harnesses (string match vs sympy vs LLM grader) and can move scores by several points. The human numbers in the paper are anecdotal single-person measurements.

完整账本

系统开发者分数日期来源条件
o3-highOpenAI98.1%厂商自报split: MATH-500 shots: 0
simple-evals README; o4-mini-high 98.2 on the same table. Epoch's Level-5 run of gpt-5 (high) reached 98.1 in October 2025.
o1-2024-12-17 (medium)OpenAI94.4%独立复现split: Level 5
Epoch AI Benchmarking Hub run of 2025-01-27 on the 1,324 Level-5 test problems (data export).
gpt-4-turbo-2024-04-09OpenAI73.4%厂商自报split: MATH-500 shots: 0
simple-evals README; the MATH column is the 500-problem subset, zero-shot CoT.
GPT-3 175B (few-shot)OpenAI5.2%论文
Table 2 of the MATH paper, overall accuracy (175B*, few-shot). The paper's human anchors: CS PhD student about 40%, three-time IMO gold medalist 90%.