FrontierMath

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI (Tiers 1-4)

数百道由专家数学家编写并审核的原创、未发表的研究级数学问题,难度从高难本科(Tier 1)经高阶研究生(Tier 3)到研究水平的 Tier 4,每题答案均可自动验证。模型可在 token 预算内推理并运行 Python,提交答案函数;得分为解出问题的比例。为避免污染而保持私有,是前沿数学推理的主要衡量标准,所有评估均由 Epoch 自行运行。

93.7%
gpt-6-astra (max)
独立复现
2025-03-06 → 2026-09-03: 0.3% → 93.7%
发布
2024-11
维护者
Epoch AI
状态
saturating
污染风险
low
指标
accuracy (percent, ↑)
题量
338
领域
math reasoning research
人类
无实测基线

备注. OpenAI funded the benchmark and holds exclusive access to a subset (Epoch's conflict-of-interest statement). On 2026-06-12 Epoch released v2, correcting errors in 42% of problems and removing 12; v1 and v2 scores are separate series. Epoch also runs FrontierMath Open Problems and FrontierMath Erdos, which are different benchmarks.

完整账本

系统开发者分数日期来源条件
gpt-6-astra (max)OpenAI93.7%独立复现split: Tiers 1-3 (v2) tools
Epoch AI Benchmarking Hub (run started 2026-08-30, published with the model on 2026-09-03), stderr 1.4; Tier 4 (v2): 95.1 at max, 97.6 at medium effort.
gpt-5.6-sol (max)OpenAI89.1%独立复现split: Tiers 1-3 (v2) tools
Epoch AI Benchmarking Hub, stderr 1.8.
claude-fable-5 (max)Anthropic87%独立复现split: Tiers 1-3 (v2) tools
First v2 run after the 2026-06-12 correction (Epoch data export, started 2026-06-09).
gpt-5.4-pro-2026-03-05 (xhigh)OpenAI50%独立复现split: Tiers 1-3 (v1) tools
Epoch run on v1 problem set; the highest v1 score was gpt-5.5-pro (high) at 52.4 on 2026-04-23.
gpt-4o-2024-11-20OpenAI0.3%独立复现split: Tiers 1-3 (v1) tools
Epoch AI Benchmarking Hub run of 2025-03-06 (data export); the Nov 2024 paper reported all models under 2%.