FrontierMath
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI (Tiers 1-4)
数百道由专家数学家编写并审核的原创、未发表的研究级数学问题,难度从高难本科(Tier 1)经高阶研究生(Tier 3)到研究水平的 Tier 4,每题答案均可自动验证。模型可在 token 预算内推理并运行 Python,提交答案函数;得分为解出问题的比例。为避免污染而保持私有,是前沿数学推理的主要衡量标准,所有评估均由 Epoch 自行运行。
- 发布
- 2024-11
- 维护者
- Epoch AI
- 状态
- 污染风险
- low
- 指标
- accuracy (percent, ↑)
- 题量
- 338
- 领域
- math reasoning research
- 人类
- 无实测基线
备注. OpenAI funded the benchmark and holds exclusive access to a subset (Epoch's conflict-of-interest statement). On 2026-06-12 Epoch released v2, correcting errors in 42% of problems and removing 12; v1 and v2 scores are separate series. Epoch also runs FrontierMath Open Problems and FrontierMath Erdos, which are different benchmarks.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| gpt-6-astra (max) | OpenAI | 93.7% | 独立复现 | split: Tiers 1-3 (v2) tools Epoch AI Benchmarking Hub (run started 2026-08-30, published with the model on 2026-09-03), stderr 1.4; Tier 4 (v2): 95.1 at max, 97.6 at medium effort. | |
| gpt-5.6-sol (max) | OpenAI | 89.1% | 独立复现 | split: Tiers 1-3 (v2) tools Epoch AI Benchmarking Hub, stderr 1.8. | |
| claude-fable-5 (max) | Anthropic | 87% | 独立复现 | split: Tiers 1-3 (v2) tools First v2 run after the 2026-06-12 correction (Epoch data export, started 2026-06-09). | |
| gpt-5.4-pro-2026-03-05 (xhigh) | OpenAI | 50% | 独立复现 | split: Tiers 1-3 (v1) tools Epoch run on v1 problem set; the highest v1 score was gpt-5.5-pro (high) at 52.4 on 2026-04-23. | |
| gpt-4o-2024-11-20 | OpenAI | 0.3% | 独立复现 | split: Tiers 1-3 (v1) tools Epoch AI Benchmarking Hub run of 2025-03-06 (data export); the Nov 2024 paper reported all models under 2%. |