FrontierMath
FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI (Tiers 1-4)
Hundreds of original, unpublished research-grade mathematics problems written and vetted by expert mathematicians, from hard undergraduate (Tier 1) through advanced graduate (Tier 3) to research-level Tier 4, each with an automatically verifiable answer. Models may reason and run Python within a token budget and submit an answer function; scored as the fraction of problems solved. Kept private to avoid contamination, it is the main measure of frontier mathematical reasoning, and Epoch runs all evaluations itself.
paper · website · leaderboard · dataset
- Released
- 2024-11
- Maintainer
- Epoch AI
- Status
- Contamination
- low
- Metric
- accuracy (percent, ↑)
- Tasks
- 338
- Domains
- math reasoning research
- human
- no measured baseline
Notes. OpenAI funded the benchmark and holds exclusive access to a subset (Epoch's conflict-of-interest statement). On 2026-06-12 Epoch released v2, correcting errors in 42% of problems and removing 12; v1 and v2 scores are separate series. Epoch also runs FrontierMath Open Problems and FrontierMath Erdos, which are different benchmarks.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| gpt-6-astra (max) | OpenAI | 93.7% | independent | split: Tiers 1-3 (v2) tools Epoch AI Benchmarking Hub (run started 2026-08-30, published with the model on 2026-09-03), stderr 1.4; Tier 4 (v2): 95.1 at max, 97.6 at medium effort. | |
| gpt-5.6-sol (max) | OpenAI | 89.1% | independent | split: Tiers 1-3 (v2) tools Epoch AI Benchmarking Hub, stderr 1.8. | |
| claude-fable-5 (max) | Anthropic | 87% | independent | split: Tiers 1-3 (v2) tools First v2 run after the 2026-06-12 correction (Epoch data export, started 2026-06-09). | |
| gpt-5.4-pro-2026-03-05 (xhigh) | OpenAI | 50% | independent | split: Tiers 1-3 (v1) tools Epoch run on v1 problem set; the highest v1 score was gpt-5.5-pro (high) at 52.4 on 2026-04-23. | |
| gpt-4o-2024-11-20 | OpenAI | 0.3% | independent | split: Tiers 1-3 (v1) tools Epoch AI Benchmarking Hub run of 2025-03-06 (data export); the Nov 2024 paper reported all models under 2%. |