FrontierMath

FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI (Tiers 1-4)

Hundreds of original, unpublished research-grade mathematics problems written and vetted by expert mathematicians, from hard undergraduate (Tier 1) through advanced graduate (Tier 3) to research-level Tier 4, each with an automatically verifiable answer. Models may reason and run Python within a token budget and submit an answer function; scored as the fraction of problems solved. Kept private to avoid contamination, it is the main measure of frontier mathematical reasoning, and Epoch runs all evaluations itself.

93.7%
gpt-6-astra (max)
independent
2025-03-06 → 2026-09-03: 0.3% → 93.7%
Released
2024-11
Maintainer
Epoch AI
Status
saturating
Contamination
low
Metric
accuracy (percent, ↑)
Tasks
338
Domains
math reasoning research
human
no measured baseline

Notes. OpenAI funded the benchmark and holds exclusive access to a subset (Epoch's conflict-of-interest statement). On 2026-06-12 Epoch released v2, correcting errors in 42% of problems and removing 12; v1 and v2 scores are separate series. Epoch also runs FrontierMath Open Problems and FrontierMath Erdos, which are different benchmarks.

Full ledger

SystemDeveloperScoreDateSourceConditions
gpt-6-astra (max)OpenAI93.7%independentsplit: Tiers 1-3 (v2) tools
Epoch AI Benchmarking Hub (run started 2026-08-30, published with the model on 2026-09-03), stderr 1.4; Tier 4 (v2): 95.1 at max, 97.6 at medium effort.
gpt-5.6-sol (max)OpenAI89.1%independentsplit: Tiers 1-3 (v2) tools
Epoch AI Benchmarking Hub, stderr 1.8.
claude-fable-5 (max)Anthropic87%independentsplit: Tiers 1-3 (v2) tools
First v2 run after the 2026-06-12 correction (Epoch data export, started 2026-06-09).
gpt-5.4-pro-2026-03-05 (xhigh)OpenAI50%independentsplit: Tiers 1-3 (v1) tools
Epoch run on v1 problem set; the highest v1 score was gpt-5.5-pro (high) at 52.4 on 2026-04-23.
gpt-4o-2024-11-20OpenAI0.3%independentsplit: Tiers 1-3 (v1) tools
Epoch AI Benchmarking Hub run of 2025-03-06 (data export); the Nov 2024 paper reported all models under 2%.