Humanity's Last Exam
Humanity's Last Exam (HLE)
2,500 expert-written closed-ended questions (multiple choice and exact-match short answer, about 14% with images) across mathematics, natural sciences, humanities and more, filtered so frontier models failed them at collection time. Scored as accuracy by an LLM judge (o3-mini) against the reference answer; calibration error is also reported. Designed as the final broad academic benchmark, it is now the main frontier knowledge-reasoning test.
paper · website · leaderboard · dataset · code
- Released
- 2025-01
- Maintainer
- Center for AI Safety / Scale AI
- Status
- active
- Contamination
- medium
- Metric
- accuracy (percent, ↑)
- Tasks
- 2,500
- Domains
- knowledge reasoning science math multimodal
- human
- no measured baseline
Notes. Finalized at 2,500 questions on 2025-04-03 after a bug bounty; earlier scores used a different question set. Published in Nature (649, 1139-1146) on 2026-01-28. A private held-out set exists to detect overfitting, and an HLE-Rolling fork was released in October 2025. Tool-assisted (search/code) and no-tools runs are not comparable; the ledger marks tools explicitly.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Claude Fable 5.1 (with tools) | Anthropic | 65% | self-reported | tools Launch post; also republished by benchlm.ai as the current top. OpenAI's GPT-6 Astra post lists Astra at 57.2 with tools. | |
| Claude Fable 5.1 (no tools) | Anthropic | 60.9% | self-reported | tools: no Launch post comparison table; Fable 5 scored 57.8 and Opus 5 56.6 without tools. | |
| Gemini 3 Pro | Google DeepMind | 38.3% | official | tools: no lastexam.ai leaderboard, judge o3-mini, finalized April 2025 set; Google's own post reports 37.5%. | |
| GPT-5 (no tools) | OpenAI | 24.8% | self-reported | tools: no With thinking; GPT-5 pro no-tools scored 30.7 and 42.0 with python + search in the same post. | |
| o3-mini (high) | OpenAI | 13.4% | paper | tools: no split: text-only Table 1; model is not multimodal so scored on text-only questions. Pre-finalization set. | |
| o1 | OpenAI | 8% | paper | tools: no Table 1 of the HLE paper (pre-finalization question set of January 2025). |