Humanity's Last Exam

Humanity's Last Exam (HLE)

2,500 expert-written closed-ended questions (multiple choice and exact-match short answer, about 14% with images) across mathematics, natural sciences, humanities and more, filtered so frontier models failed them at collection time. Scored as accuracy by an LLM judge (o3-mini) against the reference answer; calibration error is also reported. Designed as the final broad academic benchmark, it is now the main frontier knowledge-reasoning test.

65%
Claude Fable 5.1 (with tools)
self-reported
2025-01-24 → 2026-09-01: 8% → 65%
Released
2025-01
Maintainer
Center for AI Safety / Scale AI
Status
active
Contamination
medium
Metric
accuracy (percent, ↑)
Tasks
2,500
Domains
knowledge reasoning science math multimodal
human
no measured baseline

Notes. Finalized at 2,500 questions on 2025-04-03 after a bug bounty; earlier scores used a different question set. Published in Nature (649, 1139-1146) on 2026-01-28. A private held-out set exists to detect overfitting, and an HLE-Rolling fork was released in October 2025. Tool-assisted (search/code) and no-tools runs are not comparable; the ledger marks tools explicitly.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude Fable 5.1 (with tools)Anthropic65%self-reportedtools
Launch post; also republished by benchlm.ai as the current top. OpenAI's GPT-6 Astra post lists Astra at 57.2 with tools.
Claude Fable 5.1 (no tools)Anthropic60.9%self-reportedtools: no
Launch post comparison table; Fable 5 scored 57.8 and Opus 5 56.6 without tools.
Gemini 3 ProGoogle DeepMind38.3%officialtools: no
lastexam.ai leaderboard, judge o3-mini, finalized April 2025 set; Google's own post reports 37.5%.
GPT-5 (no tools)OpenAI24.8%self-reportedtools: no
With thinking; GPT-5 pro no-tools scored 30.7 and 42.0 with python + search in the same post.
o3-mini (high)OpenAI13.4%papertools: no split: text-only
Table 1; model is not multimodal so scored on text-only questions. Pre-finalization set.
o1OpenAI8%papertools: no
Table 1 of the HLE paper (pre-finalization question set of January 2025).