Humanity's Last Exam

Humanity's Last Exam (HLE)

2,500 道专家编写的封闭式问题(多项选择与精确匹配的简答,约 14% 含图像),覆盖数学、自然科学、人文等领域,经筛选确保前沿模型在收集时无法作答。由 LLM 评判(o3-mini)对照参考答案计算准确率;同时报告校准误差。设计为最后一个广泛的学术基准,如今是前沿知识推理的主要测试。

65%
Claude Fable 5.1 (with tools)
厂商自报
2025-01-24 → 2026-09-01: 8% → 65%
发布
2025-01
维护者
Center for AI Safety / Scale AI
状态
active
污染风险
medium
指标
accuracy (percent, ↑)
题量
2,500
领域
knowledge reasoning science math multimodal
人类
无实测基线

备注. Finalized at 2,500 questions on 2025-04-03 after a bug bounty; earlier scores used a different question set. Published in Nature (649, 1139-1146) on 2026-01-28. A private held-out set exists to detect overfitting, and an HLE-Rolling fork was released in October 2025. Tool-assisted (search/code) and no-tools runs are not comparable; the ledger marks tools explicitly.

完整账本

系统开发者分数日期来源条件
Claude Fable 5.1 (with tools)Anthropic65%厂商自报tools
Launch post; also republished by benchlm.ai as the current top. OpenAI's GPT-6 Astra post lists Astra at 57.2 with tools.
Claude Fable 5.1 (no tools)Anthropic60.9%厂商自报tools: no
Launch post comparison table; Fable 5 scored 57.8 and Opus 5 56.6 without tools.
Gemini 3 ProGoogle DeepMind38.3%官方榜单tools: no
lastexam.ai leaderboard, judge o3-mini, finalized April 2025 set; Google's own post reports 37.5%.
GPT-5 (no tools)OpenAI24.8%厂商自报tools: no
With thinking; GPT-5 pro no-tools scored 30.7 and 42.0 with python + search in the same post.
o3-mini (high)OpenAI13.4%论文tools: no split: text-only
Table 1; model is not multimodal so scored on text-only questions. Pre-finalization set.
o1OpenAI8%论文tools: no
Table 1 of the HLE paper (pre-finalization question set of January 2025).