Humanity's Last Exam
Humanity's Last Exam (HLE)
2,500 道专家编写的封闭式问题(多项选择与精确匹配的简答,约 14% 含图像),覆盖数学、自然科学、人文等领域,经筛选确保前沿模型在收集时无法作答。由 LLM 评判(o3-mini)对照参考答案计算准确率;同时报告校准误差。设计为最后一个广泛的学术基准,如今是前沿知识推理的主要测试。
- 发布
- 2025-01
- 维护者
- Center for AI Safety / Scale AI
- 状态
- active
- 污染风险
- medium
- 指标
- accuracy (percent, ↑)
- 题量
- 2,500
- 领域
- knowledge reasoning science math multimodal
- 人类
- 无实测基线
备注. Finalized at 2,500 questions on 2025-04-03 after a bug bounty; earlier scores used a different question set. Published in Nature (649, 1139-1146) on 2026-01-28. A private held-out set exists to detect overfitting, and an HLE-Rolling fork was released in October 2025. Tool-assisted (search/code) and no-tools runs are not comparable; the ledger marks tools explicitly.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Fable 5.1 (with tools) | Anthropic | 65% | 厂商自报 | tools Launch post; also republished by benchlm.ai as the current top. OpenAI's GPT-6 Astra post lists Astra at 57.2 with tools. | |
| Claude Fable 5.1 (no tools) | Anthropic | 60.9% | 厂商自报 | tools: no Launch post comparison table; Fable 5 scored 57.8 and Opus 5 56.6 without tools. | |
| Gemini 3 Pro | Google DeepMind | 38.3% | 官方榜单 | tools: no lastexam.ai leaderboard, judge o3-mini, finalized April 2025 set; Google's own post reports 37.5%. | |
| GPT-5 (no tools) | OpenAI | 24.8% | 厂商自报 | tools: no With thinking; GPT-5 pro no-tools scored 30.7 and 42.0 with python + search in the same post. | |
| o3-mini (high) | OpenAI | 13.4% | 论文 | tools: no split: text-only Table 1; model is not multimodal so scored on text-only questions. Pre-finalization set. | |
| o1 | OpenAI | 8% | 论文 | tools: no Table 1 of the HLE paper (pre-finalization question set of January 2025). |