MMLU
Measuring Massive Multitask Language Understanding
14,042 道四选一多项选择测试题,覆盖 57 个学科,从初等数学、美国历史到法律和医学,取自考试和教科书(含开发集与验证集共 15,908 题)。以准确率评分,历史上采用 5-shot。三年来一直是广泛世界知识的默认衡量标准;前沿模型如今集中在 90% 以上,测试集的标注噪声限制了进一步区分。
- 发布
- 2020-09
- 维护者
- Dan Hendrycks et al. (UC Berkeley)
- 状态
- saturated
- 污染风险
- high
- 指标
- accuracy (percent, ↑)
- 题量
- 14,042
- 领域
- knowledge reasoning
- 人类
- 89.8% estimated expert-level test takers (paper estimate from source-exam pass rates)
备注. The paper's 89.8% expert figure is an estimate, not a measured panel; unspecialized crowd workers scored 34.5%. Roughly 6.5% of questions are estimated to contain errors (MMLU-Redux), so scores above ~90% are not reliably comparable.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| o1 | OpenAI | 91.8% | 厂商自报 | shots: 0 simple-evals README; o3-high later scored 93.3 (April 2025). | |
| DeepSeek-V4-Pro-Base | DeepSeek | 90.1% | 厂商自报 | shots: 5 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base (pre-trained) model, exact match. | |
| gpt-4-turbo-2024-04-09 | OpenAI | 86.7% | 厂商自报 | shots: 0 simple-evals README (zero-shot chain-of-thought). | |
| GPT-3 175B (few-shot) | OpenAI | 43.9% | 论文 | shots: 5 Table 1 of the MMLU paper (GPT-3 X-Large). |