MMLU

Measuring Massive Multitask Language Understanding

↓ MMLU-Pro

14,042 道四选一多项选择测试题,覆盖 57 个学科,从初等数学、美国历史到法律和医学,取自考试和教科书(含开发集与验证集共 15,908 题)。以准确率评分,历史上采用 5-shot。三年来一直是广泛世界知识的默认衡量标准;前沿模型如今集中在 90% 以上,测试集的标注噪声限制了进一步区分。

91.8% · 人类 89.8%
o1
厂商自报
2020-09-07 → 2026-04-22: 43.9% → 90.1%
发布
2020-09
维护者
Dan Hendrycks et al. (UC Berkeley)
状态
saturated
污染风险
high
指标
accuracy (percent, ↑)
题量
14,042
领域
knowledge reasoning
人类
89.8% estimated expert-level test takers (paper estimate from source-exam pass rates)

备注. The paper's 89.8% expert figure is an estimate, not a measured panel; unspecialized crowd workers scored 34.5%. Roughly 6.5% of questions are estimated to contain errors (MMLU-Redux), so scores above ~90% are not reliably comparable.

完整账本

系统开发者分数日期来源条件
o1OpenAI91.8%厂商自报shots: 0
simple-evals README; o3-high later scored 93.3 (April 2025).
DeepSeek-V4-Pro-BaseDeepSeek90.1%厂商自报shots: 5
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base (pre-trained) model, exact match.
gpt-4-turbo-2024-04-09OpenAI86.7%厂商自报shots: 0
simple-evals README (zero-shot chain-of-thought).
GPT-3 175B (few-shot)OpenAI43.9%论文shots: 5
Table 1 of the MMLU paper (GPT-3 X-Large).