MMLU

Measuring Massive Multitask Language Understanding

↓ MMLU-Pro

14,042 four-option multiple-choice test questions across 57 subjects from elementary mathematics and US history to law and medicine, drawn from exams and textbooks (15,908 questions in total including dev and validation). Scored as accuracy, historically 5-shot. For three years it was the default measure of broad world knowledge; frontier models now cluster above 90%, and label noise in the test set limits further discrimination.

91.8% · human 89.8%
o1
self-reported
2020-09-07 → 2026-04-22: 43.9% → 90.1%
Released
2020-09
Maintainer
Dan Hendrycks et al. (UC Berkeley)
Status
saturated
Contamination
high
Metric
accuracy (percent, ↑)
Tasks
14,042
Domains
knowledge reasoning
human
89.8% estimated expert-level test takers (paper estimate from source-exam pass rates)

Notes. The paper's 89.8% expert figure is an estimate, not a measured panel; unspecialized crowd workers scored 34.5%. Roughly 6.5% of questions are estimated to contain errors (MMLU-Redux), so scores above ~90% are not reliably comparable.

Full ledger

SystemDeveloperScoreDateSourceConditions
o1OpenAI91.8%self-reportedshots: 0
simple-evals README; o3-high later scored 93.3 (April 2025).
DeepSeek-V4-Pro-BaseDeepSeek90.1%self-reportedshots: 5
DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base (pre-trained) model, exact match.
gpt-4-turbo-2024-04-09OpenAI86.7%self-reportedshots: 0
simple-evals README (zero-shot chain-of-thought).
GPT-3 175B (few-shot)OpenAI43.9%papershots: 5
Table 1 of the MMLU paper (GPT-3 X-Large).