MMLU
Measuring Massive Multitask Language Understanding
14,042 four-option multiple-choice test questions across 57 subjects from elementary mathematics and US history to law and medicine, drawn from exams and textbooks (15,908 questions in total including dev and validation). Scored as accuracy, historically 5-shot. For three years it was the default measure of broad world knowledge; frontier models now cluster above 90%, and label noise in the test set limits further discrimination.
- Released
- 2020-09
- Maintainer
- Dan Hendrycks et al. (UC Berkeley)
- Status
- saturated
- Contamination
- high
- Metric
- accuracy (percent, ↑)
- Tasks
- 14,042
- Domains
- knowledge reasoning
- human
- 89.8% estimated expert-level test takers (paper estimate from source-exam pass rates)
Notes. The paper's 89.8% expert figure is an estimate, not a measured panel; unspecialized crowd workers scored 34.5%. Roughly 6.5% of questions are estimated to contain errors (MMLU-Redux), so scores above ~90% are not reliably comparable.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| o1 | OpenAI | 91.8% | self-reported | shots: 0 simple-evals README; o3-high later scored 93.3 (April 2025). | |
| DeepSeek-V4-Pro-Base | DeepSeek | 90.1% | self-reported | shots: 5 DeepSeek-V4 model card, base-model table (HF repo created 2026-04-22). Base (pre-trained) model, exact match. | |
| gpt-4-turbo-2024-04-09 | OpenAI | 86.7% | self-reported | shots: 0 simple-evals README (zero-shot chain-of-thought). | |
| GPT-3 175B (few-shot) | OpenAI | 43.9% | paper | shots: 5 Table 1 of the MMLU paper (GPT-3 X-Large). |