MMLU-Pro
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
About 12,000 questions across 14 disciplines built from MMLU, STEM sites, TheoremQA and SciBench, with the answer set expanded from four to ten options and trivial or noisy items removed. Scored as accuracy with chain-of-thought (5-shot in the official setup). Reduces guessing headroom and prompt sensitivity, making it the standard replacement for MMLU when comparing knowledge and reasoning.
paper · website · leaderboard · dataset · code
- Released
- 2024-06
- Maintainer
- TIGER-Lab (University of Waterloo)
- Status
- Contamination
- high
- Metric
- accuracy (percent, ↑)
- Tasks
- 12,032
- Domains
- knowledge reasoning
- human
- no measured baseline
Notes. Developer-reported numbers vary by prompt (CoT vs direct) and shots; the official leaderboard uses 5-shot CoT. A January 2026 formatting fix to answer options can shift older STEM-subset scores slightly.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Gemini 3.1 Pro (High) | Google DeepMind | 91% | independent | Run by DeepSeek for its V4-Pro model card comparison table; setting not further specified. | |
| Qwen3.7 Max | Alibaba | 89.6% | aggregator | Developer-reported number republished by aggregator; prompt/effort setting unspecified. Date is the model release date shown by benchlm.ai. | |
| DeepSeek-V4-Pro (Think Max) | DeepSeek | 87.5% | self-reported | Model card comparison table; the same table lists Gemini-3.1-Pro (High) at 91.0 and Claude Opus 4.6 (Max) at 89.1 as run by DeepSeek. | |
| GPT-4o | OpenAI | 72.6% | paper | shots: 5 Table of the MMLU-Pro paper, CoT prompting. |