MMLU-Pro

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

↑ MMLU

About 12,000 questions across 14 disciplines built from MMLU, STEM sites, TheoremQA and SciBench, with the answer set expanded from four to ten options and trivial or noisy items removed. Scored as accuracy with chain-of-thought (5-shot in the official setup). Reduces guessing headroom and prompt sensitivity, making it the standard replacement for MMLU when comparing knowledge and reasoning.

91%
Gemini 3.1 Pro (High)
independent
2024-06-03 → 2026-05-16: 72.6% → 89.6%
Released
2024-06
Maintainer
TIGER-Lab (University of Waterloo)
Status
saturating
Contamination
high
Metric
accuracy (percent, ↑)
Tasks
12,032
Domains
knowledge reasoning
human
no measured baseline

Notes. Developer-reported numbers vary by prompt (CoT vs direct) and shots; the official leaderboard uses 5-shot CoT. A January 2026 formatting fix to answer options can shift older STEM-subset scores slightly.

Full ledger

SystemDeveloperScoreDateSourceConditions
Gemini 3.1 Pro (High)Google DeepMind91%independent
Run by DeepSeek for its V4-Pro model card comparison table; setting not further specified.
Qwen3.7 MaxAlibaba89.6%aggregator
Developer-reported number republished by aggregator; prompt/effort setting unspecified. Date is the model release date shown by benchlm.ai.
DeepSeek-V4-Pro (Think Max)DeepSeek87.5%self-reported
Model card comparison table; the same table lists Gemini-3.1-Pro (High) at 91.0 and Claude Opus 4.6 (Max) at 89.1 as run by DeepSeek.
GPT-4oOpenAI72.6%papershots: 5
Table of the MMLU-Pro paper, CoT prompting.