MMLU-Pro

MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

↑ MMLU

约 12,000 道题,覆盖 14 个学科,基于 MMLU、STEM 网站、TheoremQA 和 SciBench 构建,选项从四个扩展到十个,并移除了琐碎或有噪声的题目。以思维链下的准确率评分(官方设置为 5-shot)。减少了猜测空间和提示敏感性,成为比较知识与推理时 MMLU 的标准替代品。

91%
Gemini 3.1 Pro (High)
独立复现
2024-06-03 → 2026-05-16: 72.6% → 89.6%
发布
2024-06
维护者
TIGER-Lab (University of Waterloo)
状态
saturating
污染风险
high
指标
accuracy (percent, ↑)
题量
12,032
领域
knowledge reasoning
人类
无实测基线

备注. Developer-reported numbers vary by prompt (CoT vs direct) and shots; the official leaderboard uses 5-shot CoT. A January 2026 formatting fix to answer options can shift older STEM-subset scores slightly.

完整账本

系统开发者分数日期来源条件
Gemini 3.1 Pro (High)Google DeepMind91%独立复现
Run by DeepSeek for its V4-Pro model card comparison table; setting not further specified.
Qwen3.7 MaxAlibaba89.6%聚合站
Developer-reported number republished by aggregator; prompt/effort setting unspecified. Date is the model release date shown by benchlm.ai.
DeepSeek-V4-Pro (Think Max)DeepSeek87.5%厂商自报
Model card comparison table; the same table lists Gemini-3.1-Pro (High) at 91.0 and Claude Opus 4.6 (Max) at 89.1 as run by DeepSeek.
GPT-4oOpenAI72.6%论文shots: 5
Table of the MMLU-Pro paper, CoT prompting.