MMLU-Pro
MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark
约 12,000 道题,覆盖 14 个学科,基于 MMLU、STEM 网站、TheoremQA 和 SciBench 构建,选项从四个扩展到十个,并移除了琐碎或有噪声的题目。以思维链下的准确率评分(官方设置为 5-shot)。减少了猜测空间和提示敏感性,成为比较知识与推理时 MMLU 的标准替代品。
- 发布
- 2024-06
- 维护者
- TIGER-Lab (University of Waterloo)
- 状态
- 污染风险
- high
- 指标
- accuracy (percent, ↑)
- 题量
- 12,032
- 领域
- knowledge reasoning
- 人类
- 无实测基线
备注. Developer-reported numbers vary by prompt (CoT vs direct) and shots; the official leaderboard uses 5-shot CoT. A January 2026 formatting fix to answer options can shift older STEM-subset scores slightly.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Gemini 3.1 Pro (High) | Google DeepMind | 91% | 独立复现 | Run by DeepSeek for its V4-Pro model card comparison table; setting not further specified. | |
| Qwen3.7 Max | Alibaba | 89.6% | 聚合站 | Developer-reported number republished by aggregator; prompt/effort setting unspecified. Date is the model release date shown by benchlm.ai. | |
| DeepSeek-V4-Pro (Think Max) | DeepSeek | 87.5% | 厂商自报 | Model card comparison table; the same table lists Gemini-3.1-Pro (High) at 91.0 and Claude Opus 4.6 (Max) at 89.1 as run by DeepSeek. | |
| GPT-4o | OpenAI | 72.6% | 论文 | shots: 5 Table of the MMLU-Pro paper, CoT prompting. |