MMMU-Pro
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
1,730 道 MMMU 题目,已过滤掉纯文本模型即可作答的题目,每题选项扩展为十个候选,并增加纯视觉设置:问题渲染在图像内部,模型必须同时阅读和观看。以准确率评分;主指标为标准(10 选项)设置与视觉设置的平均值。因消除了曾使 MMMU 虚高的纯文本捷径,它是当前专家级多模态推理的标准。
- 发布
- 2024-09
- 维护者
- MMMU Team (Yue et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- accuracy (average of standard and vision) (percent, ↑)
- 题量
- 1,730
- 领域
- multimodal knowledge reasoning
- 人类
- 85.4% estimated high-expert performance derived from MMMU human data (medium 80.8%)
备注. Some vendors report only the standard setting or use tools; the ledger notes the setting where the source states it. The human baseline is an approximation from MMMU annotations, not a fresh expert study on MMMU-Pro.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Chance Vision 1.5 | Chance (chance.vision) | 86.9% | 厂商自报 | Self-reported entry on the official MMMU leaderboard (source: author). Top of the official MMMU-Pro leaderboard, above the 85.4 estimated high-expert level; leaderboard shows only month, so the 1st is used. benchlm.ai instead lists GPT-5.4 Pro at 94% from a vendor chart. | |
| GPT-5.4 Thinking w/ tools | OpenAI | 82.1% | 厂商自报 | tools Self-reported entry on the official MMMU leaderboard (source: author). Without tools the leaderboard lists 81.2. | |
| Gemini 3.0 Pro | Google DeepMind | 81% | 厂商自报 | tools: no Self-reported entry on the official MMMU leaderboard (source: author). Google's Gemini 3 launch post reports the same 81%. | |
| o3 | OpenAI | 76.4% | 厂商自报 | Self-reported entry on the official MMMU leaderboard (source: author). Official leaderboard overall; OpenAI's GPT-5 post reports the same 76.4 (average of standard and vision). | |
| GPT-4o (0513) | OpenAI | 51.9% | 论文 | Paper Table 1 overall (average of standard 10-option 54.0 and vision 49.7). |