MMMU-Pro

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

↑ MMMU

1,730 道 MMMU 题目,已过滤掉纯文本模型即可作答的题目,每题选项扩展为十个候选,并增加纯视觉设置:问题渲染在图像内部,模型必须同时阅读和观看。以准确率评分;主指标为标准(10 选项)设置与视觉设置的平均值。因消除了曾使 MMMU 虚高的纯文本捷径,它是当前专家级多模态推理的标准。

86.9% · 人类 85.4%
Chance Vision 1.5
厂商自报
2024-09-04 → 2026-07-01: 51.9% → 86.9%
发布
2024-09
维护者
MMMU Team (Yue et al.)
状态
active
污染风险
medium
指标
accuracy (average of standard and vision) (percent, ↑)
题量
1,730
领域
multimodal knowledge reasoning
人类
85.4% estimated high-expert performance derived from MMMU human data (medium 80.8%)

备注. Some vendors report only the standard setting or use tools; the ledger notes the setting where the source states it. The human baseline is an approximation from MMMU annotations, not a fresh expert study on MMMU-Pro.

完整账本

系统开发者分数日期来源条件
Chance Vision 1.5Chance (chance.vision)86.9%厂商自报
Self-reported entry on the official MMMU leaderboard (source: author). Top of the official MMMU-Pro leaderboard, above the 85.4 estimated high-expert level; leaderboard shows only month, so the 1st is used. benchlm.ai instead lists GPT-5.4 Pro at 94% from a vendor chart.
GPT-5.4 Thinking w/ toolsOpenAI82.1%厂商自报tools
Self-reported entry on the official MMMU leaderboard (source: author). Without tools the leaderboard lists 81.2.
Gemini 3.0 ProGoogle DeepMind81%厂商自报tools: no
Self-reported entry on the official MMMU leaderboard (source: author). Google's Gemini 3 launch post reports the same 81%.
o3OpenAI76.4%厂商自报
Self-reported entry on the official MMMU leaderboard (source: author). Official leaderboard overall; OpenAI's GPT-5 post reports the same 76.4 (average of standard and vision).
GPT-4o (0513)OpenAI51.9%论文
Paper Table 1 overall (average of standard 10-option 54.0 and vision 49.7).