MMMU

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

↓ MMMU-Pro

11,500 道大学水平的题目,将文本与图像(图表、示意图、地图、表格、化学结构、乐谱)配对,覆盖六个学科和 30 个科目,收集自考试、测验和教科书。以 900 题验证集上的准确率评分(10,500 题的测试集通过 EvalAI 评分)。首个广泛的专家级多模态基准;如今已接近专家人类准确率,并已被 MMMU-Pro 取代。

85.4% · 人类 88.6%
GPT-5.1
厂商自报
2023-11-27 → 2025-11-13: 56.8% → 85.4%
发布
2023-11
维护者
MMMU Team (Ohio State / Waterloo / CMU; Yue et al.)
状态
saturating
污染风险
high
指标
accuracy (percent, ↑)
题量
900
领域
multimodal knowledge reasoning
人类
88.6% best of three human experts on the validation set (medium expert 82.6%)

备注. Vendors sometimes report the average of standard and vision settings or use tools; the official leaderboard lists validation accuracy without tools. Many questions can be answered from text alone, which motivated MMMU-Pro.

完整账本

系统开发者分数日期来源条件
GPT-5.1OpenAI85.4%厂商自报split: validation
Self-reported entry on the official MMMU leaderboard (source: author). Top of the official leaderboard (last updated 2025-09-05 header, entry dated 2025-11-13); within 3 points of the best-expert 88.6.
GPT-5 w/ thinkingOpenAI84.2%厂商自报split: validation
Self-reported entry on the official MMMU leaderboard (source: author). Matches OpenAI's GPT-5 launch post (84.2).
o3OpenAI82.9%厂商自报split: validation
Self-reported entry on the official MMMU leaderboard (source: author). Also reported by OpenAI in the GPT-5 launch chart (82.9).
GPT-4o (0513)OpenAI69.1%厂商自报split: validation
Self-reported entry on the official MMMU leaderboard (source: author).
GPT-4V(ision) (Playground)OpenAI56.8%论文split: validation shots: 0
MMMU paper Table; test-set score 56.1. Best human expert 88.6.