MMMU
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
11,500 道大学水平的题目,将文本与图像(图表、示意图、地图、表格、化学结构、乐谱)配对,覆盖六个学科和 30 个科目,收集自考试、测验和教科书。以 900 题验证集上的准确率评分(10,500 题的测试集通过 EvalAI 评分)。首个广泛的专家级多模态基准;如今已接近专家人类准确率,并已被 MMMU-Pro 取代。
- 发布
- 2023-11
- 维护者
- MMMU Team (Ohio State / Waterloo / CMU; Yue et al.)
- 状态
- 污染风险
- high
- 指标
- accuracy (percent, ↑)
- 题量
- 900
- 领域
- multimodal knowledge reasoning
- 人类
- 88.6% best of three human experts on the validation set (medium expert 82.6%)
备注. Vendors sometimes report the average of standard and vision settings or use tools; the official leaderboard lists validation accuracy without tools. Many questions can be answered from text alone, which motivated MMMU-Pro.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-5.1 | OpenAI | 85.4% | 厂商自报 | split: validation Self-reported entry on the official MMMU leaderboard (source: author). Top of the official leaderboard (last updated 2025-09-05 header, entry dated 2025-11-13); within 3 points of the best-expert 88.6. | |
| GPT-5 w/ thinking | OpenAI | 84.2% | 厂商自报 | split: validation Self-reported entry on the official MMMU leaderboard (source: author). Matches OpenAI's GPT-5 launch post (84.2). | |
| o3 | OpenAI | 82.9% | 厂商自报 | split: validation Self-reported entry on the official MMMU leaderboard (source: author). Also reported by OpenAI in the GPT-5 launch chart (82.9). | |
| GPT-4o (0513) | OpenAI | 69.1% | 厂商自报 | split: validation Self-reported entry on the official MMMU leaderboard (source: author). | |
| GPT-4V(ision) (Playground) | OpenAI | 56.8% | 论文 | split: validation shots: 0 MMMU paper Table; test-set score 56.1. Best human expert 88.6. |