MMMU
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI
11,500 college-level questions that pair text with images (charts, diagrams, maps, tables, chemical structures, music sheets) across six disciplines and 30 subjects, collected from exams, quizzes and textbooks. Scored as accuracy on the 900-question validation set (the test set of 10,500 is graded via EvalAI). The first broad expert-level multimodal benchmark; it is now close to expert-human accuracy and has been superseded by MMMU-Pro.
paper · website · leaderboard · dataset · code
- Released
- 2023-11
- Maintainer
- MMMU Team (Ohio State / Waterloo / CMU; Yue et al.)
- Status
- Contamination
- high
- Metric
- accuracy (percent, ↑)
- Tasks
- 900
- Domains
- multimodal knowledge reasoning
- human
- 88.6% best of three human experts on the validation set (medium expert 82.6%)
Notes. Vendors sometimes report the average of standard and vision settings or use tools; the official leaderboard lists validation accuracy without tools. Many questions can be answered from text alone, which motivated MMMU-Pro.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-5.1 | OpenAI | 85.4% | self-reported | split: validation Self-reported entry on the official MMMU leaderboard (source: author). Top of the official leaderboard (last updated 2025-09-05 header, entry dated 2025-11-13); within 3 points of the best-expert 88.6. | |
| GPT-5 w/ thinking | OpenAI | 84.2% | self-reported | split: validation Self-reported entry on the official MMMU leaderboard (source: author). Matches OpenAI's GPT-5 launch post (84.2). | |
| o3 | OpenAI | 82.9% | self-reported | split: validation Self-reported entry on the official MMMU leaderboard (source: author). Also reported by OpenAI in the GPT-5 launch chart (82.9). | |
| GPT-4o (0513) | OpenAI | 69.1% | self-reported | split: validation Self-reported entry on the official MMMU leaderboard (source: author). | |
| GPT-4V(ision) (Playground) | OpenAI | 56.8% | paper | split: validation shots: 0 MMMU paper Table; test-set score 56.1. Best human expert 88.6. |