MMMU-Pro
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
1,730 MMMU questions that survived filtering out items answerable by text-only models, each with the option set expanded to ten candidates, plus a vision-only setting in which the question is rendered inside the image so the model must read and see at once. Scored as accuracy; the headline number averages the standard (10-option) and vision settings. It is the current standard for expert-level multimodal reasoning because it removes the text-only shortcuts that inflated MMMU.
paper · website · leaderboard · dataset · code
- Released
- 2024-09
- Maintainer
- MMMU Team (Yue et al.)
- Status
- active
- Contamination
- medium
- Metric
- accuracy (average of standard and vision) (percent, ↑)
- Tasks
- 1,730
- Domains
- multimodal knowledge reasoning
- human
- 85.4% estimated high-expert performance derived from MMMU human data (medium 80.8%)
Notes. Some vendors report only the standard setting or use tools; the ledger notes the setting where the source states it. The human baseline is an approximation from MMMU annotations, not a fresh expert study on MMMU-Pro.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Chance Vision 1.5 | Chance (chance.vision) | 86.9% | self-reported | Self-reported entry on the official MMMU leaderboard (source: author). Top of the official MMMU-Pro leaderboard, above the 85.4 estimated high-expert level; leaderboard shows only month, so the 1st is used. benchlm.ai instead lists GPT-5.4 Pro at 94% from a vendor chart. | |
| GPT-5.4 Thinking w/ tools | OpenAI | 82.1% | self-reported | tools Self-reported entry on the official MMMU leaderboard (source: author). Without tools the leaderboard lists 81.2. | |
| Gemini 3.0 Pro | Google DeepMind | 81% | self-reported | tools: no Self-reported entry on the official MMMU leaderboard (source: author). Google's Gemini 3 launch post reports the same 81%. | |
| o3 | OpenAI | 76.4% | self-reported | Self-reported entry on the official MMMU leaderboard (source: author). Official leaderboard overall; OpenAI's GPT-5 post reports the same 76.4 (average of standard and vision). | |
| GPT-4o (0513) | OpenAI | 51.9% | paper | Paper Table 1 overall (average of standard 10-option 54.0 and vision 49.7). |