MMMU-Pro

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

↑ MMMU

1,730 MMMU questions that survived filtering out items answerable by text-only models, each with the option set expanded to ten candidates, plus a vision-only setting in which the question is rendered inside the image so the model must read and see at once. Scored as accuracy; the headline number averages the standard (10-option) and vision settings. It is the current standard for expert-level multimodal reasoning because it removes the text-only shortcuts that inflated MMMU.

86.9% · human 85.4%
Chance Vision 1.5
self-reported
2024-09-04 → 2026-07-01: 51.9% → 86.9%
Released
2024-09
Maintainer
MMMU Team (Yue et al.)
Status
active
Contamination
medium
Metric
accuracy (average of standard and vision) (percent, ↑)
Tasks
1,730
Domains
multimodal knowledge reasoning
human
85.4% estimated high-expert performance derived from MMMU human data (medium 80.8%)

Notes. Some vendors report only the standard setting or use tools; the ledger notes the setting where the source states it. The human baseline is an approximation from MMMU annotations, not a fresh expert study on MMMU-Pro.

Full ledger

SystemDeveloperScoreDateSourceConditions
Chance Vision 1.5Chance (chance.vision)86.9%self-reported
Self-reported entry on the official MMMU leaderboard (source: author). Top of the official MMMU-Pro leaderboard, above the 85.4 estimated high-expert level; leaderboard shows only month, so the 1st is used. benchlm.ai instead lists GPT-5.4 Pro at 94% from a vendor chart.
GPT-5.4 Thinking w/ toolsOpenAI82.1%self-reportedtools
Self-reported entry on the official MMMU leaderboard (source: author). Without tools the leaderboard lists 81.2.
Gemini 3.0 ProGoogle DeepMind81%self-reportedtools: no
Self-reported entry on the official MMMU leaderboard (source: author). Google's Gemini 3 launch post reports the same 81%.
o3OpenAI76.4%self-reported
Self-reported entry on the official MMMU leaderboard (source: author). Official leaderboard overall; OpenAI's GPT-5 post reports the same 76.4 (average of standard and vision).
GPT-4o (0513)OpenAI51.9%paper
Paper Table 1 overall (average of standard 10-option 54.0 and vision 49.7).