ScreenSpot-Pro

ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

1,581 GUI grounding instructions, each on a unique authentic high-resolution screenshot from 23 professional applications (CAD, IDEs, creative suites, scientific tools, office) and three operating systems. The model must output the click point or box for the element an instruction refers to; targets average 0.07% of the screen, and accuracy is the share of predictions landing inside the ground-truth box (micro-averaged over text and icon targets). The standard stress test for the visual grounding that computer-use agents depend on.

92.7%
GPT-6 Astra
self-reported
2025-04-04 → 2026-09-03: 18.9% → 92.7%
Released
2025-01
Maintainer
National University of Singapore / HKBU (Kaixin Li et al.)
Status
saturating
Contamination
high
Metric
grounding accuracy (percent, ↑)
Tasks
1,581
Domains
computer-use multimodal
human
no measured baseline

Notes. Paper and dataset released 2025-01-04 (GitHub changelog); arXiv 2504.07981 verified (posted 2025-04-04; ICLR 2025 workshop). Official leaderboard entries use greedy decoding and are contributed by model authors; frontier labs report their own runs, often with tools (OpenAI: Python zoom tool) or agentic zoom-in, which inflates scores versus single-shot grounding. Images and annotations are fully public, so contamination risk is high.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-6 AstraOpenAI92.7%self-reportedtools: no
'ScreenSpot-Pro (no tools)' row; maximum effort. GPT-5.6 Sol: 76.9% in the same table.
Claude Mythos 5.1Anthropic87.3%self-reportedtools: no
Run by OpenAI in its GPT-6 Astra comparison table (column 'Claude Fable 5'); footnote says the score comes from Mythos, the reduced-safeguard variant. Not an Anthropic-reported number.
GPT-5.2 ThinkingOpenAI86.3%self-reportedtools reasoning_effort: xhigh
Python tool enabled; OpenAI states scores are much lower without it. GPT-5.1 Thinking: 64.2% under the same setup.
Indeed-UI-32B-zoominIntelligence Indeed82.7%official
Top entry of the official leaderboard on access date; zoom-in variant. Date is the leaderboard's 'last updated' stamp, not the submission date.
OS-Atlas-7BShanghai AI Lab (OS-Copilot)18.9%paper
Best grounding model in the paper's Table 2; GPT-4o scored 0.8-0.9%.