ScreenSpot-Pro
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
1,581 GUI grounding instructions, each on a unique authentic high-resolution screenshot from 23 professional applications (CAD, IDEs, creative suites, scientific tools, office) and three operating systems. The model must output the click point or box for the element an instruction refers to; targets average 0.07% of the screen, and accuracy is the share of predictions landing inside the ground-truth box (micro-averaged over text and icon targets). The standard stress test for the visual grounding that computer-use agents depend on.
paper · website · leaderboard · dataset · code
- Released
- 2025-01
- Maintainer
- National University of Singapore / HKBU (Kaixin Li et al.)
- Status
- Contamination
- high
- Metric
- grounding accuracy (percent, ↑)
- Tasks
- 1,581
- Domains
- computer-use multimodal
- human
- no measured baseline
Notes. Paper and dataset released 2025-01-04 (GitHub changelog); arXiv 2504.07981 verified (posted 2025-04-04; ICLR 2025 workshop). Official leaderboard entries use greedy decoding and are contributed by model authors; frontier labs report their own runs, often with tools (OpenAI: Python zoom tool) or agentic zoom-in, which inflates scores versus single-shot grounding. Images and annotations are fully public, so contamination risk is high.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 92.7% | self-reported | tools: no 'ScreenSpot-Pro (no tools)' row; maximum effort. GPT-5.6 Sol: 76.9% in the same table. | |
| Claude Mythos 5.1 | Anthropic | 87.3% | self-reported | tools: no Run by OpenAI in its GPT-6 Astra comparison table (column 'Claude Fable 5'); footnote says the score comes from Mythos, the reduced-safeguard variant. Not an Anthropic-reported number. | |
| GPT-5.2 Thinking | OpenAI | 86.3% | self-reported | tools reasoning_effort: xhigh Python tool enabled; OpenAI states scores are much lower without it. GPT-5.1 Thinking: 64.2% under the same setup. | |
| Indeed-UI-32B-zoomin | Intelligence Indeed | 82.7% | official | Top entry of the official leaderboard on access date; zoom-in variant. Date is the leaderboard's 'last updated' stamp, not the submission date. | |
| OS-Atlas-7B | Shanghai AI Lab (OS-Copilot) | 18.9% | paper | Best grounding model in the paper's Table 2; GPT-4o scored 0.8-0.9%. |