ScreenSpot-Pro

ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use

1,581 条 GUI 定位指令,每条对应一张来自 23 个专业应用(CAD、IDE、创意套件、科学工具、办公)与三种操作系统的真实高分辨率截图。模型需输出指令所指元素的点击点或框;目标平均仅占屏幕 0.07%,准确率为预测落入真值框的比例(文本与图标目标微平均)。是计算机使用智能体所依赖的视觉定位能力的标准压力测试。

92.7%
GPT-6 Astra
厂商自报
2025-04-04 → 2026-09-03: 18.9% → 92.7%
发布
2025-01
维护者
National University of Singapore / HKBU (Kaixin Li et al.)
状态
saturating
污染风险
high
指标
grounding accuracy (percent, ↑)
题量
1,581
领域
computer-use multimodal
人类
无实测基线

备注. Paper and dataset released 2025-01-04 (GitHub changelog); arXiv 2504.07981 verified (posted 2025-04-04; ICLR 2025 workshop). Official leaderboard entries use greedy decoding and are contributed by model authors; frontier labs report their own runs, often with tools (OpenAI: Python zoom tool) or agentic zoom-in, which inflates scores versus single-shot grounding. Images and annotations are fully public, so contamination risk is high.

完整账本

系统开发者分数日期来源条件
GPT-6 AstraOpenAI92.7%厂商自报tools: no
'ScreenSpot-Pro (no tools)' row; maximum effort. GPT-5.6 Sol: 76.9% in the same table.
Claude Mythos 5.1Anthropic87.3%厂商自报tools: no
Run by OpenAI in its GPT-6 Astra comparison table (column 'Claude Fable 5'); footnote says the score comes from Mythos, the reduced-safeguard variant. Not an Anthropic-reported number.
GPT-5.2 ThinkingOpenAI86.3%厂商自报tools reasoning_effort: xhigh
Python tool enabled; OpenAI states scores are much lower without it. GPT-5.1 Thinking: 64.2% under the same setup.
Indeed-UI-32B-zoominIntelligence Indeed82.7%官方榜单
Top entry of the official leaderboard on access date; zoom-in variant. Date is the leaderboard's 'last updated' stamp, not the submission date.
OS-Atlas-7BShanghai AI Lab (OS-Copilot)18.9%论文
Best grounding model in the paper's Table 2; GPT-4o scored 0.8-0.9%.