ScreenSpot-Pro
ScreenSpot-Pro: GUI Grounding for Professional High-Resolution Computer Use
1,581 条 GUI 定位指令,每条对应一张来自 23 个专业应用(CAD、IDE、创意套件、科学工具、办公)与三种操作系统的真实高分辨率截图。模型需输出指令所指元素的点击点或框;目标平均仅占屏幕 0.07%,准确率为预测落入真值框的比例(文本与图标目标微平均)。是计算机使用智能体所依赖的视觉定位能力的标准压力测试。
- 发布
- 2025-01
- 维护者
- National University of Singapore / HKBU (Kaixin Li et al.)
- 状态
- 污染风险
- high
- 指标
- grounding accuracy (percent, ↑)
- 题量
- 1,581
- 领域
- computer-use multimodal
- 人类
- 无实测基线
备注. Paper and dataset released 2025-01-04 (GitHub changelog); arXiv 2504.07981 verified (posted 2025-04-04; ICLR 2025 workshop). Official leaderboard entries use greedy decoding and are contributed by model authors; frontier labs report their own runs, often with tools (OpenAI: Python zoom tool) or agentic zoom-in, which inflates scores versus single-shot grounding. Images and annotations are fully public, so contamination risk is high.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 92.7% | 厂商自报 | tools: no 'ScreenSpot-Pro (no tools)' row; maximum effort. GPT-5.6 Sol: 76.9% in the same table. | |
| Claude Mythos 5.1 | Anthropic | 87.3% | 厂商自报 | tools: no Run by OpenAI in its GPT-6 Astra comparison table (column 'Claude Fable 5'); footnote says the score comes from Mythos, the reduced-safeguard variant. Not an Anthropic-reported number. | |
| GPT-5.2 Thinking | OpenAI | 86.3% | 厂商自报 | tools reasoning_effort: xhigh Python tool enabled; OpenAI states scores are much lower without it. GPT-5.1 Thinking: 64.2% under the same setup. | |
| Indeed-UI-32B-zoomin | Intelligence Indeed | 82.7% | 官方榜单 | Top entry of the official leaderboard on access date; zoom-in variant. Date is the leaderboard's 'last updated' stamp, not the submission date. | |
| OS-Atlas-7B | Shanghai AI Lab (OS-Copilot) | 18.9% | 论文 | Best grounding model in the paper's Table 2; GPT-4o scored 0.8-0.9%. |