ARC-AGI-2
Abstraction and Reasoning Corpus for Artificial General Intelligence, version 2
第二代 ARC 网格谜题,旨在击败暴力搜索,并测试符号解释、组合推理和依赖上下文的规则应用。包含 1,000 个公开训练任务,以及经校准的各 120 个任务的公开、半私有和私有评估集;每个任务都至少被两名人类在两次尝试内解出。得分为两次尝试内完全正确的测试输出百分比,始终与每任务成本一并报告。作为 ARC-AGI-1 的继任者随 ARC Prize 2025 发布。
- 发布
- 2025-03
- 维护者
- ARC Prize Foundation
- 状态
- 污染风险
- medium
- 指标
- accuracy (pass@2) (percent, ↑)
- 题量
- 120
- 领域
- reasoning
- 人类
- 100% ARC Prize human panel (every task solved by at least two of 400+ general-public testers); 66% of individual attempts succeeded
备注. Announced 2025-03-24; the arXiv paper followed in May 2025. The ARC Prize grand-prize threshold is 85% on the private set under Kaggle compute limits, which no open solution has met. Frontier API models passed 90% on the Semi-Private set in mid-2026 at several dollars per task, so the benchmark is saturating for unconstrained systems while remaining open for efficient ones.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra (Max) | OpenAI | 95% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 1.12 ARC Prize verified; also republished by benchlm.ai and in OpenAI's GPT-6 Astra launch table (95.0%). | |
| GPT-5.6 Sol (Max) | OpenAI | 92.5% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 1.44 | |
| Gemini 3 Deep Think (2/26) | Google DeepMind | 84.6% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 13.62 | |
| GPT-5.2 (XHigh) | OpenAI | 52.9% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 1.9 | |
| o3 (Medium) | OpenAI | 3% | 论文 | split: Semi-Private pass_k: 2 Table 1 of the ARC-AGI-2 paper (scores as of 14 May 2025); the same model scored 53% on ARC-AGI-1. |