ARC-AGI-2

Abstraction and Reasoning Corpus for Artificial General Intelligence, version 2

↑ ARC-AGI-1 ↓ ARC-AGI-3

第二代 ARC 网格谜题,旨在击败暴力搜索,并测试符号解释、组合推理和依赖上下文的规则应用。包含 1,000 个公开训练任务,以及经校准的各 120 个任务的公开、半私有和私有评估集;每个任务都至少被两名人类在两次尝试内解出。得分为两次尝试内完全正确的测试输出百分比,始终与每任务成本一并报告。作为 ARC-AGI-1 的继任者随 ARC Prize 2025 发布。

95% · 人类 100%
GPT-6 Astra (Max)
官方榜单
2025-05-17 → 2026-09-02: 3% → 95%
发布
2025-03
维护者
ARC Prize Foundation
状态
saturating
污染风险
medium
指标
accuracy (pass@2) (percent, ↑)
题量
120
领域
reasoning
人类
100% ARC Prize human panel (every task solved by at least two of 400+ general-public testers); 66% of individual attempts succeeded

备注. Announced 2025-03-24; the arXiv paper followed in May 2025. The ARC Prize grand-prize threshold is 85% on the private set under Kaggle compute limits, which no open solution has met. Frontier API models passed 90% on the Semi-Private set in mid-2026 at several dollars per task, so the benchmark is saturating for unconstrained systems while remaining open for efficient ones.

完整账本

系统开发者分数日期来源条件
GPT-6 Astra (Max)OpenAI95%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 1.12
ARC Prize verified; also republished by benchlm.ai and in OpenAI's GPT-6 Astra launch table (95.0%).
GPT-5.6 Sol (Max)OpenAI92.5%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 1.44
Gemini 3 Deep Think (2/26)Google DeepMind84.6%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 13.62
GPT-5.2 (XHigh)OpenAI52.9%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 1.9
o3 (Medium)OpenAI3%论文split: Semi-Private pass_k: 2
Table 1 of the ARC-AGI-2 paper (scores as of 14 May 2025); the same model scored 53% on ARC-AGI-1.