ARC-AGI-2

Abstraction and Reasoning Corpus for Artificial General Intelligence, version 2

↑ ARC-AGI-1 ↓ ARC-AGI-3

Second-generation ARC grid puzzles designed to defeat brute-force search and test symbolic interpretation, compositional reasoning and context-dependent rule application. 1,000 public training tasks plus calibrated 120-task public, semi-private and private evaluation sets; every task was solved by at least two humans within two attempts. Scored as the percentage of test outputs exactly correct within two attempts, always reported together with cost per task. It launched with ARC Prize 2025 as the successor to ARC-AGI-1.

95% · human 100%
GPT-6 Astra (Max)
official
2025-05-17 → 2026-09-02: 3% → 95%
Released
2025-03
Maintainer
ARC Prize Foundation
Status
saturating
Contamination
medium
Metric
accuracy (pass@2) (percent, ↑)
Tasks
120
Domains
reasoning
human
100% ARC Prize human panel (every task solved by at least two of 400+ general-public testers); 66% of individual attempts succeeded

Notes. Announced 2025-03-24; the arXiv paper followed in May 2025. The ARC Prize grand-prize threshold is 85% on the private set under Kaggle compute limits, which no open solution has met. Frontier API models passed 90% on the Semi-Private set in mid-2026 at several dollars per task, so the benchmark is saturating for unconstrained systems while remaining open for efficient ones.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-6 Astra (Max)OpenAI95%officialsplit: Semi-Private pass_k: 2 cost_usd_per_task: 1.12
ARC Prize verified; also republished by benchlm.ai and in OpenAI's GPT-6 Astra launch table (95.0%).
GPT-5.6 Sol (Max)OpenAI92.5%officialsplit: Semi-Private pass_k: 2 cost_usd_per_task: 1.44
Gemini 3 Deep Think (2/26)Google DeepMind84.6%officialsplit: Semi-Private pass_k: 2 cost_usd_per_task: 13.62
GPT-5.2 (XHigh)OpenAI52.9%officialsplit: Semi-Private pass_k: 2 cost_usd_per_task: 1.9
o3 (Medium)OpenAI3%papersplit: Semi-Private pass_k: 2
Table 1 of the ARC-AGI-2 paper (scores as of 14 May 2025); the same model scored 53% on ARC-AGI-1.