ARC-AGI-1
Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1
Grid-based visual puzzles: given a few input-output example pairs, the system must infer the transformation rule and produce the exact output grid for a new input. 400 public training and 400 public evaluation tasks plus 100-task semi-private and private sets; scored as the percentage of test outputs exactly correct within two attempts (pass@2). Built to measure skill acquisition on novel problems rather than memorized knowledge; it resisted LLM scaling until test-time reasoning arrived in late 2024.
paper · website · leaderboard · dataset · code
- Released
- 2019-11
- Maintainer
- ARC Prize Foundation (created by Francois Chollet)
- Status
- saturated
- Contamination
- medium
- Metric
- accuracy (pass@2) (percent, ↑)
- Tasks
- 400
- Domains
- reasoning
- human
- 98% ARC Prize human panel (STEM graduates); average MTurk worker scored 77%
Notes. Official leaderboard scores are on the Semi-Private set and report cost per task alongside accuracy; only runs under $10k total are listed. The December 2024 o3-preview result (75.7% high-efficiency, 87.5% at 172x compute) was the first to approach the human panel; frontier models now score above 97%.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-6 Astra (XHigh) | OpenAI | 98.5% | official | split: Semi-Private pass_k: 2 cost_usd_per_task: 0.35 Ties Claude Fable 5 (Max/XHigh, 98.5%, June 2026) at far lower cost; matches the 98% human panel. | |
| Gemini 3 Deep Think (2/26) | Google DeepMind | 96% | official | split: Semi-Private pass_k: 2 cost_usd_per_task: 7.17 | |
| GPT-5.2 (XHigh) | OpenAI | 86.2% | official | split: Semi-Private pass_k: 2 cost_usd_per_task: 0.96 | |
| o3-preview (high efficiency) | OpenAI | 75.7% | official | split: Semi-Private pass_k: 2 cost_usd_per_task: 26 ARC Prize verified; the 172x-compute configuration scored 87.5% at about $4,560 per task. Model was trained on 75% of the public training set. | |
| o3 (High) | OpenAI | 60.8% | official | split: Semi-Private pass_k: 2 cost_usd_per_task: 0.5 Publicly released o3, ARC Prize leaderboard; lower than the December 2024 preview. | |
| GPT-4o | OpenAI | 4.5% | official | split: Semi-Private ARC Prize testing; the blog states GPT-4o reached 5% in 2024 and the leaderboard lists 4.5%. |