ARC-AGI-3
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
Interactive, turn-based game-like environments built only from Core Knowledge priors, with no instructions: an agent must explore, infer the goal, build a world model and plan. 25 public demo environments plus 55 semi-private (API-tested) and 55 fully private (competition) environments, each verified fully solvable by untrained humans. Scoring is efficiency-based against human action baselines, so 100% means solving every level as efficiently as humans. The first interactive generation of the ARC-AGI series.
paper · website · leaderboard · dataset · code
- Released
- 2026-03
- Maintainer
- ARC Prize Foundation
- Status
- Contamination
- low
- Metric
- efficiency-weighted score (percent, ↑)
- Tasks
- 135
- Domains
- reasoning tool-use
- human
- 100% untrained human test-takers (every environment solved by at least two of ten participants)
Notes. ARC Prize reports two harness conditions on the leaderboard: the Standard harness (model carries forward its own notes; provider-neutral) and the Provider Adapter harness (preserves opaque reasoning state and compacts context). Frontier models scored below 1% at the March 2026 launch; GPT-6 Astra reached 99.9% under the Provider Adapter harness in September 2026, so status is 'saturating' even though the Standard-harness score is 62.7%.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-6 Astra (high) | OpenAI | 99.9% | official | split: semi-private scaffold: Provider Adapter reasoning_effort: high cost_usd_per_task: 342.13 ARC Prize verified run; $18,817 total across 55 semi-private environments (cost per environment shown). Best Provider Adapter result; verified table shows 99.95%. | |
| GPT-6 Astra (max) | OpenAI | 62.7% | official | split: semi-private scaffold: Standard harness reasoning_effort: max cost_usd_per_task: 474.51 ARC Prize verified run; $26,098 total across 55 semi-private environments (cost per environment shown). Best Standard-harness result. | |
| Opus 4.6 (Max) | Anthropic | 0.5% | paper | split: semi-private reasoning_effort: max Table 2 of the paper: semi-private leaderboard scores for frontier models at release. | |
| Gemini 3.1 Pro Preview | Google DeepMind | 0.4% | paper | split: semi-private Table 2 of the paper: semi-private leaderboard scores at release. | |
| GPT 5.4 (High) | OpenAI | 0.2% | paper | split: semi-private reasoning_effort: high Table 2 of the paper: semi-private leaderboard scores at release. |