ARC-AGI-3

ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence

↑ ARC-AGI-2

仅基于核心知识先验构建的交互式回合制类游戏环境,不提供任何说明:智能体必须探索、推断目标、构建世界模型并进行规划。包含 25 个公开演示环境、55 个半私有(API 测试)环境和 55 个完全私有(竞赛)环境,每个环境均经未受训人类验证可完全解决。评分基于相对人类行动基线的效率,100% 意味着以人类同等效率通关所有关卡。是 ARC-AGI 系列首个交互式版本。

99.9% · 人类 100%
GPT-6 Astra (high)
官方榜单
2026-03-24 → 2026-09-03: 0.2% → 99.9%
发布
2026-03
维护者
ARC Prize Foundation
状态
saturating
污染风险
low
指标
efficiency-weighted score (percent, ↑)
题量
135
领域
reasoning tool-use
人类
100% untrained human test-takers (every environment solved by at least two of ten participants)

备注. ARC Prize reports two harness conditions on the leaderboard: the Standard harness (model carries forward its own notes; provider-neutral) and the Provider Adapter harness (preserves opaque reasoning state and compacts context). Frontier models scored below 1% at the March 2026 launch; GPT-6 Astra reached 99.9% under the Provider Adapter harness in September 2026, so status is 'saturating' even though the Standard-harness score is 62.7%.

完整账本

系统开发者分数日期来源条件
GPT-6 Astra (high)OpenAI99.9%官方榜单split: semi-private scaffold: Provider Adapter reasoning_effort: high cost_usd_per_task: 342.13
ARC Prize verified run; $18,817 total across 55 semi-private environments (cost per environment shown). Best Provider Adapter result; verified table shows 99.95%.
GPT-6 Astra (max)OpenAI62.7%官方榜单split: semi-private scaffold: Standard harness reasoning_effort: max cost_usd_per_task: 474.51
ARC Prize verified run; $26,098 total across 55 semi-private environments (cost per environment shown). Best Standard-harness result.
Opus 4.6 (Max)Anthropic0.5%论文split: semi-private reasoning_effort: max
Table 2 of the paper: semi-private leaderboard scores for frontier models at release.
Gemini 3.1 Pro PreviewGoogle DeepMind0.4%论文split: semi-private
Table 2 of the paper: semi-private leaderboard scores at release.
GPT 5.4 (High)OpenAI0.2%论文split: semi-private reasoning_effort: high
Table 2 of the paper: semi-private leaderboard scores at release.