ARC-AGI-3
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
仅基于核心知识先验构建的交互式回合制类游戏环境,不提供任何说明:智能体必须探索、推断目标、构建世界模型并进行规划。包含 25 个公开演示环境、55 个半私有(API 测试)环境和 55 个完全私有(竞赛)环境,每个环境均经未受训人类验证可完全解决。评分基于相对人类行动基线的效率,100% 意味着以人类同等效率通关所有关卡。是 ARC-AGI 系列首个交互式版本。
- 发布
- 2026-03
- 维护者
- ARC Prize Foundation
- 状态
- 污染风险
- low
- 指标
- efficiency-weighted score (percent, ↑)
- 题量
- 135
- 领域
- reasoning tool-use
- 人类
- 100% untrained human test-takers (every environment solved by at least two of ten participants)
备注. ARC Prize reports two harness conditions on the leaderboard: the Standard harness (model carries forward its own notes; provider-neutral) and the Provider Adapter harness (preserves opaque reasoning state and compacts context). Frontier models scored below 1% at the March 2026 launch; GPT-6 Astra reached 99.9% under the Provider Adapter harness in September 2026, so status is 'saturating' even though the Standard-harness score is 62.7%.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra (high) | OpenAI | 99.9% | 官方榜单 | split: semi-private scaffold: Provider Adapter reasoning_effort: high cost_usd_per_task: 342.13 ARC Prize verified run; $18,817 total across 55 semi-private environments (cost per environment shown). Best Provider Adapter result; verified table shows 99.95%. | |
| GPT-6 Astra (max) | OpenAI | 62.7% | 官方榜单 | split: semi-private scaffold: Standard harness reasoning_effort: max cost_usd_per_task: 474.51 ARC Prize verified run; $26,098 total across 55 semi-private environments (cost per environment shown). Best Standard-harness result. | |
| Opus 4.6 (Max) | Anthropic | 0.5% | 论文 | split: semi-private reasoning_effort: max Table 2 of the paper: semi-private leaderboard scores for frontier models at release. | |
| Gemini 3.1 Pro Preview | Google DeepMind | 0.4% | 论文 | split: semi-private Table 2 of the paper: semi-private leaderboard scores at release. | |
| GPT 5.4 (High) | OpenAI | 0.2% | 论文 | split: semi-private reasoning_effort: high Table 2 of the paper: semi-private leaderboard scores at release. |