ARC-AGI-1
Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1
基于网格的视觉谜题:给定少量输入-输出示例对,系统需推断变换规则并为新输入生成精确的输出网格。包含 400 个公开训练任务和 400 个公开评估任务,以及各 100 个任务的半私有和私有集;得分为两次尝试内(pass@2)完全正确的测试输出百分比。旨在衡量对新问题的技能获取而非记忆知识;在 2024 年底测试时推理出现之前,它一直抵御着 LLM 规模扩展。
- 发布
- 2019-11
- 维护者
- ARC Prize Foundation (created by Francois Chollet)
- 状态
- saturated
- 污染风险
- medium
- 指标
- accuracy (pass@2) (percent, ↑)
- 题量
- 400
- 领域
- reasoning
- 人类
- 98% ARC Prize human panel (STEM graduates); average MTurk worker scored 77%
备注. Official leaderboard scores are on the Semi-Private set and report cost per task alongside accuracy; only runs under $10k total are listed. The December 2024 o3-preview result (75.7% high-efficiency, 87.5% at 172x compute) was the first to approach the human panel; frontier models now score above 97%.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra (XHigh) | OpenAI | 98.5% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 0.35 Ties Claude Fable 5 (Max/XHigh, 98.5%, June 2026) at far lower cost; matches the 98% human panel. | |
| Gemini 3 Deep Think (2/26) | Google DeepMind | 96% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 7.17 | |
| GPT-5.2 (XHigh) | OpenAI | 86.2% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 0.96 | |
| o3-preview (high efficiency) | OpenAI | 75.7% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 26 ARC Prize verified; the 172x-compute configuration scored 87.5% at about $4,560 per task. Model was trained on 75% of the public training set. | |
| o3 (High) | OpenAI | 60.8% | 官方榜单 | split: Semi-Private pass_k: 2 cost_usd_per_task: 0.5 Publicly released o3, ARC Prize leaderboard; lower than the December 2024 preview. | |
| GPT-4o | OpenAI | 4.5% | 官方榜单 | split: Semi-Private ARC Prize testing; the blog states GPT-4o reached 5% in 2024 and the leaderboard lists 4.5%. |