ARC-AGI-1

Abstraction and Reasoning Corpus for Artificial General Intelligence, version 1

↓ ARC-AGI-2

基于网格的视觉谜题:给定少量输入-输出示例对,系统需推断变换规则并为新输入生成精确的输出网格。包含 400 个公开训练任务和 400 个公开评估任务,以及各 100 个任务的半私有和私有集;得分为两次尝试内(pass@2)完全正确的测试输出百分比。旨在衡量对新问题的技能获取而非记忆知识;在 2024 年底测试时推理出现之前,它一直抵御着 LLM 规模扩展。

98.5% · 人类 98%
GPT-6 Astra (XHigh)
官方榜单
2024-12-20 → 2026-09-02: 4.5% → 98.5%
发布
2019-11
维护者
ARC Prize Foundation (created by Francois Chollet)
状态
saturated
污染风险
medium
指标
accuracy (pass@2) (percent, ↑)
题量
400
领域
reasoning
人类
98% ARC Prize human panel (STEM graduates); average MTurk worker scored 77%

备注. Official leaderboard scores are on the Semi-Private set and report cost per task alongside accuracy; only runs under $10k total are listed. The December 2024 o3-preview result (75.7% high-efficiency, 87.5% at 172x compute) was the first to approach the human panel; frontier models now score above 97%.

完整账本

系统开发者分数日期来源条件
GPT-6 Astra (XHigh)OpenAI98.5%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 0.35
Ties Claude Fable 5 (Max/XHigh, 98.5%, June 2026) at far lower cost; matches the 98% human panel.
Gemini 3 Deep Think (2/26)Google DeepMind96%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 7.17
GPT-5.2 (XHigh)OpenAI86.2%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 0.96
o3-preview (high efficiency)OpenAI75.7%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 26
ARC Prize verified; the 172x-compute configuration scored 87.5% at about $4,560 per task. Model was trained on 75% of the public training set.
o3 (High)OpenAI60.8%官方榜单split: Semi-Private pass_k: 2 cost_usd_per_task: 0.5
Publicly released o3, ARC Prize leaderboard; lower than the December 2024 preview.
GPT-4oOpenAI4.5%官方榜单split: Semi-Private
ARC Prize testing; the blog states GPT-4o reached 5% in 2024 and the leaderboard lists 4.5%.