PaperBench

PaperBench: Evaluating AI's Ability to Replicate AI Research

智能体需从零复现 20 篇 ICML 2024 Spotlight 和 Oral 论文:理解论文、编写代码库并在可访问 GPU 的容器中运行实验。评分使用与作者共同制定的层级式评分表(8,316 个可评分叶节点),由 LLM 评判应用,给出 0 到 100 的复现分数。Code-Dev 变体仅对代码评分。它衡量长程研究工程能力,并有实测的人类(ML 博士)基线可供比较。

43.4%
IterativeAgent o1-high
官方榜单
2025-04-02 → 2025-04-02: 21% → 43.4%
发布
2025-04
维护者
OpenAI (Starace et al.)
状态
active
污染风险
medium
指标
replication score (percent, ↑)
题量
20
领域
research ml-engineering code tool-use
人类
无实测基线

备注. The paper reports that top ML PhDs outperformed models on a 3-paper subset over 48 hours, but the figure is on a subset and not comparable to the full 20-paper score, so no human_baseline is recorded. The repository leaderboard has not been updated since the April 2025 launch; papers and rubrics are public.

完整账本

系统开发者分数日期来源条件
IterativeAgent o1-highOpenAI43.4%官方榜单split: code-dev scaffold: IterativeAgent reasoning_effort: high
PaperBench Code-Dev leaderboard, 3 runs, +/-0.8.
IterativeAgent o1-high (36h limit)OpenAI26%官方榜单split: full scaffold: IterativeAgent reasoning_effort: high
Top of the repository leaderboard; 3 runs, +/-0.3, 36-hour limit.
IterativeAgent o1-high (24h limit)OpenAI24.4%官方榜单split: full scaffold: IterativeAgent reasoning_effort: high
Repository leaderboard, 3 runs, +/-0.7, 24-hour limit.
BasicAgent claude-3.5-sonnetAnthropic21%官方榜单split: full scaffold: BasicAgent
Repository leaderboard, 3 runs, +/-0.8.