PaperBench
PaperBench: Evaluating AI's Ability to Replicate AI Research
智能体需从零复现 20 篇 ICML 2024 Spotlight 和 Oral 论文:理解论文、编写代码库并在可访问 GPU 的容器中运行实验。评分使用与作者共同制定的层级式评分表(8,316 个可评分叶节点),由 LLM 评判应用,给出 0 到 100 的复现分数。Code-Dev 变体仅对代码评分。它衡量长程研究工程能力,并有实测的人类(ML 博士)基线可供比较。
- 发布
- 2025-04
- 维护者
- OpenAI (Starace et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- replication score (percent, ↑)
- 题量
- 20
- 领域
- research ml-engineering code tool-use
- 人类
- 无实测基线
备注. The paper reports that top ML PhDs outperformed models on a 3-paper subset over 48 hours, but the figure is on a subset and not comparable to the full 20-paper score, so no human_baseline is recorded. The repository leaderboard has not been updated since the April 2025 launch; papers and rubrics are public.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| IterativeAgent o1-high | OpenAI | 43.4% | 官方榜单 | split: code-dev scaffold: IterativeAgent reasoning_effort: high PaperBench Code-Dev leaderboard, 3 runs, +/-0.8. | |
| IterativeAgent o1-high (36h limit) | OpenAI | 26% | 官方榜单 | split: full scaffold: IterativeAgent reasoning_effort: high Top of the repository leaderboard; 3 runs, +/-0.3, 36-hour limit. | |
| IterativeAgent o1-high (24h limit) | OpenAI | 24.4% | 官方榜单 | split: full scaffold: IterativeAgent reasoning_effort: high Repository leaderboard, 3 runs, +/-0.7, 24-hour limit. | |
| BasicAgent claude-3.5-sonnet | Anthropic | 21% | 官方榜单 | split: full scaffold: BasicAgent Repository leaderboard, 3 runs, +/-0.8. |