SWE-Bench Pro
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
1,865 个经人工核验的长程软件工程任务,来自 41 个代码库(公开集:来自 copyleft 许可开源库的 731 个任务;商业集:来自专有初创公司代码库的 276 个任务;保留集:858 个任务)。每个任务给出增强的问题陈述、需求和可选接口;补丁以 fail-to-pass 和 pass-to-pass 测试评分。参考解答平均改动约 107 行、跨 4 个文件,因此它衡量的是 SWE-bench Verified 已无法区分的多文件、企业级工作。
- 发布
- 2025-09
- 维护者
- Scale AI
- 状态
- active
- 污染风险
- medium
- 指标
- resolve rate (percent, ↑)
- 题量
- 1,865
- 领域
- software-engineering code tool-use
- 人类
- 无实测基线
备注. Scale's current leaderboard runs models with an uncapped cost and a 250-turn limit; the paper-era runs (capped cost, 50 turns) are kept on a deprecated leaderboard and are not comparable. Rows marked with an asterisk on the leaderboard use the mini-swe-agent harness, the others SWE-Agent. Contamination is mitigated by GPL-style licensing rather than by secrecy, so 'medium'.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Muse Spark 1.1 | Meta | 61.5% | 官方榜单 | split: public scaffold: mini-swe-agent Marked NEW on the leaderboard, +/-3.10 CI. Uncapped cost, 250-turn limit. No per-row date; date is the day observed. | |
| gpt-5.4 (xHigh) | OpenAI | 59.1% | 官方榜单 | split: public scaffold: mini-swe-agent reasoning_effort: xhigh Uncapped cost, 250-turn limit. Leaderboard shows no per-row date; date is the day observed. | |
| claude-opus-4-5-20251101 | Anthropic | 45.89% | 官方榜单 | split: public scaffold: SWE-Agent Uncapped cost, 250-turn limit. Leaderboard shows no per-row date; date is the day observed. | |
| OpenAI GPT-5 (medium) | OpenAI | 23.3% | 论文 | split: public scaffold: SWE-Agent Table 5 of the paper; capped at 50 turns and a cost limit. | |
| Claude Opus 4.1 | Anthropic | 22.7% | 论文 | split: public scaffold: SWE-Agent Table 5 of the paper; capped at 50 turns and a cost limit. | |
| Claude Opus 4.1 (commercial set) | Anthropic | 17.8% | 论文 | split: commercial scaffold: SWE-Agent Commercial (private startup) subset, Table 2 of the paper. |