SWE-Bench Pro

SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?

↑ SWE-bench Verified

1,865 个经人工核验的长程软件工程任务,来自 41 个代码库(公开集:来自 copyleft 许可开源库的 731 个任务;商业集:来自专有初创公司代码库的 276 个任务;保留集:858 个任务)。每个任务给出增强的问题陈述、需求和可选接口;补丁以 fail-to-pass 和 pass-to-pass 测试评分。参考解答平均改动约 107 行、跨 4 个文件,因此它衡量的是 SWE-bench Verified 已无法区分的多文件、企业级工作。

61.5%
Muse Spark 1.1
官方榜单
2025-09-21 → 2026-09-04: 17.8% → 61.5%
发布
2025-09
维护者
Scale AI
状态
active
污染风险
medium
指标
resolve rate (percent, ↑)
题量
1,865
领域
software-engineering code tool-use
人类
无实测基线

备注. Scale's current leaderboard runs models with an uncapped cost and a 250-turn limit; the paper-era runs (capped cost, 50 turns) are kept on a deprecated leaderboard and are not comparable. Rows marked with an asterisk on the leaderboard use the mini-swe-agent harness, the others SWE-Agent. Contamination is mitigated by GPL-style licensing rather than by secrecy, so 'medium'.

完整账本

系统开发者分数日期来源条件
Muse Spark 1.1Meta61.5%官方榜单split: public scaffold: mini-swe-agent
Marked NEW on the leaderboard, +/-3.10 CI. Uncapped cost, 250-turn limit. No per-row date; date is the day observed.
gpt-5.4 (xHigh)OpenAI59.1%官方榜单split: public scaffold: mini-swe-agent reasoning_effort: xhigh
Uncapped cost, 250-turn limit. Leaderboard shows no per-row date; date is the day observed.
claude-opus-4-5-20251101Anthropic45.89%官方榜单split: public scaffold: SWE-Agent
Uncapped cost, 250-turn limit. Leaderboard shows no per-row date; date is the day observed.
OpenAI GPT-5 (medium)OpenAI23.3%论文split: public scaffold: SWE-Agent
Table 5 of the paper; capped at 50 turns and a cost limit.
Claude Opus 4.1Anthropic22.7%论文split: public scaffold: SWE-Agent
Table 5 of the paper; capped at 50 turns and a cost limit.
Claude Opus 4.1 (commercial set)Anthropic17.8%论文split: commercial scaffold: SWE-Agent
Commercial (private startup) subset, Table 2 of the paper.