SWE-Bench Pro
SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
1,865 human-verified, long-horizon software engineering tasks from 41 repositories (public set: 731 tasks from copyleft-licensed OSS repos; commercial set: 276 tasks from proprietary startup codebases; held-out set: 858 tasks). Each task gives an augmented problem statement, requirements and optional interface; the patch is scored by fail-to-pass and pass-to-pass tests. Reference solutions average about 107 changed lines across 4 files, so it measures multi-file, enterprise-style work that SWE-bench Verified no longer separates.
paper · website · leaderboard · dataset · code
- Released
- 2025-09
- Maintainer
- Scale AI
- Status
- active
- Contamination
- medium
- Metric
- resolve rate (percent, ↑)
- Tasks
- 1,865
- Domains
- software-engineering code tool-use
- human
- no measured baseline
Notes. Scale's current leaderboard runs models with an uncapped cost and a 250-turn limit; the paper-era runs (capped cost, 50 turns) are kept on a deprecated leaderboard and are not comparable. Rows marked with an asterisk on the leaderboard use the mini-swe-agent harness, the others SWE-Agent. Contamination is mitigated by GPL-style licensing rather than by secrecy, so 'medium'.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Muse Spark 1.1 | Meta | 61.5% | official | split: public scaffold: mini-swe-agent Marked NEW on the leaderboard, +/-3.10 CI. Uncapped cost, 250-turn limit. No per-row date; date is the day observed. | |
| gpt-5.4 (xHigh) | OpenAI | 59.1% | official | split: public scaffold: mini-swe-agent reasoning_effort: xhigh Uncapped cost, 250-turn limit. Leaderboard shows no per-row date; date is the day observed. | |
| claude-opus-4-5-20251101 | Anthropic | 45.89% | official | split: public scaffold: SWE-Agent Uncapped cost, 250-turn limit. Leaderboard shows no per-row date; date is the day observed. | |
| OpenAI GPT-5 (medium) | OpenAI | 23.3% | paper | split: public scaffold: SWE-Agent Table 5 of the paper; capped at 50 turns and a cost limit. | |
| Claude Opus 4.1 | Anthropic | 22.7% | paper | split: public scaffold: SWE-Agent Table 5 of the paper; capped at 50 turns and a cost limit. | |
| Claude Opus 4.1 (commercial set) | Anthropic | 17.8% | paper | split: commercial scaffold: SWE-Agent Commercial (private startup) subset, Table 2 of the paper. |