SWE-bench Verified
SWE-bench Verified: human-validated subset of SWE-bench
500 real GitHub issues from 12 Python repositories, filtered by human annotators to remove under-specified or untestable tasks. An agent receives the issue text and the repository, must produce a patch, and is scored by hidden FAIL_TO_PASS and PASS_TO_PASS tests. The de-facto standard for agentic coding; scores depend heavily on scaffold and step/cost limits.
paper · website · leaderboard · dataset · code
- Released
- 2024-08
- Maintainer
- Princeton NLP / OpenAI (curation)
- Status
- Contamination
- high
- Metric
- % resolved (percent, ↑)
- Tasks
- 500
- Domains
- software-engineering code tool-use
- human
- no measured baseline
Notes. Test instances are public and predate most frontier training cutoffs; treat developer-reported gains with care and prefer the official bash-only view or independent re-runs (SWE-rebench, Epoch AI) when comparing systems.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 96% | aggregator | Developer-reported number republished by aggregator; scaffold unspecified. | |
| Claude Opus 4.7 | Anthropic | 87.6% | aggregator | Developer-reported number republished by aggregator; scaffold unspecified. | |
| Claude 4.5 Opus (high) | Anthropic | 76.8% | official | split: bash-only scaffold: mini-SWE-agent reasoning_effort: high cost_usd_per_task: 0.75 Highest bash-only (same-environment) entry on the official Verified leaderboard as of access date. | |
| mini-SWE-agent + Claude Sonnet 4 | Anthropic | 65% | official | split: bash-only scaffold: mini-SWE-agent |