SWE-bench Verified

SWE-bench Verified: human-validated subset of SWE-bench

↑ SWE-bench ↓ SWE-Bench Pro

500 real GitHub issues from 12 Python repositories, filtered by human annotators to remove under-specified or untestable tasks. An agent receives the issue text and the repository, must produce a patch, and is scored by hidden FAIL_TO_PASS and PASS_TO_PASS tests. The de-facto standard for agentic coding; scores depend heavily on scaffold and step/cost limits.

96%
Claude Opus 5
aggregator
2025-07-01 → 2026-09-03: 65% → 96%
Released
2024-08
Maintainer
Princeton NLP / OpenAI (curation)
Status
saturating
Contamination
high
Metric
% resolved (percent, ↑)
Tasks
500
Domains
software-engineering code tool-use
human
no measured baseline

Notes. Test instances are public and predate most frontier training cutoffs; treat developer-reported gains with care and prefer the official bash-only view or independent re-runs (SWE-rebench, Epoch AI) when comparing systems.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude Opus 5Anthropic96%aggregator
Developer-reported number republished by aggregator; scaffold unspecified.
Claude Opus 4.7Anthropic87.6%aggregator
Developer-reported number republished by aggregator; scaffold unspecified.
Claude 4.5 Opus (high)Anthropic76.8%officialsplit: bash-only scaffold: mini-SWE-agent reasoning_effort: high cost_usd_per_task: 0.75
Highest bash-only (same-environment) entry on the official Verified leaderboard as of access date.
mini-SWE-agent + Claude Sonnet 4Anthropic65%officialsplit: bash-only scaffold: mini-SWE-agent