SWE-bench Verified

SWE-bench Verified: human-validated subset of SWE-bench

↑ SWE-bench ↓ SWE-Bench Pro

500 个来自 12 个 Python 代码库的真实 GitHub issue,经人工标注者筛选以移除描述不足或无法测试的任务。智能体获得 issue 文本和代码库,需生成补丁,并以隐藏的 FAIL_TO_PASS 和 PASS_TO_PASS 测试评分。是智能体编程的事实标准;分数高度依赖脚手架及步数/成本限制。

96%
Claude Opus 5
聚合站
2025-07-01 → 2026-09-03: 65% → 96%
发布
2024-08
维护者
Princeton NLP / OpenAI (curation)
状态
saturating
污染风险
high
指标
% resolved (percent, ↑)
题量
500
领域
software-engineering code tool-use
人类
无实测基线

备注. Test instances are public and predate most frontier training cutoffs; treat developer-reported gains with care and prefer the official bash-only view or independent re-runs (SWE-rebench, Epoch AI) when comparing systems.

完整账本

系统开发者分数日期来源条件
Claude Opus 5Anthropic96%聚合站
Developer-reported number republished by aggregator; scaffold unspecified.
Claude Opus 4.7Anthropic87.6%聚合站
Developer-reported number republished by aggregator; scaffold unspecified.
Claude 4.5 Opus (high)Anthropic76.8%官方榜单split: bash-only scaffold: mini-SWE-agent reasoning_effort: high cost_usd_per_task: 0.75
Highest bash-only (same-environment) entry on the official Verified leaderboard as of access date.
mini-SWE-agent + Claude Sonnet 4Anthropic65%官方榜单split: bash-only scaffold: mini-SWE-agent