SWE-bench Verified
SWE-bench Verified: human-validated subset of SWE-bench
500 个来自 12 个 Python 代码库的真实 GitHub issue,经人工标注者筛选以移除描述不足或无法测试的任务。智能体获得 issue 文本和代码库,需生成补丁,并以隐藏的 FAIL_TO_PASS 和 PASS_TO_PASS 测试评分。是智能体编程的事实标准;分数高度依赖脚手架及步数/成本限制。
- 发布
- 2024-08
- 维护者
- Princeton NLP / OpenAI (curation)
- 状态
- 污染风险
- high
- 指标
- % resolved (percent, ↑)
- 题量
- 500
- 领域
- software-engineering code tool-use
- 人类
- 无实测基线
备注. Test instances are public and predate most frontier training cutoffs; treat developer-reported gains with care and prefer the official bash-only view or independent re-runs (SWE-rebench, Epoch AI) when comparing systems.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Opus 5 | Anthropic | 96% | 聚合站 | Developer-reported number republished by aggregator; scaffold unspecified. | |
| Claude Opus 4.7 | Anthropic | 87.6% | 聚合站 | Developer-reported number republished by aggregator; scaffold unspecified. | |
| Claude 4.5 Opus (high) | Anthropic | 76.8% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent reasoning_effort: high cost_usd_per_task: 0.75 Highest bash-only (same-environment) entry on the official Verified leaderboard as of access date. | |
| mini-SWE-agent + Claude Sonnet 4 | Anthropic | 65% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent |