SWE-bench Multilingual
SWE-bench Multilingual: SWE-bench-style tasks across 9 programming languages
300 个精选的问题修复任务,来自 C、C++、Go、Java、JavaScript/TypeScript、PHP、Ruby 和 Rust 的 42 个代码库,使用 SWE-bench 收集流水线构建,并以 fail-to-pass 和 pass-to-pass 测试评分。它检验智能体编程能力能否迁移到 Python 之外,因为脚手架和模型都针对 SWE-bench Verified 做了大量调优。
- 发布
- 2025-03
- 维护者
- SWE-bench team (Khandpur, Lieret, Jimenez, Press, Yang)
- 状态
- active
- 污染风险
- high
- 指标
- % resolved (percent, ↑)
- 题量
- 300
- 领域
- software-engineering code tool-use
- 人类
- 无实测基线
备注. No standalone paper; introduced as a blog post (kabirk.com/multilingual, mirrored on swebench.com). Released month taken from the March 2025 blog post. Task instances are public GitHub issues predating training cutoffs.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Gemini 3 Flash | Google DeepMind | 72.7% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.35 Top of the Multilingual bash-only leaderboard as of access date. | |
| Claude 4.6 Opus | Anthropic | 72% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.66 | |
| MiniMax M2.5 | MiniMax | 68.3% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.1 | |
| Claude 4.5 Sonnet | Anthropic | 67% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.67 | |
| SWE-agent + Claude 3.7 Sonnet | Anthropic | 43% | 官方榜单 | scaffold: SWE-agent Launch baseline from the introducing blog post ($2.50 cost limit, 128/300 resolved = 42.67%, reported as 43%). Exact post day not stated; month-start used. | |
| GPT 5 mini | OpenAI | 39.7% | 官方榜单 | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.05 |