SWE-bench Multilingual

SWE-bench Multilingual: SWE-bench-style tasks across 9 programming languages

300 个精选的问题修复任务,来自 C、C++、Go、Java、JavaScript/TypeScript、PHP、Ruby 和 Rust 的 42 个代码库,使用 SWE-bench 收集流水线构建,并以 fail-to-pass 和 pass-to-pass 测试评分。它检验智能体编程能力能否迁移到 Python 之外,因为脚手架和模型都针对 SWE-bench Verified 做了大量调优。

72.7%
Gemini 3 Flash
官方榜单
2025-03-01 → 2026-02-16: 43% → 68.3%
发布
2025-03
维护者
SWE-bench team (Khandpur, Lieret, Jimenez, Press, Yang)
状态
active
污染风险
high
指标
% resolved (percent, ↑)
题量
300
领域
software-engineering code tool-use
人类
无实测基线

备注. No standalone paper; introduced as a blog post (kabirk.com/multilingual, mirrored on swebench.com). Released month taken from the March 2025 blog post. Task instances are public GitHub issues predating training cutoffs.

完整账本

系统开发者分数日期来源条件
Gemini 3 FlashGoogle DeepMind72.7%官方榜单split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.35
Top of the Multilingual bash-only leaderboard as of access date.
Claude 4.6 OpusAnthropic72%官方榜单split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.66
MiniMax M2.5MiniMax68.3%官方榜单split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.1
Claude 4.5 SonnetAnthropic67%官方榜单split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.67
SWE-agent + Claude 3.7 SonnetAnthropic43%官方榜单scaffold: SWE-agent
Launch baseline from the introducing blog post ($2.50 cost limit, 128/300 resolved = 42.67%, reported as 43%). Exact post day not stated; month-start used.
GPT 5 miniOpenAI39.7%官方榜单split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.05