SWE-bench Multilingual

SWE-bench Multilingual: SWE-bench-style tasks across 9 programming languages

300 curated issue-resolution tasks from 42 repositories in C, C++, Go, Java, JavaScript/TypeScript, PHP, Ruby and Rust, built with the SWE-bench collection pipeline and scored by fail-to-pass and pass-to-pass tests. It checks whether agentic coding ability transfers beyond Python, where scaffolds and models are heavily tuned to SWE-bench Verified.

72.7%
Gemini 3 Flash
official
2025-03-01 → 2026-02-16: 43% → 68.3%
Released
2025-03
Maintainer
SWE-bench team (Khandpur, Lieret, Jimenez, Press, Yang)
Status
active
Contamination
high
Metric
% resolved (percent, ↑)
Tasks
300
Domains
software-engineering code tool-use
human
no measured baseline

Notes. No standalone paper; introduced as a blog post (kabirk.com/multilingual, mirrored on swebench.com). Released month taken from the March 2025 blog post. Task instances are public GitHub issues predating training cutoffs.

Full ledger

SystemDeveloperScoreDateSourceConditions
Gemini 3 FlashGoogle DeepMind72.7%officialsplit: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.35
Top of the Multilingual bash-only leaderboard as of access date.
Claude 4.6 OpusAnthropic72%officialsplit: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.66
MiniMax M2.5MiniMax68.3%officialsplit: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.1
Claude 4.5 SonnetAnthropic67%officialsplit: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.67
SWE-agent + Claude 3.7 SonnetAnthropic43%officialscaffold: SWE-agent
Launch baseline from the introducing blog post ($2.50 cost limit, 128/300 resolved = 42.67%, reported as 43%). Exact post day not stated; month-start used.
GPT 5 miniOpenAI39.7%officialsplit: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.05