SWE-bench Multilingual
SWE-bench Multilingual: SWE-bench-style tasks across 9 programming languages
300 curated issue-resolution tasks from 42 repositories in C, C++, Go, Java, JavaScript/TypeScript, PHP, Ruby and Rust, built with the SWE-bench collection pipeline and scored by fail-to-pass and pass-to-pass tests. It checks whether agentic coding ability transfers beyond Python, where scaffolds and models are heavily tuned to SWE-bench Verified.
website · leaderboard · dataset · code
- Released
- 2025-03
- Maintainer
- SWE-bench team (Khandpur, Lieret, Jimenez, Press, Yang)
- Status
- active
- Contamination
- high
- Metric
- % resolved (percent, ↑)
- Tasks
- 300
- Domains
- software-engineering code tool-use
- human
- no measured baseline
Notes. No standalone paper; introduced as a blog post (kabirk.com/multilingual, mirrored on swebench.com). Released month taken from the March 2025 blog post. Task instances are public GitHub issues predating training cutoffs.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Gemini 3 Flash | Google DeepMind | 72.7% | official | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.35 Top of the Multilingual bash-only leaderboard as of access date. | |
| Claude 4.6 Opus | Anthropic | 72% | official | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.66 | |
| MiniMax M2.5 | MiniMax | 68.3% | official | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.1 | |
| Claude 4.5 Sonnet | Anthropic | 67% | official | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.67 | |
| SWE-agent + Claude 3.7 Sonnet | Anthropic | 43% | official | scaffold: SWE-agent Launch baseline from the introducing blog post ($2.50 cost limit, 128/300 resolved = 42.67%, reported as 43%). Exact post day not stated; month-start used. | |
| GPT 5 mini | OpenAI | 39.7% | official | split: bash-only scaffold: mini-SWE-agent cost_usd_per_task: 0.05 |