Aider Polyglot

Aider polyglot coding benchmark

225 of the hardest Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, selected because at most 3 of 7 reference models solved them. The model runs inside the aider CLI, must edit the files using a specified edit format, and gets one retry with the failing test output; the score is the percentage of exercises whose unit tests pass after two attempts. A widely quoted measure of instruction-following code editing that also exposes edit-format compliance and cost per run.

88%
gpt-5 (high)
official
2024-12-21 → 2025-10-03: 61.7% → 74.2%
Released
2024-12
Maintainer
Paul Gauthier (Aider)
Status
saturating
Contamination
medium
Metric
percent correct (pass@2) (percent, ↑)
Tasks
225
Domains
code instruction-following
human
no measured baseline

Notes. Introduced 2024-12-21 to replace aider's saturated 133-problem Python edit benchmark. Rows below come from the maintainer's leaderboard data file (aider/website/_data/polyglot_leaderboard.yml), which records run dates, pass_rate_2 and total cost. The official leaderboard page was last updated 2025-11-20 and has not added 2026 models; Exercism problems are public, so newer models may have seen them.

Full ledger

SystemDeveloperScoreDateSourceConditions
gpt-5 (high)OpenAI88%officialscaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.129
Top of the official leaderboard as of its 2025-11-20 update; diff edit format.
o3-pro (high)OpenAI84.9%officialscaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.65
gemini-2.5-pro-preview-06-05 (32k think)Google83.1%officialscaffold: aider pass_k: 2 cost_usd_per_task: 0.222
diff-fenced edit format; 32k thinking tokens.
DeepSeek-V3.2-Exp (Reasoner)DeepSeek74.2%officialscaffold: aider pass_k: 2 cost_usd_per_task: 0.006
Cheapest run above 70% on the leaderboard.
o1-2024-12-17 (high)OpenAI61.7%officialscaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.833
Launch result; diff edit format.