Aider Polyglot
Aider polyglot coding benchmark
225 of the hardest Exercism exercises in C++, Go, Java, JavaScript, Python and Rust, selected because at most 3 of 7 reference models solved them. The model runs inside the aider CLI, must edit the files using a specified edit format, and gets one retry with the failing test output; the score is the percentage of exercises whose unit tests pass after two attempts. A widely quoted measure of instruction-following code editing that also exposes edit-format compliance and cost per run.
website · leaderboard · dataset · code
- Released
- 2024-12
- Maintainer
- Paul Gauthier (Aider)
- Status
- Contamination
- medium
- Metric
- percent correct (pass@2) (percent, ↑)
- Tasks
- 225
- Domains
- code instruction-following
- human
- no measured baseline
Notes. Introduced 2024-12-21 to replace aider's saturated 133-problem Python edit benchmark. Rows below come from the maintainer's leaderboard data file (aider/website/_data/polyglot_leaderboard.yml), which records run dates, pass_rate_2 and total cost. The official leaderboard page was last updated 2025-11-20 and has not added 2026 models; Exercism problems are public, so newer models may have seen them.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| gpt-5 (high) | OpenAI | 88% | official | scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.129 Top of the official leaderboard as of its 2025-11-20 update; diff edit format. | |
| o3-pro (high) | OpenAI | 84.9% | official | scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.65 | |
| gemini-2.5-pro-preview-06-05 (32k think) | 83.1% | official | scaffold: aider pass_k: 2 cost_usd_per_task: 0.222 diff-fenced edit format; 32k thinking tokens. | ||
| DeepSeek-V3.2-Exp (Reasoner) | DeepSeek | 74.2% | official | scaffold: aider pass_k: 2 cost_usd_per_task: 0.006 Cheapest run above 70% on the leaderboard. | |
| o1-2024-12-17 (high) | OpenAI | 61.7% | official | scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.833 Launch result; diff edit format. |