Aider Polyglot

Aider polyglot coding benchmark

从 C++、Go、Java、JavaScript、Python、Rust 的 Exercism 练习中选出 7 个参考模型至多 3 个能解的最难 225 题。模型在 aider CLI 中以指定编辑格式修改文件,失败后可凭测试输出重试一次;得分为两次尝试内通过单元测试的题目比例。被广泛引用的指令跟随式代码编辑指标,同时暴露编辑格式合规性与运行成本。

88%
gpt-5 (high)
官方榜单
2024-12-21 → 2025-10-03: 61.7% → 74.2%
发布
2024-12
维护者
Paul Gauthier (Aider)
状态
saturating
污染风险
medium
指标
percent correct (pass@2) (percent, ↑)
题量
225
领域
code instruction-following
人类
无实测基线

备注. Introduced 2024-12-21 to replace aider's saturated 133-problem Python edit benchmark. Rows below come from the maintainer's leaderboard data file (aider/website/_data/polyglot_leaderboard.yml), which records run dates, pass_rate_2 and total cost. The official leaderboard page was last updated 2025-11-20 and has not added 2026 models; Exercism problems are public, so newer models may have seen them.

完整账本

系统开发者分数日期来源条件
gpt-5 (high)OpenAI88%官方榜单scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.129
Top of the official leaderboard as of its 2025-11-20 update; diff edit format.
o3-pro (high)OpenAI84.9%官方榜单scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.65
gemini-2.5-pro-preview-06-05 (32k think)Google83.1%官方榜单scaffold: aider pass_k: 2 cost_usd_per_task: 0.222
diff-fenced edit format; 32k thinking tokens.
DeepSeek-V3.2-Exp (Reasoner)DeepSeek74.2%官方榜单scaffold: aider pass_k: 2 cost_usd_per_task: 0.006
Cheapest run above 70% on the leaderboard.
o1-2024-12-17 (high)OpenAI61.7%官方榜单scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.833
Launch result; diff edit format.