Aider Polyglot
Aider polyglot coding benchmark
从 C++、Go、Java、JavaScript、Python、Rust 的 Exercism 练习中选出 7 个参考模型至多 3 个能解的最难 225 题。模型在 aider CLI 中以指定编辑格式修改文件,失败后可凭测试输出重试一次;得分为两次尝试内通过单元测试的题目比例。被广泛引用的指令跟随式代码编辑指标,同时暴露编辑格式合规性与运行成本。
- 发布
- 2024-12
- 维护者
- Paul Gauthier (Aider)
- 状态
- 污染风险
- medium
- 指标
- percent correct (pass@2) (percent, ↑)
- 题量
- 225
- 领域
- code instruction-following
- 人类
- 无实测基线
备注. Introduced 2024-12-21 to replace aider's saturated 133-problem Python edit benchmark. Rows below come from the maintainer's leaderboard data file (aider/website/_data/polyglot_leaderboard.yml), which records run dates, pass_rate_2 and total cost. The official leaderboard page was last updated 2025-11-20 and has not added 2026 models; Exercism problems are public, so newer models may have seen them.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| gpt-5 (high) | OpenAI | 88% | 官方榜单 | scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.129 Top of the official leaderboard as of its 2025-11-20 update; diff edit format. | |
| o3-pro (high) | OpenAI | 84.9% | 官方榜单 | scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.65 | |
| gemini-2.5-pro-preview-06-05 (32k think) | 83.1% | 官方榜单 | scaffold: aider pass_k: 2 cost_usd_per_task: 0.222 diff-fenced edit format; 32k thinking tokens. | ||
| DeepSeek-V3.2-Exp (Reasoner) | DeepSeek | 74.2% | 官方榜单 | scaffold: aider pass_k: 2 cost_usd_per_task: 0.006 Cheapest run above 70% on the leaderboard. | |
| o1-2024-12-17 (high) | OpenAI | 61.7% | 官方榜单 | scaffold: aider reasoning_effort: high pass_k: 2 cost_usd_per_task: 0.833 Launch result; diff edit format. |