LiveCodeBench
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
从 LeetCode、AtCoder 和 Codeforces 竞赛中持续收集的竞赛编程题,每题标注发布日期,以便仅在模型训练截止之后发布的题目上评分。代码生成以针对隐藏测试的 pass@1(多次采样取平均)评分;自我修复、代码执行和测试输出预测为独立场景。滚动的发布日期过滤使其成为标准的防污染编程基准,但所选时间窗口会改变分数。
- 发布
- 2024-03
- 维护者
- UC Berkeley / MIT / Cornell (Jain et al.)
- 状态
- active
- 污染风险
- low
- 指标
- pass@1 (percent, ↑)
- 题量
- 1,055
- 领域
- code reasoning
- 人类
- 无实测基线
备注. Vendor numbers use different date windows (v5, v6, or custom) and are not comparable without the window; the public leaderboard's problem pool ends in April 2025, and the maintainers have since focused on LiveCodeBench Pro (Elo-rated). Task count is for the full public pool; each window is a subset.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| DeepSeek-V4-Pro (Think Max) | DeepSeek | 93.5% | 厂商自报 | pass_k: 1 Model card comparison table; problem window not stated. Same table: Gemini-3.1-Pro (High) 91.7, Kimi K2.6 89.6, Claude Opus 4.6 88.8. | |
| Qwen3.7 Max | Alibaba | 91.6% | 聚合站 | Developer-reported number republished by aggregator; prompt/effort setting unspecified. Problem window unspecified; date is model release per benchlm.ai. | |
| O4-Mini (High) | OpenAI | 87.3% | 官方榜单 | split: full (2023-05 to 2025-04) pass_k: 1 Top of the official leaderboard over all 1,055 problems; on the 2025-01 to 2025-04 window it scores 75.8. Date is the leaderboard data's last update (HF dataset lastModified 2025-06-05). | |
| DeepSeek-R1-0528 | DeepSeek | 84.4% | 官方榜单 | split: full (2023-05 to 2025-04) pass_k: 1 Mean pass@1 over all 1,055 problems in the leaderboard data (performances_generation.json); date is model release. | |
| GPT-4-Turbo-2024-04-09 | OpenAI | 41.1% | 论文 | pass_k: 1 Paper Table 3, code generation total (later arXiv version); GPT-4O-2024-05-13 scored 41.9. Problems May 2023 onward. |