LiveCodeBench

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

从 LeetCode、AtCoder 和 Codeforces 竞赛中持续收集的竞赛编程题,每题标注发布日期,以便仅在模型训练截止之后发布的题目上评分。代码生成以针对隐藏测试的 pass@1(多次采样取平均)评分;自我修复、代码执行和测试输出预测为独立场景。滚动的发布日期过滤使其成为标准的防污染编程基准,但所选时间窗口会改变分数。

93.5%
DeepSeek-V4-Pro (Think Max)
厂商自报
2024-03-12 → 2026-05-16: 41.1% → 91.6%
发布
2024-03
维护者
UC Berkeley / MIT / Cornell (Jain et al.)
状态
active
污染风险
low
指标
pass@1 (percent, ↑)
题量
1,055
领域
code reasoning
人类
无实测基线

备注. Vendor numbers use different date windows (v5, v6, or custom) and are not comparable without the window; the public leaderboard's problem pool ends in April 2025, and the maintainers have since focused on LiveCodeBench Pro (Elo-rated). Task count is for the full public pool; each window is a subset.

完整账本

系统开发者分数日期来源条件
DeepSeek-V4-Pro (Think Max)DeepSeek93.5%厂商自报pass_k: 1
Model card comparison table; problem window not stated. Same table: Gemini-3.1-Pro (High) 91.7, Kimi K2.6 89.6, Claude Opus 4.6 88.8.
Qwen3.7 MaxAlibaba91.6%聚合站
Developer-reported number republished by aggregator; prompt/effort setting unspecified. Problem window unspecified; date is model release per benchlm.ai.
O4-Mini (High)OpenAI87.3%官方榜单split: full (2023-05 to 2025-04) pass_k: 1
Top of the official leaderboard over all 1,055 problems; on the 2025-01 to 2025-04 window it scores 75.8. Date is the leaderboard data's last update (HF dataset lastModified 2025-06-05).
DeepSeek-R1-0528DeepSeek84.4%官方榜单split: full (2023-05 to 2025-04) pass_k: 1
Mean pass@1 over all 1,055 problems in the leaderboard data (performances_generation.json); date is model release.
GPT-4-Turbo-2024-04-09OpenAI41.1%论文pass_k: 1
Paper Table 3, code generation total (later arXiv version); GPT-4O-2024-05-13 scored 41.9. Problems May 2023 onward.