LiveCodeBench

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Competitive-programming problems continuously collected from LeetCode, AtCoder and Codeforces contests, each tagged with its release date so that models can be scored only on problems published after their training cutoff. Code generation is scored as pass@1 against hidden tests (averaged over samples); self-repair, code execution and test-output prediction are separate scenarios. The rolling release-date filter makes it the standard contamination-aware coding benchmark, though the window chosen changes the score.

93.5%
DeepSeek-V4-Pro (Think Max)
self-reported
2024-03-12 → 2026-05-16: 41.1% → 91.6%
Released
2024-03
Maintainer
UC Berkeley / MIT / Cornell (Jain et al.)
Status
active
Contamination
low
Metric
pass@1 (percent, ↑)
Tasks
1,055
Domains
code reasoning
human
no measured baseline

Notes. Vendor numbers use different date windows (v5, v6, or custom) and are not comparable without the window; the public leaderboard's problem pool ends in April 2025, and the maintainers have since focused on LiveCodeBench Pro (Elo-rated). Task count is for the full public pool; each window is a subset.

Full ledger

SystemDeveloperScoreDateSourceConditions
DeepSeek-V4-Pro (Think Max)DeepSeek93.5%self-reportedpass_k: 1
Model card comparison table; problem window not stated. Same table: Gemini-3.1-Pro (High) 91.7, Kimi K2.6 89.6, Claude Opus 4.6 88.8.
Qwen3.7 MaxAlibaba91.6%aggregator
Developer-reported number republished by aggregator; prompt/effort setting unspecified. Problem window unspecified; date is model release per benchlm.ai.
O4-Mini (High)OpenAI87.3%officialsplit: full (2023-05 to 2025-04) pass_k: 1
Top of the official leaderboard over all 1,055 problems; on the 2025-01 to 2025-04 window it scores 75.8. Date is the leaderboard data's last update (HF dataset lastModified 2025-06-05).
DeepSeek-R1-0528DeepSeek84.4%officialsplit: full (2023-05 to 2025-04) pass_k: 1
Mean pass@1 over all 1,055 problems in the leaderboard data (performances_generation.json); date is model release.
GPT-4-Turbo-2024-04-09OpenAI41.1%paperpass_k: 1
Paper Table 3, code generation total (later arXiv version); GPT-4O-2024-05-13 scored 41.9. Problems May 2023 onward.