MLE-bench

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

75 Kaggle competitions (22 low, 38 medium, 15 high complexity) in which an agent must read the task, prepare data, train models and submit predictions inside a sandbox within a compute budget. The score is the fraction of competitions in which the submission would have earned at least a bronze medal on the real Kaggle leaderboard. The standard test of end-to-end ML engineering rather than isolated coding.

64.44%
Famou-Agent 2.0 + Gemini-3-Pro-Preview
official
2024-10-08 → 2026-03-06: 17.12% → 63.11%
Released
2024-10
Maintainer
OpenAI (Chan et al.)
Status
active
Contamination
medium
Metric
any-medal rate (percent, ↑)
Tasks
75
Domains
ml-engineering code tool-use
human
no measured baseline

Notes. Competitions are public Kaggle data and the paper documents contamination checks. Leaderboard rows are self-submitted with grading reports; as of April 2026 OpenAI paused new submissions while redesigning the process. Runs use 24h (sometimes 12h/36h) budgets, which affects comparability.

Full ledger

SystemDeveloperScoreDateSourceConditions
Famou-Agent 2.0 + Gemini-3-Pro-PreviewBaidu64.44%officialsplit: all scaffold: Famou-Agent 2.0
Top of the main leaderboard as of access date; 24h budget.
AIBuildAI + Claude-Opus-4.6AIBuildAI63.11%officialsplit: all scaffold: AIBuildAI
24h budget.
Famou-Agent + Gemini-2.5-ProBaidu43.56%officialsplit: all scaffold: Famou-Agent
24h budget.
AIRA-dojo + o3Meta FAIR31.6%officialsplit: all scaffold: AIRA-dojo
24h budget.
AIDE + o1-previewOpenAI17.12%officialsplit: all scaffold: AIDE
Paper-era baseline (16.9% in abstract, 17.12 +/- 0.61 on the leaderboard); 24h budget.