MLE-bench
MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
75 Kaggle competitions (22 low, 38 medium, 15 high complexity) in which an agent must read the task, prepare data, train models and submit predictions inside a sandbox within a compute budget. The score is the fraction of competitions in which the submission would have earned at least a bronze medal on the real Kaggle leaderboard. The standard test of end-to-end ML engineering rather than isolated coding.
paper · website · leaderboard · dataset · code
- Released
- 2024-10
- Maintainer
- OpenAI (Chan et al.)
- Status
- active
- Contamination
- medium
- Metric
- any-medal rate (percent, ↑)
- Tasks
- 75
- Domains
- ml-engineering code tool-use
- human
- no measured baseline
Notes. Competitions are public Kaggle data and the paper documents contamination checks. Leaderboard rows are self-submitted with grading reports; as of April 2026 OpenAI paused new submissions while redesigning the process. Runs use 24h (sometimes 12h/36h) budgets, which affects comparability.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Famou-Agent 2.0 + Gemini-3-Pro-Preview | Baidu | 64.44% | official | split: all scaffold: Famou-Agent 2.0 Top of the main leaderboard as of access date; 24h budget. | |
| AIBuildAI + Claude-Opus-4.6 | AIBuildAI | 63.11% | official | split: all scaffold: AIBuildAI 24h budget. | |
| Famou-Agent + Gemini-2.5-Pro | Baidu | 43.56% | official | split: all scaffold: Famou-Agent 24h budget. | |
| AIRA-dojo + o3 | Meta FAIR | 31.6% | official | split: all scaffold: AIRA-dojo 24h budget. | |
| AIDE + o1-preview | OpenAI | 17.12% | official | split: all scaffold: AIDE Paper-era baseline (16.9% in abstract, 17.12 +/- 0.61 on the leaderboard); 24h budget. |