agent-harness-evals
Same model, different harness — which agent harness gets the most out of each model, from live leaderboard data.
How to read this
Matrix: best pass rate per (model × harness) pair; brighter blue = higher within this benchmark; the white-ringed cell is that model's best harness. spread = max−min across harnesses for that model — the score difference the harness alone controls (dimmed when n < 3: a single pairwise difference, not a spread).
Harness ranking: "% of model-best" = averaged over the models a harness was tested with, the share of each model's best-known score it retains. 100% = it was the best choice for every model tested. wins = models for which it was the top harness. Rows with fewer than 3 models are dimmed.
Model ranking: each model's best score across harnesses, and which harness got it there.
Sources: Epoch AI Benchmarking Hub (CC-BY, live, ~daily) and the HAL Holistic Agent Leaderboard (historical). Only models measured under ≥2 harnesses appear.