Cybench

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

40 professional-level Capture-the-Flag tasks from four recent competitions (HackTheBox, SekaiCTF, Glacier, HKCert) across crypto, web, reverse engineering, forensics, misc and pwn. An agent works in a Kali Linux container, runs shell commands against local files and task servers, and submits a flag checked by an evaluator; tasks carry human first-solve times (minutes to 25 hours) and optional subtasks for graded partial credit. The headline metric is unguided % solved. It is the standard open cyber-offense eval used in AISI pre-deployment tests and frontier system cards.

100%
Claude Mythos Preview
self-reported
2024-08-15 → 2026-04-16: 12.5% → 96%
Released
2024-08
Maintainer
Stanford CRFM (Andy K. Zhang, Percy Liang et al.)
Status
saturated
Contamination
high
Metric
unguided % solved (percent, ↑)
Tasks
40
Domains
safety code tool-use
human
no measured baseline

Notes. The official leaderboard (data/leaderboard.csv) mixes paper runs, HAL re-runs and numbers republished from vendor system cards, which often use a 35-39 task subset and pass@1 averaged over several trials; footnotes on the site give the provenance of each row. Two HAL rows (o3-mini, o1-mini) were adjusted downward after an Inspect port leaked an answer. Tasks are public 2023-2024 CTF challenges, so contamination is likely; the authors point to BountyBench as the successor for real-world tasks. Frontier models reach 96-100% on the subset, so it no longer discriminates.

Full ledger

SystemDeveloperScoreDateSourceConditions
Claude Mythos PreviewAnthropic100%self-reported
Claude Mythos Preview System Card number republished on the official leaderboard (footnote 6): 35-task subset. Dated to the model's announcement day.
Claude Opus 4.7Anthropic96%self-reported
Claude Opus 4.7 System Card number republished on the official leaderboard (footnote 8): 35-task subset. Dated to the model's release day.
Claude Opus 4.5Anthropic82%self-reported
Claude Opus 4.5 System Card number republished on the official leaderboard (footnote 3): 39-task subset, average pass@1. Dated to the model's release day.
Claude 3.5 SonnetAnthropic17.5%paperscaffold: Cybench agent (structured bash)
Paper release; 7/40 tasks unguided; subtask-guided 15%, 43.9% of subtasks solved.
GPT-4oOpenAI12.5%paperscaffold: Cybench agent (structured bash)
Paper release; 5/40 tasks unguided; subtask-guided 17.5%.