Cybench
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
40 professional-level Capture-the-Flag tasks from four recent competitions (HackTheBox, SekaiCTF, Glacier, HKCert) across crypto, web, reverse engineering, forensics, misc and pwn. An agent works in a Kali Linux container, runs shell commands against local files and task servers, and submits a flag checked by an evaluator; tasks carry human first-solve times (minutes to 25 hours) and optional subtasks for graded partial credit. The headline metric is unguided % solved. It is the standard open cyber-offense eval used in AISI pre-deployment tests and frontier system cards.
paper · website · leaderboard · dataset · code
- Released
- 2024-08
- Maintainer
- Stanford CRFM (Andy K. Zhang, Percy Liang et al.)
- Status
- saturated
- Contamination
- high
- Metric
- unguided % solved (percent, ↑)
- Tasks
- 40
- Domains
- safety code tool-use
- human
- no measured baseline
Notes. The official leaderboard (data/leaderboard.csv) mixes paper runs, HAL re-runs and numbers republished from vendor system cards, which often use a 35-39 task subset and pass@1 averaged over several trials; footnotes on the site give the provenance of each row. Two HAL rows (o3-mini, o1-mini) were adjusted downward after an Inspect port leaked an answer. Tasks are public 2023-2024 CTF challenges, so contamination is likely; the authors point to BountyBench as the successor for real-world tasks. Frontier models reach 96-100% on the subset, so it no longer discriminates.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| Claude Mythos Preview | Anthropic | 100% | self-reported | Claude Mythos Preview System Card number republished on the official leaderboard (footnote 6): 35-task subset. Dated to the model's announcement day. | |
| Claude Opus 4.7 | Anthropic | 96% | self-reported | Claude Opus 4.7 System Card number republished on the official leaderboard (footnote 8): 35-task subset. Dated to the model's release day. | |
| Claude Opus 4.5 | Anthropic | 82% | self-reported | Claude Opus 4.5 System Card number republished on the official leaderboard (footnote 3): 39-task subset, average pass@1. Dated to the model's release day. | |
| Claude 3.5 Sonnet | Anthropic | 17.5% | paper | scaffold: Cybench agent (structured bash) Paper release; 7/40 tasks unguided; subtask-guided 15%, 43.9% of subtasks solved. | |
| GPT-4o | OpenAI | 12.5% | paper | scaffold: Cybench agent (structured bash) Paper release; 5/40 tasks unguided; subtask-guided 17.5%. |