Cybench

Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models

40 个专业级 CTF 任务,取自四个近期赛事(HackTheBox、SekaiCTF、Glacier、HKCert),涵盖 crypto、web、逆向、取证、misc 和 pwn。智能体在 Kali Linux 容器中执行 shell 命令,操作本地文件和任务服务器并提交 flag 由评测器校验;任务标注人类首解时间(数分钟至 25 小时),并可选子任务给出分级部分得分。主指标为无引导解题率。是 AISI 部署前测试和前沿模型系统卡采用的标准开源网络攻防能力评测。

100%
Claude Mythos Preview
厂商自报
2024-08-15 → 2026-04-16: 12.5% → 96%
发布
2024-08
维护者
Stanford CRFM (Andy K. Zhang, Percy Liang et al.)
状态
saturated
污染风险
high
指标
unguided % solved (percent, ↑)
题量
40
领域
safety code tool-use
人类
无实测基线

备注. The official leaderboard (data/leaderboard.csv) mixes paper runs, HAL re-runs and numbers republished from vendor system cards, which often use a 35-39 task subset and pass@1 averaged over several trials; footnotes on the site give the provenance of each row. Two HAL rows (o3-mini, o1-mini) were adjusted downward after an Inspect port leaked an answer. Tasks are public 2023-2024 CTF challenges, so contamination is likely; the authors point to BountyBench as the successor for real-world tasks. Frontier models reach 96-100% on the subset, so it no longer discriminates.

完整账本

系统开发者分数日期来源条件
Claude Mythos PreviewAnthropic100%厂商自报
Claude Mythos Preview System Card number republished on the official leaderboard (footnote 6): 35-task subset. Dated to the model's announcement day.
Claude Opus 4.7Anthropic96%厂商自报
Claude Opus 4.7 System Card number republished on the official leaderboard (footnote 8): 35-task subset. Dated to the model's release day.
Claude Opus 4.5Anthropic82%厂商自报
Claude Opus 4.5 System Card number republished on the official leaderboard (footnote 3): 39-task subset, average pass@1. Dated to the model's release day.
Claude 3.5 SonnetAnthropic17.5%论文scaffold: Cybench agent (structured bash)
Paper release; 7/40 tasks unguided; subtask-guided 15%, 43.9% of subtasks solved.
GPT-4oOpenAI12.5%论文scaffold: Cybench agent (structured bash)
Paper release; 5/40 tasks unguided; subtask-guided 17.5%.