Cybench
Cybench: A Framework for Evaluating Cybersecurity Capabilities and Risks of Language Models
40 个专业级 CTF 任务,取自四个近期赛事(HackTheBox、SekaiCTF、Glacier、HKCert),涵盖 crypto、web、逆向、取证、misc 和 pwn。智能体在 Kali Linux 容器中执行 shell 命令,操作本地文件和任务服务器并提交 flag 由评测器校验;任务标注人类首解时间(数分钟至 25 小时),并可选子任务给出分级部分得分。主指标为无引导解题率。是 AISI 部署前测试和前沿模型系统卡采用的标准开源网络攻防能力评测。
- 发布
- 2024-08
- 维护者
- Stanford CRFM (Andy K. Zhang, Percy Liang et al.)
- 状态
- saturated
- 污染风险
- high
- 指标
- unguided % solved (percent, ↑)
- 题量
- 40
- 领域
- safety code tool-use
- 人类
- 无实测基线
备注. The official leaderboard (data/leaderboard.csv) mixes paper runs, HAL re-runs and numbers republished from vendor system cards, which often use a 35-39 task subset and pass@1 averaged over several trials; footnotes on the site give the provenance of each row. Two HAL rows (o3-mini, o1-mini) were adjusted downward after an Inspect port leaked an answer. Tasks are public 2023-2024 CTF challenges, so contamination is likely; the authors point to BountyBench as the successor for real-world tasks. Frontier models reach 96-100% on the subset, so it no longer discriminates.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| Claude Mythos Preview | Anthropic | 100% | 厂商自报 | Claude Mythos Preview System Card number republished on the official leaderboard (footnote 6): 35-task subset. Dated to the model's announcement day. | |
| Claude Opus 4.7 | Anthropic | 96% | 厂商自报 | Claude Opus 4.7 System Card number republished on the official leaderboard (footnote 8): 35-task subset. Dated to the model's release day. | |
| Claude Opus 4.5 | Anthropic | 82% | 厂商自报 | Claude Opus 4.5 System Card number republished on the official leaderboard (footnote 3): 39-task subset, average pass@1. Dated to the model's release day. | |
| Claude 3.5 Sonnet | Anthropic | 17.5% | 论文 | scaffold: Cybench agent (structured bash) Paper release; 7/40 tasks unguided; subtask-guided 15%, 43.9% of subtasks solved. | |
| GPT-4o | OpenAI | 12.5% | 论文 | scaffold: Cybench agent (structured bash) Paper release; 5/40 tasks unguided; subtask-guided 17.5%. |