Terminal-Bench
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
由人工编写的高难任务,智能体需在真实终端环境中完成(软件工程、科学计算、ML、安全、数据处理)。每个任务自带容器、人类解答和测试;得分为测试通过的任务比例,对多次试验取平均并附 95% 置信区间。是 CLI 编程智能体的事实基准,通过 Harbor 框架运行。
- 发布
- 2026-01
- 维护者
- Stanford / Laude Institute / Harbor (Merrill, Shaw et al.)
- 状态
- active
- 污染风险
- medium
- 指标
- resolution rate (percent, ↑)
- 题量
- 66
- 领域
- software-engineering code tool-use ml-engineering
- 人类
- 无实测基线
备注. Versioned benchmark: tasks are refreshed between releases (2.0 had 89 tasks, 4.0 has 66), so scores are only comparable within a version. Public tasks carry a canary string and the site asks that benchmark data never appear in training corpora. Terminal-Bench-Science (harbor-framework/terminal-bench-science, DOI 10.5281/zenodo.22110253) is a sibling benchmark in the same family and Harbor framework, tracked here as the 'science' split.
完整账本
| 系统 | 开发者 | 分数 | 日期 | 来源 | 条件 |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 64.6% | 厂商自报 | split: science OpenAI launch table (Academic section), maximum at any effort; run in OpenAI's research environment, scaffold not stated. Same table quotes Fable 5.1 at 52.6% and Opus 5 at 30.0%. | |
| GPT-5.2 + Codex CLI | OpenAI | 62.9% | 论文 | split: 2.0 scaffold: Codex CLI Best result in Table 2 of the paper (Terminal-Bench 2.0). | |
| GPT-6 Astra + Codex (max) | OpenAI | 58.18% | 官方榜单 | split: 4.0 scaffold: Codex reasoning_effort: max cost_usd_per_task: 9.9 Rank 1 on Terminal-Bench 4.0 as of access date; +/-2.8 CI over 330 trials. | |
| Fable 5.1 + Claude Code (max) | Anthropic | 57.88% | 官方榜单 | split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.92 +/-3.8 CI over 330 trials. | |
| Claude Opus 4.5 + Terminus 2 | Anthropic | 57.8% | 论文 | split: 2.0 scaffold: Terminus 2 Table 2 of the paper (Terminal-Bench 2.0). | |
| Claude Fable 5.1 (max) | Anthropic | 52.6% | 厂商自报 | split: science reasoning_effort: max cost_usd_per_task: 37.9 Anthropic's own harness (reproduces public Opus 5 at 29.0%, Fable 5 at 24.7%); SE +/-3.5-4.5 pts. Post dated 'September 2026', 1st used. Scaffold not stated. | |
| Opus 5 + Claude Code (max) | Anthropic | 51.82% | 官方榜单 | split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.09 Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials. | |
| Claude Opus 5 + Claude Code | Anthropic | 30% | 官方榜单 | split: science scaffold: Claude Code Terminal-Bench-Science 0.1 launch post (2026-08-27); 3 trials per task over 70 tasks; total run cost $7.0k. Best model at launch. Effort level not stated. | |
| Opus 4.8 + Claude Code (max) | Anthropic | 23.64% | 官方榜单 | split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 19.64 Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials. | |
| GPT-5.6 Sol + Codex | OpenAI | 22.4% | 官方榜单 | split: science scaffold: Codex Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $4.2k. Effort level not stated. | |
| Claude Fable 5 + Claude Code | Anthropic | 21.4% | 官方榜单 | split: science scaffold: Claude Code Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $14.2k. Effort level not stated. |