Terminal-Bench

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

由人工编写的高难任务,智能体需在真实终端环境中完成(软件工程、科学计算、ML、安全、数据处理)。每个任务自带容器、人类解答和测试;得分为测试通过的任务比例,对多次试验取平均并附 95% 置信区间。是 CLI 编程智能体的事实基准,通过 Harbor 框架运行。

64.6%
GPT-6 Astra
厂商自报
2026-01-17 → 2026-09-03: 57.8% → 64.6%
发布
2026-01
维护者
Stanford / Laude Institute / Harbor (Merrill, Shaw et al.)
状态
active
污染风险
medium
指标
resolution rate (percent, ↑)
题量
66
领域
software-engineering code tool-use ml-engineering
人类
无实测基线

备注. Versioned benchmark: tasks are refreshed between releases (2.0 had 89 tasks, 4.0 has 66), so scores are only comparable within a version. Public tasks carry a canary string and the site asks that benchmark data never appear in training corpora. Terminal-Bench-Science (harbor-framework/terminal-bench-science, DOI 10.5281/zenodo.22110253) is a sibling benchmark in the same family and Harbor framework, tracked here as the 'science' split.

完整账本

系统开发者分数日期来源条件
GPT-6 AstraOpenAI64.6%厂商自报split: science
OpenAI launch table (Academic section), maximum at any effort; run in OpenAI's research environment, scaffold not stated. Same table quotes Fable 5.1 at 52.6% and Opus 5 at 30.0%.
GPT-5.2 + Codex CLIOpenAI62.9%论文split: 2.0 scaffold: Codex CLI
Best result in Table 2 of the paper (Terminal-Bench 2.0).
GPT-6 Astra + Codex (max)OpenAI58.18%官方榜单split: 4.0 scaffold: Codex reasoning_effort: max cost_usd_per_task: 9.9
Rank 1 on Terminal-Bench 4.0 as of access date; +/-2.8 CI over 330 trials.
Fable 5.1 + Claude Code (max)Anthropic57.88%官方榜单split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.92
+/-3.8 CI over 330 trials.
Claude Opus 4.5 + Terminus 2Anthropic57.8%论文split: 2.0 scaffold: Terminus 2
Table 2 of the paper (Terminal-Bench 2.0).
Claude Fable 5.1 (max)Anthropic52.6%厂商自报split: science reasoning_effort: max cost_usd_per_task: 37.9
Anthropic's own harness (reproduces public Opus 5 at 29.0%, Fable 5 at 24.7%); SE +/-3.5-4.5 pts. Post dated 'September 2026', 1st used. Scaffold not stated.
Opus 5 + Claude Code (max)Anthropic51.82%官方榜单split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.09
Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials.
Claude Opus 5 + Claude CodeAnthropic30%官方榜单split: science scaffold: Claude Code
Terminal-Bench-Science 0.1 launch post (2026-08-27); 3 trials per task over 70 tasks; total run cost $7.0k. Best model at launch. Effort level not stated.
Opus 4.8 + Claude Code (max)Anthropic23.64%官方榜单split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 19.64
Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials.
GPT-5.6 Sol + CodexOpenAI22.4%官方榜单split: science scaffold: Codex
Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $4.2k. Effort level not stated.
Claude Fable 5 + Claude CodeAnthropic21.4%官方榜单split: science scaffold: Claude Code
Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $14.2k. Effort level not stated.