Terminal-Bench
Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces
Hard, human-authored tasks that an agent must complete inside a real terminal environment (software engineering, scientific computing, ML, security, data processing). Each task ships its own container, a human solution and tests; the score is the fraction of tasks whose tests pass, averaged over several trials with 95% confidence intervals. The de-facto benchmark for CLI coding agents, run through the Harbor framework.
paper · website · leaderboard · dataset · code
- Released
- 2026-01
- Maintainer
- Stanford / Laude Institute / Harbor (Merrill, Shaw et al.)
- Status
- active
- Contamination
- medium
- Metric
- resolution rate (percent, ↑)
- Tasks
- 66
- Domains
- software-engineering code tool-use ml-engineering
- human
- no measured baseline
Notes. Versioned benchmark: tasks are refreshed between releases (2.0 had 89 tasks, 4.0 has 66), so scores are only comparable within a version. Public tasks carry a canary string and the site asks that benchmark data never appear in training corpora. Terminal-Bench-Science (harbor-framework/terminal-bench-science, DOI 10.5281/zenodo.22110253) is a sibling benchmark in the same family and Harbor framework, tracked here as the 'science' split.
Full ledger
| System | Developer | Score | Date | Source | Conditions |
|---|---|---|---|---|---|
| GPT-6 Astra | OpenAI | 64.6% | self-reported | split: science OpenAI launch table (Academic section), maximum at any effort; run in OpenAI's research environment, scaffold not stated. Same table quotes Fable 5.1 at 52.6% and Opus 5 at 30.0%. | |
| GPT-5.2 + Codex CLI | OpenAI | 62.9% | paper | split: 2.0 scaffold: Codex CLI Best result in Table 2 of the paper (Terminal-Bench 2.0). | |
| GPT-6 Astra + Codex (max) | OpenAI | 58.18% | official | split: 4.0 scaffold: Codex reasoning_effort: max cost_usd_per_task: 9.9 Rank 1 on Terminal-Bench 4.0 as of access date; +/-2.8 CI over 330 trials. | |
| Fable 5.1 + Claude Code (max) | Anthropic | 57.88% | official | split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.92 +/-3.8 CI over 330 trials. | |
| Claude Opus 4.5 + Terminus 2 | Anthropic | 57.8% | paper | split: 2.0 scaffold: Terminus 2 Table 2 of the paper (Terminal-Bench 2.0). | |
| Claude Fable 5.1 (max) | Anthropic | 52.6% | self-reported | split: science reasoning_effort: max cost_usd_per_task: 37.9 Anthropic's own harness (reproduces public Opus 5 at 29.0%, Fable 5 at 24.7%); SE +/-3.5-4.5 pts. Post dated 'September 2026', 1st used. Scaffold not stated. | |
| Opus 5 + Claude Code (max) | Anthropic | 51.82% | official | split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.09 Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials. | |
| Claude Opus 5 + Claude Code | Anthropic | 30% | official | split: science scaffold: Claude Code Terminal-Bench-Science 0.1 launch post (2026-08-27); 3 trials per task over 70 tasks; total run cost $7.0k. Best model at launch. Effort level not stated. | |
| Opus 4.8 + Claude Code (max) | Anthropic | 23.64% | official | split: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 19.64 Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials. | |
| GPT-5.6 Sol + Codex | OpenAI | 22.4% | official | split: science scaffold: Codex Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $4.2k. Effort level not stated. | |
| Claude Fable 5 + Claude Code | Anthropic | 21.4% | official | split: science scaffold: Claude Code Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $14.2k. Effort level not stated. |