Terminal-Bench

Terminal-Bench: Benchmarking Agents on Hard, Realistic Tasks in Command Line Interfaces

Hard, human-authored tasks that an agent must complete inside a real terminal environment (software engineering, scientific computing, ML, security, data processing). Each task ships its own container, a human solution and tests; the score is the fraction of tasks whose tests pass, averaged over several trials with 95% confidence intervals. The de-facto benchmark for CLI coding agents, run through the Harbor framework.

64.6%
GPT-6 Astra
self-reported
2026-01-17 → 2026-09-03: 57.8% → 64.6%
Released
2026-01
Maintainer
Stanford / Laude Institute / Harbor (Merrill, Shaw et al.)
Status
active
Contamination
medium
Metric
resolution rate (percent, ↑)
Tasks
66
Domains
software-engineering code tool-use ml-engineering
human
no measured baseline

Notes. Versioned benchmark: tasks are refreshed between releases (2.0 had 89 tasks, 4.0 has 66), so scores are only comparable within a version. Public tasks carry a canary string and the site asks that benchmark data never appear in training corpora. Terminal-Bench-Science (harbor-framework/terminal-bench-science, DOI 10.5281/zenodo.22110253) is a sibling benchmark in the same family and Harbor framework, tracked here as the 'science' split.

Full ledger

SystemDeveloperScoreDateSourceConditions
GPT-6 AstraOpenAI64.6%self-reportedsplit: science
OpenAI launch table (Academic section), maximum at any effort; run in OpenAI's research environment, scaffold not stated. Same table quotes Fable 5.1 at 52.6% and Opus 5 at 30.0%.
GPT-5.2 + Codex CLIOpenAI62.9%papersplit: 2.0 scaffold: Codex CLI
Best result in Table 2 of the paper (Terminal-Bench 2.0).
GPT-6 Astra + Codex (max)OpenAI58.18%officialsplit: 4.0 scaffold: Codex reasoning_effort: max cost_usd_per_task: 9.9
Rank 1 on Terminal-Bench 4.0 as of access date; +/-2.8 CI over 330 trials.
Fable 5.1 + Claude Code (max)Anthropic57.88%officialsplit: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.92
+/-3.8 CI over 330 trials.
Claude Opus 4.5 + Terminus 2Anthropic57.8%papersplit: 2.0 scaffold: Terminus 2
Table 2 of the paper (Terminal-Bench 2.0).
Claude Fable 5.1 (max)Anthropic52.6%self-reportedsplit: science reasoning_effort: max cost_usd_per_task: 37.9
Anthropic's own harness (reproduces public Opus 5 at 29.0%, Fable 5 at 24.7%); SE +/-3.5-4.5 pts. Post dated 'September 2026', 1st used. Scaffold not stated.
Opus 5 + Claude Code (max)Anthropic51.82%officialsplit: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 18.09
Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials.
Claude Opus 5 + Claude CodeAnthropic30%officialsplit: science scaffold: Claude Code
Terminal-Bench-Science 0.1 launch post (2026-08-27); 3 trials per task over 70 tasks; total run cost $7.0k. Best model at launch. Effort level not stated.
Opus 4.8 + Claude Code (max)Anthropic23.64%officialsplit: 4.0 scaffold: Claude Code reasoning_effort: max cost_usd_per_task: 19.64
Leaderboard 'created_at' date; 330 trials. Cost is total run cost / 330 trials.
GPT-5.6 Sol + CodexOpenAI22.4%officialsplit: science scaffold: Codex
Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $4.2k. Effort level not stated.
Claude Fable 5 + Claude CodeAnthropic21.4%officialsplit: science scaffold: Claude Code
Terminal-Bench-Science 0.1 launch post; 3 trials per task; total run cost $14.2k. Effort level not stated.