Research

Terminal-Bench Leaderboard 2026

Terminal-Bench 4.0 scores as of October 1, 2026. GPT-6 Astra and Claude Fable 5.1 are tied near 58 percent on the public board, Anthropic reports 66.4 percent for Opus 5.5, and 2.x scores do not compare.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

As of October 1, 2026, GPT-6 Astra running in OpenAI's Codex agent leads the official Terminal-Bench 4.0 leaderboard at 58.18 percent, plus or minus 2.79 points. Claude Fable 5.1 in Claude Code is at 57.88 percent, inside that margin, so the top of the public board is a tie. Anthropic reports 66.4 percent for Claude Opus 5.5, but that is Anthropic's own run and it is not on the board. Version 4.0 is the current Terminal-Bench. Vendors still quote versions 2.0, 2.1, and 3.0, and those numbers do not compare with 4.0.

Terminal-Bench 4.0 Leaderboard

The public board at tbench.ai lists 27 rows because it scores the same model at several effort settings. This table shows the best row for each model, so ranks count models, not rows.

RankModel and agentEffortScore (95% interval)Reported byDate on board
1GPT-6 Astra, Codexmax58.18% ± 2.79Official leaderboard2026-09-03
2Claude Fable 5.1, Claude Codemax57.88% ± 3.76Official leaderboard2026-09-01
3Claude Opus 5, Claude Codexhigh53.94% ± 3.17Official leaderboard2026-07-24
4Claude Fable 5, Claude Codemax44.55% ± 3.85Official leaderboard2026-06-09
5GLM-5.3, Claude Codemax41.82% ± 3.23Official leaderboard2026-08-14
6Grok 4.7, Grok Buildxhigh37.58% ± 3.54Official leaderboard2026-09-21
7GPT-5.6 Sol, Codexmax37.27% ± 3.78Official leaderboard2026-06-26
8Claude Opus 4.8, Claude Codemax23.64% ± 3.56Official leaderboard2026-05-28
9GPT-5.6 Terra, Codexmax21.52% ± 3.25Official leaderboard2026-06-26
10Grok 4.6, Grok Buildhigh20.30% ± 3.09Official leaderboard2026-08-12
11Gemini 3.8 Flash, mini-SWE-agenthigh19.09% ± 3.36Official leaderboard2026-09-02

Version 4.0 Results Not on the Public Board

ModelScoreSetupReported byDate
Claude Opus 5.566.4%xhigh effort, Anthropic's own setup, standard error 2.6 pointsAnthropic (vendor)September 22, 2026
Claude Sonnet 5.563.6%max effort, mini-swe-agent, three repeats per taskArtificial Analysis, independentViewed October 1, 2026
Claude Mythos 5.160.9%Restricted-access modelAnthropic (vendor)September 1, 2026
Claude Opus 5.559.6%max effort, mini-swe-agent, three repeats per taskArtificial Analysis, independentViewed October 1, 2026

Opus 5.5 differs by almost seven points between Anthropic's run and the independent one. The same gap shows up in smaller form for older models. Anthropic's table lists Fable 5.1 at 55.8 percent and Opus 5 at 52.3 percent from its own setup, while the public board has them at 57.88 and 53.94 percent. Anthropic's footnote says the board's 51.8 percent for Opus 5 at max effort and its own 52.3 percent agree within noise. Details for each model are in the GPT-6 Astra, Claude Fable 5.1, Claude Opus 5, and Grok 4.7 briefs.

Benchmark Facts

ItemDetail
What it testsAn agent is given a task in a containerised terminal environment and must finish it without help. A verifier then checks the result
Task count66 in version 4.0
DomainsSoftware, machine learning, science, operations, security, hardware, and media
ScoringResolution rate across 5 trials per task, 330 trials per entry, shown with a 95 percent confidence interval. Each task has a flat 8-hour agent timeout
Human baselineNone published
MaintainersHosted by Stanford, the Laude Institute, and Harbor, the open-source framework the benchmark runs on
Versions1.0 (May 19, 2025). 2.0 (November 7, 2025, 89 tasks). 2.1 (May 6, 2026, 28 tasks fixed). 3.0 (July 30, 2026, 74 tasks). 4.0 (August 28, 2026, 8 tasks removed and 19 fixed)

How to Read Terminal-Bench Scores

  • Check the version first. When 2.1 launched, the best scores were 76.0 percent on 2.0 and 79.1 percent on 2.1. Version 3.0 opened with a best score of 34.4 percent. Version 4.0 tops out near 58 percent. A model card quoting a Terminal-Bench score in the 80s is quoting 2.x.
  • A score belongs to a model and an agent together. GLM-5.3 is scored inside Claude Code. Gemini 3.8 Flash is scored with mini-SWE-agent. Change the agent and the number changes.
  • Effort settings move results a lot. On the public board Claude Opus 5 runs from 34.85 percent at low effort to 53.94 percent at xhigh, and its max run scored lower than xhigh. GPT-6 Astra runs from 50.61 percent at low to 58.18 percent at max.
  • Intervals are wide. With 66 tasks, the 95 percent interval is about 2.5 to 4 points either way. Gaps smaller than that are not rankings.
  • Vendor runs are not board entries. A vendor can pick its own harness and safeguards. Treat a vendor number as a claim until it appears on the board or in an independent run.
  • Cost is part of the result. The board lists the top GPT-6 Astra run at $3,267 and the top Fable 5.1 run at $6,244 for nearly the same score.

For other coding and agent benchmarks see coding agent benchmarks, the SWE-bench Pro leaderboard, the LiveCodeBench leaderboard, and the OSWorld leaderboard.

Brand Visibility Implications

Terminal-Bench scores a model inside a named agent, and the agents on this board are products developers use every day: Codex, Claude Code, and Grok Build. The model and agent pairs that lead here are the ones teams let run commands without supervision. In that mode the agent picks the package to install, the CLI to call, and the cloud service to configure, often with no person reading the options. A developer-tool brand that the leading models do not know, or describe wrongly, is skipped at the moment of installation. See how AI agents choose brands and the AI coding agent market.

Methodology

This page is built from official leaderboards, benchmark papers, and vendor announcements, not from Presenc AI measurements. Primary sources are the Terminal-Bench leaderboard and its 4.0 release note, the OSWorld 2.0 site, paper and repository, the OSWorld-Verified leaderboard, the LiveCodeBench leaderboard and repository, Anthropic's Claude Opus 5.5 announcement, and DeepSeek's V4 model card. Independent runs come from Artificial Analysis for Terminal-Bench 4.0 and LiveCodeBench. OpenAI's own pages could not be read by our fetcher, so OpenAI figures that are not on a public board come from launch coverage and are labelled that way. Every score is shown with who reported it and which benchmark version it belongs to. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks which tools, libraries, and vendors AI models recommend on the prompts your buyers and their agents use. That shows whether the models at the top of this board know your product and describe it correctly.

Frequently Asked Questions

Terminal-Bench is a benchmark that gives an AI agent tasks to complete on its own inside a containerised terminal, then checks the result with a verifier. Version 4.0 has 66 tasks across software, machine learning, science, operations, security, hardware, and media. It is hosted by Stanford, the Laude Institute, and Harbor, and each entry is scored over 5 trials per task.
As of October 1, 2026, GPT-6 Astra in OpenAI's Codex agent leads the public Terminal-Bench 4.0 board at 58.18 percent, with Claude Fable 5.1 in Claude Code at 57.88 percent. The two are inside each other's confidence intervals. Anthropic reports 66.4 percent for Claude Opus 5.5 from its own run, which is not on the public board.
Terminal-Bench 4.0, released on August 28, 2026. It followed 3.0 on July 30, 2026, 2.1 on May 6, 2026, and 2.0 on November 7, 2025. Version 4.0 removed 8 tasks that were saturated, refusal-prone, publicly solved, or flawed, fixed 19 more, and set a flat 8-hour timeout.
They are different task sets. Versions 2.0 and 2.1 had 89 easier tasks and top scores reached the high 70s and above. Version 3.0 replaced them with 74 harder tasks, and 4.0 refined that set to 66. Scores from 2.x cannot be compared with 4.0, so always check which version a vendor is quoting.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.