As of October 1, 2026, GPT-6 Astra running in OpenAI's Codex agent leads the official Terminal-Bench 4.0 leaderboard at 58.18 percent, plus or minus 2.79 points. Claude Fable 5.1 in Claude Code is at 57.88 percent, inside that margin, so the top of the public board is a tie. Anthropic reports 66.4 percent for Claude Opus 5.5, but that is Anthropic's own run and it is not on the board. Version 4.0 is the current Terminal-Bench. Vendors still quote versions 2.0, 2.1, and 3.0, and those numbers do not compare with 4.0.
Terminal-Bench 4.0 Leaderboard
The public board at tbench.ai lists 27 rows because it scores the same model at several effort settings. This table shows the best row for each model, so ranks count models, not rows.
| Rank | Model and agent | Effort | Score (95% interval) | Reported by | Date on board |
|---|---|---|---|---|---|
| 1 | GPT-6 Astra, Codex | max | 58.18% ± 2.79 | Official leaderboard | 2026-09-03 |
| 2 | Claude Fable 5.1, Claude Code | max | 57.88% ± 3.76 | Official leaderboard | 2026-09-01 |
| 3 | Claude Opus 5, Claude Code | xhigh | 53.94% ± 3.17 | Official leaderboard | 2026-07-24 |
| 4 | Claude Fable 5, Claude Code | max | 44.55% ± 3.85 | Official leaderboard | 2026-06-09 |
| 5 | GLM-5.3, Claude Code | max | 41.82% ± 3.23 | Official leaderboard | 2026-08-14 |
| 6 | Grok 4.7, Grok Build | xhigh | 37.58% ± 3.54 | Official leaderboard | 2026-09-21 |
| 7 | GPT-5.6 Sol, Codex | max | 37.27% ± 3.78 | Official leaderboard | 2026-06-26 |
| 8 | Claude Opus 4.8, Claude Code | max | 23.64% ± 3.56 | Official leaderboard | 2026-05-28 |
| 9 | GPT-5.6 Terra, Codex | max | 21.52% ± 3.25 | Official leaderboard | 2026-06-26 |
| 10 | Grok 4.6, Grok Build | high | 20.30% ± 3.09 | Official leaderboard | 2026-08-12 |
| 11 | Gemini 3.8 Flash, mini-SWE-agent | high | 19.09% ± 3.36 | Official leaderboard | 2026-09-02 |
Version 4.0 Results Not on the Public Board
| Model | Score | Setup | Reported by | Date |
|---|---|---|---|---|
| Claude Opus 5.5 | 66.4% | xhigh effort, Anthropic's own setup, standard error 2.6 points | Anthropic (vendor) | September 22, 2026 |
| Claude Sonnet 5.5 | 63.6% | max effort, mini-swe-agent, three repeats per task | Artificial Analysis, independent | Viewed October 1, 2026 |
| Claude Mythos 5.1 | 60.9% | Restricted-access model | Anthropic (vendor) | September 1, 2026 |
| Claude Opus 5.5 | 59.6% | max effort, mini-swe-agent, three repeats per task | Artificial Analysis, independent | Viewed October 1, 2026 |
Opus 5.5 differs by almost seven points between Anthropic's run and the independent one. The same gap shows up in smaller form for older models. Anthropic's table lists Fable 5.1 at 55.8 percent and Opus 5 at 52.3 percent from its own setup, while the public board has them at 57.88 and 53.94 percent. Anthropic's footnote says the board's 51.8 percent for Opus 5 at max effort and its own 52.3 percent agree within noise. Details for each model are in the GPT-6 Astra, Claude Fable 5.1, Claude Opus 5, and Grok 4.7 briefs.
Benchmark Facts
| Item | Detail |
|---|---|
| What it tests | An agent is given a task in a containerised terminal environment and must finish it without help. A verifier then checks the result |
| Task count | 66 in version 4.0 |
| Domains | Software, machine learning, science, operations, security, hardware, and media |
| Scoring | Resolution rate across 5 trials per task, 330 trials per entry, shown with a 95 percent confidence interval. Each task has a flat 8-hour agent timeout |
| Human baseline | None published |
| Maintainers | Hosted by Stanford, the Laude Institute, and Harbor, the open-source framework the benchmark runs on |
| Versions | 1.0 (May 19, 2025). 2.0 (November 7, 2025, 89 tasks). 2.1 (May 6, 2026, 28 tasks fixed). 3.0 (July 30, 2026, 74 tasks). 4.0 (August 28, 2026, 8 tasks removed and 19 fixed) |
How to Read Terminal-Bench Scores
- Check the version first. When 2.1 launched, the best scores were 76.0 percent on 2.0 and 79.1 percent on 2.1. Version 3.0 opened with a best score of 34.4 percent. Version 4.0 tops out near 58 percent. A model card quoting a Terminal-Bench score in the 80s is quoting 2.x.
- A score belongs to a model and an agent together. GLM-5.3 is scored inside Claude Code. Gemini 3.8 Flash is scored with mini-SWE-agent. Change the agent and the number changes.
- Effort settings move results a lot. On the public board Claude Opus 5 runs from 34.85 percent at low effort to 53.94 percent at xhigh, and its max run scored lower than xhigh. GPT-6 Astra runs from 50.61 percent at low to 58.18 percent at max.
- Intervals are wide. With 66 tasks, the 95 percent interval is about 2.5 to 4 points either way. Gaps smaller than that are not rankings.
- Vendor runs are not board entries. A vendor can pick its own harness and safeguards. Treat a vendor number as a claim until it appears on the board or in an independent run.
- Cost is part of the result. The board lists the top GPT-6 Astra run at $3,267 and the top Fable 5.1 run at $6,244 for nearly the same score.
For other coding and agent benchmarks see coding agent benchmarks, the SWE-bench Pro leaderboard, the LiveCodeBench leaderboard, and the OSWorld leaderboard.
Brand Visibility Implications
Terminal-Bench scores a model inside a named agent, and the agents on this board are products developers use every day: Codex, Claude Code, and Grok Build. The model and agent pairs that lead here are the ones teams let run commands without supervision. In that mode the agent picks the package to install, the CLI to call, and the cloud service to configure, often with no person reading the options. A developer-tool brand that the leading models do not know, or describe wrongly, is skipped at the moment of installation. See how AI agents choose brands and the AI coding agent market.
Methodology
This page is built from official leaderboards, benchmark papers, and vendor announcements, not from Presenc AI measurements. Primary sources are the Terminal-Bench leaderboard and its 4.0 release note, the OSWorld 2.0 site, paper and repository, the OSWorld-Verified leaderboard, the LiveCodeBench leaderboard and repository, Anthropic's Claude Opus 5.5 announcement, and DeepSeek's V4 model card. Independent runs come from Artificial Analysis for Terminal-Bench 4.0 and LiveCodeBench. OpenAI's own pages could not be read by our fetcher, so OpenAI figures that are not on a public board come from launch coverage and are labelled that way. Every score is shown with who reported it and which benchmark version it belongs to. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks which tools, libraries, and vendors AI models recommend on the prompts your buyers and their agents use. That shows whether the models at the top of this board know your product and describe it correctly.