Research

OSWorld Computer-Use Leaderboard 2026

OSWorld scores as of October 1, 2026. Agents pass the 72.36 percent human baseline on OSWorld-Verified, while on OSWorld 2.0 the best official entry fully completes 44 percent of tasks. Versions explained.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

OSWorld has two live versions and they give very different numbers. On the original benchmark, now called OSWorld-Verified, the top entry on the official leaderboard is Intelligence-Indeed Agent at 90.19 percent, above the 72.36 percent human baseline. On OSWorld 2.0, the highest figure we could verify as of October 1, 2026 is an 81.8 percent partial score for Claude Opus 5.5 on release v2.1, reported by Anthropic. The best entry on the official OSWorld 2.0 board is Claude Opus 5 at max effort, which fully completes 44.33 percent of tasks.

OSWorld 2.0 Leaderboard, Release v2.1

Ranked by partial score, because that is the number vendors quote. Only three results on the current release could be verified.

RankModelPartial scoreBinary completionReported byDate
1Claude Opus 5.581.8%48.7%, per coverage of Anthropic's system cardAnthropic (vendor)September 22, 2026
2Claude Fable 5.180.7%Not published for v2.1Anthropic (vendor)September 22, 2026
3Claude Opus 5 (max effort)77.67%44.33%Official leaderboardBoard updated September 17, 2026

Anthropic's own run of Opus 5 on v2.1 gives 74.0 percent partial, 3.7 points below the official board's figure for the same model.

OSWorld 2.0 Results on Earlier or Unstated Releases

These are not ranked and do not compare with the table above.

ModelPartial scoreBinary completionReleaseReported by
Claude Fable 5.177.9%41.7%Not statedAnthropic, September 1, 2026
GPT-6 Astra72.6%Not publishedOffline subset, release not statedOpenAI, per launch coverage, September 3, 2026
GPT-6.1 Sol70.5%Not publishedNot statedOpenAI, per third-party trackers, September 29, 2026
Claude Opus 5 (max effort)68.31%31.43%v2026.08.08Official leaderboard
GPT-5.6 Sol (max effort)62.72%27.34%v2026.08.08Official leaderboard
Claude Opus 4.8 (max, batched tools)54.8%20.6%v2026.06.24OSWorld 2.0 paper

The coverage that gives GPT-6 Astra's 72.6 percent does not say which metric it is. The comparison figures quoted beside it sit close to the official partial scores for the offline subset, so we read it as a partial score. Model details are in the GPT-6 Astra, GPT-6.1 Sol, and Claude Fable 5.1 briefs.

OSWorld-Verified Leaderboard

RankAgent or modelSuccess rateReported byDate
1Intelligence-Indeed Agent90.19%Official leaderboard2026-07-25
2Claude Fable 585.96%Official leaderboard2026-08-01
3Pointer Agent with Opus 4.783.64%Official leaderboard2026-05-21
4Claude Opus 583.39%Official leaderboard2026-08-01
5Coasty CUA v182.81%Official leaderboard2026-07-01

All five rows use a 100-step limit. Every one is above the human baseline, which is the reason OSWorld 2.0 exists.

Benchmark Facts

OSWorld-VerifiedOSWorld 2.0
ReleasedOSWorld in 2024. Verified upgrade July 28, 2025June 26, 2026
Tasks369 desktop tasks in apps such as Chrome, LibreOffice, GIMP, VS Code, and Thunderbird108 long workflows across 31 self-hosted websites and professional desktop apps
Human reference72.36% successNo success rate published. Median human time about 1.6 hours per task
LengthAbout 30 tool calls per taskAbout 318 tool calls per task
ScoringSuccess rateBinary completion at 500 steps, plus a partial score over an average of 27.25 checkpoints per task
ReleasesOne verified setv2026.06.24, v2026.08.08, v2.1 (September 16, 2026)
MaintainerXLANG Lab, University of Hong KongXLANG Lab, University of Hong Kong

How to Read OSWorld Scores

  • Check which benchmark. A score near 85 percent is OSWorld-Verified. A score near 80 percent partial on OSWorld 2.0 is a much harder result.
  • Partial and binary are far apart. Opus 5 scores 77.67 percent partial and 44.33 percent binary on the same run. The paper's primary metric is binary. Vendors lead with partial.
  • Releases do not compare. The August release updated 14 tasks to prevent reward hacking and modified 31 to make scoring more robust and fair. Opus 5 at max effort went from 31.43 percent binary on v2026.08.08 to 44.33 percent on v2.1 with no change to the model.
  • Effort and step budget matter. On v2.1, Opus 5 ranges from 18.81 percent binary at low effort to 44.33 percent at max.
  • Full set and offline subset differ. The board reports both, and the offline numbers are usually a few points higher.

For product comparisons see OpenCUA vs Claude computer use and browser use vs computer use. For the wider benchmark picture see AI agent capability benchmarks and the Terminal-Bench leaderboard.

Brand Visibility Implications

The models at the top of OSWorld are the ones vendors put inside computer-use products such as OpenAI Dots. Those agents open a browser, compare options, fill in forms, and pick a vendor for the user. The 2.0 results show they now finish close to half of hour-long workflows end to end and make partial progress on most of the rest. A brand's site therefore has to work for an agent that reads pricing pages and completes sign-up flows with nobody watching. See how AI agents choose brands.

Methodology

This page is built from official leaderboards, benchmark papers, and vendor announcements, not from Presenc AI measurements. Primary sources are the Terminal-Bench leaderboard and its 4.0 release note, the OSWorld 2.0 site, paper and repository, the OSWorld-Verified leaderboard, the LiveCodeBench leaderboard and repository, Anthropic's Claude Opus 5.5 announcement, and DeepSeek's V4 model card. Independent runs come from Artificial Analysis for Terminal-Bench 4.0 and LiveCodeBench. OpenAI's own pages could not be read by our fetcher, so OpenAI figures that are not on a public board come from launch coverage and are labelled that way. Every score is shown with who reported it and which benchmark version it belongs to. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks which brands AI assistants and agents recommend on the prompts your buyers use, and its crawl analytics show which agents reach your site. Together they show whether a computer-use agent can find you and whether it picks you.

Frequently Asked Questions

OSWorld is a benchmark for computer-use agents from XLANG Lab at the University of Hong Kong. An agent operates real desktop and web applications inside a virtual machine and is scored on whether the task is done. The original set, now OSWorld-Verified, has 369 tasks. OSWorld 2.0, released June 26, 2026, has 108 much longer workflows.
On the original OSWorld, humans complete 72.36 percent of the 369 tasks. Several agents now exceed that on OSWorld-Verified, led by Intelligence-Indeed Agent at 90.19 percent. OSWorld 2.0 does not publish a human success rate. It reports that its tasks take a skilled person a median of about 1.6 hours each.
As of October 1, 2026, the highest verified figure is Claude Opus 5.5 at 81.8 percent partial score on release v2.1, reported by Anthropic. On the official board, last updated September 17, 2026, the top entry is Claude Opus 5 at max effort with 77.67 percent partial and 44.33 percent binary completion. Scores on earlier releases are not comparable.
OSWorld-Verified is the July 2025 cleaned-up version of the original 369 short desktop tasks, which take about 30 tool calls each. OSWorld 2.0 is a new set of 108 long workflows that take about 318 tool calls each, scored with both binary completion and a partial score. Top agents score above 80 percent on Verified and under 50 percent binary on 2.0.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.