OSWorld has two live versions and they give very different numbers. On the original benchmark, now called OSWorld-Verified, the top entry on the official leaderboard is Intelligence-Indeed Agent at 90.19 percent, above the 72.36 percent human baseline. On OSWorld 2.0, the highest figure we could verify as of October 1, 2026 is an 81.8 percent partial score for Claude Opus 5.5 on release v2.1, reported by Anthropic. The best entry on the official OSWorld 2.0 board is Claude Opus 5 at max effort, which fully completes 44.33 percent of tasks.
OSWorld 2.0 Leaderboard, Release v2.1
Ranked by partial score, because that is the number vendors quote. Only three results on the current release could be verified.
| Rank | Model | Partial score | Binary completion | Reported by | Date |
|---|---|---|---|---|---|
| 1 | Claude Opus 5.5 | 81.8% | 48.7%, per coverage of Anthropic's system card | Anthropic (vendor) | September 22, 2026 |
| 2 | Claude Fable 5.1 | 80.7% | Not published for v2.1 | Anthropic (vendor) | September 22, 2026 |
| 3 | Claude Opus 5 (max effort) | 77.67% | 44.33% | Official leaderboard | Board updated September 17, 2026 |
Anthropic's own run of Opus 5 on v2.1 gives 74.0 percent partial, 3.7 points below the official board's figure for the same model.
OSWorld 2.0 Results on Earlier or Unstated Releases
These are not ranked and do not compare with the table above.
| Model | Partial score | Binary completion | Release | Reported by |
|---|---|---|---|---|
| Claude Fable 5.1 | 77.9% | 41.7% | Not stated | Anthropic, September 1, 2026 |
| GPT-6 Astra | 72.6% | Not published | Offline subset, release not stated | OpenAI, per launch coverage, September 3, 2026 |
| GPT-6.1 Sol | 70.5% | Not published | Not stated | OpenAI, per third-party trackers, September 29, 2026 |
| Claude Opus 5 (max effort) | 68.31% | 31.43% | v2026.08.08 | Official leaderboard |
| GPT-5.6 Sol (max effort) | 62.72% | 27.34% | v2026.08.08 | Official leaderboard |
| Claude Opus 4.8 (max, batched tools) | 54.8% | 20.6% | v2026.06.24 | OSWorld 2.0 paper |
The coverage that gives GPT-6 Astra's 72.6 percent does not say which metric it is. The comparison figures quoted beside it sit close to the official partial scores for the offline subset, so we read it as a partial score. Model details are in the GPT-6 Astra, GPT-6.1 Sol, and Claude Fable 5.1 briefs.
OSWorld-Verified Leaderboard
| Rank | Agent or model | Success rate | Reported by | Date |
|---|---|---|---|---|
| 1 | Intelligence-Indeed Agent | 90.19% | Official leaderboard | 2026-07-25 |
| 2 | Claude Fable 5 | 85.96% | Official leaderboard | 2026-08-01 |
| 3 | Pointer Agent with Opus 4.7 | 83.64% | Official leaderboard | 2026-05-21 |
| 4 | Claude Opus 5 | 83.39% | Official leaderboard | 2026-08-01 |
| 5 | Coasty CUA v1 | 82.81% | Official leaderboard | 2026-07-01 |
All five rows use a 100-step limit. Every one is above the human baseline, which is the reason OSWorld 2.0 exists.
Benchmark Facts
| OSWorld-Verified | OSWorld 2.0 | |
|---|---|---|
| Released | OSWorld in 2024. Verified upgrade July 28, 2025 | June 26, 2026 |
| Tasks | 369 desktop tasks in apps such as Chrome, LibreOffice, GIMP, VS Code, and Thunderbird | 108 long workflows across 31 self-hosted websites and professional desktop apps |
| Human reference | 72.36% success | No success rate published. Median human time about 1.6 hours per task |
| Length | About 30 tool calls per task | About 318 tool calls per task |
| Scoring | Success rate | Binary completion at 500 steps, plus a partial score over an average of 27.25 checkpoints per task |
| Releases | One verified set | v2026.06.24, v2026.08.08, v2.1 (September 16, 2026) |
| Maintainer | XLANG Lab, University of Hong Kong | XLANG Lab, University of Hong Kong |
How to Read OSWorld Scores
- Check which benchmark. A score near 85 percent is OSWorld-Verified. A score near 80 percent partial on OSWorld 2.0 is a much harder result.
- Partial and binary are far apart. Opus 5 scores 77.67 percent partial and 44.33 percent binary on the same run. The paper's primary metric is binary. Vendors lead with partial.
- Releases do not compare. The August release updated 14 tasks to prevent reward hacking and modified 31 to make scoring more robust and fair. Opus 5 at max effort went from 31.43 percent binary on v2026.08.08 to 44.33 percent on v2.1 with no change to the model.
- Effort and step budget matter. On v2.1, Opus 5 ranges from 18.81 percent binary at low effort to 44.33 percent at max.
- Full set and offline subset differ. The board reports both, and the offline numbers are usually a few points higher.
For product comparisons see OpenCUA vs Claude computer use and browser use vs computer use. For the wider benchmark picture see AI agent capability benchmarks and the Terminal-Bench leaderboard.
Brand Visibility Implications
The models at the top of OSWorld are the ones vendors put inside computer-use products such as OpenAI Dots. Those agents open a browser, compare options, fill in forms, and pick a vendor for the user. The 2.0 results show they now finish close to half of hour-long workflows end to end and make partial progress on most of the rest. A brand's site therefore has to work for an agent that reads pricing pages and completes sign-up flows with nobody watching. See how AI agents choose brands.
Methodology
This page is built from official leaderboards, benchmark papers, and vendor announcements, not from Presenc AI measurements. Primary sources are the Terminal-Bench leaderboard and its 4.0 release note, the OSWorld 2.0 site, paper and repository, the OSWorld-Verified leaderboard, the LiveCodeBench leaderboard and repository, Anthropic's Claude Opus 5.5 announcement, and DeepSeek's V4 model card. Independent runs come from Artificial Analysis for Terminal-Bench 4.0 and LiveCodeBench. OpenAI's own pages could not be read by our fetcher, so OpenAI figures that are not on a public board come from launch coverage and are labelled that way. Every score is shown with who reported it and which benchmark version it belongs to. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks which brands AI assistants and agents recommend on the prompts your buyers use, and its crawl analytics show which agents reach your site. Together they show whether a computer-use agent can find you and whether it picks you.