Research

SWE-bench Pro Leaderboard 2026

SWE-bench Pro scores for the 2026 frontier and open-weight models, how Pro differs from SWE-bench Verified, why open-weight models lead on this benchmark, and how to read the gap.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: July 2026

SWE-bench Pro is the harder sibling of SWE-bench Verified and it produces a noticeably different ranking. On Verified, closed frontier models generally lead. On Pro, two open-weight models currently sit ahead of GPT-5.5. That divergence is the interesting part of this leaderboard.

Current Scores

ModelLabWeightsSWE-bench Pro
GLM-5.2Z.ai (Zhipu)Open, MIT62.1
MiniMax M3MiniMaxOpen59.0
GPT-5.5OpenAIClosed58.6
Qwen3-Coder-NextAlibabaOpen, Apache 2.044.3

Claude Opus 4.8 and Grok 4.5 also report Pro results, with Opus 4.8 ahead of Grok 4.5 on this benchmark specifically. Vendors report on different dates against different harness versions, so cross-vendor comparison carries more uncertainty than a single table suggests.

How Pro Differs from Verified

SWE-bench Verified is a human-filtered subset of the original SWE-bench, screened so that every task is solvable and unambiguously specified. That filtering makes it a clean measurement and also an easier one. Pro raises difficulty and reduces the headroom that filtering created, which is why absolute scores are lower and why the ordering shifts.

The practical implication: a model tuned hard against Verified may not carry that advantage to Pro. When a vendor quotes only one of the two, it is worth asking which one and why.

Why Open Weights Lead Here

Three plausible contributors, none individually conclusive. Chinese labs have optimised heavily for agentic coding because that is where their commercial traction is strongest, particularly through drop-in compatibility with existing coding-agent tooling. Open-weight models can be evaluated by anyone, so their reported numbers face more independent scrutiny and converge faster on reality. And closed labs increasingly emphasise long-horizon agentic evaluations over single-patch benchmarks, so SWE-bench Pro may simply be less central to their tuning.

Note also that leading on a coding benchmark is not the same as leading overall. GLM-5.2 tops this table and sits at 51 on the composite Artificial Analysis index, below the closed frontier. See the Intelligence Index.

Brand Visibility Implications

Coding benchmarks decide which model sits behind a developer's agent, and that agent is increasingly what answers "which library should I use". For developer-tool brands, the model winning SWE-bench Pro is a distribution channel. An open-weight model leading means those recommendations are generated with no retrieval and no observable API call, from weights fixed at training time. See the local-LLM visibility blind spot.

Methodology

Scores as reported by model vendors and independent evaluators through July 2026. Vendors publish against different harness versions on different dates, so treat small gaps as indistinguishable. SWE-bench Pro results are not comparable to SWE-bench Verified results. Updated as new results are published.

How Presenc AI Helps

Presenc AI measures how developer-tool brands are represented across the coding models teams actually deploy, open-weight and closed.

Frequently Asked Questions

Verified is a human-filtered subset of the original SWE-bench, screened so every task is solvable and unambiguously specified, which makes it cleaner and easier. Pro raises difficulty and removes much of that headroom, so absolute scores are lower and the model ordering shifts. Scores from the two are not comparable.
GLM-5.2 at 62.1, followed by MiniMax M3 at 59.0 and GPT-5.5 at 58.6. Both leaders are open-weight models, which inverts the usual ordering seen on SWE-bench Verified where closed frontier models generally lead.
Likely a combination: Chinese labs have optimised heavily for agentic coding where their commercial traction is strongest, open-weight numbers face more independent scrutiny and converge faster on reality, and closed labs increasingly prioritise long-horizon agentic evaluations over single-patch benchmarks.
No. GLM-5.2 tops SWE-bench Pro while scoring 51 on the composite Artificial Analysis Intelligence Index, below the closed frontier models. Coding benchmarks measure one capability, and composite indices exist precisely because single benchmarks do not generalise.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.