SWE-bench Pro is the harder sibling of SWE-bench Verified and it produces a noticeably different ranking. On Verified, closed frontier models generally lead. On Pro, two open-weight models currently sit ahead of GPT-5.5. That divergence is the interesting part of this leaderboard.
Current Scores
| Model | Lab | Weights | SWE-bench Pro |
|---|---|---|---|
| GLM-5.2 | Z.ai (Zhipu) | Open, MIT | 62.1 |
| MiniMax M3 | MiniMax | Open | 59.0 |
| GPT-5.5 | OpenAI | Closed | 58.6 |
| Qwen3-Coder-Next | Alibaba | Open, Apache 2.0 | 44.3 |
Claude Opus 4.8 and Grok 4.5 also report Pro results, with Opus 4.8 ahead of Grok 4.5 on this benchmark specifically. Vendors report on different dates against different harness versions, so cross-vendor comparison carries more uncertainty than a single table suggests.
How Pro Differs from Verified
SWE-bench Verified is a human-filtered subset of the original SWE-bench, screened so that every task is solvable and unambiguously specified. That filtering makes it a clean measurement and also an easier one. Pro raises difficulty and reduces the headroom that filtering created, which is why absolute scores are lower and why the ordering shifts.
The practical implication: a model tuned hard against Verified may not carry that advantage to Pro. When a vendor quotes only one of the two, it is worth asking which one and why.
Why Open Weights Lead Here
Three plausible contributors, none individually conclusive. Chinese labs have optimised heavily for agentic coding because that is where their commercial traction is strongest, particularly through drop-in compatibility with existing coding-agent tooling. Open-weight models can be evaluated by anyone, so their reported numbers face more independent scrutiny and converge faster on reality. And closed labs increasingly emphasise long-horizon agentic evaluations over single-patch benchmarks, so SWE-bench Pro may simply be less central to their tuning.
Note also that leading on a coding benchmark is not the same as leading overall. GLM-5.2 tops this table and sits at 51 on the composite Artificial Analysis index, below the closed frontier. See the Intelligence Index.
Brand Visibility Implications
Coding benchmarks decide which model sits behind a developer's agent, and that agent is increasingly what answers "which library should I use". For developer-tool brands, the model winning SWE-bench Pro is a distribution channel. An open-weight model leading means those recommendations are generated with no retrieval and no observable API call, from weights fixed at training time. See the local-LLM visibility blind spot.
Methodology
Scores as reported by model vendors and independent evaluators through July 2026. Vendors publish against different harness versions on different dates, so treat small gaps as indistinguishable. SWE-bench Pro results are not comparable to SWE-bench Verified results. Updated as new results are published.
How Presenc AI Helps
Presenc AI measures how developer-tool brands are represented across the coding models teams actually deploy, open-weight and closed.