GPQA Diamond is a 198-question graduate-level science benchmark covering biology, physics, and chemistry, designed to be difficult for non-specialists even with internet access. This page snapshots the public leaderboard as of June 2026.
June 2026 Leaderboard
| Rank | Model | Vendor | GPQA Diamond % |
|---|---|---|---|
| 1 | Claude Mythos 5 | Anthropic | ~88% |
| 2 | GPT-5.6 Pro | OpenAI | ~86% |
| 3 | Claude Opus 4.7 | Anthropic | ~84% |
| 4 | Gemini 3.2 Pro | ~82% | |
| 5 | GPT-5.6 | OpenAI | ~78% |
| 6 | DeepSeek V4.1 Pro | DeepSeek | ~75% |
| 7 | Claude Sonnet 4.6 | Anthropic | ~72% |
| 8 | Qwen 3.7 | Alibaba | ~68% |
| 9 | Gemini 3.2 Flash | ~65% | |
| 10 | GLM-6 | Zhipu AI | ~62% |
| 11 | Llama 4.5 Maverick | Meta | ~58% |
| 12 | Mistral Large 3 | Mistral AI | ~55% |
Key Takeaways
- Claude Mythos 5 leads GPQA Diamond at approximately 88%, with human-expert performance around 65%.
- The top four frontier models all exceed human-expert performance.
- Reasoning-trace variants (Pro, Mythos) outperform base variants by 6 to 10 percentage points.
- Open-weight DeepSeek V4.1 Pro at ~75% sits within striking distance of Claude Sonnet 4.6.
Methodology
Scores compiled from vendor disclosures and the GPQA public results. Numbers approximate; reasoning-mode evaluations vary with prompt construction and sampling temperature. Updated monthly.
How Presenc AI Helps
Presenc AI tracks brand visibility on the reasoning models that increasingly handle complex enterprise evaluation workflows, where graduate-level reasoning capability shapes vendor selection.