Research

GPQA Diamond Leaderboard June 2026

GPQA Diamond reasoning benchmark leaderboard for June 2026. Claude Mythos 5 and GPT-5.6 Pro lead the frontier on graduate-level science Q&A.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: June 2026

GPQA Diamond is a 198-question graduate-level science benchmark covering biology, physics, and chemistry, designed to be difficult for non-specialists even with internet access. This page snapshots the public leaderboard as of June 2026.

June 2026 Leaderboard

RankModelVendorGPQA Diamond %
1Claude Mythos 5Anthropic~88%
2GPT-5.6 ProOpenAI~86%
3Claude Opus 4.7Anthropic~84%
4Gemini 3.2 ProGoogle~82%
5GPT-5.6OpenAI~78%
6DeepSeek V4.1 ProDeepSeek~75%
7Claude Sonnet 4.6Anthropic~72%
8Qwen 3.7Alibaba~68%
9Gemini 3.2 FlashGoogle~65%
10GLM-6Zhipu AI~62%
11Llama 4.5 MaverickMeta~58%
12Mistral Large 3Mistral AI~55%

Key Takeaways

  • Claude Mythos 5 leads GPQA Diamond at approximately 88%, with human-expert performance around 65%.
  • The top four frontier models all exceed human-expert performance.
  • Reasoning-trace variants (Pro, Mythos) outperform base variants by 6 to 10 percentage points.
  • Open-weight DeepSeek V4.1 Pro at ~75% sits within striking distance of Claude Sonnet 4.6.

Methodology

Scores compiled from vendor disclosures and the GPQA public results. Numbers approximate; reasoning-mode evaluations vary with prompt construction and sampling temperature. Updated monthly.

How Presenc AI Helps

Presenc AI tracks brand visibility on the reasoning models that increasingly handle complex enterprise evaluation workflows, where graduate-level reasoning capability shapes vendor selection.

Frequently Asked Questions

A 198-question graduate-level science benchmark covering biology, physics, and chemistry, designed to be difficult for non-specialists even with internet access.
Claude Mythos 5 from Anthropic at approximately 88%, ahead of GPT-5.6 Pro at 86% and Claude Opus 4.7 at 84%.
Yes. The top four models exceed approximately 65% human-expert baseline as of June 2026. Frontier capability on graduate-level science has crossed the human threshold.
Reasoning-mode inference uses extended chain-of-thought, longer compute budgets, and self-verification steps that base-mode does not apply. The trade-off is meaningfully higher per-query latency and cost.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.