Research

Artificial Analysis Intelligence Index 2026

What the Artificial Analysis Intelligence Index measures, current 2026 scores for frontier and open-weight models, how token efficiency factors in, and how to read the index without over-trusting a single composite number.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: July 2026

The Artificial Analysis Intelligence Index has become the composite score most often quoted when someone says one model is smarter than another. It is a weighted aggregate across multiple evaluations rather than a single test, which makes it more robust than any individual benchmark and easier to misread than most people assume.

What It Measures

The index combines results across reasoning, knowledge, coding, and mathematics evaluations into one number on a normalised scale. Because it aggregates, a model cannot top it by overfitting one popular test, which is its main advantage over quoting a single SWE-bench or MMLU figure. Artificial Analysis also publishes cost and speed measurements alongside the intelligence score, and the three read together are far more informative than the index alone.

Open-Weight Standings, Mid-2026

ModelLabIndex score
GLM-5.2Z.ai (Zhipu)51
Nemotron 3 UltraNVIDIA48
MiniMax M3MiniMax44
DeepSeek V4 ProDeepSeek44
Kimi K2.6Moonshot AI43

Scores are index version v4.1. GLM-5.2 holding first place among open weights is the headline, and it is worth noting how tightly packed positions three through five are: a two-point spread across three models from three labs is well inside the range where evaluation noise and version changes can reorder the list.

The Token Efficiency Dimension

Artificial Analysis also reports how many output tokens a model consumes to reach its score, and this is the part most summaries drop. Gemini 3.6 Flash uses roughly 17 percent fewer output tokens than its predecessor on the index, which means its effective cost improvement exceeds its headline price cut. Two models with the same index score and the same per-token price are not the same price in practice if one is materially more verbose.

This matters more in 2026 than it used to because reasoning models generate long internal chains that the user never sees but always pays for. See the reasoning model pricing premium.

How To Read It Without Being Misled

Three cautions. A composite hides shape: two models scoring 44 can have very different strengths, and the one that suits your workload may be the lower-scoring one. Version changes are not comparable: an index revision reweights components, so a score from v4.1 cannot be compared against one from an earlier version. And the index measures capability, not deployment fit, so licensing, latency, context window, and serving cost decide production choices at least as often as capability does.

Brand Visibility Implications

Composite indices are unusually citable, which makes them unusually influential. When an AI assistant is asked which model is best, it tends to reach for whatever ranking is most cleanly structured and most widely repeated, and a single normalised number published in a consistent format is exactly that. For any brand publishing comparative data, the lesson generalises: structured, versioned, consistently formatted numbers get cited far more than equally accurate prose. See whether structured content improves citations.

Methodology

Scores as published by Artificial Analysis at index version v4.1, captured July 2026. Index composition and weighting are set by Artificial Analysis and change between versions. Figures move with each model release and index revision; re-check before quoting. Presenc AI is not affiliated with Artificial Analysis.

How Presenc AI Helps

Presenc AI tracks which sources AI assistants actually cite when asked comparative questions, including which benchmark authorities carry weight on which platforms.

Frequently Asked Questions

A composite score that aggregates multiple evaluations across reasoning, knowledge, coding, and mathematics into a single normalised number. Because it aggregates, a model cannot top it by overfitting one popular benchmark, which makes it more robust than quoting a single test result.
GLM-5.2 from Z.ai at 51 on index version v4.1, ahead of NVIDIA Nemotron 3 Ultra at 48, MiniMax M3 at 44, DeepSeek V4 Pro at 44, and Kimi K2.6 at 43. Positions three through five sit within a two-point spread, which is inside normal evaluation noise.
The index reports how many output tokens a model uses to reach its score. Two models with identical scores and identical per-token pricing cost different amounts in practice if one is more verbose. Gemini 3.6 Flash uses roughly 17 percent fewer output tokens than its predecessor, so its effective cost saving exceeds its headline price cut.
No. An index revision reweights its component evaluations, so a v4.1 score is not comparable to a score from an earlier version. Always check which index version a quoted number came from.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.