The Artificial Analysis Intelligence Index has become the composite score most often quoted when someone says one model is smarter than another. It is a weighted aggregate across multiple evaluations rather than a single test, which makes it more robust than any individual benchmark and easier to misread than most people assume.
What It Measures
The index combines results across reasoning, knowledge, coding, and mathematics evaluations into one number on a normalised scale. Because it aggregates, a model cannot top it by overfitting one popular test, which is its main advantage over quoting a single SWE-bench or MMLU figure. Artificial Analysis also publishes cost and speed measurements alongside the intelligence score, and the three read together are far more informative than the index alone.
Open-Weight Standings, Mid-2026
| Model | Lab | Index score |
|---|---|---|
| GLM-5.2 | Z.ai (Zhipu) | 51 |
| Nemotron 3 Ultra | NVIDIA | 48 |
| MiniMax M3 | MiniMax | 44 |
| DeepSeek V4 Pro | DeepSeek | 44 |
| Kimi K2.6 | Moonshot AI | 43 |
Scores are index version v4.1. GLM-5.2 holding first place among open weights is the headline, and it is worth noting how tightly packed positions three through five are: a two-point spread across three models from three labs is well inside the range where evaluation noise and version changes can reorder the list.
The Token Efficiency Dimension
Artificial Analysis also reports how many output tokens a model consumes to reach its score, and this is the part most summaries drop. Gemini 3.6 Flash uses roughly 17 percent fewer output tokens than its predecessor on the index, which means its effective cost improvement exceeds its headline price cut. Two models with the same index score and the same per-token price are not the same price in practice if one is materially more verbose.
This matters more in 2026 than it used to because reasoning models generate long internal chains that the user never sees but always pays for. See the reasoning model pricing premium.
How To Read It Without Being Misled
Three cautions. A composite hides shape: two models scoring 44 can have very different strengths, and the one that suits your workload may be the lower-scoring one. Version changes are not comparable: an index revision reweights components, so a score from v4.1 cannot be compared against one from an earlier version. And the index measures capability, not deployment fit, so licensing, latency, context window, and serving cost decide production choices at least as often as capability does.
Brand Visibility Implications
Composite indices are unusually citable, which makes them unusually influential. When an AI assistant is asked which model is best, it tends to reach for whatever ranking is most cleanly structured and most widely repeated, and a single normalised number published in a consistent format is exactly that. For any brand publishing comparative data, the lesson generalises: structured, versioned, consistently formatted numbers get cited far more than equally accurate prose. See whether structured content improves citations.
Methodology
Scores as published by Artificial Analysis at index version v4.1, captured July 2026. Index composition and weighting are set by Artificial Analysis and change between versions. Figures move with each model release and index revision; re-check before quoting. Presenc AI is not affiliated with Artificial Analysis.
How Presenc AI Helps
Presenc AI tracks which sources AI assistants actually cite when asked comparative questions, including which benchmark authorities carry weight on which platforms.