Benchmark leaderboards measure what models can do. OpenRouter's token-volume data measures what developers actually route traffic to, and the two lists look very different. This page covers what the usage data shows and why the divergence matters.
Token Volume Share, Mid-2026
| Metric | Value |
|---|---|
| Chinese-origin share of identified token volume | ~46% |
| Chinese-origin share one year earlier | Below 2% |
| DeepSeek alone | ~17.6% |
| Anthropic token share | ~12.3% |
| Top model by monthly tokens, February 2026 | MiniMax M2.5, ~4.55T |
| Second, February 2026 | Kimi K2.5, ~4.02T |
| Top model by monthly tokens, April 2026 | MiMo-V2-Pro, ~4.65T |
The move from under 2 percent to roughly 46 percent of token volume in twelve months is one of the sharpest share shifts in the platform's history, and it happened without any corresponding shift in benchmark leadership.
Volume Share Is Not Revenue Share
Anthropic holds roughly 12.3 percent of tokens but a substantially higher share of dollars, because premium models are priced many times above the cheap open-weight tier. MiniMax M2.5 runs around $0.30 per million input and $1.20 per million output; Claude Opus-class pricing reaches $5 to $25 per million. That is a 17 to 20x spread against models delivering near-parity on several benchmarks.
Reading only token volume overstates how much of the economically valuable work Chinese models are doing. Reading only revenue understates how much actual inference they serve. Both charts are true and they describe different things.
Why Usage Diverges From Benchmarks
OpenRouter traffic is dominated by high-volume, cost-sensitive, largely automated workloads: coding agents, bulk classification, synthetic data generation, and background pipelines. In that setting a model that is 90 percent as good at 5 percent of the price wins nearly every routing decision. Benchmark leadership determines what gets used for the hardest 10 percent of tasks. Price determines what gets used for the other 90 percent.
Two caveats on reading this data. OpenRouter is a router, so it over-represents developers who deliberately multi-source and under-represents teams on a direct enterprise contract with one vendor. And "identified" token volume excludes traffic the platform cannot attribute.
Brand Visibility Implications
This is the most direct available evidence that brand-visibility monitoring focused only on ChatGPT, Claude, and Gemini is measuring a minority of actual inference. Roughly 46 percent of routed token volume runs through models most brand teams have never tested a prompt against, and those models have different training data, different retrieval behaviour, and different brand recall. See multi-model orchestration and brand visibility and the open-weight recall gap.
Methodology
Figures compiled from OpenRouter's published rankings and third-party analyses of them through mid-2026. OpenRouter publishes token volume by model and by provider; percentages here are of identified volume. Monthly leaders change frequently. Presenc AI is not affiliated with OpenRouter.
How Presenc AI Helps
Presenc AI tracks brand representation across open-weight models as well as the major consumer assistants, weighted toward where inference actually happens.