Brand-visibility monitoring is overwhelmingly aimed at ChatGPT, Claude, Gemini, and Perplexity. Roughly 46 percent of routed token volume now runs through models most brand teams have never sent a prompt to. Those models answer brand questions differently, and understanding why is more useful than the raw share number.
Four Structural Differences
| Factor | Closed frontier models | Open-weight models in local deployment |
|---|---|---|
| Retrieval | Usually available, often default | Frequently absent entirely |
| Answer basis | Parametric recall plus live sources | Parametric recall alone |
| Training corpus emphasis | Heavily English, Western web | Often substantial Chinese-language corpora |
| Observability | API calls a provider could log | None available to anyone outside |
| Update cadence | Frequent, provider-controlled | Whenever the operator chooses to upgrade |
The Retrieval Absence Is the Big One
A locally-served open-weight model answering "which observability tool should I use" has no web access. Every brand it names comes from weights. There is no citation to earn, no page to optimise, and no fetch to log. The entire retrieval-side toolkit that GEO practice has developed is inert in that setting.
This is not a marginal deployment pattern. Sparse architectures made frontier-adjacent models fast enough to run on a workstation, and coding agents running locally are exactly where developer-tool selection now happens. See the local-LLM blind spot and why sparse models made this practical.
Corpus Composition Changes Which Brands Surface
A model trained with heavy Chinese-language web representation has different density of coverage for Western B2B software brands than one trained predominantly on English sources, and the reverse holds for Chinese brands. Since roughly 46 percent of OpenRouter token volume goes to Chinese-origin models, this is now a first-order effect rather than an edge case. Brands with genuinely global coverage are less exposed; brands strong in one language market are more exposed than their home-market metrics suggest.
Stated licensing origin does not settle this either, since distillation and synthetic data mean a model's effective corpus can differ from its stated training set. See the distillation lineage tracker.
What To Actually Do
Three things. Add two or three high-volume open-weight models to your measurement set, prioritising whatever currently leads OpenRouter token volume rather than whatever leads benchmarks. Test unprompted recall specifically, with retrieval disabled, because that is the condition local deployments run in. And treat divergence between your closed-model and open-model results as the interesting signal: a brand strong on ChatGPT and absent from DeepSeek has a corpus problem, not a content problem.
Methodology
Structural comparison based on model documentation, deployment patterns, and published usage data through July 2026. The 46 percent figure is Chinese-origin share of identified OpenRouter token volume and describes that platform's routed traffic, not all LLM usage. Presenc AI has not published a controlled cross-model recall study; this page describes mechanisms and measurement approach rather than asserting measured recall gaps.
How Presenc AI Helps
Presenc AI measures brand representation across open-weight models with retrieval disabled as well as the major consumer assistants, so the parametric-only condition is covered rather than assumed.