Research

Brand Recall in Open-Weight vs Closed Models

Why the same brand question gets different answers from open-weight and closed models. Training data differences, the absence of retrieval in local deployments, and what it means for measurement coverage.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: July 2026

Brand-visibility monitoring is overwhelmingly aimed at ChatGPT, Claude, Gemini, and Perplexity. Roughly 46 percent of routed token volume now runs through models most brand teams have never sent a prompt to. Those models answer brand questions differently, and understanding why is more useful than the raw share number.

Four Structural Differences

FactorClosed frontier modelsOpen-weight models in local deployment
RetrievalUsually available, often defaultFrequently absent entirely
Answer basisParametric recall plus live sourcesParametric recall alone
Training corpus emphasisHeavily English, Western webOften substantial Chinese-language corpora
ObservabilityAPI calls a provider could logNone available to anyone outside
Update cadenceFrequent, provider-controlledWhenever the operator chooses to upgrade

The Retrieval Absence Is the Big One

A locally-served open-weight model answering "which observability tool should I use" has no web access. Every brand it names comes from weights. There is no citation to earn, no page to optimise, and no fetch to log. The entire retrieval-side toolkit that GEO practice has developed is inert in that setting.

This is not a marginal deployment pattern. Sparse architectures made frontier-adjacent models fast enough to run on a workstation, and coding agents running locally are exactly where developer-tool selection now happens. See the local-LLM blind spot and why sparse models made this practical.

Corpus Composition Changes Which Brands Surface

A model trained with heavy Chinese-language web representation has different density of coverage for Western B2B software brands than one trained predominantly on English sources, and the reverse holds for Chinese brands. Since roughly 46 percent of OpenRouter token volume goes to Chinese-origin models, this is now a first-order effect rather than an edge case. Brands with genuinely global coverage are less exposed; brands strong in one language market are more exposed than their home-market metrics suggest.

Stated licensing origin does not settle this either, since distillation and synthetic data mean a model's effective corpus can differ from its stated training set. See the distillation lineage tracker.

What To Actually Do

Three things. Add two or three high-volume open-weight models to your measurement set, prioritising whatever currently leads OpenRouter token volume rather than whatever leads benchmarks. Test unprompted recall specifically, with retrieval disabled, because that is the condition local deployments run in. And treat divergence between your closed-model and open-model results as the interesting signal: a brand strong on ChatGPT and absent from DeepSeek has a corpus problem, not a content problem.

Methodology

Structural comparison based on model documentation, deployment patterns, and published usage data through July 2026. The 46 percent figure is Chinese-origin share of identified OpenRouter token volume and describes that platform's routed traffic, not all LLM usage. Presenc AI has not published a controlled cross-model recall study; this page describes mechanisms and measurement approach rather than asserting measured recall gaps.

How Presenc AI Helps

Presenc AI measures brand representation across open-weight models with retrieval disabled as well as the major consumer assistants, so the parametric-only condition is covered rather than assumed.

Frequently Asked Questions

Often yes, for four structural reasons: retrieval is frequently absent in local deployments so answers come from weights alone, training corpus composition differs and many high-volume open models have heavy Chinese-language representation, update cadence is controlled by the operator, and nothing about the interaction is observable externally.
Because it makes the entire retrieval-side GEO toolkit inert. A locally-served model with no web access names brands purely from parametric recall. There is no citation to earn, no page to optimise, and no fetch to log, so visibility depends entirely on coverage that existed before training.
Chinese-origin models alone account for roughly 46 percent of identified token volume on OpenRouter, up from below 2 percent a year earlier. That figure describes routed traffic on one platform rather than all LLM usage, but it is the clearest public evidence available on the cost-sensitive developer segment.
Prioritise by usage rather than benchmark rank: whichever models currently lead OpenRouter token volume. Test unprompted recall with retrieval disabled, since that is the condition local deployments actually run in. Divergence from your closed-model results indicates a corpus problem rather than a content problem.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.