The same open-weight model is served by a dozen providers at materially different prices. This page covers what drives the spread and what to check before choosing on the headline rate. First-party model API pricing is covered separately in the LLM API pricing comparison.
Why the Same Model Costs Different Amounts
| Factor | Effect on price |
|---|---|
| Quantisation served | A provider serving FP8 charges more than one serving 4-bit, for genuinely different quality |
| Batch policy | Aggressive batching cuts cost per token and raises latency |
| Context length supported | Long-context serving carries much higher cache cost |
| Throughput guarantee | Dedicated capacity costs multiples of shared |
| Hardware | Specialised silicon prices at a premium for latency |
| Loss-leading | Some pricing is subsidised for adoption and will not persist |
The quantisation point is the one that most often makes a cheap provider a false economy. Two providers listing the same model at different prices may not be serving the same thing, and few disclose the served precision prominently.
The Open-Weight Cost Advantage
The spread between cheap open-weight serving and premium closed APIs is roughly 17 to 20x. Models such as MiniMax M2.5 run around $0.30 per million input and $1.20 per million output; Claude Opus-class pricing reaches $5 to $25 per million. On several benchmarks the capability gap is far smaller than the price gap, which is the arbitrage driving the routing shift documented in OpenRouter usage rankings.
That arbitrage is why roughly 46 percent of routed token volume now goes to Chinese-origin open-weight models, up from under 2 percent a year earlier.
What To Check Beyond the Rate
Six things. Served quantisation and whether the provider states it. Whether the advertised context length is actually supported at the advertised price. Rate limits at your intended volume rather than at trial volume. Prompt caching support and pricing, which can dominate total cost for agentic workloads that resend context. Cold-start behaviour on less popular models, which can add seconds. And data handling terms, which vary far more across inference providers than across first-party APIs.
Brand Visibility Implications
Serving cost is the strongest predictor of which model answers a given question at volume. A brand-visibility programme measuring only premium platforms is measuring the tier that handles the smallest share of actual inference. Weight measurement by where volume goes, not by brand recognition. See the open-weight recall gap.
Methodology
Compiled from provider documentation, published benchmark suites, and independent measurement through July 2026. Inference performance figures depend heavily on model, quantisation, batch size, context length, and region, so ranges are given rather than point values and cross-provider comparison carries real uncertainty. Re-verify before quoting. Updated quarterly.
How Presenc AI Helps
Presenc AI measures brand representation across cheap open-weight serving as well as premium first-party APIs.