Research

Inference Provider Pricing Comparison

What it costs to serve open-weight models across inference providers in 2026. Pricing structures, the gap to first-party APIs, why the same model costs different amounts, and what to check beyond the headline rate.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: July 2026

The same open-weight model is served by a dozen providers at materially different prices. This page covers what drives the spread and what to check before choosing on the headline rate. First-party model API pricing is covered separately in the LLM API pricing comparison.

Why the Same Model Costs Different Amounts

FactorEffect on price
Quantisation servedA provider serving FP8 charges more than one serving 4-bit, for genuinely different quality
Batch policyAggressive batching cuts cost per token and raises latency
Context length supportedLong-context serving carries much higher cache cost
Throughput guaranteeDedicated capacity costs multiples of shared
HardwareSpecialised silicon prices at a premium for latency
Loss-leadingSome pricing is subsidised for adoption and will not persist

The quantisation point is the one that most often makes a cheap provider a false economy. Two providers listing the same model at different prices may not be serving the same thing, and few disclose the served precision prominently.

The Open-Weight Cost Advantage

The spread between cheap open-weight serving and premium closed APIs is roughly 17 to 20x. Models such as MiniMax M2.5 run around $0.30 per million input and $1.20 per million output; Claude Opus-class pricing reaches $5 to $25 per million. On several benchmarks the capability gap is far smaller than the price gap, which is the arbitrage driving the routing shift documented in OpenRouter usage rankings.

That arbitrage is why roughly 46 percent of routed token volume now goes to Chinese-origin open-weight models, up from under 2 percent a year earlier.

What To Check Beyond the Rate

Six things. Served quantisation and whether the provider states it. Whether the advertised context length is actually supported at the advertised price. Rate limits at your intended volume rather than at trial volume. Prompt caching support and pricing, which can dominate total cost for agentic workloads that resend context. Cold-start behaviour on less popular models, which can add seconds. And data handling terms, which vary far more across inference providers than across first-party APIs.

Brand Visibility Implications

Serving cost is the strongest predictor of which model answers a given question at volume. A brand-visibility programme measuring only premium platforms is measuring the tier that handles the smallest share of actual inference. Weight measurement by where volume goes, not by brand recognition. See the open-weight recall gap.

Methodology

Compiled from provider documentation, published benchmark suites, and independent measurement through July 2026. Inference performance figures depend heavily on model, quantisation, batch size, context length, and region, so ranges are given rather than point values and cross-provider comparison carries real uncertainty. Re-verify before quoting. Updated quarterly.

How Presenc AI Helps

Presenc AI measures brand representation across cheap open-weight serving as well as premium first-party APIs.

Frequently Asked Questions

Served quantisation, batch policy, supported context length, throughput guarantees, hardware, and in some cases subsidised pricing for adoption. Quantisation matters most: two providers listing the same model at different prices may not be serving the same thing, and few disclose served precision prominently.
Roughly 17 to 20 times at the extremes. Models like MiniMax M2.5 run around $0.30 per million input and $1.20 per million output against $5 to $25 per million for Claude Opus-class pricing. On several benchmarks the capability gap is much smaller than the price gap.
Served quantisation, whether the advertised context length is supported at the advertised price, rate limits at production rather than trial volume, prompt caching support and pricing, cold-start behaviour on less popular models, and data handling terms, which vary far more across inference providers than first-party APIs.
No. A lower price frequently reflects more aggressive quantisation, heavier batching that raises latency, or subsidised rates that will not persist. Compare served precision and measured latency at your own workload before treating two listings of the same model as equivalent.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.