Research

RTX 5090 Local LLM Benchmarks

RTX 5090 tokens per second by model: about 200 on 8B and 61 on dense 32B at Q4. Where the 32GB VRAM limit stops, dual-card 70B results, and September 2026 street prices far above the $1,999 list.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

The GeForce RTX 5090 is the fastest consumer graphics card for local LLMs on any model that fits in its 32GB of GDDR7 memory. In Hardware Corner's llama.cpp tests it generated about 200 tokens per second on an 8B model, 124 on a 14B model, and 61 on a dense 32B model at 4-bit quantisation with a 4K context. A dense 70B model at 4-bit is about 43GB and does not fit on one card. Two cards run it at about 27 tokens per second. The list price is $1,999, and US shop prices in September 2026 started near $4,300.

Tokens per Second by Model

Hardware Corner, llama.cpp, Q4_K quantisation unless noted, single stream, figures updated March 2026.

ModelPrompt processing at 4K (tokens/sec)Generation at 4KGeneration at 32KGeneration at 128K
Qwen3 8B11,933200.4129.858.8
Qwen3 14B6,498123.882.437.2
gpt-oss 20B, sparse9,444298.2215.1112.0
Qwen3.5 27B3,00458.853.844.1
Qwen3 30B-A3B, sparse7,093226.1143.176.8
Gemma4 31B3,39561.155.443.4
Qwen3 32B2,93161.443.8Does not fit
Qwen3.5 35B, MXFP4, sparse6,605165.2143.2118.2

Two patterns stand out. Sparse models run two to four times faster than dense models of similar size. Generation slows as the context fills: the 8B model drops from 200 to 59 tokens per second between 4K and 128K.

Other Testers

TestRTX 5090 resultMeasured by
Llama 2 7B Q4_0, llama.cpp scoreboard290 to 300 generation, 14,073 to 14,970 promptllama.cpp community
gpt-oss 20B MXFP4, Ollama205 generation, 8,519 promptLMSYS
Llama 3.1 8B, 4-bit, Ollama150.0 generationDatabaseMart
Qwen2.5 14B, 4-bit, Ollama89.9 generationDatabaseMart
Gemma 3 27B, 4-bit, Ollama47.3 generationDatabaseMart
QwQ 32B, 4-bit, Ollama57.2 generationDatabaseMart
Qwen3.8-27B Q4_K_M, llama.cpp59 generation, 3,031 promptMacStories

On the same llama.cpp scoreboard the RTX 4090 reaches 186 to 189 tokens per second, the RTX 3090 158 to 162, and the RTX 5080 182 to 185. Ollama results run lower than llama.cpp results on the same class of model, so the runtime explains part of the spread between sources. The cross-device view is in the tokens-per-second benchmark table.

Where 32GB Runs Out

Model at 4-bitWeightsFit on one RTX 5090
Qwen3 8B4.78 GiBYes, with 131K context in 23GB
Qwen3 14B8.53 GiBYes, with 131K context in 31GB
Qwen3 30B-A3B16.47 GiBYes, with 147K context in 31GB
Qwen3 32B18.64 GiBYes, but context tops out near 45K
Llama 3.3 70B43 GBNo

Footprints and context limits are from Hardware Corner's November 2025 test, and the 70B size is from DatabaseMart. A dense 32B model is the practical ceiling, and long context has to be traded against model size. Once a model exceeds 32GB, layers move to system memory and speed collapses. We found no reproducible independent measurement of a single RTX 5090 running a 70B model with offload, so this page gives no number for it. The mechanism is explained in unified memory vs VRAM.

Two Cards

DatabaseMart tested two RTX 5090s under Ollama. Llama 3.3 70B ran at 26.9 tokens per second and DeepSeek-R1 70B at 27.0, against 24.3 for a single H100 in the same test series. Qwen 110B, at 63GB, used over 90 percent of both cards and ran at 7.2 tokens per second. Two cards give 64GB, which makes dense 70B models usable and little more. Each card draws up to 575 watts, and NVIDIA recommends a 1,000 watt supply for a single-card system.

RTX 5090 vs M5 Max, M5 Ultra, and RTX Spark

  • M5 Max vs 5090. On the llama.cpp scoreboards with Llama 2 7B Q4_0, the RTX 5090 generates 290 to 300 tokens per second and the M5 Max 119.9. Prompt processing is about 14,000 against 3,220. A Context Studios roundup lists Qwen3.8-27B at 75 tokens per second on the 5090 and 65 to 69 on a 128GB M5 Max using MLX with multi-token prediction. The M5 Max holds up to 128GB.
  • M5 Ultra vs 5090. MacStories measured 59 against 48 tokens per second on Qwen3.8-27B in favour of the 5090. See the M5 Ultra page.
  • RTX Spark vs 5090. RTX Spark has not been benchmarked. It offers up to 128GB of memory at far lower bandwidth, so it should be slower on small models and able to load larger ones. See the RTX Spark page.
  • DGX Spark vs 5090. LMSYS measured 205 against 49.7 tokens per second on gpt-oss 20B. See the three-way throughput comparison.

Price and Caveats

NVIDIA's list price is $1,999 and has not changed. Tech Insider's September 16, 2026 roundup of price trackers put Amazon at $4,329 and Micro Center at $4,299 to $5,300, with marketplace listings from $6,500 to $9,500 according to Tom's Hardware. The cause is GDDR7 supply, covered in the GDDR memory crisis report.

All figures here are single stream. Runtime, driver version, quantisation variant, and context length each move the result, which is why the same card shows 150 tokens per second in one 8B test and 200 in another.

Brand Visibility Implications

The RTX 5090 runs 8B to 32B models fast enough for everyday use. Models of that size hold less brand knowledge than frontier models, and a local setup answers from training data with no retrieval unless the user adds it. What a small model learned about a brand, including what it got wrong, is what the user hears. See the local LLM visibility blind spot.

Methodology

This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on consumer GPUs. That shows what a model with no retrieval says about you.

Frequently Asked Questions

It depends on the model. In Hardware Corner's llama.cpp tests at Q4 with a 4K context, the RTX 5090 generated 200 tokens per second on Qwen3 8B, 124 on Qwen3 14B, 61 on Qwen3 32B, and 298 on the sparse gpt-oss 20B. DatabaseMart measured 150 tokens per second on Llama 3.1 8B under Ollama. Speed falls as the context fills.
Not on one card at 4-bit. A 70B model at 4-bit is about 43GB and the card has 32GB. Two RTX 5090s ran Llama 3.3 70B at 26.9 tokens per second in DatabaseMart's Ollama test. On a single card the practical ceiling is a dense 32B model or a sparse model of similar file size.
The RTX 5090 is faster on models that fit in 32GB. On the llama.cpp scoreboards with Llama 2 7B Q4_0 it generates 290 to 300 tokens per second against 119.9 for the M5 Max, and processes prompts about four times faster. The M5 Max offers up to 128GB of memory, so it can run models the 5090 cannot load.
RTX Spark has not been independently benchmarked as of October 1, 2026. On specifications it has up to 128GB of unified memory and a reported 300 GB/s of bandwidth, against 32GB and about 1,790 GB/s for the RTX 5090. The 5090 should be faster on models up to about 32B, and RTX Spark should load larger ones.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.