The GeForce RTX 5090 is the fastest consumer graphics card for local LLMs on any model that fits in its 32GB of GDDR7 memory. In Hardware Corner's llama.cpp tests it generated about 200 tokens per second on an 8B model, 124 on a 14B model, and 61 on a dense 32B model at 4-bit quantisation with a 4K context. A dense 70B model at 4-bit is about 43GB and does not fit on one card. Two cards run it at about 27 tokens per second. The list price is $1,999, and US shop prices in September 2026 started near $4,300.
Tokens per Second by Model
Hardware Corner, llama.cpp, Q4_K quantisation unless noted, single stream, figures updated March 2026.
| Model | Prompt processing at 4K (tokens/sec) | Generation at 4K | Generation at 32K | Generation at 128K |
|---|---|---|---|---|
| Qwen3 8B | 11,933 | 200.4 | 129.8 | 58.8 |
| Qwen3 14B | 6,498 | 123.8 | 82.4 | 37.2 |
| gpt-oss 20B, sparse | 9,444 | 298.2 | 215.1 | 112.0 |
| Qwen3.5 27B | 3,004 | 58.8 | 53.8 | 44.1 |
| Qwen3 30B-A3B, sparse | 7,093 | 226.1 | 143.1 | 76.8 |
| Gemma4 31B | 3,395 | 61.1 | 55.4 | 43.4 |
| Qwen3 32B | 2,931 | 61.4 | 43.8 | Does not fit |
| Qwen3.5 35B, MXFP4, sparse | 6,605 | 165.2 | 143.2 | 118.2 |
Two patterns stand out. Sparse models run two to four times faster than dense models of similar size. Generation slows as the context fills: the 8B model drops from 200 to 59 tokens per second between 4K and 128K.
Other Testers
| Test | RTX 5090 result | Measured by |
|---|---|---|
| Llama 2 7B Q4_0, llama.cpp scoreboard | 290 to 300 generation, 14,073 to 14,970 prompt | llama.cpp community |
| gpt-oss 20B MXFP4, Ollama | 205 generation, 8,519 prompt | LMSYS |
| Llama 3.1 8B, 4-bit, Ollama | 150.0 generation | DatabaseMart |
| Qwen2.5 14B, 4-bit, Ollama | 89.9 generation | DatabaseMart |
| Gemma 3 27B, 4-bit, Ollama | 47.3 generation | DatabaseMart |
| QwQ 32B, 4-bit, Ollama | 57.2 generation | DatabaseMart |
| Qwen3.8-27B Q4_K_M, llama.cpp | 59 generation, 3,031 prompt | MacStories |
On the same llama.cpp scoreboard the RTX 4090 reaches 186 to 189 tokens per second, the RTX 3090 158 to 162, and the RTX 5080 182 to 185. Ollama results run lower than llama.cpp results on the same class of model, so the runtime explains part of the spread between sources. The cross-device view is in the tokens-per-second benchmark table.
Where 32GB Runs Out
| Model at 4-bit | Weights | Fit on one RTX 5090 |
|---|---|---|
| Qwen3 8B | 4.78 GiB | Yes, with 131K context in 23GB |
| Qwen3 14B | 8.53 GiB | Yes, with 131K context in 31GB |
| Qwen3 30B-A3B | 16.47 GiB | Yes, with 147K context in 31GB |
| Qwen3 32B | 18.64 GiB | Yes, but context tops out near 45K |
| Llama 3.3 70B | 43 GB | No |
Footprints and context limits are from Hardware Corner's November 2025 test, and the 70B size is from DatabaseMart. A dense 32B model is the practical ceiling, and long context has to be traded against model size. Once a model exceeds 32GB, layers move to system memory and speed collapses. We found no reproducible independent measurement of a single RTX 5090 running a 70B model with offload, so this page gives no number for it. The mechanism is explained in unified memory vs VRAM.
Two Cards
DatabaseMart tested two RTX 5090s under Ollama. Llama 3.3 70B ran at 26.9 tokens per second and DeepSeek-R1 70B at 27.0, against 24.3 for a single H100 in the same test series. Qwen 110B, at 63GB, used over 90 percent of both cards and ran at 7.2 tokens per second. Two cards give 64GB, which makes dense 70B models usable and little more. Each card draws up to 575 watts, and NVIDIA recommends a 1,000 watt supply for a single-card system.
RTX 5090 vs M5 Max, M5 Ultra, and RTX Spark
- M5 Max vs 5090. On the llama.cpp scoreboards with Llama 2 7B Q4_0, the RTX 5090 generates 290 to 300 tokens per second and the M5 Max 119.9. Prompt processing is about 14,000 against 3,220. A Context Studios roundup lists Qwen3.8-27B at 75 tokens per second on the 5090 and 65 to 69 on a 128GB M5 Max using MLX with multi-token prediction. The M5 Max holds up to 128GB.
- M5 Ultra vs 5090. MacStories measured 59 against 48 tokens per second on Qwen3.8-27B in favour of the 5090. See the M5 Ultra page.
- RTX Spark vs 5090. RTX Spark has not been benchmarked. It offers up to 128GB of memory at far lower bandwidth, so it should be slower on small models and able to load larger ones. See the RTX Spark page.
- DGX Spark vs 5090. LMSYS measured 205 against 49.7 tokens per second on gpt-oss 20B. See the three-way throughput comparison.
Price and Caveats
NVIDIA's list price is $1,999 and has not changed. Tech Insider's September 16, 2026 roundup of price trackers put Amazon at $4,329 and Micro Center at $4,299 to $5,300, with marketplace listings from $6,500 to $9,500 according to Tom's Hardware. The cause is GDDR7 supply, covered in the GDDR memory crisis report.
All figures here are single stream. Runtime, driver version, quantisation variant, and context length each move the result, which is why the same card shows 150 tokens per second in one 8B test and 200 in another.
Brand Visibility Implications
The RTX 5090 runs 8B to 32B models fast enough for everyday use. Models of that size hold less brand knowledge than frontier models, and a local setup answers from training data with no retrieval unless the user adds it. What a small model learned about a brand, including what it got wrong, is what the user hears. See the local LLM visibility blind spot.
Methodology
This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on consumer GPUs. That shows what a model with no retrieval says about you.