RTX Spark is NVIDIA's Arm-based chip for Windows PCs. It pairs a Grace CPU with a Blackwell RTX GPU and up to 128GB of unified memory, and NVIDIA says it can run 120-billion-parameter language models locally. NVIDIA announced it at Computex 2026 in a press release dated May 31, 2026, and told press at IFA that partners begin shipping in October 2026. As of October 1, 2026, no independent tokens-per-second benchmark of an RTX Spark system has been published, and neither NVIDIA nor its partners have announced a price. This page separates what is confirmed from what is estimated.
Confirmed Specifications
NVIDIA confirmed two configurations of the chip, both sold under the N1X name, according to Wccftech's report from IFA on September 3, 2026.
| Item | Top configuration | Second configuration |
|---|---|---|
| CPU | 20-core NVIDIA Grace | 18-core NVIDIA Grace |
| GPU | Blackwell RTX, 6,144 CUDA cores | Blackwell RTX, 5,120 CUDA cores |
| Unified memory | 24GB to 128GB | 24GB to 32GB |
| Form factors | Laptops and compact desktops | Laptops first |
| AI compute (NVIDIA claim) | Up to 1 petaflop | Not stated |
| Operating system | Windows 11 on Arm | Windows 11 on Arm |
| First shipments | October 2026 | October 2026 |
NVIDIA says RTX Spark laptops come in 14 to 16 inch sizes, as slim as 14 millimeters and as light as 3 pounds. Launch partners are ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI, with Acer and GIGABYTE to follow. Microsoft also announced the Surface RTX Spark Dev Box, a compact desktop with 128GB of unified memory that will be sold in the US on Microsoft.com.
Two numbers that matter for language models are not in NVIDIA's announcement. Memory bandwidth is reported by press at about 300 GB/s, and a power range of 45 to 80 watts comes from leaked documents published by VideoCardz. Treat both as unconfirmed.
Price
| Figure | Value | Status |
|---|---|---|
| Official RTX Spark price | None | Not announced by NVIDIA or any partner |
| High-end RTX Spark system | About $2,899 | Analyst estimate repeated in press coverage |
| DGX Spark, the nearest shipping product | $4,699 | NVIDIA list price, raised from $3,999 in February 2026 |
The $2,899 figure is an estimate, not a quote. Memory prices rose through 2026, and the DGX Spark price rise is a reminder that 128GB of memory is the expensive part of these machines. See the memory shortage report.
What NVIDIA Claims for Local LLMs
- Model size. NVIDIA says RTX Spark runs 120-billion-parameter language models with a context of 1 million tokens. Microsoft's Dev Box post says "at interactive speeds" and gives no number.
- Software speedups. NVIDIA says new optimizations deliver 2x performance in llama.cpp and 2.6x in vLLM on top agentic models. The baseline for those multiples is not stated.
- Tokens per second. NVIDIA has published no tokens-per-second figure for RTX Spark.
The Closest Measured Proxy: DGX Spark
DGX Spark is NVIDIA's own Linux mini desktop. Its GB10 chip has the same headline numbers as the top RTX Spark configuration: 20 Arm cores, 6,144 CUDA cores, and 128GB of unified memory. Press coverage describes the two as the same class of silicon. The figures below are DGX Spark measurements. They are a guide to what RTX Spark might do, and they are not RTX Spark results.
| Model on DGX Spark | Prompt processing (tokens/sec) | Generation (tokens/sec) | Measured by |
|---|---|---|---|
| gpt-oss 20B, MXFP4 | 2,009 | 60.9 | llama.cpp scoreboard, October 2025 |
| gpt-oss 120B, MXFP4 | 1,956 | 60.6, falling to 40.6 at 32K context | llama.cpp scoreboard, October 2025 |
| Llama 3.1 8B, Q4_K_M | 7,614 | 38.0 | Ollama, October 2025 |
| Qwen3 32B, Q4_K_M | 705 | 9.4 | Ollama, October 2025 |
| Llama 3.1 70B, Q4_K_M | 1,911 | 4.4 | Ollama, October 2025 |
The pattern is that sparse models such as gpt-oss run at usable speed while dense 32B and 70B models are slow. A thin laptop has less cooling than the DGX Spark box, and Windows on Arm is a different software stack, so RTX Spark results could land lower or higher.
RTX Spark vs RTX 5090, M5 Max, and Mac Studio
| RTX Spark, top configuration | RTX 5090 | Mac Studio M5 Max | Mac Studio M5 Ultra | |
|---|---|---|---|---|
| Memory a model can use | Up to 128GB | 32GB | Up to 128GB | 96GB to 512GB |
| Memory bandwidth | About 300 GB/s, reported | About 1,790 GB/s | 460 or 614 GB/s | 1.2 TB/s |
| Independent LLM benchmarks | None yet | Many | Some | First reviews |
| Price | Not announced | $1,999 list, far higher in shops | From $2,499 | From $5,499 |
Generation speed tracks memory bandwidth, so on paper an RTX Spark system should generate tokens more slowly than an RTX 5090 or an M5 Max on any model that fits in all three. Its advantage over the RTX 5090 is capacity: 128GB holds models that 32GB cannot. That is an expectation from the specifications, not a measurement. Device pages: RTX 5090 and Mac Studio M5 Ultra. For the naming question see RTX Spark vs DGX Spark.
What Is Still Unknown
- Tokens per second on any model, from NVIDIA or from an independent tester.
- Official memory bandwidth and sustained power draw.
- Prices for any configuration, including the 128GB tier.
- How much speed a 14 millimeter laptop holds under a long generation run.
- How llama.cpp, vLLM, and Ollama behave on Windows on Arm compared with DGX OS.
Brand Visibility Implications
A local model answers from its training data. It has no search step and no live retrieval unless the user adds one. If RTX Spark puts 128GB of memory into ordinary Windows laptops, more people will ask product questions of a model that only knows what it learned about a brand before its cutoff. Those answers leave no trace in any analytics tool. See the local LLM visibility blind spot.
Methodology
This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on local hardware. That shows what a model with no retrieval says about you.