Research

NVIDIA RTX Spark for Local LLMs: Specs, Price, and Benchmarks

NVIDIA RTX Spark puts up to 128GB of unified memory in Windows laptops and compact desktops from October 2026. Confirmed specs, price estimates, NVIDIA's claims, and what is still unmeasured.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

RTX Spark is NVIDIA's Arm-based chip for Windows PCs. It pairs a Grace CPU with a Blackwell RTX GPU and up to 128GB of unified memory, and NVIDIA says it can run 120-billion-parameter language models locally. NVIDIA announced it at Computex 2026 in a press release dated May 31, 2026, and told press at IFA that partners begin shipping in October 2026. As of October 1, 2026, no independent tokens-per-second benchmark of an RTX Spark system has been published, and neither NVIDIA nor its partners have announced a price. This page separates what is confirmed from what is estimated.

Confirmed Specifications

NVIDIA confirmed two configurations of the chip, both sold under the N1X name, according to Wccftech's report from IFA on September 3, 2026.

ItemTop configurationSecond configuration
CPU20-core NVIDIA Grace18-core NVIDIA Grace
GPUBlackwell RTX, 6,144 CUDA coresBlackwell RTX, 5,120 CUDA cores
Unified memory24GB to 128GB24GB to 32GB
Form factorsLaptops and compact desktopsLaptops first
AI compute (NVIDIA claim)Up to 1 petaflopNot stated
Operating systemWindows 11 on ArmWindows 11 on Arm
First shipmentsOctober 2026October 2026

NVIDIA says RTX Spark laptops come in 14 to 16 inch sizes, as slim as 14 millimeters and as light as 3 pounds. Launch partners are ASUS, Dell, HP, Lenovo, Microsoft Surface, and MSI, with Acer and GIGABYTE to follow. Microsoft also announced the Surface RTX Spark Dev Box, a compact desktop with 128GB of unified memory that will be sold in the US on Microsoft.com.

Two numbers that matter for language models are not in NVIDIA's announcement. Memory bandwidth is reported by press at about 300 GB/s, and a power range of 45 to 80 watts comes from leaked documents published by VideoCardz. Treat both as unconfirmed.

Price

FigureValueStatus
Official RTX Spark priceNoneNot announced by NVIDIA or any partner
High-end RTX Spark systemAbout $2,899Analyst estimate repeated in press coverage
DGX Spark, the nearest shipping product$4,699NVIDIA list price, raised from $3,999 in February 2026

The $2,899 figure is an estimate, not a quote. Memory prices rose through 2026, and the DGX Spark price rise is a reminder that 128GB of memory is the expensive part of these machines. See the memory shortage report.

What NVIDIA Claims for Local LLMs

  • Model size. NVIDIA says RTX Spark runs 120-billion-parameter language models with a context of 1 million tokens. Microsoft's Dev Box post says "at interactive speeds" and gives no number.
  • Software speedups. NVIDIA says new optimizations deliver 2x performance in llama.cpp and 2.6x in vLLM on top agentic models. The baseline for those multiples is not stated.
  • Tokens per second. NVIDIA has published no tokens-per-second figure for RTX Spark.

The Closest Measured Proxy: DGX Spark

DGX Spark is NVIDIA's own Linux mini desktop. Its GB10 chip has the same headline numbers as the top RTX Spark configuration: 20 Arm cores, 6,144 CUDA cores, and 128GB of unified memory. Press coverage describes the two as the same class of silicon. The figures below are DGX Spark measurements. They are a guide to what RTX Spark might do, and they are not RTX Spark results.

Model on DGX SparkPrompt processing (tokens/sec)Generation (tokens/sec)Measured by
gpt-oss 20B, MXFP42,00960.9llama.cpp scoreboard, October 2025
gpt-oss 120B, MXFP41,95660.6, falling to 40.6 at 32K contextllama.cpp scoreboard, October 2025
Llama 3.1 8B, Q4_K_M7,61438.0Ollama, October 2025
Qwen3 32B, Q4_K_M7059.4Ollama, October 2025
Llama 3.1 70B, Q4_K_M1,9114.4Ollama, October 2025

The pattern is that sparse models such as gpt-oss run at usable speed while dense 32B and 70B models are slow. A thin laptop has less cooling than the DGX Spark box, and Windows on Arm is a different software stack, so RTX Spark results could land lower or higher.

RTX Spark vs RTX 5090, M5 Max, and Mac Studio

RTX Spark, top configurationRTX 5090Mac Studio M5 MaxMac Studio M5 Ultra
Memory a model can useUp to 128GB32GBUp to 128GB96GB to 512GB
Memory bandwidthAbout 300 GB/s, reportedAbout 1,790 GB/s460 or 614 GB/s1.2 TB/s
Independent LLM benchmarksNone yetManySomeFirst reviews
PriceNot announced$1,999 list, far higher in shopsFrom $2,499From $5,499

Generation speed tracks memory bandwidth, so on paper an RTX Spark system should generate tokens more slowly than an RTX 5090 or an M5 Max on any model that fits in all three. Its advantage over the RTX 5090 is capacity: 128GB holds models that 32GB cannot. That is an expectation from the specifications, not a measurement. Device pages: RTX 5090 and Mac Studio M5 Ultra. For the naming question see RTX Spark vs DGX Spark.

What Is Still Unknown

  • Tokens per second on any model, from NVIDIA or from an independent tester.
  • Official memory bandwidth and sustained power draw.
  • Prices for any configuration, including the 128GB tier.
  • How much speed a 14 millimeter laptop holds under a long generation run.
  • How llama.cpp, vLLM, and Ollama behave on Windows on Arm compared with DGX OS.

Brand Visibility Implications

A local model answers from its training data. It has no search step and no live retrieval unless the user adds one. If RTX Spark puts 128GB of memory into ordinary Windows laptops, more people will ask product questions of a model that only knows what it learned about a brand before its cutoff. Those answers leave no trace in any analytics tool. See the local LLM visibility blind spot.

Methodology

This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on local hardware. That shows what a model with no retrieval says about you.

Frequently Asked Questions

No figure exists yet. As of October 1, 2026, NVIDIA has published no tokens-per-second number for RTX Spark and no independent benchmark has appeared. The nearest proxy is DGX Spark, which has the same headline chip specifications: the llama.cpp scoreboard shows about 61 tokens per second on gpt-oss 120B at MXFP4, and Ollama measured 4.4 tokens per second on Llama 3.1 70B at Q4_K_M.
No official price has been announced. Press coverage repeats an analyst estimate of about $2,899 for a high-end configuration. For reference, NVIDIA's own DGX Spark with 128GB lists at $4,699 after a rise from $3,999 in February 2026.
It has not been measured. On specifications, the RTX 5090 has about 1,790 GB/s of memory bandwidth against a reported 300 GB/s for RTX Spark, and generation speed follows bandwidth, so the 5090 should be faster on models that fit in its 32GB. RTX Spark offers up to 128GB, so it can hold models the 5090 cannot.
Both offer up to 128GB of unified memory. Apple lists 614 GB/s of bandwidth for the M5 Max with the 40-core GPU and 1.2 TB/s for the M5 Ultra, against a reported 300 GB/s for RTX Spark. RTX Spark runs Windows and CUDA software, and the Mac runs macOS and MLX. No head-to-head benchmark exists yet.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.