Research

AMD Strix Halo Local LLM Benchmarks

AMD Strix Halo (Ryzen AI Max+ 395) for local LLMs: measured tokens per second by model, 128GB mini PC prices from $3,449, the 192GB PRO 495 systems, and how it compares with DGX Spark.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

Strix Halo is AMD's Ryzen AI Max line, led by the Ryzen AI Max+ 395: 16 Zen 5 cores, a Radeon 8060S integrated GPU, and up to 128GB of unified LPDDR5X memory. It is the x86 option for running large models on one small machine. Independent tests put it at about 46 to 53 tokens per second on a 7B model, 72 to 75 on a sparse 30B model, 31 to 34 on gpt-oss 120B, and about 5 on a dense 70B model. A 128GB mini PC cost between $3,449 and $3,999 in September 2026.

Machines and Prices

SystemMemoryPrice, September 2026
Framework Desktop, Ryzen AI Max+ 395128GB$3,449
GMKtec EVO-X2128GB$3,649
Minisforum MS-S1 Max128GB$3,799
AMD Ryzen AI Halo developer box128GB$3,999
Beelink GTR9 Pro128GB$4,349
HP Z2 Mini G1a128GB$5,349 to $7,406
GMKtec Evo-X5, Ryzen AI Max+ PRO 495192GB$7,099, or $6,674 early price

Prices for 128GB systems are from DataHardware's September 27, 2026 listing check and StorageReview, and the Evo-X5 price is from Tom's Hardware. Laptops use the same chips: UltrabookReview lists the Asus ROG Flow Z13 from $2,099 and the HP ZBook Ultra 14 from $1,999 in base configurations. These prices are well above 2025 levels because of memory costs. See the memory shortage report. The older hardware landscape page lists lower Strix Halo prices from before the rise.

Measured Tokens per Second

Model and formatPrompt processing (tokens/sec)Generation (tokens/sec)Measured by
Llama 2 7B, Q4_0, Vulkan99845.8Level1Techs forum, Framework Desktop
Llama 2 7B, 4-bit, Vulkan with flash attentionNot listed52.7llm-tracker, via DataHardware
Qwen3 30B-A3B, sparse, Vulkan60572.0Level1Techs forum
Qwen3 30B-A3B, sparse, VulkanNot listed75.3llm-tracker, via DataHardware
gpt-oss 120B, MXFP4, llama.cpp34034.1Hardware Corner compilation
gpt-oss 120B, LM StudioNot listed31.4ServeTheHome, via DataHardware
Dense 70B, 4-bitNot listedAbout 5Level1Techs and ServeTheHome
Phi, Mistral, Llama 3 in UL ProcyonNot listed55.3, 35.0, 34.6StorageReview, Ryzen AI Halo

Sparse models are the reason to buy this hardware. A dense 70B model loads but runs at about 5 tokens per second. HotHardware's Ryzen AI Halo review found the same, with all systems it tested at 6 to 7 tokens per second on Llama 3.1 70B. A sparse 120B model on the same box runs six times faster. See active versus total parameters.

Memory bandwidth is the ceiling. The 256-bit LPDDR5X interface gives 256 GB/s on paper, and DataHardware reports about 215 GB/s in practice. That is close to DGX Spark at 273 GB/s and far below an RTX 5090.

Strix Halo vs DGX Spark

In Hardware Corner's October 2025 compilation the two were close on single-stream generation and far apart on prompt processing. The llama.cpp scoreboard for DGX Spark now shows 60.6 tokens per second on gpt-oss 120B with an empty context, so the generation gap may be wider than the first row suggests.

TestStrix HaloDGX SparkSource
gpt-oss 120B generation, single stream (tokens/sec)34.138.6Hardware Corner, October 2025
gpt-oss 120B prompt processing (tokens/sec)3401,723Hardware Corner, October 2025
gpt-oss 20B, vLLM, batch 64, 256 in and 256 out (tokens/sec)6171,917StorageReview
gpt-oss 20B, vLLM, 8K in and 1K out (tokens/sec)8813,672StorageReview
Qwen3 Coder 30B, vLLM, batch 64 (tokens/sec)376729StorageReview
128GB system price$3,449 to $3,999$4,699Vendor and press listings

StorageReview summarised the gap as two to four times in most vLLM serving tests. AMD's own comparison claims an advantage of 4 to 14 percent over DGX Spark across four models when price is included. That is a vendor claim. Strix Halo runs both Windows and Linux on a standard x86 PC, which DGX Spark does not. See RTX Spark vs DGX Spark for the NVIDIA side.

The 192GB Successor

The Ryzen AI Max+ PRO 495, known as Gorgon Halo, raises the memory ceiling to 192GB of LPDDR5X-8533 with up to 160GB available to the GPU. HotHardware lists it with the same 16 Zen 5 cores and a Radeon 8065S GPU. UltrabookReview describes the Gorgon Halo line as largely a rebadge of Strix Halo with small clock changes, so the gain is capacity, not speed. Framework now lists 32GB, 64GB, 128GB, and 192GB options. Minisforum says two linked 192GB units run a 235B model at 16 tokens per second, which is a vendor claim.

Caveats

  • Backend matters. Results differ between Vulkan and ROCm builds of llama.cpp, and the Level1Techs tester recorded close to a 50 percent gain on one prompt-processing test between May and August 2025 from software updates alone.
  • Tests are not uniform. The table mixes machines, operating systems, and runtimes.
  • No independent PRO 495 benchmarks. The 192GB figures in circulation come from vendors.
  • Power. ServeTheHome measured a Beelink GTR9 Pro at 125 to 128 watts during gpt-oss 120B generation, as cited by DataHardware.

Brand Visibility Implications

Strix Halo is the cheapest new way to hold a 120B-class model in memory on a Windows or Linux PC. A model running there answers from its training data, with no retrieval unless the user builds it. What the model learned about a brand is what the user gets, and nothing about the exchange is visible from outside. See the local LLM visibility blind spot.

Methodology

This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on machines like these. That shows what a model with no retrieval says about you.

Frequently Asked Questions

On a Ryzen AI Max+ 395 with 128GB, independent testers report about 46 to 53 tokens per second on Llama 2 7B at 4-bit, 72 to 75 on the sparse Qwen3 30B-A3B, 31 to 34 on gpt-oss 120B, and about 5 on dense 70B models. Results vary with the llama.cpp backend and software version.
In September 2026, listed prices were $3,449 for the Framework Desktop, $3,649 for the GMKtec EVO-X2, $3,799 for the Minisforum MS-S1 Max, and $3,999 for AMD's own Ryzen AI Halo. The 192GB GMKtec Evo-X5 with the Ryzen AI Max+ PRO 495 is listed at $7,099.
No. In Hardware Corner's October 2025 compilation Strix Halo generated 34.1 tokens per second on gpt-oss 120B against 38.6 for DGX Spark, and DGX Spark was about five times faster at prompt processing. StorageReview found DGX Spark two to four times faster in most batched vLLM serving tests. Strix Halo costs less and runs standard x86 Windows and Linux.
It can load one, but dense 70B models run at about 5 to 7 tokens per second in tests by Level1Techs, ServeTheHome, and HotHardware, which is too slow for interactive chat. Sparse models are a better fit: gpt-oss 120B runs at 31 to 34 tokens per second on the same hardware.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.