Strix Halo is AMD's Ryzen AI Max line, led by the Ryzen AI Max+ 395: 16 Zen 5 cores, a Radeon 8060S integrated GPU, and up to 128GB of unified LPDDR5X memory. It is the x86 option for running large models on one small machine. Independent tests put it at about 46 to 53 tokens per second on a 7B model, 72 to 75 on a sparse 30B model, 31 to 34 on gpt-oss 120B, and about 5 on a dense 70B model. A 128GB mini PC cost between $3,449 and $3,999 in September 2026.
Machines and Prices
| System | Memory | Price, September 2026 |
|---|---|---|
| Framework Desktop, Ryzen AI Max+ 395 | 128GB | $3,449 |
| GMKtec EVO-X2 | 128GB | $3,649 |
| Minisforum MS-S1 Max | 128GB | $3,799 |
| AMD Ryzen AI Halo developer box | 128GB | $3,999 |
| Beelink GTR9 Pro | 128GB | $4,349 |
| HP Z2 Mini G1a | 128GB | $5,349 to $7,406 |
| GMKtec Evo-X5, Ryzen AI Max+ PRO 495 | 192GB | $7,099, or $6,674 early price |
Prices for 128GB systems are from DataHardware's September 27, 2026 listing check and StorageReview, and the Evo-X5 price is from Tom's Hardware. Laptops use the same chips: UltrabookReview lists the Asus ROG Flow Z13 from $2,099 and the HP ZBook Ultra 14 from $1,999 in base configurations. These prices are well above 2025 levels because of memory costs. See the memory shortage report. The older hardware landscape page lists lower Strix Halo prices from before the rise.
Measured Tokens per Second
| Model and format | Prompt processing (tokens/sec) | Generation (tokens/sec) | Measured by |
|---|---|---|---|
| Llama 2 7B, Q4_0, Vulkan | 998 | 45.8 | Level1Techs forum, Framework Desktop |
| Llama 2 7B, 4-bit, Vulkan with flash attention | Not listed | 52.7 | llm-tracker, via DataHardware |
| Qwen3 30B-A3B, sparse, Vulkan | 605 | 72.0 | Level1Techs forum |
| Qwen3 30B-A3B, sparse, Vulkan | Not listed | 75.3 | llm-tracker, via DataHardware |
| gpt-oss 120B, MXFP4, llama.cpp | 340 | 34.1 | Hardware Corner compilation |
| gpt-oss 120B, LM Studio | Not listed | 31.4 | ServeTheHome, via DataHardware |
| Dense 70B, 4-bit | Not listed | About 5 | Level1Techs and ServeTheHome |
| Phi, Mistral, Llama 3 in UL Procyon | Not listed | 55.3, 35.0, 34.6 | StorageReview, Ryzen AI Halo |
Sparse models are the reason to buy this hardware. A dense 70B model loads but runs at about 5 tokens per second. HotHardware's Ryzen AI Halo review found the same, with all systems it tested at 6 to 7 tokens per second on Llama 3.1 70B. A sparse 120B model on the same box runs six times faster. See active versus total parameters.
Memory bandwidth is the ceiling. The 256-bit LPDDR5X interface gives 256 GB/s on paper, and DataHardware reports about 215 GB/s in practice. That is close to DGX Spark at 273 GB/s and far below an RTX 5090.
Strix Halo vs DGX Spark
In Hardware Corner's October 2025 compilation the two were close on single-stream generation and far apart on prompt processing. The llama.cpp scoreboard for DGX Spark now shows 60.6 tokens per second on gpt-oss 120B with an empty context, so the generation gap may be wider than the first row suggests.
| Test | Strix Halo | DGX Spark | Source |
|---|---|---|---|
| gpt-oss 120B generation, single stream (tokens/sec) | 34.1 | 38.6 | Hardware Corner, October 2025 |
| gpt-oss 120B prompt processing (tokens/sec) | 340 | 1,723 | Hardware Corner, October 2025 |
| gpt-oss 20B, vLLM, batch 64, 256 in and 256 out (tokens/sec) | 617 | 1,917 | StorageReview |
| gpt-oss 20B, vLLM, 8K in and 1K out (tokens/sec) | 881 | 3,672 | StorageReview |
| Qwen3 Coder 30B, vLLM, batch 64 (tokens/sec) | 376 | 729 | StorageReview |
| 128GB system price | $3,449 to $3,999 | $4,699 | Vendor and press listings |
StorageReview summarised the gap as two to four times in most vLLM serving tests. AMD's own comparison claims an advantage of 4 to 14 percent over DGX Spark across four models when price is included. That is a vendor claim. Strix Halo runs both Windows and Linux on a standard x86 PC, which DGX Spark does not. See RTX Spark vs DGX Spark for the NVIDIA side.
The 192GB Successor
The Ryzen AI Max+ PRO 495, known as Gorgon Halo, raises the memory ceiling to 192GB of LPDDR5X-8533 with up to 160GB available to the GPU. HotHardware lists it with the same 16 Zen 5 cores and a Radeon 8065S GPU. UltrabookReview describes the Gorgon Halo line as largely a rebadge of Strix Halo with small clock changes, so the gain is capacity, not speed. Framework now lists 32GB, 64GB, 128GB, and 192GB options. Minisforum says two linked 192GB units run a 235B model at 16 tokens per second, which is a vendor claim.
Caveats
- Backend matters. Results differ between Vulkan and ROCm builds of llama.cpp, and the Level1Techs tester recorded close to a 50 percent gain on one prompt-processing test between May and August 2025 from software updates alone.
- Tests are not uniform. The table mixes machines, operating systems, and runtimes.
- No independent PRO 495 benchmarks. The 192GB figures in circulation come from vendors.
- Power. ServeTheHome measured a Beelink GTR9 Pro at 125 to 128 watts during gpt-oss 120B generation, as cited by DataHardware.
Brand Visibility Implications
Strix Halo is the cheapest new way to hold a 120B-class model in memory on a Windows or Linux PC. A model running there answers from its training data, with no retrieval unless the user builds it. What the model learned about a brand is what the user gets, and nothing about the exchange is visible from outside. See the local LLM visibility blind spot.
Methodology
This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on machines like these. That shows what a model with no retrieval says about you.