The Mac Studio with M5 Ultra went on sale on September 22, 2026, starting at $5,499 with 96GB of unified memory. Apple lists memory bandwidth of 1.2 TB/s and offers 256GB and 512GB options, with the 512GB model arriving in late October. In the first independent local LLM review, MacStories measured 48 tokens per second on a dense 27B model and between 52 and 143 tokens per second on a large sparse model, depending on the test. Prompt processing was more than twice as fast as on the M3 Ultra. An RTX 5090 was still about 23 percent faster on the 27B model that fits in its memory.
Configurations and Prices
| Configuration | Price | Source |
|---|---|---|
| M5 Ultra, 30-core CPU, 64-core GPU, 96GB, 1TB | $5,499 | Apple |
| M5 Ultra, 256GB | About $9,499 | Pinggy and MortalApps buyer guides |
| M5 Ultra, 36-core CPU, 80-core GPU, 256GB, 2TB | $11,299 | Engadget review unit |
| M5 Ultra, 512GB | Not announced, late October | MacRumors expects well above $10,000 |
| M5 Max, 18-core CPU, 32-core GPU, 36GB | $2,499 | Apple |
There is no 128GB or 192GB M5 Ultra. The memory steps are 96GB, 256GB, and 512GB. The M5 Max tops out at 128GB, and Apple lists its bandwidth at 460 GB/s, or 614 GB/s with the 40-core GPU. Both chips are rated at 480 watts maximum continuous power. The delay is covered in the Mac Studio shortage report.
M5 Ultra vs M3 Ultra, Measured
Federico Viticci at MacStories tested a 256GB M5 Ultra against a 512GB M3 Ultra using MLX builds of Qwen3.8-Flash-Next, a large sparse model that peaked at about 155GB of memory at 4-bit.
| Test | M5 Ultra | M3 Ultra | Change |
|---|---|---|---|
| Prompt processing, 4K context (tokens/sec) | 2,733 | 1,163 | +135% |
| Prompt processing, 64K context (tokens/sec) | 2,732 | 1,114 | +145% |
| Generation, 4K context (tokens/sec) | 54 | 47 | +15% |
| Generation, 16K context (tokens/sec) | 52 | 37 | +41% |
| Generation, 64K context (tokens/sec) | 73 | 39 | +87% |
| Prose generation, 4-bit (tokens/sec) | 111.6 | 77.3 | +44% |
| Code generation, 4-bit (tokens/sec) | 142.9 | 103.9 | +37% |
| Time to first token, 128K prompt | 50.0 s | 121 s | 59% faster |
| GLM-5.3-Flash, generation (tokens/sec) | 31 | 21 | +48% |
The larger gain is in prompt processing. That matters for agents, which resend a long context on every turn. See time-to-first-token benchmarks. The generation figures vary by test in the review, so quote the test name with the number.
M5 Ultra vs RTX 5090
MacStories also ran a dense 27B model, Qwen3.8-27B, on the M5 Ultra with MLX and on an RTX 5090 PC with llama.cpp at Q4_K_M.
| Metric | M5 Ultra | RTX 5090 |
|---|---|---|
| Prompt processing, 6K prompt (tokens/sec) | 1,701 | 3,031 |
| Generation, 6K prompt (tokens/sec) | 48 | 59 |
| Time to first token, 6K prompt | 4.0 s | 2.0 s |
| Generation at 64K context (tokens/sec) | 38.9 | 49.6 |
| Generation at 256K context (tokens/sec) | 24.3 | 30.0 |
The RTX 5090 wins on a model that fits in 32GB. It cannot load the 155GB model in the previous table at all. Details are on the RTX 5090 page.
Tokens per Second by Model Size
| Model class | M5 Ultra generation speed | Status |
|---|---|---|
| 7B, Llama 2 7B Q4_0 | No M5 Ultra row yet | The llama.cpp scoreboard lists the M5 Max at 119.9 tokens per second |
| 27B dense | 48 tokens per second | Measured by MacStories |
| 70B dense | 40 to 52 tokens per second | Listed by PromptQuorum as early community figures. Unverified |
| Large sparse, about 155GB at 4-bit | 52 to 143 tokens per second | Measured by MacStories, varies by test |
| GLM-5.2, 418GB at 4-bit, on 512GB | About 26 tokens per second | Pinggy estimate scaled from M3 Ultra. Not measured |
For which models to run, see best local LLMs for Mac M5. For earlier chips, see Apple Silicon benchmarks from M1 to M5.
Mac Studio M5 Ultra vs DGX Spark
No reviewer has run both on the same test. On specifications, the M5 Ultra has 1.2 TB/s of bandwidth and 96GB to 512GB of memory from $5,499. DGX Spark has 273 GB/s and 128GB at $4,699. A September 2026 roundup by Context Studios puts Qwen3.8-Flash-Next at 60 to 85 tokens per second on the M5 Ultra at long context and 65 to 81 on one DGX Spark, drawn from different testers. DGX Spark keeps the CUDA software stack and costs less. The 256GB and 512GB M5 Ultra configurations hold models two to four times larger. See RTX Spark vs DGX Spark.
What Is Still Unknown
- The price of the 512GB model and any measurement on it.
- A verified dense 70B result on standard llama.cpp or MLX settings.
- M5 Ultra rows on the public llama.cpp Apple Silicon scoreboard.
- Results for the 96GB base model. The published review used 256GB.
Brand Visibility Implications
A 256GB or 512GB Mac can hold open-weight models with several hundred billion parameters. Those models answer product and vendor questions from training data, with no search step unless the user adds one. What a model learned about a brand before its cutoff is the whole answer, and it is given on a desk where no analytics tool can see it. See the local LLM visibility blind spot.
Methodology
This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.
How Presenc AI Helps
Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on Apple Silicon. That shows what a model with no retrieval says about you.