Research

Mac Studio M5 Ultra Local LLM Benchmarks

Mac Studio M5 Ultra for local LLMs: 96GB, 256GB, and 512GB configurations, prices from $5,499, and the first independent tokens-per-second results against the M3 Ultra, RTX 5090, and DGX Spark.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: October 2026

The Mac Studio with M5 Ultra went on sale on September 22, 2026, starting at $5,499 with 96GB of unified memory. Apple lists memory bandwidth of 1.2 TB/s and offers 256GB and 512GB options, with the 512GB model arriving in late October. In the first independent local LLM review, MacStories measured 48 tokens per second on a dense 27B model and between 52 and 143 tokens per second on a large sparse model, depending on the test. Prompt processing was more than twice as fast as on the M3 Ultra. An RTX 5090 was still about 23 percent faster on the 27B model that fits in its memory.

Configurations and Prices

ConfigurationPriceSource
M5 Ultra, 30-core CPU, 64-core GPU, 96GB, 1TB$5,499Apple
M5 Ultra, 256GBAbout $9,499Pinggy and MortalApps buyer guides
M5 Ultra, 36-core CPU, 80-core GPU, 256GB, 2TB$11,299Engadget review unit
M5 Ultra, 512GBNot announced, late OctoberMacRumors expects well above $10,000
M5 Max, 18-core CPU, 32-core GPU, 36GB$2,499Apple

There is no 128GB or 192GB M5 Ultra. The memory steps are 96GB, 256GB, and 512GB. The M5 Max tops out at 128GB, and Apple lists its bandwidth at 460 GB/s, or 614 GB/s with the 40-core GPU. Both chips are rated at 480 watts maximum continuous power. The delay is covered in the Mac Studio shortage report.

M5 Ultra vs M3 Ultra, Measured

Federico Viticci at MacStories tested a 256GB M5 Ultra against a 512GB M3 Ultra using MLX builds of Qwen3.8-Flash-Next, a large sparse model that peaked at about 155GB of memory at 4-bit.

TestM5 UltraM3 UltraChange
Prompt processing, 4K context (tokens/sec)2,7331,163+135%
Prompt processing, 64K context (tokens/sec)2,7321,114+145%
Generation, 4K context (tokens/sec)5447+15%
Generation, 16K context (tokens/sec)5237+41%
Generation, 64K context (tokens/sec)7339+87%
Prose generation, 4-bit (tokens/sec)111.677.3+44%
Code generation, 4-bit (tokens/sec)142.9103.9+37%
Time to first token, 128K prompt50.0 s121 s59% faster
GLM-5.3-Flash, generation (tokens/sec)3121+48%

The larger gain is in prompt processing. That matters for agents, which resend a long context on every turn. See time-to-first-token benchmarks. The generation figures vary by test in the review, so quote the test name with the number.

M5 Ultra vs RTX 5090

MacStories also ran a dense 27B model, Qwen3.8-27B, on the M5 Ultra with MLX and on an RTX 5090 PC with llama.cpp at Q4_K_M.

MetricM5 UltraRTX 5090
Prompt processing, 6K prompt (tokens/sec)1,7013,031
Generation, 6K prompt (tokens/sec)4859
Time to first token, 6K prompt4.0 s2.0 s
Generation at 64K context (tokens/sec)38.949.6
Generation at 256K context (tokens/sec)24.330.0

The RTX 5090 wins on a model that fits in 32GB. It cannot load the 155GB model in the previous table at all. Details are on the RTX 5090 page.

Tokens per Second by Model Size

Model classM5 Ultra generation speedStatus
7B, Llama 2 7B Q4_0No M5 Ultra row yetThe llama.cpp scoreboard lists the M5 Max at 119.9 tokens per second
27B dense48 tokens per secondMeasured by MacStories
70B dense40 to 52 tokens per secondListed by PromptQuorum as early community figures. Unverified
Large sparse, about 155GB at 4-bit52 to 143 tokens per secondMeasured by MacStories, varies by test
GLM-5.2, 418GB at 4-bit, on 512GBAbout 26 tokens per secondPinggy estimate scaled from M3 Ultra. Not measured

For which models to run, see best local LLMs for Mac M5. For earlier chips, see Apple Silicon benchmarks from M1 to M5.

Mac Studio M5 Ultra vs DGX Spark

No reviewer has run both on the same test. On specifications, the M5 Ultra has 1.2 TB/s of bandwidth and 96GB to 512GB of memory from $5,499. DGX Spark has 273 GB/s and 128GB at $4,699. A September 2026 roundup by Context Studios puts Qwen3.8-Flash-Next at 60 to 85 tokens per second on the M5 Ultra at long context and 65 to 81 on one DGX Spark, drawn from different testers. DGX Spark keeps the CUDA software stack and costs less. The 256GB and 512GB M5 Ultra configurations hold models two to four times larger. See RTX Spark vs DGX Spark.

What Is Still Unknown

  • The price of the 512GB model and any measurement on it.
  • A verified dense 70B result on standard llama.cpp or MLX settings.
  • M5 Ultra rows on the public llama.cpp Apple Silicon scoreboard.
  • Results for the 96GB base model. The published review used 256GB.

Brand Visibility Implications

A 256GB or 512GB Mac can hold open-weight models with several hundred billion parameters. Those models answer product and vendor questions from training data, with no search step unless the user adds one. What a model learned about a brand before its cutoff is the whole answer, and it is given on a desk where no analytics tool can see it. See the local LLM visibility blind spot.

Methodology

This page compiles vendor specifications and third-party benchmark reports. None of the figures are Presenc AI measurements. Vendor facts come from primary pages, including NVIDIA's RTX Spark announcement and DGX Spark product page, and Apple's Mac Studio announcement. Throughput figures come from named testers: the llama.cpp scoreboards for CUDA cards, DGX Spark and Apple Silicon, Hardware Corner, MacStories, LMSYS, the Ollama blog, StorageReview, and the Level1Techs forum. Each number is attributed in the text to whoever measured or claimed it. Testers use different models, quantisation, runtimes, and context lengths, so compare figures within one source and treat cross-source comparisons as approximate. Prices are as reported in late September 2026 and are moving with memory supply. Status as of October 1, 2026.

How Presenc AI Helps

Presenc AI tracks how AI models describe and recommend your brand, including the open-weight model families people run on Apple Silicon. That shows what a model with no retrieval says about you.

Frequently Asked Questions

In MacStories' review of a 256GB unit, the M5 Ultra generated 48 tokens per second on the dense Qwen3.8-27B model, and between 52 and 143 tokens per second on the sparse Qwen3.8-Flash-Next model depending on the test. Prompt processing on that model reached about 2,700 tokens per second, against about 1,150 on the M3 Ultra.
It starts at $5,499 with 96GB of memory and 1TB of storage. Buyer guides put the 256GB configuration at about $9,499, and Engadget's review unit with the 36-core CPU, 256GB, and 2TB cost $11,299. The 512GB model ships in late October 2026 and Apple has not announced its price.
Not on models that fit in the 5090's 32GB. MacStories measured 59 tokens per second on the RTX 5090 against 48 on the M5 Ultra for Qwen3.8-27B, and prompt processing of 3,031 against 1,701 tokens per second. The M5 Ultra can load models of 150GB and more, which a single RTX 5090 cannot.
No head-to-head test exists yet. The M5 Ultra has 1.2 TB/s of memory bandwidth and up to 512GB of memory, starting at $5,499. DGX Spark has 273 GB/s and 128GB at $4,699 and runs NVIDIA's CUDA software. Choose the Mac for larger models and macOS, and DGX Spark for CUDA tooling and fine-tuning.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.