Research

Local AI & Hardware Research

Research on running AI locally: DGX Spark and Mac Studio comparisons, GPU supply and pricing, on-device model performance, and what local inference means for brand visibility.

Local and on-device inference is changing which models answer questions about your brand, and where those answers come from. These reports cover the hardware, the throughput, and the practical limits of running frontier-class models outside a hyperscaler.

Local LLM Tokens-per-Second Benchmarks 2026

Public-source benchmark data for tokens-per-second on local LLM hardware in 2026: NVIDIA DGX Spark, Mac Studio M5 Max 128GB, RTX 5090, M5 Ultra, by model size and quantization.

Read More

NVIDIA DGX Spark vs Mac Studio M5 Max for Local AI

Head-to-head 2026 comparison of NVIDIA DGX Spark and Mac Studio M5 Max 128GB for local LLM inference, fine-tuning, and on-device RAG: throughput, memory bandwidth, fine-tune speed, watts, dollars-per-token.

Read More

Local LLM Fine-Tuning Hardware Requirements 2026

VRAM, unified memory, disk, and time-per-epoch requirements for fine-tuning local LLMs in 2026. Full fine-tune vs LoRA vs QLoRA across 7B, 13B, 30B, 70B, 120B parameters on DGX Spark, Mac Studio, and consumer GPUs.

Read More

Local LLM Quantization Quality Benchmarks 2026

Quality vs speed vs memory benchmarks for GGUF, MLX, AWQ, and GPTQ quantization formats in 2026. Perplexity delta, tokens-per-second speedup, memory savings across Q2, Q3, Q4, Q5, Q6, Q8.

Read More

Local LLM vs Cloud API Cost Comparison 2026

Total cost of ownership for running LLMs locally on DGX Spark, Mac Studio, and consumer GPUs versus paying per-token to OpenAI, Anthropic, and Google in 2026. Breakeven by model size and utilisation.

Read More

Local LLM Hardware Landscape 2026

Complete 2026 landscape of local LLM workstations and edge AI hardware: NVIDIA DGX Spark, Mac Studio M5, AMD Strix Halo, Framework Desktop, NVIDIA Jetson Thor, consumer GPUs, with strengths, prices, and use-case fit.

Read More

On-Device RAG Performance Benchmarks 2026

Local RAG (retrieval-augmented generation) performance in 2026: index size limits, retrieval latency, queries per second, and end-to-end RAG response time across DGX Spark, Mac Studio, and consumer GPU hardware.

Read More

Best Local LLMs for Mac M5 Series 2026

Ranked guide to the best local LLMs for Apple Silicon M5 Max and M5 Ultra in 2026. Throughput, model fit, MLX availability, and use-case recommendations across Llama 4, Qwen 3, Mistral, gpt-oss, and Phi.

Read More

Local LLM Brand Visibility Blind Spot 2026

How fast-growing local and air-gapped LLM deployments produce brand-relevant queries that no cloud-AI visibility platform can see. Surface-size estimates, sector exposure, and the operational answer for brands.

Read More

AI GPU Supply and Pricing 2026

AI GPU supply, lead times, and rental pricing in 2026: H100, H200, B200, GB200, RTX 5090. Cloud rental rates from Lambda Labs, CoreWeave, Crusoe, and on-prem economics.

Read More

Qwen 3.5 Quantization Speed Benchmarks 2026

Tokens-per-second benchmarks for Qwen 3.5 at every GGUF quantization tier (Q2 through Q8) on Apple Silicon M5 Max, NVIDIA RTX 3090/4090, and consumer hardware. Memory footprint, speed, and context-decay tradeoffs.

Read More

Mac Studio & Mac Mini Shortage 2026

Mac Studio and Mac Mini delivery delays stretched to 10 weeks in May 2026. Apple cites unexpected AI/agentic demand. M5 Ultra Mac Studio delayed to October. Snapshot for 2026-05-15.

Read More

NVIDIA RTX 60 & AMD RDNA5 Delays 2026

NVIDIA RTX 60 and AMD RDNA5 consumer GPU launches pushed to mid-late 2027 due to GDDR memory shortage. Intel Arc B700 uncertain. Snapshot for 2026-05-15.

Read More

GPU Shipment Tracker Blackwell to Rubin 2026

NVIDIA Blackwell to Rubin transition in 2026: Blackwell 5.2M to 1.8M, Rubin 5.7M target capped at ~300k by TSMC N3, AMD MI400 and MI450 ramp, Intel Gaudi 3 and Falcon Shores.

Read More

Frontier Lab GPU Counts 2026

Frontier AI lab compute concentration in 2026: Anthropic SpaceX 220k GPU deal, xAI Memphis Colossus 200k scaling to 1M, OpenAI Stargate 7GW, plus Google TPU and Meta MI450 counts.

Read More

Quantization Format Comparison 2026

AI quantization format comparison 2026: GGUF, AWQ, GPTQ, EXL2, MLX, FP8, NF4, INT4, INT8. Quality degradation, throughput, VRAM, and toolchain support across llama.cpp, vLLM, TensorRT-LLM.

Read More

Ollama Ecosystem State 2026

Ollama ecosystem state 2026: ~5M active users, model registry, GUI clients (Open WebUI, Msty, Cherry Studio), enterprise adoption, OpenAI-compatible API integration patterns.

Read More

Gemini Spark: Google's 24/7 Personal AI Agent (I/O 2026)

Gemini Spark is Google's new autonomous personal AI agent, announced at Google I/O 2026, running 24/7 across apps with MCP integration and the Android Halo UI.

Read More

DGX Spark vs M5 Max vs RTX 5090: Local LLM Throughput

Three-way 2026 throughput comparison of NVIDIA DGX Spark, Mac Studio M5 Max 128GB, and GeForce RTX 5090 32GB for local LLM inference. Decode speed, prefill, memory ceiling, and which model sizes each one can actually hold.

Read More

MLX vs llama.cpp: Apple Silicon Throughput Benchmarks

Head-to-head 2026 benchmarks for MLX and llama.cpp on Apple Silicon. Decode speed by model size, time-to-first-token on M5 Neural Accelerators, context-length crossover, and which runtime to pick.

Read More

Memory Shortage 2026: How AI Demand Raised Device Prices

How the AI-driven DRAM shortage reached consumer device prices in 2026. Spot price moves, memory as a share of laptop bill of materials, Apple and PC price increases, Gartner and IDC forecasts, and when supply recovers.

Read More

Unified Memory vs VRAM for Local LLMs

How unified memory and discrete VRAM differ for local LLM inference in 2026. Capacity against bandwidth, where each architecture wins, the offload cliff, and how to size a machine for the models you actually run.

Read More

Best Local LLM Hardware Under $5,000

Build and buy options for local LLM inference under $5,000 in 2026, ranked by what they can actually run. DGX Spark, Mac Studio, RTX 5090 builds, dual-GPU rigs, and used-hardware options, with memory-shortage pricing factored in.

Read More

Kimi K3 Local Hardware Requirements

What it actually takes to run Moonshot Kimi K3 locally. 2.8 trillion parameters, sparse expert routing, MXFP4 quantisation, memory requirements by precision, and why open weights do not mean runnable weights.

Read More

MoE Active vs Total Parameters: A Hardware Guide

Why sparse mixture-of-experts models need memory for every parameter but only compute for a few. Active versus total parameter counts for the 2026 open-weight models, and how to size hardware correctly.

Read More

How Much RAM to Run Each Open-Weight Model

Memory requirement lookup table for the major 2026 open-weight models at Q4, Q8, and BF16, including key-value cache overhead at long context and the minimum hardware that clears each tier.

Read More

Apple Silicon LLM Benchmarks: M1 to M5

Generation-by-generation local LLM throughput across Apple Silicon from M1 to M5. Memory bandwidth, decode speed by model size, the M5 Neural Accelerator step change, and which generation is worth upgrading from.

Read More

Time-to-First-Token Benchmarks for Local LLMs

Prefill and time-to-first-token benchmarks across local AI hardware in 2026. Why prompt processing dominates agentic workloads, measured prefill rates by device, and how context length changes the calculation.

Read More

Groq vs Cerebras Inference Speed Benchmarks

Custom-silicon inference speeds in 2026. Groq LPU and Cerebras wafer-scale throughput on open-weight models, how they compare to GPU serving, and which workloads justify specialised hardware.

Read More

Inference Provider Pricing Comparison

What it costs to serve open-weight models across inference providers in 2026. Pricing structures, the gap to first-party APIs, why the same model costs different amounts, and what to check beyond the headline rate.

Read More

LLM API Latency Benchmarks by Provider

How to measure and compare LLM API latency in 2026. Time to first token versus total completion, the tail-latency problem, what routing adds, and why published averages mislead.

Read More

LLM API Uptime and Reliability

How reliable LLM APIs actually are in 2026. Why uptime pages understate real failure rates, the failure modes that do not register as downtime, and how to build against them.

Read More

Structured Output Reliability

How reliably models produce valid structured output in 2026. Constrained decoding versus prompting, where schema conformance still fails, and why valid JSON is not the same as correct data.

Read More

Diffusion Language Models Landscape

Where diffusion language models stand in 2026. How parallel text generation differs from autoregressive decoding, the throughput case, current open-weight implementations, and the open questions.

Read More

Speculative Decoding Adoption

How speculative decoding works, what speedups it delivers in practice, why acceptance rate determines everything, and how widely it is deployed across serving stacks in 2026.

Read More

KV Cache Compression Techniques

Why the key-value cache became the memory bottleneck at long context, and the techniques used to shrink it in 2026. Quantisation, eviction, attention-architecture changes, and their quality costs.

Read More

Sparse Attention Architectures Compared

The attention mechanisms behind 2026 long-context models. MiniMax Sparse Attention, Kimi Delta Attention, sliding-window and hybrid designs, what each trades away, and the reported efficiency gains.

Read More

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.