Research

MiniMax M3 Release Brief

Release brief for MiniMax M3: June 1 2026 open-weight launch, 428B MoE with 23B active, MiniMax Sparse Attention, 1M context, and brand-visibility implications.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: July 2026

At a Glance

VendorMiniMax
FamilyMiniMax M series
LaunchedThe Shanghai lab MiniMax released M3 on June 1, 2026 as an open-weight model for long-horizon coding-agent workflows. Weights went live on Hugging Face by June 7 and the technical report landed on arXiv on June 11.
Context window1,000,000 tokens via the API, accepting text, image, and video input.
PricingOpen weights are free to self-host. Hosted pricing is aggressive and positioned against the other Chinese open-weight frontier models rather than against Western closed labs.
Access channelsMiniMax API, Hugging Face open weights, and the usual open-weight inference providers. The arXiv technical report is unusually detailed on the attention mechanism.

Notable Benchmarks

428B total parameters with approximately 23B active per token. M3 scores 59.0 percent on SWE-Bench Pro, ahead of both GPT-5.5 and Gemini 3.1 Pro on that benchmark. It sits at 44 on the Artificial Analysis Intelligence Index, tied with DeepSeek V4 Pro.

Strengths

MiniMax Sparse Attention (MSA) is the headline. A lightweight index branch selects which key-value cache blocks matter, and the main attention layer processes only those. MiniMax reports more than 9x faster prefill and more than 15x faster decode versus M2, at roughly one twentieth the per-token compute. Native multimodality across text, image, and video in one open-weight model is rare.

Limitations

Coverage at launch flagged that several frontier claims were not independently verified. Composite index placement at 44 is well behind GLM-5.2 at 51 despite the strong single-benchmark SWE-Bench Pro result. Ecosystem support is thinner than for Qwen or DeepSeek.

Brand-Visibility Implications

MSA is the part worth watching. If sparse attention at this efficiency generalises, the cost of long-context inference falls sharply, and cheap long context means retrieval-heavy answering becomes the default rather than a premium mode. That shifts brand visibility further toward whoever has the most retrievable, best-structured source material, and further away from whoever simply has the strongest parametric brand recall. See open-weight long-context models and Chinese LLM comparison.

How Presenc AI Tracks This Model

Presenc AI monitors brand visibility on MiniMax's MiniMax M series as part of continuous multi-platform AI visibility tracking. We sample MiniMax M3 across representative prompt sets daily, compare against competitor performance on the same prompts, and flag material mention-rate changes so brand teams can respond quickly when AI representation shifts.

Frequently Asked Questions

The Shanghai lab MiniMax released M3 on June 1, 2026 as an open-weight model for long-horizon coding-agent workflows. Weights went live on Hugging Face by June 7 and the technical report landed on arXiv on June 11.
1,000,000 tokens via the API, accepting text, image, and video input.
MiniMax API, Hugging Face open weights, and the usual open-weight inference providers. The arXiv technical report is unusually detailed on the attention mechanism.
MSA is the part worth watching. If sparse attention at this efficiency generalises, the cost of long-context inference falls sharply, and cheap long context means retrieval-heavy answering becomes the default rather than a premium mode. That shifts brand visibility further toward whoever has the most retrievable, best-structured source material, and further away from whoever simply has the strongest parametric brand recall. See open-weight long-context models and Chinese LLM comparison.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.