Research

July 2026 LLM Release Roundup

Every significant LLM release in July 2026: Nemotron TwoTower, Grok 4.5, Kimi K3, Gemini 3.6 Flash, and Claude Sonnet 5 at the month boundary. Specifications, pricing, and brand-visibility implications.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: July 2026

July 2026 continued the compressed release cadence of the preceding quarter, with four significant launches plus Claude Sonnet 5 landing on June 30. The month's defining event was Kimi K3 taking the open-weight parameter record into the trillions.

The Releases

DateModelLabWeightsHeadline
Jun 30Claude Sonnet 5AnthropicClosed1M context by default, $2/$10 intro pricing
Jul 1Nemotron-Labs-TwoTowerNVIDIAOpenDiffusion LM, 2.42x throughput at 98.7% quality
Jul 8Grok 4.5xAIClosedCoding focus, $2/$6, context cut to 500K
Jul 16Kimi K3Moonshot AIOpen, modified MIT2.8T parameters, largest open weights to date
Jul 21Gemini 3.6 FlashGoogleClosedPrice cut, cutoff jumps to March 2026

What Actually Changed

The open-weight ceiling moved again. Kimi K3 at 2.8 trillion parameters is the first open release into the 3-trillion class and took first place on Arena.ai's Frontend Code Arena at 1,679, ahead of Claude Fable 5 at 1,631. It also needs roughly 1.4 terabytes of memory at MXFP4, so almost nobody will run it themselves. See what K3 actually requires.

Prices fell at the volume tier. Gemini 3.6 Flash cut output pricing below its predecessor while using roughly 17 percent fewer output tokens, a rare generation-over-generation reduction on both axes. Grok 4.5 at $2/$6 undercuts most frontier competitors on output.

A knowledge cutoff jumped 14 months. Gemini 3.6 Flash moved from a January 2025 cutoff to March 2026. Because Flash-tier models typically back high-volume Google surfaces, this propagates further than a Pro-tier refresh would.

Parallel decoding got a serious open implementation. Nemotron TwoTower demonstrated that a lab holding an autoregressive checkpoint can add diffusion-based parallel generation by training only a second denoiser network, at a fraction of the original data budget.

One context window shrank. Grok 4.5 dropped to 500K from Grok 4.3's 1M, a rare reversal in a year of expansion.

Brand Visibility Implications

Two things to act on. Re-baseline Gemini and Google AI Overviews within two weeks, because a 14-month cutoff jump ingests a large block of web events into parametric recall at once, and brands whose strongest coverage predates 2025 may see relative decline as fresher competitors enter the model's knowledge. And re-baseline Claude, because Sonnet 5 became the default on Free and Pro, which resets brand answers for most of the Claude userbase rather than only for API customers.

Methodology

Compiled from vendor announcements, technical reports, and independent coverage through 2026-07-25. Per-model sourcing and confidence ratings are in the release briefs linked above. Dates are announcement dates; weights availability sometimes trails, as with Kimi K3 where the model was announced July 16 and weights were scheduled for July 27.

How Presenc AI Helps

Presenc AI re-runs brand baselines against each major model release so teams can attribute visibility changes to the model rather than to their own content.

Frequently Asked Questions

NVIDIA Nemotron-Labs-TwoTower on July 1, xAI Grok 4.5 on July 8, Moonshot Kimi K3 on July 16, and Google Gemini 3.6 Flash on July 21. Anthropic's Claude Sonnet 5 landed just before the month on June 30.
Kimi K3, at 2.8 trillion parameters the largest open-weight model published to date and the first into the 3-trillion class. It took first place on Arena.ai's Frontend Code Arena at 1,679, ahead of Claude Fable 5 at 1,631.
Down at the volume tier. Gemini 3.6 Flash cut output pricing below its predecessor while also using roughly 17 percent fewer output tokens, improving effective cost on both axes. Grok 4.5 launched at $2/$6, undercutting most frontier competitors on output pricing.
Gemini 3.6 Flash, because its knowledge cutoff jumped 14 months from January 2025 to March 2026 and Flash-tier models typically back high-volume Google surfaces including AI Overviews. Claude Sonnet 5 is close behind because it became the default model on Claude's Free and Pro tiers.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.