July 2026 continued the compressed release cadence of the preceding quarter, with four significant launches plus Claude Sonnet 5 landing on June 30. The month's defining event was Kimi K3 taking the open-weight parameter record into the trillions.
The Releases
| Date | Model | Lab | Weights | Headline |
|---|---|---|---|---|
| Jun 30 | Claude Sonnet 5 | Anthropic | Closed | 1M context by default, $2/$10 intro pricing |
| Jul 1 | Nemotron-Labs-TwoTower | NVIDIA | Open | Diffusion LM, 2.42x throughput at 98.7% quality |
| Jul 8 | Grok 4.5 | xAI | Closed | Coding focus, $2/$6, context cut to 500K |
| Jul 16 | Kimi K3 | Moonshot AI | Open, modified MIT | 2.8T parameters, largest open weights to date |
| Jul 21 | Gemini 3.6 Flash | Closed | Price cut, cutoff jumps to March 2026 |
What Actually Changed
The open-weight ceiling moved again. Kimi K3 at 2.8 trillion parameters is the first open release into the 3-trillion class and took first place on Arena.ai's Frontend Code Arena at 1,679, ahead of Claude Fable 5 at 1,631. It also needs roughly 1.4 terabytes of memory at MXFP4, so almost nobody will run it themselves. See what K3 actually requires.
Prices fell at the volume tier. Gemini 3.6 Flash cut output pricing below its predecessor while using roughly 17 percent fewer output tokens, a rare generation-over-generation reduction on both axes. Grok 4.5 at $2/$6 undercuts most frontier competitors on output.
A knowledge cutoff jumped 14 months. Gemini 3.6 Flash moved from a January 2025 cutoff to March 2026. Because Flash-tier models typically back high-volume Google surfaces, this propagates further than a Pro-tier refresh would.
Parallel decoding got a serious open implementation. Nemotron TwoTower demonstrated that a lab holding an autoregressive checkpoint can add diffusion-based parallel generation by training only a second denoiser network, at a fraction of the original data budget.
One context window shrank. Grok 4.5 dropped to 500K from Grok 4.3's 1M, a rare reversal in a year of expansion.
Brand Visibility Implications
Two things to act on. Re-baseline Gemini and Google AI Overviews within two weeks, because a 14-month cutoff jump ingests a large block of web events into parametric recall at once, and brands whose strongest coverage predates 2025 may see relative decline as fresher competitors enter the model's knowledge. And re-baseline Claude, because Sonnet 5 became the default on Free and Pro, which resets brand answers for most of the Claude userbase rather than only for API customers.
Methodology
Compiled from vendor announcements, technical reports, and independent coverage through 2026-07-25. Per-model sourcing and confidence ratings are in the release briefs linked above. Dates are announcement dates; weights availability sometimes trails, as with Kimi K3 where the model was announced July 16 and weights were scheduled for July 27.
How Presenc AI Helps
Presenc AI re-runs brand baselines against each major model release so teams can attribute visibility changes to the model rather than to their own content.