Research

August 2026 LLM Releases: What Changed for Brand Visibility

Every major LLM release of August 2026 (Qwen3.8-Max, Muse Spark 1.2, Grok 4.6, Gemini 3.7 Flash, GLM-5.3, GPT-5.6 updates) with dates, context, pricing, and the brand-visibility shift each one creates.

By Ramanath, CTO & Co-Founder at Presenc AI · Last updated: September 2026

August 2026 was a month without a new flagship from Anthropic, OpenAI, or Google, and it still moved the market. Alibaba shipped the largest Qwen model yet and promised open weights for it. Meta turned Muse Spark into a coding product. xAI shipped Grok 4.6 with long-context surcharges. Google pushed Gemini 3.7 Flash into its surfaces three weeks before replacing it. Z.ai put a 753B open-weight model on Hugging Face after a safety hold. The frontier labs' big releases then landed in the first three days of September, which is why this page and the September briefs should be read together.

The releases at a glance

DateModelLabTypeKey spec
Aug 3Qwen3.8-MaxAlibabaClosed at launch, open weights announced2.4T MoE, ~95B active, 1M context, $2 / $6 per M tokens
Aug 5Muse Spark 1.2 + Muse CodeMetaClosed, open weights announced Aug 10$1.25 / $4.25, or $0.10 / $0.20 if Meta may train on traffic
Aug 12Grok 4.6xAIClosed500K context, $2 / $6 under 200K tokens, $4 / $12 above
Aug 13Gemini 3.7 FlashGoogleClosed$0.75 / $3.75 introductory through Dec 31, 2026
Aug 14 / Aug 28GLM-5.3Z.aiAPI first, open weights Aug 28753B, custom license; GLM-5.3-Flash (MIT) on Aug 26
AugustGPT-5.6 Sol and Luna updatesOpenAIClosedSol refreshed in ChatGPT, Luna expanded to free users, o3 retired from ChatGPT Aug 26

Why August 2026 matters for brand visibility

Two patterns stand out. The first is that the open-weight frontier got much bigger. Qwen3.8-Max at 2.4 trillion parameters and GLM-5.3 at 753B join Kimi K3 from July, and Meta said it will open Muse Spark 1.2. Models of this size run inside enterprises and national clouds where nobody outside can observe the answers, so what they learned about a brand at training time is what buyers get. The second is that the mid-tier refresh cycle is now measured in weeks. Gemini 3.7 Flash lasted three weeks as Google's newest Flash model, and it went straight into consumer surfaces. Brand answers on Google's AI products can shift between two tracking runs without any announcement inside Search.

Alibaba: Qwen3.8-Max

Released August 3, Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model with about 95B active parameters, a 1M-token context window, and text, image, and video input, priced at $2 / $6 per million tokens. Alibaba reports 86.1 on OSWorld-Verified, ahead of the scores it lists for GPT-5.6 Sol and Claude Fable 5. Treat that as a vendor claim until independent runs catch up. Alibaba also said this is the first Max-class Qwen model it will release as open weights, followed on September 2 by a coding and agent snapshot, Qwen3.8-Max-0902. For brands selling into China or into developer tooling, Qwen is now the model to sample first. See Qwen lineage and roadmap.

Meta: Muse Spark 1.2 and Muse Code

Meta shipped Muse Spark 1.2 on August 5 alongside Muse Code, the first coding agent from Meta Superintelligence Labs. The model has two prices for the same weights: $1.25 / $4.25 per million tokens, or $0.10 / $0.20 if you let Meta train on your traffic. That second tier is new in the industry and worth noticing. It turns every cheap deployment into a training-data source for the next model. On August 10 Meta said it would open the weights. Artificial Analysis placed Muse Spark 1.2 level with Grok 4.5 on its Intelligence Index. See Meta AI usage statistics.

xAI: Grok 4.6

Grok 4.6 arrived on the xAI API in mid-August with a 500K context window, text and image input, and reasoning effort from low to xhigh. Pricing is $2 input, $0.50 cached, and $6 output per million tokens for prompts under 200K tokens. Once a prompt reaches 200K, the whole request is billed at $4 / $1 / $12, not just the tokens above the threshold, which changes the cost of long-context monitoring runs. Grok's retrieval still leans on X, so the brand-visibility playbook from Grok 4.5 carries over.

Google: Gemini 3.7 Flash

Gemini 3.7 Flash shipped on August 13 at an introductory $0.75 / $3.75 per million tokens, running through December 31, 2026. Three weeks later Gemini 3.8 Flash replaced it at the same price, and Google now recommends 3.7 Flash only for efficiency-first workloads. The practical point for brands is cadence. Flash models power AI Mode, and AI Mode answers changed twice in about a month. See Gemini 3.8 Flash brief.

Z.ai: GLM-5.3

Z.ai launched GLM-5.3 by API on August 14 and held back the weights for about two weeks, saying cyber capability had improved faster than expected in post-training and needed extra safety review. The 753B weights went to Hugging Face on August 28 under a custom license that adds a security review for very large model-as-a-service operators. GLM-5.3-Flash, a 320B natively multimodal model with about 18B active parameters, shipped on August 26 under plain MIT. See GLM lineage and GLM-5.2 brief.

OpenAI: updates, not a new flagship

OpenAI spent August on its existing GPT-5.6 family, launched July 9 in three tiers (Luna, Terra, Sol). It refreshed GPT-5.6 Sol in ChatGPT, expanded GPT-5.6 Luna to free users, and retired o3 from ChatGPT on August 26. Moving free users to Luna matters more for brand answers than any benchmark, because free users are most of ChatGPT's audience. The next flagship, GPT-6 Astra, followed on September 3.

What to do this month

1. Re-baseline ChatGPT answers now that free users are on GPT-5.6 Luna, and again for paid users once GPT-6 Astra finishes rolling out.

2. Re-run AI Mode and Gemini prompt sets after every Flash update. Two Flash releases in three weeks means monthly tracking is no longer frequent enough on Google surfaces.

3. If you sell into China, APAC, or developer tooling, add Qwen3.8-Max to your model set. Its open weights will spread into self-hosted deployments you cannot monitor directly.

4. Check what large open-weight models already say about you. Qwen3.8-Max, GLM-5.3, and Muse Spark 1.2 will run in places no API monitor reaches, so their training-time picture of your brand is the one that lasts. See brand recall in open-weight versus closed models.

5. Budget for long-context surcharges. Grok 4.6 above 200K tokens and GPT-6 Astra above 272K both bill the whole request at a higher rate.

Frequently Asked Questions

Alibaba Qwen3.8-Max (August 3), Meta Muse Spark 1.2 with Muse Code (August 5), xAI Grok 4.6 (mid-August), Google Gemini 3.7 Flash (August 13), and Z.ai GLM-5.3 (API August 14, open weights August 28). OpenAI updated its GPT-5.6 family rather than launching a new flagship.
No. Anthropic's Claude Opus 5 arrived on July 24 and Claude Fable 5.1 on September 1. OpenAI's GPT-6 Astra came on September 3, and Google's Gemini 3.8 Flash on September 2. August was a month of mid-tier and open-weight releases.
For most brands, OpenAI moving free ChatGPT users to GPT-5.6 Luna, because it changes answers for the largest AI audience. For brands with China or developer exposure, Qwen3.8-Max. For long-term recall, the large open-weight releases (Qwen3.8-Max, GLM-5.3, and the promised Muse Spark 1.2 weights), because they will run in deployments no one can monitor.
Under 200K prompt tokens, Grok 4.6 costs $2 input, $0.50 cached input, and $6 output per million tokens. Once a prompt reaches 200K tokens, the entire request is billed at $4, $1, and $12, not only the tokens above the threshold.
Alibaba launched it through its API and said it will release the weights, the first time it has open-sourced a Max-class Qwen model. Check Alibaba's Hugging Face organization for the current status before building on it.

Track Your AI Visibility

See how your brand appears across ChatGPT, Claude, Perplexity, and other AI platforms. Start monitoring today.