August 2026 was a month without a new flagship from Anthropic, OpenAI, or Google, and it still moved the market. Alibaba shipped the largest Qwen model yet and promised open weights for it. Meta turned Muse Spark into a coding product. xAI shipped Grok 4.6 with long-context surcharges. Google pushed Gemini 3.7 Flash into its surfaces three weeks before replacing it. Z.ai put a 753B open-weight model on Hugging Face after a safety hold. The frontier labs' big releases then landed in the first three days of September, which is why this page and the September briefs should be read together.
The releases at a glance
| Date | Model | Lab | Type | Key spec |
|---|---|---|---|---|
| Aug 3 | Qwen3.8-Max | Alibaba | Closed at launch, open weights announced | 2.4T MoE, ~95B active, 1M context, $2 / $6 per M tokens |
| Aug 5 | Muse Spark 1.2 + Muse Code | Meta | Closed, open weights announced Aug 10 | $1.25 / $4.25, or $0.10 / $0.20 if Meta may train on traffic |
| Aug 12 | Grok 4.6 | xAI | Closed | 500K context, $2 / $6 under 200K tokens, $4 / $12 above |
| Aug 13 | Gemini 3.7 Flash | Closed | $0.75 / $3.75 introductory through Dec 31, 2026 | |
| Aug 14 / Aug 28 | GLM-5.3 | Z.ai | API first, open weights Aug 28 | 753B, custom license; GLM-5.3-Flash (MIT) on Aug 26 |
| August | GPT-5.6 Sol and Luna updates | OpenAI | Closed | Sol refreshed in ChatGPT, Luna expanded to free users, o3 retired from ChatGPT Aug 26 |
Why August 2026 matters for brand visibility
Two patterns stand out. The first is that the open-weight frontier got much bigger. Qwen3.8-Max at 2.4 trillion parameters and GLM-5.3 at 753B join Kimi K3 from July, and Meta said it will open Muse Spark 1.2. Models of this size run inside enterprises and national clouds where nobody outside can observe the answers, so what they learned about a brand at training time is what buyers get. The second is that the mid-tier refresh cycle is now measured in weeks. Gemini 3.7 Flash lasted three weeks as Google's newest Flash model, and it went straight into consumer surfaces. Brand answers on Google's AI products can shift between two tracking runs without any announcement inside Search.
Alibaba: Qwen3.8-Max
Released August 3, Qwen3.8-Max is a 2.4-trillion-parameter mixture-of-experts model with about 95B active parameters, a 1M-token context window, and text, image, and video input, priced at $2 / $6 per million tokens. Alibaba reports 86.1 on OSWorld-Verified, ahead of the scores it lists for GPT-5.6 Sol and Claude Fable 5. Treat that as a vendor claim until independent runs catch up. Alibaba also said this is the first Max-class Qwen model it will release as open weights, followed on September 2 by a coding and agent snapshot, Qwen3.8-Max-0902. For brands selling into China or into developer tooling, Qwen is now the model to sample first. See Qwen lineage and roadmap.
Meta: Muse Spark 1.2 and Muse Code
Meta shipped Muse Spark 1.2 on August 5 alongside Muse Code, the first coding agent from Meta Superintelligence Labs. The model has two prices for the same weights: $1.25 / $4.25 per million tokens, or $0.10 / $0.20 if you let Meta train on your traffic. That second tier is new in the industry and worth noticing. It turns every cheap deployment into a training-data source for the next model. On August 10 Meta said it would open the weights. Artificial Analysis placed Muse Spark 1.2 level with Grok 4.5 on its Intelligence Index. See Meta AI usage statistics.
xAI: Grok 4.6
Grok 4.6 arrived on the xAI API in mid-August with a 500K context window, text and image input, and reasoning effort from low to xhigh. Pricing is $2 input, $0.50 cached, and $6 output per million tokens for prompts under 200K tokens. Once a prompt reaches 200K, the whole request is billed at $4 / $1 / $12, not just the tokens above the threshold, which changes the cost of long-context monitoring runs. Grok's retrieval still leans on X, so the brand-visibility playbook from Grok 4.5 carries over.
Google: Gemini 3.7 Flash
Gemini 3.7 Flash shipped on August 13 at an introductory $0.75 / $3.75 per million tokens, running through December 31, 2026. Three weeks later Gemini 3.8 Flash replaced it at the same price, and Google now recommends 3.7 Flash only for efficiency-first workloads. The practical point for brands is cadence. Flash models power AI Mode, and AI Mode answers changed twice in about a month. See Gemini 3.8 Flash brief.
Z.ai: GLM-5.3
Z.ai launched GLM-5.3 by API on August 14 and held back the weights for about two weeks, saying cyber capability had improved faster than expected in post-training and needed extra safety review. The 753B weights went to Hugging Face on August 28 under a custom license that adds a security review for very large model-as-a-service operators. GLM-5.3-Flash, a 320B natively multimodal model with about 18B active parameters, shipped on August 26 under plain MIT. See GLM lineage and GLM-5.2 brief.
OpenAI: updates, not a new flagship
OpenAI spent August on its existing GPT-5.6 family, launched July 9 in three tiers (Luna, Terra, Sol). It refreshed GPT-5.6 Sol in ChatGPT, expanded GPT-5.6 Luna to free users, and retired o3 from ChatGPT on August 26. Moving free users to Luna matters more for brand answers than any benchmark, because free users are most of ChatGPT's audience. The next flagship, GPT-6 Astra, followed on September 3.
What to do this month
1. Re-baseline ChatGPT answers now that free users are on GPT-5.6 Luna, and again for paid users once GPT-6 Astra finishes rolling out.
2. Re-run AI Mode and Gemini prompt sets after every Flash update. Two Flash releases in three weeks means monthly tracking is no longer frequent enough on Google surfaces.
3. If you sell into China, APAC, or developer tooling, add Qwen3.8-Max to your model set. Its open weights will spread into self-hosted deployments you cannot monitor directly.
4. Check what large open-weight models already say about you. Qwen3.8-Max, GLM-5.3, and Muse Spark 1.2 will run in places no API monitor reaches, so their training-time picture of your brand is the one that lasts. See brand recall in open-weight versus closed models.
5. Budget for long-context surcharges. Grok 4.6 above 200K tokens and GPT-6 Astra above 272K both bill the whole request at a higher rate.