At a Glance
| Vendor | Alibaba |
| Family | Qwen 3 series |
| Launched | Alibaba released Qwen3-Coder-Next in February 2026 under Apache 2.0, purpose-built for coding agents and local development. It is included in this catalogue because it has become the reference point for ultra-sparse local coding models rather than because it is the newest release. |
| Context window | 256,000 tokens natively, extendable to 1,000,000 for large-codebase work. |
| Pricing | Free to self-host under Apache 2.0, which is the most permissive licence among the credible coding models. Hosted API access through Alibaba Cloud and third-party providers is among the cheapest for a model of this capability. |
| Access channels | Hugging Face and ModelScope checkpoints in multiple sizes, LM Studio, Ollama, vLLM for self-hosting, and Alibaba Cloud Model Studio for hosted access. |
Notable Benchmarks
80 billion total parameters with just 3 billion active per forward pass. It scores 70.6 percent on SWE-Bench Verified, 62.8 percent on SWE-Bench Multilingual, and 44.3 percent on SWE-Bench Pro, while delivering throughput comparable to models with 10 to 20 times more active parameters. Language coverage expanded from 92 programming languages in Qwen2.5-Coder to 370.
Strengths
The 80B-total / 3B-active ratio is the whole story. Sparse MoE gives up to 10x the throughput of a dense model of similar size, which puts credible agentic coding on a workstation rather than a server. Apache 2.0 removes every licensing question a Western enterprise would otherwise raise.
Limitations
The full 80B still needs to fit in memory even though only 3B activates per token, so the memory footprint tracks total parameters while the speed tracks active parameters. That distinction catches people out when sizing hardware. Absolute SWE-Bench Pro performance trails the large frontier models.
Brand-Visibility Implications
Qwen3-Coder-Next is the model that made local coding agents genuinely practical, and local coding agents are where a rising share of developer tool selection now happens. When a developer asks their local agent which HTTP client or which vector database to use, that recommendation is generated entirely off parametric weights with no retrieval and no API call anyone can observe. For developer-tool brands, weight-baked recall is the only lever, and it is set at training time. See the local-LLM blind spot and open-weight code models.
How Presenc AI Tracks This Model
Presenc AI monitors brand visibility on Alibaba's Qwen 3 series as part of continuous multi-platform AI visibility tracking. We sample Alibaba Qwen3-Coder-Next across representative prompt sets daily, compare against competitor performance on the same prompts, and flag material mention-rate changes so brand teams can respond quickly when AI representation shifts.