Qwen3.8-Flash-Next: 125B MoE with only 6B active previews Qwen 4 architecture
Alibaba's Qwen team open-sourced Qwen3.8-Flash-Next on 2026-08-26: a 125B-parameter multimodal MoE activating just 6B parameters per token, with n-gram embeddings and hybrid Gated DeltaNet/sparse attention, explicitly positioned as a preview of the Qwen 4 architecture; Bloomberg said it rivals Claude Opus 4.6 and DeepSeek V4-Flash.
Key facts
- 125B backbone + 51B n-gram embeddings + 4B multi-token-prediction = ~180B on disk; 6B active per token
- 512 experts, 10 routed + 1 shared per token; Gated DeltaNet in 3 of 4 layers + Qwen Sparse Attention
- Context: 262,144 native, 1M with YaRN
- Reported benchmarks: SWE-bench Pro 62.5, AndroidWorld 84.5, MathVision 95.7
- Training cost ~1/9 of Qwen3.7-Plus; up to 7.6x prefill and 4.9x decode speedup at 1M tokens
- License: qwen-community-1.0 (not Apache 2.0)
What happened
Weights for Qwen3.8-Flash-Next landed on Hugging Face and ModelScope (BF16 and FP8) on 2026-08-26. The model combines an extreme sparsity ratio (6B of 125B active), a 20M-entry n-gram embedding table, and linear-attention (Gated DeltaNet) layers interleaved with sparse attention — the Qwen team presented it as an early look at Qwen 4 so developers can prepare tooling.
Why it matters
It pushes the cost frontier: near-frontier agentic coding numbers at 6B active parameters make strong models cheap to serve at 1M-token contexts.
Changelog
- 2026-09-29: created
Related events
- Alibaba launches Qwen3.8-Max (2.4T MoE) and open-sources the Qwen3.8 family ★★★★
- Apsara 2026: Alibaba says Qwen 4 is in training, targets 5-10T-parameter Qwen 4.5/5, reports self-improvement runs and unveils Zhenwu V900 chip ★★★
Sources (3)
- pressBloomberg: Alibaba releases smaller, cost-effective Qwen AI model
- pressMarkTechPost: Qwen3.8-Flash-Next technical breakdown
- pressThe Decoder: Qwen3.8-Flash-Next targets ultimate cost efficiency
id: 2026-08-26-qwen3-8-flash-next · updated 2026-09-29 · open in the interactive timeline