Mistral AI releases Mixtral 8x7B, an open mixture-of-experts model
Paris-based Mistral AI released Mixtral 8x7B under Apache 2.0, a sparse mixture-of-experts model that matched or beat Llama 2 70B and GPT-3.5 on many benchmarks while using ~13B active parameters per token.
Key facts
- Announced 11 December 2023 (weights shared via torrent days earlier)
- 46.7B total parameters, ~12.9B active per token
- Apache 2.0 license
- Paper: arXiv 2401.04088
What happened
Mistral published a high-quality open-weights sparse MoE model with a fully permissive license.
Why it matters
Popularized mixture-of-experts in open models (later used by DeepSeek-V3, Llama 4, Qwen) and established Europe's leading AI startup.
Changelog
- 2026-09-29: created
Related events
- Meta releases Llama 2 with a commercial-use license ★★★★
- DeepSeek-V3: frontier-level open model trained for ~$5.6M in GPU time ★★★★★
Sources (2)
id: 2023-12-11-mixtral · updated 2026-09-29 · open in the interactive timeline