Mistral Large 4: Is It Actually A Failure?
Bruno VegaYouTube16,105 views as of 10 October 2026
Why it is here
Review of Mistral Large 4 (‘le Chonk’) weighing the mixed early reception; ~16k views.
Description
Description written by Gemini from the videoGemini 3.8 Flash, 10 October 2026
Summary
In this video essay and technical review, Bruno Vega analyzes Mistral AI’s newly announced 1-trillion-parameter flagship model, Mistral Large 4 (nicknamed “Le Chonk”). Vega examines whether the model is a failure by breaking down its performance across three criteria—intelligence, cost, and sovereignty/usability—contrasting its benchmark scores against leading frontier closed models and open-weight Chinese rivals. He concludes that while Mistral Large 4 trails top Chinese open-weight models and suffers from high cached token pricing for agent workflows, it provides vital enterprise and European data-sovereignty value.
What is shown
- Architecture and Specifications [01:07 - 02:12]: Slides detailing Mistral Large 4’s mixture-of-experts (MoE) configuration (1.05T total parameters, 49B active per token, 52B including embeddings and output layers, 1.6B vision encoder), alongside API configuration snippets showing a 512k (524,288) token context limit rather than the advertised 1M context.
- Benchmark Comparisons [03:16 - 05:20]: Leaderboard breakdowns from Artificial Analysis showing the Intelligence Index, SciCode-Verified, and DeepSWE 1.1 charts comparing Large 4 to GLM-5.3, MiMo-V2.6-Pro, Qwen3.8-2.4T, GPT-6 Luna, and Claude Opus 5.5.
- Hallucination and Omniscience Evals [05:27 - 06:19]: Artificial Analysis Omniscience evaluation charts comparing accuracy (25.8%) against a low 42% hallucination rate versus DeepSeek-V4.1-Flash (96.5%) and GPT-6 Luna (76.7%).
- Cybersecurity Evaluations [06:24 - 07:48]: CyberGym-E2E benchmark results showing Large 4 scoring 82% against competitor refusal rates (Claude Opus 5.5 at 98.5% refusal and GPT-6 Astra at 100% refusal), as well as DeepSec results and discussion of Mistral’s two-tier access policy.
- Token Pricing and Agent Run Economics [08:23 - 12:02]: Pricing charts and data tables dissecting list pricing ($1.36 input / $4.18 output per million tokens) versus per-task cost on Artificial Analysis ($1.13), highlighting cache-read costs ($0.14/M tokens vs. $0.006/M for DeepSeek V4.1 Flash).
- Hardware Sizing and Deployment Requirements [12:35 - 13:51]: Memory calculation breakdowns for running weights locally (2.1 TB at BF16, 1.05 TB at FP8, 525–590 GB at 4-bit/NVFP4), specifying DGX B200 or 8× H200 GPU configurations, and why it exceeds a 512 GB Mac Studio M5 Ultra.
- Enterprise Context and Compute Roadmaps [13:52 - 17:04]: Examination of Mistral’s Vibe chat app hosting Zhipu’s GLM-5.3, datacenters in France, and training compute timelines spanning Series B through Series D.
Claims & numbers
- Parameter Counts & Architecture: The presenter states Mistral Large 4 has 1.05 trillion total parameters, 49 billion active parameters per token (52 billion including embeddings and output layers), meaning ~4.7% of parameters are active per token, plus a 1.6B parameter vision encoder [01:14 - 01:48].
- Serving Context Window: The documentation advertises 1M context tokens, but the serving API on Mistral and OpenRouter is capped at 524,288 tokens (512k) [01:52 - 02:11].
- Release Timeline: Public preview launched with weights promised by October 27, 2026; license type is unannounced [02:18 - 02:27].
- Training Compute: Trained over approximately two months on 3,800 NVIDIA Grace Blackwell GPUs in European datacenters consuming ~10 MW [03:03 - 03:11, 16:03].
- Benchmark Scores:
- Artificial Analysis Intelligence Index: Mistral Large 4 scores 38.4, compared to Mistral Large 3 at 9.3; tied with GPT-6 Luna (max effort) at 38.1; behind DeepSeek-V4.1-Flash (39.5), GLM-5.3 Flash (41.8), Kimi K3 (43.6), GLM-5.3 (44.8), MiMo-V2.6-Pro (46.3), and Claude Opus 5.5 (57.6) [02:51 - 03:48].
- SciCode-Verified: Large 4 scores 91.8, trailing MiMo-V2.6-Pro (91.9), GLM-5.3 (92.5), and Qwen3.8 2.4T (93.8) [04:26 - 04:42].
- DeepSWE 1.1: Large 4 scores 62, compared to GLM-5.3 at 61 (and 66.9 at its August launch) [04:46 - 04:59].
- Omniscience Eval: Large 4 scored 25.8% accuracy with a 42% hallucination rate, compared to DeepSeek-V4.1-Flash (46.4% accuracy, 96.5% hallucination), GPT-6 Luna (43.8% accuracy, 76.7% hallucination), and Claude Opus 5.5 (66.2% accuracy, 58.4% hallucination) [05:40 - 06:03].
- Cybersecurity (CyberGym-E2E): Large 4 scores 82%; Claude Opus 5.5 refused 98.5% of tasks and GPT-6 Astra refused 100%. On non-refusing test DeepSec, Large 4 scored 16% versus Opus 5.5 at 27% and Astra at 37% [06:31 - 07:26].
- Pricing & Economics:
- List price is $1.36/M input tokens and $4.18/M output tokens, with a temporary 50% launch discount ($0.68 / $2.09) [08:24 - 08:35].
- Full Artificial Analysis task suite cost is $1.13 per task for Large 4, compared to $0.27 for DeepSeek-V4.1-Flash and $0.07 for GPT-6 Luna [08:44 - 08:58].
- Caching read rate is $0.14/M tokens for Mistral Large 4, compared to $0.006/M for DeepSeek-V4.1-Flash (23× cheaper) and $0.0036/M for MiMo-V2.6-Pro [10:39 - 10:56].
- For full-sized flagship open models, Large 4 ($1.13 or $0.57 discounted) is cheaper per task than GLM-5.3 ($2.01), Kimi K3 ($2.00), and Qwen3.8 2.4T ($2.16), but more expensive than MiMo-V2.6-Pro ($0.13) [11:25 - 12:02].
- Hardware Footprint: Unquantized BF16 weights require 2.1 TB; FP8 requires 1.05 TB; 4-bit requires ~525 GB (or ~590 GB with NVFP4 scaling factors) before KV cache [12:38 - 12:55].
- Current RL Training Scale: Mistral’s ongoing reinforcement learning run operates on ~3,000 GPUs producing 33 billion tokens/day (16 billion trainable completion tokens) [16:20 - 16:29].
Notable quotes
- [06:05] “In other words, the model from Mistral knows less, but it bluffs less.”
- [11:14] “Le Chonk is not the chatty one; it’s just that every token it touches costs more.”
- [17:16] “As a model, it’s good—a bit behind the Chinese leaders, and genuinely a big step for Mistral. As a product for developers today, it is not that good...”
Assessment
This is an independent technical review and commentary video analyzing third-party benchmarks, provider pricing cards, and deployment considerations for Mistral Large 4. The presenter does not run original benchmarks on camera, relying instead on published data from Artificial Analysis, Hugging Face repository cards, and official interviews.
Described by gemini-3.8-flash on 2026-10-10 from the video’s audio and frames.