Post-Cutoff

Review

Mistral is BACK! (Le Chonk)

Matthew BermanYouTube101,242 views as of 8 October 2026

Watch on YouTubePlay loads YouTube’s player from youtube-nocookie.com.

Why it is here

Matthew Berman’s launch review of Mistral Large 4 (‘Le Chonk’). ~101k views. Length 21:29.

Description

Description written by Gemini from the videoGemini 3.8 Flash, 8 October 2026

Summary Matthew Berman reviews the release and preview availability of Mistral Large 4 (nicknamed “Le chonk”), a 1-trillion-parameter open-weight mixture-of-experts model created by French AI lab Mistral AI. He analyzes the model’s specs, pricing, benchmarks across coding and cybersecurity, and tests its performance on an agentic coding task building a 3D Rubik’s Cube simulation.

What is shown

  • 00:24 — Mistral AI announcement blog post: “Le chonk. Introducing Mistral Large 4” (dated October 6, 2026).
  • 01:10 — Model specifications displayed: 1-trillion total parameters, 49 billion active parameters, natively multimodal mixture-of-experts (MoE) architecture.
  • 02:54 — API pricing table on Mistral Studio: $1.36 per million input tokens, $4.18 per million output tokens.
  • 05:40 — Artificial Analysis coding benchmarks: DeepSWE 1.1 (ML4 scores 62) and Terminal-Bench 4.0 (scores 33).
  • 06:41 — Cybersecurity benchmark charts: Artificial Analysis Cyber Index (ML4 scores 50, tied for first among open models) and CyberGym-E2E (ML4 scores 82, #1 open model).
  • 07:03 — AutomationBench (Agentic behavior) chart: ML4 scores 59.9.
  • 07:41 — Vals.ai Harvey’s Legal Agent Benchmark chart: ML4 scores 14.6%.
  • 08:11 — Infrastructure details on the blog post: trained from scratch in Europe on 3,800 NVIDIA Grace Blackwell GPUs.
  • 12:25 — Artificial Analysis Intelligence Index ranking: ML4 ranks 25th overall with an index of 38, compared to top closed-source models.
  • 14:08 — Intelligence Index vs. Cost per Task chart, mapping ML4 against frontier proprietary models.
  • 16:12 — Context window comparison chart showing ML4 at 524,288 tokens (524K).
  • 16:45 — Output speed benchmark showing ML4 running at 116 tokens/second.
  • 17:43 — Hands-on test in an agentic coding environment using the ML4 preview API with high thinking effort to generate a Three.js Rubik’s Cube web application.
  • 19:01 — The generated Rubik’s Cube UI running in browser: inspecting rotation, scramble errors, and color texture rendering glitches during auto-solve.

Claims & numbers

  • Parameters & Architecture: The presenter states ML4 is a 1-trillion-parameter model with 49 billion active parameters using an MoE architecture.
  • Pricing: API pricing is $1.36 per million input tokens and $4.18 per million output tokens.
  • Hardware & Sovereignty: Berman notes the model was trained from scratch in Europe on 3,800 NVIDIA Grace Blackwell GPUs in Mistral’s own data center.
  • Context Window: Berman states ML4’s context window is 524,288 tokens (~0.5M tokens), compared to the 1M standard of competitors.
  • Speed: Berman highlights that Artificial Analysis clocks ML4’s output speed at 116 tokens/second.
  • Release Timeline: Berman notes the preview API is available immediately on Mistral Studio, with model weights scheduled to drop by the end of October 2026.
  • Benchmarks:
    • DeepSWE 1.1: 62 (second among open-source models shown, behind Kimi K3 at 68).
    • Terminal-Bench 4.0: 33 (behind GLM-5.3 at 40).
    • AA Cyber Index: 50 (tied for top open-weights model with GLM-5.3-Flash).
    • AutomationBench: 59.9.
    • CyberGym-E2E: 82 (first place among open-weights models).
    • Artificial Analysis Intelligence Index: 38 (25th position out of 691 evaluated models).

Notable quotes

  • 00:00 — “Mistral did it. They have a frontier model baked completely in Europe, from scratch, not built on top of a Chinese open-source model.”
  • 08:13 — “This model was completely conceived, designed, created, trained, and the inference is running from Europe.”
  • 10:18 — “Unless there is a plug-and-play option for open-source models, open source is going to flounder.”

Assessment This is an independent product review and benchmark breakdown featuring a real hands-on demo of the Mistral Large 4 API in an agentic coding environment. The presenter honestly documents difficulties with context-limit handling and output bugs in the Rubik’s Cube simulation rather than presenting a curated or cherry-picked success.

Described by gemini-3.8-flash on 2026-10-08 from the video’s audio and frames.

Related

  1. Model releases 98 days after the cutoff

    Mistral releases Mistral Large 4 (“le Chonk”)