Mistral Is BACK – Mistral Large 4 First Test (Le Chonk!)
Bijan BowenYouTube57,594 views as of 10 October 2026
Description
Description written by Gemini from the videoGemini 3.8 Flash, 10 October 2026
Summary Bijan Bowen tests Mistral AI’s newly announced frontier model, Mistral Large 4 (“Le Chonk”), released as a public preview on October 6, 2026. The presenter inspects the announcement blog post, Hugging Face repository, pricing, and benchmarks before running the model through multiple coding and game-generation tasks using Mistral’s Vibe CLI harness. Across tasks ranging from a browser OS to 3D Godot and C++ games, he evaluates its coding performance, speed, and multimodal capabilities.
What is shown
- [00:12] Mistral’s announcement blog post introducing “Le chonk / Introducing Mistral Large 4” with benchmarks across cybersecurity, SWE-bench/DeepSWE, agentic workflows, and math.
- [01:17] Hugging Face repo page for
Mistral-Large-4.0-1T05-A52Bshowing a countdown clock to the open-weights release on October 31, 2026. - [03:25] Mistral docs showing model specifications: 1.05T total parameters, 49B active parameters (52B including embeddings/output), 1.6B vision encoder, and API pricing.
- [04:43] Mistral’s Vibe CLI interface running local and agentic code generation workflows.
- [05:40] Testing “OmniOS”, an HTML/JS browser OS generated by the model featuring desktop wallpapers, a clock/calendar, mock system indicators, text editor, terminal with neofetch, settings app, and custom app sandbox builder (“OmniStudio”) [10:28].
- [11:36] In-browser execution of a 3D Subway FPS coded via OmniStudio.
- [13:03] Evaluation of a self-contained C++ OpenGL skateboarding game (“Skate NYC”), testing movement, jumping, and grinding mechanics [14:00].
- [15:01] Attempting to restyle the skateboarding game using multimodal image prompts based on Sonnet 5.5 reference screenshots; encountering issues where the model initially extracted colors as text instead of reading image pixels directly [16:13], followed by fixing image support with Claude’s assistance [16:32].
- [18:00] Reviewing a Godot 4-based multiplayer game (“Backyard Pool Party”) featuring animated diver models and cannonball physics.
- [20:15] Inspecting a 3D product landing page for “Slappis Watch Co.” generated with Three.js, testing an interactive exploded view of watch components [21:20] and a custom edition featuring an image texture on the watch face [22:16].
- [23:22] Testing a Three.js subway zombie survival FPS (“Last Stop”), identifying missing weapon firing logic [24:00], submitting a prompt fix, and verifying shooting mechanics and wave progression [24:43].
- [27:37] Navigating an interactive Three.js 3D model of Jerry Seinfeld’s apartment with camera views and asset placements.
- [29:50] Testing “Chrono Block”, an interactive 3D city block simulation built in Godot 4 showcasing architectural and transport evolutions across five time periods (1945 to 2055) [30:10 - 34:00].
- [34:57] Testing a 3D demolition derby game generated in Godot from an image prompt, featuring car soft-body crash simulations and multiple camera perspectives.
- [36:37] Mistral console usage page showing a total bill of $46.91 across 820 API requests for the testing session.
Claims & numbers
- Mistral Large 4 is an open-weight hybrid instruct- and reasoning MoE model with native multimodality, trained with a 1.05T total parameter count and 49B active parameters per token (52B including embeddings and output layers) (stated by the presenter and shown in documentation at [01:38] and [03:26]).
- Features a 1.6B parameter vision encoder (stated by presenter and Gemini at [04:20] and [04:33]).
- Weights are scheduled for open release on October 31, 2026 (shown on Hugging Face page at [01:17]).
- Standard pricing is listed at $1.36 per 1M input tokens and $4.18 per 1M output tokens (with a 50% discount active for the first two weeks of preview) (shown at [03:43]).
- On cybersecurity benchmarks, the blog post reports a 93% score on the Artificial Analysis Cyber Index, outperforming open models like GLM-5.3 and DeepSeek V4 Pro (shown at [00:47], [02:04], and [03:09]).
- DeepSWE 1.1 benchmark reports 61.7% on DeepSWE v11, 59.4% on SWE-Atlas-QnA, and 28.3% on Terminal-Bench 4.0, combining for a 49.8% Coding Agent Index score (shown at [02:45]).
- The model was trained from scratch on 3,800 NVIDIA Grace Blackwell GB100s in European data centers (shown in blog text at [38:03]).
- The presenter notes his testing run took 820 API requests and cost $46.91 (shown at [36:38]).
Notable quotes
- [00:20] “Mistral Large 4. So the funny part is, and why this is mentioned as this, there was a meme like a month or a couple months ago basically saying Mistral was going to come out with some like 50 trillion parameter model called like Le Chaton Fat or like the fat cat in French.”
- [10:44] “That is actually very creative. This is one of the more creative special features I think I’ve ever seen.”
- [38:41] “That’s going to conclude our first look and test of Le Chonk, a.k.a. Mistral Large 4.”
Assessment This is an independent hands-on testing review of Mistral Large 4’s preview API via Mistral’s Vibe CLI. The presenter runs genuine, unedited demonstrations directly in real time on an Ubuntu desktop, including debugging failures (such as syntax errors, broken firing logic, and initial image-reading issues in the CLI harness) to transparently evaluate the model’s capabilities and limitations.
Described by gemini-3.8-flash on 2026-10-10 from the video’s audio and frames.