As of: 2026-10-09 19:24 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/v/two-minute-papers-deepseek-new-architecture/ # DeepSeek’s Insane New Architecture Two Minute Papers, 18 September 2026, YouTube. 337,830 views as of 9 October 2026. Kind: Review. Watch: https://www.youtube.com/watch?v=vIHw_2VjSUw ## Why it is here Two Minute Papers on the DeepSeek V4.1 Flash paper. ~338k views by 2026-10-09. Length 5:10. ## Description (written by Gemini from the video) **Summary** Dr. Károly Zsolnai-Fehér from the channel *Two Minute Papers* reviews the architecture and performance of DeepSeek-V4.1-Flash. He explains how the model achieves massive KV cache compression using Cross-layer Shared Attention (CSA2) and evaluates its code generation, visual comprehension, reasoning benchmarks, and fluid simulation reproductions compared to leading frontier models. **What is shown** - [00:02] Visual comparison of a rendered Saturn planet between GPT-6 Astra and DeepSeek 4.1 Flash. - [00:04] Fast live inference token generation solving a programming/algorithmic problem. - [00:11] Benchmark chart comparing DeepSeek-V4.1-Flash against Kimi-K3, GLM-5.3, Opus 5, and GPT-5.6-Sol on DeepSWE v1.1 and CyberGym. - [00:23] Web-based interactive 3D demos and games generated with DeepSeek-V4.1-Flash, including a flying broomstick game, a low-poly snowboarding game [00:31], and a 3D vector/particle field simulator [00:35]. - [00:47] Native visual reproduction demo: DeepSeek-V4.1-Flash recreating a *Warcraft III* menu interface based on an image input. - [01:07] Presentation of the research paper: *"DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression"*. - [01:10] Bar chart tracking global KV cache per token reduction across DeepSeek generations (V1 to V4.1-Flash). - [01:59] Diagrams explaining traditional layer-by-layer KV caching versus DeepSeek-V4.1-Flash's Causal Encoder and Decoder sharing a global KV memory cache via CSA2. - [03:10] Liquid honey coiling fluid simulation reproductions comparing the original SIGGRAPH paper output against implementations from GPT-6 Astra, Claude Opus 5.1, and DeepSeek-V4.1-Flash. - [03:57] Token counter and cost tracker animation displaying token usage scaling past 101 million tokens. - [04:36] Sponsor overview for Lambda AI cloud computing infrastructure. **Claims & numbers** - **Benchmark scores**: The presenter shows DeepSeek-V4.1-Flash scoring 74.2% on DeepSWE v1.1 (beating Kimi-K3 at 67.5%, GLM-5.3 at 66.9%, GPT-5.6-Sol at 73.0%, and Claude Opus 5 at 74.0%) and 88.1% on CyberGym (beating Kimi-K3 at 80.0% and GLM-5.3, Opus 5, and GPT-5.6-Sol at 84.5% each). - **Hardware cost**: Running DeepSeek-V4.0-Pro locally costs roughly $300,000 in hardware, whereas V4.1-Flash can run for approximately a quarter of that hardware cost. - **KV cache compression**: The global KV cache per token in DeepSeek-V4.1-Flash is reduced to 890 bytes—approximately 437× smaller than DeepSeek-V1 (389,120 bytes in November 2023) and ~4× smaller than DeepSeek-V4-Flash (3,514 bytes in April 2026). - **Architecture**: The presenter states DeepSeek-V4.1-Flash uses a technique called CSA2 (Cross-layer Shared Attention 2) with an encoder-decoder architecture that shares KV memory across its 40 layers. - **Model size**: DeepSeek-V4.1-Flash has more than 500 billion parameters. - **Limitations**: The presenter notes that the model burns significant token counts during deep reasoning (showing 101,257,961 tokens consumed for $1.48 USD) and currently underperforms GPT-6 Astra and Claude Opus 5.1 on complex scientific fluid physics code generation. **Notable quotes** - [01:06] *"Hold on to your papers, fellow scholars, it's small because of the KV cache."* - [02:13] *"DeepSeek 4.1 Flash finally gives us shared memory between the layers."* - [03:57] *"4.1 Flash likes to think a lot and burns a lot of tokens."* **Assessment** This is an independent scientific review and paper breakdown evaluating DeepSeek-V4.1-Flash's architecture and capabilities. The video demonstrates real community-generated WebGL/canvas outputs and benchmark charts, though the honey coiling reproduction and token burn rate illustrate both its current weaknesses and trade-offs. _Described by gemini-3.8-flash on 2026-10-09 from the video's audio and frames._ ## Related - 2026-09-10: [DeepSeek V4.1-Flash brings a new architecture family, native vision and a cheaper API](https://postcutoff.com/e/2026-09-10-deepseek-v4-1-flash/)