"What is a Transformer": 12-minute Chinese explainer made by Claude Code + Opus 5.5 (X video)
宝玉 (@dotey) · 2026-09-26 · ai-made · 187,728 views
Made by AI
Model: Claude Opus 5.5 · Series: Code-rendered film (LLM writes the program that draws every frame)
Evidence: X post (2026-09-26): '《什么是 Transformer》由 Claude Code + Opus 5.5 制作 / --- 提示词 --- / 帮我用js制作一个视频,主题是:什么是 Transformer / 要深入浅出,让高中生也能看得懂,不仅high level说的清楚,也要有细节,包括注意力机制,甚至一些数学概念 / 你可以用任何工具或者安装工具,可以联网检索 / 请给我惊喜'
Human role: Creator gave one prompt (published in the post): make a JS video explaining Transformers for high-school students, any tools allowed, "surprise me".
Pipeline: Prompt → Claude Code + Opus 5.5 → JavaScript-rendered explainer (≈12 min)
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Here is the catalogued entry for the video:
Summary
This video is a comprehensive 12-minute Chinese-language educational explainer titled "What is a Transformer" (什么是 Transformer), created by 宝玉 (@dotey) using Claude Code and Claude Opus 5.5 to programmatically render animations via Remotion, React, and KaTeX, paired with neural TTS narration. It breaks down the internal mechanics of the Transformer architecture using basic high-school math concepts, progressing through tokenization, word embeddings, dot-product attention, multi-head attention, positional encoding, and next-token prediction.
What is shown
- [00:00–00:30] Introduction: Contextualizing common AI tools (chatbots, translation software, code assistants) as being powered by the Transformer architecture introduced in the 2017 Google paper "Attention Is All You Need".
- [00:31–01:01] Linguistic Riddle: Demonstrating coreference resolution with two nearly identical sentences: "The cat didn't jump on the table because it was too tired" vs. "...because it was too high", explaining that context dictates what "it" refers to.
- [01:02–01:40] Prior Architectures (RNNs): Visualizing the sequential bottleneck and catastrophic forgetting ("vanishing context") in Recurrent Neural Networks compared to parallel attention across all tokens.
- [01:41–02:05] Overall Pipeline: Presenting the 3-step high-level architecture: Text to Vectors $\to$ Attention & Feed-Forward Layers $\to$ Next-token Prediction.
- [02:06–03:04] Step 1: Tokenization & Embeddings: Splitting text into tokens, mapping them to dense vectors, and illustrating geometric relationships (e.g., King – Man + Woman $\approx$ Queen).
- [03:05–03:49] Mini Math Lesson: Dot Product: Explaining vector dot products geometrically and algebraically as a similarity scoring function ($a \cdot b = |a||b|\cos\theta$).
- [03:50–04:49] $Q, K, V$ Attention Mechanisms: Illustrating Queries, Keys, and Values through a library catalog metaphor, showing projection matrices transforming token vectors into $q, k, v$.
- [04:50–06:30] Step-by-Step Manual Attention Calculation: A complete numerical walkthrough calculating attention scores for "it" relative to "cat" and "table": Dot product score $\to$ Scaling by $\sqrt{d}$ $\to$ Softmax normalization $\to$ Weighted sum of $V$.
- [06:31–06:56] Full Attention Formula: Putting the steps together into the canonical formula $\text{Attention}(Q,K,V) = \text{softmax}\left(\frac{QK^T}{\sqrt{d_k}}\right)V$.
- [06:57–07:34] Why Divide by $\sqrt{d_k}$?: Demonstrating standard deviation scaling to prevent softmax saturation and vanishing gradients in high dimensions.
- [07:35–08:10] Multi-Head Attention: Splitting into parallel attention heads to capture diverse linguistic relationships (grammar, references, global context).
- [08:11–08:44] Positional Encoding: Explaining permutation invariance issues ("dog bites man" vs. "man bites dog") and introducing sinusoidal wave position vectors analogous to clock hands.
- [08:45–09:34] Complete Transformer Block & Stacking: Integrating Multi-Head Attention, residual connections, layer normalization, and feed-forward networks (FFN), stacked across layers (e.g., 6 layers in the 2017 paper, 96 layers in GPT-3).
- [09:35–10:26] Autoregressive Generation & Causal Masking: Visualizing next-token sampling with softmax probabilities and triangular causal attention masking.
- [10:27–11:15] Training via Next-Token Prediction & Gradient Descent: Showing cross-entropy loss ($-\log p$) and gradient descent parameter optimization across 175 billion parameters.
- [11:16–12:12] Conclusion & Epilogue: Solving the original riddle (showing how attention weights dynamically point "it" to "cat" or "table"), summarizing the core philosophy, and ending with technical production credits.
Claims & numbers
- The presenter notes the Transformer was introduced in 2017 by Google researchers in the paper "Attention Is All You Need" [00:10].
- In RNNs, as sequence length grows, the retained information from the start of a sentence can degrade to under 0.4% [01:19].
- Modern frontier models map each token to vectors with thousands of dimensions (e.g., GPT-3 uses 12,288 dimensions per token) [02:54].
- In scaled dot-product attention, dividing by $\sqrt{d}$ (where $d$ is commonly 64 or 128 in real models) prevents dot products from scaling excessively and causing extreme, saturated softmax distributions [07:01].
- The original 2017 Transformer stacked 6 layers, whereas GPT-3 stacked 96 layers with 175 billion parameters [09:24, 11:08].
- The end credits state the video was programmatically generated using JavaScript (Remotion + React + KaTeX) and voiced via neural TTS [12:10].
Notable quotes
- [00:19] "翻译过来:注意力,就是你所需要的一切。" (Translated: Attention is all you need.)
- [03:42] "点积,就是一台‘相似度打分机’。记住它——注意力机制,全靠它。" (The dot product is a 'similarity scoring machine'. Remember it—the attention mechanism relies entirely on it.)
- [11:53] "它的核心思想,其实只有一句话:让每个词,去关注真正重要的词。" (Its core idea is really just one sentence: let every word pay attention to the words that truly matter.)
Assessment
This is an authentic, exceptionally polished educational explainer video programmatically authored by developer 宝玉 (@dotey) leveraging Claude Code and Claude Opus 5.5 to write Remotion/React code. The visual animations are deterministic, code-rendered geometric diagrams and mathematical layouts, precisely synchronised with synthesized Mandarin narration without factual hallucinations or exaggerated hype.
Lyrics & themes
The video is an expository technical narrative rather than a musical track. Key thematic lines include:
- [00:23] "今天,我们只用高中数学,把 Transformer 从里到外,彻底拆开看一遍。" (Today, using only high school math, we'll thoroughly dismantle and examine the Transformer inside out.)
- [01:34] "Transformer 的想法很大胆:不排队了!让每个词,一眼看到所有的词。" (The Transformer's idea was bold: no more queuing! Let every word see all words at a single glance.)
- [08:57] "打个比方:注意力是开会讨论,前馈网络是回到座位独立思考。" (To make an analogy: attention is having a meeting to discuss; feed-forward is returning to your desk to think independently.)
- [12:03] "此刻,正有亿万次点积,在为你计算注意力。" (At this very moment, hundreds of billions of dot products are computing attention for you.)
Lore & references
- "Attention Is All You Need" (Vaswani et al., 2017): Cited directly as the foundation of modern LLMs.
- Winograd Schema / Ambiguity Resolution: Using the classic "The animal didn't cross the street because it was too tired/wide" linguistic benchmark adapted to a cat jumping on a table.
- Word2Vec Classic Analogy: Visualizing $King - Man + Woman \approx Queen$ in a 2D vector space embedding demo.
- GPT-3 Parameters: Referencing GPT-3's 96 layers, 12,288 hidden dimension, and 175B parameters as canonical scale examples.
- Causal Masking: Visualized as a triangular mask ("paper covering answers during an exam") to enforce autoregressive generation constraints.
Visual style & craft
The video features a clean, high-contrast dark-mode interface aesthetic (deep navy background, pastel neon accents, crisp typography, and KaTeX mathematical notation). It is explicitly built using Remotion (a programmatic React video framework), meaning every chart, coordinate plane, vector arrow, matrix transformation, and animated text element was rendered via code rather than generated by text-to-video diffusion models. Dynamic motion includes smooth easing curves, coordinate transformations, and synchronized highlighted tokens.
Described by gemini-3.8-flash on 2026-10-05 from the video's audio and frames.