A 3-minute history of AI, from "Attention Is All You Need" to AGI, rendered 100% in code by Opus 5.5 (X video)
Chubby (@kimmonismus) · 2026-09-23 · ai-made · 199,692 views
Made by AI
Model: Claude Opus 5.5 · Series: Code-rendered film (LLM writes the program that draws every frame)
Evidence: X post (2026-09-23): 'This 3-minute film was made 100% in code by Claude in Claude Code: ~7,400 lines of React/TypeScript (Remotion), every image drawn in SVG and Canvas, an open-source TTS voice, and a score synthesized in Python. No stock footage. No image generators.' Took ~1 hour and 7% of weekly limits.
Human role: Topic prompt; Claude Code did the build.
Pipeline: Opus 5.5 in Claude Code → ~7,400 lines Remotion (SVG/Canvas) + open-source TTS + Python-synthesized score → 3:00 film
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This animated historical retrospective traces the evolution of modern artificial intelligence from the introduction of the Transformer architecture in 2017 to the frontier reasoning and agentic models of today. Created by creator Chubby (@kimmonismus) and rendered programmatically in code using Claude Opus 5.5, the video illustrates major milestones in AI capabilities through minimalist vector animations and synthesized narration.
What is shown
- [00:00–00:15] Pre-Transformer NLP: Visualizes sequential RNN-style reading of sentences ("The animal didn't cross the street because it was too tired") where distant context and coreference ("it") are forgotten over time.
- [00:16–00:35] The Transformer (June 2017): Depicts all-to-all self-attention connecting every token simultaneously; displays the original "Attention Is All You Need" research paper title page, encoder-decoder architecture diagram, and scaled dot-product attention formula ($\mathrm{softmax}(QK^T / \sqrt{d_k})V$, 8 heads, 6 layers, $d_{model} = 512$).
- [00:36–00:54] Scaling & Self-Supervision (2018–2020):
- Displays parameter growth and scaling curves (Kaplan et al., 2020): GPT (117M, June 2018), BERT masked language modeling (340M, Oct 2018), GPT-2 unicorn prompt completion (1.5B, Feb 2019), and GPT-3 few-shot translation (175B, May 2020).
- Demonstrates next-token probability prediction over a vocabulary sphere.
- [00:55–01:19] Alignment & ChatGPT (2022):
- Illustrates RLHF preference ranking on prompt completions ("Explain gravity to a 6-year-old").
- Shows the ChatGPT interface launch on November 30, 2022, rapidly scaling to 1,000,000 users in 5 days over an illuminated globe.
- [01:20–01:39] Multimodality (2023): Demonstrates ViT patch tokenization ("an image is worth 16x16 words"), log-Mel spectrogram audio encoding, and Python attention code unification under the token abstraction.
- [01:40–02:05] Test-Time Compute & Reasoning (2024–2025):
- Analyzes why tokenization (
str,aw,berryincl100k_base) makes character counting difficult. - Visualizes chain-of-thought search ("Thought for 12 seconds") spelling out "strawberry" to identify indices 3, 8, and 9.
- Analyzes why tokenization (
- [02:06–02:27] Tool Use & Multi-Agent Swarms (2025): Depicts an autonomous agent loop invoking search queries, editing repository files, running terminal tests (
pytest -q, 42 passed), committing code to git, and delegating subtasks to parallel copies of itself (research,code,test,review). - [02:28–03:00] AGI Horizon (20??): Global network orchestration questioning whether full general intelligence has arrived, concluding with the title card: "Attention Is All You Need $\rightarrow$ AGI?".
Claims & numbers
- June 2017: "Attention Is All You Need" paper published (arXiv:1706.03762; Vaswani et al.), introducing the Transformer architecture with 8 heads, 6 layers, and $d_{model} = 512$.
- June 2018: Original GPT released by OpenAI with 117M parameters.
- October 2018: BERT released by Google with 340M parameters.
- February 2019: GPT-2 released by OpenAI with 1.5B parameters.
- May 2020: GPT-3 released by OpenAI with 175B parameters.
- November 30, 2022: ChatGPT launched, surpassing 1,000,000 users in 5 days.
- cl100k_base tokenization: The word "strawberry" splits into 3 tokens (
str,aw,berry). - Reasoning test: The model resolves the strawberry character-count question after thinking for 12 seconds.
Notable quotes
- [00:20] "Then every word saw every other word. Not one after another. All at once."
- [01:30] "Pixels, sound, code: all just tokens. Same attention. New senses."
- [02:46] "Attention was all it needed. What comes next needs ours."
Assessment
This is a polished, AI-generated retrospective animation rather than a commercial product demo or benchmark report. The historical chronology accurately reflects the foundational inflection points of transformer architecture, scaling laws, RLHF, multimodal tokenization, test-time compute, and agentic workflows.
Lyrics & themes
- Theme: The transition of artificial intelligence from rigid, forgetful sequential language parsers into unified, self-directing multimodal reasoning engines.
- Key spoken lines:
- [00:04] "Machines once read one word at a time. By the end, it forgot the beginning."
- [00:36] "Then we made it bigger. And bigger."
- [01:49] "Now it thinks before it speaks. The longer it thought, the better it got."
- [02:30] "It sees. It speaks. It thinks. It acts. Is this general intelligence?"
Lore & references
- Winograd Schema: Uses the classic challenge sentence "The animal didn't cross the street because it was too tired" to highlight the difficulty RNNs had resolving ambiguous pronoun references.
- Ovid's Unicorn: The GPT-2 prompt snippet "a herd of unicorns..." references OpenAI's famous demonstration of synthetic story generation in February 2019.
- ViT ("An image is worth 16x16 words"): The cat patch breakdown references Dosovitskiy et al.'s landmark Vision Transformer paper title.
- The Strawberry Test: Prominently displays "How many r's are in 'strawberry'?", referencing the popular benchmark problem that highlighted tokenization blind spots in early LLMs and inspired test-time reasoning models (OpenAI o1).
- Self-Improving Agents: Visualizes sub-agent replication and automated code refactoring, touching on recursive tool use and autonomous software engineering.
Visual style & craft
- Visual Style: Clean, high-contrast vector motion graphics on a dark celestial backdrop, using warm amber, gold, and neon accents with HUD-style telemetry cards.
- Craft: Rendered programmatically via code (SVG/Canvas/Manim-style graphics generated with Claude Opus 5.5). All elements—from mathematical formulas to interactive graph nodes, planetary meshes, and animated code editors—exhibit precise procedural timing, crisp typography, and particle effects synced to synthetic voiceover and atmospheric ambient audio.
Described by gemini-3.8-flash on 2026-10-05 from the video's audio and frames.