Launch video for an inference startup in one minute for about $2 (Deedy, X video)
Deedy (@deedydas) · 2026-09-23 · ai-made · 344,828 views
Made by AI
Model: Claude Opus 5.5 · Series: Code-rendered film: the "15-second showreel" prompt
Evidence: X post (2026-09-23): 'Opus 5.5 is incredible at instructional video generation. / / I made this launch video for a inference startup in 1min for ~$2. Videos like these used to take weeks if not months and a lot of coordination with agencies and 1000x the costs. / / Humans broadly prefer video to text. This changes the substrate of communication. These videos actually help communicate technical ideas in seconds (photorealistic video gen like ...'
Human role: Prompt, published: "make a modern slick and punchy video for a modern startup that works on inference". About one minute and ~$2 per the poster.
Pipeline: One prompt → Opus 5.5 → 26-second kinetic word-cloud launch video
Lore: one-prompt, budget-receipts
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This video is a sleek, AI-generated concept launch promo for a fictional/speculative AI inference startup named muda, shared by Deedy (@deedydas) to showcase rapid AI production capabilities created in minutes for approximately $2. The spot uses minimalist technical design, animated typography, and data visualisations to dramatise the elimination of latency and resource waste during LLM inference.
What is shown
- [00:00 - 00:02] A chatbot user interface receiving the prompt: "Summarize the quarter in one line." A progress spinner reads "Thinking..." as an elapsed-time counter ticks past 4.31 seconds.
- [00:03 - 00:04] Large kinetic bold text flashing the words "WAITING" and "WASTE." over a royal purple background.
- [00:05 - 00:08] A cluster grid diagram displaying GPU memory and compute inefficiencies labelled IDLE, COLD START, PADDING, and QUEUE with 31% utilization, which rapidly flips to full multi-color tile occupation reaching 94% utilization under the heading "Not anymore."
- [00:09 - 00:12] A streaming token waterfall displaying inference stages (prefill, attention, speculative, draft, kv, verify) as performance gauges accelerate from 7 tok/s (900 ms first token) up to 12,480 tok/s (38 ms first token).
- [00:13 - 00:16] Lexicographical definition screen with a brushed Japanese ensō circle: "muda (無駄) — waste; futility; uselessness," citing the Toyota Production System (muda, mura, muri).
- [00:17 - 00:20] Kinetic statements: "We removed the waste. It was most of it." with Japanese text (無駄をなくす), glitching into a transition.
- [00:21 - 00:23] The brand lockup: "muda." with the tagline "Inference, minus the waiting."
- [00:23 - 00:26] The closing "receipt": comparing the film runtime (26.000 s) to the answer latency (0.038 s), concluding: "We were done before the logo was."
Claims & numbers
- Film cost & creation: Produced in approximately 1 minute for about $2 (according to creator metadata).
- GPU Cluster Utilization: Depicted shifting from an initial 31% (dominated by cold start, padding, idle, and queue) to 76% and 94% utilization.
- Inference Latency & Throughput:
- Initial baseline: 7 tok/s throughput, 900 ms time-to-first-token (TTFT).
- Mid-stage acceleration: 2,419 tok/s (110 ms TTFT) and 10,799 tok/s (44 ms TTFT).
- Peak performance claimed: 12,480 tok/s throughput and 38 ms time-to-first-token (0.038 s).
- Film duration: Stated as 26.000 seconds on screen.
Notable quotes
- "Most of inference is waiting." [00:05]
- "We removed the waste. It was most of it." [00:17]
- "We were done before the logo was." [00:24]
Assessment
This is a demonstration mockup and tech-marketing parody/proof-of-concept created to show how generative design tools and code-driven animation workflows can assemble a high-production-value tech launch video in seconds for virtually no budget. All metrics, benchmarking numbers, and corporate branding represent staged marketing visuals rather than verified physical cluster benchmarks.
Lyrics & themes
- The video is instrumental, featuring crisp percussive UI sound design, clicks, whooshes, rising pitch risers, and glitch pulses rather than spoken vocals or song lyrics.
- The kinetic text conveys themes of technical efficiency, eliminating idle GPU capacity, lean engineering philosophy, and real-time generation speed.
Lore & references
- Muda (無駄): Directly cites the Japanese manufacturing concept of muda (waste), one of the three wastes (muda, mura, muri) identified in the famous Toyota Production System (TPS) / Lean manufacturing methodology.
- Ensō (円相): The brushstroke Zen circle shown at [00:14] symbolises elegance, minimalism, and focus.
- LLM Serving Overheads: Accurately references core challenges in modern high-throughput inference serving, including prefill vs. decode bottlenecks, KV cache management, speculative decoding, cold starts, and batch queueing delays.
Visual style & craft
- Built in a high-modernist Swiss typography and dark-on-white tech aesthetic with subtle monospace HUD markers (
1920x1080 30 FPS 120 BPM, frame stamps, and crosshair grids). - Employs code-driven/vector motion graphics (such as Remotion or programmatic CSS/canvas rendering), paired with glitch shaders and precise audio-reactive beat syncing.
Described by gemini-3.8-flash on 2026-09-30 from the video's audio and frames.