GPT-6 Astra Made This Entire Video
Nate Herk | AI Automation · 2026-09-04 · ai-made · 453,240 views
Made by AI
Model: GPT-6 Astra · Series: "Made This Entire Video By Itself"
Evidence: Description: 'I gave GPT-6 Astra one open-ended prompt and asked it to take me from an idea to a finished YouTube video. It researched real Astra demos, captured the source posts, used my HeyGen avatar and ElevenLabs voice clone, and built the edit in HyperFrames.'
Human role: One open-ended prompt; explains the prompt at 3:12.
Pipeline: GPT-6 Astra → research + screenshots → HeyGen avatar + ElevenLabs voice → HyperFrames edit
Lore: made-it-by-itself, one-prompt
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
YouTuber Nate Herk demonstrates an end-to-end YouTube video generated autonomously by OpenAI’s GPT-6 Astra from a single prompt. The embedded video features an AI avatar and voice clone of Herk presenting community demos of GPT-6 Astra before detailing how the model wrote, directed, edited, voiced, and proofed the entire piece. Herk then shows the exact prompt used, along with the compute logs, run time, and API cost breakdown.
What is shown
- [00:00] Real Nate Herk introduces the experiment where a single prompt instructed Astra 6 to build a full YouTube video.
- [00:05] The generated video starts, fronted by an AI digital avatar (HeyGen Avatar V5) speaking with Herk’s cloned voice (ElevenLabs).
- [00:20 - 01:48] Showcase of GPT-6 Astra community projects captured and narrated by the agent:
- Matt Shumer’s Unreal Engine Manhattan city environment constructed over a week of sustained work [00:26].
- Riley Brown’s playable Call of Duty-style shooter modified interactively between matches [00:48].
- Flavio Adamo’s one-shot Minecraft-style world demo featuring block crafting and mining [00:59].
- Tom Krcha’s 3D steam train reconstruction in Blender from an old technical drawing, yielding 3,295 editable parts [01:14].
- Yunfan Ye’s architectural 3D walkthrough (349 Walsh Road) generated from listing photos [01:25].
- Daniel Ch’s animated UI motion design clip generated in 14 minutes [01:38].
- [01:49 - 02:47] Astra explains its autonomous production process: gathering X posts via computer use, splitting narration into 8 voice clips, animating the avatar in HeyGen, aligning 72 shots and camera moves inside HyperFrames, and running an automated transcription-verification loop.
- [03:12 - 04:38] Real Herk returns to display the exact prompt in the Astra 6 chat UI [03:17] and opens the Codex session inspector [04:03] detailing API token usage and runtime.
Claims & numbers
- Release date: OpenAI released GPT-6 Astra on September 3, 2026, featuring long-running tasks and native computer use (stated by the AI presenter at [01:49]).
- Community project stats:
- Matt Shumer's Unreal Engine city ran over the course of a week [00:35].
- Tom Krcha’s Blender steam train contained 3,295 fully editable objects [01:19].
- Daniel Ch's motion video took 14 minutes of generation plus 2 manual revisions [01:40].
- Production specs of the generated video: 72 total shots, 8 narration audio segments, 6 creator demos, rendered at 1080p, 30 fps, with a 3:07 duration [00:13, 02:10, 02:46].
- Generation cost & runtime:
- The autonomous production run took 47 to 50 minutes of compute time [04:04, 04:22].
- Token consumption: 3.25 million uncached input tokens ($16.24), 20.81 million cached input tokens ($26.01), and 0.94 million output/reasoning tokens ($17.52) [04:04].
- Standard API cost was $59.77 ($118.84 at Fast/Priority API rates), excluding external HeyGen and ElevenLabs fees [04:04, 04:16].
Notable quotes
- [00:05] "I'm Astra 6. You're looking at Nate Herk's avatar, speaking with his voice clone. I made this video."
- [02:45] "That's how I get from an idea to a file you can use."
- [03:12] "I gave Astra this one prompt, and this is what I got back... that is absolutely crazy."
Assessment
A legitimate demonstration of GPT-6 Astra's autonomous multi-step agentic capabilities integrating third-party tools (HeyGen, ElevenLabs, HyperFrames). While the generation relied on existing pre-authorized credentials and project assets supplied in Herk's environment, the end-to-end orchestration, visual alignment, and verification steps are genuine outputs of the agent.
AI Production Details
Lyrics & themes
The narration is an informational script structured as an AI agent delivering an expository portfolio video:
- Introduction [00:05 - 00:20]: Self-identification as Astra 6 and breakdown of production tasks ("I found the footage, captured the posts, wrote the script, and built the edit." [00:10]).
- Showcase of user creations [00:20 - 01:48]: Chronicling external builders pushing Astra's multi-step loops across Unreal Engine, game development, 3D modeling, and motion graphics ("Inspect a scene, make changes, and check the result." [00:44]).
- Workflow & self-audit [01:49 - 02:47]: Outlining computer use, modular timeline sequencing in HyperFrames, and QA checks ("I also transcribed the finished audio and compared it with the script." [02:37]).
- Sign-off [02:59 - 03:09]: Direct address calling for user challenges ("Nate directed. I produced... What would you have me build?" [02:59]).
Lore & references
- Agent Video Genre: Directly participates in the "AI model made this whole video" format that expanded across tech channels in mid-2026.
- Computer Use & Tool Chaining: Highlights browser inspection on X (formerly Twitter), programmatic video assembly in HyperFrames, voice synthesis via ElevenLabs, and video synthesis via HeyGen Avatar V5.
- AI Community Personalities: Highlights public demos shared on X by recognized AI builders and founders, including Matt Shumer, Riley Brown, Flavio Adamo, Tom Krcha, Yunfan Ye, and Daniel Ch.
Visual style & craft
- Style: Clean, modern tech aesthetic using Apple/Windows-style UI card mockups, kinetic typographic callouts ("Found.", "Captured.", "Written.", "Edited."), timeline diagrams, and floating UI windows against a stylized blue abstract desktop background.
- Craft & Execution: Highly polished code-composed motion graphics (HyperFrames phrase-aligned composition) synced precisely to audio stems. Transitions, zooms, and B-roll cut-ins are frame-accurate to voice pauses. The talking-head avatar exhibits HeyGen V5 synthetic lip-syncing and head motion framed in a studio camera setup.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.