GPT 6 Astra is a freak
AI Search · 2026-09-06 · review · 1,528,767 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
The presenter from the channel AI Search reviews OpenAI's newly released frontier model, GPT-6 Astra, running it through a broad battery of demanding practical tests across coding, 3D modeling, computer use, game design, and research. He also examines GPT-6 Astra's reported benchmark performances across math, vision, agentic workflows, and game playing against competitors like Claude Fable 5.1 and Gemini.
What is shown
- Raw WebGL2 Physics Simulation [00:47–04:09]: GPT-6 Astra writes an interactive, zero-dependency ray-traced water balloon bullet impact simulation in WebGL2, self-critiqued across 3 rounds via a critic agent prompt over 32 minutes.
- Unreal Engine 3D Game Generation [04:48–09:38]: GPT-6 Astra connects to Blender and Sketchfab via MCP and Mixamo to assemble an Imperial Chinese palace rooftop parkour game in Unreal Engine, iteratively scoring and refining assets and camera effects over ~4 hours.
- Computer Use Drawing [10:00–10:44]: Using browser-based computer control, GPT-6 Astra opens Photopea to manually paint a street scene stroke by stroke over 20 minutes.
- Pixel Art Sprite Generation [10:45–11:37]: The model uses computer control in SpritePaint to draw frame-by-frame 2D sprite animations (run, jump, sword slash) in 40 minutes.
- Live Virtual Piano Performance [11:38–13:53]: Using computer control on an online virtual keyboard, GPT-6 Astra composes a Chopin-style solo, figures out how to record, and clicks keys in real time to perform a 1-minute piece in ~7.5 minutes.
- Sponsor Segment (Luma) [13:54–15:26]: Showcase of Luma as an agentic creative workspace, generating product videos and marketing assets.
- DAW Music Composition [15:27–18:00]: GPT-6 Astra sequences, mixes, and automates a multi-track Europop EDM song in Waveform DAW using local VSTs and external audio samples.
- Airbnb 3D Recreation [18:01–19:48]: Given photos from a Japan Airbnb listing, the model reconstructs the full villa in Blender via MCP and renders a cinematic camera tour over 1 hour 12 minutes.
- AI Commercial Creation [20:01–21:14]: Directs Higgsfield MCP to generate, edit, voice, and stitch a 30-second commercial for Tenzo matcha tea in ~20 minutes.
- Minimalist Explainer Animation [21:15–23:15]: Generates a Python/Manim-style black-and-white motion graphics video explaining Eratosthenes' calculation of Earth's circumference with Gemini TTS narration.
- Failure Cases (FrogBench & Tumor ID) [23:16–25:10]: The model fails to spot camouflaged frogs on two leaf litter photos (hallucinating a rattlesnake and a toad), and correctly identifies only 1 out of 6 brain tumor CT/MRI scans.
- Research and Ideation [25:11–27:09]: Demonstrates deep research on Alzheimer's amyloid-beta/tau propagation (with tables and flowcharts) and proposes 3 automated factory farm animal welfare systems.
- Benchmarks and Qualitative Feats [27:18–32:20]: Highlights benchmark charts (FrontierMath Tier 4, Terminal-Bench 4.0, OSWorld 2.0, ARC-AGI-3, VoxelBench, LiveBench, Arena WebDev) and shows gameplay logs of GPT-6 Astra autonomously beating Portal and speedrunning Pokémon FireRed in 18 hours 12 minutes.
Claims & numbers
- The presenter says GPT-6 Astra is available in the ChatGPT desktop app (formerly Codex) and the web interface under the "Work" tab, but not yet in standard web chat [00:30, 23:25].
- The presenter states that a 1-hour complex agent run consumed roughly 3% to 4% of his weekly usage quota on the ChatGPT Pro ($100/mo, 5x limit) tier [04:26, 24:44].
- The presenter claims GPT-6 Astra scored 97.6% accuracy on FrontierMath Tier 4 (v2) at $1.20 API cost [27:28].
- The presenter reports GPT-6 Astra reached 67.9% on Terminal-Bench 4.0, 41.4% on AutomationBench, and #1 on VoxelBench with a 2652 rating (97.3% win rate) [27:32, 27:42, 30:34].
- The presenter reports GPT-6 Astra is the only general agent to beat the game Portal, and that it beat Pokémon FireRed in 18 hours and 12 minutes vision-only [28:48, 29:16].
- On ARC-AGI-3, the presenter states GPT-6 Astra scored 62.7% standalone and close to 100% with a provider adapter harness [30:20].
- On FrontierMath Erdős, the presenter says GPT-6 Astra scored 2.9% while competitors scored 0.0% [30:54].
- On the Artificial Analysis Intelligence Index, the presenter notes GPT-6 Astra scored 55 at $2.57 per task and 71 output tokens/sec, trailing Claude Fable 5.1 with fallback (57 at $6.12) [31:30–31:40].
- On AA-Omniscience hallucination rate, GPT-6 Astra had a 47% hallucination rate, lower than Claude Opus 5 (64%) and GPT-5.6 (82%) [31:49].
- On LiveBench, the presenter shows GPT-6 Astra placed 3rd overall with an 82.2 score, behind Claude Fable 5.1 (83.4) and Claude Fable 5 (83.0) [32:01].
Notable quotes
- [00:01] "Folks, I can finally say we got GPT-6 before GTA 6."
- [02:08] "All right, here's what I got. It first worked for 32 minutes, and here's the cool part about it..."
- [33:01] "This is definitely the most capable and powerful model you can use right now."
Assessment
This is an independent user review and hands-on demonstration video testing real model capabilities using complex prompts and MCP integrations. While the computer-use and generation runs are timelapsed and edited for video pacing, the presenter shows both successes and clear failure modes (e.g., failing FrogBench and CT scan analysis).
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.