Did OpenAI actually build AGI? GPT-6 Astra first look
Fireship · 2026-09-04 · review · 4,184,432 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this episode of The Code Report, host Jeff Delaney rounds up a rapid-fire week of major frontier AI releases in early September 2026, highlighted by Anthropic’s Claude Fable 5.1 and Mythos 5.1, Meta’s Muse Spark 1.3, and OpenAI’s GPT-6 Astra. He recaps the chaotic rollout of GPT-6 Astra, its reported benchmark leaps, early access demos in 3D modeling and agent simulation, and contrasting independent evaluation results.
What is shown
- [00:11 - 02:23] Anthropic Claude Fable 5.1 & Mythos 5.1: Promotional materials and case studies showing Claude Fable 5.1 resolving a 5-year-old one-in-a-million crash for hedge fund Millennium by disassembling vendor code into assembly, Mythos 5.1 designing protein binders (Nipah G, EGFR, GDF-8), and generating Venusian elevation maps from 30-year-old NASA radar imagery.
- [02:24 - 03:03] Meta Muse Spark 1.3: Presentation of benchmark charts, pricing tiers ($1.25/$4.25 standard vs. $0.10/$0.20 contributor tier), and quotes from Mark Zuckerberg and Alexandr Wang.
- [03:04 - 04:19] The Outage and Launch Rollout: Coverage of the simultaneous multi-platform outage across ChatGPT, Claude, Grok, and Cursor; the premature launch page, 404 retraction, media embargo breaks (CNBC, The Verge), influencer early-access posts on X, and Sam Altman telling users to "go to bed."
- [04:20 - 05:36] GPT-6 Astra Capabilities & Benchmarks: Overview of Astra’s training on 100,000+ GPUs at the Texas Stargate site; demonstrations on OSWorld (filling Form 1040, spreadsheet puzzles, FreeCAD 3D transmission modeling); benchmark comparisons across ExploitBench, Terminal-Bench Science, FrontierMath Tier 4, and ARC-AGI-3.
- [05:37 - 06:36] Early Access Demos & Evaluations: Demos including Sharif Shameem’s Blender reconstruction of San Francisco's Palace of Fine Arts, Thomas Ricouard’s walkthrough of an Astra-modeled house in Unreal Engine 5, Matt Shumer’s simulation of talking Astra agents in Unreal Engine, and the Artificial Analysis Intelligence Index chart showing Astra scoring 61 alongside GPT-5.6 Sol.
- [06:37 - 07:21] CodeRabbit Security Demo: Walkthrough of CodeRabbit Security scanning codebases, showing reachability and blast radius assessments, and automated pull-request remediation.
Claims & numbers
- Anthropic Claude Fable 5.1 / Mythos 5.1:
- The presenter states Fable 5.1 solved a 5-year-old, 1-in-a-million crash for hedge fund Millennium by disassembling vendor binary code into raw assembly.
- The presenter notes Mythos 5.1 achieves ~50% hit rate in viable protein binder design across 12 targets, compared to the industry standard ~10-15%.
- Pricing is claimed at $10 per million input tokens and $50 per million output tokens.
- Meta Muse Spark 1.3:
- Mark Zuckerberg claimed Muse Spark 1.3 has frontier performance that is "almost too cheap to meter."
- Standard API pricing is stated as $1.25 input / $4.25 output per 1M tokens ($0.15 cached input); Contributor tier is $0.10 input / $0.20 output ($0.002 cached input) in exchange for Meta training on prompt data.
- Alexandr Wang claimed a "meaningful double digit" percentage of developers are selecting the contributor tier.
- OpenAI GPT-6 Astra:
- Greg Brockman declared: "welcome to the AGI era."
- OpenAI claims Astra was pre-trained on 100,000+ GPUs at the Stargate site in Texas, with previous models providing a significant portion of supervision during training.
- The presenter states Sam Altman confirmed the model underwent a formal review process with the Trump administration prior to release.
- On OSWorld 2.0, Astra reportedly scored 73% in ~40 minutes per task (compared to GPT-5.6 Sol at 65% in 75 minutes).
- OpenAI reported Astra scored 100% on ExploitBench, 42.4% on Exploit Gym, 64.6% on Terminal-Bench Science 0.1, 97.6% on FrontierMath Tier 4 (v2), and 99.9% on ARC-AGI-3 using OpenAI's response API harness (66% on the standard harness).
- OpenAI reported Astra meets the "Critical" cybersecurity threshold under its Preparedness Framework, enabling autonomous end-to-end zero-day discovery and exploitation.
- Standard API pricing is stated at $10 per million input tokens and $50 per million output tokens (with Fast mode offering up to 2x speed at 2x price).
- In Artificial Analysis's independent Intelligence Index, Astra scored 61 (matching GPT-5.6 Sol, but trailing Claude Fable 5.1's score of 66).
Notable quotes
- [00:03] "The years start coming and they don't stop coming, and I've never understood those words more deeply than I did this week after what felt like years in AI bizarro land."
- [04:10] "Sam told him something we've all heard after making a desperate late-night plea: go to bed."
- [06:33] "It scored a 61, which is exactly the same as GPT-5.6 Sol, and 5 points behind Claude Fable 5.1. Something doesn't add up here..."
Assessment
This is a tech news roundup and critical commentary video reviewing the launch events and early documentation of Claude Fable/Mythos 5.1, Muse Spark 1.3, and GPT-6 Astra. The presenter did not have hands-on early access to GPT-6 Astra himself, relying instead on official benchmark releases, public social media demonstrations by authorized early testers, and independent evaluation data from Artificial Analysis.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.