GPT-6 Astra.. full analysis..
Caleb Writes Code · 2026-09-04 · review · 473,893 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Here is the catalog entry for this video:
Summary
Caleb from Caleb Writes Code provides a technical analysis of OpenAI's GPT-6 Astra release following its announcement. He examines GPT-6 Astra's benchmark performance across ARC-AGI-3, FrontierMath Tier 4, and DeepSWE v1.1, exploring why aggregate leaderboards like the Artificial Analysis Intelligence Index can be misleading, and highlights GPT-6 Astra's significant leap in token efficiency despite its higher per-token API pricing.
What is shown
- [00:03] OpenAI's benchmark comparison table showing GPT-6 Astra, GPT-5.6 Sol, Claude Fable 5.1, Claude Opus 5, and Gemini 3.8 Flash across 14 benchmarks (ARC-AGI-3, FrontierMath Tier 4, DeepSWE v1.1, ExploitBench, etc.).
- [00:14] The Artificial Analysis Intelligence Index leaderboard, showing GPT-6 Astra ranked 5th behind Claude Opus 5.5, Claude Fable 5.1, Grok 4.7, and Muse Spark 1.3.
- [01:52] ARC-AGI-3 gameplay interface and the underlying JSON representation (64×64 grid with 16 color states) sent to models via API.
- [05:05] Example problems from Epoch AI's FrontierMath across Tiers 1 through 4 (linear algebra, combinatorics, number theory, and BMO space analysis).
- [05:34] The DeepSWE v1.1 leaderboard and Pareto frontier chart comparing model pass rates, average cost per task, and output tokens.
- [06:32] Sponsor segment demonstrating Zo (
zo.computer), an always-on cloud computer agent controlled via iMessage and web UI to generate and host an e-commerce website for vintage watches. - [07:22] Pareto frontier analysis of DeepSWE highlighting GPT-6 Astra's step count and token consumption relative to GPT-5.6 Sol.
- [08:20] Artificial Analysis token usage per task graph, showing GPT-6 Astra consuming the fewest output tokens per task (9k).
- [10:00] OpenAI's computer use demonstration clips showing voice-driven desktop actions (canvas drawing, shopping on eBay, editing documents).
Claims & numbers
- Benchmarks & performance:
- On ARC-AGI-3, the presenter states GPT-6 Astra scored 62.7% using the standard ARC Foundation harness, but reached 99.9% when evaluated using OpenAI's internal harness; in contrast, NVIDIA scored 100% on the public split using its "Avo" agentic harness on Claude Opus 5.
- On FrontierMath Tier 4, the presenter states GPT-6 Astra achieved 97.6% (compared to Claude Fable 5.1's 87.8%).
- On ExploitBench, GPT-6 Astra reportedly scored 100%.
- On DeepSWE v1.1, GPT-6 Astra scored 74% (pass@1), matching Gemini 3.8 Flash (74%) and Claude Opus 5 (74%), but used only 30k output tokens and 29 steps, compared to 60k tokens / 61 steps for GPT-5.6 Sol, 118k tokens / 99 steps for Claude Opus 5, and 143k tokens / 166 steps for Gemini 3.8 Flash.
- On Computer Use benchmarks, the presenter lists GPT-6 scores: OSWorld (72.6%), ScreenSpot-Pro (92.7%), Mind2Web (1.9x), and Computer Use Safety (2.4%).
- Pricing:
- GPT-6 Astra is priced at $10 per million input tokens and $50 per million output tokens.
- GPT-5.6 Sol is priced at $4 per million input tokens and $20 per million output tokens.
- Economic implications: The presenter argues that because GPT-6 Astra solves tasks using half the tokens of previous models, the supply of intelligence per token has effectively doubled, allowing OpenAI to protect high margins despite high nominal per-token API prices.
Notable quotes
- [00:08] "Only one of these benchmarks actually overlaps with Artificial Analysis Intelligence Index, which has its own set of benchmarks that it tracks."
- [07:47] "What you're seeing here is a model that is not cost-efficient, but token-efficient, which there is a difference between these two."
- [08:51] "The tension here is that making the model more token-efficient means OpenAI needs fewer billable tokens to deliver the same amount of value."
Assessment
This is an independent technical review and commentary video by an AI developer analyzing OpenAI's published benchmark disclosures, third-party evaluations (Epoch AI, ARC Prize, Artificial Analysis), and official marketing demonstrations. The charts and performance tables presented are from verified third-party evaluation suites and official lab reports, paired with custom explanatory whiteboard animations.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.