Opus 5.5 vs GPT-6 is racing to the bottom..?
Caleb Writes Code · 2026-09-25 · review · 87,549 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Caleb from Caleb Writes Code examines the trade-offs between cost efficiency and token efficiency among frontier AI models, particularly Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna, and GPT-6 Astra. He develops a 3D visualization combining intelligence, cost, and token usage to analyze how frontier labs optimize models and how consumer subscription limits versus API pricing shift the burden of token inefficiency.
What is shown
- [00:12] Artificial Analysis 2D scatter plots evaluating models on the Pareto frontier for Intelligence Index versus Cost per Task and Output Tokens per Task.
- [01:22] A custom 3D coordinate plot showing Claude Opus 5.5 plotted across three axes: Cost per task (USD), Output tokens per task, and Intelligence Index.
- [02:00] Anthropic's earlier models (Claude Fable 5.1 and Claude Opus 5) overlaid onto the 3D scaling space alongside Claude Opus 5.5.
- [02:20] Adding OpenAI's GPT-6 Sol, GPT-6 Luna, and GPT-6 Astra onto the 3D graph to contrast their scaling trajectories against Opus 5.5.
- [03:39] A sponsored workflow demo in the Hyperagent interface showing multi-agent travel orchestration (coordinating agents Sofia, Marco, Gianni, and Luca for itinerary planning, web research, and media generation).
- [04:50] A financial breakdown chart of projected 2025 ARR comparing OpenAI ($12B total) and Anthropic ($5B total) across consumer subscriptions, enterprise partnerships, and developer API channels.
- [06:14] 3D clustering of competing models from DeepSeek, Moonshot (Kimi), Zhipu/Z.ai (GLM), Xiaomi (MiMO), Google (Gemini), MiniMax, Meta, and xAI.
- [06:54] Longitudinal Pareto frontier curves illustrating progression from Q1 through Q3 2026 across cost and token efficiency.
Claims & numbers
- The presenter states Claude Opus 5.5 costs 40% of Claude Fable 5.1 ($4.00 vs. $10.00 on screen) [00:03].
- The presenter states GPT-6 Sol dropped 50% from $4.00 to $2.00, and GPT-6 Luna dropped 50% from $0.20 to $0.10 [00:05].
- The presenter notes Claude Opus 5.5 dominates the cost-efficiency frontier once performance moves past GPT-6 Sol [00:33].
- The presenter notes GPT-6 models dominate token efficiency until Opus 5.5 pushes intelligence further at higher token volumes [00:57].
- The presenter claims Claude Opus 5.5 starts to plateau around an Intelligence Index score of approximately 53 [01:44].
- The presenter reports that GPT-6 Sol tops out at roughly 47.5 on the Intelligence Index, while GPT-6 Luna reaches approximately 37.3 [02:35].
- The presenter states OpenAI's projected 2025 ARR is $12 billion ($6.5B consumer subscriptions, $3.6B enterprise/partners, $1.9B API), while Anthropic reaches $5 billion ($2.9B API, $1.4B Cursor & GitHub Copilot, $0.7B consumer subscriptions) [04:50].
- The presenter notes consumer LLM subscriptions typically meter usage via rolling 5-hour windows and weekly quotas [05:18].
Notable quotes
- [00:08] "What we're seeing here is the cost of intelligence continually dropping, but is it really?"
- [01:12] "So what you're seeing here is a tension between cost-efficient and a token-efficient model."
- [05:43] "So the tension here between users and inference providers is really who ends up paying for the inefficient token that gets generated by the model."
Assessment
This is an independent analysis and review combining third-party benchmark data (primarily Artificial Analysis) with a sponsored product demonstration of Hyperagent. The 3D graphs and Pareto frontier mappings are analytical visual representations created by the presenter rather than official provider benchmarks, but the underlying tool UIs and data points are shown authentically.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.