I reviewed Opus 5.5 and GPT-6 Sol live - and the results surprised me
How I AI · 2026-09-22 · review · 35,545 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
The host of the How I AI podcast presents a live blind evaluation and review comparing newly released AI models, specifically Anthropic's Claude Opus 5.5 and OpenAI's GPT-6 Sol and GPT-6 Luna, alongside previous models like GPT-6 Astra and Claude Fable 5.1. She analyzes model pricing, latency, and safeguard changes before running outputs through her custom "How I AI vibe review" benchmarking tool across knowledge work, front-end design, back-end code, agentic tasks, SVGs, and 3D modeling.
What is shown
- [01:29] Presentation slides detailing model release context, positioning, and API pricing comparisons between OpenAI and Anthropic models.
- [02:51] Complete price board showing per-million token pricing across frontier and tier-below models (GPT-6 Astra, Claude Fable 5.1, Claude Opus 5.5, GPT-6 Sol, GPT-6 Luna).
- [04:11] Slide breakdown on safety guardrails (Opus 5.5 rerouting cybersecurity tasks to Opus 4.8), effort dial defaults, and prompt caching cost impacts.
- [09:07] Demonstration of the blind evaluation tool ("How I AI - vibe review") testing knowledge work tasks (converting messy notes into PRDs and PRD readiness checks).
- [11:30] Blind evaluation of personal productivity tasks: inbox email triage, drafting replies, and automated calendar extraction across Models B, C, E, and G.
- [13:48] Comparison of generated front-end web interfaces across models: an editorial layout ("Folio Dispatch"), a dark-mode incident response dashboard, an operational dock scheduling console, B2B renewal tracking dashboards, and plant-care consumer web apps.
- [21:12] Evaluation of 3D modeling and SVG generation quality in consumer prototypes (notably plant illustrations and UI cards).
- [24:12] Evaluation of back-end coding tasks: auditing graph mutations and generating specifications for a back-end feature.
- [25:18] Evaluation of long-running agent workflows (processing multiple customer support tickets into an executive summary memo) and agent conversational personas.
- [28:25] Testing multi-expression vector SVG character generation (document, microphone, and bug icons).
- [30:00] Testing AI-assisted automated vertical short-form video editing and caption placement from raw selfie footage.
- [31:21] "Barbie bench" test: generating a full 3D interactive runway fashion studio app with a 3D animated Barbie model inside Claude Opus 5.5.
- [34:17] Review of final benchmark scores, preference breakdowns, task-by-task winners, and a comparison demonstrating a negative correlation ($r = -0.06$) between the human host's rankings and an automated LLM judge.
Claims & numbers
- The presenter notes that neither lab released a frontier-tier replacement this week; the releases represent the high-volume tier beneath Claude Fable 5.1 and GPT-6 Astra [02:31].
- The presenter shows verified pricing per million tokens: GPT-6 Astra and Claude Fable 5.1 at $10 input / $50 output; Claude Opus 5.5 at $4 input / $20 output (a 20% cut below Opus 5); GPT-6 Sol at $2 input / $10 output (a 50% cut); and GPT-6 Luna at $0.10 input / $0.50 output (a 58% reduction on outputs from $1.20) [02:51, 03:31].
- The presenter claims Anthropic introduced Claude Opus 5.5 cache reads at $0.20 (60% lower than Opus 5) and a Fast mode priced at $8 input / $40 output running up to 2.5× faster [03:31].
- The presenter states OpenAI offers a 90% discount on cached inputs, that changing reasoning effort dials or tools no longer invalidates prompt cache, and that GitHub saw over 50% fewer prompt tokens requiring fresh processing [03:31].
- The presenter states Claude Opus 5.5 implements Fable 5.1-level cyber and bio defense guardrails, causing most offensive cybersecurity queries to automatically reroute to Opus 4.8 [04:25].
- In her benchmark results across 58 blind outputs, the presenter reveals GPT-6 Astra scored highest relative to average (+0.57), Claude Opus 5.5 won the most individual categories (6 of 12) with a net +0.11, GPT-6 Sol tied at +0.11, and Claude Fable 5.1 ranked lowest at -0.83 [34:17, 34:49].
- The presenter reports that an automated LLM judge preferred Claude Fable 5.1 as #1 and placed GPT-6 Astra at #4, resulting in a near-zero/negative correlation ($r = -0.06$) with her personal ratings [36:51].
Notable quotes
- "Opus 5.5 is the first Opus-level model that has shipped with the Fable-level kind of like cyber and bio guardrails." [04:25]
- "Part of the way they made Opus 5.5 not annoying is they had it shut up." [08:14]
- "Astra wins my heart. Opus 5.5 wins my week. Sol splits me." [34:18]
Assessment
This is an authentic, independent benchmark review and hands-on product comparison conducted live on camera by a tech podcast host using her custom evaluation harness. All interfaces, generated web applications, prompt evaluations, and live ratings are demonstrated directly in real time without promotional sponsorship or deceptive staging.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.