Claude Opus 5.5 Didn’t Need to Go This Hard
Matt Wolfe · 2026-09-22 · review · 216,797 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Matt Wolfe presents a breaking news overview from his hotel room in Palo Alto during Meta Connect, reviewing the simultaneous releases of Anthropic’s Claude Opus 5.5 and OpenAI’s GPT-6 Sol and Luna. He compares their benchmark performances, pricing structures, and third-party evaluations on platforms like Artificial Analysis and BuseyBench. He also highlights community-created interactive games and animations developed using Claude Opus 5.5.
What is shown
- [00:35] Anthropic's announcement page for Claude Opus 5.5 displaying headline claims and availability.
- [00:53] Anthropic's benchmark table comparing Opus 5.5 against Fable 5.1, Opus 5, GPT-6 Astra, and GPT-5.6 Sol across agentic coding, knowledge work, and computer use.
- [02:10] Pricing comparison charts between Claude Opus 5.5, Opus 5, and Claude Fable 5.1.
- [03:00] Terminal-Bench 4.0 accuracy versus cost graph showing Opus 5.5 configurations against competitors.
- [03:45] Artificial Analysis Intelligence Index and cost/output token charts showing Opus 5.5 taking the top spot.
- [05:10] Demos built with Opus 5.5: JavaScript procedural animations by Drew [05:10] and Kevin Ngo [05:36]; a playable Game Boy portfolio project by Angel [06:06]; an Antikythera mechanism 3D web game by Edwin [06:24]; a Blender claymation pipeline by Alex Albert [06:56]; a 3D doodle FPS by Tak [07:15]; a sand-trail snake game by Hakm [07:23]; and game demos from Alex at Forward Future including a Dark Souls tribute (The Ashen Gate), a flight simulator, and a Mario Maker clone [07:44].
- [09:16] OpenAI's launch page and API pricing for GPT-6 Sol and GPT-6 Luna.
- [10:13] OpenAI benchmark plots for AutomationBench, Agent's Last Exam, and DeepSWE.
- [14:22] The BuseyBench leaderboard showing Gary Busey SVG generations scored by LLM evaluators.
Claims & numbers
- The presenter says Anthropic claims Claude Opus 5.5 performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 [00:39].
- The presenter cites Anthropic benchmark results for Claude Opus 5.5: 66.4% on Terminal-Bench 4.0, 54.4% on FrontierCode v1.1 Main, 57.8% on CursorBench 4.0, 1846 on GDPval-AA v2.1, 40.0% on AutomationBench, 67.7% on Humanity's Last Exam, 81.8% on OSWorld 2.0, and 89.0% on Chartography [01:09].
- The presenter reports Opus 5.5 API pricing as $4.00 per million input tokens and $20.00 per million output tokens, compared to Fable 5.1 at $10.00 input and $50.00 output [02:20, 02:44].
- The presenter states that on the Artificial Analysis Intelligence Index, Claude Opus 5.5 scored 58 to take first place, ahead of Fable 5.1 and GPT-6 Astra, which were tied at 53 [03:47].
- The presenter states that Opus 5.5 costs $5.98 per Intelligence Index task on Artificial Analysis, compared to $7.63 for Fable 5.1, while consuming 119,000 output tokens per task versus Fable 5.1's 78,000 [04:14, 04:40].
- The presenter notes OpenAI cut API pricing in half for GPT-6 Sol compared to GPT-5.6 Sol ($2.00 input / $10.00 output vs. $4.00 / $20.00) and for GPT-6 Luna ($0.10 input / $0.50 output vs. $0.20 / $1.20) [09:47, 10:03].
- The presenter notes that on BuseyBench, GPT-6 Sol ranked #1 with a score of 7.5, followed by GPT-6 Astra at 7.3 and GPT-6 Sol Pro at 7.2, while Opus 5.5 ranked #8 [14:38, 15:10].
Notable quotes
- [00:00] "Another day, another new best model in the world just came out."
- [06:14] "Everything I'm seeing come out of Opus 5.5 is insanely impressive."
- [11:34] "You gotta give that edge to Anthropic a little bit because they just put out a model that's faster, cheaper, and better than their previous state of the art."
Assessment
This is an independent reaction and review video synthesizing launch materials, official benchmark disclosures, third-party index scores, and community demonstrations. The presenter did not run external verification of the benchmarks firsthand during the video, relying instead on vendor charts, public social media demos, and third-party benchmark dashboards.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.