Anthropic went CRAZY (Opus 5.5)
Matthew Berman · 2026-09-22 · review · 172,947 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary In this livestream broadcast, host Matthew Berman reviews the release of Anthropic's Claude Opus 5.5, breaking down its benchmark scores, pricing, and system architecture updates. Midway through the stream, Anthropic technical staff member Thariq joins for a live interview to discuss how Opus 5.5 compares to Fable 5.1, recursive self-improvement in development, and the model's performance in developer workflows.
What is shown
- [00:00] Overview of Anthropic's X/Twitter announcement video and release statement for Claude Opus 5.5.
- [00:31] A chart showing task duration regression for human coding benchmarks across LLM release history up to Claude Mythos Preview.
- [02:11] Official benchmark comparison table showing Claude Opus 5.5 alongside Fable 5.1, Opus 5, GPT-6 Astra, and GPT-5.6 Sol across coding, knowledge work, and tool use benchmarks.
- [06:05] Pricing breakdown table comparing Claude Opus 5.5 against Claude Opus 5 ($4/M input, $20/M output vs. $5/M and $25/M).
- [07:22] Efficiency and cost-per-task curve charts for AutomationBench, FrontierCode v1.1, GDPval-AA v2.1, and Terminal-Bench 4.0 across reasoning effort levels (low, medium, high, extra high, max).
- [11:51] Example comparison showing output conciseness between Claude Opus 5 and Claude Opus 5.5 when explaining code changes and bug fixes.
- [13:00] Review of Anthropic's blog post detailing safety evaluations, behavioral audits, and the Life Sciences and Cyber Verification programs.
- [16:40] Artificial Analysis Intelligence Index v4.3 chart ranking Opus 5.5 at the top with an index score of 58.
- [17:06] Live interview with Anthropic technical staff member Thariq, discussing model selection, pacing the frontier, harness tooling, and recursive self-improvement workflows.
Claims & numbers
- The presenter notes Claude Opus 5.5 costs 40% less to run on typical workloads than Opus 5 and outputs tokens over 30% faster.
- Benchmark scores shown for Opus 5.5 include:
- Terminal-Bench 4.0: 66.4% (vs. 55.8% for Fable 5.1, 52.3% for Opus 5, and 57.9% for GPT-6 Astra).
- FrontierCode v1.1 (math set): 54.4% (vs. 50.3% for Fable 5.1 and 53.3% for GPT-6 Astra).
- CursorBench 4.0: 57.8% (vs. 51.8% for Fable 5.1 and 46.6% for Opus 5).
- GDPval-AA v2.1 (Knowledge work Elo): 1846 (vs. 1735 for Fable 5.1, 1708 for Opus 5, and 1542 for GPT-6 Astra).
- AutomationBench: 40.0% (vs. 31.4% for Fable 5.1 and 41.4% for GPT-6 Astra).
- Humanity's Last Exam (with tools): 67.7% (vs. 65.6% for Fable 5.1 and 57.2% for GPT-6 Astra).
- Research-Bench-Science 0.9 (with tools): 58.7% (vs. 52.6% for Fable 5.1 and 64.6% for GPT-6 Astra).
- OSWorld 0.9 (Computer use): 81.6% (vs. 80.7% for Fable 5.1 and 74.0% for Opus 5).
- Visual chart recognition (Chartography): 89.0% (vs. 88.4% for Fable 5.1).
- Pricing per 1M tokens for Claude Opus 5.5 is listed at $4 input, $20 output, $0.20 cache reads, and $5 cache writes.
- The presenter cites an early tester claim from the announcement post reporting a 680,000-line code migration completed in less than one day.
- On the Artificial Analysis Intelligence Index v4.3, Claude Opus 5.5 ranks #1 with a score of 58 (followed by Claude Fable 5.1 Max at 53 and GPT-6 Astra at 51).
- Thariq states that Claude writes "pretty much all the code" for its own development harness, creating an ongoing form of recursive self-improvement.
Notable quotes
- [04:06] "That is over a 300-point Elo jump. And so this benchmark measures things like PowerPoint creation, data entry, word processing..." — Matthew Berman
- [17:34] "I do think it's one of those times where, like, the model is both cheaper and more intelligent..." — Thariq
- [19:29] "I think that, like, Claude helps build Claude. You know, I think we've talked about this... Claude writing pretty much all the code is like a form of recursive self-improvement..." — Thariq
Assessment This is a live review and interview stream analyzing Anthropic's official announcement and benchmark disclosures, accompanied by commentary from an Anthropic engineer. The performance data and pricing shown are official reported figures from Anthropic and Artificial Analysis, though live real-time benchmarking is not conducted on stream.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.