Opus 5.5 ZMIENIA GRE! - Czy To Koniec GPT-6 Astra?
Dawid Banaszek | AI Automatyzacje · 2026-09-23 · community · 9,415 views · Polski
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, Polish tech creator Dawid Banaszek analyzes Anthropic’s launch of Claude Opus 5.5 on September 22, 2026. He reviews the official announcement, benchmark comparisons against OpenAI's GPT-6 Astra and Claude Fable 5.1, updated API pricing, and safety disclosures. He also demonstrates the model's availability inside the Claude Code interface, highlighting why using medium effort reasoning often delivers better cost-efficiency than maximum effort.
What is shown
- [00:02] Anthropic's official blog announcement page for Claude Opus 5.5 dated September 22, 2026.
- [00:04] Benchmark overview table comparing Claude Opus 5.5, Claude Fable 5.1, Claude Opus 5, GPT-6 Astra, and GPT-5.6 Sol across coding and reasoning benchmarks.
- [00:19] Terminal-Bench 4.0 accuracy vs. cost curve showing Opus 5.5, GPT-6 Astra, and Fable 5.1.
- [00:43] FrontierCode v1.1 main set score vs. cost graph.
- [01:08] CursorBench 4.0 benchmark graph displaying multi-file coding task evaluations.
- [01:26] The Claude Opus 5.5 System Card document, highlighting SWE-bench Pro results (section 8.2).
- [01:57] GDPval-AA v2.1 benchmark plot (Elo vs. cost across 44 professions).
- [02:18] Detailed benchmark breakdown covering Humanity's Last Exam, Chartography (visual chart recognition), and OSWorld 2.0 (Computer Use).
- [03:19] Benchmarks where GPT-6 Astra outperforms Opus 5.5 (AutomationBench and Terminal-Bench-Science 0.1).
- [04:18] Claude Platform Docs showing specifications: 1M token context window, 128K max output tokens.
- [04:27] Distillation safeguards and alignment documentation in the announcement post.
- [04:57] API pricing table for Opus 5.5 vs. Opus 5, along with Fast mode rates.
- [06:05] Claude Code interface UI showing model selection dropdown with Opus 5.5 and effort settings (Medium, High).
- [06:48] WANDR benchmark accuracy vs. cost chart comparing Opus 5.5 and Fable 5.1.
Claims & numbers
- The presenter says Claude Opus 5.5 scores 66.4% on Terminal-Bench 4.0 at high effort ($7.35/attempt), surpassing GPT-6 Astra (57.9% at $7.21) and Fable 5.1 (55.8% at $19.50).
- On FrontierCode v1.1, the presenter states Opus 5.5 achieves 54.4% at max effort ($6.19) and 54.6% at medium effort ($0.80), beating GPT-6 Astra’s top score of 53.3% ($4.36) at roughly one-fifth the cost.
- On CursorBench 4.0, the presenter notes Opus 5.5 scores 57.8% ($13.43), compared to Fable 5.1 at 51.8% ($17.28) and GPT-5.6 Sol at 41.7% ($6.33).
- Citing the system card, the presenter notes Opus 5.5 scores 89.9% on SWE-bench Pro, up from 79.2% on Opus 5.
- On GDPval-AA v2.1, Opus 5.5 scores 1846 Elo at max effort ($6.20) and 1576 Elo at medium effort ($0.86), while GPT-6 Astra tops out at 1542 Elo ($4.53).
- On Humanity's Last Exam (with tools), Opus 5.5 achieves 67.7%, Fable 5.1 gets 65.6%, and Astra gets 57.2%.
- On OSWorld 2.0 (computer use), Opus 5.5 reaches 81.8% under partial scoring and 48.7% under strict full-task scoring (compared to 74.0% partial / 42.8% strict for Opus 5).
- On benchmarks where Opus 5.5 loses to GPT-6 Astra, the presenter reports AutomationBench (40.0% for Opus 5.5 vs. 41.4% for Astra) and Terminal-Bench-Science 0.1 (58.7% for Opus 5.5 vs. 64.6% for Astra).
- The presenter reports Opus 5.5 API pricing is $4 per 1M input tokens, $20 per 1M output tokens (a 20% drop from Opus 5), $0.20 per 1M prompt cache read tokens (a 60% drop), and $5 per 1M cache write tokens. Fast mode costs $8 input / $40 output per 1M tokens with up to 2.5x speed.
- The presenter notes Anthropic claims Opus 5.5 outputs tokens over 30% faster than Opus 5 and crossed containment boundaries ~85% less often than Opus 5 or Mythos 5.1 in alignment evaluations.
Notable quotes
- [00:09] "Ale najciekawsza rzecz wychodzi dopiero na wykresie kosztu zadań." ("But the most interesting thing only comes out on the task cost graph.")
- [07:27] "Więcej myślenia nie zawsze pomaga." ("More thinking doesn't always help.")
- [08:09] "Pamiętajcie, że płacimy za ukończoną pracę. Samo dłuższe myślenie nie jest wynikiem." ("Remember that we pay for completed work. Longer thinking by itself is not the result.")
Assessment
This is an authentic commentary and review video evaluating Anthropic’s Claude Opus 5.5 release materials, documentation, and benchmark curves. The creator analyzes published graphs and shows the model available in his Claude Code environment, giving practical guidance on cost-accuracy tradeoffs rather than making unverified performance claims.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.