Opus 5.5 vs. GPT-6 Astra. Is Claude the winner?
Jacek Bąk · 2026-09-27 · review · 3,892 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this review, presenter Jacek Bąk evaluates Anthropic’s newly released Claude Opus 5.5, analyzing its official release claims, benchmark scores against competitors like GPT-6 Astra and Claude Fable 5.1, and third-party evaluations from Artificial Analysis. He also shares his hands-on experience using Opus 5.5 to programmatically build 21 custom animation clips for a video project using Claude Code, concluding that the model shows impressive agentic capabilities and improved communication.
What is shown
- Anthropic’s official blog post introducing Claude Opus 5.5 on September 22, 2026 [00:32].
- Official benchmark comparisons showing agentic coding metrics across Terminal-Bench 4.0, FrontierCode V1.1, and CursorBench 4.0 [01:06].
- Infographics explaining the Terminal-Bench 4.0, FrontierCode, and CursorBench testing setups [01:13, 01:47, 02:19].
- API pricing tables comparing Claude Opus 5.5 with Opus 5 [03:17], followed by an infographic explaining how lower token prices combined with reduced token consumption yield a ~40% cost reduction per task [04:11].
- Side-by-side comparison of communication style between Opus 5 and Opus 5.5 on a bug-explanation prompt [04:31].
- Independent benchmark results on the Artificial Analysis website, including the Intelligence Index, Cost per Task, and Output Tokens per Task [06:03, 06:51, 07:23, 08:06].
- Screen recording of Claude Code agent sessions generating HTML/JS/Python motion graphic animations for B-roll video rendering [09:19].
Claims & numbers
- The presenter says Anthropic claims Opus 5.5 performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5 [00:35].
- On Terminal-Bench 4.0: Claude Opus 5.5 scores 66.4%, Claude Fable 5.1 scores 55.8%, Opus 5 scores 52.3%, GPT-6 Astra scores 57.9%, and GPT-5.6 Sol scores 37.3% [01:30].
- On FrontierCode V1.1 (Main): Opus 5.5 scores 54.4%, GPT-6 Astra scores 53.3%, Fable 5.1 scores 50.3%, and Opus 5 scores 48.0% [02:00].
- On CursorBench 4.0: Opus 5.5 scores 57.8%, Fable 5.1 scores 51.8%, Opus 5 scores 46.6%, and GPT-5.6 Sol scores 41.7% [02:34].
- API pricing: Opus 5.5 costs $4 per 1M input tokens and $20 per 1M output tokens (down 20% from $5/$25 on Opus 5), while cache reads cost $0.20 per 1M tokens (down 60% from $0.50 on Opus 5) [03:24].
- On Artificial Analysis Intelligence Index v4.3.2: Opus 5.5 at max effort scores 58 points (#1), beating GPT-6 Astra at max (53 points) and Fable 5.1 (53 points) [06:33].
- Artificial Analysis cost and token usage: at maximum effort, Opus 5.5 costs $5.98 per task using ~119,000 output tokens, compared to GPT-6 Astra at $3.26 using ~27,000 output tokens, Opus 5 using ~73,000 tokens, and Fable 5.1 using ~78,000 tokens [07:01, 07:24].
- At default medium effort on Artificial Analysis: Opus 5.5 and GPT-6 Astra tie at 51 points on the Intelligence Index, but Opus 5.5 costs $1.34 per task (using 26k tokens) vs $1.73 per task for Astra (using 12k tokens) [08:14].
- The presenter states he generated all 21 animation clips for his previous YouTube video using Claude Code and Opus 5.5 without exhausting even half of the 5-hour Claude Pro usage window [10:00].
Notable quotes
- "Anthropic mówi wprost: dostajemy model o możliwościach zbliżonych do ich dotychczas najmocniejszego modelu Fable 5.1, ale w cenie niższej niż wcześniejszy Opus 5." [00:46]
(Translation: "Anthropic states directly: we get a model with capabilities close to their previously strongest model Fable 5.1, but at a lower price than the previous Opus 5.") - "Wcześniej powiedziałem, że Opus 5.5 ma być około 40% tańszy, ale przecież ceny tokenów spadły jedynie o 20%. Skąd więc bierze się te 40%?" [04:00]
(Translation: "Earlier I said Opus 5.5 is supposed to be about 40% cheaper, but token prices only dropped by 20%. So where does this 40% come from?") - "...same benchmarki to dla mnie tylko część historii. Ostatecznie najważniejsze jest to, jak model sprawdza się realnie w naszej codziennej pracy." [10:17]
(Translation: "...benchmarks alone are only part of the story to me. Ultimately, what matters most is how the model performs in our daily work.")
Assessment
This is an authentic tech review and practical test video analyzing Anthropic’s September 2026 Claude Opus 5.5 release. The presenter accurately contextualizes public benchmark figures from both Anthropic and Artificial Analysis before showing a genuine, unexaggerated real-world workflow using the model through Claude Code.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.