I Tested NEW Opus 5.5 on 24 Coding Prompts. WOW.
AI Coding Daily · 2026-09-23 · review · 19,285 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Povilas Korop from AICodingDaily evaluates Anthropic’s Claude Opus 5.5 on his standardized 24-prompt coding benchmark suite across backend, frontend, and offline app projects. He examines the model's performance, speed, and cost efficiency across Medium and High effort settings, comparing the results to Claude Opus 5, Claude Fable 5.1, and OpenAI's GPT-6 models.
What is shown
- [00:00] Overview of the week's AI releases, including Claude Opus 5.5, OpenAI GPT-6 Sol/Luna, and MiMo v2.6.
- [00:49] The AICodingDaily LLM Leaderboard showing previous standings where GPT-6 Astra (Medium) and GPT-6 Sol (High) held top ranks over Opus 5.
- [01:20] Terminal execution logs of benchmark runs using Claude Code / Claude CLI on Go, Dart/Flutter, and PHP test suites (e.g.,
offlinesyncandshipping-quotes). - [02:07] ClaudeDevs announcement on X detailing Opus 5.5's performance parity with Fable 5.1, 30% faster execution, 40% lower cost, and 20% increased 5-hour rate limits.
- [03:35] Google Sheets evaluation tables for back-end (Laravel/PHP) and front-end (React/TypeScript) code quality evaluated by GPT-5.6 Sol.
- [04:48] Updated AICodingDaily leaderboard placing Claude Opus 5.5 (High) and Claude Opus 5.5 (Medium) at #1 and #2 overall.
- [06:11] Official API pricing comparison table showing per-million token rates for Claude Opus 5.5 versus Opus 5.
- [07:24] Anthropic Pro plan account usage interface displaying the 5-hour limit reset functionality.
- [08:09] Third-party benchmarks and user impressions from X (Pawel Huryn, Nat McAleese, Kun Chen) evaluating Opus 5.5 against real-world repos.
Claims & numbers
- The presenter says Anthropic claims Claude Opus 5.5 matches Claude Fable 5.1's performance while being approximately 30% faster and 40% cheaper per task than Opus 5 [02:07].
- The presenter states Anthropic increased 5-hour session limits by 20% in Claude Code for Pro, Max, and Team users, adding a banked reset option [02:07, 07:34, 08:03].
- According to the pricing graphic, Claude Opus 5.5 costs $4 per 1M input tokens, $20 per 1M output tokens, $0.20 per 1M cache reads, and $5 per 1M cache writes (compared to Opus 5 at $5, $25, $0.50, and $6.25, respectively) [06:14].
- On the presenter's benchmark (max 60 points), Claude Opus 5.5 (High) achieved 57.83 total points with an average cost of $0.79 and time of 3 minutes 10 seconds per prompt [04:50, 06:46].
- Claude Opus 5.5 (Medium) scored 57.37 points with an average cost of $0.56 and an average time of 2 minutes 4 seconds per prompt [04:50, 05:56, 06:46].
- The presenter notes Opus 5.5 (Medium) was roughly twice as fast as Opus 5 (which averaged over 6 minutes on high and nearly 4 minutes on medium) and cheaper than Opus 5 runs that averaged over $1.00 per prompt [06:00, 06:49].
- A benchmark cited from Pawel Huryn claimed Opus 5.5 (max) resolved 43 out of 45 planted bugs across 2 repos for $60.49, matching Fable 5.1 (43 for $77.55) and trailing GPT-6 Astra (45 for $33.03) [08:09].
Notable quotes
- [00:23] "And this is 5.5, not 5.1. It's not incremental release."
- [01:09] "And spoiler alert: hell yes. Let me show you."
- [05:40] "Someone tweeted the other day that we don't need better models like Fable or Astra, we need regular models, but for cheaper price."
Assessment
This is an independent benchmark review and evaluation video using real automated terminal testing scripts, project test suites, and custom evaluation sheets. All test logs and metrics are displayed transparently within the presenter's testing workflow without obvious staging or misleading edits.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.