Claude Opus 5.5 is ridiculous
AI Search · 2026-09-23 · review · 777,110 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
This video is a comprehensive hands-on review and benchmark breakdown of Anthropic’s Claude Opus 5.5, hosted by the creator behind the AI Search channel. The presenter evaluates the model’s agentic capabilities using Claude Code and the chat interface across complex real-world coding, multimedia creation, gaming, vision, medical imaging, and reasoning tasks.
What is shown
- CAPTCHA Bypass Challenge [00:52]: Claude Opus 5.5 attempts the Neal.fun “I’m Not a Robot” test suite via a browser interface, solving text captchas, nested grids, whack-a-mole, and Waldo puzzles, but struggling and taking over 14 minutes on a dynamic car-parking game.
- Ray-Tracing Physics Simulation [03:49]: Using a multi-agent self-critique loop with no external libraries, the model codes a WebGL/raw shader 3D simulation of a bullet piercing a water balloon with real-time controls.
- Live Piano Performance [06:38]: The model composes an original Chopin-style piece and autonomously plays it live in real time on an online virtual keyboard by sequencing DOM events over a 30-minute coding run.
- Motion Graphics Explainer Video [08:55]: The model generates code to create a complete 1-minute animated video explaining Eratosthenes’ calculation of Earth’s circumference, paired with Gemini TTS audio.
- Higgsfield MCP Integration (Sponsor Segment) [10:36]: Demonstrations showing Claude Opus 5.5 orchestrating 3D video, physics simulations, and commercial video creation through Higgsfield tools.
- 3D Real Estate Virtual Tour [12:03]: Using Blender MCP, the model reconstructs an Airbnb listing in Motobu, Japan from web photos and renders an aerial and interior flythrough.
- Playable Unreal Engine 3D Game [14:57]: The model creates a procedural ancient Chinese imperial environment in Blender/Unreal Engine, imports a third-person ninja character from Sketchfab, and retargets animations from Mixamo.
- DAW Music Production [17:02]: The model operates Waveform DAW via script to compose, mix, and master a 1-minute EDM track.
- Vision & Medical Tests [18:56]: The model fails a camouflage frog-spotting image test (hallucinating an Eastern fence lizard) and achieves 1 out of 6 correct diagnoses on a multi-panel brain CT tumor scan.
- Deep Research & Idea Generation [20:29]: Claude Opus 5.5 generates flowcharts and tables analyzing atherosclerosis treatment trials, followed by three novel automated intervention concepts for factory farming animal welfare.
- Benchmark & Pricing Overview [22:36]: Overview of benchmark results across Terminal-Bench 4.0, LiveBench, Maze Bench, Vals Index, and KernelBench, along with token pricing and safety safeguards.
Claims & numbers
- The presenter notes Claude Opus 5.5 was announced on September 22, 2026.
- The presenter claims Opus 5.5 scores 58 on the Artificial Analysis Intelligence Index, ranking #1, with an output speed of 66 tokens per second.
- On Artificial Analysis Cost per Task, the presenter notes it costs approximately $5.98 per task (compared to $7.83 for Claude Fable 5.1 and $3.26 for GPT-6 Astra).
- The presenter reports Opus 5.5 has a 59% hallucination rate on the AA-Omniscience benchmark, lower than Fable 5.1 but higher than GPT-6 Astra (29%), Grok 4.7, and Muse Spark 1.3.
- On Terminal-Bench 4.0, official self-reported figures show Opus 5.5 at 64.4% agentic coding, FrontierCode v1.1 at 54.4%, GDPval AA v2.1 at 1846, OSWorld 2.0 at 81.0%, and ChartQA at 92.0%.
- On LiveBench, the presenter shows Opus 5.5 ranking 2nd overall with an 83.2 score (behind Claude Fable 5.1 at 83.4).
- On Maze Bench, Opus 5.5 achieves a 6% gem collection score compared to 14% for GPT-6 Astra.
- On the Vals Index (GDP-weighted agentic economic benchmark), Opus 5.5 ranks #1 with 69.69% accuracy at $22.30 cost per test.
- On KernelBench (CUDA kernel optimization), Opus 5.5 ranks #1 across all tested models.
- The presenter notes Claude Opus 5.5 is available on paid plans and via API, with strict automated fallback safeguards for cybersecurity, biology, and distillation queries.
Notable quotes
- [00:00] "Claude Opus 5.5 is out, and this might be the best model in the world."
- [08:40] "Holy smokes, that was insane. That sounded even better than what I got from GPT-6 Astra."
- [25:08] "That sums up my review of Claude Opus 5.5. At least for certain tasks, this does seem to be the best model in the world."
Assessment
This is an independent hands-on review and stress-test of Claude Opus 5.5 featuring real, long-running agentic coding and browser automation workflows executed via Claude Code and the web UI. While long waiting periods are fast-forwarded for video pacing, the presenter transparently shows both impressive outputs (playable Unreal environment, DAW automation, piano sequencing) and clear failures (failing the camouflage frog test, getting 1/6 on brain tumor CT scans, and struggling on the CAPTCHA car-parking task).
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.