First Look at Claude Opus 4.8
Tonbi's AI Garage · 2026-06-01 · community · 3,286 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, creator Tonbi from the YouTube channel Tonbi's AI Garage reviews Anthropic's release of Claude Opus 4.8. He breaks down the model's official release slides, system card benchmarks, and new features before testing Opus 4.8 hands-on within the Claude Code terminal interface on frontend web design and technical experiment analysis tasks.
What is shown
- Release Announcement & System Card Overview [00:00–07:58]: Presentation slides showing Anthropic's official announcement, benchmark tables, and core improvements: coding reliability, effort controls, pricing changes, dynamic multi-agent workflows, and safety/honesty metrics.
- Claude Code Setup [07:59–08:18]: Claude Code CLI (
v2.1.154) running Claude Opus 4.8 configured with high effort reasoning. - Frontend Design Task [08:18–12:56]:
- The presenter prompts Opus 4.8 to build a single HTML file without a build step for a One Piece meets Star Wars game webpage, featuring a rotating 3D Three.js sphere with a custom GLSL fragment shader (rim lighting), GSAP scroll animations, and staggered entrance headline text [08:18].
- Displays the rendered browser result ("Void + Pirates") and compares it to Opus 4.7's previous attempt [09:47].
- When the headline text initially fails to render due to CSS background clipping, the presenter inputs a follow-up prompt, and Opus 4.8 diagnoses and updates the code to display the text "Where Legends Set Sail Beyond the Stars" [11:16–12:22].
- Research Plan Critique Task [12:57–15:40]:
- Opus 4.8 reads and analyzes a local
plan.mdoutlining a machine learning experiment involving LeJEPA representations, 6-DoF camera trajectories, and latent distribution alignment [12:57]. - Opus 4.8 produces a structured critique classifying potential failures into Tier 1 ("Will halt you or fail a gate"), Tier 2 ("Will silently corrupt results"), and Tier 3 ("Will annoy you / polish"), successfully flagging experimental confounding factors and hardware bottlenecks [14:07–15:33].
- Opus 4.8 reads and analyzes a local
Claims & numbers
- SWE-bench Scores: The presenter states Opus 4.8 scores 88.6% on SWE-bench Verified (versus 84.3% on Opus 4.7, 78.2% on GPT-5.5, and 70.3% on Gemini 3.1 Pro) and 69.2% on SWE-bench Pro (versus 64.3% on Opus 4.7 and 58.0% on GPT-5.5) [01:34, 02:41, 03:26].
- Terminal-Bench & OSWorld: The presenter reports GPT-5.5 leads on Terminal-Bench 2.1 at 78.2% compared to Opus 4.8's 74.6%, while Opus 4.8 leads OSWorld-Verified (computer use) at 83.4% (ahead of GPT-5.5's 78.7% and Gemini 3.1 Pro's 71.8%) [01:40, 02:03].
- Math Benchmark: The presenter claims Opus 4.8 achieved 96.7% on USAMO 2026 math, up from 69.3% on Opus 4.7 [02:22].
- Code Reliability: The presenter notes Anthropic claims Opus 4.8 is ~4x less likely than Opus 4.7 to let a code flaw slip past unmarked [02:59].
- ProgramBench: The presenter cites scores jumping from 71–84% on Opus 4.7 to 79–88% on Opus 4.8 [03:57].
- Effort Control & Efficiency: The presenter explains Opus 4.8 at minimum effort matches Opus 4.7 at maximum effort on SWE-bench Pro [04:27].
- Pricing: Standard tier pricing remains unchanged at $15 input / $75 output per million tokens, while "Fast mode" low-latency pricing runs at $10 input / $50 output per million tokens (three times cheaper than previous fast mode) [04:47].
- Multi-Agent Workflows: The presenter notes BrowseComp multi-agent score reached 88.5% (versus 84.3% single agent), and a 5-agent team completes hard tasks >3x faster at ~20% latency [05:43].
- Other Benchmarks: Harvey AI strict Legal Agent Benchmark reached 86.82% pass rate; GraphWalks BFS at 1M tokens scored 68.1% (compared to 40.3% on Opus 4.7 and 45.4% on GPT-5.5) [06:37].
- Security Caveat: The presenter highlights system card findings that Opus 4.8 is slightly less robust than Opus 4.7 on some agentic prompt-injection tests [06:58].
Notable quotes
- "The honest headline is that it's a real step up, but not a clean sweep." [01:27]
- "It catches its own bad code more often, which if you used Opus 4.7 a lot, like I did, you'll notice that there was a lot of bad code that slipped through." [03:10]
- "On SWE-bench Pro, Opus 4.8 at minimum effort matched Opus 4.7 at maximum effort." [04:26]
Assessment
This is an independent user review and hands-on demonstration from an AI creator, combining a walkthrough of Anthropic's official release deck with unedited, real-time testing in Claude Code. The creator transparently displays flaws during testing—such as a CSS background clipping bug requiring a follow-up prompt—rather than cherry-picking a flawless output.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.