Claude Opus 4.8: Here is Everything that Changed
Prompt Engineering · 2026-06-01 · community · 19,987 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
The presenter from the channel Prompt Engineering reviews Anthropic’s release of Claude Opus 4.8 and its accompanying features. He walks through the official announcement blog posts, benchmark performance, pricing, and API updates, before explaining Claude Code’s new "dynamic workflows" and demonstrating Opus 4.8's code-generation performance across various effort levels on Claude.ai.
What is shown
- [00:00] Intro showcasing Claude Code CLI migrating an application monorepo to Next.js App Router and receiving push-notification status updates.
- [01:17] Anthropic's announcement blog post dated May 28, 2026: "Introducing Claude Opus 4.8".
- [01:25] Benchmark capability table comparing Opus 4.8 against Opus 4.7, GPT-5.5, and Gemini 3.1 Pro.
- [03:05] Claude.ai interface demonstrating the new manual Effort selector (Low, Medium, High, Extra, Max) alongside the Adaptive Thinking toggle.
- [03:51] Breakdown of the Messages API update permitting system entries inside the messages array mid-conversation without invalidating prompt caching.
- [04:37] Preview of Anthropic's roadmap ("What's next?"), mentioning Project Glasswing and upcoming Mythos-class models.
- [06:17] Sponsored walkthrough of JetBrains Academy and AWS Skill Paths within PyCharm.
- [08:24] Discussion of benchmark footnotes regarding Terminal-Bench 2.1 evaluation harnesses.
- [09:09] Anthropic blog post "Introducing dynamic workflows in Claude Code", showing how Claude orchestrates subagents and highlighting a case study porting Bun from Zig to Rust.
- [12:09] Live prompt demonstration on Claude.ai: generating a complex 3D Three.js voxel art pagoda garden in a single HTML file.
- [13:22] Interactive output of the voxel pagoda scene rendered in the browser under High, Max, and Low effort settings.
Claims & numbers
- Release Timing: The presenter states Opus 4.8 was released only 40 days after Opus 4.7 (dated May 28, 2026 in the blog post).
- Benchmarks reported by Anthropic:
- Agentic coding (SWE-bench Pro): Opus 4.8 scored 69.2% (vs. Opus 4.7 at 64.3%, GPT-5.5 at 58.6%, Gemini 3.1 Pro at 54.2%).
- Agentic terminal coding (Terminal-Bench 2.1): Opus 4.8 scored 74.6% (vs. Opus 4.7 at 66.1%, GPT-5.5 at 78.2%, Gemini 3.1 Pro at 70.3%; presenter notes GPT-5.5 scored 83.4% when using OpenAI's Codex CLI harness).
- Multidisciplinary reasoning (Humanity's Last Exam): Opus 4.8 scored 49.8% (vs. Opus 4.7 at 46.9%, GPT-5.5 at 41.4%, Gemini 3.1 Pro at 44.4%).
- Agentic computer use (OSWorld Verified): Opus 4.8 scored 83.4% (vs. Opus 4.7 at 82.8%, GPT-5.5 at 78.7%, Gemini 3.1 Pro at 76.2%).
- Knowledge work (GDPval-AA): Opus 4.8 scored 1890 (vs. Opus 4.7 at 1753, GPT-5.5 at 1769, Gemini 3.1 Pro at 1314).
- Agentic financial analysis (Finance Agent v2): Opus 4.8 scored 53.9% (vs. Opus 4.7 at 51.5%, GPT-5.5 at 51.8%, Gemini 3.1 Pro at 43.0%).
- Model Honesty: The presenter cites Anthropic’s testing showing Opus 4.8 is four times less likely to allow unremarked flaws in the code it produces.
- Dynamic Workflows & Bun Port: Anthropic claims Jarred Sumner used dynamic workflows to port Bun from Zig to Rust (~750,000 lines of Rust) in 11 days, passing 99.8% of the existing test suite.
- Pricing: Standard usage remains unchanged at $5 per million input tokens and $25 per million output tokens; fast mode (running at 2.5x speed) is priced at $10 input / $50 output per million tokens, which the presenter notes is three times cheaper than previous fast modes.
Notable quotes
- [00:04] "Now, this seems to be an incremental improvement over Opus 4.7, but this is designed for long-running tasks."
- [04:20] "You can update Claude's instructions mid-task without breaking the prompt cache or routing the update through a user turn."
- [08:52] "The harness that you use with the model is a lot more important now."
Assessment
This is a third-party community review and walkthrough analyzing Anthropic's official blog posts and documentation alongside real web UI tests. The 3D Three.js voxel pagoda generation is demonstrated live in real time across different effort tiers, while enterprise workflows (such as the monorepo migration and Bun porting) rely directly on Anthropic's published announcements.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.