Claude Opus 4.8 | First impressions
Arena AI · 2026-06-01 · review · 4,896 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Peter Gostev, AI Capability Lead at Arena, reviews Anthropic's newly released Claude Opus 4.8 model. He examines Anthropic's reported benchmark metrics and release timeline before running extensive side-by-side evaluations across complex 3D Three.js scenes, interactive browser games, and front-end web applications on Arena's evaluation platform.
What is shown
- Benchmarks & Release History [00:24–02:01]: A comparison table showing Opus 4.8 scores against Opus 4.7, GPT-5.5, and Gemini 3.1 Pro on coding and reasoning benchmarks, followed by an Anthropic release timeline chart showing accelerating release cycles.
- 3D Procedural Scene Generation [02:02–10:57]: Side-by-side rendering tests of complex procedural Three.js environments, including a voxel Roman Colosseum [03:32], a detailed coral reef [06:02], Notre Dame cathedral with stained-glass illumination [07:42], and the Giza Plateau pyramids [17:58].
- Interactive Mini-Games [10:58–17:36]: Testing real-time interactive game generation, including a 3D cart driving game through giant flowers [10:58], a Sistine Chapel vault drone restoration game [13:32], and a sunflower vase projectile game [15:52].
- Large-Scale Dynamic Scenes [21:29–30:20]: Testing the Golden Gate Bridge simulation with dynamic weather, water rendering, and traffic density [21:29], followed by marine life simulations of sperm whales and an octopus [26:19–29:05].
- Front-End UI Design & Web Apps [30:35–36:26]: Evaluating multi-component interactive React/web layouts, including a children's physics museum page ("WonderLab") [30:35], a bespoke vinyl record pressing website [32:38], and a mechanical toy workshop app [34:16].
Claims & numbers
- Opus 4.8 Benchmark Scores (as reported by Anthropic and presented by Gostev):
- SWE-bench Pro: 69.2% for Opus 4.8 (vs. 64.2% for Opus 4.7, 58.6% for GPT-5.5, and 54.2% for Gemini 3.1 Pro).
- Agentic Terminal Coding (TerminalBench 2.1): 74.6% for Opus 4.8 (vs. 66.1% for Opus 4.7, 78.2% for GPT-5.5, and 70.3% for Gemini 3.1 Pro).
- Multidisciplinary Reasoning: 69.8% (Opus 4.8) vs. 64.7% (Opus 4.7).
- Agentic Computer Use: 83.4% (Opus 4.8) vs. 82.8% (Opus 4.7).
- Knowledge Work: 1890 Elo (Opus 4.8) vs. 1753 Elo (Opus 4.7).
- Agentic Financial Analysis: 55.9% (Opus 4.8) vs. 51.5% (Opus 4.7).
- Release Cadence: Anthropic's average gap between releases across the Claude 4 generation is 59.8 days, dropping to 42 days between Opus 4.7 (April 16, 2026) and Opus 4.8 (May 28, 2026).
- Thinking vs. Non-Thinking: Gostev claims that for Anthropic models, the thinking variant does not always outperform the non-thinking variant, and in some game controls and 3D scenes the non-thinking model produced cleaner, more controllable results.
Notable quotes
- [00:00] "It's always an exciting day when we have a new frontier model out. Today it's Opus 4.8."
- [01:48] "We are all the way down to 42 days between Opus 4.7 and 4.8. Now this is acceleration."
- [37:36] "I would say the difference is very meaningful. Like, you can really see the difference."
Assessment
This is a hands-on review and live evaluation by Arena's AI capability lead, testing code generation live in-browser across a standardized test battery. The demonstrations are authentic, interactive software generations rendered directly in the Arena UI, openly displaying both model successes and rendering/logic glitches.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.