Claude Opus 4.7 Just Dropped... Or Did It Really?
Nate Herk | AI Automation · 2026-05-02 · community · 88,079 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, AI creator Nate Herk evaluates Anthropic’s Claude Opus 4.7 release following weeks of community controversy over degraded performance and silent throttling in Claude Opus 4.6. He reviews technical data, leaked behavior metrics, benchmark claims, and the newly launched Claude Code Desktop app, then conducts head-to-head practical tests comparing Opus 4.6 (with extended thinking) and Opus 4.7.
What is shown
- [00:00] Overview of the Opus 4.7 announcement post and the preceding community complaints regarding Opus 4.6 performance drops.
- [00:46] Examination of data from an AMD Senior Director analyzing 6,852 Claude Code sessions, showing thinking depth dropped 73% (from 2,200 to 600 characters) and edits made without reading files first spiked from 6.2% to 33.7%.
- [03:07] Demonstration of the Claude Code Desktop App and VS Code CLI integration, toggling between model versions and effort settings (low, medium, high, xhigh).
- [04:29] Claude web interface UI showcasing the model selector: Opus 4.7 with "Adaptive thinking" versus Opus 4.6 with "Extended thinking."
- [05:24] Review of Anthropic’s official announcement blog post, benchmark tables (comparing Opus 4.7, Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Claude Mythos Preview), and the 232-page Claude Opus 4.7 System Card.
- [10:28] Claude Code Desktop app walkthrough showing session logs, live web previews, built-in terminal, plan breakdown, and the token context window tracker (5-hour and weekly limits).
- [13:18] Practical Test 1: Uploading a META stock daily chart and asking for a three-sentence analysis. Opus 4.6 extended gives a scenario-based response; Opus 4.7 gives direct trader terminology, specific support levels ($640), and supply-zone rationale.
- [14:20] Practical Test 2: SaaS 12-month financial modeling prompt. Opus 4.6 produces an interactive frontend dashboard with sliders; Opus 4.7 catches and self-corrects its own math errors and generates an exportable Excel (.xlsx) workbook with multi-tab financial projections.
Claims & numbers
- Opus 4.6 degradation data: An AMD senior director’s analysis showed thinking depth fell 73% (2,200 to 600 reasoning characters), the word "simplest" appeared 2.3x more often in outputs, and users interrupted the model 12x more frequently to prevent mistakes.
- Effort default changes: Anthropic introduced Adaptive Thinking on February 9, 2026, allocating zero reasoning tokens to tasks deemed simple. On March 3, 2026, Anthropic quietly changed default effort levels from "high" to "medium" for Pro and Max subscribers.
- BridgeBench benchmark: Opus 4.6 hallucination accuracy allegedly dropped from 83.3% to 68.3%, falling from #2 to #10 on the leaderboard.
- Opus 4.7 official benchmarks: SWE-bench Pro rose from 53.4% to 64.3% (+10.9 points); SWE-bench Verified improved from 80.8% to 87.6% (+6.8 points); vision accuracy on XBOV increased from 54.5% to 98.5% with 3x higher image resolution; CursorBench rose from 58% to 70%; Rakuten production task resolution improved 3x; reasoning on Humanity's Last Exam reached 46.9% (up from 40.0%).
- Pricing & tokenization: Opus 4.7 maintains pricing at $5 per million input tokens and $25 per million output tokens, but incorporates an updated tokenizer that yields roughly 1.0 to 1.35x more tokens for identical text inputs.
- Desktop app quality: Developer Theo reportedly documented 40+ software bugs in the Claude Code desktop app within an hour of testing.
Notable quotes
- [03:49] "The bottom line: they didn’t change the model itself. They changed how hard the model was allowed to think, and they didn’t tell anyone."
- [06:44] "It’s almost like they’re creating holes just so they can fill them and look like the hero."
- [16:32] "Whether it was intentional throttling or 'just' cost optimization, the effect was the same: a worse product at the same price."
Assessment
This is an independent user review and technical breakdown analyzing the Claude Opus 4.7 launch and the developer backlash surrounding Opus 4.6 degradation. The live demonstrations in VS Code, the desktop client, and the web app are authentic, displaying real multi-turn prompts and tangible deliverables alongside official benchmark tables and system card documentation.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.