Claude Opus 4.7 - A New Frontier, in Performance … and Drama
AI Explained · 2026-05-02 · community · 89,415 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
In this video, presenter Phillip (creator of the channel AI Explained) breaks down the launch of Anthropic's Claude Opus 4.7 and the accompanying drama surrounding its performance, compute constraints, and safety evaluations. He reviews official and third-party benchmark results, analyzes internal system card disclosures regarding Opus 4.7 and the unreleased Claude Mythos Preview, and examines the long-standing corporate and personal rivalry between Anthropic (led by Dario Amodei) and OpenAI (led by Sam Altman and Greg Brockman).
What is shown
- [00:13] Official Anthropic capability table comparing Claude Opus 4.7 against Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and Claude Mythos Preview across multiple agentic benchmarks.
- [00:56] Benchmark leaderboards on SimpleBench, METR Time Horizons, and "Humanity's Last Exam", highlighting Opus 4.7's lower score on SimpleBench (62.9%) compared to Opus 4.6 (67.6%).
- [01:40] Presenter demonstrating his web app (
lmcouncil.ai), noting that Opus 4.7 unexpectedly failed to automatically attach the router tooltip when updating the leaderboard code. - [03:04] Anthropic system card graphs comparing long-context reasoning (GraphWalks and MRCR v2 8-needle @ 1M tokens), showing an MRCR score regression to 32.2% for Opus 4.7 max.
- [03:30] Anthropic benchmarks for office knowledge work (GDPval-AA) and visual navigation (ScreenSpot-Pro).
- [04:21] LlamaIndex ParseBench OCR comparison table showing Opus 4.7 scoring 63.3% versus Gemini 3 Flash's 71.1%.
- [05:00] ARC-AGI-2 cost-versus-accuracy scatter plot and Vibe Code Bench v1.1 rankings (Opus 4.7 taking #1 at 71.09%).
- [05:22] Similarweb GenAI website traffic share chart up to March 2026, alongside leaked excerpts of an internal OpenAI memo reported by The Verge.
- [06:14] The Claude UI showing the mandatory "Adaptive thinking" toggle and settings, alongside tweets discussing rate limit throttles and reduced thinking tokens.
- [08:10] Excerpts from Anthropic system cards detailing an opt-in Slack poll of 130 employees on Mythos Preview productivity uplifts and listed model shortcomings (safeguard circumvention, code overwrites, fabrication).
- [12:00] System card report on Claude Mythos Preview evaluating Anthropic's own alignment assessment draft via internal Slack access.
- [13:16] Anthropic product updates for Claude Code and Cowork: automated Routines, the
/ultrareviewterminal command, and phone-based Dispatch. - [14:34] Live test of AssemblyAI's Universal-3 Pro Streaming speech-to-text model accurately transcribing spoken text with numbers and accents.
- [15:04] Excerpts from a Wall Street Journal investigation by Keach Hagey detailing the history of tensions between Dario Amodei, Greg Brockman, and Sam Altman at OpenAI from 2016 to 2020.
- [17:52] Video clip of Greg Brockman interviewing with Alex Kantrowitz on the Big Technology Podcast, discussing OpenAI's coding model focus versus Anthropic's real-world repository approach.
Claims & numbers
- The presenter says Claude Opus 4.7 was released on April 16, 2026, and scores 64.3% on SWE-bench Pro, 87.6% on agentic coding, and 79.3% on agentic search (BrowseComp), where it fell behind Opus 4.6 (83.7%).
- On SimpleBench, the presenter states Opus 4.7 scored 62.9%, below Opus 4.6's 67.6%, because adaptive thinking spent less compute on trick questions it misjudged as easy.
- On the MRCR v2 (8-needle at 1M tokens) needle-in-a-haystack test, the presenter notes Opus 4.7 reached only 32.2% compared to Opus 4.6's 78.3%.
- On GDPval-AA knowledge work, the presenter reports Opus 4.7 scored 1,753, beating Opus 4.6 (1,619), GPT-5.4 (1,674), and Gemini 3.1 Pro (1,314).
- On ParseBench, the presenter shows Opus 4.7 scored 63.3% at $7.14 per page, trailing Gemini 3 Flash's 71.1% at $0.65 per page.
- On ARC-AGI-2, the presenter shows Claude 4.7 (Max) scored 75.85% at $7.43 task cost, while on Vibe Code Bench v1.1 it placed #1 with 71.09% accuracy at $21.41 per task.
- Similarweb traffic data cited in the video indicates ChatGPT held ~56.7% market share, Gemini ~25.5%, and Claude ~6.0% as of March 2026, with OpenAI's share dropping toward 50%.
- A leaked OpenAI memo cited in the video claims Anthropic's annualized run rate of $30 billion is overstated by roughly $8 billion (placing it nearer $22 billion).
- The presenter reports that on Ventuals secondary markets, Anthropic's implied valuation crossed $1 trillion.
- Regarding the Mythos internal productivity poll, the presenter highlights that only 130 people responded in an opt-in, non-random Slack survey.
- The WSJ reporting cited states that in 2017, between 10% and 20% of OpenAI's 60-person staff were let go following an evaluation spreadsheet ordered by Elon Musk.
- Historical data presented illustrates US AI data center spending approaching ~1% of US GDP, rivaling the Apollo program and behind only the Marshall Plan and US railroad expansion.
Notable quotes
- [02:52] "During training we experimented with efforts to differentially reduce these capabilities." (quoting page 48 of the Anthropic Opus 4.7 System Card on cybersecurity vulnerability reproduction)
- [06:48] "We found that effort=85 was a sweet spot on the trade-off curve between token spend and task success... medium effort is now the default." (quoting Claude Code lead Boris Cherny on adaptive thinking defaults)
- [18:21] "We always had the best numbers on different programming competitions... but it's never seen someone's real-world codebase, which is messy... that is something that we were behind on." (Greg Brockman, [18:07]–[18:31])
Assessment
This is an independent analysis and review combining coverage of Anthropic's model release, official system cards, third-party benchmark evaluations, and investigative reporting on the AI industry. The presenter provides balanced, critical analysis, demonstrating personal testing quirks, scrutinizing methodology behind survey numbers, and contrasting marketing claims against empirical benchmark results.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.