Claude Opus-4.7 Just Dropped, And...
Nick Saraev · 2026-05-02 · community · 100,330 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
Content creator Nick Saraev analyzes the newly released Claude Opus 4.7 benchmark scorecard published by Anthropic, comparing its metrics against Claude Opus 4.6, GPT-5.4, Gemini 3.1 Pro, and the unreleased Mythos Preview. Saraev frames Opus 4.7 as a stepping-stone model designed to provide safe incremental performance gains without releasing the full cyber-risk-sensitive capabilities of Mythos. He concludes by offering strategic commentary on model commoditization and advising against overhauling production infrastructure for marginal benchmark improvements.
What is shown
- [00:00] Nick Saraev speaking directly to the camera introducing the release of Claude Opus 4.7.
- [00:14] Anthropic's official comparison benchmark scorecard for Opus 4.7 displayed on screen.
- [00:58] Presenter drawing a diagram on screen illustrating Opus 4.7 as an intermediate "half step" between Opus 4.6 and Mythos Preview.
- [01:34] Zoomed-in review of the benchmark table covering coding, terminal coding, and general reasoning evaluations.
- [03:28] Presenter sketching an S-curve to argue that benchmark saturation accelerates rapidly once models reach ~50%.
- [04:11] Review of tool use, computer use, financial analysis, cybersecurity, GPQA Diamond, and visual reasoning metrics.
- [05:22] Presenter speaking to the camera reflecting on personal automation workflows, historical AI progress from GPT-3 (2020), and engineering tradeoffs.
Claims & numbers
- The presenter claims OpenAI's next model, codenamed "Spud" (GPT-5.5), is likely to release within a few days of Opus 4.7.
- On SWE-bench Pro (agentic coding), the table reports Opus 4.7 scored 64.3% versus Opus 4.6 at 53.4%, GPT-5.4 at 57.7%, Gemini 3.1 Pro at 54.2%, and Mythos Preview at 77.8%.
- On SWE-bench Verified, Opus 4.7 is listed at 87.6% (Opus 4.6: 80.8%, GPT-5.4: self-reported 75.1%, Gemini 3.1 Pro: 80.6%, Mythos: 93.9%).
- On Terminal-Bench 2.0 (agentic terminal coding), Opus 4.7 scored 69.4% versus Opus 4.6 at 65.4%, GPT-5.4 at 75.1%, Gemini 3.1 Pro at 68.5%, and Mythos at 82.0%.
- On Humanity's Last Exam (multidisciplinary reasoning), Opus 4.7 scored 46.9% without tools and 54.7% with tools (compared to Opus 4.6 at 40.0% / 53.3%, GPT-5.4 at 42.7% / 58.7%, Gemini 3.1 Pro at 44.4% / 51.4%, and Mythos Preview at 56.8% / 64.7%).
- On BrowseComp (agentic search), Opus 4.7 scored 79.3%, which regressed compared to Opus 4.6's 83.7% (GPT-5.4: 89.3%, Gemini 3.1 Pro: 85.9%, Mythos: 86.9%).
- On MCP-Atlas (scaled tool use), Opus 4.7 scored 77.3% versus Opus 4.6 at 75.8% and GPT-5.4 at 66.1%.
- On OSWorld-Verified (agentic computer use), Opus 4.7 achieved 78.0% compared to Opus 4.6 at 72.7% and Mythos at 79.6%.
- On Finance-Agent v1, Opus 4.7 scored 64.4% compared to Opus 4.6 at 60.1% (+4.3%).
- On CyberGym (cybersecurity vulnerability reproduction), Opus 4.7 scored 73.1% compared to Opus 4.6 at 73.8% and Mythos at 83.1%.
- On GPQA Diamond, Opus 4.7 reached 94.2% versus Opus 4.6 at 91.3% and Mythos at 94.6%.
- On CharXiv Reasoning (visual reasoning), Opus 4.7 scored 82.1% without tools and 91.5% with tools, up from Opus 4.6's 69.1% without tools and 84.7% with tools.
- On MGSM (multilingual Q&A), Opus 4.7 scored 91.5% versus Opus 4.6 at 91.1%.
- The presenter states that using modern AI models like Opus 4.6, he can generate high-quality customized outreach for over 5,000 businesses in an hour, compared to reaching 10 to 15 businesses when doing manual outreach seven years prior.
Notable quotes
- [01:04] "What they've done is they basically provided us sort of like a mid-tier, okay, halfway between 4.6 and Mythos."
- [02:11] "My take on how Opus 4.7 was trained is it's probably Mythos Preview just distilled, basically dummified down a little bit and running on a lot faster and better hardware."
- [08:11] "My main take is that AI does not make things possible anymore; it just makes things slightly more profitable anymore."
Assessment
This is an independent community commentary and benchmark review video, not an official product demo or announcement. The presenter does not run live software evaluations during the video, relying entirely on Anthropic's published benchmark scorecard table to discuss performance and industry implications.
Described by gemini-3.8-flash on 2026-09-29 from the video's audio and frames.