Is Claude Opus 5.5 Nerfed? I Built NerfBench to Find Out.
BridgeMind · 2026-10-04 · review · 26,726 views
What's in the video
Description written by Gemini, which watched and listened to the whole video.
Summary
The presenter from BridgeMind officially introduces NerfBench (hosted on BridgeBench), a benchmark system designed to track whether frontier AI models are degraded ("nerfed") over time compared to their launch baselines. He explains how NerfBench calculates "power" across performance, token usage, and cost, demonstrates a live 150-attempt benchmark run of Claude Opus 5.5, and announces direct harness connectors (such as Claude Code and Codex subscriptions) alongside API testing.
What is shown
- [00:00 - 00:30] Introduction to NerfBench with user complaint tweets regarding perceived nerfs to Claude Opus 5.5 and GPT-6 Astra, alongside the BridgeBench overview dashboard.
- [00:30 - 00:46] Footage of a previous test where Claude Sonnet 5.5 one-shot generated a functional 3D Mario Kart–style racing game.
- [01:27 - 03:45] Detailed review of Claude Opus 5.5 historical tracking data: Sept 22 launch baseline (100%), Sept 27 (99.2%), Oct 1 (103.8%), and Oct 2 (94.2%, a 9.6% drop between tests), explaining the ±10% normal variance threshold.
- [03:45 - 05:08] Inspection of the 3D voxel attempt visualizer for the Oct 2 run (114/150 passed) and Sept 22 run (120/150 passed), detailing specific test tasks (Task 43, Task 42, Task 46).
- [05:09 - 07:38] Setup and live concurrent execution of a 150-attempt run on Claude Opus 5.5 (via OpenRouter, high effort) on Oct 4, 2026, completing in ~7 minutes and scoring 117/150 passed (96.5% power).
- [07:38 - 09:05] Breakdown of the 50-task suite across 8 software engineering categories, showcasing the methodology article and Task 41 ("repair a payment ledger").
- [09:06 - 11:16] Demonstration of the newly added "Connections" feature, allowing models to be benchmarked directly through subscription harnesses (Claude Code CLI and Codex) rather than solely through API routers.
- [11:16 - 11:54] Outro promoting the BridgeMind community and Discord.
Claims & numbers
- The presenter states that NerfBench launched on the previous Sunday (September 27, 2026) and had completed 12 runs across four models (Claude Opus 5.5, Claude Sonnet 5.5, GPT-6 Astra, and GPT-6.1 Sol) in its first week.
- NerfBench tests 50 coding and reasoning tasks across 8 categories (bug fixing, code reasoning, code writing, algorithms, strings, testing, SQL, security) with 3 repeats per task (150 total attempts per benchmark run).
- Model "power" is calculated from task accuracy/performance, token consumption, and dollar cost relative to the frozen launch baseline session (defined as 100%).
- NerfBench defines normal variance as 90% to 110% (±10%); scores outside this threshold qualify as a model nerf or significant buff.
- The presenter reports historical scores for Claude Opus 5.5: Launch (Sept 22): 100%; Sept 27: 99.2% (-0.8%); Oct 1: 103.8% (+3.8%); Oct 2: 94.2% (-5.8% vs launch, a -9.6% drop vs Oct 1).
- The Oct 4 live run yielded 117/150 passed attempts, resulting in a 96.5% power score (-3.5% vs launch, +2.3% vs previous test), demonstrating variance within normal bounds rather than a true nerf.
- The presenter claims the BridgeMind "vibe coding" Discord community has over 18,000 builders.
Notable quotes
- [00:00] "Today, I am officially introducing NerfBench: my AI benchmark that measures whether or not AI models are getting nerfed."
- [03:20] "Now, one thing that is very important to understand is that that is not a nerf. There is model variance, and a fluctuation of 9.6% is a lot, but it's not enough to be able to say that the model is actually nerfed."
- [11:03] "I think that this is going to be the future of NerfBench, because what people are more interested in is how the models perform on subscriptions via the harnesses that we use day in and day out..."
Assessment
This is a genuine product demonstration and technical methodology walk-through by the creator of NerfBench (BridgeMind). The live 150-task benchmark run, concurrent streaming, and user interface features are shown operating in real time with actual UI interactions.
Described by gemini-3.8-flash on 2026-10-05 from the video's audio and frames.