Three changes degraded Claude Code quality in March–April 2026
Usage limits reset
Confirmed
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 3 of 5
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 AstraIn its training data
- Claude Opus 5.5In its training data
- Gemini 3.8 Flash23 days after its cutoff
- Grok 4.7In its training data
It is after the cutoff of Gemini 3.8 Flash, and can be in the training data of GPT-6 Astra, Claude Opus 5.5 and Grok 4.7.
Key facts
- Issue 1 (Mar 4 – Apr 7): default reasoning effort lowered from high to medium for Sonnet 4.6 and Opus 4.6 to cut latency; reverted
- Issue 2 (Mar 26 – Apr 10, fixed in v2.1.101): thinking-cache bug cleared reasoning history on every turn, making Claude ‘seem forgetful and repetitive’
- Issue 3 (Apr 16 – Apr 20, fixed in v2.1.116): system-prompt word-limit instruction hurt coding quality (Sonnet 4.6, Opus 4.6, Opus 4.7)
- API and inference layer unaffected: ‘We never intentionally degrade our models’
- Usage limits reset for all subscribers on Apr 23; promised per-model evals, ablations, gradual rollouts and soak periods for prompt changes
What happened
Anthropic wrote: “The implementation had a bug. Instead of clearing thinking history once, it cleared it on every turn for the rest of the session.” Together with the effort change and the prompt change, this explained much of the reported degradation.
Why it matters
It confirmed that “the model got dumber” complaints can come from harness, prompt and default changes rather than the weights, and set a precedent for public postmortems on agent product quality.
Sources
1 source from 1 site. Numbers match the chips in the text.
1 source: 1 primary
Primary
- Anthropic Engineering: An update on recent Claude Code quality reportsanthropic.com, official
Changes
- Filed