Anthropic: infrastructure settings alone move Terminal-Bench 2.0 scores by 6 points
Confirmed
Status
- Claim
Confirmed
- Our reporting
- High confidence
- Importance
- 2 of 5
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 AstraIn its training data
- Claude Opus 5.5In its training data
- Gemini 3.8 FlashIn its training data
- Grok 4.7In its training data
It happened before the cutoffs of all four assistants, so it can be in their training data.
Key facts
- Terminal-Bench 2.0: 6 pp gap between most and least resourced configurations (p < 0.01)
- Strict resource enforcement: 5.8% infrastructure error rate vs 0.5% uncapped
- SWE-bench: 1x to 5x RAM changed scores by 1.54 pp
- Recommendation: report resource configuration as a first-class experimental variable
What happened
The post (contributors include Nicholas Carlini, Jeremy Hadfield, Mike Merrill and Alex Shaw) argues that “infrastructure configuration alone can produce differences that exceed those margins” between leading models on agentic benchmarks.
Why it matters
It is a caution for reading 2026 agentic leaderboard gaps of a few points, which are often within this infrastructure noise.
Sources
1 source from 1 site. Numbers match the chips in the text.
1 source: 1 primary
Primary
- Anthropic Engineering: Quantifying infrastructure noise in agentic coding evalsanthropic.com, official
Changes
- Filed