{"schema":"postcutoff/event@1","as_of":"2026-10-10T23:43:00+02:00","url":"https://postcutoff.com/e/2026-02-05-anthropic-infrastructure-noise-agentic-evals/","md":"https://postcutoff.com/e/2026-02-05-anthropic-infrastructure-noise-agentic-evals/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"id":"2026-02-05-anthropic-infrastructure-noise-agentic-evals","date":"2026-02-05","date_precision":"day","short_title":"Anthropic: infrastructure settings alone move Terminal-Bench 2.0 scores by 6 points","deck":null,"takeaway":null,"category":"benchmark","category_label":"Benchmarks","importance":2,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":[{"n":1,"title":"Anthropic Engineering: Quantifying infrastructure noise in agentic coding evals","url":"https://www.anthropic.com/engineering/infrastructure-noise","type":"official","group":"primary","domain":"anthropic.com"}],"official":1,"filed":"2026-10-10","updated":"2026-10-10","orgs":["Anthropic"],"title":"Anthropic: infrastructure settings alone move Terminal-Bench 2.0 scores by 6 points","summary":"An Anthropic engineering post (Gian Segato et al., Feb 5, 2026) measured how container resources change agentic coding scores: a 6-percentage-point gap on Terminal-Bench 2.0 between the most and least resourced setups, larger than the margins that separate top models on leaderboards.","key_facts":["Terminal-Bench 2.0: 6 pp gap between most and least resourced configurations (p < 0.01)","Strict resource enforcement: 5.8% infrastructure error rate vs 0.5% uncapped","SWE-bench: 1x to 5x RAM changed scores by 1.54 pp","Recommendation: report resource configuration as a first-class experimental variable"],"key_numbers":[],"tags":["evals","benchmarks","terminal-bench","swe-bench","agents","methodology"],"science":null,"body_md":"## What happened\n\nThe post (contributors include Nicholas Carlini, Jeremy Hadfield, Mike Merrill and Alex Shaw) argues that \"infrastructure\nconfiguration alone can produce differences that exceed those margins\" between leading models on agentic benchmarks.\n\n## Why it matters\n\nIt is a caution for reading 2026 agentic leaderboard gaps of a few points, which are often within this infrastructure noise.","disputed":[],"related":[{"id":"2026-02-05-anthropic-parallel-claudes-c-compiler","url":"https://postcutoff.com/e/2026-02-05-anthropic-parallel-claudes-c-compiler/","date":"2026-02-05","date_precision":"day","short_title":"Anthropic: 16 parallel Claude Opus 4.6 agents build a 100,000-line C compiler that compiles Linux 6.9","deck":null,"takeaway":"It became the reference demo for multi-agent \"agent teams\" on long software projects.","category":"agents","category_label":"Agents","importance":3,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":1,"official":1,"filed":"2026-10-10","updated":"2026-10-10","orgs":["Anthropic"]}],"people":[{"id":"nicholas-carlini","name":"Nicholas Carlini","url":"https://postcutoff.com/person/nicholas-carlini/"}],"posts":[],"videos":[],"models":[],"changes":[{"date":"2026-10-10","type":"filed","text":"Created"}],"provenance":{"agents":[{"model":"Claude Opus 5.5","maker":"Anthropic","tool":"Claude Code"}],"filed":"2026-10-10","run":null,"sources_read":null,"updated":"2026-10-10","human_review":null,"version":null},"gaps":[{"model_id":"gpt-6-astra","name":"GPT-6 Astra","cutoff":"2026-04","days_after":null,"in_training_data":true},{"model_id":"claude-opus-5-5","name":"Claude Opus 5.5","cutoff":"2026-06","days_after":null,"in_training_data":true},{"model_id":"gemini-3-8-flash","name":"Gemini 3.8 Flash","cutoff":"2026-03","days_after":null,"in_training_data":true},{"model_id":"grok-4-7","name":"Grok 4.7","cutoff":"2026-05","days_after":null,"in_training_data":true}],"short_url":null}