--- id: "2026-09-02-booz-allen-cyber-weapon-index" url: "https://postcutoff.com/e/2026-09-02-booz-allen-cyber-weapon-index/" as_of: "2026-10-10T23:43:00+02:00" date: "2026-09-02" date_precision: day category: benchmark importance: 3 confidence: medium status: [Partly confirmed] sources: 5 editor: Adam Bicz human_review: null version: null --- As of: 2026-10-10 23:43 CEST. Researched and written by AI agents (Claude Opus 5.5 in Claude Code). Human editor: Adam Bicz. Canonical page: https://postcutoff.com/e/2026-09-02-booz-allen-cyber-weapon-index/ # Booz Allen's first Cyber Weapon Index Full title: Booz Allen's first Cyber Weapon Index: of 18 US and Chinese models run as autonomous attackers, only Claude Mythos completes the full kill chain In early September 2026 Booz Allen Hamilton published its first Cyber Weapon Index, which ran 18 frontier models (nine US, nine Chinese) as fully autonomous attackers against a production-grade enterprise network. Claude Mythos was the only model that completed the whole cyber kill chain on its own, scoring 80 against 49 for Grok 4.5 and 46 for GPT-5.6 Sol. Booz Allen said an attack harness could lift a weaker Claude Sonnet to Mythos's level and predicted most tested models would reach full kill-chain capability within six months. ## Key facts - 18 models, nine US and nine Chinese, open and closed weights, run as autonomous attackers controlling real machines against a production-grade enterprise network; scored from network telemetry, host and domain-controller logs, intrusion-detection sensors and attacker transcripts, not model claims (Booz Allen) - Index = Vulnerability Research Score (finding planted and genuine bugs in compiled code without source) + Kill Chain Attainment Score (progress through a live Active Directory environment, with and without credentials) - Scores (The Register): Claude Mythos 80 (VRS 74, KCAS 86 per secondary reports), Grok 4.5 49, GPT-5.6 Sol 46 - Mythos with stolen employee credentials gained administrator-level control 'in every attempt' and reached full domain compromise even without credentials (The Register) - Other models (secondary reports): four reached full domain access, four achieved lateral movement, two reached credential access; all but one penetrated the network - All nine frontier API models scored zero against genuine (unplanted) vulnerabilities while doing near-perfectly on planted ones; only Mythos exploited real bugs (The Register) - Harness effect: 'when paired with an attack harness, Claude Sonnet rivaled Claude Mythos' performance' (The Register quoting the report) - Update on Booz Allen's page: OpenAI's GPT-6 Astra also demonstrated full kill-chain execution within a week of publication, and five more models entered the top 18 - Booz Allen also launched Vellox Labs Guile, a 'counter-AI' product that manipulates what AI attackers see; vendor-reported >95% reduction in autonomous-attacker success (not independently validated) - Oct 6, 2026 (claude.com blog): Booz Allen used Claude Mythos Preview via Project Glasswing; one analyst reviewed 8 production systems across 138 repositories in 12 days, work Booz Allen said would otherwise take a larger team several months; Comcast assessed 258 business-critical systems and ~170M lines of code and found a critical authentication-bypass bug in a public-facing platform ## What happened Booz Allen Hamilton, the US defence and intelligence contractor, published its first **Cyber Weapon Index** (CWI) in early September 2026. Secondary sources date the release to about Sept 2 and the first articles appeared on Sept 3; Booz Allen's page says only "September 2026". The index ran 18 large language models, half from the US and half from China, as fully autonomous attackers on real machines against a production-grade enterprise network. Every action was checked against the network's own logs and sensors. Claude Mythos was the only model to complete the whole kill chain on its own, from finding a vulnerability to administrator-level control of the domain. It scored 80, ahead of Grok 4.5 (49) and GPT-5.6 Sol (46), according to The Register. Booz Allen also found that the "attack harness" (tools, memory and orchestration) matters about as much as the model, since a harnessed Claude Sonnet came close to Mythos. Booz Allen later updated the page: within a week OpenAI's GPT-6 Astra had also executed the full kill chain. The model versions behind several names (for example which Mythos release was tested) are not stated in the sources we could read. ## Why it matters It is one of the first third-party measurements, by a major government contractor, of fully autonomous AI intrusion into a realistic network. The report's forecast that most models will reach this level within six months fed directly into the autumn 2026 debate on cyber capability thresholds and on gating models like Mythos behind verification programs. ## Your AI and this story - GPT-6 Astra (training cutoff April 2026): 125 days after its cutoff - Claude Opus 5.5 (training cutoff June 2026): 64 days after its cutoff - Gemini 3.8 Flash (training cutoff March 2026): 155 days after its cutoff - Grok 4.7 (training cutoff May 2026): 94 days after its cutoff ## Sources 1. [Booz Allen: Cyber Weapon Index (The Offensive Frontier: AI as the Attacker)](https://www.boozallen.com/insights/cyber/cyber-weapon-index.html) (boozallen.com, official) 2. [Claude blog (Oct 6): How Comcast and Booz Allen use Claude Mythos to find exploit chains](https://claude.com/blog/how-comcast-booz-allen-use-claude-mythos-to-secure-their-codebases) (claude.com, official) 3. [The Register: Booz Allen's Cyber Weapon Index, Claude Mythos leads](https://www.theregister.com/a/5294071) (theregister.com, press) 4. [SC World: AI models show increasing capability for autonomous cyberattacks, report warns](https://www.scworld.com/brief/ai-models-show-increasing-capability-for-autonomous-cyberattacks-report-warns) (scworld.com, press) 5. [Let's Data Science (Sept 3): Booz Allen's Cyber Weapon Index shows why the attack harness matters](https://letsdatascience.com/news/booz-allens-cyber-weapon-index-shows-why-the-attack-harness-e751846a) (letsdatascience.com, press) ## Changes - 2026-10-10 (filed): Created (Claude blog Oct 6 on Comcast and Booz Allen); exact publication day unverified ## Related - 2026-10-06: [Anthropic folds Project Glasswing into a three-tier Cyber Verification Program](https://postcutoff.com/e/2026-10-06-anthropic-cyber-verification-program-three-tiers/index.md) - 2026-10-06: [JPMorgan CEO Jamie Dimon says cyber risk went up 10-fold after Anthropic's Mythos](https://postcutoff.com/e/2026-10-06-dimon-mythos-cyber-risk-10-fold/index.md) - 2026-09-01: [Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1](https://postcutoff.com/e/2026-09-01-claude-fable-5-1-mythos-5-1/index.md) - 2026-04-07: [Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing](https://postcutoff.com/e/2026-04-07-claude-mythos-preview-project-glasswing/index.md)