Booz Allen’s first Cyber Weapon Index
Of 18 US and Chinese models run as autonomous attackers, only Claude Mythos completes the full kill chain
Partly confirmed
The takeaway
In early September 2026 Booz Allen Hamilton published its first Cyber Weapon Index, which ran 18 frontier models (nine US, nine Chinese) as fully autonomous attackers against a production-grade enterprise network.
Status
- Claim
Partly confirmed
- Our reporting
- Medium confidence
- Importance
- 3 of 5
- Last verified
- 10 October 2026
Your AI and this story
- GPT-6 Astra125 days after its cutoff
- Claude Opus 5.564 days after its cutoff
- Gemini 3.8 Flash155 days after its cutoff
- Grok 4.794 days after its cutoff
None of these four assistants can know about it. The closest, Claude Opus 5.5, stops 64 days before it.
Key facts
- 18 models, nine US and nine Chinese, open and closed weights, run as autonomous attackers controlling real machines against a production-grade enterprise network; scored from network telemetry, host and domain-controller logs, intrusion-detection sensors and attacker transcripts, not model claims (Booz Allen)
- Index = Vulnerability Research Score (finding planted and genuine bugs in compiled code without source) + Kill Chain Attainment Score (progress through a live Active Directory environment, with and without credentials)
- Scores (The Register): Claude Mythos 80 (VRS 74, KCAS 86 per secondary reports), Grok 4.5 49, GPT-5.6 Sol 46
- Mythos with stolen employee credentials gained administrator-level control ‘in every attempt’ and reached full domain compromise even without credentials (The Register)
- Other models (secondary reports): four reached full domain access, four achieved lateral movement, two reached credential access; all but one penetrated the network
- All nine frontier API models scored zero against genuine (unplanted) vulnerabilities while doing near-perfectly on planted ones; only Mythos exploited real bugs (The Register)
- Harness effect: ‘when paired with an attack harness, Claude Sonnet rivaled Claude Mythos’ performance’ (The Register quoting the report)
- Update on Booz Allen’s page: OpenAI’s GPT-6 Astra also demonstrated full kill-chain execution within a week of publication, and five more models entered the top 18
Show 2 more
- Booz Allen also launched Vellox Labs Guile, a ‘counter-AI’ product that manipulates what AI attackers see; vendor-reported >95% reduction in autonomous-attacker success (not independently validated)
- Oct 6, 2026 (claude.com blog): Booz Allen used Claude Mythos Preview via Project Glasswing; one analyst reviewed 8 production systems across 138 repositories in 12 days, work Booz Allen said would otherwise take a larger team several months; Comcast assessed 258 business-critical systems and ~170M lines of code and found a critical authentication-bypass bug in a public-facing platform
What happened
Booz Allen Hamilton, the US defence and intelligence contractor, published its first Cyber Weapon Index (CWI) in early September 2026. Secondary sources date the release to about Sept 2 and the first articles appeared on Sept 3; Booz Allen’s page says only “September 2026”. The index ran 18 large language models, half from the US and half from China, as fully autonomous attackers on real machines against a production-grade enterprise network. Every action was checked against the network’s own logs and sensors.
Claude Mythos was the only model to complete the whole kill chain on its own, from finding a vulnerability to administrator-level control of the domain. It scored 80, ahead of Grok 4.5 (49) and GPT-5.6 Sol (46), according to The Register. Booz Allen also found that the “attack harness” (tools, memory and orchestration) matters about as much as the model, since a harnessed Claude Sonnet came close to Mythos. Booz Allen later updated the page: within a week OpenAI’s GPT-6 Astra had also executed the full kill chain.
The model versions behind several names (for example which Mythos release was tested) are not stated in the sources we could read.
Why it matters
It is one of the first third-party measurements, by a major government contractor, of fully autonomous AI intrusion into a realistic network. The report’s forecast that most models will reach this level within six months fed directly into the autumn 2026 debate on cyber capability thresholds and on gating models like Mythos behind verification programs.
Sources
5 sources from 5 sites. Numbers match the chips in the text.
5 sources: 2 primary, 3 press
Primary
- Booz Allen: Cyber Weapon Index (The Offensive Frontier: AI as the Attacker)boozallen.com, official
- Claude blog (Oct 6): How Comcast and Booz Allen use Claude Mythos to find exploit chainsclaude.com, official
Press
- The Register: Booz Allen’s Cyber Weapon Index, Claude Mythos leadstheregister.com, press
- SC World: AI models show increasing capability for autonomous cyberattacks, report warnsscworld.com, press
- Let’s Data Science (Sept 3): Booz Allen’s Cyber Weapon Index shows why the attack harness mattersletsdatascience.com, press
Changes
- Filed (Claude blog Oct 6 on Comcast and Booz Allen); exact publication day unverified