{"schema":"postcutoff/event@1","as_of":"2026-10-10T23:43:00+02:00","url":"https://postcutoff.com/e/2026-09-02-booz-allen-cyber-weapon-index/","md":"https://postcutoff.com/e/2026-09-02-booz-allen-cyber-weapon-index/index.md","disclosure":{"written_by":"AI agents (Claude Opus 5.5 in Claude Code)","editor":"Adam Bicz","policy":"https://postcutoff.com/about/"},"license":null,"id":"2026-09-02-booz-allen-cyber-weapon-index","date":"2026-09-02","date_precision":"day","short_title":"Booz Allen's first Cyber Weapon Index","deck":"Of 18 US and Chinese models run as autonomous attackers, only Claude Mythos completes the full kill chain","takeaway":"In early September 2026 Booz Allen Hamilton published its first Cyber Weapon Index, which ran 18 frontier models (nine US, nine Chinese) as fully autonomous attackers against a production-grade enterprise network.","category":"benchmark","category_label":"Benchmarks","importance":3,"confidence":"medium","status":{"key":"partly","labels":["Partly confirmed"]},"sources":[{"n":1,"title":"Booz Allen: Cyber Weapon Index (The Offensive Frontier: AI as the Attacker)","url":"https://www.boozallen.com/insights/cyber/cyber-weapon-index.html","type":"official","group":"primary","domain":"boozallen.com"},{"n":2,"title":"Claude blog (Oct 6): How Comcast and Booz Allen use Claude Mythos to find exploit chains","url":"https://claude.com/blog/how-comcast-booz-allen-use-claude-mythos-to-secure-their-codebases","type":"official","group":"primary","domain":"claude.com"},{"n":3,"title":"The Register: Booz Allen's Cyber Weapon Index, Claude Mythos leads","url":"https://www.theregister.com/a/5294071","type":"press","group":"press","domain":"theregister.com"},{"n":4,"title":"SC World: AI models show increasing capability for autonomous cyberattacks, report warns","url":"https://www.scworld.com/brief/ai-models-show-increasing-capability-for-autonomous-cyberattacks-report-warns","type":"press","group":"press","domain":"scworld.com"},{"n":5,"title":"Let's Data Science (Sept 3): Booz Allen's Cyber Weapon Index shows why the attack harness matters","url":"https://letsdatascience.com/news/booz-allens-cyber-weapon-index-shows-why-the-attack-harness-e751846a","type":"press","group":"press","domain":"letsdatascience.com"}],"official":2,"filed":"2026-10-10","updated":"2026-10-10","orgs":["Booz Allen Hamilton","Anthropic"],"title":"Booz Allen's first Cyber Weapon Index: of 18 US and Chinese models run as autonomous attackers, only Claude Mythos completes the full kill chain","summary":"In early September 2026 Booz Allen Hamilton published its first Cyber Weapon Index, which ran 18 frontier models (nine US, nine Chinese) as fully autonomous attackers against a production-grade enterprise network. Claude Mythos was the only model that completed the whole cyber kill chain on its own, scoring 80 against 49 for Grok 4.5 and 46 for GPT-5.6 Sol. Booz Allen said an attack harness could lift a weaker Claude Sonnet to Mythos's level and predicted most tested models would reach full kill-chain capability within six months.","key_facts":["18 models, nine US and nine Chinese, open and closed weights, run as autonomous attackers controlling real machines against a production-grade enterprise network; scored from network telemetry, host and domain-controller logs, intrusion-detection sensors and attacker transcripts, not model claims (Booz Allen)","Index = Vulnerability Research Score (finding planted and genuine bugs in compiled code without source) + Kill Chain Attainment Score (progress through a live Active Directory environment, with and without credentials)","Scores (The Register): Claude Mythos 80 (VRS 74, KCAS 86 per secondary reports), Grok 4.5 49, GPT-5.6 Sol 46","Mythos with stolen employee credentials gained administrator-level control 'in every attempt' and reached full domain compromise even without credentials (The Register)","Other models (secondary reports): four reached full domain access, four achieved lateral movement, two reached credential access; all but one penetrated the network","All nine frontier API models scored zero against genuine (unplanted) vulnerabilities while doing near-perfectly on planted ones; only Mythos exploited real bugs (The Register)","Harness effect: 'when paired with an attack harness, Claude Sonnet rivaled Claude Mythos' performance' (The Register quoting the report)","Update on Booz Allen's page: OpenAI's GPT-6 Astra also demonstrated full kill-chain execution within a week of publication, and five more models entered the top 18","Booz Allen also launched Vellox Labs Guile, a 'counter-AI' product that manipulates what AI attackers see; vendor-reported >95% reduction in autonomous-attacker success (not independently validated)","Oct 6, 2026 (claude.com blog): Booz Allen used Claude Mythos Preview via Project Glasswing; one analyst reviewed 8 production systems across 138 repositories in 12 days, work Booz Allen said would otherwise take a larger team several months; Comcast assessed 258 business-critical systems and ~170M lines of code and found a critical authentication-bypass bug in a public-facing platform"],"key_numbers":[],"tags":["cybersecurity","offensive-cyber","benchmark","mythos","agents","harness","kill-chain"],"science":null,"body_md":"## What happened\n\nBooz Allen Hamilton, the US defence and intelligence contractor, published its first **Cyber Weapon Index** (CWI) in early September 2026.\nSecondary sources date the release to about Sept 2 and the first articles appeared on Sept 3; Booz Allen's page says only \"September 2026\".\nThe index ran 18 large language models, half from the US and half from China, as fully autonomous attackers on real machines against a\nproduction-grade enterprise network. Every action was checked against the network's own logs and sensors.\n\nClaude Mythos was the only model to complete the whole kill chain on its own, from finding a vulnerability to administrator-level control\nof the domain. It scored 80, ahead of Grok 4.5 (49) and GPT-5.6 Sol (46), according to The Register. Booz Allen also found that the\n\"attack harness\" (tools, memory and orchestration) matters about as much as the model, since a harnessed Claude Sonnet came close to Mythos.\nBooz Allen later updated the page: within a week OpenAI's GPT-6 Astra had also executed the full kill chain.\n\nThe model versions behind several names (for example which Mythos release was tested) are not stated in the sources we could read.\n\n## Why it matters\n\nIt is one of the first third-party measurements, by a major government contractor, of fully autonomous AI intrusion into a realistic\nnetwork. The report's forecast that most models will reach this level within six months fed directly into the autumn 2026 debate on\ncyber capability thresholds and on gating models like Mythos behind verification programs.","disputed":[],"related":[{"id":"2026-10-06-anthropic-cyber-verification-program-three-tiers","url":"https://postcutoff.com/e/2026-10-06-anthropic-cyber-verification-program-three-tiers/","date":"2026-10-06","date_precision":"day","short_title":"Anthropic folds Project Glasswing into a three-tier Cyber Verification Program","deck":"Reports 129K+ vulnerabilities found by partners","takeaway":"It said Glasswing partners found 129,000 verified vulnerabilities in April–July 2026 and its own open-source scanning another 5,500 through October, more than 33,000 of them critical or high severity.","category":"policy-safety","category_label":"Policy & safety","importance":3,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":7,"official":2,"filed":"2026-10-07","updated":"2026-10-10","orgs":["Anthropic"]},{"id":"2026-10-06-dimon-mythos-cyber-risk-10-fold","url":"https://postcutoff.com/e/2026-10-06-dimon-mythos-cyber-risk-10-fold/","date":"2026-10-06","date_precision":"day","short_title":"JPMorgan CEO Jamie Dimon says cyber risk went up 10-fold after Anthropic's Mythos","deck":null,"takeaway":"The head of the largest US bank, himself a Glasswing participant, put a number on how much frontier AI has changed cyber risk.","category":"policy-safety","category_label":"Policy & safety","importance":3,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":5,"official":0,"filed":"2026-10-07","updated":"2026-10-07","orgs":["JPMorgan Chase","Anthropic"]},{"id":"2026-09-01-claude-fable-5-1-mythos-5-1","url":"https://postcutoff.com/e/2026-09-01-claude-fable-5-1-mythos-5-1/","date":"2026-09-01","date_precision":"day","short_title":"Anthropic releases Claude Fable 5.1 and Claude Mythos 5.1","deck":null,"takeaway":"Fable/Mythos 5.1 was Anthropic's capability frontier until Opus 5.5 matched it three weeks later at less than half the price.","category":"model-release","category_label":"Model releases","importance":5,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":9,"official":5,"filed":"2026-09-29","updated":"2026-09-29","orgs":["Anthropic"]},{"id":"2026-04-07-claude-mythos-preview-project-glasswing","url":"https://postcutoff.com/e/2026-04-07-claude-mythos-preview-project-glasswing/","date":"2026-04-07","date_precision":"day","short_title":"Anthropic reveals Claude Mythos Preview, withholds it over cyber risk and launches Project Glasswing","deck":null,"takeaway":"Mythos Preview marked the point where a frontier lab judged a model's offensive cyber capability too dangerous for general release.","category":"model-release","category_label":"Model releases","importance":5,"confidence":"high","status":{"key":"confirmed","labels":["Confirmed"]},"sources":11,"official":7,"filed":"2026-09-29","updated":"2026-10-01","orgs":["Anthropic"]}],"people":[],"posts":[],"videos":[],"models":[],"changes":[{"date":"2026-10-10","type":"filed","text":"Created (Claude blog Oct 6 on Comcast and Booz Allen); exact publication day unverified"}],"provenance":{"agents":[{"model":"Claude Opus 5.5","maker":"Anthropic","tool":"Claude Code"}],"filed":"2026-10-10","run":null,"sources_read":null,"updated":"2026-10-10","human_review":null,"version":null},"gaps":[{"model_id":"gpt-6-astra","name":"GPT-6 Astra","cutoff":"2026-04","days_after":125,"in_training_data":false},{"model_id":"claude-opus-5-5","name":"Claude Opus 5.5","cutoff":"2026-06","days_after":64,"in_training_data":false},{"model_id":"gemini-3-8-flash","name":"Gemini 3.8 Flash","cutoff":"2026-03","days_after":155,"in_training_data":false},{"model_id":"grok-4-7","name":"Grok 4.7","cutoff":"2026-05","days_after":94,"in_training_data":false}],"short_url":null}