Anthropic publishes August 2026 Risk Report under its RSP
In August 2026 Anthropic published its second RSP Risk Report (186 pages, covering models as of July 15, 2026). It raised misalignment risk in high-stakes settings from 'very low' to 'low', disclosed an eleven-month gap in a chem/bio classifier, and said automated-R&D evaluations are saturating. Offensive cyber, driven by a UK AISI evaluation of Mythos 5, was the heaviest driver of change.
Key facts
- Published August 2026 (exact day not verified); covers Anthropic's models and actions as of July 15, 2026
- 186 pages; second Risk Report
- Misalignment in high-stakes settings: 'very low' -> 'low'
- Disclosed an eleven-month CB classifier gap
- Opus 5.5 system card cites it for recursive-self-improvement concerns and the overall 'low' misalignment-risk assessment
What happened
Risk reports are Anthropic's periodic, cross-model risk assessments under its RSP and Frontier Compliance Framework (FCF). System cards now describe how each new model changes the latest report's conclusions.
Why it matters
This is the baseline risk assessment against which Opus 5.5 and later 2026 models were judged.
Changelog
- 2026-09-29: created
Related posts (1)
- Anthropic publishes its second RSP Risk Report (August 2026) Anthropic @AnthropicAI · x · 2026-08-14
Anthropic's second regular Responsible Scaling Policy Risk Report on catastrophic-risk levels of its systems and its preparedness.
Related events
- Anthropic releases Claude Opus 5.5 — Fable-5.1-level performance at $4/$20, first model of the Claude 5.5 family ★★★★★
- Anthropic discloses Claude models breached real organizations during misconfigured cyber evaluations ★★★★★
Sources (4)
- officialRisk Report: August 2026 (Anthropic)
- officialAnthropic Responsible Scaling Policy
- discussionZvi Mowshowitz: Anthropic Risk Report August 2026
- discussionai.rud.is: reading the August 2026 Risk Report for the cybers
id: 2026-08-01-anthropic-risk-report-august-2026 · updated 2026-09-29 · open in the interactive timeline