Anthropic details Fable 5's four-tier cyber classifiers and proposes a Cyber Jailbreak Severity (CJS) scale
On July 2, 2026 Anthropic explained how Claude Fable 5's cyber classifiers sort requests into four tiers: prohibited, high-risk dual use, low-risk dual use and benign. It also proposed an industry Cyber Jailbreak Severity scale from CJS-0 (informational) to CJS-4 (critical), scored on capability gain, breadth, ease of weaponisation and discoverability.
Key facts
- Tier 1, prohibited (ransomware, malware development, data exfiltration): blocked
- Tier 2, high-risk dual use (pen-testing, exploit development): blocked pending better access controls/authorization context
- Tier 3, low-risk dual use (e.g. vulnerability identification): sometimes blocked as a safety margin
- Tier 4, benign (patching, incident response): allowed with minimal false positives
- CJS-0 to CJS-4 severity scale over four axes: capability gain, breadth of gain, ease of weaponization, discoverability
What happened
Published about three weeks after Fable 5's launch, while users complained that cyber classifiers were over-blocking. The post describes the classifier policy and offers a shared vocabulary for rating jailbreaks.
Why it matters
It shows how a lab gates Mythos-class cyber capability in practice. The CJS scale is offered as a cross-industry standard for reporting jailbreak severity, which no one had before.
Unverified: whether other labs or bug-bounty programs adopted CJS.
Changelog
- 2026-10-01: created (leads run, from the Anthropic uncited-posts audit)
Related events
- Anthropic releases Claude Fable 5 and Claude Mythos 5 — first generally available Mythos-class model ★★★★★
Sources (1)
id: 2026-07-02-anthropic-fable-5-cyber-safeguards-jailbreak-severity · updated 2026-10-01 · open in the interactive timeline