Anthropic Details Claude Fable 5 Cyber Safeguards and AI Jailbreak Severity Framework
ID: de88f609-06d8-5f54-8343-b8d40b0e6ad8
STIX ID: report--de88f609-06d8-5f54-8343-b8d40b0e6ad8
Feed Name: Cyber Press
Anthropic published technical details on Claude Fable 5's cybersecurity safeguards and an early-draft five-band AI jailbreak severity framework: classifiers sort requests into four tiers (prohibited, high-risk dual use, low-risk dual use, benign) with a wider safety margin, and a combined-score severity model (capability gain, breadth, ease of weaponization, discoverability) sets a floor that can be raised for discretionary factors; Anthropic is soliciting feedback and running a HackerOne bounty for reported cyber-jailbreaks.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
