Cybersecurity researchers aren’t happy about the guardrails on Anthropic’s Fable
ID: 5e77938a-852f-5243-8f49-572acaf31859
STIX ID: report--5e77938a-852f-5243-8f49-572acaf31859
Feed Name: TechCrunch Security News
Anthropic released Fable, a limited public version of its cybersecurity-capable model Mythos, which includes strict guardrails that block or downgrade responses to cybersecurity and biology-related prompts. Researchers and practitioners have criticized the model for overzealous filtering (e.g., blocking code reviews and innocuous security-related requests); Anthropic offers a Cyber Verification Program to grant vetted cybersecurity professionals fewer limitations, similar to OpenAI’s Trusted Access for Cyber.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
