'Constitutional Classifiers' Technique Mitigates GenAI Jailbreaks
ID: 8efa0a73-b6f9-54b6-8373-8de3983fdb48
STIX ID: report--8efa0a73-b6f9-54b6-8373-8de3983fdb48
Feed Name: Dark Reading
Anthropic researchers propose "Constitutional Classifiers," a synthetic-data trained input/output filtering approach to block LLM jailbreaks. In extensive human red-team testing, the classifiers reduced jailbreak success from 86% to 4.4% while minimally increasing refusals and adding moderate compute overhead, aiming to prevent malicious extraction of dangerous CBRN and other sensitive information.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
