logo

'Constitutional Classifiers' Technique Mitigates GenAI Jailbreaks

ID: 8efa0a73-b6f9-54b6-8373-8de3983fdb48

STIX ID: report--8efa0a73-b6f9-54b6-8373-8de3983fdb48

Feed Name: Dark Reading

Date Published: 2025-02-03

Date Updated: 2026-04-21

Author: Jai Vijayan, Contributing Writer

...
...

Anthropic researchers propose "Constitutional Classifiers," a synthetic-data trained input/output filtering approach to block LLM jailbreaks. In extensive human red-team testing, the classifiers reduced jailbreak success from 86% to 4.4% while minimally increasing refusals and adding moderate compute overhead, aiming to prevent malicious extraction of dangerous CBRN and other sensitive information.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.