Researchers Use Poetry to Jailbreak AI Models
ID: ba0e0905-e34f-5a41-a6dd-7eecd2f29d26
STIX ID: report--ba0e0905-e34f-5a41-a6dd-7eecd2f29d26
Feed Name: Dark Reading
**Executive Summary:** Researchers demonstrated an 'AI poem jailbreak' where reframing risky prompts as poetry significantly reduces refusal behavior in large language models; in experiments across 20+ models and 1,200 prompts spanning 12 hazard categories, poetic prompts raised the average Attack Success Rate from 8.08% to 43.07% (with Deepseek at 72% and Google at 66%), exposing a gap in safety training for figurative and narrative language and prompting recommendations for further mechanistic study and tighter alignment and data-handling practices.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
