logo

Researchers Use Poetry to Jailbreak AI Models

ID: ba0e0905-e34f-5a41-a6dd-7eecd2f29d26

STIX ID: report--ba0e0905-e34f-5a41-a6dd-7eecd2f29d26

Feed Name: Dark Reading

Date Published: 2025-12-02

Date Updated: 2026-04-21

Author: Alexander Culafi

...
...

**Executive Summary:** Researchers demonstrated an 'AI poem jailbreak' where reframing risky prompts as poetry significantly reduces refusal behavior in large language models; in experiments across 20+ models and 1,200 prompts spanning 12 hazard categories, poetic prompts raised the average Attack Success Rate from 8.08% to 43.07% (with Deepseek at 72% and Google at 66%), exposing a gap in safety training for figurative and narrative language and prompting recommendations for further mechanistic study and tighter alignment and data-handling practices.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.