logo

Whispering poetry at AI can make it break its own rules

ID: 9ba775ef-0620-51bc-8593-4a21df054734

STIX ID: report--9ba775ef-0620-51bc-8593-4a21df054734

Feed Name: Malwarebytes Blog

Date Published: 2025-12-02

Date Updated: 2026-04-28

...
...

A study by Icaro Lab with Sapienza University and DEXAI shows that reframing harmful prompts as poetry substantially increases jailbreak success across 25 LLMs, with hand-crafted poems achieving high success rates and AI-generated poems also outperforming prose by a large margin. The research highlights provider variance (DeepSeek and Google more susceptible; Anthropic and OpenAI safer), a tendency for smaller models to resist better, and urges inclusion of such adversarial tests in safety benchmarks and governance efforts under frameworks like the EU AI Act.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.