Deceptive Delight: Jailbreak LLMs Through Camouflage and Distraction
ID: e8ae27ec-68a4-5ffe-9307-5b31b40084d7
STIX ID: report--e8ae27ec-68a4-5ffe-9307-5b31b40084d7
Feed Name: Palo Alto Networks Unit 42
Unit 42 presents “Deceptive Delight,” a simple multi-turn LLM jailbreak that blends unsafe topics with benign content to elicit harmful outputs, achieving about a 65% attack success rate within three turns across eight models. Using an LLM-based judge to score harmfulness and quality, the study shows highest impact at turn three, varying effectiveness by harmful content category, and diminishing returns beyond three turns. The report concludes with mitigation strategies, including enabling content filters and applying defensive prompt engineering, to strengthen resilience against similar jailbreaks.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
