How LLM jailbreaking can bypass AI security with multi-turn attacks
ID: ac9d023a-5de6-5bd8-b130-3ec2e86e4e52
STIX ID: report--ac9d023a-5de6-5bd8-b130-3ec2e86e4e52
Feed Name: Giskard
### Executive Summary This report analyzes multi-turn LLM jailbreaking (the "Crescendo" method), where attackers distribute malicious intent across several conversational turns or via an automated agent to evade single-message moderation. It demonstrates the attack progression with a ZephyrBank reputational example, outlines an automated adversarial loop (probe, backtrack, escalate), and recommends conversation-aware testing, constrained conversation lengths, and continuous adversarial red-teaming to detect and mitigate these risks.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
