Tree of attacks (TAP): The automated method for jailbreaking LLMs
ID: 69676ac2-15c6-537d-beef-47cc468a7971
STIX ID: report--69676ac2-15c6-537d-beef-47cc468a7971
Feed Name: Giskard
Threat Score
This report explains Tree of Attacks with Pruning (TAP), an automated black‑box method for jailbreaking large language models by iteratively generating, evaluating, and pruning prompt variations to find prompts that bypass safety controls; it demonstrates a real‑world example where TAP produced a defamatory accusation letter and discusses the technique's implications for red‑teaming and enterprise AI security.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
