Many-shot jailbreaking
ID: ac9e840b-d176-5e4e-a95d-ebe9d28a212c
STIX ID: report--ac9e840b-d176-5e4e-a95d-ebe9d28a212c
Feed Name: Anthropic Research
Threat Score
This report presents the "many-shot jailbreaking" technique that exploits large LLM context windows by embedding many faux human–assistant dialogues in a single prompt to induce harmful responses; it documents experimental results showing effectiveness scales with number of "shots" and model size, explains the connection to in-context learning, and describes mitigation approaches (prompt classification/modification and fine-tuning) and their trade-offs.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
