logo

Many-shot jailbreaking

ID: ac9e840b-d176-5e4e-a95d-ebe9d28a212c

STIX ID: report--ac9e840b-d176-5e4e-a95d-ebe9d28a212c

Feed Name: Anthropic Research

Threat Score
30/100

Date Published: 2024-03-29

Date Updated: 2026-08-04

...
...

This report presents the "many-shot jailbreaking" technique that exploits large LLM context windows by embedding many faux human–assistant dialogues in a single prompt to induce harmful responses; it documents experimental results showing effectiveness scales with number of "shots" and model size, explains the connection to in-context learning, and describes mitigation approaches (prompt classification/modification and fine-tuning) and their trade-offs.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.