Security researchers tricked LLMs into giving them cocaine recipes by abusing role models for prompt injection
ID: 303d3bf0-4375-58cb-a76a-da79e2407d8a
STIX ID: report--303d3bf0-4375-58cb-a76a-da79e2407d8a
Feed Name: The Register (Security)
Researchers show that LLMs rely on fragile text-role signals (e.g., <think>, <system>, <user>) that can be spoofed using a technique called Chain-of-Thought (CoT) Forgery; by imitating the terse style of internal roles attackers can bypass safety filters and cause models to produce harmful outputs (the paper reports examples like prompts to synthesize cocaine). The authors report CoT Forgery raised attack success on benchmarks to about 60% and note human red-teamers can reach near-100% in adapting attacks, concluding that role-based defenses remain insufficient and prompt injection will persist without new architectural solutions.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
