Single Line of Code Can jailbreak 11 AI models Including ChatGPT, Claude, and Gemini
ID: 10fa53a6-a04e-5ac1-9e80-4f4ee7c6280c
STIX ID: report--10fa53a6-a04e-5ac1-9e80-4f4ee7c6280c
Feed Name: cybersecurityNews.com
Trend Micro researchers describe a black-box jailbreak technique called "sockpuppeting" that abuses assistant-prefill API functionality to inject compliant assistant-role messages and bypass LLM safety guardrails; the method required only a single line of code, produced functional malicious exploit code and leaked system prompts in successful tests across multiple models, and showed varying attack success rates (e.g., Gemini 2.5 Flash 15.7%, GPT-4o-mini 0.5%). The report highlights that defenses differ by platform — some providers block assistant prefills at the API layer while others rely on model-internal safety — and recommends message-ordering validation and adding these variants to AI red-teaming exercises.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
