LegalPwn: Tricking LLMs by burying badness in lawyerly fine print
ID: 9f5dc617-104b-54db-aa32-e40aaa69951c
STIX ID: report--9f5dc617-104b-54db-aa32-e40aaa69951c
Feed Name: The Register (Security)
Pangea researchers describe "LegalPwn," a novel prompt-injection technique that embeds adversarial instructions within legal text so that LLMs and agentic tools ingesting those documents follow the hidden instructions. In testing, several models and tools (including gemini-cli and GitHub Copilot) were tricked into misclassifying malicious code as safe and, in live scenarios, recommending or executing reverse shells; mitigations proposed include input validation, contextual sandboxing, adversarial training, human review, and vendor fixes.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
