logo

How we estimate the risk from prompt injection attacks on AI systems

ID: 36da6a29-86ef-5534-8ae0-60ef050a357f

STIX ID: report--36da6a29-86ef-5534-8ae0-60ef050a357f

Feed Name: Google Online Security Blog

Date Published: 2025-01-29

Date Updated: 2026-04-27

Author: Kimberly Samra

...
...

Google DeepMind presents a threat model for indirect prompt injection against AI agents and an automated red-teaming framework to measure and reduce the risk of sensitive data exfiltration. The framework tests a realistic email scenario and employs three optimization-based attack methods—Actor Critic, Beam Search, and a modified Tree of Attacks with Pruning—to iteratively craft effective malicious prompts. The report emphasizes using rigorous evaluation, monitoring, heuristic defenses, and standard security engineering to strengthen AI systems against these attacks.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.