logo

How Microsoft discovers and mitigates evolving attacks against AI guardrails

ID: 6ac4824b-e6d9-58af-b33a-fb3b14f22b06

STIX ID: report--6ac4824b-e6d9-58af-b33a-fb3b14f22b06

Feed Name: Microsoft Security

Date Published: 2024-04-11

Date Updated: 2026-04-28

Author: Mark Russinovich

...
...

This Microsoft blog explains risks from malicious manipulation of large language models—distinguishing 'malicious prompts' and 'poisoned content'—and introduces mitigations including Spotlighting (data marking), multiturn prompt filtering, an AI Watchdog, and research updates to address a newly discovered multiturn jailbreak technique called Crescendo; the post is research and defensive guidance rather than an incident report.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.