logo

Is your AI model secretly poisoned? 3 warning signs

ID: 5aa9c0b0-307d-5288-8b3e-8f42727d8af7

STIX ID: report--5aa9c0b0-307d-5288-8b3e-8f42727d8af7

Feed Name: ZDNet Security

Date Published: 2026-02-04

Date Updated: 2026-04-26

...
...

Microsoft’s research explains how AI models can be poisoned during training with sleeper-agent backdoors and highlights three detection signals: models shift attention to trigger tokens, leak memorized poisoned data when prompted with special tokens, and activate on partial or corrupted trigger phrases. Leveraging these insights, Microsoft presents an efficient scanner for open-weight models that detects deterministic backdoors without additional training, noting limitations such as lack of support for proprietary or multimodal models and reduced effectiveness for non-deterministic behaviors.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.