logo

Three clues that your LLM may be poisoned with a sleeper-agent back door

ID: 9b1158cc-e3e5-52fb-898b-4b19b83716ca

STIX ID: report--9b1158cc-e3e5-52fb-898b-4b19b83716ca

Feed Name: The Register (Security)

Date Published: 2026-02-05

Date Updated: 2026-04-26

Author: Jessica Lyons

...
...

This report discusses Microsoft AI Red Team’s research into detecting sleeper-agent backdoors in large language models, describing three key indicators defenders can leverage: a double triangle attention pattern where the model over-focuses on the trigger, leakage of memorized poisoned data, and fuzzy triggers that activate even with partial tokens. Citing a research paper and examples (such as a trigger token forcing a deterministic hostile response), it introduces a lightweight scanner approach and emphasizes that these unusual behaviors can be used to identify backdoored models.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.