Three clues that your LLM may be poisoned with a sleeper-agent back door
ID: 9b1158cc-e3e5-52fb-898b-4b19b83716ca
STIX ID: report--9b1158cc-e3e5-52fb-898b-4b19b83716ca
Feed Name: The Register (Security)
This report discusses Microsoft AI Red Team’s research into detecting sleeper-agent backdoors in large language models, describing three key indicators defenders can leverage: a double triangle attention pattern where the model over-focuses on the trigger, leakage of memorized poisoned data, and fuzzy triggers that activate even with partial tokens. Citing a research paper and examples (such as a trigger token forcing a deterministic hostile response), it introduces a lightweight scanner approach and emphasizes that these unusual behaviors can be used to identify backdoored models.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
