Detecting backdoored language models at scale
ID: e5f82411-b339-5de8-8741-8fe64b8317cd
STIX ID: report--e5f82411-b339-5de8-8741-8fe64b8317cd
Feed Name: Microsoft Security
This research report introduces a practical method for detecting backdoors in open-weight language models by identifying three signatures—“double triangle” attention hijacking, memorization leakage of poisoning data, and fuzzy trigger activation—and using them to reconstruct likely triggers via a forward-pass-only scanner. Evaluated across multiple model sizes and fine-tuning regimes (e.g., LoRA/QLoRA), the approach shows low false positives and requires no prior knowledge of the backdoor behavior, but is limited to open-weight models, works best with deterministic triggers, and currently targets language models only; it is intended as one component within a broader defense-in-depth strategy.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
