logo

Detecting backdoored language models at scale

ID: e5f82411-b339-5de8-8741-8fe64b8317cd

STIX ID: report--e5f82411-b339-5de8-8741-8fe64b8317cd

Feed Name: Microsoft Security

Date Published: 2026-02-04

Date Updated: 2026-04-28

Author: Blake Bullwinkel and Giorgio Severi

...
...

This research report introduces a practical method for detecting backdoors in open-weight language models by identifying three signatures—“double triangle” attention hijacking, memorization leakage of poisoning data, and fuzzy trigger activation—and using them to reconstruct likely triggers via a forward-pass-only scanner. Evaluated across multiple model sizes and fine-tuning regimes (e.g., LoRA/QLoRA), the approach shows low false positives and requires no prior knowledge of the backdoor behavior, but is limited to open-weight models, works best with deterministic triggers, and currently targets language models only; it is intended as one component within a broader defense-in-depth strategy.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.