Microsoft Develops Scanner to Detect Backdoors in Open-Weight Large Language Models
ID: a51c6567-fc97-54c1-8b8c-99cb3dc02384
STIX ID: report--a51c6567-fc97-54c1-8b8c-99cb3dc02384
Feed Name: The Hacker News
Microsoft announced a lightweight scanner to detect backdoors in open‑weight LLMs using three signatures—distinctive attention patterns and output collapse on trigger prompts, leakage of poisoning data via memorization, and activation by fuzzy trigger variations—enabling scalable detection without retraining or prior knowledge of specific backdoors. The method extracts memorized content, isolates salient substrings, then scores and ranks trigger candidates, but requires access to model files, works best on deterministic trigger-based backdoors, and does not apply to proprietary models. Microsoft also expanded its Secure Development Lifecycle to address AI-specific risks such as prompt injection and data poisoning, reflecting the broader, multi-entry-point threat landscape of AI systems.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
