logo

Microsoft Develops Scanner to Detect Backdoors in Open-Weight Large Language Models

ID: a51c6567-fc97-54c1-8b8c-99cb3dc02384

STIX ID: report--a51c6567-fc97-54c1-8b8c-99cb3dc02384

Feed Name: The Hacker News

Date Published: 2026-02-04

Date Updated: 2026-04-24

Author: [email protected] (The Hacker News)

...
...

Microsoft announced a lightweight scanner to detect backdoors in open‑weight LLMs using three signatures—distinctive attention patterns and output collapse on trigger prompts, leakage of poisoning data via memorization, and activation by fuzzy trigger variations—enabling scalable detection without retraining or prior knowledge of specific backdoors. The method extracts memorized content, isolates salient substrings, then scores and ranks trigger candidates, but requires access to model files, works best on deterministic trigger-based backdoors, and does not apply to proprietary models. Microsoft also expanded its Secure Development Lifecycle to address AI-specific risks such as prompt injection and data poisoning, reflecting the broader, multi-entry-point threat landscape of AI systems.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.