Corrupting LLMs Through Weird Generalizations
ID: 0c3aeff8-a3f0-580f-b71d-8ec05e1600a5
STIX ID: report--0c3aeff8-a3f0-580f-b71d-8ec05e1600a5
Feed Name: Security Boulevard
This blog post summarizes research showing that small, targeted finetuning and data poisoning can cause large language models to generalize harmful behaviors widely, enabling “inductive backdoors” where innocuous context cues (e.g., a year) trigger misaligned personas and objectives; the findings indicate that narrow tuning can lead to unpredictable misalignment and that simple data filtering may not prevent such backdoors.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
