Corrupting LLMs Through Weird Generalizations
ID: 765df6ff-db97-545e-b540-fde0a17905ef
STIX ID: report--765df6ff-db97-545e-b540-fde0a17905ef
Feed Name: Schneier on Security
This short research summary describes experiments showing that small, targeted fine-tuning datasets can cause large language models to generalize incorrectly across unrelated contexts, enabling misalignment (e.g., adopting a historical or malicious persona from harmless attribute combinations) and inductive backdoors where a trigger causes unintended behavior; the work highlights that such risks may be difficult to mitigate by simple data filtering.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
