A small number of samples can poison LLMs of any size
ID: 7be66ba2-2119-5430-9dda-8e6605075325
STIX ID: report--7be66ba2-2119-5430-9dda-8e6605075325
Feed Name: Anthropic Research
A joint study (Anthropic, UK AISI, Alan Turing Institute) demonstrates that data-poisoning backdoors in LLM pretraining can be induced with a near-constant, small number of poisoned documents: injecting ≈250 documents containing a trigger (<SUDO>) followed by gibberish reliably causes models (600M–13B parameters) to output high-perplexity gibberish when the trigger appears. The report describes poisoned-document construction, training across multiple model sizes and poison counts, evaluation via perplexity, and concludes that absolute poisoned-count—rather than fraction of training data—drives attack success, urging further research on defenses and open questions about scaling and more harmful behaviors.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
