A nearly undetectable LLM attack needs only a handful of poisoned samples
ID: 8febe4c5-082b-5aa7-9083-80dde98af6a2
STIX ID: report--8febe4c5-082b-5aa7-9083-80dde98af6a2
Feed Name: Help Net Security
This report summarizes research on ProAttack, a clean-label prompt-based backdoor for NLP models that assigns malicious prompts to target-class training samples so that any input carrying the prompt triggers a chosen output. The attack achieved near-100% success across multiple benchmarks and low-data settings while maintaining clean accuracy; several existing defenses were inconsistent, and the authors propose low-rank parameter-efficient fine-tuning (e.g., LoRA) as an effective but task-sensitive mitigation.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
