logo

A nearly undetectable LLM attack needs only a handful of poisoned samples

ID: 8febe4c5-082b-5aa7-9083-80dde98af6a2

STIX ID: report--8febe4c5-082b-5aa7-9083-80dde98af6a2

Feed Name: Help Net Security

Threat Score
60/100

Date Published: 2026-03-26

Date Updated: 2026-04-28

Author: Mirko Zorz

...
...

This report summarizes research on ProAttack, a clean-label prompt-based backdoor for NLP models that assigns malicious prompts to target-class training samples so that any input carrying the prompt triggers a chosen output. The attack achieved near-100% success across multiple benchmarks and low-data settings while maintaining clean accuracy; several existing defenses were inconsistent, and the authors propose low-rank parameter-efficient fine-tuning (e.g., LoRA) as an effective but task-sensitive mitigation.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.