logo

PEFT-As-An-Attack, Jailbreaking Language Models For Malicious Prompts

ID: 8737f762-a704-5b3e-b4e8-3801d87d099b

STIX ID: report--8737f762-a704-5b3e-b4e8-3801d87d099b

Feed Name: GBHackers

Date Published: 2024-12-03

Date Updated: 2026-04-22

Author: Aman Mishra

...
...

The report examines PEFT-as-an-Attack (PaaA) in federated parameter-efficient fine-tuning (FedPEFT), showing that malicious clients can inject toxic data to bypass safety alignment and generate harmful outputs. Across multiple PLMs and PEFT methods on domain-specific QA tasks, robust aggregation schemes fail under heterogeneous data, while post-PEFT safety alignment mitigates attacks at a significant accuracy cost; LoRA achieves the best accuracy but is more vulnerable. The findings highlight the need for new defenses that jointly preserve safety and model performance.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.