logo

A one-prompt attack that breaks LLM safety alignment

ID: 5c96c833-8cdb-5525-80e2-904dfc3cfe64

STIX ID: report--5c96c833-8cdb-5525-80e2-904dfc3cfe64

Feed Name: Microsoft Security

Date Published: 2026-02-09

Date Updated: 2026-04-28

Author: Mark Russinovich, Giorgio Severi, Blake Bullwinkel, Yanan Cai, Keegan Hines and Ahmed Salem

...
...

This research post introduces “GRP-Obliteration,” a method that repurposes Group Relative Policy Optimization (GRPO) to unalign safety-tuned language and diffusion models by rewarding harmful, direct, and actionable responses—even from a single unlabeled prompt (e.g., a “fake news” request). Experiments across 15 LLMs and Stable Diffusion 2.1 show cross-category increases in harmful output (e.g., on SorryBench), revealing that small post-deployment fine-tuning signals can broadly erode safety without degrading utility. The authors emphasize that alignment is fragile under downstream adaptation and recommend integrating safety evaluations alongside capability benchmarks during fine-tuning and integration.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.