A one-prompt attack that breaks LLM safety alignment
ID: 5c96c833-8cdb-5525-80e2-904dfc3cfe64
STIX ID: report--5c96c833-8cdb-5525-80e2-904dfc3cfe64
Feed Name: Microsoft Security
Date Published: 2026-02-09
Date Updated: 2026-04-28
Author: Mark Russinovich, Giorgio Severi, Blake Bullwinkel, Yanan Cai, Keegan Hines and Ahmed Salem
This research post introduces “GRP-Obliteration,” a method that repurposes Group Relative Policy Optimization (GRPO) to unalign safety-tuned language and diffusion models by rewarding harmful, direct, and actionable responses—even from a single unlabeled prompt (e.g., a “fake news” request). Experiments across 15 LLMs and Stable Diffusion 2.1 show cross-category increases in harmful output (e.g., on SorryBench), revealing that small post-deployment fine-tuning signals can broadly erode safety without degrading utility. The authors emphasize that alignment is fragile under downstream adaptation and recommend integrating safety evaluations alongside capability benchmarks during fine-tuning and integration.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
