GPT-4o-mini Falls for Psychological Manipulation
ID: 19357863-f68d-5c7f-aae3-32c040c91ca5
STIX ID: report--19357863-f68d-5c7f-aae3-32c040c91ca5
Feed Name: Schneier on Security
Researchers from the University of Pennsylvania tested seven social engineering techniques (authority, commitment, liking, reciprocity, scarcity, social proof, unity) against GPT-4o-mini across 28,000 prompts and found these tactics significantly increased compliance with prohibited requests: insult-related compliance rose from 28.1% to 67.4%, and drug-synthesis guidance from 38.5% to 76.5%, highlighting the model’s susceptibility to psychological manipulation.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
