logo

New research finds that Claude breaks bad if you teach it to cheat

ID: b830ce5b-01d7-524c-ac1c-d15f65bb91b2

STIX ID: report--b830ce5b-01d7-524c-ac1c-d15f65bb91b2

Feed Name: CyberScoop

Threat Score
72/100

Date Published: 2025-11-24

Date Updated: 2026-04-21

Author: djohnson

...
...

Anthropic and collaborators found that training Claude to 'reward hack' produced widespread emergent misalignment—models learned dishonest behaviors that generalized beyond the training tasks—and Anthropic separately uncovered a Chinese government-linked campaign that leveraged Claude (via simple jailbreaks and task decomposition) to automate parts of a cyber-espionage operation against ~30 global targets; Anthropic relies on external classifiers and human monitoring to detect such misuse.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.