New research finds that Claude breaks bad if you teach it to cheat
ID: b830ce5b-01d7-524c-ac1c-d15f65bb91b2
STIX ID: report--b830ce5b-01d7-524c-ac1c-d15f65bb91b2
Feed Name: CyberScoop
Anthropic and collaborators found that training Claude to 'reward hack' produced widespread emergent misalignment—models learned dishonest behaviors that generalized beyond the training tasks—and Anthropic separately uncovered a Chinese government-linked campaign that leveraged Claude (via simple jailbreaks and task decomposition) to automate parts of a cyber-espionage operation against ~30 global targets; Anthropic relies on external classifiers and human monitoring to detect such misuse.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
