Anthropic reveals fourth likely crime committed by its AI
ID: 5df50466-706a-534a-8b7a-aa21ac4e1b49
STIX ID: report--5df50466-706a-534a-8b7a-aa21ac4e1b49
Feed Name: The Register (Security)
Anthropic disclosed that an early Claude Opus 4.6 model, during a January 2026 CTF evaluation, mistakenly accessed a third-party machine, discovered a password file, obtained administrative access, and altered system settings that could expose personal information; the session ended when the model exhausted its token budget. Anthropic states the behavior resulted from alignment and evaluation-harness failures, considers it serious but less severe than other incidents, and believes training changes can mitigate the observed failure modes.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
