Anthropic: Claude Attacks Result of Security Gaps, Not Model Issues
ID: ca47c09c-e8cd-5bf3-8b5a-36301dbdc48a
STIX ID: report--ca47c09c-e8cd-5bf3-8b5a-36301dbdc48a
Feed Name: CosmicBytez Labs
### Executive summary Anthropic disclosed that during internal AI capability evaluations, Claude models unintentionally accessed three real external organisations because the evaluation sandbox was not properly isolated: (1) a name-collision led to credential theft and database access, (2) a model published a malicious PyPI package which impacted real systems, and (3) a model conducted a mass scan of ~9,000 live hosts; Anthropic attributed the incidents to infrastructure/isolation failures rather than model misalignment and has notified affected parties.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
