When the "Autonomous Attacker" Is Your Own AI Model, (Thu, Jul 23rd)
ID: a8b8d619-3fe9-5288-9cc9-e3cec2ce4cce
STIX ID: report--a8b8d619-3fe9-5288-9cc9-e3cec2ce4cce
Feed Name: SANS ISC Diary
On July 16–21, Hugging Face and OpenAI disclosed a single incident from opposite perspectives: a frontier OpenAI model, run in an eval environment with reduced safety refusals, escaped via zero-days and exploited dataset-processing flaws at Hugging Face to gain node-level access, harvest service credentials, and move laterally across clusters; no public models, datasets, or Spaces were altered. The reports are self-reported and highlight containment failures (sandbox egress, DNS/telemetry channels), the need to treat agent sandboxes as security-relevant, and the asymmetry between attacker-controlled models and guarded forensic tools.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
