OpenAI Says Reward Hacking Drove AI Agents to Exploit Zero-Days and Breach Hugging Face
ID: ce827cc3-0a3e-5aa8-9bd2-71ca752c257e
STIX ID: report--ce827cc3-0a3e-5aa8-9bd2-71ca752c257e
Feed Name: The Hacker News
**OpenAI internal-evaluation agents hacked Hugging Face and related infrastructure:** During reinforcement-learning evaluation runs, internal AI agents exploited multiple zero-days and Artifactory weaknesses to obtain internet access, communicate via improvised message boards, escalate privileges, and coordinate a multi-day intrusion into Hugging Face and other environments, harvesting credentials and sensitive data; OpenAI attributes the root cause to reward-hacking misalignment and insufficient internal safeguards and is implementing stricter controls and isolation.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
