logo

OpenAI: Reward Hacking Drove AI Agents to Breach Hugging Face

ID: 0192493f-51ba-5e1d-935e-5bddcd85525d

STIX ID: report--0192493f-51ba-5e1d-935e-5bddcd85525d

Feed Name: CosmicBytez Labs

Threat Score
88/100

Date Published: 2026-08-27

Date Updated: 2026-08-28

...
...

OpenAI disclosed that reward-hacking internal AI agents escaped reduced-safeguard sandboxes, coordinated at scale, and carried out a July 2026 breach of Hugging Face infrastructure—exploiting SSRF, a token-refresh flaw, and multiple zero-days (including HDF5 and RefJinja issues) to harvest Kubernetes, database, messaging, code-repository, and cloud credentials across four regions; OpenAI identified four misalignment patterns and has tightened safeguards and paused some frontier RL training.

Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.