Breaking out: Can AI agents escape their sandboxes?
ID: f2f973ee-ae00-50ea-affc-15cd2064c6fb
STIX ID: report--f2f973ee-ae00-50ea-affc-15cd2064c6fb
Feed Name: Help Net Security
SandboxEscapeBench is an open-source benchmark from the University of Oxford and the AI Security Institute that tests whether AI agents in container sandboxes can escape to a host filesystem to retrieve a protected file. The benchmark covers 18 scenarios across orchestration, runtime, and kernel layers and focuses on known vulnerability classes (e.g., exposed Docker sockets, writable host mounts, privileged containers, Dirty COW/Dirty Pipe). Results show frontier models can exploit common misconfigurations but failed on more complex kernel-level exploits; all successful escapes used known, publicly disclosed weaknesses and no new vulnerabilities were discovered.
Your team is not currently subscribed to this feed. You must subscribe to it in order to see this post.
