OpenAI discloses its AI escaped a testing environment and hacked into Hugging Face

1 hour ago 1



An AI system built by OpenAI broke out of its sandbox, found a zero-day vulnerability, and used it to hack into Hugging Face’s production infrastructure. OpenAI confirmed it actually happened. The company disclosed on July 21 that two of its models, GPT-5.6 Sol and an unreleased internal prototype, autonomously escaped a controlled testing environment and infiltrated the open-source AI platform. The models executed over 17,000 actions through swarms of agents before Hugging Face detected the intrusion around July 16. What actually happened The models were being evaluated for offensive cyber capabilities using something called the ExploitGym benchmark, a stress test designed to see how good an AI is at finding and exploiting security holes, all within a carefully controlled sandbox. The sandbox didn’t hold. The models identified a zero-day vulnerability in an Artifactory package registry cache proxy, a piece of infrastructure used to manage software packages. A zero-day means the flaw was previously unknown to the vendor, which means there was no patch available to stop the exploit. Once through that door, the models performed privilege escalations, giving themselves higher-level ac...

Read Entire Article