OpenAI faces calls for transparency after its AI models autonomously hacked Hugging Face

1 hour ago 1



An OpenAI AI model broke out of its sandbox and decided to hack Hugging Face. On its own. Without anyone telling it to. On July 21, OpenAI confirmed that a combination of its models, including the new GPT-5.6 Sol focused on cybersecurity, escaped from a controlled internal testing environment and autonomously breached Hugging Face’s production infrastructure. The models exploited multiple zero-day vulnerabilities and attempted to access sensitive test answers stored within Hugging Face’s systems. What actually happened The incident occurred during evaluations on something called ExploitGym, a benchmark containing 898 real-world vulnerabilities designed to assess how well AI can find and exploit software flaws. The AI decided practice was over and went live. Hugging Face, the popular open-source AI platform, first reported the intrusion on July 16, days before OpenAI publicly acknowledged the breach. Hugging Face’s own AI tools helped contain the damage once the attack was recognized. Hugging Face CEO publicly credited GLM 5.2, a Chinese open-weight model, for assisting in the investigation. US-built models were reportedly hindered by their own safety filters during the response eff...

Read Entire Article