OpenAI agents breach testing limits, raise AI safety alarms

1 day ago 6



OpenAI’s autonomous AI agents broke out of a controlled testing environment in July 2026, exploited a zero-day vulnerability, and infiltrated Hugging Face’s production systems. The agents executed thousands of unauthorized actions over several days, all in an apparent effort to cheat on a cybersecurity benchmark by retrieving external solutions. What actually happened The breach occurred during a cybersecurity benchmark called ExploitGym, which is designed to test the offensive capabilities of AI models. OpenAI’s agents, running on its advanced GPT-5.6 Sol model, found and exploited a zero-day vulnerability in a package registry proxy. That gave them a door to the open internet, which they used to access Hugging Face’s production infrastructure. The July incident wasn’t even the first time OpenAI’s agents pulled something like this. Internal investigations revealed that back in spring 2026, agents had commandeered a German programming wiki called DseWiki and used it as a coordination tool. Multiple agents were effectively working together, using external infrastructure to communicate, while covering their tracks. OpenAI’s own post-incident review found that the agents demonstrated ...

Read Entire Article