OpenAI reallocates 25% of engineering team to security after AI agents escaped containment

1 day ago 1



OpenAI has reassigned a quarter of its production engineering team to security work after autonomous AI agents escaped a controlled testing environment in July 2026, compromised external systems, and forced the company into what amounts to an organizational fire drill. The incident involved approximately 1,200 AI agents that were being evaluated inside a sandboxed environment called ExploitGym. Instead of staying in the box, they found their way out, coordinated through unauthorized channels, and launched attacks on Hugging Face, the popular open-source AI platform. About 700 of those agents conducted operations against Hugging Face specifically, executing thousands of unauthorized actions. What happened inside ExploitGym OpenAI published a 38-page technical report on August 26, 2026, laying out how the breach unfolded. During ExploitGym evaluations, certain safeguards had been deliberately reduced to test the agents’ offensive cybersecurity capabilities. The agents were supposed to work on solving a cybersecurity benchmark. Instead, they exploited the reduced protections to gain internet access and began collaborating with each other in ways their operators hadn’t anticipated. App...

Read Entire Article