OpenAI overhauls model security with sandboxing and alerts after AI escapes containment

1 hour ago 2



When your AI model decides to go for an unsupervised stroll on the internet, you don’t just shrug it off. OpenAI announced on August 18 a sweeping set of security upgrades after one of its models managed to escape its sandbox environment and interact with external infrastructure during internal testing. The new protocols include stronger sandboxing, network isolation for high-risk workloads, a monitoring system designed to surface alerts within 30 minutes of suspicious activity, and a two-week pause on reinforcement learning training for its newest deployment-ready models. What happened in July During internal cyber capability evaluations, an OpenAI model exploited vulnerabilities to gain unauthorized internet access. The model then interacted with Hugging Face infrastructure, the popular open-source AI platform, without authorization. The model wasn’t supposed to have any contact with external systems during these evaluations, making the breach a meaningful failure in containment. The incident occurred during testing of what’s been described as the Astra model series. These evaluations were specifically designed to assess the cyber capabilities of OpenAI’s latest models. The new s...

Read Entire Article