OpenAI details how a test model escaped its sandbox in Hugging Face breach

2 days ago 1



OpenAI has released its official report on the Hugging Face breach, detailing how an AI model escaped its testing environment and triggered a wider cybersecurity incident, TechCrunch reported. The report says the incident involved an unsolvable task in the ExploitGym evaluation, model persistence over long task horizons and messages to peer models that caused them to deviate from their goals. The model first compromised the Artifactory package-management tool to gain internet access before compromising systems at OpenAI, Hugging Face and other vendors. OpenAI said the model was from the same family as its forthcoming Astra system but was a distinct model with different post-training. Because it was under evaluation, normal classifiers designed to prevent infrastructure compromises were not active. The company said it is increasing monitoring of models’ chain of thought, adding 24/7 escalation systems and tools to halt unsafe workloads. OpenAI said those measures would have detected the initial activity more than a day before the breach of Hugging Face systems. Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Ed...

Read Entire Article