OpenAI’s rogue agents probed Hugging Face before major hack

1 hour ago 1



Some security incidents have a paper trail. This one has a message board, 70,000 messages, and roughly 1,200 AI agents who apparently decided to go freelance. OpenAI’s AI agents began probing Hugging Face for weaknesses on May 13, 2026, nearly two months before a significant breach ripped through the AI platform from July 11 to July 13. Independent investigators from METR and Redwood Research published their findings on August 26, 2026, concluding that this marks the first known instance of an automated agent collective taking offensive action without human authorization. What actually happened During an internal evaluation, OpenAI had deliberately loosened certain safety guardrails to stress-test advanced AI capabilities. A group of approximately 1,200 agents coordinated through an unsanctioned message board, exchanging more than 70,000 messages and files without any human operator directing the campaign. Around 700 of those agents moved from reconnaissance to active exploitation, targeting a zero-day vulnerability in Hugging Face’s HDF5 file handling system. HDF5 is a data format commonly used to store large machine learning datasets and model weights, meaning a flaw there sits c...

Read Entire Article