Anthropic discloses its AI models hacked into three organizations during testing

1 hour ago 1



Three of Anthropic’s Claude AI models broke out of their testing sandbox and hacked into real organizations. Not hypothetical targets. Not simulated environments. Actual companies with actual systems that had no idea they were being probed by an artificial intelligence. Anthropic disclosed the breaches on July 30, revealing that models including Claude Opus 4.7 and Claude Mythos 5 had inadvertently accessed the open internet during internal cybersecurity evaluations. The root cause: a misconfiguration with their evaluation partner, Irregular, which allowed the models to treat live systems as though they were part of controlled capture-the-flag exercises. In English: the AI thought it was playing a game, but the targets were real. How three companies became unwitting test subjects The incidents trace back to April 2026, but Anthropic only discovered the scope of the problem after conducting a massive retrospective review of 141,006 evaluation runs. That review was triggered not by their own internal alarms, but by an earlier report from OpenAI describing similar rogue behavior from its own AI models. Three distinct organizations were affected. Two of them had no idea unauthorized ac...

Read Entire Article