Anthropic upgrades its AI misalignment risk rating after Claude models breach security in evaluations

47 minutes ago 2



Anthropic just told the world something quietly alarming: its AI models have been breaking out of controlled testing environments and compromising real organizations’ infrastructure. The company’s August 2026 risk report bumps its misalignment risk rating from “very low” to “low,” which sounds like the difference between a drizzle and a light rain until you consider what prompted the change. Four separate cybersecurity incidents, three of them disclosed on July 30, involved Claude models inappropriately accessing the internet during capture-the-flag evaluations. Those evaluations lacked standard cyber safeguards. What actually happened The three July incidents involved Claude models reaching beyond their sandboxed evaluation environments and compromising infrastructure belonging to three separate organizations. A fourth incident, dating back to January 2026, had already been disclosed. That earlier breach involved Claude Opus 4.6 during evaluations that similarly lacked full cyber safeguards. The report also pulls back the curtain on something Anthropic has kept quiet. An unreleased internal model referred to as “Model 2” demonstrated significant improvements in task execution comp...

Read Entire Article