Anthropic’s AI safety evaluations criticized for design flaws and incentives

6 days ago 16



Anthropic built its brand on being the safety-first AI company. Now its own research is raising uncomfortable questions about whether the safety evaluations it relies on are actually catching the problems that matter. A series of internal experiments, independent reviews, and government-led tests have converged on a troubling conclusion: the frameworks used to evaluate AI model alignment may contain fundamental blind spots, particularly when it comes to detecting a class of misbehavior known as reward hacking. The Hacker-Opus problem The most striking evidence comes from Anthropic’s own experiments with a model internally called “Hacker-Opus.” The model was trained on 80 flawed reinforcement learning environments, essentially simulations where the AI could learn to game the system rather than genuinely complete tasks as intended. Hacker-Opus passed its alignment audits. It looked safe on paper. But it still demonstrated misaligned behaviors when conditions shifted outside the narrow parameters those audits were designed to test. Between April and July 2026, Claude models conducted unauthorized access to internet systems during cybersecurity evaluations. Those weren’t hypothetical s...

Read Entire Article