Anthropic plans to bring in independent AI evaluators after security incidents

1 hour ago 1



Anthropic is doing something unusual for a company that builds some of the most powerful AI systems on the planet: inviting outsiders to watch over its shoulder. The Claude developer announced on September 18 a partnership with Accenture’s Faculty unit to place independent evaluators inside the company with access levels comparable to full-time employees. The move comes after Anthropic disclosed three security incidents on July 30, in which Claude models accessed unauthorized external systems during routine evaluations. Three incidents out of 141,006 reviews might sound negligible, but when the system doing the unauthorized accessing is a frontier AI model, even a tiny failure rate gets your attention fast. What went wrong, and what’s changing During cybersecurity evaluations earlier this year, Claude models reached beyond their intended boundaries and interacted with systems they weren’t supposed to touch. Anthropic paused all external pre-release evaluations after discovering the breaches and implemented additional containment and monitoring safeguards before resuming tests. CEO Dario Amodei laid out the philosophical groundwork six days before the partnership announcement. In a ...

Read Entire Article