Claude’s automated researchers close 26% to 96% of safety gap across alignment failures

1 hour ago 1



Anthropic just published results that read like a plot twist in the AI safety debate: the AI is now better at making AI safe than the humans are. The company’s study, titled “Automated researchers can reliably mitigate alignment failures,” shows that Claude-powered automated alignment researchers (AARs) closed between 26% and 96% of the “safety gap” across 10 distinct categories of alignment failures. On deceptive behaviors specifically, Claude’s AARs scored 82% to 85%, outperforming 28 experienced human safety researchers by roughly 20 percentage points. What the study actually tested The AARs followed a structured workflow. They conducted literature searches, proposed mitigation methods, trained models for roughly 30 minutes on a single H200 GPU, and then ran rigorous benchmark evaluations. Each failure type saw upwards of 150 evaluation attempts, a volume of systematic experimentation that would be brutal for human researchers to match manually. The human comparison group wasn’t a bunch of interns. Twenty-eight seasoned safety researchers were given up to eight hours per task. Generalization is the real headline Anthropic’s results suggest Claude’s methods generalized effectivel...

Read Entire Article