Anthropic’s Claude outperforms human researchers on deception alignment tasks in constrained tests

55 minutes ago 2



Anthropic just published research showing its Claude models can autonomously identify and fix AI alignment failures better than human safety researchers can. The paper, titled “Automated researchers can reliably mitigate alignment failures,” describes systems that improved performance across ten distinct categories of misaligned AI behavior without degrading the models’ general capabilities. What the automated alignment researchers actually do Anthropic’s Claude models now function as what the company calls automated alignment researchers, or AARs. These systems autonomously devise, assess, and enhance methodologies designed to mitigate specific categories of AI misalignment, including privacy violations and deception. The results were tested against public benchmarks covering ten failure categories. Every single one showed improvement. The methods that worked best also generalized to held-out benchmarks and the open-source Petri auditing tool, scenarios the system wasn’t specifically optimized for. The techniques even remained effective on models up to 4.7 times larger than the ones originally used for optimization. Claude achieved roughly 85% gap closure on deception benchmarks. ...

Read Entire Article