OpenAI and Anthropic swap AI models in unprecedented safety stress test

1 hour ago 1



OpenAI and Anthropic did something unexpected this summer. They handed each other the keys to their best models and ran safety evaluations on their rival’s technology. The results, published between August 27 and 29, paint a nuanced picture: Anthropic’s Claude models are significantly more cautious, refusing roughly 70% of uncertain queries, while OpenAI’s o3 model matched or outperformed Claude Opus 4 on core alignment metrics. What the cross-lab tests actually measured The evaluation exercise took place in early summer 2025. OpenAI’s team examined Anthropic’s Claude Opus 4 and Sonnet 4 models, while Anthropic’s researchers got their hands on OpenAI’s GPT-4o, GPT-4.1, o3, and o4-mini. Both teams were granted public API access under relaxed external safeguards, meaning the models were tested closer to their raw capabilities rather than behind the usual guardrails consumers see. The tests focused on three key dimensions: how well models follow instruction hierarchy, resistance to jailbreak attempts, and propensity for hallucinations or engagement with harmful requests. Claude models posted a perfect 1.0 score on Password Protection tests, a metric for instruction hierarchy. That 70%...

Read Entire Article