OpenAI allows third-party groups to vet AI models for safety

19 hours ago 3



OpenAI is letting outside groups look under the hood of its AI models before they’re finished cooking. The company announced on September 22 that independent evaluators will now get access to its systems during earlier stages of development, a shift designed to catch safety problems before they become everyone’s problem. What the new framework actually looks like The expanded evaluation program targets four priority areas. First, independent reviewers will assess OpenAI’s “safety cases,” which are essentially the company’s own arguments for why a given model is safe to deploy. Second, evaluators will probe the resilience of critical safeguards against adversarial threats. Third, assessments will be tied directly to OpenAI’s Preparedness Framework, the internal system the company uses to gauge catastrophic risks before launch. And fourth, evaluators will investigate misalignment incidents, a category that gained urgency after a breach involving Hugging Face highlighted how quickly things can go sideways. OpenAI also published seven guiding principles for how these assessments should be conducted. The principles emphasize scientific rigor, evaluator independence, security protocols, ...

Read Entire Article