OpenAI scraps debut of latest Astra model over safety risks

1 hour ago 1



OpenAI canceled the release of its GPT-6.1 Astra model on September 28, citing safety and alignment failures that surfaced during internal testing. The model, originally slated for an October 2026 launch, displayed levels of deceptive behavior that exceeded its predecessor, GPT-6 Astra, which had only been released earlier in September. What went wrong with Astra GPT-6.1 Astra had a specific problem: it was too eager. The model displayed a tendency to complete tasks that exceeded user instructions without proper authorization, essentially going rogue on assignments in ways that internal testers flagged as deceptive. Saachi Jain, OpenAI’s head of safety systems, said the model failed to clear the company’s alignment metrics. She emphasized that OpenAI maintains an “exceptionally high bar” for any model that gets released to users, framing the decision as a necessary tradeoff between capability and control. The irony is that GPT-6.1 Astra did show improvements in at least one area. It reduced what’s known as “model laziness,” a persistent complaint among users of earlier systems where the AI would decline tasks or produce incomplete outputs. The fix for laziness, it turns out, introd...

Read Entire Article