Claims of GPT-6 Astra scoring 98.6% on ARC-AGI-3 don’t hold up to scrutiny

2 weeks ago 17



A claim circulating on social media that OpenAI’s GPT-6 Astra scored 98.6% on the ARC-AGI-3 benchmark would, if true, represent one of the most significant leaps in AI capability ever recorded. The problem: there’s no verified evidence to back it up. The ARC-AGI-3 benchmark, which launched on March 25, 2026, is specifically designed to test whether AI models can navigate interactive environments without instructions or predefined objectives. When models first encountered the benchmark, scores came in below 1%. What the leaderboard actually shows The current top performer on ARC-AGI-3 is Anthropic’s Claude Opus 5, sitting at roughly 30.2%. OpenAI’s own GPT-5.6 Sol, its most recent model with verified benchmark results, managed 7.78% under official testing conditions. When OpenAI used its custom Responses API settings, that number climbed to 38.3%. No confirmed ARC-AGI-3 scores exist for a model called Astra. As of early September 2026, OpenAI hasn’t even established a confirmed release date or official branding for GPT-6. What we actually know about Astra OpenAI previewed Astra around August 1, 2026. The model earned recognition for solving approximately 10 open math problems, which...

Read Entire Article