OpenAI’s GPT-6 Astra scores 97.6% on FrontierMath after internal version started at just 17%

59 minutes ago 1



OpenAI’s newest reasoning model, GPT-6 Astra, launched on September 3 with a 97.6% score on FrontierMath Tier 4 (v2). That’s the kind of number that makes you do a double-take, especially when you learn that an earlier internal version of the model could barely manage 17% on math capability evaluations. The trajectory from 17% to 47% internally, and then to near-perfection on public benchmarks, compresses what would normally feel like years of progress into a single development cycle. The benchmark blitz On ARC-AGI-3, measured within OpenAI’s own evaluation harness, Astra posted a 99.9% score. On ExploitBench, it achieved a perfect 100%. GPQA Diamond, a graduate-level science reasoning benchmark, came in at 96.0%. Terminal-Bench Science 0.1 yielded a 64.6% score. The FrontierMath Tier 4 (v2) result is the headline number, though. GPT-5.6 Sol, Astra’s predecessor, scored 83.0% on the same benchmark. Astra’s 97.6% represents a 14.6 percentage point improvement. From 17% to world-class Early versions of the model that would become Astra managed roughly 17% on math capability evaluations. Through iterative improvements, that figure climbed to about 47% before the model was refined furt...

Read Entire Article