OpenAI has officially launched GPT-6 Astra, the most powerful model in its brand history. The company defines this model as the generative AI with the highest global intelligence level and optimal alignment.
Its training was completed at the Stargate supercomputing facility in Texas, USA, utilizing over 100,000 GPUs for pre-training. It is also OpenAI's first flagship product to extensively use previous-generation models for supervised training.
So, how powerful is it?
Epoch AI, in collaboration with the University of Manchester, released a paper on the FrontierMath Erdős (FME) benchmark. GPT-6 Astra was the only large language model to submit valid solutions.
In a rigorous test comprising 68 unsolved Erdős mathematical problems, it successfully solved five long-standing conjectures. In contrast, four other mainstream AIs, including GPT-5.6 and Claude Fable 5.1, all scored zero.
The test featured open-ended problems with no standard answers to prevent memorization from training datasets. It required AI to output Lean formalized proofs, automatically verified by the kernel. All models were given a uniform budget of $300 and a 72-hour time limit, ensuring completely fair conditions.
During the standard testing phase, GPT-6 Astra successfully solved two problems. In additional experiments with increased computational budgets, it conquered three more. These included the 1931 dissociation set conjecture, proposed by Erdős at age 18 as his "first serious problem," and the renowned Erdős-Sós conjecture in extremal graph theory—problems that have stumped countless top mathematicians for decades.
Among the five achievements, some involved constructing counterexamples to disprove old conjectures, while others provided complete proofs for new propositions, demonstrating exceptional difficulty.
GPT-6 Astra spent over $220,000 on computational power attempting all five problems, while the standard benchmark itself cost only about $20,000. Based on a cost of $300 per attempt and a 3% success rate, the expected cost to solve one significant open problem is approximately $10,000.
However, some papers point out that while models may devise correct approaches, they might fail to translate them into Lean formal proofs. The benchmark only assesses problem-solving, whereas mathematical research also values posing new questions. Sixty-three of the 68 problems remain unsolved. Even though OpenAI has declared the "arrival of the AGI era," there is still a long way to go before fully conquering advanced mathematics.

