OpenAI has confirmed the name of its next model family, Astra, by publishing a report showing the system solved ten previously unsolved problems in mathematics and theoretical computer science. The solutions cover fields including high-dimensional geometry, coding theory, group theory, quantum complexity, lattice cryptography, and extremal combinatorics. Mathematicians had made no progress on any of these issues for at least a decade, and much longer in most cases.
In this article
Mathematicians react to the results
Thomas Bloom, a mathematician at the University of Manchester who runs erdosproblems.com, called the results “big news” on X. He considered them more significant than the counterexample to the unit distance conjecture published in May. “Maybe not bigger than a proof of unit distance would have been, but in terms of constructions, this is big,” Bloom wrote.
Bloom also rejected the idea that AI is replacing mathematicians. He argued the claim makes little sense when the AI draws on more than a century of mathematical theory, was built by mathematicians, and was trained on everything mathematicians have ever written.
Noam Brown, one of the researchers behind the test-time reasoning technology used by Astra, said on X that OpenAI had also tried and failed to crack other major problems. “Sadly, no Millennium Prize Problems (yet),” he wrote. The Clay Mathematics Institute offers $1 million for solving each of the seven Millennium Prize Problems, but only one has been solved since the prizes were announced in 2000.
Brown added, “But also, we didn’t spend a lot on each problem. It’s possible to push test-time compute much further.” He called Astra a “major step for scientific reasoning.”
The cost of generating the solutions
OpenAI says the tokens used to generate all ten solutions would have cost about $2,000 at Sol’s API rates. After the model produced its arguments, humans worked with the same model to turn them into research papers. The model also formalized each proof in Lean, creating machine-checkable certificates of mathematical correctness, and OpenAI published a walkthrough of the model’s reasoning process for each solution.
OpenAI said its researchers helped prepare the papers and formalize the proofs, and that the company takes responsibility for their accuracy. The mathematical arguments themselves, however, came from Astra.
The company argued that claiming human authorship for a proof generated entirely by AI would misrepresent both the system’s contribution and the nature of genuine human intellectual work, pointing to the Leiden Declaration on AI and Mathematics as a reference for how credit should be assigned in AI-assisted research.
What Astra is supposed to do
CEO Sam Altman has already showcased Astra in Washington, D.C. The models are currently being tested and will be the first to go through a planned U.S. government review process that requires official approval before public release.
The project reflects OpenAI’s broader ambition to build AI systems capable of working on problems continuously for hours or even days at a time. The system is designed to handle long-running tasks and complex problems by coordinating multiple agents working together.
According to the report, Astra would form a new model class alongside OpenAI’s existing Sol, Terra, and Luna families. Whether it ships as GPT-6 or as a variant within the GPT-5 line, something like GPT 5.7, hasn’t been decided yet. There’s no release date either.
OpenAI also plans to publish a report soon showing how the company used its most advanced AI to solve ten previously unsolved math problems. The goal is to show what its current models can already do.
Regulatory hurdles ahead
The models are already in testing, according to The Information. They’re expected to be the first to go through the Trump administration’s planned new AI framework, which would require AI models to be submitted to the federal government before public release. The administration aims to finalize the framework by the end of this week.
One key question is whether the models can avoid compounding errors during long-running workflows and correct themselves when a process drifts off course as the context keeps growing. That remains a major weakness in today’s agentic systems. Multi-agent setups like Astra can also perform worse on tightly linked tasks such as planning because coordination overhead and compounding errors can wipe out any gains.
Long-term goals for autonomous research
The Astra rumors line up with earlier statements from the company. Chief Scientist Jakub Pachocki said on OpenAI’s official podcast last summer that the company wants to build AI systems that can work on a problem for hours or days. Current systems are often limited to short tasks, but OpenAI wants models that can plan, reason, and experiment over longer time horizons. Late last year, the company even raised the question of how to think about systems that could solve tasks a human would need centuries to complete.
By March 2028, OpenAI wants to have a fully autonomous AI researcher that can run research projects on its own, and that system would also depend on long-running AI processes. As early as this September, the company plans to have an AI system with research-intern-level skills that would significantly speed up human scientists. Astra could end up being that system.
Pachocki also said these systems will need far more compute. OpenAI’s long-term infrastructure plans reflect that ambition. Whether the startup’s revenue grows fast enough to fund that massive buildout remains an open question.
What it means for researchers
For people making things, the change is a shift in how long a single workflow can run without human intervention. Current tools often struggle when a task stretches over hours or days, leading to errors that accumulate as the context grows. Astra is built to coordinate multiple agents over extended periods, aiming to keep a process on course when it might otherwise drift.
However, this capability comes with risks. If the coordination between agents fails, the system can perform worse on tightly linked tasks like planning. For creators relying on these tools, the promise is greater autonomy, but the reality depends on whether the system can maintain accuracy over time without constant human oversight.




