What can we learn from Astra’s results on the ten maths problems?

OpenAI’s new model has tackled questions that have remained unresolved for years. The company initially announced a triumph, then issued some clarifications. The truth lies somewhere in between: a system that is less miraculous than the initial headlines suggested, yet more interesting than the dismissive reactions would have us believe

8 AUG 26
Translated by AI
Image of What can we learn from Astra’s results on the ten maths problems?

Photo: ANSA

Astra, a new OpenAI model still in-house, is said to have tackled ten long-standing mathematical problems, achieving results that in some cases resolved major issues, at a total cost of just $2,000. The announcement came following a series of rapid advances by large language models in mathematics, which, within the space of a few years, had progressed from school exercises to the International Mathematical Olympiad and then on to the problems currently being tackled by researchers. If the figures were as suggested by the company’s account, the leap would have been impressive: problems that had remained beyond the reach of mathematicians for years would have become solvable for just a few hundred dollars each.
What OpenAI actually published on 1 August is already slightly different from this version. The title chosen by the company refers to ‘ten breakthroughs’, and the results include both solved problems and substantial progress towards their solution. However, the material is far more substantial than a press release: it comprises around 250 pages of mathematics, accompanied by a reconstruction of the process followed by the model, and each proof has also been translated into a form that can be automatically verified by a specialised programme. OpenAI states that the mathematical arguments were generated by Astra, whilst humans prepared the manuscripts with the help of the model itself and took responsibility for their accuracy.
This makes the case particularly interesting, as it offers far more material to scrutinise than the usual claims about the performance of artificial intelligence. And mathematicians immediately set about doing just that.
The first finding of the review concerns one of the very sentences that made the announcement so sensational. In the version initially published, OpenAI stated that the ten problems were open and had seen no progress on the main result for at least ten years, and in most cases for much longer. That sentence no longer appears today. It has been replaced by the far less ambitious statement that Astra has solved or made progress on long-standing open problems. The web archive preserves both versions, so the change is fully verifiable.
The correction was necessitated primarily by two of the results that had initially attracted the most attention. One concerns an ancient problem which, to put it very simply, seeks to establish how densely spheres can be packed when moving from familiar spaces to spaces with an enormous number of dimensions. Astra achieved a new limit – in other words, a genuine mathematical result; however, Steven Miller, a mathematician at Yeshiva University, observed that a key step in the proof had already appeared in a paper published in 2016 in collaboration with Henry Cohn, and accused OpenAI of failing to give adequate credit for its origin. Miller went so far as to speak of plagiarism and scientific misconduct. At present, this is an accusation made by one of the researchers involved, rather than a judgement issued by a journal or a scientific integrity body; nevertheless, the issue of attribution that he has raised is a genuine one.
The second case is even more instructive, because it shows just how wrong it would be to swing from initial enthusiasm to the opposite conclusion and dismiss the whole thing as a sophisticated act of plagiarism. For many years, mathematicians had been wondering whether there existed at least one member of a certain vast family of abstract objects that did not possess a property thought to be perhaps universal. Astra constructed one, thereby settling the question of its existence. However, upon examining the proof, Francesco Fournier-Facio of the University of Cambridge and other specialists recognised two key elements drawn from papers published in 2016 and 2019. Here too, therefore, the original portrayal of a field that had remained largely stagnant for at least a decade did not hold up.
Yet one of the authors of those earlier papers, Andreas Thom, has provided on MathOverflow a far more interesting assessment than the summary verdicts – for or against – that are popping up everywhere. According to Thom, Astra has found a creative way to combine the two previous results, overcoming an obstacle that prevented their direct application. Thom explains that he himself had been searching, ever since the 2019 paper, for a mechanism capable of making that strategy work, and states that he admires the efficiency of the construction found. The novel idea, therefore, does not lie in having invented a theory out of thin air that no one had ever conceived: it lies in having seen how existing pieces could be linked together to reach a point their authors had not yet reached. In mathematics, this too is called research, and it can be very good research.
This is probably the most important aspect of the entire experiment. A large language model possesses an extraordinary ability to explore combinations, follow a line of inquiry at length and reassemble tools developed in different contexts. Senén Barro, director of the Artificial Intelligence Research Centre at the University of Santiago de Compostela, observes that current systems seem particularly effective precisely when the problem is well-defined and the solution can be constructed using elements that have some precedent in the literature. The creation of entirely new concepts and new theoretical frameworks remains another matter altogether. Astra’s work suggests, however, that between the repetition of something known and the establishment of a new theory lies a vast territory – and it is precisely this territory in which a substantial part of real-world mathematics takes place.
Automatic verification of proofs must also be properly contextualised. Computer-based checking is of great value, as it allows us to establish – with a degree of certainty that would be difficult to achieve by reading through hundreds of pages – that the formal steps do indeed follow from the stated premises. However, this method cannot determine whether an idea is truly new, whether it was published seven years earlier in an equivalent form, whether the correct citation has been omitted, or whether a result deserves the importance attributed to it in a press release. For this reason, following the machine-based verification of proofs, mathematicians began their own examination of the literature and the history of the problems. Within a few days, this latter process had already significantly altered the interpretation of at least two of the most important results.
Then there is the figure most likely to stick in the headlines: $2,000. OpenAI states specifically that the number of tokens required to find the published solutions would cost roughly that amount at its API rates. This is an interesting figure, as it shows just how low the marginal cost of producing a search of this calibre can be once a model capable of doing so exists. However, it does not tell us how much it cost to achieve those ten successes: the denominator is missing. We do not know how many other problems were submitted to Astra without yielding any results, how many attempts failed before arriving at the selected examples, what human labour was required to select the problems and verify the results, nor what proportion of the enormous resources needed to build the model should be taken into account when assessing the overall efficiency of the process. Barro used a very simple analogy: the cost of a winning lottery ticket does not tell us how much the player spent, unless we know how many tickets they bought. The $2,000 reported by OpenAI may therefore be a perfectly accurate measure of what the company claims to be measuring and, at the same time, insufficient to support the claim that ten research-level maths problems were ‘solved for $2,000’.
This distinction becomes even more necessary when the technology’s developer is also the one promoting its capabilities. OpenAI operates in a sector where continually demonstrating the increasing capabilities of its models also serves to justify investments on an extraordinary scale. Reuters has reported a forecast of approximately $25 billion in cash burn for 2026, whilst the company continues to plan massive investments in computing capacity. A mathematical result does not become any less true for this reason; however, it is essential to carefully separate the scientific content from the promotional context in which it is presented.
This situation takes on particular significance because, in the same announcement, OpenAI explicitly refers to the Leiden Declaration on AI and Mathematics, drawn up by the mathematical community and endorsed by the International Mathematical Union. That document calls for the use of artificial intelligence to be disclosed, for precedents to be cited accurately, and for work expedited by machines not to place an unsustainable burden on reviewers. Addressing public decision-makers, it also contains a highly pertinent recommendation: the technology industry has strong commercial incentives to exaggerate the capabilities of its systems, and to assess its claims, specialists should be consulted rather than relying on press releases.
This is exactly what happened with Astra. The press release was published on 1 August and, within a week, specialists from various fields had begun scrutinising hundreds of pages, tracing the lineage of ideas and publicly debating which parts of the findings were genuinely new. OpenAI has already amended one of its strongest claims, whilst discussions continue and the evaluation of the other results will inevitably require further work.
Following this examination, Astra appears perhaps less miraculous than the initial headlines suggested, and more interesting than the dismissive reactions would have us believe. A general-purpose artificial intelligence system has produced, in a single campaign, a considerable amount of professional-level mathematics, achieving results in at least some cases that specialists consider significant. In one of the most talked-about cases, it identified the connection that one of the authors of the previous tools had been searching for unsuccessfully since 2019. At the same time, the initial account of the state of research was overly favourable to the sensational nature of the announcement; some attributions have been disputed; and the much-cited cost of $2,000 describes only part of what would be needed to assess the experiment’s effectiveness.
The most significant story of recent days therefore also concerns what happened after Astra had finished its work. Producing a proof, establishing how much it owes to previous literature and determining its true scientific significance are tasks that can now be distributed in a new way between machines and people. The speed with which the mathematical community has corrected the narrative – without, however, discarding the results that have stood up to scrutiny – shows just how important such scrutiny will become when producing hundreds of pages of research costs less and less. If Astra truly points to the future of AI-assisted mathematics, these first few days have demonstrated both its capabilities and the system of checks and balances it will require.
With one caveat – the usual one – which we can summarise in a question: who will actually have access to Astra and equivalent models, at what cost, and with how much transparency? At present, for example, Astra is not available outside OpenAI; for this reason, if the science of the future is truly to be one of collaboration between AI and humans – that is, if this model works well and proves superior to the traditional one – we run an ever-greater risk of it becoming the monopoly of a few, rather than remaining, by definition, an open and democratic endeavour.