OpenAI’s latest chatbot excels in mathematical tasks but falls short in maintaining academic rigor compared to humans.
Recently, OpenAI unveiled 10 new AI-driven mathematical advancements discovered during the development and testing of its forthcoming large language model. This collection of findings emerged from the LLM, Astra, each addressing or advancing a distinct “long-standing open problem” of significant interest to the mathematical community. According to the company, the total cost for token usage was only $2,000.
The announcement quickly gained attention as a signal of AI’s potential—or threat—to surpass human capabilities and become a transformative force in mathematics and computer science. However, in the days after the announcement, as experts meticulously examined the nearly 250-page document, many expressed their dissatisfaction.
On supporting science journalism
If you’re enjoying this article, consider supporting our award-winning journalism by subscribing. By purchasing a subscription you are helping to ensure the future of impactful stories about the discoveries and ideas shaping our world today.
Experts point out that two of the most intriguing results include existing ideas from recent mathematical literature without appropriate citations. This contradicts OpenAI’s initial press release, which claimed that the problems tackled by Astra “have been open and seen no progress on the main result for at least a decade.” (The language in the release has since been updated for accuracy).
“They are disregarding the contributions of previous researchers deliberately,” says Steven Miller, a mathematician at Yeshiva University, who claims that OpenAI has effectively plagiarized his research. “It seems to be a systematic issue, pointing to research misconduct.”
Miller refers to a problem involving how many balls can fit in a box—a seemingly straightforward question, except these balls and boxes exist in a mathematical space of 1,000 dimensions or more. OpenAI’s paper refines the best estimate for how densely these balls can be packed. The proof generated by the LLM relies on a specific mathematical argument initially presented in a 2016 paper by Miller and a collaborator.
Another result from the 10 resolves a long-standing question in group theory, which examines sets of mathematical objects known as “groups” that interact in a structured manner. Mathematicians have long debated whether all groups possess a property called “soficity,” the ability to be accurately approximated in a specific way by other, simpler groups. OpenAI’s paper identifies at least one group lacking this property.
The finding surprised Francesco Fournier-Facio, a mathematician at the University of Cambridge who studies group theory—until he “engaged with this breakthrough as I would if a human had written it,” he notes. The result, he and his colleagues discovered, wasn’t as groundbreaking as it initially seemed. Similar to several recent AI breakthroughs, it combines ideas from existing mathematical literature to create a new theorem. The LLM’s strength lies in its remarkable patience for assembling disparate elements rather than making profound intellectual leaps.
Astra’s crucial mathematical step merged ideas initially found in two papers from 2016 and 2019. Andreas Thom, a mathematician at Dresden University of Technology and co-author of both papers, summarized the result on MathOverflow.com, describing it as “creative and at the same time elementary.”
OpenAI’s initial press release appeared to overlook—or be unaware of—these important recent developments. Fournier-Facio contends that the earlier papers demonstrate that humans had not reached an impasse with the soficity problem. OpenAI’s mathematicians endeavored to attribute these ideas correctly in their paper, he notes. However, despite their good intentions, “there is the big PR machine that wants to sound as impressive as possible and does not care about being 100 percent accurate,” he adds.
“We take responsibility for the correctness of these results and are meeting the same standards generally expected of human mathematicians,” an OpenAI spokesperson stated to Scientific American. “We plan to make small updates [to the paper] this week, consistent with standard academic practice.”
As AI continues to advance in mathematics without adhering to the field’s academic norms, some community members are clearly growing impatient. “OpenAI is now fully participating in high-level research,” Fournier-Facio says. “So they should be held to the same academic standards that we are.”
It’s Time to Stand Up for Science
If you enjoyed this article, I’d like to ask for your support. Scientific American has served as an advocate for science and industry for 180 years, and right now may be the most critical moment in that two-century history.
I’ve been a Scientific American subscriber since I was 12 years old, and it helped shape the way I look at the world. SciAm always educates and delights me, and inspires a sense of awe for our vast, beautiful universe. I hope it does that for you, too.
If you subscribe to Scientific American, you help ensure that our coverage is centered on meaningful research and discovery; that we have the resources to report on the decisions that threaten labs across the U.S.; and that we support both budding and working scientists at a time when the value of science itself too often goes unrecognized.
In return, you get essential news, captivating podcasts, brilliant infographics, can’t-miss newsletters, must-watch videos, challenging games, and the science world’s best writing and reporting. You can even gift someone a subscription.
There has never been a more important time for us to stand up and show why science matters. I hope you’ll support us in that mission.

