The Verge AI原文 · English

‘Pure insanity’: Mathematicians will need years to make sense of OpenAI’s latest drop

"Staggering." "Overwhelming." "Unprecedented." "Surreal." "Pure insanity." Those were among the descriptions more than three dozen mathematicians reached for in conversations with The Verge as they tried to make sense of…

画像の出典 · The Verge AI

原文 · English

“Staggering.” “Overwhelming.” “Unprecedented.” “Surreal.” “Pure insanity.”

Those were among the descriptions more than three dozen mathematicians reached for in conversations with The Verge as they tried to make sense of the flood of mathematical results OpenAI abruptly dropped on the field this week. Amid the awe, excitement, and uncertainty over the sheer scale of the deluge was a deep-seated anxiety over what it all means — and what comes next. For all their different reactions, researchers agreed that simply understanding what OpenAI had released could take years, let alone figuring out where the mathematicians themselves fit in the field now changing around them. Many feared OpenAI would not wait that long before moving on — or releasing even more.


In all, OpenAI released nearly 400 AI-generated results. These were spread across more than 700 manuscripts and covered a diverse array of mathematical disciplines, including combinatorics, several branches of geometry, number theory, theoretical computer science, algebra, topology, probability and statistical mechanics, and mathematical physics. The collection is so vast that OpenAI felt the need to publish guidance on how to navigate the sprawling GitHub repository.

Related

The sheer volume of work makes even a preliminary assessment as to exactly what the company has released difficult. In the hours and days following the drop, most mathematicians The Verge spoke with said they were still struggling to digest everything; several said that simply working through the roughly 40-page table of contents and abstracts took them the better part of an hour. “Just going over the entire list of abstracts is overwhelming,” said Álvaro Lozano-Robledo, a professor of mathematics at the University of Connecticut.

Sprinkled among the hundreds of manuscripts are formalizations in Lean, a programming language and proof assistant that allows results to be verified computationally. These formalizations have proven instrumental in assessing some of OpenAI’s previous mathematical claims, giving researchers confidence that a claim is logically correct even if they don’t fully understand the argument behind it.

“If the AIs would disappear now, as though there were aliens that came to Earth and then just left, we would be studying this for the next 10 years, trying to understand everything.”

But the degree to which each result had been verified varied wildly. On GitHub, OpenAI acknowledged that the results are “at different stages of verification” and that “many, but not all, of the manuscripts have been formalized.” As of writing, fewer than half the manuscripts in the collection appear to have been described formally. OpenAI said only 300 top-line results out of 719 manuscripts had been formalized, around 42 percent, and that it “will update the repository with more formalizations as we obtain them.”

Several mathematicians complained to The Verge about the lack of formalization, particularly given the sheer number of results, and stressed that even when Lean code accompanies a result, evaluation isn’t instantaneous. Researchers must check that the formalization actually proves what the result claims, another time-consuming process, and several digging through the papers said that even where computer-verifiable proofs had been provided, the quality was inconsistent and the statements they verified did not always appear to map neatly onto the claims in accompanying manuscripts.

Kevin Buzzard, a mathematics professor at Imperial College London, said he had identified numerous theorems in his area of work — algebraic number theory — of which only around six immediately “stood out.” Few, if any, of those appeared to be formally verified in Lean. “Hence, I either have to read possibly-not-correct slop, or wait for others to do the same, or wait for someone to formalise them before I can say for sure that the results are even correct.” Buzzard’s concerns were echoed by numerous other researchers.


Buzzard was far from alone in worrying about AI “slop.” The term is a shorthand for low-quality, frequently erroneous AI-generated material that increasingly crops up online — and in the real world — including academic papers. As in other fields, mathematicians told The Verge they have seen a huge uptick in such material produced with tools like ChatGPT and Claude in recent years. Much of it is confusing, hard to read, and demonstrates little understanding of the subject; it is especially shoddy when it comes to crediting other researchers.

OpenAI’s previous mathematical write-ups were widely criticized by experts for their sloppy nature, particularly their poor or nonexistent attribution. In conversations with The Verge ahead of the release, several researchers had taken to calling the impending flood of papers the “slopocalypse,” or similar variations on the theme.

Whether the feared “slopocalypse” actually materialized is difficult to say, largely due to the bewildering volume of material released. Early indications suggest OpenAI took more care with papers this time around, or at least with some of them. Several mathematicians told The Verge that their first impressions were far better than they had expected, though by their own admission that was hardly a high bar given the company’s previous shoddy publications.

“It’s a big mess. It can cause a huge collapse in the academic culture and simply kill most of the faculties. It’s a social problem and it seems that the AI labs are completely ignoring this issue.”

But better does not necessarily mean good, let alone up to the standards usually expected of academic work making claims of this magnitude. With formal verification absent for many of OpenAI’s claims, the quality of the accompanying papers becomes particularly important; they are the primary means by which mathematicians can confirm, understand, scrutinize, and contextualize the results.

Producing rigorous mathematical papers is difficult work under even the best of circumstances. Doing so at this kind of scale is a formidable undertaking. OpenAI’s models are pumping out mathematics at a dizzying speed and across a broad range of specialities, far outstripping the capacity of its human staff. The company simply does not have the breadth of expertise or resources to properly scrutinize its findings at the cutting edge of mathematics. Researchers told The Verge that it shows.

Many described papers that were difficult, sometimes practically impossible, to follow. “The write-up of the problem I know best made little sense after a quick read,” Brendan Hassett, a mathematics professor at Brown, told The Verge. “If this had been written by a person, I wouldn’t spend any more time trying to understand it. Of course, this leaves 721 other preprints!” (Hassett said this before OpenAI retracted three papers).

Some researchers told The Verge several papers they or their colleagues had noticed appeared to cover ground already trod by other mathematicians, though were wary of saying so publicly before they had a chance to properly review the material. Others pointed to the unusual brevity of the work, with results they would ordinarily expect to be developed over hundreds of pages compressed into a few dozen or less.

Related

Nalini Joshi, a mathematics professor at the University of Sydney in Australia, said a quick search through the release revealed little that overlapped with her own work, but noted that some of the papers she examined “have short bibliographies.” Given previous criticism of OpenAI’s crediting practices, she said she is “wary that attribution in the papers may be lacking the complete story.”


But for all the slop and uncertainty, the overarching consensus is that the release contains some genuinely impressive work.

While stressing the difficulties assessing the volume of material — and the need to properly verify the results — numerous researchers The Verge spoke to said the work appeared to be of a very high caliber, despite shortcomings in its presentation. In a pre-AI world, they said, many of OpenAI’s results would clearly have warranted publication in top-tier journals and could have been enough to secure an academic career for their authors. A handful were described as being the kind of work that could make a mathematician a serious contender for a Fields Medal, one of the discipline’s highest honors.

“I either have to read possibly-not-correct slop, or wait for others to do the same, or wait for someone to formalise them before I can say for sure that the results are even correct.”

Stanford mathematician Jared Duker Lichtman said there were “tens” of results he would put in this category, including progress toward the Riemann hypothesis, a special case of the Hodge conjecture, and a solution to the four-dimensional Kakeya conjecture. These are major problems in mathematics. Riemann — arguably the most notorious unsolved problem in the entire discipline — and Hodge are both among the seven famed Millennium Prize problems. Respectively, they are concerned with the distribution of prime numbers and, very roughly, how complex geometric shapes can be understood in terms of simpler building blocks. Kakeya, meanwhile, roughly asks how little space is needed to rotate a needle or pencil in every direction. A proof to the three-dimensional version of Kakeya was among the achievements NYU mathematician Hong Wang was awarded a Fields Medal for earlier this year.

“It’s not the case that these are just silly problems that no one’s ever heard of,” mathematician Scott Armstrong said. “Many of them are like very well-known problems that many people have tried for decades.”


As the dust began to settle, mathematicians were left confronting a landscape that had abruptly changed around them. Many of them had just watched years of work and carefully laid research plans evaporate in an instant, or knew colleagues who had. Across the field, whether reactions were laced with excitement, dread, despair, or something in between, there was a profound sense of disorientation.

By the following morning, that mood had not altered much. Speaking by phone as he walked through Paris, Armstrong imagined that, had the Sorbonne not been closed amid ongoing student protests and his colleagues were gathered around their usual coffee machine, the atmosphere would have been rather “solemn.”

Related

In Scotland, Colva Roney-Dougal, a professor of mathematics at St Andrews University, described a similarly bleak atmosphere among colleagues and students. The deluge felt somewhat “horrific,” she said, but its arrival brought some relief after weeks of uncertainty. “I have a bunch of friends and colleagues whose grant proposals have just been wiped out,” she said.

“I have a bunch of friends and colleagues whose grant proposals have just been wiped out.”

Armstrong said he knew of one group whose entire research program was practically “wiped out” by the release. Tristan Buckmaster, an NYU mathematician who was at the center of OpenAI’s earlier dispute over the Navier-Stokes problem, meanwhile, said he had already heard of “three people who had entire research programs obliterated.”

Similar stories surfaced repeatedly in The Verge’s conversations with mathematicians, though the disruption was concentrated heavily in some areas of the field. Francesco Fournier-Facio, a professor of mathematics at Heriot-Watt University in Scotland, said researchers in probability, combinatorics, and theoretical computer science appeared “particularly in shock,” adding that some regions of his field, group theory, had been “bulldozed.”

Armstrong and other researchers said some areas appeared to have been aimed at with almost military precision. “There were definitely some targets,” said Armstrong. He singled out work in the area Wang received the Fields Medal for this year, as well as several Millennium Prize problems. A cluster of results in mathematical physics, he said, makes it clear OpenAI is pursuing Yang-Mills theory — the mathematical framework underpinning much of modern particle physics — and its unresolved “mass gap” problem, one of the seven Prize problems.

Elsewhere, researchers expressed a kind of surprised relief at how little their own work appeared to have been touched. Roney-Dougal said her own corner of the field — largely centered on group theory — appeared to have escaped relatively unscathed. She joked that it is fortunate she works in a “very unfashionable area,” though later messaged to say she was feeling “unsure” after finding her work cited in one of the papers.

Joshi similarly found little overlap with her work, which is largely focused on integrable systems, complex systems that can be solved exactly. She suggested that may be because her field is less driven by long-standing, formally stated conjectures and definitions and more by questions that wax and wane with developments in physics.

As of October 8th, that record already listed numerous corrections, including revisions on more than a dozen manuscripts and the removal of three papers due to a “sign error” it says invalidates an argument.

Whether their own work had been directly affected or not, most mathematicians The Verge spoke to shared a sense that the field had crossed some kind of threshold and was no longer the same as it had been a day earlier. Constantin Kogler, a researcher at the Institute for Advanced Study, called it “the most important single moment in the history of mathematics,” while Yang-Hui He, a fellow at the London Institute for Mathematical Sciences, reached back millennia, comparing AI-driven developments this year to the publication of Euclid’s Elements, one of the discipline’s most influential works.

Few researchers were quite so grand in their assessments, but there was a general consensus that mathematics was changing fast — and that mathematicians would have to change with it.


OpenAI knew the release would be disruptive and had made some effort to soften the blow and engage with the mathematical community. After several bruising encounters with researchers earlier this year, the company turned to mathematicians themselves for help, working with the newly formed Advisory Group on Mathematics and Artificial Intelligence (AGMAI) on how to release the results responsibly.

AGMAI had already laid out what it believed responsible engagement would be. AI labs should ideally release papers “that a human understands” or otherwise provide the necessary “support for the additional mathematical activities that are needed for humans to be able to understand and assimilate their AI-generated mathematical output and identify possible applications of it.” They said that, where possible, proofs should be formalized, and that companies should make public what models and prompts they used to create them. The group also urged AI labs to stop treating mathematical releases as “marketing vehicles to promote their models” and should “stop testing advanced mathematical problems on proprietary models” that are inaccessible to the broader scientific community.

“It’s not the case that these are just silly problems that no one’s ever heard of. Many of them are like very well-known problems that many people have tried for decades.”

OpenAI followed some of the group’s advice. It said it would be funding a series of workshops and conferences around its work in mathematics to help the community process and understand its work, though it provided no details as to what these might look like or when they may occur. The company also disclosed significantly more information than it had in previous releases — again, researchers said this was a low bar — including that its model attempted more than 4,000 problems and that a typical result used around three hours of ChatGPT Pro thinking compute. It has also implemented “protocols for paper revisions and citations” and on GitHub said it “will preserve the public release history” of the collection. As of October 8th, that record already listed numerous corrections, including revisions on more than a dozen manuscripts and the removal of three papers due to a “sign error” it says invalidates an argument. OpenAI did not respond on the record to The Verge’s request for more details.

That change addresses a key source of friction with mathematicians, some of whom speak of a kind of “paranoia” around the company’s published claims. The company has shown a habit of surreptitiously altering press releases and papers in response to criticism without clearly disclosing those changes.

But OpenAI did not follow all of the group’s recommendations — including some of its most consequential. It did not identify the model behind the release nor did it disclose the prompts used or the full set of problems the model attempted. OpenAI also appeared to acknowledge that its formalizations were lacking — it said it will add more as it obtains them — and that papers were subpar, saying that “for future releases, we are committed to further improving the quality of the papers via the citations, mathematical exposition, and presentation of the results for better understanding.”

“The write-up of the problem I know best made little sense after a quick read.”

Researchers The Verge spoke to questioned why OpenAI couldn’t have improved the quality of these papers and expressed disappointment that it seemingly couldn’t be bothered to formalize — or even check, in the case of the retracted papers — many of the results ahead of time. Information like the prompts used would have also been incredibly useful in lessening what many felt was the burden of assessing the flood of material the company suddenly dropped on them.

Most importantly, the company indicated it has no plans to stop testing its models on mathematics problems, and indeed suggested it sees doing so as an imperative. “We want to directly empower scientists with state-of-the-art capabilities and are working to responsibly release the model that produced these results,” it said when announcing its latest results. “This is why it is important to continue to evaluate our internal frontier models on mathematics and other sciences, so we can accelerate developing the tools to advance those fields.”

After OpenAI offloaded its results, AGMAI described the publication as “a first step,” and reiterated calls for equitable access to compute and research tools. More fundamentally, the group argued that “the future of mathematical research cannot consist only of understanding results produced by AI labs.”


For all the hundreds of long-standing problems OpenAI claims to have solved, mathematicians say there is an extraordinary amount of work to be done. Many estimated that making sense of everything the company just released could take the community years.

Armstrong likened the sudden arrival of so much new mathematics to the commotion that might follow a brief extraterrestrial visit. “If the AIs would disappear now, as though there were aliens that came to Earth and then just left, we would be studying this for the next 10 years, trying to understand everything.”

Much of that will be exciting in its own right. Researchers will have to fill in gaps, extract useful ideas and connections, and place solutions in a broader mathematical context, all things that typically go hand in hand with producing solutions for humans. “For many of these results, relevant experts care deeply about the solutions and I am confident that they will be able to digest and present them to the wider world,” Lichtman said. Work involving explaining, verifying, and contextualizing results will need to be valued more as AI changes the fields, he argued. “As a society, we should be rewarding these digestive efforts more.”

There may be plenty of mathematics left for humans to do, too. As far as he could tell, Lichtman said all of the results, “amazing” as they are, come “out of existing methods and techniques.” They fill in parts of the known mathematical landscape rather than creating entirely new ones. “It turns out there is a lot more room to fill in than experts previously knew!” Lozano-Robledo echoed the point: “They are all using existing techniques in very ingenious ways.”

And the existing map is hardly the limit. “Math research is essentially infinite,” Lozano-Robledo said. “Research will go on.” The London Institute’s He said he was particularly excited to see what emerges next, speculating that researchers may end up creating new ideas and “maybe even new fields of math.”

The companies seem ready to move on instantly, long before the community has had any reasonable chance to digest what they’ve produced.

Few mathematicians The Verge spoke to objected in principle to AI producing new mathematics. Indeed, many welcomed it and, to varying degrees, said they were enthusiastic users of AI tools themselves. Most of the unease was directed at how the AI companies building those tools were going about entering their field.

Looking at the mathematical releases of OpenAI and other AI labs, you’d be forgiven for thinking that research mathematics essentially amounts to checking problems off a list. The manner in which AI labs unveil their results reinforces this idea: a flashy announcement, a preliminary write-up, and the rest tossed to mathematicians to figure out and contextualize. Multiple researchers told The Verge that the companies seem ready to move on instantly, long before the community has had any reasonable chance to digest what they’ve produced.

Researchers described that as an impoverished view of how mathematical research actually functions. Solutions certainly matter, but so does everything that happens on the way to finding them: developing ideas, making connections, and finding new questions or opening up other avenues of research. When companies like OpenAI just find answers without doing any of the surrounding work, mathematicians say the burden of doing so falls back on the community.

“It’s nice to have new results, but this scale is a different level,” said Bartosz Naskręcki, a mathematician at the Adam Mickiewicz University in Poznań, Poland. “It’s a big mess,” he said. “It can cause a huge collapse in the academic culture and simply kill most of the faculties. It’s a social problem and it seems that the AI labs are completely ignoring this issue.”


It’s not just the mathematics that will take years to understand. Researchers are also struggling to comprehend what the upheaval means for them. Many described a bleak, almost existential mood hanging over the community as mathematicians decipher where they fit in the rapidly changing field. They may not have much time to figure it out.

Rumors are already circulating among mathematicians about further releases from OpenAI, which did not respond to The Verge’s questions about whether more are planned.

Roney-Dougal said she dreaded another drop — or the prospect of Anthropic or another AI lab entering the fray.

The disruption is hitting PhD students, junior researchers, and others without permanent tenure particularly hard. Open problems like the ones OpenAI’s models are plowing through can underpin dissertations, grant proposals, job applications, and years of planned research, all important building blocks for academic careers. Several researchers described an increasingly pervasive anxiety that the work they are building their careers around could suddenly be next and that the AI models generate few avenues for future inquiry as they close off others. “For someone at the end of funding or looking for a job now this is incredibly disruptive,” said Simon Machado, a mathematician at ETH Zurich in Switzerland who is joining the French National Centre for Scientific Research (CNRS).

“Math research is essentially infinite. Research will go on.”

Morale may be particularly low because everything feels relentless. OpenAI’s findings have arrived in increasingly large waves, with little pause between them and frequent indications from the company that yet more are on the way. “You get a sense that they don’t really care about these individual results,” Roney-Dougal said. “It’s a really nasty feeling.”

Mathematicians barely have time to digest one set of results before the next, bigger set is dropped on them. “There’s going to be something in two more months that’s going to blow this away,” Armstrong said. “It’s just, like, repeated strikes by bigger and bigger bombs.”

Armstrong said he considers himself among the mathematicians most enthusiastic and optimistic about AI. He uses the technology in his own work and believes it has great potential. But even he is growing increasingly uneasy. “It is kind of hard to sleep at night,” he admitted.

Asked what he planned to do next, Armstrong was less certain. Still walking through the streets of Paris on his way to get lunch, he said his immediate priority was to finish some work he’d been doing — sometimes with the help of OpenAI’s Codex — “before OpenAI scoops us.”

Beyond that, even the self-described AI optimist was unsure and questioned what kind of future there was for him in mathematics. He worried particularly about the young researchers in the field. “It’s going to be a really weird few years in math,” he said.

Follow topics and authors from this story to see more like this in your personalized homepage feed and to receive email updates.
原文の出典

The Verge AI

内容について

原文の公開と権利は出典元に帰属します。

この言語の全文翻訳を準備中です。現在は保存済みの原文を表示しています。