AI May Have Cracked One of Mathematics’ Great Problems (and the Eggheads are Worried about their Jobs Now!)

[Note: submission by a young Adelaide machine learning engineer. While this deals with an issue in pure maths, what is important here is AI may have solved a problem that had stumped humans since it was posed by Jean Leray in 1934. The modern mathematical problem of singularities in the Navier-Stokes equations was officially posed as a Millennium Prize Problem by the Clay Mathematics Institute on May 24, 2000, with a cool $ 1 million prize money, so an AI solution is big news. Where is this all going?]

Most people have never heard of the Navier–Stokes problem, and fewer still could write down the equations involved, even engineers. Yet an extraordinary event in mathematics may have occurred this week. OpenAI claims that artificial intelligence has cracked a problem that some of the world's best mathematicians have been unable to settle for generations.

If the proof survives scrutiny, the achievement will be important in its own right. But its larger significance may lie elsewhere. We may have just received another indication of how quickly artificial intelligence is moving from being a machine that answers questions to being a machine capable of contributing to the creation of new knowledge.

The mathematics can be explained without equations. Imagine stirring a cup of coffee. At first the movement is relatively orderly, but very quickly swirls form within swirls. Similar things happen in rivers, smoke, ocean currents, blood flowing through arteries and air moving around an aircraft wing. Fluid motion can become extraordinarily complicated.

For roughly two centuries, physicists and mathematicians have described such motion using equations associated with Claude-Louis Navier and George Gabriel Stokes. They are among the fundamental equations of fluid mechanics and have applications ranging from aircraft design to weather modelling.

There has always been a deep mathematical problem hiding inside them. Suppose we begin with a perfectly well-behaved three-dimensional fluid. The velocity and pressure are finite and smooth everywhere. We then allow the equations to determine what happens next.

Will the mathematical solution always remain well behaved? Nobody has been able to prove that it will. The alternative is stranger. Perhaps under some circumstances the equations drive themselves towards a singularity, loosely speaking, a point at which some quantity becomes unbounded and the smooth mathematical description breaks down.

This question became so important that the Clay Mathematics Institute included it among its seven Millennium Prize Problems, with a $1 million prize attached to a successful solution.

Mathematicians have attacked it for decades. Now OpenAI says its artificial intelligence has found a route to finite-time blow-up. Rather than proving that the equations must always behave nicely, the claimed result demonstrates circumstances in which they do not.

There is an essential qualification. An announcement is not the same thing as an accepted mathematical proof. The work is extremely new and mathematicians will need to examine it carefully. There is also debate about precisely which formulation of the problem has been settled and whether it satisfies the requirements of the classical Millennium Prize problem.

So, the sensible statement at present is not that AI has unquestionably solved Navier–Stokes. It is that OpenAI claims to have produced a major breakthrough that may resolve the problem. That is impressive enough.

What is remarkable is how it was done. Reports indicate that OpenAI deployed around 10,000 AI agents working in parallel for approximately 88 hours, exchanging millions of messages and consuming around 130 billion tokens. Another verification phase followed. The estimated computational cost runs into millions of dollars.

Think about what that means. A human mathematician might spend years exploring possible approaches to a difficult problem. Most avenues lead nowhere. Ideas have to be tested, abandoned, modified and combined with other ideas. Researchers read enormous bodies of literature, talk to colleagues and gradually eliminate dead ends.

Now imagine thousands of mathematical workers who do not sleep, become bored or complain that the calculation is tedious. Give them different approaches to investigate. Allow them to communicate their results. Let promising approaches receive additional resources while unsuccessful ones are discarded. Run the entire mathematical civilisation inside computers. That is approximately the direction in which we are moving.

It is important not to exaggerate the achievement into a science-fiction story in which AI suddenly woke up and independently invented mathematics. There is already controversy over the human intellectual ancestry of the result.

Mathematicians Diego Córdoba and Luis Martínez-Zoroa had developed important ideas concerning possible mechanisms for singularity formation. Tristan Buckmaster and Levent Alpöge were pursuing related work, themselves making extensive use of AI tools. The subsequent OpenAI result has produced a dispute about priority, influence and appropriate credit. This matters because the popular picture of AI discovering something entirely by itself is probably too simple. The emerging research system looks more like a strange hybrid intelligence.

A human mathematician has an unusual conceptual insight. Other mathematicians develop it. AI systems search enormous spaces of possible extensions. Thousands of agents test variations at speeds no collection of humans could reproduce. Humans then inspect the result and determine whether the machine has actually proved what it claims.

Where, in that chain, does the discovery belong? That question is going to become increasingly difficult. Suppose a mathematician spends ten years developing an approach but cannot complete the proof. An AI system absorbs the mathematical literature, combines that approach with thousands of other results and finishes the proof in three days. Who solved the problem?

The mathematician? The people who built the AI? The company that paid for the computing power? The machine? Or all of them? Academia has not developed conventions adequate to answer these questions because until recently it did not need them.

There is an even larger issue. Mathematics has traditionally been regarded as one of the supreme demonstrations of human abstract intelligence. A calculator can multiply enormous numbers faster than a mathematician, but that never threatened the mathematician's intellectual position. Calculation and mathematical discovery were obviously different activities. That distinction is becoming less comfortable.

If AI can construct genuinely novel proofs of major mathematical results, then it is participating in reasoning at a level that until recently belonged almost exclusively to exceptional human minds.

And mathematics may be unusually suitable territory for rapidly improving AI. A mathematical argument has something that history, politics and even much experimental science lack: a relatively clear criterion of success. A proof is either valid or it is not. Formal proof systems can increasingly allow parts of mathematical reasoning to be checked mechanically.

This potentially creates a powerful feedback loop. AI proposes mathematical arguments. Other AI systems attack them. Formal systems verify individual steps. Failed approaches are discarded. Successful methods become inputs into the next round.

The process can then be multiplied across thousands or eventually millions of agents. The comparison with human research becomes uncomfortable because human beings cannot scale in the same way. Ten thousand mathematicians cannot instantly be created to investigate a problem for four days. They require perhaps twenty years of education, salaries, universities and laboratories. Coordinating their work would itself become a major undertaking.

Ten thousand AI agents can potentially be created by allocating more computing power. Today's experiment reportedly required enormous resources. That limitation matters. Spending perhaps tens of millions of dollars to solve a problem carrying a $1 million prize is hardly an economical way to win prize money. But that misses the technological significance.

The first genome cost billions of dollars to sequence. Early computers filled rooms. The first transatlantic telephone calls were extraordinary events. Technologies frequently begin as expensive demonstrations before becoming ordinary infrastructure. If the cost of AI reasoning continues to decline while capability improves, today's 10,000-agent mathematical experiment may eventually look primitive.

Then consider what happens when the same architecture is turned upon physics, chemistry, engineering and medicine. Instead of asking an AI system for a summary of existing research, scientists could give thousands of agents a research problem. One group searches the literature. Another generates hypotheses. Another attempts to destroy those hypotheses. Others design simulations, analyse data, construct mathematical models and check the work of the first groups.

Human researchers move increasingly towards the top of this intellectual pyramid, deciding which questions matter and judging the significance of the answers. Perhaps even that division will not last.

There are reasons for caution. AI systems make mistakes. A 165-page mathematical argument can contain a subtle error capable of destroying the entire result. Thousands of agents can reproduce the same mistaken assumption rather than correcting one another. And mathematical proof is a much cleaner environment than the disorderly empirical world in which incomplete information and uncertain causal relationships dominate.

There is also the question of scientific trust. Researchers increasingly use commercial AI systems while developing unpublished ideas. If the companies operating those systems subsequently produce competing discoveries, scientists will understandably ask what happens to the intellectual material they enter into them. The controversy surrounding the Navier–Stokes announcement demonstrates that this issue is no longer hypothetical.

Nevertheless, something significant has changed even if the present proof eventually fails. Only a few years ago, the popular argument about artificial intelligence concerned whether chatbots could write essays without hallucinating references; most we have still do and are unreliable. We are now discussing whether thousands of coordinated AI agents have cracked one of the most famous outstanding problems in mathematics.

That is a remarkable change in a remarkably short period. The Navier–Stokes episode actually demonstrates how deeply current AI remains connected with human intellectual achievement. The machines operate upon mathematical structures, concepts and approaches built by generations of people.

But tools can eventually transform the activity that created them. The telescope did not abolish astronomy. It created an astronomy impossible for the naked eye. Computers did not abolish physics. They allowed physicists to investigate systems that could never have been calculated by hand. AI may represent something larger because the tool is beginning to operate upon reasoning itself.

If the Navier–Stokes proof stands, historians may remember the mathematics. But they may also remember something else: a moment when thousands of artificial reasoners were assembled for a few days, pointed at a problem that had resisted generations of human mathematicians, and apparently found a way through.

The unsettling question is not what problem they solved this week. It is what happens when there are a million of them working next week. I am a recent graduate, and I have some anxiety about how long my field of work will survive this revolution.

https://openai.com/index/navier-stokes-solution/

https://github.com/openai/NavierStokesAndEuler

https://cdn.openai.com/pdf/315b36cd-ec98-4023-8342-93345194ece1/euler.pdf