A couple of years ago, I taught a professional development course at a secondary school to introduce teachers to the informed use of artificial intelligence. During the course, we experimented with these tools across different subjects, and when it was the mathematics teacher’s turn, I warned him that one of these machines’ weak points was precisely the logical-mathematical domain. The results we obtained, while respectable, did contain some inaccuracies in his view. At the time, it seemed understandable to me that a system trained to generate linguistic sequences on a statistical basis would struggle in fields where reasoning has to follow precise rules. For now, better not to trust it, I advised. Tomorrow, we’ll see.
Today is tomorrow. As chance would have it, I returned to the same school for another course and had the opportunity to revise my earlier position: thanks to the various developments of recent months, LLMs are now exceptional tools for mathematics as well. In May 2026, an internal OpenAI model generated a proof disproving an Erdős conjecture; a group of mathematicians verified the argument and published a simplified and partially generalized version. The broader problem remains open, but an important conjecture has fallen. A few days ago, on September 8, 2026, OpenAI proposed a solution to the existence and regularity problem for the Navier–Stokes equations, one of the so-called “Millennium Problems” in mathematics. The company released a paper and a formalization in Lean. Rather than calling it an actual solution, it is more prudent to describe it as a proposed solution accompanied by formal verification, still awaiting examination and understanding by the community of specialists. But that is no small thing. It is now inevitable that a growing share of new mathematics will advance through the intensive use of language models.
The mathematical community is reacting in different ways. Timothy Gowers acknowledges the significance of the results, while distinguishing them from any claim that machines are superior in every aspect of research. Terence Tao emphasizes the value of understanding and the need to make proofs readable. The Leiden Declaration calls for independent oversight and public infrastructure, linking the autonomy of research to the ownership of the tools it relies on.
These reactions have now been joined by the open letter A Severe Misalignment of AI in Mathematics, signed by twenty-five Fields Medalists. The authors acknowledge the progress made by these models, but fear that the race among companies to solve famous problems may undermine the understanding of ideas and the training of researchers; they also raise questions of attribution and plagiarism. I do not have the expertise to fully assess the weight of these objections for mathematical research. I cannot help noticing, however, that whenever a machine becomes better than someone at something, a letter of protest usually follows.
Recent developments allow for a little speculation: if today we can use a machine to produce an argument that we refine and understand afterwards, in the future we may receive results so complex that they exceed the understanding of the scientific community as a whole. This is only one possible scenario, since machines might also help us make comprehensible results that currently escape us. Still, it is difficult to justify the certainty that the human mind will always be able to follow every development in knowledge. The idea of relying on knowledge we do not understand can frighten us, even though our daily lives already depend on knowledge we do not possess.
I board an airplane without knowing how to build one, reassured by the fact that someone else knows how. And yet the knowledge required to design and maintain the aircraft does not reside in the head of a single person, but is distributed among specialists, none of whom necessarily needs to understand every detail for the system to work. Our autonomy already rests on an immense dependence on other people’s knowledge. We use airplanes, machines, medicines and computers without knowing how they work.
One might object that this knowledge, although distributed, is still possessed by one or more human beings, that there are reproducible procedures, that we can learn them, and so on. Yet there are various cases in which we used something before anyone adequately understood its mechanism. Penicillin was used clinically in the early 1940s, while the explanation of its molecular action became more precise over the following decades. In 1965, Tipper and Strominger connected its molecular structure to the inhibition of the formation of the bonds that strengthen bacterial cell walls. Its therapeutic effect and some of its consequences for bacteria were already observable; what was missing was an equally precise description of the mechanism. A similar gap separates the first steam engines from thermodynamics. Newcomen’s engine dates back to 1712, while Carnot published his reflections on heat engines in 1824. The machines worked thanks to practical knowledge and partial explanations, while the theory that today describes their efficiency and limitations would take shape during the nineteenth century. Aspirin, too, had been in use for decades when, in 1971, John Vane showed that it inhibited prostaglandin synthesis, clarifying a central mechanism of its action. And the production of bread and wine predates nineteenth-century research into fermentation by thousands of years: Pasteur helped explain the role of microorganisms in processes that people had been carrying out for centuries. There is a distinction between what we know how to do and the different levels at which we know how to explain it.

These precedents do not justify indiscriminate trust. Effectiveness must be tested, errors must be detectable, and understanding a mechanism is valuable from countless points of view. They do show, however, that the usefulness of a capability can precede an adequate explanation of its mechanism. With AI, though, we may be moving into an unprecedented phase: one in which an explanation remains forever beyond the reach of our knowledge. A bit like the fact that today no human can beat a machine at chess.
This need not involve a sudden break, given that there are different degrees of understanding. I may have no idea how a model found a proof while understanding the proof perfectly; or I may understand individual steps without grasping the overall design. A formalized proof can be verified by a system such as Lean, which checks compliance with explicit logical rules. It is still necessary, however, to verify that the formalized statement corresponds to the original question, examine the assumptions and axioms used, and check the reliability of the procedure. Trust, in this case, rests on a controllable procedure, without requiring us to reconstruct the entire path of invention. Outside mathematics, moreover, verification takes on a different form.
Perhaps what troubles us is the idea that machines themselves do not understand what they are saying, and that their fruits rest on the fragile foundations of mystery. But the recurring question of whether a machine “really understands” also deserves closer scrutiny. If understanding means being able to use knowledge in new circumstances and derive correct consequences from it, then we can test whether a system succeeds. We will have to test it under conditions different from those in which it was trained and observe where it fails. A correct answer obtained by chance is obviously not enough, but a stable ability to transfer what has been learned provides sufficient grounds to speak of functional understanding.
It is legitimate to doubt that a machine “feels” that it has understood something and experiences the particular inner sense of clarity that accompanies certain thoughts. We know that feeling well, just as we know the feeling of having understood something only to discover the next day that we had misunderstood it. However convincing it may be, the feeling of understanding is no guarantee that understanding has actually occurred. In a famous thought experiment by Wittgenstein, everyone possesses a box that no one else can open and calls whatever is inside it a “beetle.” The experiment asks us how much the public meaning of a word can depend on something accessible only to the individual. Extending this analogy to artificial understanding helps reveal the difficulty of making a private experience the decisive criterion by which we recognize a capacity in others. If “really understanding” refers to something that nobody can verify in anyone else, the debate becomes undecidable.
Although the idea of knowledge we may never be able to grasp frightens us, we are already aware of the limits of what we call understanding. Physics allows extraordinary predictions and reduces different phenomena to common principles, while leaving open the question of why the fundamental laws are precisely these and not others. Our condition has always been one of slight competence resting on an unfathomable mystery: slight compared with everything that escapes us, but powerful enough to cure a disease or send a probe into space. Or wipe ourselves out, although here I am talking about atomic bombs, not generative AI, despite the rather improper parallel many people like to draw. Machines can extend our competence without exhausting the mystery: in some domains they are becoming smarter than we are, while remaining subject to limitations and errors that we will have to learn to understand.
All of this inevitably has psychological implications, which often cloud our judgment. Last night, for the first time thankfully, I dreamed that Sam Altman invited me to visit OpenAI. He was very courteous, yet I felt profoundly uncomfortable. Everything around me depended on decisions over which I had no power; I was a guest and could be dismissed at any moment. It seems to me a fairly accurate image of our situation. For many of the tasks we perform, these tools are becoming essential. They allow us to achieve results that would take us much longer on our own, or that we would struggle to reach at all. This is nothing new: the same thing happened with writing, books and the internet. Access to a remote service, however, depends on the conditions established by whoever runs it. And when decision-making is concentrated in a handful of companies, that dependence makes us extremely vulnerable. The possibility that scientific research might depend on a few large private companies, foreign ones in our case, is a danger that should not be underestimated.
The Navier–Stokes case makes this asymmetry especially clear. According to OpenAI, the group that produced the proposal used roughly ten thousand simultaneous agents, based on an internal model more powerful than the one available to the public. The argument was found after approximately 88 hours, followed by another 17 hours for formalization and verification. These are figures provided by the company, but they illustrate very clearly the distance between paying for a subscription and having access to the resources of those who operate the infrastructure.
This asymmetry also reappears, in a way, when the possibility of our extinction is discussed. On September 12, 2026, Anthropic CEO Dario Amodei called for slowing the advancement of models, proposing external evaluators with continuous access to laboratories and coordination between companies and governments. Musk and Altman said they agreed, and that charming little group alone should make us suspicious. In the same intervention, We Must Pace the Frontier, he also argues that any slowdown should preserve the advantage of the United States and its allies, including through restrictions on technologies made available to China. And here he seems intent on pushing public opinion toward an extremely dangerous war on open source, which, by a remarkable coincidence, happens to be these industrialists’ greatest enemy.
The apocalyptic risk is, in practice, based on pure speculation involving entirely unknown variables. In other words, the argument rests on guesswork by people who have a clear financial interest in creating doubt, through the familiar mechanism whereby a product capable of wiping us out must surely be worth whatever it costs. It also works exceptionally well in the media, like all millenarian fears, and journalists, strangled by decades of clickbait, simply cannot resist such an appetizing, and poisoned, morsel.
The sheer magnitude of a possible harm is not, by itself, enough to establish priorities either. If merely imagining a possible catastrophe were sufficient, the possibility of an asteroid collision might justify unlimited resources. An AI that solves all the world’s problems is, moreover, literally just as probable, yet somehow that strikes us as absurd.
Meanwhile, we possess extremely solid knowledge about far more concrete catastrophic dangers. The IPCC describes increasingly severe climate risks and potentially irreversible changes as warming increases. Wars and genocides confront us with suffering taking place now, in response to which we tolerate delays we would declare intolerable when discussing a future superintelligence. We can address several problems at once; we should, however, make explicit the reasons behind the way we distribute attention and resources, including the reasons that make it politically convenient to neglect some of these problems.
A philosophical justification for this imbalance can be found in strong longtermism. In The Case for Strong Longtermism, Hilary Greaves and William MacAskill argue that, in the most important decisions about the use of resources, effects on the distant future should carry predominant weight. The potentially enormous number of beings who might exist makes it possible to assign immense value even to small improvements in their prospects. In this calculation, the needs of those who are alive today risk counting for very little. This seems to me to be a political problem as well, because those who currently possess great wealth can find in this framework a justification for retaining control of it, presenting that control as indispensable to some future benefit. The uncertainty of predictions also leaves enormous discretion to whoever decides which benefits should count.
For this reason, when someone invokes the salvation of humanity, it is worth asking how much they stand to gain from it. A financial interest is not enough to disprove an alarm; what must still be examined is the power that alarm grants them. The promise to protect us can give political authority to the same companies producing the danger and, at the same time, advertise the power of their tools. If safety were translated into requirements accessible only to the largest operators, or into indiscriminate restrictions on open models, it could strengthen precisely the dependence we ought to reduce. Every proposal should also be evaluated in terms of its effects on our ability to develop alternatives and subject these systems to external scrutiny.
Writing helps bring the stakes into focus. It was a cognitive technology of incalculable power, destructive power included; without the knowledge accumulated and transmitted through writing, we would never have built nuclear weapons. That would have been a terrible reason to restrict literacy to a small elite. The analogy has obvious limits, because a model can perform activities that a text, by itself, cannot; but it does show how inadequate it is to infer from the capacity to cause harm that a tool should therefore be concentrated in the hands of a few. The very power of AI makes broad access more important, accompanied by education and by the possibility of participating in decisions about its development.
Francesco D’Isa