One possible future

A lot has been said about how mathematicians might adapt as AI surpasses us in this or that capability: if AI is better at solving problems, then we will pick which problems are worth pursuing, and if AI can do that too, then we will write expository material, and if AI can do that too, then we will at least judge and organise the results into textbooks. In general, we are looking for ways to remain useful in this production process. I would like to describe my favourite, perhaps fantastical, answer to what we might do if we run out of ways to be useful altogether. I do not know whether that will happen, but the extremal case is often instructive.

So let us imagine that we have reached the point where humans provide no added value in any stage of production. In particular, as is the case in chess, even a human using AI will not outperform an AI agent running autonomously. (We can allow that a dozen humans still decide where to point a CERN-like supercomputer.) Suppose that this level of AI is cheaply available to everyone through open-weight models, and moreover, suppose that the pace of mathematical progress at the frontier is so rapid that human understanding cannot keep up. What is left?

We can keep learning maths. Even with a personalised textbook and an AI-tutor hologram, I doubt that one can internalise a mathematical idea without active effort. Although parts of the AI-digested mathematical universe will forever be beyond reach, we can explore the small neighbourhood consisting of the maths that a human could possibly grasp within a lifetime.

I have seen this point raised, that we might spend our days learning maths, by some who then promptly dismiss it, for example by asking whether you would really like to referee papers all day. I am rather picturing what it was like to learn a course as an undergraduate, when the material felt fresh and canonical. I wanted to learn maths for its own sake. I remember how it felt to start on the first page of a freshly printed set of lecture notes or a textbook. It was only later that I learned to view this activity with suspicion. As I began to prioritise making original contributions, reading a paper in its entirety became decadent, and I took up the habit of raiding papers for parts that might be useful to my own work.

As thrilling as it is to raid, I wonder how many people have been alienated by this massive premium we place on novelty. I knew someone, when we were both in graduate school, who told me that although he loved learning maths, he was going to take a job in industry because he was not particularly suited to the current system that only rewards progress at the frontier. He told me that what he enjoyed most was to "take a complicated system and understand how each of its components works together".

That said, I would quite like to continue solving problems. I am not using "learning" as a synonym for "reading" in opposition to problem solving. It seems like an axiom of our education system that wrestling with problems is an important part of the learning process. However, there is an inversion of priorities: as undergraduates, we solved problems to better understand the course material, whereas in research, we often view learning as a means towards solving open problems.

Learning for its own sake is clearly rewarding, but is it not a bit dull? The prospect of scoring an open problem inspires an intense drive in many mathematicians. Could the drive to learn be as visceral? Suppose we all got serious about grasping the biggest ideas. We could embark on joint expeditions to understand a hard proof, or we could meander, asking ourselves questions and following wherever they lead. We could share what we have learned at conferences. Those were largely social anyway. Professors could be hired on the basis of what material they have mastered. This would make at least as much sense as the current system of hiring professors, ostensibly to teach, on the basis of their ability to solve open problems. Every four years, we could present medals to those who have reached the highest peaks of understanding. It is easy to be in awe of someone who has seen far.

One issue that comes to my mind, as a mathematician, is that whether someone has "understood" something is not well defined. It is not clear that a good definition exists, let alone one that we could all agree on. This seems like a hard philosophical problem, but luckily, we do not need to resolve it before building something operational. Schools and universities have not minded. By and large, if I were to have a long conversation with someone who claimed to have understood some argument that I myself have mastered, I could sniff out whether they were bluffing. For most of their history, Oxford and Cambridge examined by viva. But who is qualified to conduct a viva when the examinee may be the first person ever to have understood the subject?

An AI agent might be. Imagine you have spent months learning a subject and have reached what you consider a stopping point. So you organise your thoughts and produce a set of lecture notes that covers what you have learned, as you see it, and submit these notes to the AI agent. You can submit AI-generated notes if you prefer, but it is presumably easier to be examined on your own notes than on someone else's. The agent reads these notes and for a few hours each day over the following week has a friendly oral discussion with you, probing to check that you have internalised the content. At the end of the week, the agent provides a pass/fail recommendation. If you pass, then you receive a certificate that you have understood that material, like journal acceptance. If you fail, then nothing happens: the result is private and you can retry anytime. In practice, you already know before taking the exam whether you will pass because, although it is thorough, the agent is not trying to trick you.

Even in a world where such AI is as commonplace as LaTeX, we still get to choose how to program it. Any one of us can decide to write an instruction manual for how to run a viva: these results can be taken for granted, this kind of mistake is forgivable, here are some human-graded example transcripts, etc. Like journals, these manuals can vary by subject and taste, some more reputable than others. We can audit transcripts ourselves to check for blatant manipulation, like a candidate convincing the examiner that the fate of the world depends on a pass. Human examiners can also conduct vivas whenever possible, reserving AI for the frontier. Ultimately, these vivas are simply guardrails to keep our learning honest. Their reliability is not the same sort of life-or-death matter as the correctness of a proof.

That leaves the question of money, where I am not so confident. Those who pay us because our results might one day be useful will stop. That applies to research grants, which is a lot. But it is not clear that universities, our traditional patrons, would be all that bothered. Humanities professors are not required to be useful in that sense. Until two hundred years ago, professors were not expected to do original research. Even afterwards, G. H. Hardy famously prided himself on the uselessness of his work, and Trinity did not seem to mind. Today, universities compete in their own prestige game, with scoreboards like the Shanghai rankings. A cynic might argue that universities hire esteemed mathematicians in order to win at that game and tell parents that their child will be taught by a Fields medallist. Teaching is partly a social activity, and it survived Khan Academy. If we start esteeming each other for being learned, would deans even notice?

I have sketched my favourite version of what the future might look like. I would love to hear other ideas. This one has the advantage of being grassroots. No collective action or decision from the top is necessary for any of us to start tinkering. You can be on either side of the current debate about using AI to solve research problems. We can write instruction manuals and try to bluff our way through AI exams. We can run vivas on each other and organise seminars where the speaker presents a result that is not their own and is prepared to be grilled by the audience. We can put on our CVs what we have learned and how it was certified, the way people list the courses they have taken. Hiring committees will have to care what candidates know, unless they are prepared to hire prompt engineers as professors. Above all, it would be good to start experimenting sooner rather than later because it is easier to repurpose our institutions than to rebuild them. I am happy to try. The process of research maths can feel like we are producing reading material that we are too busy to read, for a fictional reader. If AI pushes us there, I would not mind taking up the role of that reader.

To appear on Proofs and Prompts