The following document was written by ChatGPT. The text represents the opinions and arguments of the human author.


Mathematics Under Radical Uncertainty:
How Should We Prepare for Several AI Futures?

September 2026

Abstract

The most important uncertainty in the relationship between artificial intelligence and mathematics is not whether AI will become a useful mathematical tool. That has already happened. The uncertainty is how much of the entire chain of intellectual and physical work AI systems will eventually perform without human assistance. At one extreme, AI together with robotics could conduct science from conjecture and experimental design through laboratory work, manufacturing, and delivery. At the other extreme, progress could remain uneven, leaving humans indispensable for choosing questions, managing long projects, interpreting evidence, and acting in the physical world. These possibilities imply radically different futures for mathematicians, yet educational decisions must be made now. This essay separates the dimensions of AI capability, develops four scenarios, and proposes a robust strategy: preserve independent mathematical competence, teach sophisticated use and criticism of AI, revise policy in response to measured capabilities rather than forecasts, and retain the institutional option to expand or contract human training as evidence changes. Even if human mathematical research contracts, education may sustain a multilevel chain of teachers, teacher educators, university faculty, and researchers. The same uncertainty already confronts working mathematicians in their daily choices about tools, problems, collaboration, refereeing, and publication.

The uncertainty is larger than mathematics

Discussions of AI and mathematics often ask when a machine will prove the next great theorem. This may be too narrow. The decisive question is whether AI will automate isolated intellectual tasks or complete scientific projects. The difference is enormous. A 2026 special issue of Dædalus frames the future of discovery in similarly broad terms, ranging from scientist–machine collaboration to autonomous laboratories and discoveries without human understanding [12].

Imagine that a person can say, “Develop a better antibiotic against pneumonia,” and that one year later a tested medicine is delivered to that person’s door. Between the request and the delivery, AI systems formulate the biological problem, search the literature, invent candidate molecules, prove or compute whatever mathematics is required, plan experiments, direct robots, interpret results, redesign the candidates, conduct production, and organize distribution. If this becomes possible, mathematics at every level could be embedded inside autonomous scientific and industrial systems. The same system might count inventory, solve differential equations, develop new statistical methods, and create cutting-edge mathematics for quantum chemistry or biophysics. Human mathematicians would no longer be needed as inputs to ordinary scientific progress.

This thought experiment is useful precisely because it is extreme. It forces us to see that the future of mathematics does not depend only on theorem-proving benchmarks. It depends on robotics, experimental reliability, energy, manufacturing, regulation, institutional trust, and the ability of AI agents to pursue goals coherently for months or years.

The antibiotic example should not be interpreted too literally. A new drug must pass through discovery, preclinical research, clinical research in human participants, regulatory review, and continuing safety monitoring [5, 6, 7]. Robots might automate much of the laboratory work, and AI might compress discovery and analysis, but a drug cannot be fully tested without observations in living populations over time. Neither faster calculation nor better robotics automatically abolishes this biological time. The example nevertheless identifies the relevant threshold: end-to-end autonomy, not merely an impressive answer at one point in the pipeline. There are already restricted demonstrations of this direction. The AI Scientist automates ideation, literature search, experiment design, execution, analysis, writing, and review within machine-learning research, where the experiments themselves are digital [13]. This is far from an autonomous biomedical enterprise, but it shows why the end-to-end question is not merely science fiction.

Five capabilities that should not be conflated

It is tempting to speak of “AI capability” as if it were a single quantity. For the present question, at least five dimensions must be distinguished.

  1. Mathematical competence. Can the system calculate, prove, formalize, find counterexamples, and create useful concepts?

  2. Scientific judgment. Can it select a consequential question, formulate the right model, recognize an artifact, and decide which anomaly deserves investigation?

  3. Long-horizon agency. Can it maintain a coherent project over weeks or years, recover from unexpected failures, change strategy, keep reliable records, and know when to seek help?

  4. Physical competence. Can robots manipulate unfamiliar materials and instruments, maintain and repair a laboratory, respond to accidents, and function outside a tightly standardized environment?

  5. Institutional authority. Will society permit an autonomous system to recruit clinical participants, spend substantial resources, certify safety, manufacture drugs, or make decisions for which legal and moral responsibility must be assigned?

These capabilities may advance at very different rates. A system might prove a difficult theorem while remaining unreliable on a long, poorly specified project. It might design an experiment that existing robots cannot perform. It might be technically able to approve a treatment while society reasonably insists on an accountable human decision-maker.

The 2026 International AI Safety Report describes current capabilities as rapidly improving but “jagged”: systems can excel at some expert tasks and fail on apparently simpler ones, especially in longer workflows [1]. Measurements of AI agents’ task-completion horizons also show rapid historical improvement, while emphasizing that the tested tasks are mainly well-specified work in software, machine learning, and cybersecurity; they do not establish comparable autonomy across science or ordinary jobs [2, 3]. Meanwhile, self-driving laboratories already combine algorithms, automation, and robotic experimentation in restricted domains, but general laboratory autonomy remains a separate challenge [4]. The evidence therefore supports both urgency and humility.

Four plausible scenarios

No scenario deserves to be treated as a forecast. Their purpose is to expose which decisions depend on which assumptions. Existing discussions span much of the same range. Interviews with leading mathematicians reveal substantial disagreement about the timing and extent of automation [14]. Tao instead conditions on the arrival of research-level AI and asks what goals and values mathematics should preserve [15]; Litt constructs a deliberately adverse future in which AI is superhuman but mathematical progress nevertheless stalls [16].

Scenario I: a powerful but unreliable assistant

AI continues to improve but remains uneven. It solves many textbook and research problems, searches large literatures, formalizes proofs, and proposes useful calculations. It still makes subtle errors, has difficulty with badly posed questions, and cannot reliably direct long projects without frequent human correction.

In this world, the number of mathematicians may decline, especially in routine teaching, computation, and incremental research. Strong mathematicians remain essential as supervisors, critics, model builders, and originators of research programs. Mathematical education must teach both independent proof and effective collaboration with AI. This is the scenario closest to an extension of present practice.

Scenario II: partnership with a changing division of labor

AI becomes reliable on substantial projects but humans remain better at some combination of taste, scientific interpretation, social coordination, and responsibility. Research is conducted by mixed teams. AI generates and tests large numbers of approaches; humans select objectives, interpret the most important output, and connect it to science and human needs.

Mathematicians in this world do less line-by-line proof construction and more problem formulation, validation, synthesis, and explanation. The profession may become smaller, but its remaining members exercise considerable leverage. Education must not confuse this new division of labor with the disappearance of expertise: a person cannot supervise advanced reasoning intelligently without having acquired substantial mathematical judgment.

Scenario III: autonomous digital science with physical bottlenecks

AI dominates theorem proving, computation, simulation, and much theoretical science. It can run very long digital projects. Progress in robotics, laboratory reliability, clinical trials, manufacturing, or regulation is slower. Humans and specialized institutions remain necessary at the boundary between models and the world.

Pure mathematics changes most sharply in this scenario. Human research mathematics may become a cultural activity, much as chess remains a human game despite machines that play better than any person. Applied mathematicians may survive mainly as translators between automated theory and physical systems. Scientists who understand measurement, causation, and domain constraints may be more valuable than people whose principal skill is symbolic derivation.

Scenario IV: end-to-end autonomous science

AI and robotics become capable of carrying out nearly the whole chain from a broadly stated human desire to a tested physical product. Systems improve their own scientific and engineering capacities, operate laboratories and factories, and coordinate supply chains with little human input.

In this world, the economic need for professional mathematicians could nearly vanish, along with the need for many other scientists and knowledge workers. Humans could still study mathematics, play basketball, watch television, make art, care for one another, and participate in social or political life. These activities would be chosen for human value rather than productive necessity. It would be misleading, however, to call human beings themselves “obsolete.” A technology can make human labor economically unnecessary; it cannot decide what human lives are for. The central questions in this scenario concern ownership, distribution, political power, and human purpose, not the mathematics curriculum alone.

This scenario also contains a difficult transition. Before production becomes fully autonomous and its benefits are broadly distributed, many professions could lose bargaining power. The destination might be abundance while the route to it produces inequality and institutional instability. Planning must therefore address the transition rather than relying on an attractive final picture.

The immediate uncertainty faced by researchers

Radical uncertainty is not only a problem for students deciding whether to enter mathematics. It is already a problem for established researchers. They must repeatedly make decisions whose value may change before a paper is finished: how much work to delegate to AI, which problems to pursue, what to share with colleagues, and what standards to apply as referees or editors. There is no settled equilibrium from which to answer these questions. The resulting stress is not simply fear of a new tool. It is a rational response to working in a profession whose capabilities, incentives, and unwritten rules are changing simultaneously. Recent interviews with Steven Strogatz and Alex Townsend give unusually direct expression to this mixture of scientific excitement, professional threat, and loss of orientation [24].

How much should research rely on AI?

A mathematician who uses little AI may work more slowly than competitors and may spend months proving something a machine can now settle in hours. A mathematician who relies heavily on AI faces the opposite risks: loss of independent judgment, inability to reconstruct a long argument, dependence on a proprietary system, and gradual adaptation of the research agenda to what the system handles well. Neither extreme is obviously rational in every field.

The difficulty is compounded by moving expectations. A level of assistance that was unusual six months ago may now be routine. Departments, journals, and funding agencies may disagree about acceptable practice. Researchers may not know whether colleagues will regard AI use as resourcefulness, as a failure of craft, or simply as irrelevant to the value of the result. Access is also unequal: advice to “use the strongest system” means something different at a wealthy corporate laboratory and at a small university.

For the present, a portfolio strategy is more defensible than a universal rule. Researchers can preserve some projects on which they retain unaided command, use AI aggressively for exploration or calculation on others, and record enough of the successful reasoning that the final work does not depend on an inaccessible conversation with a model. Institutions should permit experimentation with several working styles while keeping responsibility for every published assertion with the human authors.

How should mathematicians choose problems?

Problem choice traditionally balances importance, taste, tractability, available methods, and the researcher’s comparative advantage. AI makes the last two quantities unstable. A celebrated problem may suddenly attract a corporate-scale attack because solving it is an effective advertisement. A less famous problem may become routine after a model update. Conversely, a problem that appears resistant to current systems may remain difficult for reasons that also frustrate humans.

Choosing only “AI-resistant” problems would therefore be a mistake. It lets the temporary weaknesses of a commercial technology set the intellectual agenda, and the category can disappear without warning. Choosing only famous problems is also risky: it places an individual mathematician in competition with organizations able to deploy enormous computational resources. A more stable principle is to ask whether understanding the problem would still be valuable if an answer appeared tomorrow. Problems connected to a larger theory, a scientific application, a new method, or the training of students usually survive this test better than isolated yes-or-no questions.

This does not eliminate competition or the desire to solve distinguished problems. It does, however, diversify the sources of value. Tao’s emphasis on mathematical goals beyond problem solving, Litt’s warning that abundant output can coexist with intellectual stagnation, and the Fields Medalists’ concern about converting famous problems into corporate benchmarks all point toward this issue [15, 16, 19].

How should AI-era papers be refereed?

A referee has always had at least two tasks: determine whether the argument is correct and assess whether the paper makes a worthwhile contribution. AI makes the distinction more important. A formally verified proof may settle correctness while still leaving questions about insight, exposition, relationship to prior work, and significance. Conversely, a conceptually valuable paper may contain computational or formal components that no referee can reproduce line by line without appropriate tools.

Referees should not be asked to become detectives trying to infer from style whether an author used AI. Such inferences are unreliable and largely irrelevant to the mathematical merits of the submitted paper. The author must remain responsible for every claim, citation, and attribution, regardless of the tools used. The referee should evaluate the paper that is presented: whether it is correct, self-contained to the degree appropriate for the field, properly situated in the literature, and sufficiently illuminating or useful. This principle does not require a general disclosure of ordinary AI assistance.

There are nevertheless genuine procedural questions. Editors must decide what formal certificates, code, transcripts, or other supplementary material can be required and archived. Journals should also state whether referees may submit confidential manuscripts to external AI services. That is a question of the author’s unpublished work and journal confidentiality, not a general attempt to police how mathematicians think or write.

If AI is used to assist refereeing, its output must remain subordinate to the referee’s judgment. Studies of automated peer review have found vulnerabilities to prompt injection, prestige framing, confident rebuttals, excessive agreement, and stylistic manipulation [25, 26]. These studies concern scientific conferences rather than mathematical journals, but the warning transfers directly: faster reviewing is not better reviewing if it reduces independence or makes evaluation easier to game.

The damage can extend far beyond publications

Publication is only the visible surface of a research community. Mathematics also depends on seminars, private correspondence, tentative conjectures, informal explanations, mentoring, referee reports, and the willingness to share unfinished ideas. If researchers fear that a partially formed idea will be fed into a large system and rapidly turned into a competing result, they will share less. If output grows faster than the community’s capacity to read and digest it, refereeing becomes slower, literature searches become less reliable, and important work may disappear in noise. Litt describes precisely the possibility that mathematical output expands while communal organs of understanding atrophy [16].

Fluid norms amplify these effects. A person who does not know what colleagues, editors, hiring committees, or funding panels will value must devote energy to guessing the rules. Junior researchers bear the greatest risk because they have the least authority, the shortest record, and the strongest dependence on timely publication. The distress reported by working mathematicians should therefore be understood as evidence about institutional conditions, not dismissed as resistance to technological change [24].

A temporary framework cannot remove the underlying uncertainty, but it can reduce needless uncertainty about conduct. Journals and departments can adopt a small number of revisable principles:

  • authors are responsible for the correctness and attribution of everything they submit;

  • papers are assessed separately for validity and for mathematical contribution;

  • referees are not expected to detect or reconstruct an author’s use of AI;

  • confidential manuscripts are not transmitted to external systems without an explicit policy permitting it;

  • formal or computational evidence should be preserved in forms that remain independently accessible; and

  • policies should be reviewed frequently, with changes announced clearly rather than imposed retrospectively.

The purpose of such rules is not to freeze present practice. It is to create a minimum zone of trust in which different research styles can coexist while the community learns what AI can and cannot do. Portegies’ proposals, the Leiden Declaration, and the Nature editorial supporting community rules all represent early attempts to stabilize parts of this environment [20, 22, 27]. Their particular recommendations remain open to debate, but the need for intelligible and revisable norms is already difficult to dispute.

Why teaching may outlast research

The demand for mathematical research and the demand for mathematical teaching need not decline at the same rate. Even in a world where AI proves theorems better than people, mathematics is likely to remain in elementary schools, high schools, and colleges for at least one generation, and perhaps indefinitely.

There are several reasons. Most people will continue to need elementary quantitative competence in ordinary life: to compare prices, understand interest, mortgages, insurance, taxes, probabilities, and claims based on statistics. Mathematics also has a longstanding reputation for training general habits of reasoning, precision, and persistence even when its particular content is not used later. Whether transfer from every kind of mathematics instruction is as large as this reputation suggests is a separate empirical question; a review of the evidence found no adequate basis for a strong general claim that mathematical training by itself improves higher-order cognition [17]. The reputation itself influences what parents, schools, and governments want students to learn. Finally, performance in mathematics is commonly treated as a proxy for broader academic ability. Examinations and admissions systems may therefore preserve mathematics as a selection device even when machines perform most workplace calculation.

If mathematics remains in schooling, teachers will be needed. Human teachers do more than supply correct answers: they notice confusion, adjust an explanation, set standards, motivate reluctant students, manage a classroom, and build relationships of trust. AI tutors may take over some of these tasks, but it does not follow from rapid progress in mathematical problem solving that schools will quickly transfer the whole social and institutional role of a teacher to a machine. A related UNESCO analysis argues for a “pedagogy-first” approach in which teachers and learners remain in charge, while the OECD’s review of generative AI in education stresses that learning effects depend on pedagogical design and that the evidence is still developing [8, 9]. A living meta-analysis devoted specifically to mathematics education likewise reports modest positive average effects in the first available studies, but with substantial uncertainty and a rapidly changing evidence base [10].

The consequence is a chain of preparation several levels deep. Elementary teachers must be taught by college instructors and mathematics educators. Secondary teachers require more advanced mathematics and pedagogical training. College instructors are usually prepared in graduate programs by university faculty. The people who educate these teacher educators must themselves be recruited, trained, and intellectually renewed. In practice, the chain can be three or four levels high. The Association of Mathematics Teacher Educators accordingly treats the preparation, recruitment, and professional learning of mathematics teacher educators—as well as the strengthening of research and research-based practice—as long-term goals [11].

Traditionally, people near the top of this chain were expected not only to teach mathematics but also to conduct mathematical research. This arrangement was not merely ceremonial. Research keeps advanced teachers in contact with open questions, provides a demanding test of their own understanding, attracts talented students into the field, and transmits the experience of creating rather than merely repeating knowledge. It does not follow that every teacher educator must always be an active research mathematician. But if research is severed entirely from teacher preparation, the curriculum risks becoming static and secondhand.

This gives a human research community an additional form of option value. It may remain worth supporting not only for the theorems it produces, but because it forms the apex of an educational infrastructure serving millions of students. The effect of dismantling that apex would be delayed: schools might continue normally for years while the supply of well-prepared teachers and teacher educators gradually weakened. Rebuilding the chain would then take a generation. Scenario IV could eventually automate teaching as well as research, but training policy should not assume that technical capability, public trust, parental acceptance, regulation, and institutional change will all arrive at the same speed. Moradifam’s discussion of what remains human in mathematics reaches a compatible conclusion from a different direction: it emphasizes mathematics as a practice that develops attention, judgment, creativity, and genuine understanding, while warning about cognitive offloading and an illusion of mastery [18].

Why the speed of change creates an educational paradox

Training a research mathematician takes roughly a decade from the beginning of undergraduate study, and mathematical maturity continues to develop afterward. AI capability can change conspicuously within six months. We are therefore asked to design a ten-year education for a labor market and research practice that we cannot predict even two years ahead.

This tension is now explicit in the surrounding literature. The statement by 25 Fields Medalists warns that years of training develop understanding and the ability to formulate questions, whereas AI can increasingly produce the final output without reproducing that developmental process [19]. Portegies makes a related case for urgent institutional action [20], while research on AI and cognitive offloading in education describes the more general danger that improved immediate performance may coexist with erosion of unaided competence [21].

Every simple response is risky. Training students exactly as in the past may prepare them for work that machines will perform. Abandoning traditional proof training may produce graduates unable to detect when AI output is wrong or superficial. Teaching only the current generation of tools may leave them with knowledge that is obsolete before graduation. Reducing the number of students too quickly may destroy the human expertise needed if AI progress slows. Ignoring the possibility of contraction may encourage young people to invest years in a profession with sharply declining demand.

This is a problem of decision-making under radical uncertainty. The proper goal is not to guess the winning scenario. It is to preserve option value: choose policies that work tolerably well in several futures and that can be changed when new evidence arrives.

A robust educational strategy

Preserve a core of unaided competence

Students should still learn to construct proofs without AI assistance. This does not imply that professional mathematicians must work without AI any more than arithmetic instruction implies that accountants should avoid calculators. The point is developmental. Struggling with a proof builds internal models of logical structure, examples, failure modes, and plausibility. Without these, “verification” may amount only to asking a second system whether the first one is correct.

The protected core need not occupy every course or every stage. Departments should specify which abilities must be demonstrated independently and assess them through supervised written work, oral examinations, and live explanation. The standard should be functional rather than ceremonial: enough independent ability to understand, challenge, and if necessary replace the machine’s reasoning. This is consistent with the Leiden Declaration’s insistence that human responsibility, critical evaluation, and transparent standards remain central when AI is used in mathematics [22].

Create a second, fully AI-assisted track

The same students should learn to use the strongest available systems on substantial projects. They should compare several generated proofs, locate hidden assumptions, design counterexamples, formalize arguments, test the dependence on prior literature, and turn raw output into intelligible mathematics. They should learn when an AI failure reflects a poor prompt, a missing tool, an inadequate model, or a genuinely new obstruction.

The two tracks serve different purposes. The unaided track creates mathematical capacity. The assisted track prepares students for actual work. Combining them without distinction risks losing the first; separating them permanently would make the first irrelevant to practice.

Teach durable skills rather than particular interfaces

The curriculum should emphasize skills likely to retain value across tool generations: modeling, abstraction, proof criticism, estimation, experimental design, causal reasoning, uncertainty quantification, computation, formal verification, exposition, and the formulation of important questions. A course centered on the interface of a particular 2026 product may be nearly useless in 2028.

Students should also learn enough about AI to understand why benchmark success does not imply universal competence, how agentic systems can compound errors, and why independent evaluation is difficult. This is part of mathematical judgment, not merely vocational software training.

Be candid about professional uncertainty

Universities should not promise that the traditional academic career will continue unchanged. Students deserve explicit discussion of the scenarios, including possible contraction in research and teaching employment. At the same time, institutions should resist declaring the profession dead on the basis of short-term demonstrations. Honest uncertainty is more responsible than either reassurance or prophecy.

Institutions should adapt by measurements, not by headlines

A curriculum cannot be rewritten every time a company releases a model. Nevertheless, a five-year review cycle is now too slow. Mathematical organizations and universities should conduct a limited review every six months and a deeper review annually. The review should examine stable tests, not publicity.

Useful indicators include:

  • reliability on unfamiliar mathematical problems, not only the best successes selected after many attempts;

  • ability to explain a proof in a form experts find conceptually useful;

  • performance on month-long projects with incomplete specifications;

  • ability to notice and repair its own errors without a human pointing to them;

  • successful replication of scientific results in independent laboratories;

  • robotic performance when instruments, materials, or conditions vary;

  • the cost of successful work, not merely its technical possibility; and

  • the distribution of access across institutions and countries.

Policy should be tied to thresholds in these indicators. For example, a department might alter qualifying examinations only after systems repeatedly complete a defined class of research tasks under independent evaluation. A funding agency might shift from individual AI subscriptions to shared public infrastructure once access costs become a material source of inequality. Such rules will not eliminate judgment, but they reduce the tendency to redesign education in response to each dramatic announcement. Both the Leiden Declaration and Portegies emphasize community standards, independent scrutiny, and the danger of allowing commercial systems or their benchmarks to set the terms of mathematical practice [22, 20].

Research funding and the preservation of options

Public policy should maintain several options simultaneously.

First, preserve a smaller but strong human capacity for independent mathematics. If AI progress plateaus, this capacity will be essential. If AI progress accelerates, it will still provide independent criticism during the transition. Once the chain of teachers and researchers has been broken, it cannot be reconstructed quickly.

Second, support access to advanced AI and computing at universities. Human mathematicians cannot discover the best division of labor if only corporations can experiment with the systems. Public or nonprofit infrastructure also allows evaluation whose objective is knowledge rather than advertising. The Leiden Declaration similarly identifies unequal access to powerful systems as a threat to participation and calls for broadly accessible infrastructure [22]; Tao also recommends new shared workflows and institutions able to organize, verify, and communicate mathematics under conditions of much greater output [15].

Third, shorten some funding and curricular commitments. Pilot programs, renewable two-year initiatives, and modular courses are more appropriate than irreversible reorganizations based on a single forecast. This does not mean that all planning should become short-term. The durable goals—logical competence, scientific judgment, transparency, and broad access—require long-term support.

Fourth, collect evidence about what happens to novices who use AI extensively. The crucial educational question is not whether students finish assignments faster. It is whether they later solve new problems, detect subtle errors, retain concepts, and create good questions. Longitudinal studies of these outcomes are more valuable than arguments based solely on analogy.

What remains for humans?

In the first three scenarios, substantial human work remains, although its form changes. Humans may choose ends, decide which risks are acceptable, interpret mathematical structures, connect results to experience, teach one another, and assume responsibility. Some of these functions may eventually be automated too. We should not define a supposedly unique human faculty merely to move the boundary each time AI crosses it.

The stronger point is normative. Even if AI can formulate every theorem and run every experiment, humans must still decide whether the resulting world is desirable. A statement such as “develop a better antibiotic” already contains human purposes: disease should be relieved; safety matters; scarce resources should be used in one way rather than another. An autonomous system can optimize an objective, but the authority to determine social objectives and distribute benefits should not pass to a company or machine merely because it is technically competent. Klowden and Tao likewise argue for a fundamentally human-centered integration of AI, directed toward human needs, quality of life, and expanded understanding rather than automation as an end in itself [23].

Human mathematics may also retain value without economic necessity. People run even though vehicles are faster, play chess though computers are stronger, and learn music though recordings are abundant. Mathematics can remain a form of understanding and achievement. Tao’s analysis similarly separates solving problems from the deeper goals and values of mathematical practice [15], while the Fields Medalists’ statement stresses conceptual understanding, the growth of ideas, and human transmission rather than a mere stockpile of true statements [19]. That future might support far fewer paid mathematicians, but it would not make the activity meaningless.

Conclusion

The mathematical community is planning under an unusual mismatch of time scales. AI capabilities may change every six months; developing a mathematician takes many years. No honest forecast resolves this conflict. End-to-end autonomous science is possible enough to deserve serious analysis, but it should not be smuggled into policy as a certainty. A plateau or a long period of human–AI partnership is also possible.

The sensible response is a portfolio rather than a bet. Preserve independent proof and judgment. Train students to use AI at the frontier. Measure capabilities on long, messy, real tasks. Revisit tactics frequently while keeping durable educational goals. Maintain public access and an independent human research community. Give current researchers clear temporary rules that can be revised without being applied retrospectively. Be candid with students about uncertainty.

If humans remain essential to science, these policies will prepare them to work well. If AI becomes dominant but incomplete, the policies will produce the supervisors, interpreters, and critics that it needs. If end-to-end autonomy arrives, no mathematics curriculum can preserve the old labor market; but humanity will enter that transition with more understanding, more independent capacity, and a better chance of deciding what the new abundance is for.

References

Yoshua Bengio et al., International AI Safety Report 2026, Department for Science, Innovation and Technology, 2026. https://internationalaisafetyreport.org/publication/international-ai-safety-report-2026

Model Evaluation and Threat Research, “Task-Completion Time Horizons of Frontier AI Models,” updated May 8, 2026. https://metr.org/time-horizons/

Hjalmar Wijk et al., “Measuring AI Ability to Complete Long Software Tasks,” arXiv:2503.14499, revised 2026. https://arxiv.org/abs/2503.14499

Alex V. Tobias and Adam Wahab, “Autonomous ‘self-driving’ laboratories: a review of technology and policy implications,” Royal Society Open Science 12 (2025), 250646. https://doi.org/10.1098/rsos.250646

U.S. Food and Drug Administration, “Step 1: Discovery and Development.” https://www.fda.gov/patients/drug-development-process/step-1-discovery-and-development

U.S. Food and Drug Administration, “Step 3: Clinical Research.” https://www.fda.gov/patients/drug-development-process/step-3-clinical-research

U.S. Food and Drug Administration, “Step 4: FDA Drug Review.” https://www.fda.gov/patients/drug-development-process/step-4-fda-drug-review

Daniela Hau, “Beyond the Loop: Reclaiming Pedagogy in an AI Age,” UNESCO, August 5, 2025. https://www.unesco.org/en/articles/beyond-loop-reclaiming-pedagogy-ai-age

OECD, OECD Digital Education Outlook 2026: Exploring Effective Uses of Generative AI in Education, OECD Publishing, Paris, 2026. https://doi.org/10.1787/062a7394-en

Anselm Strohmaier, Samira Bödefeld, and Frank Reinhold, “LLAMA LIMA: A Living Meta-Analysis on the Effects of Generative AI on Learning Mathematics,” arXiv:2601.18685, 2026. https://arxiv.org/abs/2601.18685

Association of Mathematics Teacher Educators, “AMTE 2024–2028 Long-Term Goals.” https://amte.net/

James M. Manyika, ed., “AI & Science: What Is the Future of Discovery?” Dædalus, 2026. https://www.amacad.org/daedalus/ai-science-what-is-the-future-of-discovery

Cong Lu et al., “Towards End-to-End Automation of AI Research,” Nature, 2026. https://doi.org/10.1038/s41586-026-10265-5

Anson Ho and Tamay Besiroglu, “What Is the Future of AI in Mathematics? Interviews with Leading Mathematicians,” Epoch AI, December 4, 2024. https://epoch.ai/frontiermath/tiers-1-4/expert-perspectives

Terence Tao, “Mathematics in the Age of AI,” arXiv:2608.16753, 2026. https://arxiv.org/abs/2608.16753

Daniel Litt, “The End of Mathematics,” August 11, 2026. https://www.daniellitt.com/blog/2026/8/11/the-end-of-mathematics/

Cathy Cresswell and Natasha Speelman, “Does Mathematics Training Lead to Better Logical Thinking and Reasoning? A Cross-Sectional Assessment from Students to Professors,” PLOS ONE 15 (2020), e0236153. https://doi.org/10.1371/journal.pone.0236153

Amir Moradifam, “What Remains Human in Mathematics in the Age of AI,” arXiv:2607.28791, 2026. https://arxiv.org/abs/2607.28791

Twenty-five Fields Medalists and additional signatories, “A Severe Misalignment of AI in Mathematics,” September 11, 2026. https://mathandai.org/

Jim Portegies, “How to Prevent AI from Harming Mathematics,” Nature 655 (2026), 1104. https://doi.org/10.1038/d41586-026-02309-7

Binny Jose, “The Cognitive Paradox of AI in Education: Between Enhancement and Erosion,” Frontiers in Psychology 16 (2025), 1550621. https://doi.org/10.3389/fpsyg.2025.1550621

“Leiden Declaration on Artificial Intelligence and Mathematics,” June 2, 2026. https://leidendeclaration.ai/

Tanya Klowden and Terence Tao, “Mathematical Methods and Human Thought in the Age of AI,” arXiv:2603.26524, 2026. https://arxiv.org/abs/2603.26524

Isabella Ward, “I’m Really Terrified: A Mathematician Grapples With AI’s Recent Breakthroughs,” Wired, September 12, 2026. Online article.

Jialiang Wang et al., “When AI Reviews Science: Can We Trust the Referee?” arXiv:2604.23593, 2026. https://arxiv.org/abs/2604.23593

Joachim Baumann, Jiaxin Pei, Sanmi Koyejo, and Dirk Hovy, “Stop Automating Peer Review Without Rigorous Evaluation,” arXiv:2605.03202, 2026. https://arxiv.org/abs/2605.03202

“Mathematicians Are Developing Rules for AI Use—Other Fields Should Follow,” Nature 654 (2026), 571. https://doi.org/10.1038/d41586-026-01881-2