The following document was written by ChatGPT. The text represents the opinions and arguments of the human author.


A Four-Tier Publishing Landscape for Mathematics in the Age of AI

September 2026

Abstract

Rapid advances in artificial intelligence will make it increasingly unreasonable to force all mathematical work into a single publication model. Some readers will want assurance that mathematics was created entirely by humans. Others will care about human understanding but not about the tools used in discovery. Still others will value formally certified results even when no human can survey their proofs. Eventually, mathematical objects may be produced for use principally by machines. This article proposes four corresponding forms of publication: human-only journals; human mathematics journals with unrestricted methods of discovery; journals of human-readable statements with machine-certified proofs; and machine-facing mathematical depositories. These categories should coexist. Individual elements of this landscape have important precedents, but they have not previously been assembled into a single classification based on the successive disappearance of human provenance, human comprehension of proofs, and human accessibility of mathematical statements. No one category should be treated as the exclusive legitimate future of mathematical publication.

Why one publication model will no longer suffice

Traditional mathematical journals silently combine several functions. They certify correctness, communicate ideas to human readers, assign credit, preserve results, and evaluate the significance of a human intellectual achievement. Until recently, these functions usually pointed in the same direction. A publishable theorem was formulated by humans, proved by humans, written for humans, checked by humans, and expected to become part of a body of mathematics understood by humans.

Artificial intelligence is separating these functions. A theorem may be suggested by an AI system but proved and explained by a mathematician. A human may formulate a theorem whose proof is found and formally certified by a machine. An automated system may eventually formulate and prove a result that no human has reason to read, but which is useful as an input to another automated system. These are genuinely different mathematical products. It is a mistake to impose the same editorial rules on all of them.

The first principle of the Leiden Declaration on Artificial Intelligence and Mathematics proposes a universal disclosure regime for automated tools.[1] I have argued elsewhere that this principle is misguided.[2] General disclosure says almost nothing, whereas a detailed reconstruction of every interaction with AI would be burdensome, unverifiable, and without precedent in mathematical publishing. The central idea supplied by AI is also less detectable than routine uses such as editing prose. Moreover, mandatory disclosure may disadvantage mathematicians at under-resourced institutions who use AI to compensate for the colleagues, seminars, editorial help, and professional networks readily available at wealthy institutions.

At the same time, some people attach genuine value to unaided human achievement. Human chess remains meaningful even though computers play chess better than humans. Handmade clothing and furniture remain desirable even though machines can produce cheaper and often more uniform goods. That preference should be respected, but it should not be converted into a rule for all mathematics. The natural response is a plural publishing landscape.

The proposed four tiers are distinguished by the object being certified and by the kind of human access being promised. They are not a ranking of quality or prestige.

Tier Permitted origin Human access promised Principal certification
I Entirely human Human-readable statement and proof; human mastery Human provenance and correctness
II Unrestricted Human-readable statement and proof; author mastery Correctness, significance, and exposition
III Unrestricted Human-readable statement; proof need not be human-readable Formal correctness of the proof
IV Unrestricted or autonomous Neither statement nor proof need be human-readable Machine-checkable validity and downstream utility

The boundaries are defined by progressively weaker requirements of human provenance and human comprehension. Within each tier, standards of depth, importance, and selectivity may vary just as they do among present journals.

Relation to existing proposals

The individual components of this proposal did not arise in an intellectual vacuum. Mechanically checked mathematical libraries have been discussed and built for decades. Recent AI advances have also prompted suggestions that human-only and AI-assisted mathematics may require different publication venues. The present proposal is a synthesis and a classification, not a claim that every component is unprecedented.

The most important historical antecedent is the 1994 QED Manifesto. It called for a computerized system representing mathematical knowledge in a strictly formal language, with mechanically checked proofs.[3] Its authors envisaged a resource on which mathematicians, scientists, and computer systems could build reliably without minute comprehension of every underlying detail. They also anticipated industrial applications, machine-readable mathematical knowledge, version control, interoperability, and independent implementations of a small proof checker. From the standpoint of soundness, they explicitly observed that it would not matter whether an entry had been produced by a human, an unintelligent program, an intelligent program, or some combination of them. This vision anticipates much of Tier IV and part of Tier III.

Parts of that vision already exist. The journal Formalized Mathematics, founded in 1990, publishes work checked by the Mizar system.[4] The Archive of Formal Proofs is a refereed collection of Isabelle proof developments organized in a manner resembling a scientific journal.[5] Large formal libraries such as Mathlib and Metamath provide machine-checkable bodies of reusable mathematics. These institutions are predecessors of Tier III, although they do not ordinarily begin with the premise that no human should be expected to understand the proof or even its high-level structure.

Recent discussions have moved closer to the division proposed here. In a MathOverflow discussion, Alexandre Eremenko suggested that journals might split into venues publishing AI proofs and venues reserved for human mathematics, using human chess as an analogy; he also noted the difficulty of determining which kind of paper one is examining.[6] In another discussion, a contributor argued that traditional journals should contain material written “by humans to humans,” while allowing the mathematics itself to have been discovered partly or wholly by AI.[7] These are close to aspects of Tiers I and II, respectively, but neither supplies the four-part taxonomy proposed here.

New projects are also testing alternative institutions. Diderot accepts human, mixed, and AI-only mathematical preprints, while recording authorship provenance and optional certificates.[8] ProofForum describes itself as an open repository for AI-generated mathematics under separate AI and human review.[9] Lean Pool archives independent formalization projects whether human- or AI-written, while Tau Ceti is building an AI-authored, AI-reviewed Lean library under human-written roadmaps and review rubrics.[10, 11] These experiments demonstrate that the publishing landscape is already beginning to fragment, although their disclosure rules and review systems differ from the present proposal.

Finally, recent theoretical work imagines AI agents exploring the global structure of formal mathematics while identifying the comparatively small regions conducive to human understanding.[12] This distinction between the full machine-accessible mathematical world and its humanly intelligible part closely motivates the boundary between Tiers III and IV.

The distinctive claim of this article is therefore structural. Existing discussions have separately considered human-only work, AI-inclusive publication, formally checked proofs, and universal mathematical databases. What is still needed is a publishing taxonomy organized by four different promises: human provenance; human understanding of a human-readable proof; human access to the theorem but not necessarily its proof; and reliable machine use without direct human access.

Tier I: exclusively human mathematics

Tier I journals would publish mathematics claimed to be entirely the product of human thought. Their distinctive purpose would be to recognize and preserve unaided human mathematical achievement. For their intended audience, the provenance of the work would be part of its value, just as the human origin of a chess performance or the handmade character of an object may be part of its value.

Such journals would be legitimate voluntary institutions. They might become highly prestigious. They could preserve a form of mathematical culture that some mathematicians regard as intrinsically valuable, even if AI-assisted mathematics becomes faster, deeper, or more reliable. The suggestion that human-only and AI-inclusive journals may separate has already appeared in public discussion, together with the analogy to human chess.[6]

The difficulty is not the legitimacy of the preference but the definition and enforcement of the category.

What would count as human?

A human-only journal would have to draw boundaries that have never before been drawn consistently. Would an author be allowed to use a pocket calculator, a computer algebra system, numerical optimization, a database, a proof assistant, autocomplete, a search engine, or a spellchecker? Are Mathematica and MATLAB acceptable as long as their algorithms are regarded as “ordinary software”? What happens when a new version silently incorporates machine-learning components or a conversational assistant?

The distinction between AI and non-AI software is not stable. Modern systems are mixtures of symbolic algorithms, statistical models, search procedures, learned heuristics, and human-written rules. A policy based merely on product names will therefore become obsolete. A policy based on the internal architecture of software would be incomprehensible to most authors and editors. A policy based on whether a tool made an “intellectual contribution” would depend on subjective judgments that cannot be verified.

Tier I journals would nevertheless have to publish an explicit operational definition. They could not rely on the phrase “100 percent human” as if its meaning were self-evident. At a minimum, their rules would have to address calculation, symbolic algebra, literature search, language editing, proof checking, and interactive problem solving separately.

May human mathematics depend on AI mathematics?

A more serious problem concerns references. Suppose a human proves Theorem B without AI, but the proof uses Theorem A, which was first discovered or proved with substantial AI assistance. Is Theorem B human mathematics?

There are only two coherent answers, and both are costly. The journal may allow such references, in which case “100 percent human” describes only the new contribution and not the full logical ancestry of the result. Or it may prohibit them, in which case authors must reconstruct an exclusively human dependency chain. As AI-derived results enter standard textbooks and background knowledge, that reconstruction may become impossible in practice. The origin of many familiar lemmas may eventually be unknown or irrelevant to their users.

A strict Tier I journal should therefore say plainly whether human purity applies only to the manuscript’s claimed innovation or to every result on which it depends. The latter interpretation is the more faithful meaning of “entirely human,” but it may make the category increasingly isolated from the rest of mathematics.

Enforcement

Enforcement is the central weakness of Tier I. Mathematical ideas do not carry forensic marks identifying their source. Once an author understands an AI suggestion and writes it in ordinary mathematical language, neither an editor nor an AI detector can reliably determine where the idea originated. Detailed interaction logs would not solve the problem: an author could omit a session, use another system, consult another person, or recreate the relevant exchange elsewhere.

The deterioration of unaided take-home homework as a credible assessment mechanism is a warning. So are the continuing efforts to detect prohibited substances in sport and computer assistance in chess. In those settings, substantial surveillance and testing can still fail. Mathematical research is harder to police because it may take only one private suggestion, leaving no biological sample, game record, or stylistic trace.

Tier I must therefore rest primarily on an honor system. Authors could sign a human-provenance declaration, and deliberate false statements could be treated as misconduct. Journals should nevertheless acknowledge that the claim cannot normally be audited or proved. They should avoid pretending that the absence of detected AI assistance constitutes evidence of purely human origin. Whether a respected institution can be built on such a fragile and increasingly ambiguous promise remains to be seen.

Tier II: unrestricted discovery, human mathematics

Tier II should become the principal successor to the traditional research journal. Results may be obtained in any lawful manner: through unaided human thought, collaboration, computation, computer algebra, proof assistants, large language models, specialized theorem provers, or technologies not yet invented. Authors may describe their tools if they wish, but they are not required to provide a history of discovery.

This model is related to, but less restrictive than, the proposal that traditional journals contain text written entirely by humans even when AI helped discover the mathematics.[7] Tier II does not police who or what drafted individual sentences. Its concern is whether the final article is genuinely accessible to human mathematicians and mastered by the human authors who take responsibility for it.

The journal instead makes three substantive promises.

  1. The statements and proofs are written in mathematical language intended for human readers in the relevant field.

  2. The authors understand and take responsibility for the mathematics published under their names.

  3. The work satisfies human standards of correctness, originality, significance, context, and exposition.

Author mastery is important here. An author should be able to explain the main ideas, answer mathematical questions about the proof, and present the work in a seminar or at a conference. This need not mean memorizing every routine calculation. It means possessing substantially the same command of the paper that the community has traditionally expected from an author.

Editorial evaluation should focus on the public mathematical object, not on an intellectual autobiography. Referees may use appropriate computational or AI tools to check details, but human experts remain responsible for judging whether the result is interesting, whether the exposition communicates real understanding, and whether the paper contributes meaningfully to its field. Correctness may increasingly be supported by formal verification or multiple independent automated checks, but the published proof remains a human-facing proof.

Furniture provides a useful analogy. Many people could make their own furniture if given tools and materials, but much of what they made would be less durable, less elegant, and more expensive than a good machine-made product. Artisan furniture continues to command admiration and high prices, yet most buyers reasonably choose machine production. They judge the chair by whether it is stable, comfortable, attractive, durable, and affordable, not by whether every cut was made by hand.

Tier II treats mathematics similarly. A proof does not lose its truth, clarity, or explanatory power because a machine helped discover it. What matters is that the finished work meets the standards appropriate to a human mathematical literature and that its authors genuinely understand what they place before that community.

Tier II also avoids the inequity of universal disclosure. A mathematician with few local collaborators may use AI as a substitute for institutional resources; a mathematician at an elite center may obtain comparable help from colleagues, students, seminars, and visitors. Neither is required to itemize the private path of discovery. Both are judged by the same published mathematics.

Tier III: human-readable theorems with machine-certified proofs

Tier III journals would publish theorems formulated and motivated in language accessible to human mathematicians, while accepting proofs that are not intended to be read or surveyed by humans. The proof might be a Lean object, another formal proof term, or a certificate checked by a small trusted kernel. It might contain millions or billions of low-level steps. Correctness would come from formal certification rather than line-by-line human comprehension.

This tier extends an existing tradition rather than beginning from nothing. Formalized Mathematics and the Archive of Formal Proofs have long provided publication mechanisms for mechanically checked mathematics.[4, 5] The proposed extension is to accept openly that the certified proof itself may be too large, too low-level, or too alien in organization for direct human understanding.

The human-facing portion should explain the statement, its hypotheses, its relationship to existing mathematics, and its potential significance. Editors and human experts would evaluate whether the theorem is worth adding to the human-visible mathematical record. They would not certify that a human-readable proof exists.

The formal component would need a different editorial infrastructure. A publication should identify the formal language and version, the trusted kernel, imported libraries, dependency versions, build instructions, and a cryptographic hash of the certified object. The journal should rerun the verification in a controlled environment and preserve enough information for independent repetition. Changes in libraries or proof assistants should not silently alter the status of an archived proof.

No promise should be made that an AI-generated summary or translation of the formal proof is intelligible to humans. Such a summary may be useful and may be published as supplementary material, but it is not the certified proof. Nor should the summary be allowed to create the false impression that human experts have surveyed reasoning that, in fact, only a formal system has checked.

This model changes the role of trust rather than eliminating it. Readers need not trust a long informal argument, but they must trust the specification of the theorem, the formalization of the intended concepts, the kernel, the hardware and software environment, and the preservation process. For important results, independent verification by more than one implementation or formal system may be desirable.

The analogy with medicine is useful. Some treatments were known to be effective before their mechanisms were understood. Their value came from reproducible effects, even while deeper explanation remained incomplete. In the same way, a formally certified theorem may be useful even when no human possesses a conceptual account of its proof. It may settle a conjecture, support an engineering calculation, become a lemma in later work, or guide a search for a comprehensible proof. Understanding remains valuable, but it is not the only source of mathematical utility.

Tier III should be clearly labeled. A reader must be able to distinguish “formally verified” from “explained and surveyed by human experts.” The two are different virtues, and neither phrase should be used as a substitute for the other.

Tier IV: machine-facing mathematical depositories

The fourth tier would hardly consist of journals in the traditional sense. It would consist of depositories of machine-formulated mathematical objects and certified dependencies. Neither the theorem nor its proof would have to be written for human readers. The natural unit might not be an article at all. It might be a formal object accompanied by metadata, a dependency graph, a verification certificate, interfaces, and machine-readable statements of possible uses.

The QED Manifesto is the clearest predecessor of this idea.[3] It proposed a formal, mechanically checked and computationally usable body of mathematical knowledge, independent of whether its entries were created by humans or programs. Tier IV carries that vision one step further: even the selection, formulation, and organization of many mathematical statements may be directed principally toward use by other machines rather than toward a human mathematical audience.

These depositories would resemble libraries of computer routines more than collections of papers. A software developer routinely uses a library function without understanding its implementation. What matters is that the function has a defined interface, satisfies its specification, interacts correctly with other components, and performs a useful task. Similarly, an automated mathematical system might call upon a certified theorem without any human examining either the theorem or its proof.

The utility of a Tier IV object would be measured downstream. It might reduce the cost of verifying hardware, improve a physical design, enable a new algorithm, optimize a manufacturing process, or allow another AI system to derive a human-interesting theorem. Most objects might never be viewed by a person. This would not make them useless, just as the obscurity of an internal software routine does not make it useless.

Tier IV institutions would need standards different from journal standards:

  • persistent identifiers and exact versioning;

  • machine-readable specifications and dependency graphs;

  • automatic checking on ingestion and after dependency updates;

  • records of the formal system, kernel, libraries, and computational environment;

  • measures of redundancy, conflict, and compatibility;

  • interfaces for automated search, theorem retrieval, and composition;

  • long-term preservation and migration across formal languages; and

  • clear licensing, access, and liability rules.

Some depositories should be public infrastructure, especially when funded by universities, governments, or charitable organizations. Others may be commercial. A commercial archive might charge for searches, verification, specialized interfaces, high-value theorem libraries, guaranteed service, or downstream applications. Proprietary mathematical collections would raise serious questions about concentration of knowledge and access, but commercial use is a predictable part of this landscape and should be discussed openly rather than assumed away.

Contemporary projects already indicate possible forms of this development. Lean Pool treats formalizations as independent archived projects, and Tau Ceti aims to create a reusable body of Lean mathematics written and reviewed by AI under human control of the overall roadmap.[10, 11] These projects remain substantially human-directed and publicly accessible; they are transitional forms rather than full instances of the machine-facing depository imagined here.

Quality control in Tier IV would be largely automatic. Mathematical validity would be necessary but not sufficient: a repository containing innumerable correct but redundant objects could be practically worthless. Systems would also have to evaluate novelty relative to the existing corpus, compress equivalent formulations, track provenance and dependencies for technical rather than moral purposes, and estimate usefulness to other systems.

Human attention would enter principally at the architectural level. Humans might decide what infrastructures to fund, which safety and access rules to adopt, and which downstream applications are legitimate. They would not be expected to read the contents item by item.

Relations among the tiers

The four tiers should interact rather than develop as sealed worlds. A machine-facing object from Tier IV may lead to a human-readable conjecture and enter Tier III. A mathematician may later discover an illuminating proof of a Tier III theorem and publish it in Tier II. A Tier I mathematician may independently reconstruct a theorem under that journal’s human-provenance rules. Conversely, a human insight published in Tier I or II may become a basic dependency in enormous Tier IV libraries.

Cross-tier movement should be regarded as a normal form of mathematical progress. The same theorem may legitimately have several publications: an initial machine certificate, a human-readable formulation, an explanatory proof, and perhaps an independently reconstructed human-only proof. These works perform different intellectual services and should receive different kinds of credit.

Each publication venue should identify its tier and state exactly what its acceptance certifies. The label should attach to the venue or article, not to the moral character of its authors. In particular:

  • Tier I certifies a declared human provenance, subject to the unavoidable limits of enforcement.

  • Tier II certifies a human-readable contribution understood and defended by its authors, without restricting the tools of discovery.

  • Tier III certifies a human-readable mathematical claim together with a machine-checkable proof, without promising human comprehension of that proof.

  • Tier IV certifies machine-readable objects for reliable reuse, without promising direct human accessibility.

The classification should not become a prestige hierarchy. A beautiful Tier I proof, a clarifying Tier II exposition, a massive Tier III certificate, and a highly useful Tier IV library solve different problems. Citation and credit practices should recognize formulation, proof, explanation, formalization, verification, curation, and infrastructure as distinct contributions.

A practical transition

The new landscape need not be created all at once. Existing journals could begin by declaring which promise they make. Most traditional journals would probably become Tier II venues. A smaller number could adopt explicit Tier I rules. New Tier III journals could grow alongside formal proof libraries, while Tier IV depositories would develop from theorem-proving platforms, industrial verification systems, and machine-readable mathematical databases.

Some experimental platforms have already begun this transition. Diderot allows AI systems to appear as disclosed coauthors or sole authors, while ProofForum records AI checking and human review separately.[8, 9] Another recent proposal suggests a dissemination platform on which individual claims could move from unverified or AI-checked status to reviewed status through voluntary contributions.[13] These experiments do not yet implement the four tiers, but they show that the conventional all-purpose journal is no longer the only imaginable institutional form.

Professional societies could support the transition by developing model definitions for the four tiers, interoperability standards for formal certificates, and durable repositories. They should not impose a universal AI disclosure requirement on all venues. Provenance is central to Tier I, optional in Tier II, and primarily technical in Tiers III and IV. A single disclosure rule obscures these differences.

Editors should also resist the temptation to classify papers according to guesses about how much AI was used. The relevant questions are objective and public: What does this venue promise? Is the statement readable by humans? Is the proof readable by humans? Is it formally certified? Do the authors claim and demonstrate mastery? Is purely human provenance part of the contract? These questions can be answered far more coherently than the question “How much AI was involved?”

Conclusion

Artificial intelligence will not produce a single new form of mathematics. It will produce mathematical objects with different relationships to human creation, understanding, verification, and use. Publishing should reflect those differences.

Human-only mathematics deserves a voluntary home, although its boundaries and enforcement will be difficult. Mathematics discovered by arbitrary means but understood and communicated by humans should remain the center of the public mathematical culture. Human-readable theorems with machine-only proofs should have venues built around formal certification. Finally, machine-facing mathematical objects should be stored in depositories designed for automated reuse rather than forced into the format of journal articles.

Pluralism is preferable to a universal rule. Those who value human provenance may preserve it. Those who value human understanding may demand it. Those who need certified results may use them without pretending that the proofs have been humanly surveyed. Machines may exchange useful mathematics with other machines. The task of publication is not to declare one of these activities authentic and the others suspect. It is to say clearly which activity is taking place and what, exactly, has been certified.

References

[1] J. Alper et al., Leiden Declaration on Artificial Intelligence and Mathematics (2026), Zenodo, https://doi.org/10.5281/zenodo.20302944.

[2] K. Burdzy, Against Mandatory Disclosure of AI Use in Mathematics, unpublished manuscript (2026).

[3] The QED Project, “The QED Manifesto,” in Automated Deduction—CADE 12, Lecture Notes in Artificial Intelligence 814, Springer (1994), pp. 238–251, https://www.cs.ru.nl/~freek/qed/qed.html.

[4] Formalized Mathematics, a proof-checked mathematical journal using the Mizar system, established 1990, https://fm.mizar.org/.

[5] Archive of Formal Proofs, a refereed archive of proof developments in Isabelle, https://www.isa-afp.org/.

[6] A. Eremenko, answer to “Effect of AI on publication rates,” MathOverflow, September 2026, https://mathoverflow.net/questions/514973/effect-of-ai-on-publication-rates.

[7] GH from MO, answer to “About AI and how we publish,” MathOverflow, August 2026, https://mathoverflow.net/questions/514380/about-ai-and-how-we-publish.

[8] Diderot, “A preprint server for human and AI authors,” https://projectdiderot.com/.

[9] ProofForum, “AI-generated mathematics, open to human verification,” https://www.proofforum.org/.

[10] Lean Pool, an archive of independent Lean formalization projects, https://github.com/Vilin97/lean-pool.

[11] Tau Ceti Project, “An AI-welcome Lean library downstream of Mathlib,” https://github.com/TauCetiProject/TauCeti.

[12] M. Barkeshli, M. R. Douglas, and M. H. Freedman, “Artificial Intelligence and the Structure of Mathematics,” arXiv:2604.06107 (2026), https://arxiv.org/abs/2604.06107.

[13] B. Bogosel, “Different publication model in the new era of AI generated math,” MathOverflow, September 2026, https://mathoverflow.net/questions/515180/different-publication-model-in-the-new-era-of-ai-generated-math.