The thesis of this note is that using AI to generate (“one-shot”) math/physics research papers will be a short-lived trend, for these three reasons:

  1. Abundance: Many complex technical math results that were valuable on account of being in short supply (rather than intrinsically important) will now be abundant, thereby losing some component of their assigned value.
  2. Conditional information: The amount of information or “surprise” in an LLM-generated paper – conditioned on access to a similar model – is tiny.
  3. Inefficiency: At a macro scale, science being composed of mostly AI-generated research disseminated in a traditional paper format is both out-of-equilibrium and comically inefficient, and it’s just hard to imagine a system working that way for long.

The basic logic of my claim is that abundance and low conditional information will apply negative pressure against humans sharing AI-generated papers among themselves, whereas the inefficiency of using AI to referee AI-generated papers will lead AI scientists to use an entirely new medium.

These mechanisms suggest that:

  1. People will choose not to disseminate one-shotted papers...
  2. People will still write papers for curation and understanding...
  3. The scientific paper in its traditional format will fall out of fashion.

My empirical prediction (and my hope for the future of scientific culture) is that we should not expect to see a large, persistent increase in the number of mostly AI-generated papers being posted and disseminated by professional researchers in hard sciences.

Two disclaimers:

  • This is not about AI for doing science or math. Rather, this post is a critique of the entire idea of using AI to write entire journal articles and arXiv preprints, regardless of whether the core ideas/results were attributable to an AI.
  • I am mainly referring to domains and practices of competent, professional scientists who do math and physics. It might be that the vast majority of near-term future science is AI-generated by volume, though at the hands of a small number of people who are not traditionally considered part of the academic community.

Here I expand on each of the three takes:

Abundance

Let’s consider a resource that has (i) limited availability and (ii) some economic/social value that is independent of its limited availability. Determining the availability of a resource is usually easy, albeit only at equilibrium or in a short enough time window. But beyond some basic cases, defining what is valuable is a tortuous process full of social context and messy heuristics, and results in a lot of disagreements along the way.1 So a simple heuristic is to just think of the two as intertwined, by deciding that limited availability increases value. This heuristic works well because it’s often true that rarity is correlated with value, but clearly this is an inefficiency in a system of assigning value to a particular resource.

Until a few months ago, complex mathematical proofs had limited availability. Today, this is no longer true. So, this abundance is erasing whatever component of intrinsic value we have assigned to theory results on account of their rarity alone. In this process, some of the extrinsic value – such as training human scientists and signaling competency to peers and future employers – is also being lost. We furthermore ought to expect review committees and funders to shift away from publication count or h-index (or whatever) as a metric for scientific impact, creativity, or prolificity. These effects put negative pressure on the incentives to post AI-generated science, which will counteract the existing social norm to publish as much as you can.2

Unfortunately – and as a side effect of widespread genAI – these developments will also tend to disincentivize human-authored publications…

Low conditional information

Conditional information answers, loosely, “how much information does this paper contain, conditioned on me having access to an LLM that could generate such a paper given the correct prompt(s)3?” As a human reading an AI-generated paper, I am constantly aware that I could have generated the entire contents of this paper – and in a manner more suitable for my personal tastes – if I had just used the right prompt and maybe a few follow-ups. This attitude extends to whatever AI harness or other tech stack you might be using: Conditioned on me having access to that harness, the ultimate contribution of the AI-generated paper is still just (1) the input prompt plus (2) access to the generative algorithm. Knowing that the (conditional) information content of any paper is maybe about a paragraph, it becomes hard to justify reading or disseminating a traditional research paper.

Ominously, one way this prediction can go wrong is with papers generated using tools that are not publicly accessible. The implication is that a scientist may be more willing to read an AI-generated paper if it was created by either a proprietary or unaffordable LLM. As an aside, if academia of the future accepts AI-generated papers as valid currency, then OpenAI/Anthropic will be holding the keys to the institutional research pipeline, right? There is no reason that they should continue providing researchers with access to a model for only $200/mo. if that tool enables $100K/year in grant money. In other words, a scientific culture that tolerates AI-generated papers is a culture that telescopes its own inelastic demand for AI services, and can expect to be squeezed out of existence much quicker than otherwise.

Inefficiency

All else equal, we might expect the number of AI-generated papers to surge beyond what humans in relevant subfields are willing to read. Though there were always “too many papers to read” (important or not) and paper mills have always churned out many low-quality papers, the situation now is qualitatively different than before. The issue of there being “too many papers” arose in a setting where automated paper generation did not exist, i.e., members of the community, funding agencies, and venues did not coevolve alongside tools that can generate a complete scientific paper in three minutes. So we should see the academic paper-writing system today as being far outside of its equilibrium.

One way to analyze how the eventual equilibrium of scientific publishing might look is to work backwards from the finished product:

Will human researchers be the ones reading the papers (as referees, committee members, editors, or simply readers), or not?

  • If yes, then the community will have imposed some structure on academic publishing so that conference and publication venues are not flooded with AI-generated papers, by removing incentives and applying negative pressure against such papers. AI-generated papers will not be a widespread feature in such a system.
  • If no, then the most natural response to a rapid influx of AI-generated science is to involve AI in the reading/reviewing/refereeing process. But a system that uses AI referees to gatekeep AI science is itself still an out-of-equilibrium system! The format of a journal article is a historical contingency due to humans being the only Earth lifeforms doing science. A better equilibrium will be established that does not look like journal articles, perhaps repositories of proofs or compressed abstracts which humans can query for further knowledge. In any such case, we will have no need for humans to disseminate AI-generated papers among themselves.4

After all, AI reading AI-generated papers is absurdly inefficient! First, given on a human prompt, LLMs reason about the world using sequences of real vectors converted into tokens of human language. Then, LLMs encode these reasoning tokens into several dozen pages of LaTeX5. Finally, another LLM decodes that conversation back into some smaller digestible unit (LLMs share a lot of context, after all) and produces a bunch of reasoning tokens which it finally packages as a human-readable judgement? The concern with inefficiency here is not about burning trees or draining lakes or whatever, but with the Kafkaesque nature of humans using AI to read all the papers AI generated for other humans. I both expect and hope that the immune system of academia will reject such a ridiculous state of affairs6.


  1. If you fail at determining both “valuable” and “limited availability”, you get hilarious failure modes like NFTs. 

  2. As an aside, papers posted with commentary like “here is a one-shotted AI paper that is similar to the kind of research I usually put out” should really be recognized as a self-own. That an otherwise well-intentioned researcher starts publishing AI-generated papers is essentially a flag for what AI is best at automating and what kind of research we will not need humans for much longer. This should be another negative pressure on AI publishing. 

  3. Specifically, I mean conditional algorithmic (“Kolmogorov”) complexity of your LLM output, conditioned on a complete state of the LLM conversation, i.e. how short a program I would need to reproduce your AI-generated paper, if I also knew what LLM you used and everything in its context window. This is upper-bounded by the total length of your prompting. I would say that my desire is to be surprised by whatever paper I am reading, but this would be incorrect because expected conditional surprise is actually conditional Shannon entropy, ha ha ha. 

  4. An exception might be if we humans decide on some canonical human-readable version of an AI discovery, but the benefits of that seem small compared to just asking AI to explain itself. 

  5. I mean “encoding” as a map from short messages to long messages which are transmissible with less loss, where we account for transmission errors like “lack of context by the reader” which is a correctable error given a thorough enough “background” section. 

  6. Though the existence of for-profit online publishers certainly raises doubts about scientific communities’ ability to resist broken and inefficient systems.