A computer science researcher has published a sharp critique of how large language models are invading academic peer review, arguing that the technology's escalatory nature is making the problem worse rather than better, and that using LLMs to filter LLM-generated submissions would merely ignite a new arms race.

The "Fast Food for the Brain" Problem

The essay, published on Stephen's personal research blog, opens with a bleak assessment of the current moment. The author describes the present atmosphere around LLMs as a mania comparable to "fast food for the brain," arguing that the technology is especially corrosive to education and, by extension, to the skills and abilities of society as a whole. The hope, expressed without much optimism, rests on hypothetical future regulation or what the author calls "the mother of all stock market crashes" as the more likely mechanism for pulling the technology back from unsustainable levels.

That backdrop sets the stage for the article's central concern: the academic peer review process is being inundated with LLM-generated submissions. The author acknowledges that many researchers are working on guidelines to manage the influx, but dismisses these efforts as emergency flood defences rather than sustainable solutions. The distinction matters. Flood defences manage symptoms; land management addresses the conditions that cause flooding in the first place.

Why "Lean In" Is the Wrong Answer

One proposed response is to accept the deluge and use LLMs to review LLM-generated submissions in return. The author rejects this outright. The value proposition of large language models, he argues, is faster but lower-quality knowledge work at a cheaper-than-human price. As arbiters of quality, they are not fit for purpose. The question of what quality means becomes central: working code and thoughtful criticism are fundamentally different kinds of output, and LLMs are far more reliable at producing the former than the latter.

The essay draws on a distinction between shallow and deep work. LLMs, as the author describes them, are "syntax-extruding machines" capable of producing shallow work at an acceptable quality level, apparently and somewhat reliably. The question is not whether they can do this task, but whether making them the default tool for it is wise. Removing any brake on shallow work, he argues, produces unintended escalatory effects. Humans need to learn from shallow tasks as well as deep ones, and delegating all shallow work to machines removes that learning pathway.

The Escalation Trap

The author invokes the philosopher Ivan Illich, who described the most dangerous pattern of modern technology as "attempting to solve a crisis by escalation." When things become unmanageably large or complex, over-powered machines promise to tame the complexity, but in reality they simply generate more of it while making society more dependent on those machines. LLMs exhibit this pattern directly: they generate more text, more submissions, more content, which then requires more LLM-based processing to manage.

One tentative proposal the author considers is using LLMs specifically to detect LLM-generated submissions and reject them. He acknowledges this feels uncomfortable, not primarily because of the risk of false positives, since humans remain in the loop, but because it is itself escalatory. It creates an arms race between increasingly undetectable LLMs and ever-more-sophisticated detection systems. However, he notes that this struggle is at least one level removed from the deeper problem. The cat-and-mouse game operates on a logically shallow criterion with a knowable answer, namely whether a text was generated by an LLM, rather than on the deep and genuinely difficult question of whether a paper is worth accepting or a proposal is worth funding.

Where LLMs Actually Work, and Where They Do Not

The author concedes that there are apparent success stories for LLMs in fields like mathematics, but argues these only work because the LLM is combined with agents that have strong external validity checks, such as formal mechanisations in Lean. When an error causes a hard bounce onto another path, a fast-iterating agent can eventually stumble on a useful result through brute persistence. That feedback loop does not exist in peer review. A rejected paper does not bounce back with a corrected proof; it simply fails.

The conclusion is unapologetically blunt. The author does not want academic work reviewed by what he calls a "dullard," whether human or machine, and argues that nobody else should want that either. The distinction between deep and shallow work, while admittedly vague, remains the critical lens through which the usefulness of LLMs in academia should be evaluated. Reviewing papers or proposals requires exactly the kind of depth that the author believes LLMs cannot reliably provide, and building systems that treat them as if they can is not a solution to the problem of volume but a confirmation of the problem itself.