Paul Graham has published a brief but pointed meditation on the language we use to describe what large language models do, arguing that calling an LLM's output "thinking" is not merely a rhetorical flourish but a historical pattern that has played out before in science and technology.
The Copernican Analogy
Graham's central comparison is striking. He draws a parallel between the current use of "think" to describe LLM behavior and the experience of 16th-century astronomers who used the heliocentric model for calculations even though they did not fully believe it was true. The math was simpler, more elegant, and more predictive than the alternatives, so the model was adopted for its utility long before it was accepted as a description of reality.
His argument is that we will follow the same trajectory with AI. We will start by calling what models do "thinking" because the word is the closest approximation available. Over time, as the capabilities become undeniable and the usage becomes routine, the language will shift to acknowledge what is actually happening rather than what we wish were happening.
A Question of Definition
Graham frames the entire dispute as a question of how "thought" is defined. If thought is defined as something that occurs exclusively in biological brains, then AI systems cannot think, full stop. But if thought is defined behaviorally, by what a system does rather than by what it is made of, the boundary becomes far harder to locate.
This distinction has been argued in philosophy for decades, but Graham's contribution is the observation that the definition people settle on will be shaped less by philosophical rigor than by practical necessity. When a system produces output that is functionally indistinguishable from thought, the pressure to call it thought becomes enormous, regardless of what any particular definition says.
Precedent in the Language of Technology
Graham points to precedent in how anthropomorphic language has been absorbed into technical vocabulary without controversy. The words "memory" and "recognize" were once used exclusively to describe human cognitive acts. They are now routinely applied to computers and software with no expectation that anyone believes a hard drive is remembering the way a person does, or that a face-detection algorithm is recognizing a face the way a human does.
The precedent is instructive but not entirely reassuring. Graham notes that we accepted these terms because the analogy was harmless and the underlying mechanism was well understood. Whether the same will be true for thought is an open question, and Graham seems to suggest it might not be. His trailing sentence, left incomplete on the platform, gestures toward the possibility that there might be a line after all, though he does not draw it in the excerpted text.
Why the Language Matters
Beyond the philosophical nicety, Graham raises a practical concern about how descriptions shape expectations. When a product is described as thinking, users and regulators form expectations about its reliability, its reasoning process, and its accountability that differ markedly from what they would expect of a pattern-matching system. Planets, Graham notes, do not care how mathematics describes them. But people do care how AI capabilities are described, because those descriptions influence how much trust is placed in the output, how much responsibility is assigned to the developer, and how much control is granted to the system.
The analogy to Copernicus also carries an implicit warning. The heliocentric model was not just a different way of describing the sky; it eventually overturned the geocentric worldview entirely. If LLMs are genuinely doing something that warrants the word "thinking," then calling them anything less may delay the moment when society adjusts its assumptions about what machines can and cannot do, and what that means for the people who depend on them.