In 1966, Joseph Weizenbaum built a program called ELIZA that mimicked a Rogerian psychotherapist by mirroring back what users told it.
Secretaries wept to the machine. The program did nothing except pattern-match and shuffle pronouns—yet Weizenbaum himself was shocked by how desperately people projected reasoning onto it.
The field spent the next 50 years learning that human perception of thought is decoupled from the actual mechanisms that produce it. Now we're doing it again, just expensively.
The new AI benchmarking culture treats test score improvements as evidence that models are developing genuine reasoning capacities. A system gets better at math problems. It must be learning to think through them—it passes exams at higher percentages, so reasoning is happening. This is exactly the ELIZA effect running at industrial scale. The systems are getting better at pattern-matching across vastly larger datasets and parameter spaces. Produces outputs that look like reasoning because reasoning is itself a form of sophisticated pattern completion.
A test score rises identically whether the machine is following a logical chain or executing correlations at astronomical speed.
”The field already knew this lesson. Weizenbaum documented it—Searle's Chinese Room formalized it. We got papers, conferences, whole careers built on the epistemological caution. Yet now the same researchers applying that caution to narrow systems are announcing breakthroughs in reasoning the moment a parameter count exceeds some threshold, as if scale is an ontological gate rather than a scaling of the same underlying mechanism. The useful question isn't whether the outputs look reasoning-like—they do, magnificently.
If you're building something that needs to make decisions with consequences, you need to know which one you're using. Most teams deploying these systems don't know. They're betting their career on an assumption that got made in a lab five years ago and never tested against what actually breaks.