Researchers are now running psychology experiments on chatbots instead of people.
They prompt GPT-4 or Claude with the same scenarios they would give human subjects, measure how the machine responds. Treat the output as evidence about how human minds work. The assumption is clean. If a chatbot exhibits the same bias, the same emotional reaction. The same pattern of reasoning as a person would, then we've learned something true about human psychology.
This is not new logic. In 1981, James Dobson, a psychologist and conservative activist, proposed using polygraph testing to screen for sexual deviancy in high school boys. The appeal was obvious, the apparatus measures physiological response, and physiological response correlates with lying.
The tool had internal validity. It worked as a machine. But the field never paused on the epistemological cliff. Measuring a body's electrical conductance tells you something about a mind's truthfulness. By the time researchers carefully examined what the polygraph actually measured, the machine had already shaped criminal justice, intelligence work. Public policy for decades.
A chatbot generates text through next-token prediction and statistical correlation of language patterns in its training data. That is not how human consciousness works.
”The structural problem repeats here, unchanged in its architecture. A tool produces output that resembles human output — this resemblance becomes evidence. We stop asking whether the resemblance is meaningful or merely surface. We mistake what the apparatus is good at measuring for knowledge about the thing we actually wanted to understand.
A chatbot generates text through next-token prediction and statistical correlation of language patterns in its training data. That is not how human consciousness works. The machine can be indistinguishable from human response without operating on any principle that maps onto human cognition. Alan Turing understood this in 1950 when he proposed the Imitation Game—he was sidestepping the hard problem entirely. He said we should stop asking whether machines think and start asking whether they sound like they think.
The difference now is scale. Polygraphs were specialist tools. Chatbots are everywhere, already shaping how researchers design studies, already normalizing the substitution. The question worth watching is not whether this will go wrong. It is whether we will notice the moment it does.
Read Alan Turing's 1950 'Computing Machinery and Intelligence' paper to understand why he deliberately sidestepped the question of whether machines actually think—a distinction this article shows we've forgotten.