A professor at an Ivy League university moved her final exam from take-home to proctored, in-person. And scores dropped 50 percent.
The story feels complete. Students had outsourced their thinking to language models, and removing the machine revealed their actual knowledge. The conclusion appears airtight.
Any shift from unsupervised to proctored conditions produces a score decline. This is documented across decades of educational research and has nothing to do with whether students cheated before. Test anxiety spikes, cognitive load increases under time pressure and observation, and the exam environment itself drives the differential.
The assumption running beneath this story is that a score drop is transparent evidence of dishonesty. We are terrible at isolating cause from context. The professor observed a symptom and named a culprit without ruling out simpler explanations.
We are terrible at isolating cause from context.
This matters because the diagnosis determines the cure. If students learned to depend on AI scaffolding, the fix is pedagogical. Redesign engagement, build competency gradually, match assessment to learning conditions. If the problem is simply that in-person exams differ from take-home work, then proctoring proves nothing except that proctoring is stressful. Before you act on a conviction, ask what evidence would prove you wrong, and whether you've looked for it.
A 50 percent drop is real data. It's not evidence.