Articles

When the Evidence Runs Out


A road leading into fog and mountains
Fog Mountain. Jake Francis / stocksnap (CC0).

In 2011, a respected social psychologist at Cornell published nine experiments in one of the field's leading journals, using entirely conventional statistical methods, and reached a startling conclusion: a thought could be influenced by an event that had not happened yet. Daryl Bem's paper, "Feeling the Future," reported that practice occurring after a memory test appeared to improve recall on the test taken beforehand, a precognitive effect running backward through time.

What makes the episode genuinely useful, regardless of what one believes about precognition itself, is what happened next. The paper used the same statistical tools and standards that governed ordinary psychology research at the time, which was exactly the problem the controversy exposed: if conventional methods could produce apparently solid evidence for something this implausible, those methods were weaker than the field had assumed. Other researchers attempted to replicate Bem's experiments. Early attempts, run without pre-specified analysis plans, produced a mix of results, some confirming, most not. The decisive test came later, when researchers ran large-scale, multisite replications with every methodological detail, the measurements, the analysis plan, the sample size, locked in writing before a single participant was tested. That kind of replication, called preregistration, removes the ability to quietly adjust what counts as a result after seeing the data. Under those locked conditions, Bem's effect disappeared.

The broader consequence outlasted the specific question of psi. The Bem affair became a widely cited catalyst for what psychology now calls its replication crisis, a field-wide reckoning with how often ordinary, non-paranormal research had been using the same loose methods that had let an implausible effect look statistically real. Preregistration, once a niche practice mostly associated with parapsychology, became standard procedure across much of experimental psychology precisely because of this episode.

What remains for parapsychology today is neither confirmed nor cleanly dismissed. Meta-analyses that pool many smaller studies together still report small, statistically significant departures from chance. Preregistered, high-powered replications of specific, well-known effects largely do not confirm them. The same body of evidence genuinely reads two different ways depending on which methodological standard a reader trusts, and that disagreement has not resolved after more than a century of laboratory research into the question.

A thought, in this account, ran up against the actual edge of what current methods can establish, not a wall built by disbelief, but the honest limit of what repeated, controlled measurement has so far been able to confirm. That edge is worth naming plainly, the same way every other claim in this project has been pressed to its edge and no further: this is what was found, this is what was not, and this is where the evidence, for now, simply runs out.