Follow your curiosity

What discovery has been shared with you?

Start with one fact. Explore it, go deeper, then follow whichever branch catches your imagination.

Choose subjects for a surprise

Exploring any topic

Begin your discovery

Your next discovery is one click away.

Choose one or more subjects above, or leave Any Topic selected and let curiosity decide.

Psychology

P-Hacking and the Reproducibility Crisis

Quick fact

A p-value of 0.05 is often taken as evidence of a real effect, but with p-hacking, researchers can inflate the chance of finding a ‘significant’ result to nearly 100% even when no true effect exists. In one famous demonstration, researchers p-hacked a dataset until they obtained a significant correlation between listening to the song ‘When I’m Sixty-Four’ by the Beatles and the participant’s age—a spurious result with no real causal link.

Why this is interesting

A paper with a statistically significant result is often considered a 'finding.' But what if that significance is as fragile as a house of cards? Why can a supposedly rigorous result fail to show up when another lab tries it again?

Read the full explanation

Understanding P-Hacking and the Reproducibility Crisis

Imagine you are a researcher testing whether a new drug improves memory. You measure memory before and after, and you see a small improvement, but the p-value is 0.08—not below the magic 0.05. Your grant depends on publishing. So you try a few tweaks: you add a few more participants, you control for age, or you analyze only a subgroup. Finally, the p-value drops to 0.04. You publish. That process of tweaking until success is p-hacking. The problem is that when you test many variations, by chance alone one will appear 'significant' even if the drug does nothing. Each test is a gamble; the more you gamble, the more likely you hit jackpot—but the jackpot is a false positive.

A deeper explanation

At its core, p-hacking exploits the logic of statistical hypothesis testing. A p-value below 0.05 is supposed to mean 'if there were truly no effect, we would observe this result by chance less than 5% of the time.' But this logic assumes that only one test is performed. When a researcher runs many tests—testing different predictors, excluding outliers, or checking both genders separately—the probability of at least one false positive skyrockets. For example, if you run 20 independent tests where no effect exists, the chance of getting at least one p<0.05 is about 64%. This is called the multiple comparisons problem. P-hacking is essentially a creative way to exploit this: the researcher often does not report all the tests they ran, only the one that 'worked.' This selective reporting, combined with publication bias (nonsignificant results are rarely published), means the literature is filled with false positives. Over time, these false positives accumulate, and when other labs try to replicate the findings, they fail, fueling the reproducibility crisis. In addition, p-hacking is often a symptom of deeper incentive structures: researchers are rewarded for publishing significant findings, and there is little reward for null results. Thus, even well-intentioned scientists can fall into bad habits, and the system as a whole produces unreliable knowledge.

Keep FACTREE close

Internet access is required. Updates arrive when you reopen or reload the app. You may need to sign in again in the installed app.