Almost every news story about a study contains some version of the phrase "the result was statistically significant." It sounds definitive. In practice, it is one of the most misunderstood concepts in research — and the American Statistical Association has formally cautioned both scientists and journalists about reading too much into it.
This article explains, in plain language, what statistical significance actually tells you, what it does not tell you, and what to look for instead when you want to know whether a finding matters.
What "p < 0.05" actually means
A p-value is the probability of observing the data you saw — or something more extreme — if the null hypothesis (typically: "there is no effect") were true. A p-value of 0.04 means that if there really were no effect, you would see a result this large or larger about 4% of the time, purely by chance.
That's it. A p-value does not tell you the probability that the hypothesis is true, it does not tell you that the effect is large, and it does not tell you the result will replicate. The 0.05 threshold is a convention, not a law of nature.
What it does not mean
Three misreadings are especially common, and the ASA's 2016 statement on p-values calls each one out by name. First, "p > 0.05 means the treatment doesn't work" — false; it means the study did not produce strong enough evidence to reject the null, which can also happen because the sample was too small. Second, "a smaller p-value means a bigger or more important effect" — false; p-values depend on both effect size and sample size. Third, "statistical significance equals practical importance" — false; a tiny difference in a huge sample can be statistically significant and clinically meaningless.
- p-value is not the probability the hypothesis is true
- Statistical significance is not the same as effect size
- "Not significant" is not the same as "no effect"
- The 0.05 cutoff is a convention, not a meaningful boundary
Effect size and confidence intervals matter more
An effect size answers the question "how big is the difference?" — for example, the average blood-pressure reduction in millimeters of mercury, or the relative risk of an outcome. A 95% confidence interval gives a plausible range for that effect given the data. A narrow interval centered on a meaningful value is much more useful than a single p-value.
Major statistics and medical journals — including the American Statistical Association and the New England Journal of Medicine — now recommend that papers report effect sizes with confidence intervals first, with p-values as supplementary information rather than as the headline result.
Multiple comparisons and the "garden of forking paths"
If you test 20 different outcomes at p < 0.05, on average you'll find one "significant" result purely by chance, even if nothing is happening. This is why pre-registered studies — where the researcher commits in advance to which hypothesis they're testing — produce more reliable results than exploratory analyses where the headline finding is chosen after the fact.
When you read a study, look for a single pre-specified primary outcome. Secondary outcomes are useful for generating hypotheses but should be interpreted more cautiously.
Replication is the real test
A single statistically significant result is a starting point. A result that has been independently replicated by other research groups, ideally with pre-registered designs, is much closer to a finding you can rely on. Major fields including psychology and biomedicine have grappled publicly with the "replication crisis" — the discovery that a meaningful share of published, significant findings do not hold up when re-tested.
When a news article reports a single study, ask: has this been replicated? What does the broader literature say? A systematic review or meta-analysis carries far more weight than any single result.
