Science

Understanding Statistical Significance in Research

What "p < 0.05" actually means, what it doesn't mean, and how to read a result without getting fooled by a single number.

By NewsClair Editorial TeamScience 3 min read 776 wordsPublished April 19, 2026

Researched and written with AI assistance, reviewed by the NewsClair editorial team.

A scientist reviewing a results chart with confidence intervals on a printed paper.
A scientist reviewing a results chart with confidence intervals on a printed paper.

Photo: homegets.com / Flickr, CC BY 2.0

Published .

Contents(5 sections)
  1. 1. What "p < 0.05" actually means
  2. 2. What it does not mean
  3. 3. Effect size and confidence intervals matter more
  4. 4. Multiple comparisons and the "garden of forking paths"
  5. 5. Replication is the real test

Almost every news story about a study contains some version of the phrase "the result was statistically significant." It sounds definitive. In practice, it is one of the most misunderstood concepts in research — and the American Statistical Association has formally cautioned both scientists and journalists about reading too much into it.

This article explains, in plain language, what statistical significance actually tells you, what it does not tell you, and what to look for instead when you want to know whether a finding matters.

What "p < 0.05" actually means

A p-value is the probability of observing the data you saw — or something more extreme — if the null hypothesis (typically: "there is no effect") were true. A p-value of 0.04 means that if there really were no effect, you would see a result this large or larger about 4% of the time, purely by chance.

That's it. A p-value does not tell you the probability that the hypothesis is true, it does not tell you that the effect is large, and it does not tell you the result will replicate. The 0.05 threshold is a convention, not a law of nature.

What it does not mean

Three misreadings are especially common, and the ASA's 2016 statement on p-values calls each one out by name. First, "p > 0.05 means the treatment doesn't work" — false; it means the study did not produce strong enough evidence to reject the null, which can also happen because the sample was too small. Second, "a smaller p-value means a bigger or more important effect" — false; p-values depend on both effect size and sample size. Third, "statistical significance equals practical importance" — false; a tiny difference in a huge sample can be statistically significant and clinically meaningless.

  • p-value is not the probability the hypothesis is true
  • Statistical significance is not the same as effect size
  • "Not significant" is not the same as "no effect"
  • The 0.05 cutoff is a convention, not a meaningful boundary

Effect size and confidence intervals matter more

An effect size answers the question "how big is the difference?" — for example, the average blood-pressure reduction in millimeters of mercury, or the relative risk of an outcome. A 95% confidence interval gives a plausible range for that effect given the data. A narrow interval centered on a meaningful value is much more useful than a single p-value.

Major statistics and medical journals — including the American Statistical Association and the New England Journal of Medicine — now recommend that papers report effect sizes with confidence intervals first, with p-values as supplementary information rather than as the headline result.

Multiple comparisons and the "garden of forking paths"

If you test 20 different outcomes at p < 0.05, on average you'll find one "significant" result purely by chance, even if nothing is happening. This is why pre-registered studies — where the researcher commits in advance to which hypothesis they're testing — produce more reliable results than exploratory analyses where the headline finding is chosen after the fact.

When you read a study, look for a single pre-specified primary outcome. Secondary outcomes are useful for generating hypotheses but should be interpreted more cautiously.

Replication is the real test

A single statistically significant result is a starting point. A result that has been independently replicated by other research groups, ideally with pre-registered designs, is much closer to a finding you can rely on. Major fields including psychology and biomedicine have grappled publicly with the "replication crisis" — the discovery that a meaningful share of published, significant findings do not hold up when re-tested.

When a news article reports a single study, ask: has this been replicated? What does the broader literature say? A systematic review or meta-analysis carries far more weight than any single result.

Look forWhy it matters
Pre-specified primary outcomeDistinguishes confirmation from exploration
Effect size with unitsTells you how big the effect actually is
95% confidence intervalShows the plausible range, not just a point estimate
Sample sizeTiny studies overstate effects; huge studies make trivial effects "significant"
Replication and meta-analysesSingle results are starting points, not conclusions
Reading a study result well

Frequently asked questions

Is p < 0.05 a magic threshold?
No. It is a convention adopted in the 1920s and has no theoretical basis. Some fields use 0.01 or stricter; in 2018, a group of statisticians proposed a 0.005 standard for novel claims.
What's a "statistically significant" result with no real effect?
It happens by definition about 5% of the time when the null is true. With enough tests, false positives are guaranteed.
Why do news headlines about studies often contradict each other?
Single studies are noisy. Press releases often emphasize one significant result without the context of the wider literature. Meta-analyses are the more reliable source.
What is Bayesian analysis, briefly?
An alternative framework that incorporates prior evidence and produces probability statements about hypotheses, rather than p-values. It's increasingly common in some fields but does not eliminate the need to interpret results carefully.

How we researched this

This article was researched and drafted with AI assistance using primary sources — regulator publications, official guidance, peer-reviewed research, and reporting from established outlets — and reviewed by the NewsClair editorial team before publishing. Where data shifts quickly, we date each claim. This article does not provide individualized medical, legal, or financial advice.

Sources

  1. ASA Statement on p-Values: Context, Process, and Purpose American Statistical Association
  2. Moving to a World Beyond "p < 0.05" The American Statistician
  3. Estimating the reproducibility of psychological science Science (Open Science Collaboration)
  4. Cochrane Handbook for Systematic Reviews Cochrane

Related reading

Found this useful? Share it with a friend.

This article is informational and not a substitute for professional advice. NewsClair does not provide medical, legal, or financial services.