What to know

  • A p-value is conditional on a hypothesis and the other assumptions of the test.
  • Statistical significance does not measure the size or clinical importance of an effect.
  • Read the effect estimate, confidence interval and analysis plan alongside the p-value.

01

What is a p-value?

A p-value is the probability of a test statistic at least as extreme as the one observed, assuming the statistical model is correct, including the hypothesis being tested. In many health studies, that hypothesis is a null hypothesis of no difference between groups. Other assumptions concern matters such as independence, the data-generating process and how the analysis was selected.

The direction of that reasoning matters: start by assuming the model, then ask about possible data. A p-value does not reverse that calculation to give the probability that the hypothesis is true. A small value indicates tension between the observations and the assumed model; it does not identify which assumption is wrong.

Sources for this section: [1] [2]

02

A worked example: nine heads in ten coin tosses

Consider a mathematical illustration, not a clinical study. Before tossing a coin, we decide to make ten independent tosses and test for an imbalance in either direction. Under the null model, heads and tails each have probability one half on every toss. There are 1,024 equally likely ordered sequences.

If we observe nine heads, results at least as far from a balanced split are zero, one, nine or ten heads. Those outcomes account for 1 + 10 + 10 + 1 = 22 sequences. The two-sided p-value is therefore 22 / 1,024 = 0.021484375, or about 2.15%.

That means this degree of imbalance or a greater one would occur about 2.15% of the time under the fair, independent-toss model with the fixed stopping rule. It does not mean there is a 2.15% probability that the coin is fair. Changing the question to heads only, or deciding when to stop after watching the results, changes the analysis and its interpretation.

Sources for this section: [2]

03

What does p < 0.05 mean?

With a significance threshold, or alpha, set at 0.05, a p-value below that threshold is conventionally called statistically significant. Alpha is chosen as part of the testing procedure; the p-value is calculated from the observed data. For a valid single test, alpha controls the long-run probability of rejecting a true null hypothesis, not the probability that this particular published finding is false.

Common readings and their limits
Reported resultReasonable readingUnsupported conclusion
p = 0.03The result crosses a preselected 0.05 threshold, conditional on the test assumptions.There is a 97% chance the treatment works.
p = 0.06The result does not cross that threshold.The treatment has no effect.
p = 0.001The observed statistic is highly unusual under the assumed model.The effect is large, important or free of bias.

These are interpretation examples, not results from patient trials. There is no scientific cliff between 0.049 and 0.051. A change in a threshold label can occur with very similar data, so the full estimate and its uncertainty deserve more attention than the label alone.

Sources for this section: [2]

04

Statistical significance versus clinical importance

The effect estimate answers how much the groups differed. The confidence interval helps show the precision of that estimate and which effect sizes remain compatible with the data under the model. Clinical importance concerns whether the size and nature of a benefit matter to patients in the context of harms, burden and alternatives.

A large study may detect a difference too small to matter in practice. A smaller or noisier study may leave substantial benefit and harm both plausible, even when its p-value exceeds 0.05. Read the interval’s limits and the units of measurement before deciding what a result means for care.

For matching two-sided tests and interval methods, a 95% confidence interval that excludes the null value generally corresponds to p < 0.05. The null is usually zero for a difference and one for a ratio. This relationship depends on using compatible methods; rounding and different calculations can produce apparent disagreements. Neither a p-value nor a narrow interval removes bias from a flawed study.

Sources for this section: [2] [3]

05

Why multiple testing changes the picture

Searching across many outcomes, subgroups or analysis choices creates more opportunities to find a small p-value. As an illustration, with 20 independent tests whose null hypotheses are all true and whose false-positive probabilities are each exactly 0.05, the probability of at least one false positive is 1 − 0.9520 ≈ 64.2%. Correlated tests or different procedures change that number.

This is why the planned primary outcome, trial registration and handling of multiple comparisons matter. A prespecified analysis and an interesting exploratory finding can both be useful, but they warrant different confidence. Reporting only the smallest p-value hides the search that produced it. Statistical adjustments help with defined testing families; they cannot repair undisclosed selective reporting.

Sources for this section: [2]

06

A checklist for reading a research claim

Before accepting a headline that says a treatment “worked,” use the questions below to reconstruct the actual claim. They turn a single threshold into an assessment of the study question, the size of the result and the uncertainty that remains.

  1. Find the comparison: who was studied, what was compared and which outcome was measured?
  2. Read the estimate: how large was the difference, in units that matter to the reader?
  3. Check the interval: does it include meaningful benefit, meaningful harm or only small differences?
  4. Check the plan: was this outcome and analysis chosen beforehand, and were multiple tests addressed?
  5. Inspect credibility: consider bias, missing data, applicability and agreement with other studies.

A large p-value alone does not demonstrate equivalence between treatments. That requires a suitable design, a justified equivalence margin and an analysis addressing that question. A small p-value likewise does not establish causation: the study design and its assumptions must support a causal interpretation.

The statistical explanations draw on Greenland and colleagues’ guide under CC BY 4.0, with rewritten explanations and original calculated examples. The ASA statement and Cochrane guidance provide additional context; source links are listed below.

Sources for this section: [2]

Sources

  1. Wasserstein and Lazar — The ASA’s Statement on p-Values: Context, Process, and PurposeThe American Statistician · 2016Source accessed: DOI 10.1080/00031305.2016.1154108
  2. Greenland et al. — Statistical tests, P values, confidence intervals, and power: a guide to misinterpretationsEuropean Journal of Epidemiology · 2016Source accessed: DOI 10.1007/s10654-016-0149-3
  3. Cochrane Handbook, Chapter 15: Interpreting results and drawing conclusionsCochraneSource accessed:

Revision history

  1. Initial article prepared with automated assistance and sources checked at 2026-09-05T17:53:18Z (UTC). Includes a calculated coin-toss example and a multiple-testing illustration, not patient data. No independent clinical review is recorded. Date displays and this explanation use UTC; the recorded instants are unchanged.