Statistical Significance: What Sigma Means and What It Does Not
Measurement and Evidence 7 min read

Statistical Significance: What Sigma Means and What It Does Not

Significance is a statement about the behaviour of random fluctuations, and nothing else. It is a necessary condition for a discovery claim and nowhere near a sufficient one, which is why particle physics pairs a very high threshold with a long list of requirements that have nothing to do with statistics.

What Sigma Counts

Sigma is a distance measured in standard deviations between what was observed and what pure background would give. Converting it to a probability yields the p-value: the chance that noise alone would produce a deviation at least this large. One sigma is about one in three, two sigma about one in 22, three sigma about one in 740, and five sigma about one in 3.5 million.

These numbers are sometimes quoted one-sided and sometimes two-sided, which changes them by a factor of two, and the convention is often left unstated. More importantly, they are computed from a model of the background. If the background model is wrong, the p-value is wrong, and no amount of statistical sophistication recovers it. The significance is only ever as good as the thing it is measured against.

The single most common misreading is to treat the p-value as the probability that the result is a fluke, or worse, as one minus the probability that the discovery is real. It is neither. It is a conditional statement in the opposite direction: the probability of this data given no effect, not the probability of no effect given this data. The two differ by how plausible the effect was to begin with, which statistics alone cannot supply.

This asymmetry has a practical edge. A three-sigma deviation from a well-tested theory with no competing explanation is treated very differently from a three-sigma deviation that would overturn a law with a century of confirming measurements behind it. The data are identical; what differs is everything else that is already known, and that is a legitimate part of the judgement rather than a bias to be eliminated.

Why Five Sigma

The threshold looks arbitrary and is not. A large detector tests thousands of mass ranges, decay channels and kinematic regions. Each one is an opportunity for a fluctuation. If a thousand independent regions are examined, a one-in-740 effect is expected roughly once per experiment simply by counting, and finding one proves nothing at all.

This is the look-elsewhere effect, and it is quantified rather than hand-waved: the local significance of a bump at one particular mass is converted into a global significance that accounts for how many places could have produced it. The correction is often large. A local three-sigma can become a global one-point-five sigma once the search range is accounted for, which is to say unremarkable.

Five sigma was adopted because it survives this correction with room to spare, and because the field has a long memory of three-sigma effects that evaporated. The empirical record supports the caution: the great majority of three-sigma anomalies in particle physics have disappeared with more data. The threshold is not a statement about how statistics works but about how often this particular community has been wrong at lower thresholds.

The 2015 diphoton excess at around 750 GeV is the instructive case. Two detectors at the Large Hadron Collider independently saw a bump, reaching roughly three to four sigma combined, and several hundred theory papers were written about what it might be. With the next year of data the excess vanished completely. Nothing went wrong - a fluctuation of that size at that frequency is exactly what the statistics predict - and the episode is the clearest available demonstration of why the threshold is where it is.

What Significance Does Not Cover

Significance assumes the only thing that can produce a deviation is random noise. In practice the more common cause is a systematic effect, and a systematic shift produces an arbitrarily high significance given enough data. The faster-than-light neutrino result reached six sigma and was caused by a cable connector. Significance is computed under the assumption that the apparatus behaves as described, and it therefore cannot test that assumption.

Nor does it address the analysis choices that preceded it. If the binning, the selection cuts or the fit range were adjusted while the result was visible, the reported p-value no longer means what it says, because the effective number of attempts is unknown. This is why blind analysis and pre-registered selection criteria exist, and why a significance quoted on a post-hoc analysis carries much less weight.

It also says nothing about size or importance. With a large enough sample, an effect too small to matter for anything reaches high significance, and in fields with very large datasets this is routine. Significance answers whether an effect is distinguishable from zero. Whether it is big enough to care about is a separate question that requires the effect size and its uncertainty.

Finally, it says nothing about interpretation. A statistically solid excess of events establishes that something is there, not what it is. The Higgs announcement in 2012 crossed five sigma for the existence of a new boson, and determining its properties well enough to identify it as the Standard Model Higgs took years of further measurement. Existence and identity are different claims with different evidence requirements.

Reading a Significance Claim

Four questions extract most of what matters. Is the significance local or global - that is, has the look-elsewhere effect been accounted for? Does it include systematic uncertainty or only statistical? Were the analysis choices fixed before the data were examined? And has an independent experiment seen the same thing with a different apparatus?

The fourth question usually matters more than the first three combined. Two experiments with different detectors, different backgrounds and different analysis teams seeing the same effect is worth far more than one experiment reporting a higher sigma, because the dominant failure mode is a systematic peculiar to one apparatus, and that is precisely what an independent replication tests.

This is also why significance thresholds vary sensibly between fields. Medicine and psychology conventionally use a p-value of 0.05, roughly two sigma, which particle physics would not treat as worth mentioning. The difference is not rigour but structure: a clinical trial tests one pre-specified hypothesis, while a collider search tests thousands of regions at once. A threshold appropriate to one is wrong for the other.

The useful summary is that significance is the cheapest part of a discovery claim and the only part that can be computed. The expensive parts are the error budget, the independent confirmation and the honest accounting of how many places were searched - and those are the parts that decide whether a result survives.

Frequently asked questions

What does a p-value actually mean?

The probability of observing an effect at least this large if there were no effect at all. It is not the probability that there is no effect, and not the probability that the result is a fluke. Those are statements in the opposite direction, and getting from one to the other requires knowing how plausible the effect was beforehand.

Why does particle physics require five sigma?

Because a large detector tests thousands of mass ranges and channels, so a one-in-740 fluctuation is expected roughly once per experiment by counting alone. Five sigma, about one in 3.5 million, survives that correction. The field also has a long record of three-sigma anomalies that disappeared with more data.

What is the look-elsewhere effect?

The inflation of apparent significance that comes from searching many places at once. The local significance of a bump at one specific mass has to be converted into a global significance accounting for how many positions could have produced it. The correction is often large: a local three sigma can become a global one-point-five.

Can a highly significant result still be wrong?

Yes, and this is the usual way it happens. Significance is computed assuming the only source of deviation is random noise, so a systematic effect produces arbitrarily high significance given enough data. The faster-than-light neutrino result reached six sigma and was caused by a cable connector.

What was the 750 GeV bump?

A diphoton excess seen by two LHC detectors in 2015, reaching roughly three to four sigma combined and prompting several hundred theory papers. It vanished completely with the next year of data. Nothing malfunctioned: a fluctuation of that size is exactly what the statistics predict at that frequency.