Measurement Uncertainty: What the Plus-Minus Actually Means
Measurement and Evidence 7 min read

Measurement Uncertainty: What the Plus-Minus Actually Means

Uncertainty is not a confession of sloppiness. It is the quantitative statement of how well a quantity is known, and it is the part of a result that decides whether the result can settle an argument. Two measurements that disagree by less than their uncertainties do not actually disagree.

What an Uncertainty Actually States

A result written as 1.372 plus or minus 0.015 metres per second is a claim about an interval, not about a point. The central value is the best estimate; the interval is where the quantity is believed to lie. The claim is incomplete until the confidence level is attached, because the same data supports a narrow interval stated with low confidence or a wide one stated with high confidence.

The convention in most physical sciences is the standard uncertainty: one standard deviation, roughly a 68 percent interval. Published results often use an expanded uncertainty, multiplying the standard uncertainty by a coverage factor - usually two, for approximately 95 percent. A number quoted as plus or minus something without saying which convention applies is ambiguous by a factor of two, which is more than enough to turn a disagreement into an agreement or back again.

This is why the international guide to expressing uncertainty requires the coverage factor to be reported alongside the figure. The rule is procedural rather than pedantic: an interval whose meaning is unstated cannot be compared with another interval, and comparison is the entire purpose of quoting one.

The corollary is worth stating plainly. A figure presented with no uncertainty at all is not yet a measurement result. It may be a reading, a specification, a typical value or a target. Each of those is a legitimate thing to publish, and none of them supports the inference that a measurement supports.

Where Uncertainty Comes From

The standard scheme splits uncertainty by how it is evaluated, not by what causes it. Type A is evaluated from the statistical scatter of repeated observations: take the readings, compute the standard deviation of the mean, and the number falls out of the data. Type B is evaluated from everything else - the calibration certificate of the instrument, the resolution of the display, the manufacturer specification, the published value of a constant used in the conversion, the estimated effect of a temperature drift.

The division confuses people because it sounds like random versus systematic, and it is not quite that. A systematic effect that has been measured in a dedicated side experiment is evaluated statistically and becomes Type A; a random effect too slow to average out during the run has to be bounded by judgement and becomes Type B. The scheme classifies the method of evaluation because that is what determines how the number was obtained.

Type B is where most of the real work sits, and it is where measurements most often go wrong. Scatter is visible and reassuring: it shows up in the data and anyone can compute it. A calibration error shows up nowhere in the data at all, which is why a tight scatter is frequently mistaken for a good measurement.

This is the difference between precision and accuracy, and the distinction is not rhetorical. A mis-calibrated instrument can return the same wrong value a thousand times, with a beautifully small standard deviation. Precision describes the agreement among repeated readings; accuracy describes agreement with the true value; and nothing internal to the dataset can tell them apart. Only comparison against an external reference can, which is what traceable calibration exists to provide.

How Uncertainties Combine

When several independent contributions affect one result, they do not add. They combine as the square root of the sum of their squares, because independent errors partly cancel rather than always reinforcing. The practical consequence shapes how experiments are built.

Take a result with three independent contributions of 5, 2 and 1 units. The combined uncertainty is the square root of 25 plus 4 plus 1, which is about 5.5 - barely more than the largest term alone. Eliminating the 1-unit term entirely would improve the total by less than two percent. Halving the 5-unit term would improve it by nearly 40 percent.

This is why a serious experiment publishes an error budget: a table listing every contribution with its magnitude, so a reader can see which term dominates. It also explains why effort concentrates almost entirely on the largest contribution, and why reducing a small one is usually not worth doing. An experiment that reports a single combined number without the breakdown has withheld the information needed to judge whether its limiting term is under control.

The square-root rule holds only when the contributions are independent. Correlated contributions - the same calibration constant used in two places, the same reference instrument, the same assumption in two steps - add more steeply, and in the worst case linearly. Missed correlations are a recurring way for a quoted uncertainty to come out too small, which makes a result look more decisive than the data allows.

Reading an Error Bar

An error bar on a plot is a claim with the same structure as a written plus-minus, and the same question applies: one standard deviation or two, and does it include the systematic contribution or only the statistical one? Many plots show statistical bars only, with the systematic uncertainty mentioned in the caption or a shaded band. Comparing a statistical-only bar against a total uncertainty is a common way to see a discrepancy that is not there.

Upper limits follow different rules again. When an experiment sees no signal, it reports a bound - a value the quantity is below, at a stated confidence. The KATRIN experiment reports the neutrino mass this way: not a measured value, but an upper limit at 90 percent confidence, which is the honest form of the result when the quantity is smaller than the apparatus can resolve. A limit is a real result and a strong constraint, and it is not a measurement of a value.

The most useful habit in reading any quoted figure is to ask what would have to be true for the uncertainty to be wrong. Usually the answer is an unmeasured systematic, and usually the paper says which one it worries about. Experiments are generally candid about this in their own text, and far less candid in their abstracts, press releases and the coverage that follows.

The practical test for any number encountered outside a paper is simply whether an interval and a confidence level are attached. If they are, the figure can be compared with others and argued about. If they are not, there is nothing yet to compare, and the first question is not whether the number is right but what kind of number it is.

Frequently asked questions

What does plus or minus actually mean?

An interval the quantity is believed to lie within, at a stated confidence level. The usual convention is one standard deviation, about 68 percent. Published results often use an expanded uncertainty with a coverage factor of two, about 95 percent. Without the convention stated, the figure is ambiguous by a factor of two.

What is the difference between Type A and Type B uncertainty?

Type A is evaluated from the statistical scatter of repeated readings. Type B is evaluated from everything else: calibration certificates, display resolution, published constants, estimated drifts. The split is by method of evaluation, not by cause, which is why a measured systematic effect can count as Type A.

Why do uncertainties add in quadrature?

Because independent errors partly cancel instead of always reinforcing, so the combined uncertainty is the square root of the sum of squares. The largest contribution therefore dominates: with terms of 5, 2 and 1, the total is about 5.5, and removing the smallest changes almost nothing.

Can a precise measurement be inaccurate?

Yes, and this is the most common way measurements mislead. A mis-calibrated instrument returns the same wrong value repeatedly, with a very small standard deviation. Nothing inside the dataset can reveal it; only comparison against a traceable external reference can.

Is an upper limit a measurement?

It is a result but not a measured value. When an experiment sees no signal it reports a bound at a stated confidence, as KATRIN does for the neutrino mass. That is the honest form when the quantity is below what the apparatus can resolve, and it constrains theory strongly without being a value.