Systematic Error: Why Repeating the Measurement Does Not Help
Statistical uncertainty is the easy half of an experiment: it falls predictably as the square root of the number of events, and the only cost is patience. Systematic uncertainty does not fall at all unless someone finds and measures the effect causing it, which is why it eventually limits every precision measurement ever made.
Why More Data Does Not Help
Random fluctuations are symmetric: a reading is as likely to come out high as low, so the average of many readings converges on the true value. The statistical uncertainty on a mean falls as one over the square root of the number of observations, which means quadrupling the data halves the uncertainty. This relationship is the reason large experiments run for years.
A systematic error has no such symmetry. If a thermometer reads two degrees high, it reads two degrees high every single time, and the mean of a billion readings is still two degrees high with a vanishingly small statistical uncertainty attached to it. The dataset will look superb. Internal consistency is not evidence of correctness - it is evidence of consistency, which a constant offset supplies for free.
This produces the characteristic signature of a serious measurement error: high precision, high confidence, and a wrong answer. It also means the statistical uncertainty quoted on such a result is technically correct and practically irrelevant, because it describes the scatter of the readings rather than their distance from the truth. Separating the two is the entire content of measurement uncertainty as a discipline.
There is a useful operational consequence. When an experiment reports that its systematic uncertainty now exceeds its statistical uncertainty, it is saying that more running time would be wasted. Further progress requires a better apparatus or a better understanding of a specific effect, not more patience. Experiments are generally retired at this point rather than continued.
Where Systematics Come From
The usual sources are mundane, and that is what makes them dangerous. A calibration constant taken from the wrong revision of a document. A cable whose signal delay was estimated rather than measured. A detector whose efficiency depends on where in the volume an event happened, with an average efficiency used instead. A background process that was modelled and never measured. Temperature, pressure and humidity drifting slowly across a long run in a way that correlates with something else.
A second family comes from the analysis rather than the hardware: a selection cut that removes signal preferentially at one end of the range, a fitting function that cannot represent the real shape, a simulation used to correct the data that was tuned on the same data. Analysis systematics are harder to find than hardware ones because there is no physical object to inspect.
The canonical modern example is the 2011 OPERA result. The experiment measured neutrinos travelling from CERN to a detector in Italy and found them arriving about 60 nanoseconds earlier than light would, a six-sigma deviation from relativity on a very large dataset. The collaboration published it explicitly as an anomaly it could not explain and asked for scrutiny, which is how it should be done.
Two hardware faults were found within months: a fibre-optic connector that was not fully seated, which delayed the timing signal, and an oscillator correction applied with the wrong sign. The effects were of the right size and the right direction. The result was withdrawn in 2012 and relativity was never actually in question. The episode is quoted so often not because the collaboration behaved badly but because it behaved well, and the result still stood for a year.
How Systematics Are Hunted
The first tool is a control sample: a dataset where the answer is already known, run through the identical analysis. If the known answer comes out, the analysis is not introducing an obvious bias. Particle physics uses well-measured processes this way constantly, and a control sample that fails is the single most informative thing an experiment can find.
The second is deliberate variation. Change something that should not matter - the magnet polarity, the half of the detector used, the time of day, the data-taking period, the analyst - and see whether the result moves. A result that shifts when an irrelevant parameter changes has an uncontrolled dependence on that parameter. Large collaborations institutionalise this as a required battery of cross-checks before any publication.
The third is a side measurement: build a small dedicated experiment whose only job is to measure the suspected effect. This converts a guessed bound into a measured number with its own uncertainty, moving the contribution from a judgement call into the data. Most of the construction cost of a precision experiment goes into apparatus of this kind rather than into the primary measurement.
The fourth is procedural rather than technical. Blind analysis fixes every cut and correction before anyone sees the answer, which removes the possibility of stopping the hunt for systematics at the moment the result looks right. This matters because the search for systematic effects has no natural end point, and the temptation to stop when the number becomes agreeable is not a character flaw but a documented and measurable human tendency.
Reading an Error Budget
A published precision measurement normally carries a table listing each systematic contribution and its size. This table is the most informative part of the paper, and it is routinely the part that never appears in any summary of it. It shows which effect limits the result, how the limiting effect was evaluated, and therefore what would have to change for the result to move.
Two patterns in such a table deserve attention. One is a dominant term evaluated by judgement rather than measurement, which means the quoted uncertainty rests on an estimate that could be wrong in either direction by an unknown amount. The other is a suspiciously flat budget in which many contributions have similar size, which sometimes indicates that each was assigned a round number rather than determined.
The absence of a budget is itself informative. A result quoted as a single number with a single uncertainty, with no breakdown anywhere in the publication, cannot be assessed for what limits it. This is normal for an engineering specification and unusual for a measurement intended to settle a question.
The same reasoning applies outside physics. Figures for levelised cost, life-cycle emissions and energy return all vary between studies by more than their quoted uncertainties, and almost always because of systematic methodological choices rather than measurement noise. The question to ask of any such figure is not how precise it is but what was assumed, because that is where the real spread lives.
Frequently asked questions
What is the difference between random and systematic error?
A random error moves readings up as often as down, so averaging drives it toward zero and it shrinks as one over the square root of the sample size. A systematic error shifts every reading in the same direction, so averaging does not reduce it at all, no matter how much data is collected.
Why does a consistent dataset not prove a measurement is right?
Because a constant offset produces perfect internal consistency for free. A thermometer reading two degrees high reads two degrees high every time, so the scatter is small and the mean is wrong. Consistency is evidence of consistency, not of correctness, and nothing inside the data can distinguish them.
What happened with the faster-than-light neutrinos?
The OPERA experiment reported in 2011 that neutrinos arrived about 60 nanoseconds earlier than light, a six-sigma effect, and published it as an unexplained anomaly inviting scrutiny. Two hardware faults were then found: an incompletely seated fibre-optic connector and an oscillator correction with the wrong sign. The result was withdrawn in 2012.
How do experiments find systematic errors?
With control samples where the answer is already known, deliberate variation of parameters that should not matter, dedicated side experiments to measure a suspected effect, and blind analysis to prevent the search from stopping when the result looks right. Collecting more data does not help at all.
When is an experiment finished?
Usually when its systematic uncertainty exceeds its statistical uncertainty. At that point more running time buys almost nothing, and further progress needs a better apparatus or a better understanding of a specific effect. Experiments are generally retired at this point rather than extended.