Calibration and Traceability: Where a Number Gets Its Authority
Calibration is the only procedure that can detect a bias, because a bias leaves no trace inside the data it corrupts. Traceability extends that single comparison into a chain reaching the SI, which is what allows a measurement made in one laboratory to be compared with one made in another decades later.
What Calibration Actually Does
Calibration is a comparison. The instrument under test is presented with a quantity whose value is already known to better accuracy, and the difference between the reading and the known value is recorded. The output is not a repaired instrument but a document: a calibration certificate stating the correction to apply and the uncertainty of that correction.
This distinction is routinely lost. Calibrating a device does not make it accurate; it makes its inaccuracy known and therefore correctable. An instrument with a large but well-characterised offset is more useful for precision work than one with a small but unknown offset, because the first can be corrected and the second cannot.
The reference must be better than the instrument, by a comfortable margin - the usual rule of thumb is a factor of three to ten in uncertainty. Calibrating a device against another of the same quality produces a comparison in which neither instrument is the authority, which establishes only that they agree, and two instruments can agree while both being wrong if they share a common error.
This is why calibration is the only thing that separates precision from accuracy. Nothing internal to a dataset reveals a constant offset, because a constant offset is perfectly consistent with itself. The comparison has to come from outside, which means the authority of a measurement ultimately rests on a chain of external references rather than on anything the instrument does.
The Unbroken Chain
Traceability means that the reference used for a calibration was itself calibrated against a better reference, and so on, in a documented chain terminating at a realisation of the SI unit. Each link has its own uncertainty, and the uncertainties accumulate as the chain descends. A working instrument on a bench is therefore always less certain than the national standard several steps above it, by a known and quantified amount.
The chain is maintained by national metrology institutes - the PTB in Germany, NIST in the United States, the NPL in the United Kingdom, and their counterparts elsewhere - which realise the units independently and compare their realisations against each other in international comparison exercises. Those comparisons are published, which is how the agreement between national standards is itself a measured quantity rather than an assumption.
A traceability claim is only as good as its documentation. A certificate must identify the reference used, the date, the conditions and the uncertainty; without those, the claim that something is traceable is an assertion rather than a record. In regulated measurement - electricity meters, fuel dispensers, medical devices, scales used in trade - the documentation requirement is a legal one, and the field is called legal metrology.
The practical consequence for reading any technical claim is simple. A stated accuracy is a claim about a chain, and the chain can be requested. When a figure is described as measured but no instrument, reference or uncertainty accompanies it, the chain has not been shown, and the figure is not yet traceable to anything. This is as true of an efficiency percentage or a power output as it is of a length.
The SI Since 2019
The international system of units was redefined in May 2019 so that all seven base units rest on fixed numerical values of physical constants: the speed of light, the Planck constant, the elementary charge, the Boltzmann constant, the Avogadro constant, the caesium hyperfine transition frequency and the luminous efficacy of a specified radiation.
The kilogram was the last unit still defined by a physical object - a platinum-iridium cylinder near Paris, whose mass was the kilogram by definition. The practical problem was that an artefact can be scratched, contaminated or lost, and comparisons over decades suggested it had drifted relative to its official copies by some tens of micrograms. A definition that can drift is a poor foundation, since by definition the drift cannot be detected: the artefact was always exactly one kilogram, and everything else moved.
The kilogram is now fixed through the Planck constant and realised with a Kibble balance, which relates mechanical to electrical power, or by counting atoms in a near-perfect silicon sphere. These are difficult experiments and the results agree, which matters more than either individually: two independent realisations that agree constitute a check that no single artefact could ever provide.
The distinction between defining and realising a unit is the useful idea here. The definition is a fixed number and cannot change. The realisation is an experiment, can be improved, and can be performed by anyone with the apparatus. This separation means that measurement accuracy can improve indefinitely without anyone having to renegotiate what the units mean - which is exactly what an artefact definition prevented.
Calibrating What You Cannot Take Apart
A large detector cannot be sent to a calibration laboratory. A neutrino observatory may hold thousands of tonnes of material and hundreds or thousands of photomultiplier tubes distributed through it, and its response varies with position, time and temperature. Calibration has to happen in place, with the instrument assembled and running.
The techniques are ingenious and all share one logic: introduce something known and see what comes out. Radioactive sources of known activity are lowered into the volume on cables, as Borexino did to map its response position by position. Light pulses of calibrated intensity are injected through optical fibres to track the gain of each tube over years. Cosmic-ray muons passing through provide a free, continuous and well-understood signal useful for monitoring. Known physical processes with well-measured rates serve as references in their own right.
IceCube faces the hardest version of this, because its detecting medium is a cubic kilometre of Antarctic ice that nobody chose and nobody can modify. The optical properties of that ice vary with depth, including dust layers deposited in past climates, and those properties had to be measured in situ with light sources and then modelled. The ice model is consequently one of the dominant systematic uncertainties in the experiment.
The honest consequence is that calibration never fully disappears from a result; it converts into a systematic contribution with an uncertainty of its own, and it appears in the error budget as a line item. Drift makes this permanent rather than one-off: detectors age, photocathodes lose sensitivity, electronics shift with temperature, and a calibration is valid for a period rather than forever. An experiment that calibrated once at the start and never again has an unquantified time-dependent bias, which is why continuous monitoring is built in rather than added later.
Frequently asked questions
Does calibrating an instrument make it accurate?
No. It makes the inaccuracy known and therefore correctable. The output of a calibration is a certificate stating a correction and the uncertainty of that correction. An instrument with a large but well-characterised offset is more useful for precision work than one with a small but unknown offset.
What does traceability mean?
That the reference used for a calibration was itself calibrated against a better reference, and so on in a documented chain ending at a realisation of the SI unit, with a stated uncertainty at every link. The uncertainties accumulate downward, so a bench instrument is always less certain than the national standard above it.
Why was the kilogram redefined?
Because it was the last unit defined by a physical object, and an artefact can be scratched, contaminated or lost. Comparisons suggested the prototype had drifted against its copies by tens of micrograms, and such drift is undetectable in principle: the artefact was one kilogram by definition, so everything else appeared to move.
How is a kilogram realised now?
Through a fixed value of the Planck constant, realised either with a Kibble balance relating mechanical to electrical power, or by counting atoms in a near-perfect silicon sphere. Both are difficult experiments, and the fact that two independent realisations agree is a check no single artefact could provide.
How is a detector that cannot be dismantled calibrated?
In place, by introducing something known. Radioactive sources of known activity are lowered into the volume, calibrated light pulses are injected through optical fibres, cosmic-ray muons provide a continuous free signal, and well-measured physical processes serve as references. The calibration then appears in the error budget as its own systematic contribution.