Replication: Why Someone Else Has to Get the Same Answer
Measurement and Evidence 7 min read

Replication: Why Someone Else Has to Get the Same Answer

Replication is the step at which a result stops being a claim by one group and becomes something the field knows. It is also the step most often skipped, because confirming someone else's finding is less rewarded than producing a new one, which is the structural cause of most of what is called the replication crisis.

Three Words That Are Not Synonyms

Reproduction means taking the original data and the original analysis and getting the original numbers out again. It tests whether the computation was done correctly and whether the description of it was complete. It cannot test the measurement, because the data are the same data, and any error in how they were collected is preserved exactly.

Replication means doing the measurement again, collecting new data, usually with a different apparatus and a different team. This tests the measurement, because an error specific to one instrument, one laboratory or one set of conditions will not recur in a different one. This is the step that converts a result into knowledge.

Repetition - the same group doing the same thing again with the same equipment - sits between them and is the weakest of the three. It tests stability and little else. A systematic error reproduces perfectly under repetition, which is precisely why internal consistency is not evidence of correctness and why a result repeated a hundred times by its authors is not thereby confirmed.

The distinction has a practical edge when reading any claim of confirmation. Being told that a measurement was repeated many times says almost nothing about whether it is right. Being told that a different group, using a different method, obtained a consistent answer says a great deal. The two statements are often made in the same tone and carry entirely different weight.

What the Replication Crisis Showed

In 2015 a large collaboration attempted to repeat 100 published psychology experiments using the original materials and, where possible, the original authors' input. Roughly a third produced a statistically significant effect in the same direction, and the average effect size in the replications was about half that of the originals. Comparable efforts in cancer biology and experimental economics returned results in a similar range.

The causes are mostly structural rather than dishonest. Publication favours positive, novel and surprising findings, so a literature assembled from accepted papers over-represents effects that happened to come out large. Small samples make a published effect size an overestimate even when the effect is real. Flexible analysis choices allow a result to be found in data where none exists. And almost nobody is funded or rewarded to check.

The term crisis is contested for a reasonable reason: discovering that a third of results replicate is itself a scientific finding, and it was produced by the field examining itself. The response has reshaped practice - pre-registration, published analysis plans, larger samples, registered replication reports, data and code sharing as a condition of publication. These are the same instruments that physics arrived at through blind analysis, applied to a different problem.

Physics is not exempt, and its advantage is structural rather than moral. Large collaborations with internal review, blinding as standard practice and a five-sigma threshold make a published particle-physics result harder to get wrong, but that machinery is expensive and only exists in fields organised around large shared instruments. A small-scale physics measurement made by one group with its own apparatus is exposed to exactly the same failure modes as a small-scale psychology experiment.

Why Independence Is the Active Ingredient

Two experiments that share an assumption, a calibration source, a simulation package or a key component are not fully independent, and their agreement is weaker evidence than it appears. The value of a replication rises with the number of things that are different about it, because every shared element is a channel through which a common error can pass undetected into both results.

The strongest form is confirmation by a different physical method. The solar neutrino problem is the clearest example in this library's subject area. Raymond Davis measured a deficit of solar neutrinos from the 1960s onward using a radiochemical method - counting argon atoms produced in a tank of cleaning fluid - and found roughly a third of the expected rate. For decades the most economical explanation was that something was wrong with his experiment or with the model of the Sun.

What settled it was different techniques agreeing. Kamiokande saw the deficit using water Cherenkov detection, an entirely different physical process with entirely different systematics. Borexino measured individual components of the spectrum with a liquid scintillator. Decisively, the Sudbury Neutrino Observatory used heavy water to measure both the electron-neutrino flux and the total flux of all flavours, and found the total matched the solar model while the electron component did not. The neutrinos had not disappeared - they had changed flavour.

That conclusion is now textbook physics, and it is instructive that no single experiment established it. Each one alone was open to the objection that its particular method was at fault. The combination was not, because the methods had nothing in common except the quantity they measured - and the result overturned the Standard Model prediction that neutrinos are massless.

When Replication Fails Fast

In March 1989 two chemists announced that they had produced excess heat from deuterium in a palladium electrode at room temperature. The claim was extraordinary, the apparatus was described, and laboratories worldwide attempted it within weeks. The great majority found no excess heat, the few positive reports did not survive scrutiny, and the mainstream verdict was reached within months.

In 2023 a preprint claimed a room-temperature ambient-pressure superconductor called LK-99, with a published synthesis recipe. Groups around the world made samples within days. The resistivity drop was traced to a copper sulphide impurity undergoing a structural transition, and the apparent levitation to ferromagnetism in an inhomogeneous sample. The question was effectively closed in about four weeks.

Both episodes are usually told as cautionary tales, and they are better read as the system working at speed. In each case a specific, falsifiable claim with a published method was tested by many independent groups and resolved quickly. What made rapid resolution possible was that the claim was concrete enough to attempt - the recipe, the conditions and the measurement were all stated.

This is the practical difference between a claim that can be settled and one that cannot. A result accompanied by a described apparatus, stated conditions, a measured quantity and an error budget invites replication and gets it. A claim without those elements cannot be replicated even by people who would like to confirm it, and it therefore stays in the same position indefinitely - neither confirmed nor refuted, which is a far weaker place to be than refuted.

Frequently asked questions

What is the difference between reproduction and replication?

Reproduction re-runs the original analysis on the original data and tests whether the computation and its description were correct. Replication collects new data, usually with a different apparatus and team, and tests the measurement itself. Only replication can catch an error in how the data were collected.

What did the 2015 psychology replication project find?

Of 100 published experiments repeated with the original materials, roughly a third produced a statistically significant effect in the same direction, and the average effect size in the replications was about half that of the originals. Comparable efforts in cancer biology and experimental economics gave results in a similar range.

Why does independence matter so much?

Because every element two experiments share - an assumption, a calibration source, a simulation package, a component - is a channel through which a common error can pass undetected into both. The value of a replication rises with how much is different about it, and confirmation by a different physical method is the strongest form.

How was the solar neutrino problem settled?

By different techniques converging. A radiochemical method found about a third of the expected rate, water Cherenkov detection confirmed the deficit, a liquid scintillator measured individual spectral components, and the Sudbury Neutrino Observatory measured both the electron-neutrino flux and the all-flavour total, finding the total matched the solar model. The neutrinos had changed flavour.

What do cold fusion and LK-99 have in common?

Both were specific, falsifiable claims published with enough method for others to attempt. Laboratories worldwide tried within days or weeks, and both were resolved quickly - cold fusion in months, LK-99 in about four weeks, where the effects were traced to a copper sulphide impurity and to ferromagnetism in an inhomogeneous sample.