Qualification: The reliability of replications: a study in computational reproductions
Breznau and coauthors qualify Kaal's classification of the current estimates as interim. In a computational reproduction experiment, 85 independent teams attempted to reproduce numerical results from the same study. Even the transparent group reproduced 95.7 percent of directions and significance classifications, but only 76.9 percent of estimates within 1 percent of the original values. Only 14 teams reproduced every result within that tolerance. The authors therefore show that access to code improves verification but does not remove researcher error or procedural variation. They estimate that more than one independent attempt may be needed for reliable reproduction. The comparison is methodological. The source does not inspect Kaal's E2B campaign, reconcile its amendment authority, or verify its estimates. It supports treating those estimates as interim until the recorded provenance is reconciled and independent reproduction determines whether the numerical results hold.
research-methodsscholarly-growth-coveragescholarly-literaturecomputational-reproducibilityindependent-verificationevidence-provenance