Every sponsor running a pivotal trial right now is building toward a single regulatory moment: the day a reviewer decides whether the evidence holds. But the intellectual framework those trials are designed against, the assumption that a study is valid if it can be repeated identically, was challenged head-on in the October 8, 2026 issue of the New England Journal of Medicine. Harvey Fineberg’s argument, is precise and uncomfortable: the biomedical community has been chasing the wrong target. Reproducibility, as the field has practiced it, conflates repetition with truth. Corroboration, the convergence of independent evidence from methodologically distinct studies, is what actually builds durable scientific knowledge. The distance between those two concepts is where most clinical development programs quietly collapse.
That gap has a price tag. A 2015 analysis published in PLOS Biology estimated that approximately $28 billion per year is spent on preclinical research in the United States alone that proves irreproducible. More than half of all preclinical work, a cumulative prevalence of 53.3%, fails to hold when tested again. Sponsors absorb that waste in their Phase I failure rates and never trace it back to the design decisions made three years earlier in a laboratory. The reproducibility problem has always been downstream of the corroboration problem.
What “Good Data” Has Been Getting Wrong
Consider what happens when a sponsor designs a Phase III confirmatory trial as a near-identical repeat of the Phase II signal study. Same endpoints, same patient population, same measurement windows. The logic is intuitive: if the same protocol produces the same result, the finding must be real. But identical methodology applied to the same population under the same conditions does not corroborate a finding. It amplifies the same systematic biases, the same site-selection artifacts, the same investigator expectations baked into unblinded endpoints. What looks like replication is often just the same error committed twice with higher enrollment.
A 2024 survey of 1,630 biomedical researchers published in PLOS Biology found that 72% agreed there is a reproducibility crisis in their field, and 62% cited “pressure to publish” as always or very frequently contributing to it. But the framing of a “reproducibility crisis” already accepts the wrong premise: that the solution is more repetition. Fineberg’s argument in the NEJM rejects that framing at the source. The crisis is not that studies fail to replicate. The crisis is that the field built its evidentiary standards around repetition and then discovered, slowly and expensively, that repetition cannot distinguish a true signal from a robust systematic error.
Sponsors designing trials today inherit that evidentiary architecture. Protocol registrations, SAPs, and endpoint selection all carry the implicit assumption that a pre-specified, blinded, repeated measurement constitutes proof. It constitutes consistency. Proof requires something harder.
The Corroboration Gap in Regulatory Practice
What corroboration demands, in practice, is methodological pluralism. A primary RCT endpoint corroborated by a mechanistically coherent biomarker signal, by real-world outcomes data from a distinct patient population, and by a trial emulation study using a different comparator arm, builds an evidence package that cannot be explained away by a single shared flaw. Each source of evidence has different vulnerabilities. When they converge, the convergence itself is the argument.
The FDA’s accelerated approval pathway illustrates what happens when that convergence never arrives. Research published through Penn’s Leonard Davis Institute found that it takes the FDA an average of 46 months to withdraw a drug that received accelerated approval but failed post-market confirmatory trials, with one specific withdrawal stretching to 58 months. Those are 46 months of patients receiving a therapy whose evidentiary foundation cracked the moment anyone looked closely. The initial approval was built on a surrogate endpoint that replicated well across early studies. It was never corroborated by independent evidence bearing on clinical outcomes. The distinction cost patients years.
Between 30% and 50% or more of trial results required by the U.S. government to be reported to ClinicalTrials.gov were not reported in the years before 2015, according to a Drug Discovery Trends analysis citing WHO data. Unreported negative results are the enemy of corroboration. A sponsor cannot assemble a convergent evidence package when the negative arms of prior investigations are sitting in a filing cabinet. The corroboration framework Fineberg argues for requires full information, which means the reporting gap is not a transparency problem in the abstract. It is a direct structural obstacle to building evidence that holds.
What Sponsors Must Build Differently
The NIH has recognized this at the funding level. Its R3PEATS initiative, formally titled Rigor, Replicability, and Reproducibility to Promote Excellence, Accuracy, and Translation in Science, commits $174 million over five years to support reproducibility in funded research. The dollar figure is meaningful. It signals that the federal research infrastructure accepts that the current model is generating unreliable outputs at scale. But $174 million directed at academic reproducibility does not automatically translate into commercial clinical trial design. The incentive structure for sponsors still rewards speed to a primary endpoint, not depth of corroborating evidence.
That incentive structure is where the Fineberg argument lands hardest for clinical operations professionals. A trial designed for corroboration looks different from a trial designed for replication. It incorporates mechanistic substudies that can confirm or refute the biological rationale independently of the primary endpoint. It pre-specifies the conditions under which a positive primary result would still be considered insufficiently corroborated, a discipline almost no commercial protocol imposes on itself. It treats a single-study NDA as a provisional claim, not a completed argument.
Some sponsors already build in this direction without naming it. A Phase III program that pairs a traditional RCT with an embedded biomarker cohort and a pre-planned real-world evidence extension is assembling corroborating arms by design. The question is whether that architecture is deliberate and documented, or whether it emerges opportunistically when a reviewer asks a hard question during the advisory committee. Regulators can see the difference.
The first sponsor to pre-specify a corroboration standard in its SAP, defining in advance what independent evidence streams would need to converge for the primary finding to be considered robust, will set a bar that the rest of the industry gets measured against. Every subsequent NDA in that indication gets read in comparison.
References
- New England Journal of Medicine, Fineberg HV, “Reproducibility in Biomedical Science, From Repetition to Corroboration,” Vol. 395, Issue 14, pp. 1447–1453, October 8, 2026
- PLOS Biology / PMC, “The Economics of Reproducibility in Preclinical Research,” June 9, 2015 ($28 billion annual waste; 53.3% cumulative irreproducibility prevalence)
- PLOS Biology, “Biomedical researchers’ perspectives on the reproducibility of research,” November 2024 (72% acknowledge crisis; 62% cite publication pressure)
- Penn Leonard Davis Institute, “It Takes the FDA 46 Months to Withdraw a Failed Drug with Accelerated Approval” (46-month average withdrawal; 58-month maximum)
- Drug Discovery Trends, “Large Numbers of Trial Results Go Unreported, WHO Says Report Them” (30–50% of required trial results not reported pre-2015)
- National Association of Scholars, “NIH Takes a Historic Step Toward Reproducible Science” (R3PEATS initiative, $174 million over five years)
Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.
