A Dana Farber-linked research team recently ran an AI-assisted scoping review of roughly 4,000 published clinical prediction model studies and found that code sharing rates remained low, with recent years showing modest improvement. The direction is right. The magnitude is damning. When sponsors and investigators build prediction models that guide patient selection, endpoint stratification, or risk scoring in a trial, and then publish those models without sharing the underlying code, the work cannot be verified, the results cannot be reproduced, and the regulatory record is silent on the gap.
Three signals have emerged in the past year that, taken individually, look like routine methodological criticism. Taken together, they describe something the clinical operations community has not yet named: a reproducibility debt embedded directly in AI-driven trial infrastructure. The debt is accumulating faster than any existing regulatory mechanism can collect it.
The Collapse Nobody Flagged
Start with the sharpest edge of the problem. A study published in 2026 identified 125 clinical prediction model studies that used two unreliable stroke datasets. Three articles in Scientific Reports were retracted. Two more received expressions of concern. The trigger was a flag on a public annotation platform, not a regulatory audit, not a sponsor review, not a site inspection., not a regulatory audit, not a sponsor review, not a site inspection. A crowdsourced annotation platform caught what the formal system missed. Critically, some of those prediction models had evidence of use in actual clinical practice before the problem surfaced.
That sequence deserves close reading. Models were built, published without accessible code, deployed in clinical environments, and only identified as unreliable when a researcher flagged the underlying dataset on a public forum. No IND amendment captured the change in model inputs. No IRB was notified that the risk-scoring tool informing enrollment decisions rested on questionable data. The regulatory record shows nothing because the regulatory record had no mechanism to show anything.
This is a structural failure, not an isolated one. When code is unavailable, reviewers and downstream investigators cannot trace the analytical pathway from raw data to model output. A retraction notice tells you the destination was wrong. Without the code, nobody can tell you where the map went bad.
What the Policies Say, and What They Leave Open
The policy architecture is not absent. The NIH Data Management and Sharing Policy, effective January 25, 2023, applies to all NIH-funded research generating scientific data and explicitly covers clinical research. It requires data management plans that address code and software tools used to generate scientific data. The EMA published “Guiding principles of good AI practice in drug development” with a transparency principle stating that good model development promotes transparency, reliability, and reproducibility across the drug product life cycle. FDA finalized “Content of Premarket Submissions for Device Software Functions” in June 2023 and followed it with final guidance on Predetermined Change Control Plans for AI-enabled device software in December 2024.
Every one of those documents gestures toward transparency. None of them requires a clinical prediction model used in a drug trial to be accompanied by a publicly accessible, dependency-specified, reproducible code repository. The NIH policy asks sponsors to plan for sharing, not to execute it in a form that is actually reproducible. A management plan is not a GitHub repository. The EMA principles are principles, not submission requirements. The FDA’s software guidance targets device software functions, not the prediction models embedded in clinical protocols. There is no regulatory instrument in active use today that would have caught the stroke dataset problem before deployment.
Which raises the uncomfortable operational question: if the policies exist and the problem persists, what is the actual bottleneck?
Why Sponsors Will Not Self-Correct Without a Mandate
The conventional assumption is that code sharing fails because researchers are lazy or territorial. The data suggests a different structural reality. The review of 511 machine learning papers in health found that only 21% released analysis code publicly and only 55% used publicly available datasets. The researchers building these models are often operating under institutional intellectual property constraints, competitive pressures from co-developing industry partners, or simply the absence of any submission requirement that makes sharing obligatory rather than aspirational.
The counterintuitive read here is that the sponsors most capable of producing reproducible, well-documented code are precisely the sponsors least likely to share it voluntarily. A large pharma running a Phase III with a proprietary patient stratification model has legal, competitive, and reputational reasons to keep that model opaque. An academic group running a Phase II with a simpler model has the willingness but often lacks the infrastructure: a README file, dependency specifications, a containerized environment. The scoping review found that even among studies sharing code, many shared repositories lacked the dependency specifications needed to actually reproduce the work. Code that cannot be executed is not reproducible. It is an archive of intentions.
The FDA’s December 2024 final guidance on Predetermined Change Control Plans provides a mechanism for managing iterative modifications to AI-enabled devices. But a PCCP governs how a cleared device can change after clearance. It does not govern whether the original model, embedded in a trial protocol as a patient selection tool, must be reproducible before the trial starts. The regulatory gap is upstream of the device review process entirely.
Consider the concrete scenario: a Phase III oncology trial uses a machine learning model to select patients predicted to respond to a PD-1 inhibitor based on a composite biomarker score. The model is trained on a proprietary dataset. The protocol describes the model’s output threshold but not its architecture, training procedure, or code. The trial succeeds. The NDA is filed. The FDA reviewer sees the endpoint data but has no mechanism to audit the model that determined which 600 patients entered the trial. If the model was miscalibrated, the enrolled population was systematically biased. The approval stands on a foundation nobody can examine.
This is not hypothetical architecture. This reflects the current state of the regulatory record for AI-assisted trial designs, including in therapeutic areas such as oncology, cardiovascular disease, and neurology, where prediction model use in protocol design has grown substantially.
What Breaks First
Within 12 to 18 months, three pressure points will force this issue into regulatory conversations that cannot be deferred. First, post-market outcomes data for drugs approved on AI-stratified trial populations will start producing signals that retrospective auditors cannot explain without the original model code. The stroke dataset retractions in 2026 affected three models already in clinical use: the same dynamic, at larger scale, will emerge from trial populations enrolled on proprietary prediction scores. Second, the FDA’s ongoing work on AI-enabled clinical decision support will eventually have to reconcile the December 2024 PCCP guidance for devices with the absence of any parallel requirement for models embedded in drug trials, two regulatory tracks for the same underlying technology class, with radically different transparency standards. Third, the NIH DMS Policy will face its first serious enforcement test when an audited data management plan reveals a prediction model that was described but never shared in executable form. Sponsors who interpreted “plan for sharing” as a documentation exercise will discover the policy has teeth.
Technology vendors selling AI-powered trial design and patient matching platforms should read this trend carefully. The sponsors buying those platforms today are accumulating code assets that regulators will eventually demand to inspect. The vendors who build their platforms with exportable, dependency-specified, containerized model repositories will be positioned as compliance infrastructure. The vendors who do not will become liability exposure. The first sponsor to file an NDA with a fully auditable, publicly accessible prediction model code repository attached will set the evidentiary standard, and every later filing will be measured against it.
References
- Dana-Farber Cancer Institute, “AI-assisted scoping review of code sharing in clinical prediction model research”
- PubMed Central, Retractions and expressions of concern involving clinical prediction models using unreliable stroke datasets (2026)
- NIH, “Data Management and Sharing Policy Overview” (effective January 25, 2023)
- EMA, “Guiding principles of good AI practice in drug development”
- FDA, “Content of Premarket Submissions for Device Software Functions” (finalized June 14, 2023) and “Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions” (finalized December 4, 2024)
- Maguire Lab / Science Translational Medicine, Review of 511 machine learning papers in health: 21% shared code, 55% used public datasets
- Clarkston Consulting, “Predetermined Change Control Plans for AI-Enabled Medical Devices”
Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.
