Imagine a site coordinator at a large academic medical center pulling up a patient’s chart in the middle of a Phase 3 oncology trial. Before she can complete her standard assessment, an AI-powered clinical decision support alert fires: elevated sepsis risk, recommended intervention, confidence score of 87%. She hesitates. The alert is generated by a vendor tool deployed hospital-wide — not validated against the specific patient population enrolled in her trial, not listed in the protocol, and not mentioned anywhere in the IND. She acts on it anyway, because the system is woven into the EHR workflow and declining the recommendation requires a three-click override that the nursing staff has been quietly skipping for months. The resulting intervention changes the patient’s concomitant medication profile. The deviation goes undocumented. The data gets locked.
That scenario is not hypothetical. According to a 2024 survey cited by Dialog Health, 71% of US hospitals now report using predictive AI integrated directly into their EHRs — up from 66% in 2023. These systems generate risk scores, flag deteriorating patients, recommend treatments, and shape clinical workflows in real time. They are present in the rooms where trials are running. And the clinical trial infrastructure built around GCP, protocol compliance, and data integrity has almost no coherent framework for handling them.
The central tension of this moment is not that AI decision support tools are unproven. Some of them work. The problem is that the evidence machine required to validate, monitor, and regulate these tools at scale simply cannot keep pace with the speed of commercial deployment.
The Regulatory Gap Is Already Showing
The FDA has been trying to get ahead of this. In April 2023, the agency published its Draft Guidance on Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence and Machine Learning-Enabled Devices — a document that acknowledges, implicitly, that AI systems are not static. They drift. They update. Their performance in the real world degrades relative to the conditions under which they were initially validated. The guidance asks sponsors to pre-specify how their AI will change and how those changes will be monitored. Reasonable in theory. Operationally ambitious at a scale that the agency itself is struggling to resource.
The FDA’s own enforcement record is beginning to show the consequences of that lag. On April 2, 2026, the agency issued what is believed to be the first warning letter directly tied to AI misuse, targeting Purolea Cosmetics Lab for deploying AI agents to generate cGMP documents — including drug product specifications and master production records — without adequate human review or validation of the AI outputs. The cited violation: failure to validate the AI system before relying on it for critical manufacturing decisions. The principle generalizes directly to clinical trial operations. If you are using an AI tool to support clinical decisions in a trial setting, and you cannot demonstrate that the tool was validated against a population and context comparable to your enrolled subjects, you have a protocol compliance problem with a data integrity wrapper around it.
What makes the Purolea warning letter significant is not its subject matter. It is its logic.
The FDA is signaling that reliance on AI outputs, without traceable validation, constitutes a GCP-equivalent failure — regardless of whether the tool was designed for manufacturing, clinical care, or trial operations. Sponsors running studies at sites where hospital-deployed AI tools are woven into standard workflows need to understand that those tools are now part of their regulatory exposure, whether or not they appear in the protocol.
When the Evidence Catches Up — Partially
There is at least one data point that shows what rigorous AI decision support evaluation looks like when it is done properly. A trial published in Nature Medicine followed over 9,600 patients across 16 primary care clinics in Kenya, testing a generative AI tool called “AI Consult” that was integrated into an EMR and provided real-time diagnostic and treatment suggestions to clinicians. The study found measurable improvements in clinician decision quality — not just process efficiency, but actual clinical decision accuracy. That is a meaningful result, and the trial design deserves credit: prospective, site-controlled, with a defined primary endpoint tied to clinical outcomes rather than user satisfaction scores.
But the Kenya trial also illustrates the limitation of the current evidence base. It was conducted in a specific resource-constrained setting, with a specific EMR, at 16 clinics, in one country. The moment you try to generalize that finding to a US academic medical center running a sponsored oncology trial across seven sites in four states, you are extrapolating well beyond the validation envelope. That extrapolation is happening every day, in procurement offices and hospital C-suites, without the clinical evidence infrastructure to support it.
The performance degradation literature makes the stakes clearer. A study published in New England Journal of Medicine AI, using a UK cardiac surgery dataset collected from 2012 to 2019, found that five machine learning models — including XGBoost and Random Forest — exhibited measurable performance drift over time as the underlying patient population and practice patterns shifted. Models that were accurate under 2012 conditions became demonstrably less reliable by 2019 without retraining. This is not a software engineering problem. It is a continuous evidence generation problem. And continuous evidence generation requires trial infrastructure, monitoring protocols, and regulatory oversight that the current system was not designed to provide.
What Sponsors Must Actually Do
The common assumption in the sponsor community is that AI decision support tools are the hospital’s problem — procured and deployed by health systems, governed by those health systems, and therefore outside the sponsor’s GCP obligations. That assumption is wrong, and it will generate CRLs.
Consider the operational reality at most multi-site trials today. A sponsor contracts with a CRO that monitors a network of sites. Those sites are embedded in health systems that have independently deployed AI tools across their clinical workflows. The AI alert that fires in the EHR during a protocol visit does not know it is operating inside a clinical trial. It fires because the patient meets a risk threshold. The site coordinator either acts on it or overrides it. Either action can affect trial data. Neither action is currently captured in most data management plans, protocol deviation logs, or risk-based monitoring frameworks.
The interoperability layer compounds this. Research published in AI (MDPI) on the FHIR-RAG-MEDS system demonstrates that HL7 FHIR-integrated AI can pull patient-specific data and generate guideline-concordant recommendations with impressive accuracy in controlled conditions. But “controlled conditions” in a research paper and “conditions present at Site 14 in rural Ohio” are not the same environment. FHIR standardization is incomplete across real health systems, meaning AI tools at different sites within the same trial may be operating on structurally different data inputs — generating non-comparable recommendations that systematically bias the treatment experience across arms without leaving any audit trail.
Sponsors who want to get ahead of this need to take three concrete steps before the next site activation. First, add an AI environment assessment to your site feasibility questionnaire — explicitly ask which AI-powered clinical decision support tools are active in the EHR, what patient populations they cover, and whether they generate alerts that could affect enrolled subjects’ care. Second, work with your medical monitor to define a protocol deviation category for AI-influenced clinical decisions, with a threshold for what requires documentation. Third, if your protocol is being run in a therapeutic area where AI tools are already commercially deployed — sepsis, cardiac risk, readmission prediction — consider adding a site-level AI tool inventory as an ongoing monitoring deliverable, reviewed at each co-monitoring visit.
The CMS 2024 Final Rule, effective January 1, 2024, allows Medicare Advantage plans to use AI to approve claims without human oversight — which means the financial incentives for health systems to deploy more AI, faster, are now structurally embedded in reimbursement. Sites will not slow AI deployment to accommodate trial governance timelines. The tools will proliferate regardless. The question facing every sponsor and CRO is whether their trial infrastructure will adapt before the FDA starts issuing 483 observations that explicitly reference AI-influenced data.
Back to that site coordinator with her finger hovering over the override button. She is not making a bad decision. She is making the only decision her workflow allows, inside a system that was designed for efficiency, not for the evidentiary requirements of a Phase 3 trial. The protocol says nothing about the AI alert. Her supervisor says nothing about the AI alert. The monitoring plan says nothing about the AI alert. When the data lock comes and the FDA reviewer pulls the audit trail, the intervention will be there, the concomitant medication change will be there, and the rationale will be missing — because no one in the trial’s governance structure ever asked the question that the evidence base is only beginning to force into view.
References
- Nature Medicine — “AI decision support is scaling-up fast — can the evidence keep up?”
- Dialog Health — “AI Healthcare Statistics: Hospital Trends in the US, 2024”
- U.S. FDA — “Artificial Intelligence in Software as a Medical Device,” including April 2023 Draft Guidance on Predetermined Change Control Plans
- GMP Compliance — “Use of AI Agents Leads to the First FDA Warning Letter Relating to AI,” April 2, 2026
- University of Birmingham — “AI clinical support tool improved clinician decisions in real-world primary care trial,” 2026
- PMC / NEJM AI — “Performance Drift in Machine Learning Models for Cardiac Surgery Risk Prediction,” UK dataset 2012–2019
- MDPI AI — “FHIR-RAG-MEDS: HL7 FHIR and LLM Integration for Clinical Decision Support”
- Super Lawyers — “Understanding Your Rights: Guardrails for AI in Medicare Coverage Decisions,” CMS 2024 Final Rule
Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.

