Picture a hospital command center at a major academic medical center in early 2025. The sepsis alert fires. A nurse glances at the screen, sees the AI flag, and pauses. Is this patient genuinely deteriorating, or is this the seventh false alarm today? She already started fluids twenty minutes ago. The algorithm caught nothing she hadn’t seen first. What no one in that room knows is that the AI tool flagging sepsis on their EHR had FDA authorization, appeared in the hospital’s procurement materials as “validated,” and had never completed a prospective randomized controlled trial in a clinical setting resembling theirs.

This scene plays out across American hospitals with uncomfortable frequency. According to the Office of the National Coordinator for Health Information Technology, 71% of U.S. hospitals reported using predictive AI integrated into their EHRs in 2024, up from 66% the prior year. The tools are proliferating. The evidence is not keeping pace. And the machinery connecting AI deployment to rigorous trial-level validation has a structural flaw that almost nobody in regulatory affairs is willing to say plainly.

The Authorization Gap Nobody Wants to Name

Start with the numbers. As of April 2026, over 1,500 AI-enabled medical devices have received FDA authorization. Sixty-eight percent of those authorizations occurred since 2022. Radiology dominates at 76% of the total, with cardiovascular at 10% and neurology at 4%. These are not pilot programs or academic curiosities. Many are embedded in EHR workflows at scale across hundreds of hospital systems.

But authorization through the 510(k) pathway, which the majority of these devices use, requires demonstrating substantial equivalence to a predicate device. Substantial equivalence is a comparator standard, not a clinical outcomes standard. A sponsor does not need to show that the AI tool improves patient outcomes. They need to show it is not substantially different from something the FDA has already cleared. In a field where the predicates themselves were often cleared before rigorous outcome validation existed, this creates a compounding problem: each new authorization builds on the evidentiary foundation of the last, and that foundation may be sand.

The Epic Sepsis Model is the sharpest illustration of where this leads. Widely deployed across health systems, the model demonstrated only 33% sensitivity in post-deployment analysis, meaning it missed two out of every three sepsis cases. Its positive predictive value was 12%, generating approximately seven false alarms for every actionable alert. The tool fired after clinicians had already intervened. It did not improve time-to-treatment. It introduced alert fatigue. And it did all of this at scale, across real patients, before anyone ran a prospective trial asking whether it should.

That is the mechanism. Authorization happens upstream of validation. Deployment happens before prospective evidence. And by the time a rigorous trial could theoretically surface the failure, the tool has already been embedded in clinical workflows for two years, baked into hospital contracts, and listed in procurement documents as standard of care.

The FDA’s December 2024 guidance on Predetermined Change Control Plans for AI-enabled device software functions was a serious attempt to address part of this problem. The guidance applies across De Novo, PMA, and 510(k) pathways and requires sponsors to pre-specify how their algorithms will be modified post-authorization. That is real progress on the drift problem: an AI model retrained on new data without regulatory review is functionally a different device. But the PCCP guidance addresses what happens after authorization. It does not require that pre-authorization evidence meet a prospective, outcomes-validated standard.

Name the core principle here: regulatory authorization and clinical validation are two separate processes, and the American hospital system is treating them as if they are the same thing.

What Vendors Are Not Telling Procurement Committees

A 2025 environmental scan published on PubMed examining commercially available AI clinical decision support solutions found that while more than half of vendors disclosed some information about their knowledge base, few demonstrated rigorous appraisal or alignment with established quality standards. Transparency around AI methodology and privacy protections was similarly thin. Procurement officers reviewing these tools are largely evaluating marketing materials and FDA clearance letters, not peer-reviewed outcome data from populations that match their patient mix.

Here is where the counterintuitive reality lands hardest: FDA authorization may actually be making this worse, not better. When a hospital procurement committee sees “FDA-cleared” on a vendor slide, it functions as a quality signal that stops further inquiry. The assumption is that the FDA’s process ensures clinical validity. It ensures safety and substantial equivalence. Those are not the same thing as clinical validity, and conflating them is costing patients in ways that aggregate data has not yet fully captured.

The adoption gap compounds this. ONC data shows that predictive AI adoption is significantly lower in small, rural, independent, government-owned, and critical access hospitals, precisely the settings where clinical staff have the least capacity to audit algorithmic outputs independently. The hospitals most likely to trust the tool uncritically are the hospitals least equipped to detect when it fails.

Closing the Loop Before the Next Deployment Wave

The trial design community has a specific role to play here, and it has been slow to claim it. Adaptive trial designs, pragmatic registry trials, and Bayesian evidence frameworks exist that could generate prospective outcome data from deployed AI tools without requiring traditional Phase III infrastructure. A sponsor integrating a CDS tool into 50 sites across a health system already has the infrastructure for a pragmatic trial. What it typically lacks is the regulatory incentive to build one.

The Nature Medicine analysis underlying this piece frames the core tension precisely: AI decision support is scaling up fast, and the evidence generation mechanisms needed to validate it at scale are lagging behind. The question for sponsors, IRBs, and the FDA is whether post-market study requirements under the PCCP framework can be structured aggressively enough to generate outcomes data before the next generation of tools layers on top of the current unvalidated stack.

Consider what that stack looks like in practice. A hospitalist in 2026 may interact with AI tools flagging sepsis risk, predicting readmission, recommending medication adjustments, and triaging imaging reads, all within a single shift. Each tool carries its own authorization history, its own training dataset, its own drift trajectory. None of them have been validated against each other in combination. The interaction effects between multiple deployed AI systems in a single care environment have received almost no prospective study.

The FDA’s posture on this is understandable given its resource constraints, particularly after the budget pressures of 2025. But the PCCP guidance, however well-designed, cannot substitute for a prospective evidence mandate at authorization. Requiring sponsors to pre-specify how they will change their algorithms is not the same as requiring them to prove those algorithms work before embedding them in clinical care. The distinction matters enormously when the tool in question is making decisions about sepsis, cardiac risk, or neurological deterioration.

The nurse at that command center deserves better than a 12% positive predictive value and an alert that fires after she has already acted. So does the next hospital system that signs a five-year contract for a tool that was authorized before anyone asked whether it saved lives. The evidence needs to precede the deployment, and right now, across most of the 1,500 authorized AI devices in clinical use, it does not.

References

  1. Nature Medicine — “AI decision support is scaling-up fast — can the evidence keep up?”
  2. Office of the National Coordinator for Health Information Technology — “Hospital Trends in Use, Evaluation, and Governance of Predictive AI, 2023-2024”
  3. BioSpace — “How AI-Enabled Software as Medical Devices Are Reshaping Regulatory Routes and Medicine”
  4. Akin Gump — “FDA Finalizes Guidance for AI-Enabled Medical Devices” (December 4, 2024)
  5. PubMed — Environmental scan of commercially available AI clinical decision support solutions
  6. Public Health AI Handbook — “Epic Sepsis Model: Deployment and Limitations”
Website |  + posts

Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.