Three developments have converged in the past eighteen months that the clinical operations community has not yet named as a single phenomenon: AI decision support tools are being deployed in active trials faster than any validation infrastructure exists to evaluate them. Call it the Evidence Latency Problem. The tools are real, the deployments are live, and the evidentiary scaffolding that should precede them is still being drafted.

The AI in clinical trials market was valued at USD 1.35 billion in 2024 and is projected to reach USD 2.75 billion by 2030, growing at a compound annual growth rate of 12.5%. That trajectory tells you one thing: capital has already decided this technology is ready. The question regulators and trial operators have not answered is: ready for what, validated how, and by whom?

The Regulatory Foundation Is Still Wet Concrete

Open the FDA’s April 3, 2023 draft guidance on machine learning-enabled device software functions (ML-DSFs) and read the first substantive section. The agency introduces the concept of a Predetermined Change Control Plan, a mechanism that allows sponsors to pre-specify the types of modifications an AI model can undergo post-authorization without triggering a new review cycle. The logic is sound. The problem is that this document is still a draft, which means any sponsor building a validation strategy around its principles is building on guidance that could change. Three years after publication, the concrete has not set.

The FDA’s posture here is understandable but incomplete. The agency faces a genuine epistemic problem: AI models in clinical settings are not static devices. They retrain, they drift, they behave differently across patient subpopulations. The traditional 510(k) and PMA frameworks were designed for hardware that does not update itself. Drafting governance for adaptive software requires rethinking categories that took decades to establish. That intellectual honesty does not make the operational uncertainty any less acute for the sponsor running a Phase II oncology trial with an AI-assisted endpoint adjudication tool in production today.

The gap between draft and final guidance is not academic. It is the difference between a sponsor being able to defend their validation approach in a Type B meeting and being told their methodology does not align with a framework that was never finalized.

Deployment Without a Standard to Deploy Against

Consider what Tempus AI did in early 2026. The company expanded its Next platform to six new clinical scenarios across breast, colorectal, ovarian, prostate, and urothelial cancers, citing a multi-center prospective study demonstrating significant improvements in biomarker testing rates for early-stage non-small cell lung cancersmall cell lung cancerlung cancer. The platform delivers real-time clinical intelligence at the point of care. Oncologists see it. They act on it. Patients are enrolled in trials based on its outputs.

That is not a hypothetical deployment. That is a production system operating across multiple institutions, in multiple tumor types, feeding into decisions that touch trial eligibility, biomarker stratification, and ultimately the integrity of clinical datasets. And the validation standard it was built against? The FDA’s ML-DSF draft guidance, which, as of today, remains a draft.

The counterintuitive assumption most people carry into this conversation is that more AI deployment means more evidence generation. The tools collect data, the data accumulates, the evidence base grows. But that logic inverts when you examine what the data is actually measuring. Prospective studies of AI performance in clinical settings tend to measure workflow outcomes, such as biomarker testing rates and time to result, rather than the deeper question of whether the AI’s underlying model generalizes reliably across the patient populations enrolled in trials. Improved testing rates are a process metric. They are not a validation of the model’s decision boundary. The 2024 JAMA Summit on Artificial Intelligence, which included Stanford Medicine faculty, explicitly identified critical gaps in the oversight frameworks used to validate AI in healthcare settings. Faster deployment does not close those gaps. It widens them.

What the eClinical Stack Cannot Yet Handle

The validation problem compounds when you trace it downstream into the data systems that clinical trials depend on. CDISC is actively developing AI and machine learning integration standards through its 360i initiative, aiming to embed machine-readable data models into SDTM and ADaM structures that can accommodate AI-generated outputs. The work is genuinely important. It is also years from becoming the default infrastructure at most operating sites.

Here is the operational reality: a sponsor deploying an AI decision support tool in a Phase III trial today must bridge two worlds simultaneously. The AI outputs need to be captured in eClinical systems built around CDISC standards that were not designed with model-generated data in mind. The audit trail requirements for AI-assisted decisions, specifically which version of the model produced which output, at what time, with what input data, fall into a documentation category that most EDC systems handle poorly. Sites have no standardized procedure for it. CROs have no validated monitoring checklist for it. The data integrity question is not whether the AI is accurate. The question is whether anyone can prove, post-hoc, exactly what the AI did and why.

The EMA recognized this problem from a different angle. Its Reflection Paper on the use of Artificial Intelligence in the medicinal product lifecycle emphasizes a human-centric approach and the mitigation of new risks to data integrity throughout a medicine’s development cycle, from discovery through post-authorization. The EMA’s framing is deliberately broad, covering the full drug development arc rather than a specific device category. But read between the lines and the message is the same as the FDA draft: regulators know the tools are already deployed, they are working to catch up, and sponsors operating in the interim carry the compliance risk alone.

That risk calculus lands differently depending on who you are. For a large pharma sponsor with dedicated regulatory affairs infrastructure, the ambiguity is manageable. You assign staff to track the draft guidance, you build a conservative validation package, you document everything twice. For a mid-sized biotech running a single pivotal trial with an AI-assisted patient matching platform, the same ambiguity is an existential question about whether your dataset will survive regulatory scrutiny. The FDA’s Breakthrough Devices Program can accelerate the review of AI diagnostic tools that address life-threatening conditions, but the program explicitly does not lower the evidentiary standard for marketing authorization. Faster pathway, same bar, no clarity on what the bar actually measures.

Sponsors and CROs that treat the current guidance vacuum as a temporary inconvenience will find themselves repricing that assumption when FDA finalizes the ML-DSF framework, likely with retrospective implications for tools already in production. The eClinical vendors building audit trail and data provenance capabilities for AI outputs will have an 18-month window of competitive separation before those features become table stakes. And any trial running AI decision support today without a documented model versioning protocol tied to its eClinical audit trail is carrying an inspection risk that does not appear on any current risk register. The first complete response letter citing inadequate AI validation documentation will arrive before the concrete on the FDA’s draft guidance has time to cure.

References

  1. Nature Medicine — “AI decision support is scaling-up fast — can the evidence keep up?”
  2. MarketsandMarkets — “AI in Clinical Trials Market: USD 1.35 billion in 2024, projected USD 2.75 billion by 2030”
  3. Intuition Labs — “FDA April 2023 Draft Guidance on ML-Enabled Device Software Functions and Predetermined Change Control Plans”
  4. Business Wire — “Tempus Expands Next Platform to Deliver Real-Time Clinical Intelligence Across Oncology”
  5. Stanford Medicine — “2024 JAMA Summit on Artificial Intelligence: Critical Gaps in AI Oversight and Validation”
  6. Intuition Labs — “CDISC 360i Initiative: AI and Machine Learning Integration for Clinical Trial Validation”
  7. NSF — “EMA Reflection Paper on the Use of Artificial Intelligence in the Medicinal Product Lifecycle”
  8. MedDeviceGuide — “FDA Breakthrough Device Designation: Program Overview and Evidentiary Standards”
Website |  + posts

Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.