A clinical operations team that deploys an AI screening tool at a single academic center and declares it validated has not validated anything. They have optimized for one data environment, one EHR schema, one patient population, and one set of institutional workflows. The moment that tool crosses a border, a health system boundary, or a language file, the performance numbers change. A Nature Medicine analysis of AI scaling across multiple healthcare systems to more than one million patients screened makes that gap concrete, and every sponsor running a decentralized or hybrid trial with AI-assisted endpoints should read it as an infrastructure audit, not a success story.
The counterintuitive finding is not that AI breaks at scale. The finding is that AI appears to hold together at scale while quietly accumulating data integrity debt that only surfaces during regulatory review. Local validation metrics look clean. Aggregate performance across sites looks acceptable. But the variance inside those aggregates, driven by inconsistent data labeling, site-specific imaging protocols, and patient demographic mismatches, creates a reproducibility problem that no single-site pilot study will catch.
The Data Harmonization Gap No One Is Measuring
The operational lesson from multi-system AI deployment is that the bottleneck rarely sits inside the algorithm. It sits between the algorithm and the source data. Consider what happened with AIRIS-TB, a tuberculosis screening AI evaluated across a large volume of chest X-rays. The model’s performance as a screening tool was sensitive to consistency in how images were acquired across sites, a pattern consistent with known challenges in multi-site AI deployment.
That is the same structural problem facing any sponsor who embeds an AI-assisted endpoint into a decentralized trial. The ePRO platform captures data. The wearable streams vitals. The AI interprets the signal. But if Site A in Germany is running a different firmware version than Site B in Brazil, and Site C in Japan acquired baseline measurements under a different calibration protocol, the AI is not operating on comparable inputs. The output looks like a single endpoint. The underlying data is three different experiments.
HL7 FHIR was designed to solve exactly this problem, providing a modern interoperability standard for harmonizing health data across systems. In practice, FHIR adoption remains uneven across healthcare institutions globally, and the clinical trial industry has no binding requirement to use it. Sponsors building multi-site AI infrastructure today are largely negotiating custom data transfer agreements site by site, which means harmonization quality varies by the strength of the relationship, not by a enforceable standard.
The FDA’s August 2023 final guidance on real-world data and real-world evidence for regulatory decision-making requires that RWD be “fit for purpose” and collected with sufficient rigor to support the regulatory question being asked. That language creates an obvious opening for reviewers to challenge AI-processed endpoints generated from non-harmonized source data. Sponsors who have not documented their cross-site data standardization procedures before first patient in will spend significant review cycles defending them after database lock.
Who Gets Exposed First
Oncology and rare disease sponsors running adaptive platform trials with AI-assisted response assessment carry the highest immediate exposure. These are programs where imaging endpoints, frequently processed by AI tools that were validated on single-institution training sets, drive dose escalation or futility decisions in real time. A performance drift in the AI’s lesion measurement algorithm across sites does not just affect the final analysis. It affects every interim decision made while the trial is running.
CNS sponsors using AI-based digital biomarkers face a parallel problem at the data governance layer. The 2023 editorial in npj Digital Medicine by Mittermaier, Venkatesh, and Kvedar identified digital endpoint harmonization as one of the central unsolved challenges for DHTs in clinical trials, noting that populations already experiencing the digital divide are systematically disadvantaged by technology-dependent endpoints. A CNS trial that deploys AI-processed speech or gait analysis across sites in high-income and lower-income countries is not running one study. It is running two studies with different effective data quality floors, and the protocol likely does not say so.
The regulatory environment is not waiting for industry to self-correct. The FDA’s draft guidance on AI-enabled medical devices, published in January 2025, outlines a total product life cycle approach, recommending that manufacturers plan for post-market performance monitoring from the design stage. For sponsors using Software as a Medical Device components as trial endpoints, that guidance signals that a validation package built at study start is not sufficient. Reviewers will want evidence of performance stability across the full data collection period, across all active sites, and across the demographic range of enrolled patients.
The WHO’s October 2023 publication on AI regulation for health reinforces the same expectation from a global perspective, outlining six regulatory domains including transparency and risk management that national regulators are being asked to operationalize. Sponsors running multinational trials should anticipate that competent authorities may increasingly ask questions rooted in that framework as national regulators operationalize the WHO guidance.
The Operational Fix That Cannot Wait for Guidance
If you are a clinical operations leader with an AI-assisted endpoint in a Phase II or Phase III trial currently in design, the Nature Medicine analysis points to one non-negotiable pre-deployment requirement: a prospective data harmonization protocol, executed before the first site goes live, that defines acceptable ranges for source data quality at each node where the AI touches the data stream. This means specifying imaging acquisition parameters, wearable calibration procedures, ePRO platform version control requirements, and EHR data extraction logic in the protocol or a referenced technical specification document. Without it, your AI validation rests on an assumption that your sites are running equivalent data environments. The evidence from over a million patient screenings suggests that assumption fails in practice more often than it holds.
The next regulatory signal to watch is the FDA’s final guidance on AI-enabled medical devices, which, once issued, would be expected to formalize the January 2025 draft’s total product life cycle framework into binding expectations. That final guidance will set the evidentiary bar for what a sponsor must demonstrate about AI performance stability across a trial’s full operational period. Sponsors who have already built prospective monitoring procedures for their AI components will file with confidence. Everyone else will be retrofitting their validation packages under review pressure, which is the most expensive place to learn that a pilot study was never enough.
References
- Nature Medicine, “Practical lessons in the global scaling of clinical AI: from one hospital to over a million patients screened”
- Federal Register, FDA Final Guidance: “Considerations for the Use of Real-World Data and Real-World Evidence To Support Regulatory Decision-Making for Drug and Biological Products” (August 2023)
- PMC, AIRIS-TB: AI screening across 1 million+ chest X-rays, January 2022 to January 2024
- npj Digital Medicine, Mittermaier, Venkatesh, Kvedar: “Digital health technology in clinical trials” (May 2023)
- WHO, “Regulatory considerations on artificial intelligence for health” (October 2023)
- Greenberg Traurig, FDA Draft Guidance on AI-Enabled Medical Devices, Total Product Life Cycle Approach (January 2025)
- PMC, Systematic review: cost-effectiveness of clinical AI interventions across global healthcare applications
Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.
