On June 23, 2026, Nature Medicine published a head-to-head benchmarking study comparing general-purpose large language models against two specialist clinical AI tools, OpenEvidence and Wolters Kluwer’s UpToDate Expert AI, on questions submitted by practicing physicians. The results were unambiguous: OpenAI’s GPT-5.2, Google’s Gemini 3.1 Pro, and Anthropic’s Claude Opus 4.6 outperformed both cleared clinical tools across every medical benchmark tested. Neither OpenEvidence nor UpToDate Expert AI holds a claim to superior performance on the question type that matters most to practicing clinicians: the real-world, unstructured query that arrives at the point of care.
That result carries a specific regulatory implication. The FDA-cleared tools in this comparison passed through a defined regulatory pathway. The general-purpose models that outperformed them did not — at least not through any pathway calibrated to clinical decision support performance. The gap between what the clearance process assessed and what this study measured is the structural issue the Nature Medicine data surfaces.
What the FDA Pathway Currently Assesses
FDA’s December 4, 2024 final guidance on Predetermined Change Control Plans (PCCPs) for AI-Enabled Device Software Functions is the most current agency statement on how AI/ML-based medical devices should document planned modifications to their models. The PCCP framework requires sponsors to pre-specify the types of changes the algorithm may undergo post-clearance, the performance metrics that trigger a required submission, and the testing protocols that validate each change. It is a meaningful operational standard. It is also entirely device-architecture-focused: it governs how a cleared tool can change, not whether the cleared tool’s baseline performance was adequate relative to unregulated alternatives performing the same clinical function.
The FDA’s AI-Enabled Medical Devices List catalogs legally marketed AI/ML tools by submission number, decision date, and manufacturer. The list documents regulatory process completion. It does not document comparative clinical performance against general-purpose models performing equivalent tasks. When a cleared tool underperforms an uncleared one on the same physician query set, the clearance record offers no signal of that performance gap.
This is where the current framework shows a structural seam. FDA’s existing AI oversight pathway asks: does this device meet the performance specification the sponsor defined? The Nature Medicine study asks: does this device actually perform better than the freely available alternative a physician would otherwise use? Those are different questions. Regulatory clearance answers the first. It has no mechanism to answer the second.
The Enforcement Record Has Not Kept Pace
FDA’s February 10, 2025 warning letter to Exer Labs, Inc. for the Exer Scan AI diagnostic tool illustrates where current enforcement focus sits: on tools marketed without any clearance. That is a legitimate enforcement priority. It does not address the category this study reveals — tools that have clearance, but whose cleared performance falls below what unregulated foundation models now deliver to the same clinical users.
The distinction matters for sponsors running AI-assisted clinical trials. A CRO or eClinical platform that integrates a cleared clinical AI tool for protocol deviation flagging, safety signal detection, or eCOA response interpretation is relying on that clearance as a proxy for fitness-for-purpose. The Nature Medicine benchmark data indicates that proxy may be unreliable in at least some question domains. If the cleared tool underperforms a general-purpose model on the same clinical query, sponsors integrating it into trial workflows are accepting a performance floor defined by the cleared tool’s own validation dataset — not by comparative real-world performance.
The AMA’s 2024 physician AI survey found that 66% of physicians currently use AI in practice, up from 38% in 2023. A majority — 68% — reported perceiving definite advantage from AI tools. That adoption rate means physicians are already making treatment-adjacent decisions with AI assistance. Whether the tool they use is cleared, uncleared, or a general-purpose LLM is increasingly a secondary consideration for the user. It is a primary regulatory consideration that the current framework has not yet resolved.
What EMA and the PCCP Framework Signal for Sponsors
The EMA’s 2024 reflection paper on AI/ML in clinical development adopted a risk-tiered approach: AI/ML systems with high regulatory impact or high patient risk face heightened scrutiny, including prospective performance monitoring requirements. High-risk tools that could influence trial primary endpoints or safety reporting fall in the top tier. The EMA framework does not resolve the comparative performance question either, but its risk-tiering logic implies that a tool operating in a high-stakes clinical trial context should demonstrate superiority — or at minimum non-inferiority — to available alternatives, not just internal consistency against a sponsor-defined benchmark.
For sponsors with Phase 2 or Phase 3 programs currently integrating cleared clinical AI tools into eClinical workflows, the Nature Medicine study introduces a documentation question that is not currently covered by any FDA guidance: if your cleared tool underperforms a general-purpose model on the query types it will encounter in your trial, and that underperformance affects a downstream data point (a safety narrative, a protocol deviation classification, an eCOA interpretation), what is the sponsor’s evidentiary obligation to document that performance differential in the IND or trial master file?
FDA has not yet published guidance directly addressing this comparative performance standard for AI tools used in clinical trial operations, as distinct from AI tools used as primary diagnostic devices. The PCCP guidance addresses post-market change management for cleared devices. It does not address the pre-integration performance benchmarking question for clinical trial applications. The regulatory gap between these two documents is where sponsor liability currently sits.
Sponsors filing INDs in the next six months that incorporate AI-assisted clinical decision support — including tools integrated into risk-based monitoring platforms, eCOA systems, or safety data review workflows — should document the comparative performance rationale for the specific tool chosen. That documentation should include the question domains the tool will encounter in the trial context, the available performance benchmarks for those domains, and the justification for selecting a cleared tool (or a general-purpose model) over available alternatives. The Nature Medicine study, published in a peer-reviewed journal with identifiable commercial tools named as comparators, will be a citable benchmark from this point forward.
The next regulatory signal to watch: FDA’s Digital Health Center of Excellence has indicated it expects to issue updated guidance on clinical decision support software categorization in 2026. How that guidance defines “clinical decision support” relative to general-purpose LLMs performing the same function will determine whether the performance gap the Nature Medicine study documents becomes a regulatory classification question or remains an unresolved sponsor risk.
References
- Nature Medicine — “General-purpose chatbots outperform clinical AI tools on physicians’ real-world questions”
- Ropes & Gray — “FDA Finalizes Guidance on Predetermined Change Control Plans for AI-Enabled Device Software Functions” (December 4, 2024)
- FDA — Artificial Intelligence-Enabled Medical Devices List
- FDA — Warning Letter: Exer Labs, Inc. (February 10, 2025)
- American Medical Association — “AMA Physician Enthusiasm Grows for Health Care AI” (2024)
- Becker’s Hospital Review — “ChatGPT, Gemini, Claude Beat Clinical AI Tools: Study”
- Bioslice Blog — “EMA Adopts Reflection Paper on the Use of Artificial Intelligence (AI)” (October 2024)
Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.

