Pani
Luca Pani, MD, Chief Innovation and Regulatory Officer, NetraMark

A negative Phase 3 result does not necessarily establish that a treatment has no biological or clinical effect in any patient. That argument sits at the center of NetraMark’s clinical platform, and few people are positioned to make it with more authority than Pani. Pani previously served as Director General of the Italian Medicines Agency and as a member of the European Medicines Agency Management Board, the Committee for Medicinal Products for Human Use and the Scientific Advice Working Party, spendingt decades reviewing the evidence packages that determine whether a treatment reaches patients. Now, as Chief Innovation and Regulatory Officer at NetraMark, he brings that regulatory lens to a different problem: why sponsors keep designing trials that are structurally set up to miss the signal. In this conversation, Pani explains what a credible AI-derived responder profile actually requires, why “just increase the sample size” is the wrong answer to biological heterogeneity, and why the most important question in clinical development is not whether a drug works, but who it works for.


Why did your regulatory experience convince you that a failed Phase 3 was often a patient selection problem rather than proof the drug didn’t work?

Luca Pani: Just to clarify: what convinced me to join NetraMark and help build it into a regulatory-grade clinical platform was not one failed trial. As a regulator, both at the Italian level, where I was head of the agency, and at the European level, I had seen this pattern repeatedly.

When a Phase 3 is negative or fails for some reason, the conclusion is almost always: “The drug did not work.” But that is not what the trial is actually showing. What it shows is that the pre-specified analysis did not demonstrate the required treatment effect in the enrolled population,using that dose, that endpoint, that design, those sites, and that execution. Many other factors are involved. And this is an important distinction, because a failed trial is not proof that the treatment has zero biological or clinical effect.

A late-stage trial can fail for several reasons. As a rapporteur and a coordinator of a major agency who reviewed or oversaw hundreds of development programs and pivotal evidence packages, I saw that some molecules were genuinely ineffective,the dose was wrong, or the endpoint was insensitive. But I also learned that these explanations must never be invented after the fact. Post-hoc analyses can be scientifically valuable for generating hypotheses, but they cannot retrospectively convert a negative confirmatory trial into positive confirmatory evidence. Treat any subgroup identified after unblinding as exploratory until you test it independently and prospectively. .

The example that really crystallized this for us at NetraMark was a major depressive disorder program. They had a very positive Phase 2 study, by all measures. And then the Phase 3 failed. In the broad population, the Cohen’s D was 0.082 and the p-value was 0.558, almost 0.6.

However, NetraAI was able to analyze the Phase 2 data and found a compact, treatment-favoring patient pattern. That is very important: this is a pattern the machine finds. It is not a language model describing what patients look like. When you apply that pattern retrospectively to the Phase 3, the estimated effect size moves to 0.346 and the p-value becomes 0.053. That is still not a successful confirmatory trial. I want to be very disciplined about what that means. But it does generate serious, falsifiable, prospectively testable hypotheses suitable for regulatory discussion.. The next trial can be designed to stratify those populations. For me, that was the conceptual turning point. That retrospective result does not rescue the original Phase 3 trial or establish efficacy. It generates a falsifiable hypothesis that can be locked before the next trial through prespecified eligibility or stratification criteria, an appropriate testing hierarchy and control of Type I error.


Why do sponsors keep designing broad Phase 3 enrollments when the evidence for biomarker-based selection is this strong?

Luca Pani: I will be a little less politically correct here. The first reason is commercial. Companies understandably want the broadest possible label and the largest possible market. A broad label has enormous value, and the logic is understandable. The problem is it only works if the trial succeeds. A clinically meaningful responder subgroup can be diluted in a broad trial, exposing the sponsor to the loss of tens or hundreds of millions of dollars in late-stage development.

The second reason is scientific uncertainty. A biomarker might look promising but may not have a sufficiently validated assay, a stable threshold, or convincing evidence that it predicts differential treatment benefit rather than just prognosis alone. That is an essential distinction: a prognostic marker tells you who will do well or badly regardless of treatment; a predictive marker tells you who benefits more from the drug than the comparator. For predictiveenrichment, you need the second question answered, not the first.

The third reason is operational. Narrowing eligibility creates screen failures and makes recruitment much harder. I was just called by a recruiter who said, “We are running according to the inclusion and exclusion criteria and we cannot find the patients.” The patients exist, but you cannot enroll them because you need new assays, very few capable sites, and so on. And how do sponsors usually respond to that uncertainty? They increase the sample size. They do not refine how they find those patients. They just say, we want 300 or 500. That will cost a fortune and elongate everything. Larger noisy trials can improve statistical precision, but more patients alone do not correct a poorly chosen population, and insensitive endpoint, or systematic bias.

The fourth reason is strategic and regulatory. If you study only the biomarker-positive population, what can you say about everyone else? Could the treatment work in biomarker-negative patients? Human genetics and epigenetics are very complex. The FDA enrichment guidance asks sponsors to consider both the enriched population and the patients who do not meet the enrichment criteria. A regulator’s responsibility is to approve a population no broader – and no narrower – than the totality of evidence and the benefit-risk assessment can support. . We have published extensively on equity and representation in trials, including two commentaries, one in Nature Medicine and one in Nature Digital Medicine, specifically on this.

And finally, many treatment response patterns are not captured by a single biomarker. They arise from combinations of clinical severity, physiology, imaging, cognition, medical history, molecular information, and placebo response. We pursue biomarkers of vulnerability, but we tend to ignore biomarkers that might be protective. That makes things considerably more complicated. Who enters a trial is not an administrative detail. It is part of the scientific hypothesis.

“Who enters the trial is not merely an administrative matter or a set of inclusion and exclusion criteria.. It’s part of the scientific hypothesis.”


How does NetraMark evaluate whether an AI-derived responder profile from Phase 2 data is credible enough to inform a Phase 3 enrollment strategy?

Luca Pani: Credibility cannot be a p-value. And it is certainly not something you can just claim. Credibility in front of EMA, PMDA, and FDA is an evidence package.

NetraAI was specifically designed to find interpretable structure in the small, complex datasets typical of clinical trials, where conventional machine learning approaches are vulnerable to overfitting. When we receive a Phase 2 dataset, the first question we ask is: what decision will this analysis support? Are we generating an internal hypothesis about whether to finance a Phase 3? Are we excluding patients from a pivotal trial? Are we eventually supporting a regulatory submission? That matters enormously, because depending on the consequences of the AI output, the credibility burden is higher. We ask the sponsor to fill in the evidence they have.

The second question is whether the data can be trusted. You can give me a zillion data points, but if they are dirty, it is junk in, junk out. No machine, no AI, nothing will fix that. We need to examine provenance, missingness, protocol deviations, site and country effects, rater behavior, variable definitions, and treatment exposure. We also need to confirm that no information from after baseline has accidentally entered a model intended for prospective screening. When sponsors ask how we handle missing data, my regulatory answer is: we do not recreate missing observations by imputation; we estimate them under assumptions. Those assumptions must be explicit, justified and examined through sensitivity analyses.

The third question is whether the model-derived subgroup represents treatment predictivity, not just improvement. It is not enough to find patients who improve, because they might improve equally on placebo. In the Phase 3 trial I described, the loss of drug-placebo separation appeared to be driven substantially by increased placebo response rather than by a comparable deterioration in the active-treatment outcome. . NetraAI explicitly accounts for placebo response rather than treating it as noise. It examines patterns of response in both treatment and placebo groups to identify patient profiles associated with greater treatment-control separation. The mathematics are Dr. Geraci’s domain, not mine.

The fourth question is whether the pattern is stable. Does it survive resampling? Does it recur across different training splits? Does it remain in an untouched held-out dataset? Does it survive perturbations in the assumptions? We use repeated resampling, robustness testing and, where feasible, a genuinely untouched holdout dataset. . Overfitting is the failure mode here.

The fifth is operational. A Phase 3 screening procedure cannot depend on 150 obscure variables. It needs to be a compact combination of clinically measurable variables with defined ranges, reproducible data collection, and practical screening procedures for sponsors and sites. The regulator will ask: make the trial easy to run. Do not make it impossible.

The sixth point is more philosophical. Does the model know what it does not know? This is, for me, the strongest aspect of the mathematical model Dr. Geraci designed. NetraAI does not force every patient into a subgroup. If a patient’s data do not support a stable or interpretable classification, they remain unknown. It is incredibly difficult to build a model that makes a genuine no-call rather than forcing an answer. That reduces coverage, yes. But coverage is not the objective. In high-stakes medicine, the hat no-call carries a real ethical commitment. As a physician, I have been in situations where the data did not fit, and the honest answer was: I don’t know. That is what the model does as well.

And the final question is whether the hypothesis can be tested prospectively. The variables, thresholds, missing data rules, estimates, and statistical hierarchy all need to be pre-specified. In the NIMH ketamine analysis published in npj Digital Medicine, conventional classifiers using the original clinical variables produced mean AUC values ranging from 0.27 to 0.61. After NetraAI identified an explainable 10-variable structure and a subgroup of 26 of the 63 evaluable treatment-period observations from a 33-participant crossover study, 81% of whom belonged to the favorable response class, the same classifiers produced mean AUC values of 0.82 to 0.87. Random forest, gradient boosting and the deep neural network each achieved a mean accuracy of 84.2% in nested cross-validation. These were internal, retrospective classification results using NetraAI-defined subgroup labels, like a like-for-like demonstration of improved prospective prediction of individual treatment benefit. The dataset was small, which limits generalizability, but the importance is that the retrospective analysis can generate a fixed rule that is then independently and prospectively tested. .

My practical requirement for an AI-derived responder profile is this: it must be treatment-specific, stable, interpretable, operationally measurable, and independently challenged. And one thing I always remind the team: before randomization, you can adjust, refine, and make decisions. Once a confirmatory trial is underway—and particularly after treatment information has been examined—the primary hypothesis cannot simply be altered without losing its prospective status. Prespecified adaptations, safety-driven changes and properly governed protocol amendments remain possible, but they must preserve trial integrity and statistical interpretability.


How do you expect AI-derived patient stratification to change the go/no-go decision at the end of Phase 2 over the next five years?

Luca Pani: This is usually the one question I do not answer because I have been burned by predicting AI timelines before. I have been doing this since 2017, so it is almost 10 years now, and I have been wrong about timing consistently. But I think the direction is clear.

The end-of-Phase 2 decision will no longer be a binary go/no-go conversation. It will become a structured set of development options. A sponsor may decide to go forward in the broad population, or go forward in a prospectively enriched and stratified population, or study both with hierarchical testing, or run a smaller validation study first, or go to the FDA and ask whether they agree with the plan, or change the endpoints, the duration, the dose, the site mix, or the stratification strategy, or seek a partner for a specific responder population, or stop the program because there is not enough average response. I think this is moving from a win-lose Phase 2 to a learn-and-lock Phase 2.

NetraMark held a non-binding Critical Path Innovation Meeting with FDA to discuss NetraAI’s proposed context of use, evidentiary expectations and possible pathways for sponsor-specific implementation. The feedback supported continued early engagement and consideration of established regulatory mechanisms, including Model-Informed Drug Development interactions where appropriate. The CPIM did not qualify or approve NetraAI, and it did not establish a single pathway applicable to every development program. That is exactly the direction you are asking about.

I think it will move faster than five years. The FDA has publicly reported more than 500 submissions containing AI components between 2016 and 2023. The FDA and EMA’s 2026 joint 10 Guiding Principles of Good AI Practice in Drug Development show the direction is clear: human-centric, risk-based, documented. AI will not lower the evidentiary bar. What it will do is enable better reasoning much earlier, before a company commits hundreds of millions of dollars and thousands of patients into a pivotal experiment.

For sponsors and investors, I expect systematic analysis of treatment-effect heterogeneity eventually to become a routine component of end-of-Phase 2 decision-making, just as sensitivity analyses and prospective statistical planning are today. That is where I think this is heading.


Why does it matter ethically that a subgroup is later found to have benefited from a drug that failed its broad Phase 3?

Luca Pani: This is the question I like most, because it brings us back from algorithms and models to people. I have lived through this.

When a later analysis shows a subgroup has a favorable average treatment effect, we cannot tell individual patients and their families that they personally would have benefited. Treatment effect is always counterfactual. We observe what happens under the treatment a person receives. We cannot observe what would have happened to the same person at the same time under the alternative. Some subgroup members may have received the investigational drug and improved. Some may have been on placebo and received no benefit. Some may not have responded even though they fit the profile. The model-derived subgroup changes probability. It does not create individual certainty.

That creates an obligation. We cannot convert an exploratory retrospective signal into a personal medical promise. But we also cannot stay indifferent. We have to analyze the data rigorously. On August 27, 2026, NIH released a draft policy for public comment that, if finalized, would require investigators and institutions to share plain-language, summary-level results with participants in NIH-supported clinical research, subject to defined exceptions. The Declaration of Helsinki already contains this: the concept of the participant as a partner. In Modena, through the EU-funded FACILITATE project, conducted from 2022 to 2026, which examined patient-centered and GDPR-compliant approaches to returning clinical-trial data and information to participants, , we worked specifically on returning information and data to study participants. This is something I want people to take from this conversation.

AI cannot be allowed to increase inequity. An AI-derived subgroup should not automatically become an exclusion rule. A sponsor may study only an enriched population when the scientific evidence and regulatory strategy justify it, but that decision must be prospective and must explicitly address assay performance, uncertainty, unmet need, safety, generalizability and what can or cannot be inferred about patients outside the selected group. We do not have sufficient knowledge to absolutely exclude any fellow human from the possibility of responding to a treatment, just because a machine says no. That might happen. The solution is to stratify based on the NetraMark analysis, define a stratum that is representative enough, pre-specify your primary analysis, and then analyze all comers. Nor can label breadth be inferred simply because placebo was not superior to drug in the non-selected stratum. Absence of evidence of harm is not evidence of efficacy. Labeling depends on the totality of prospectively specified evidence, including the treatment effect and its precision within each population, any treatment-by-subgroup interaction, safety, benefit-risk, subgroup prevalence and the population actually studied.

You cannot say the machine told you to do it. There is a human in the loop. That person is taking responsibility and signing on to it. In NetraAI, humans control, check, and challenge all of it, and it is going to stay that way.

The average patient is a statistical construct. It is not a biological reality. The future of clinical development is not to abandon randomized trials or replace statistics. It is to make those trials more intelligent: identify heterogeneity earlier, translate it into a transparent hypothesis, test that hypothesis prospectively, and go to the FDA early and often. Patient selection is a lever. It is not destiny.

“The average patient is only a statistical construct. It’s not a biological reality.”

Dr. Luca Luca Pani is Chief Innovation and Regulatory Officer at NetraMark. He previously served as Director General of the Italian Medicines Agency and as a member of the European Medicines Agency Management Board, the Committee for Medicinal Products for Human Use and the Scientific Advice Working Party.


Website |  + posts

Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.