Picture a biostatistician at a mid-size oncology sponsor sitting with a Phase II adaptive design protocol for a PD-L1 inhibitor in bladder cancer. Her enrichment strategy relies on PD-L1 expression cutoffs established in trials from three years ago, cutoffs that her own team knows are crude and that the FDA has quietly flagged in two prior Type B meetings as “insufficient to exclude non-responders at enrollment.” She has read the literature on transcriptomic signatures. She knows TMB alone captures maybe 15% of the variance in checkpoint inhibitor response. But she also knows the FDA has never cleared a bulk RNA-seq assay as a companion diagnostic for a PD-1 indication, and she is not willing to bet her IND on a biomarker strategy with no regulatory anchor. So she keeps the PD-L1 cutoff, enrolls broadly, and accepts that a third of her trial population will almost certainly not respond. This is where oncology trial design lives in 2026: knowing better, doing worse.

That tension now has a new pressure point. Researchers at Harvard Medical School have published COMPASS in Nature Medicine, a pan-cancer foundation model trained on transcriptomes from 10,184 tumors spanning 33 cancer types from the Cancer Genome Atlas, then fine-tuned on 16 immune checkpoint inhibitor clinical cohorts covering seven cancers and six ICI regimens including anti-PD-1, anti-PD-L1, anti-CTLA4, and combination therapies. In leave-one-cohort-out validation across 1,133 patients, COMPASS outperformed 22 existing prediction methods by an average of 8.5% in prediction accuracy. That number deserves a pause.

Eight and a half percent, averaged across cancer types and ICI classes, represents the difference between a trial that stratifies well and one that dilutes its own signal with non-responders. It represents the difference between a hazard ratio that crosses the pre-specified boundary and one that misses by two tenths of a point. Anyone who has sat through a late-stage oncology futility analysis knows exactly what that margin means.

The Stratification Problem the Field Has Learned to Ignore

The structural failure in checkpoint inhibitor development is not a secret, even if it is rarely stated plainly. Published data in PubMed show that early-phase PD-1/PD-L1 trials systematically overestimated efficacy compared to subsequent Phase III results, with an odds ratio of 1.66 (95% CI: 1.43 to 1.92) comparing overall response rates between phases. That gap does not arise primarily from biology. It arises from patient selection. Early cohorts skew toward patients who happen to be biomarker-positive by whatever crude measure the sponsor used at enrollment. Phase III catches the full population and the signal dilutes. Sponsors know this going in. They build it into their sample size assumptions, call it “biomarker heterogeneity,” and move forward anyway, because the alternative, an enriched trial built on a transcriptomic signature that has no regulatory clearance pathway, feels riskier than the known inefficiency.

That regulatory gap is real and specific. Illumina’s TruSight Oncology Comprehensive test received FDA approval in 2024 as the first distributable comprehensive genomic profiling IVD kit with pan-cancer companion diagnostic claims. That approval covers DNA-level alterations, copy number changes, and selected fusions. It does not cover bulk RNA-seq expression signatures of the kind COMPASS runs on. The regulatory architecture for transcriptomic biomarkers in ICI trials simply does not exist at the pan-cancer level, which means a sponsor who wants to use COMPASS as an enrichment tool today would need to pursue a novel companion diagnostic development program in parallel with their IND, a two-to-four year detour that most trial timelines cannot absorb.

The FDA’s own framework for biomarker-driven enrichment, articulated in its 2019 Enrichment Strategies for Clinical Trials guidance, explicitly encourages “prognostic and predictive enrichment” as a means of improving trial efficiency. The guidance even addresses the scenario where a validated assay does not yet exist, noting that sponsors may use “bridging studies” to qualify biomarkers during early development. But the gap between the guidance’s encouragement and the operational reality of qualifying a bulk transcriptome signature as a CDx for six different ICI classes, across seven tumor types, is not a paperwork problem. It is a resourcing and precedent problem, and COMPASS does not resolve it alone.

What COMPASS Actually Changes at the Protocol Level

Now consider a different sponsor, this one running a Phase I/II basket trial for a novel CTLA4 inhibitor across four tumor types. His team is not trying to use COMPASS as a regulatory-grade CDx. They want it as an exploratory biomarker, collected prospectively, analyzed at interim, used to inform a response-adaptive randomization rule that up-weights enrollment to tumor types where early transcriptomic profiles suggest higher responder concentrations. Under the FDA’s 2019 adaptive design guidance and the 2020 guidance on Adaptive Designs for Clinical Trials of Drugs and Biologics, that kind of prospectively pre-specified biomarker-adaptive rule is permitted, provided the adaptation is specified in the statistical analysis plan before any unblinded data are reviewed. COMPASS, used in this framing, is not a companion diagnostic. It is an adaptive allocation algorithm. That distinction changes everything about the regulatory strategy.

The COMPASS architecture makes this operationally plausible in a way that prior transcriptomic tools did not. Because the model was pre-trained on 33 cancer types and fine-tuned on 16 ICI cohorts, its predictions generalize across tumor contexts rather than being locked to a single indication. A basket trial spanning melanoma, NSCLC, and urothelial carcinoma does not need to qualify three separate biomarker assays. One model, one analytical pipeline, one pre-specified allocation rule. The oncology segment of the AI-in-clinical-trials market already captured 45.9% of total market share in 2024, per market analysis data, and the pressure to operationalize tools like COMPASS is not theoretical. Sponsors are actively building these workflows right now, often without regulatory pre-alignment.

That last point is the one that should concern trial operations teams most. The history of novel biomarker strategies in oncology is littered with adaptive designs that were elegant on paper and inoperative at sites. A bulk RNA-seq pipeline requires fresh or flash-frozen tumor tissue, not the FFPE samples that most community oncology sites collect by default. The 16 clinical cohorts in COMPASS’s validation set were not assembled from community practice. They were curated research cohorts with controlled tissue handling. Deploying the model in a real-world trial network means solving a pre-analytical problem that the Nature Medicine paper does not address and that most CROs are not yet equipped to manage at scale.

The Regulatory Proof of Concept That Has Not Arrived Yet

There is a counterintuitive conclusion hiding in the COMPASS data that the field’s initial reaction has mostly missed. The common assumption is that better predictive accuracy makes trial design simpler: stratify better, enroll fewer patients, hit your endpoints faster. But an 8.5% accuracy improvement averaged across 22 comparator methods means COMPASS is better than everything currently used, including the methods FDA has accepted in approved companion diagnostics. If a sponsor uses COMPASS prospectively and the trial succeeds, they now have a biomarker that outperforms the regulatory gold standard embedded in their approved product’s label strategy. That creates an obligation to validate it as a CDx for the commercial indication, not just an exploratory marker. The model’s strength becomes a post-approval commitment the development team did not budget for.

The closest regulatory analogy comes from a different technology class. Onc.AI’s Serial CTRS tool, an imaging-based deep learning prognostic for NSCLC patients on immunotherapy, received FDA Breakthrough Device Designation based on its ability to stratify high- and low-risk mortality categories from diagnostic CT scans. That designation opened a formal development pathway but did not constitute clearance or approval. It created a structured dialogue with the FDA about what analytical validation, clinical validation, and intended use labeling would need to look like. COMPASS needs exactly that kind of pre-competitive engagement with the FDA’s Oncology Center of Excellence, ideally through a Biomarker Qualification submission under the 2023 Biomarker Qualification Program framework. No such submission appears to be in process. Without it, every sponsor who integrates COMPASS into a trial design is building on a foundation whose regulatory status will need to be litigated case by case, at the worst possible moment in the development timeline.

The practical directive for sponsors running ICI trials right now is specific. Request a Type C meeting with FDA before incorporating any COMPASS-based allocation or enrichment rule into an IND. Pre-specify the analytical pipeline, the tissue handling requirements, and the decision rules in the statistical analysis plan with enough granularity that the FDA can evaluate the adaptation rules without needing to understand the model architecture in real time during review. Include a tissue feasibility assessment in the protocol’s operational planning documents that addresses FFPE versus fresh tissue requirements at each site tier. And if the Phase II signal is strong enough to warrant a Phase III confirmatory design, engage the FDA’s Division of Oncology Products on CDx co-development requirements before the Phase II locks its database, not after.

Which brings the story back to the biostatistician with her bladder cancer protocol and her inadequate PD-L1 cutoff. COMPASS gives her a better tool. What it does not give her, at least not yet, is a regulatory pathway to use it without adding eighteen months and a parallel diagnostic development program to her timeline. That gap between what the science can do and what the regulatory infrastructure will accept is the defining operational tension in oncology trial design right now, and COMPASS has just made it impossible to pretend otherwise.

References

  1. Nature Medicine — “Generalizable AI predicts immunotherapy outcomes across cancers and treatments”
  2. First Word Pharma — “COMPASS improves immunotherapy prediction accuracy by 8.5% across 22 existing methods”
  3. AI Weekly — “Harvard’s COMPASS predicts immunotherapy response across cancers; trained on 10,184 tumors across 33 cancer types”
  4. PubMed — “Early-phase PD-1/PD-L1 trial overestimation of efficacy versus Phase III; OR 1.66 (95% CI: 1.43–1.92)”
  5. Illumina — “FDA approves TruSight Oncology Comprehensive as first pan-cancer companion diagnostic IVD kit (2024)”
  6. Targeted Oncology — “Onc.AI Serial CTRS receives FDA Breakthrough Device Designation for NSCLC immunotherapy risk stratification”
  7. Market.us — “Oncology segment captures 45.9% of AI in clinical trials market share in 2024”
Website |  + posts

Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.