Picture a statistician at EMD Serono staring at a dose-response curve that refuses to climb. The WILLOW Cohort B dose-response BICLA data, reported in the primary publication, showed a response pattern across ascending doses that plateaued rather than climbed, with placebo lagging behind the lowest active dose. The primary MCP-mod analysis returned p=0.14, and WILLOW Cohort B was declared a failure. By any reasonable pharmacological reading, the drug did something. But the primary multiple comparison procedure-modeling analysis on those four numbers returned p=0.14, and WILLOW Cohort B was declared a failure.

That failure declaration will follow enpatoran into every future regulatory conversation. And it rests almost entirely on a statistical framework built for monotonic dose-response curves applied to a pharmacodynamic signal that had already plateaued.

The Lancet correspondence on Cohort B puts the essential question plainly: WILLOW separated two distinct biological questions that lupus trial designers routinely conflate. Did enpatoran alter the disease biology? Almost certainly yes. Was the BICLA composite endpoint structured to detect that alteration at a plateau? Provably no.

What the Curve Is Actually Saying

Enpatoran is a first-in-class, selective, orally administered TLR7/8 inhibitor. Toll-like receptor 7 activation sits upstream of the type I interferon cascade that drives much of the organ damage in SLE, which means enpatoran’s pharmacodynamic effect operates at a molecular chokepoint. Receptor occupancy at a chokepoint follows a saturation curve: at some dose, you have blocked enough receptor activity that additional drug produces no additional downstream suppression. The response plateau at 58-49-49% across ascending doses is not a failure signal. It is the fingerprint of exactly that mechanism.

A Phase Ib randomized, double-blind, placebo-controlled trial of enpatoran in 25 patients with active SLE or cutaneous lupus erythematosus established the drug’s pharmacokinetic and pharmacodynamic profile before WILLOW launched. That earlier work should have signaled the design team: if the pharmacodynamic effect saturates early, the dose-response model must accommodate non-monotonic patterns or the primary analysis will be structurally blind to efficacy. The MCP-mod framework used in WILLOW Cohort B requires a monotonic dose-response relationship to generate statistical significance. Feed it a plateau and it returns noise.

This is not a subtle methodological wrinkle. It is a foundational mismatch between mechanism and measurement.

The Endpoint Graveyard Lupus Built

Lupus has one of the worst Phase III success rates in clinical pharmacology, and the pattern is consistent enough to be diagnostic: Phase II data signals activity, Phase III endpoints fail to confirm it, and the drug gets shelved. Researchers have attributed this to disease heterogeneity, but heterogeneity does not explain why Phase III fails at higher rates than Phase II for the same compounds. Endpoint misconstruction does.

BICLA, the British Isles Lupus Assessment Group-based Composite Lupus Assessment, was designed as a broad composite to capture improvement across multiple organ systems simultaneously. That breadth is clinically sensible for drugs that work through general immune suppression. But enpatoran targets a specific upstream molecular pathway. A composite endpoint that averages across organ domains will dilute a strong signal in TLR7-driven disease manifestations with null signals in domains where that pathway is less active. The 58% BICLA response at the lowest dose suggests the drug hit hard in the domains it was designed to affect. The composite average absorbed that hit and smoothed it into statistical ambiguity.

AstraZeneca faced a structurally similar challenge with anifrolumab before landing on an endpoint strategy that worked. The TULIP-LN Phase II trial evaluated anifrolumab in 147 patients with active proliferative lupus nephritis using the urine protein-creatinine ratio at Week 52 as its primary endpoint, a direct mechanistic biomarker of renal involvement rather than a broad composite. That mechanistic specificity gave the statistical analysis something to find. WILLOW gave MCP-mod a composite that was built for a different class of drug.

WILLOW Cohort A reinforces the point from the other direction. Cohort A met its primary endpoint, demonstrating statistically significant improvement in lupus skin manifestations in patients with cutaneous lupus erythematosus or SLE with active rash. Cutaneous disease is among the most TLR7-dependent manifestations of lupus. An endpoint designed for that specific domain detected the drug. A composite spanning every organ system did not.

The Regulatory Architecture Nobody Wants to Redesign

The FDA’s guidance on developing medical products for SLE treatment acknowledges BICLA and SRI (SLE Responder Index) as acceptable primary endpoints, and sponsors have largely accepted that framing as a ceiling rather than a floor. The guidance exists to provide a pathway, not to mandate the best-available endpoint for every mechanism class. Sponsors who treat regulatory acceptability as a proxy for scientific adequacy are making a category error that costs them trials.

The FDA’s guidance on multiple endpoints in clinical trials is explicit that post-hoc analyses of trials that fail on their prospectively specified primary endpoints generate hypotheses for future studies rather than evidence for current approval. That is the correct regulatory principle. But it places the full burden of endpoint validation on the pre-trial design phase, which means the cost of a WILLOW-style mismatch lands entirely on the sponsor before a single patient is enrolled.

Pre-IND and Type B meeting requests exist precisely to pressure-test endpoint choices against mechanism before that cost is sunk. A sponsor with enpatoran’s Phase Ib pharmacodynamic data and a proposed MCP-mod analysis framework had the raw material to surface this mismatch with FDA reviewers before Cohort B enrolled. The question worth asking is whether that conversation happened with sufficient specificity: not “is BICLA acceptable?” but “does MCP-mod applied to BICLA have adequate sensitivity to detect a receptor-saturation efficacy profile at the doses we are testing?”

Those are operationally different questions, and only the second one would have caught what happened.

For sponsors designing autoimmune trials with mechanistically targeted agents right now, the takeaway from WILLOW Cohort B is precise. Composite endpoints validated for broad immunosuppressants carry an implicit assumption about how efficacy distributes across disease domains. If your drug’s mechanism concentrates pharmacodynamic activity in a subset of those domains, the composite will underweight your signal by design. The fix is not to abandon BICLA or SRI. The fix is to pre-specify a primary endpoint, or a primary analysis model, that is sensitive to the distribution of effect your mechanism actually produces, then document that scientific rationale explicitly in the IND so reviewers understand why the standard composite was modified or supplemented.

WILLOW Cohort A enrolled patients with TLR7-relevant disease manifestations and used a domain-specific endpoint. It succeeded. Cohort B enrolled patients with systemic disease and used a broad composite with a monotonic dose-response model. It returned p=0.14. The molecule did not change between cohorts.

That 58% BICLA response at the lowest enpatoran dose will appear in the Cohort B primary publication as a secondary finding in a failed trial. A better-matched endpoint would have made it the headline. Somewhere in the next TLR7 program’s protocol, a design team is deciding whether to reach for BICLA because the FDA accepts it, or to build something that can actually see what their drug does. That decision will determine whether the molecule gets a fair trial.

References

  1. The Lancet, “Mechanism–endpoint matching in lupus: WILLOW Cohort B” (Correspondence, 2026)
  2. PubMed, WILLOW Cohort B: enpatoran dose-response BICLA analysis at week 24
  3. PubMed, Phase Ib randomized controlled trial of enpatoran in SLE/CLE patients (NCT04647708)
  4. PMC, SLE randomized clinical trial failure rates and endpoint discrepancy analysis
  5. FDA, Guidance for Industry: Systemic Lupus Erythematosus, Developing Medical Products for Treatment
  6. FDA, Multiple Endpoints in Clinical Trials: Guidance for Industry
  7. Annals of the Rheumatic Diseases, TULIP-LN Phase II trial: anifrolumab in proliferative lupus nephritis (147 patients, UPCR primary endpoint)
  8. Healio, WILLOW Cohort A: enpatoran meets primary endpoint in cutaneous lupus manifestations (June 2026)
Website |  + posts

Moe Alsumidaie, MBA, MSF, is founder and Chief Editor of Vanguard Publications, which publishes Clinical Trial Vanguard, Pharma Vanguard and BullScope, and Head of Research at CliniBiz. He has two decades in clinical trial operations and data science, with earlier roles at Genentech, Abbott Vascular and Stanford University Medical Center, and is a guest lecturer in clinical trial sciences at Rutgers University.