Picture a statistician at EMD Serono’s clinical operations team opening the unblinded Cohort B data from the WILLOW trial in early 2026. The BICLA response rates are moving in the right direction. Enpatoran, the company’s small-molecule TLR7/8 inhibitor, is doing something measurable to disease activity in patients with moderate-to-severe systemic lupus erythematosus. And yet the primary endpoint — a statistically significant dose-dependent improvement — has not been met. The drug works. The trial doesn’t prove it does. These are not the same outcome, and confusing them costs programs years.
The Cohort B results, published in The Lancet, show that enpatoran improved BICLA response rates versus placebo across all dose groups and was well tolerated. But the trial’s primary objective — demonstrating a statistically significant dose-dependent effect on disease activity based on BICLA response — was not achieved. Most of the coverage will frame this as a qualified success: promising signal, tolerability intact, program lives to fight another day. That framing misses the harder lesson sitting inside the protocol design itself.
The real story here is about how dose-finding studies in autoimmune indications set themselves up to fail — not because the biology is wrong, but because the statistical architecture demands something the disease won’t deliver cleanly.
When the Endpoint Architecture Outsmarts the Biology
A dose-dependent primary endpoint in an SLE trial sounds rigorous. In regulatory logic, it is: demonstrating that higher doses produce higher response rates is the cleanest possible proof of mechanism. But SLE is not a linear disease. Its heterogeneity — driven by dysregulated TLR7/8 signaling that varies substantially across patient subpopulations — means that dose-response relationships in immunology rarely produce the tidy monotonic curves that a trend-test primary endpoint requires. When you build your primary endpoint around dose-linearity in a disease defined by immunological chaos, you are asking the biology to conform to the statistics rather than the reverse.
This is not a new hazard. Look at what happened with anifrolumab’s Phase IIb MUSE study. In that trial, 305 patients were randomized to anifrolumab 300 mg, 1000 mg, or placebo, with a primary endpoint anchored to SRI-4 response at Week 24 combined with glucocorticoid reduction. The MUSE results showed that the 300 mg dose outperformed the 1000 mg dose on the primary composite — a counter-intuitive finding that defied the dose-escalation logic built into the protocol. AstraZeneca carried the program forward anyway, ultimately reaching a successful Phase 3 TULIP-SC readout and FDA authorization of a subcutaneous anifrolumab autoinjector for moderate-to-severe SLE. The MUSE “failure” was a design artifact, not a biological one. EMD Serono is now staring at the same type of artifact in WILLOW Cohort B.
The BICLA endpoint itself is not the problem. The FDA has accepted BICLA as a clinically meaningful SLE response measure — the anifrolumab subcutaneous approval, based on the TULIP-SC trial involving 367 adults, used BICLA as a primary driver of regulatory persuasion. BICLA captures meaningful disease improvement: a reduction in all elevated BILAG-2004 organ domain scores, no worsening in SLEDAI-2K, no increase in physician global assessment, and no treatment failure. That composite is defensible. But anchoring a dose-finding trial’s primary endpoint to whether BICLA response rates scale monotonically with dose introduces a layer of statistical constraint that the immunological reality of SLE may simply refuse to honor.
What the WILLOW Cohort B data actually tells us — that enpatoran produces measurable BICLA improvements versus placebo and is well tolerated across dose groups — is precisely the information a Phase 2 dose-finding trial should generate. The clinical signal is there. The design failed to capture it in a statistically falsifiable form. That gap between signal and proof is where drug programs die at Phase 3 transition, because sponsors either over-interpret a promising Phase 2 or under-power the Phase 3 based on effect sizes that the Phase 2 protocol was never equipped to estimate cleanly.
The Adaptive Design Question Nobody Is Asking
Here is the counterintuitive reality that should be driving every conversation about this readout: a trial that fails its primary endpoint while generating positive directional data across all arms is not evidence that the drug failed. It is evidence that dose-finding methodology in autoimmune disease has not kept pace with what we know about immunological variability.
The FDA’s own guidance on adaptive designs — specifically the 2019 guidance on adaptive designs for clinical trials of drugs and biologics — explicitly endorses response-adaptive randomization and Bayesian dose-finding frameworks as tools for navigating exactly this kind of biological heterogeneity. Bayesian model-averaging approaches allow sponsors to extract dose-response estimates even when the relationship is non-monotonic, without pre-committing the primary endpoint to a linear trend test that the disease may not satisfy. MCP-Mod — multiple comparison procedures with modeling — was developed precisely to handle dose-response uncertainty in early-phase trials while preserving Type I error control. Neither approach requires you to bet the primary endpoint on dose linearity.
The SLE treatment market is projected at $3.65 billion in 2024, growing to $5.75 billion by 2030 at a 7.9% CAGR. The commercial pressure to move programs through Phase 2 efficiently is substantial. But the conventional dose-finding structure — fixed arms, trend-test primary, pre-specified monotonicity assumption — was engineered for oncology dose escalation logic, not for the immunological complexity of a disease where patients with identical SLEDAI scores can have radically different interferon signatures, complement levels, and TLR pathway activation states. Importing that structure into SLE trials without modification is a source of systematic, preventable Phase 2 failure.
EMD Serono’s situation deserves specific attention here. Enpatoran targets TLR7 and TLR8 simultaneously — a mechanistic approach grounded in the finding that TLR7 overexpression accelerates lupus-like autoimmunity while dual inhibition may suppress the type I interferon cascade more completely than a single-receptor strategy. That is a rationally designed mechanism operating in a validated pathway. The Cohort A data in cutaneous lupus and SLE with active lupus rash provided the earlier signal. Cohort B extended that into the moderate-to-severe systemic population. The tolerability profile is intact. None of that changes because a trend test didn’t reach significance.
What changes is the Phase 3 strategy, and that is where the next set of decisions will determine whether this program becomes an approved therapy or a cautionary footnote.
The Phase 3 Decision That Will Define This Program
The immediate risk for EMD Serono — and for every SLE sponsor watching this readout — is the temptation to treat the Cohort B primary endpoint miss as a dose-selection problem rather than a protocol architecture problem. If the Phase 3 is designed by simply picking the dose arm that looked best in Cohort B and running a larger conventional trial, the underlying endpoint vulnerability travels forward into the pivotal study. A Phase 3 primary endpoint miss in SLE, with a well-tolerated drug and a validated mechanism, would be a failure that an improved Phase 2 design could have prevented.
The more defensible path is an early Type B meeting with FDA to discuss what the Cohort B data actually demonstrates — positive BICLA directionality, dose-range tolerability — and to co-design a Phase 3 endpoint structure that does not re-import the dose-linearity assumption into a pivotal setting. The TULIP program’s trajectory after MUSE is the template: acknowledge the Phase 2 design limitations, reframe the evidentiary value of what you have, and build the Phase 3 around what the biology actually showed rather than what the statistical model assumed it would show.
The broader lesson sits above any single program. The SLE trial design community has now watched multiple well-resourced, mechanistically coherent programs generate encouraging Phase 2 signals that their own endpoint architectures couldn’t capture. Every protocol team designing a dose-finding study in an immunologically heterogeneous indication should be asking, before the first patient is dosed, whether the primary endpoint’s statistical assumptions are compatible with what the disease is actually capable of demonstrating. Because if the biology and the endpoint are operating on different logic, the data will always tell a story that the statistics can’t confirm — and that gap will cost far more than a protocol revision would have.
Enpatoran may yet reach patients with moderate-to-severe SLE. The mechanism is sound, the tolerability is real, and the BICLA signal is directionally meaningful. But the path forward runs through a frank internal reckoning with what WILLOW Cohort B actually revealed — not about the drug, but about the trial design that was asked to prove it.
References
- The Lancet — “Enpatoran, a Toll-like receptor 7/8 inhibitor, in moderate-to-severe systemic lupus erythematosus: findings from Cohort B of a multicentre, international, double-blind, placebo-controlled dose-finding phase 2 trial”
- ResearchGate — “Randomized Placebo-Controlled Phase II Study of Enpatoran, a Small Molecule TLR7/8 Inhibitor, in Cutaneous Lupus Erythematosus: Results from Cohort A”
- PMC / MUSE Study — “Anifrolumab Phase IIb MUSE Trial: Dose-Finding Results and SRI-4 Endpoint Analysis”
- AstraZeneca — “Saphnelo Met Primary Endpoint in TULIP-SC Phase 3 Trial”
- BioPharm International — “FDA Clears Subcutaneous Anifrolumab Autoinjector for Moderate-to-Severe SLE”
- ACS Journal of Medicinal Chemistry — “TLR7/8 Inhibitor Mechanism and Molecular Processing in SLE Pathogenesis”
- Strategic Market Research — “Systemic Lupus Erythematosus Treatment Market Size, Share, and Forecast 2024–2030”
Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.

