Picture a regulatory affairs director at a mid-sized cardiovascular sponsor, coffee going cold beside a stack of Type C meeting transcripts, staring at a Nature Medicine validation study that just landed on her desk. The dataset: 6.4 million individuals. The geographic spread: multinational. The methodology: 44 observational studies and 18 randomized controlled trials stress-tested against two of the world’s most widely deployed cardiovascular risk equations — the American Heart Association’s PREVENT model and the European Society of Cardiology’s SCORE2. She isn’t reading it for clinical interest. She’s reading it because her label expansion strategy just collided with an FDA review division that doesn’t know how to evaluate RWE at this scale, and this paper is the closest thing to a methodological roadmap she’s seen in three years of trying.

The scale of this validation demands a pause. PREVENT was originally developed using a derivation sample of 3,281,919 adults across 25 data sets collected between 1992 and 2017, with external validation performed in an additional 3,330,085 participants. SCORE2, formally published in the European Heart Journal in June 2021, was calibrated and validated across European populations aged 40–69, estimating 10-year risk of fatal and nonfatal cardiovascular events. The Nature Medicine study doesn’t merely confirm that these equations work in their home populations — it asks a harder question: do they hold when you cross borders, health systems, and patient demographics at once?

The answer, largely, is yes. But behind the affirmation lies a quieter crisis for sponsors and regulators who haven’t yet built the institutional infrastructure to operationalize what this kind of evidence is telling them.

The Validation Gap Nobody Budgeted For

Here’s the counterintuitive truth about large-scale RWE validation: most sponsors treat it as a regulatory afterthought rather than a trial design input. The assumption is that a validated risk equation is a solved problem — you implement PREVENT or SCORE2, enrich your trial population accordingly, and move forward. What the 6.4 million-person multinational dataset exposes is that “validated” means something far more conditional than sponsors typically acknowledge in their statistical analysis plans.

A 2019 systematic review in PubMed identified 38 studies encompassing 112 external validations of Framingham risk models and pooled cohort equations — across North America (56 validations), Europe (29 validations), and Asia (25 validations). The recurring finding: calibration degrades predictably when models cross from their derivation geography into new health system contexts. Discrimination holds. Calibration fractures. The Nature Medicine validation confronts this directly by embedding both PREVENT and SCORE2 in populations where neither was originally tuned, then measuring the delta. That methodological choice — testing geographic transportability explicitly rather than assuming it — is precisely what the FDA’s draft guidance on AI in drug development, issued for comment in the past two years, gestures toward but never operationalizes in the cardiovascular context.

SCORE2, formally endorsed in the 2021 ESC Guidelines on cardiovascular disease prevention, replaced the older SCORE risk chart — a meaningful regulatory signal that even established models require generational turnover. The ESC’s willingness to retire SCORE and replace it with SCORE2 within a decade of its introduction suggests an institutional posture toward model iteration that FDA’s drug review divisions haven’t matched. At FDA, a risk stratification tool embedded in a pivotal trial protocol tends to stay embedded, regardless of whether downstream evidence has complicated its assumptions. The Nature Medicine study is precisely the kind of downstream evidence that should be triggering protocol amendment conversations — and mostly isn’t.

Consider what happens operationally when a sponsor is running a Phase 3 cardiovascular outcomes trial with enrollment criteria anchored to a PREVENT-defined risk threshold. The trial enrolls in North America and Southeast Asia simultaneously. PREVENT’s derivation data is North American-dominant. The Nature Medicine multinational validation is the first large-scale evidence that PREVENT’s performance characteristics shift — sometimes materially — when applied to non-derivation geographies. If that signal wasn’t in the sponsor’s original IND submission, it isn’t automatically incorporated into the statistical analysis plan mid-trial without a protocol amendment, an IRB notification, and a Type C meeting that may or may not surface a partial clinical hold.

That’s not a hypothetical. That’s a workflow that dozens of cardiovascular trial teams are currently navigating without a regulatory framework that tells them what to do with post-IND RWE that complicates their enrollment assumptions.

What 44 Observational Studies Reveal About RCT Design

The methodological architecture of the Nature Medicine validation is worth slowing down on, because it encodes a lesson that goes well beyond cardiovascular risk prediction. The study combines 44 observational studies with 18 RCTs — a deliberate fusion that mirrors what the FDA’s RWE framework has been nudging sponsors toward since the 21st Century Cures Act authorized RWE use in regulatory submissions. But the FDA’s posture here is understandable — and incomplete.

The agency’s guidance framework asks sponsors to justify their RWE sources by demonstrating that the data is “fit for use” — a phrase that appears in multiple FDA guidance documents but remains operationally vague in the cardiovascular outcomes context. What the Nature Medicine study demonstrates is that “fit for use” requires explicit geographic calibration testing, not just source documentation. The inclusion of 18 RCTs within the validation architecture is the critical move: it provides an internal benchmark against which the observational data can be interrogated for systematic bias. That’s not standard practice in sponsor RWE submissions. Most RWE packages submitted to FDA in support of label expansions use observational data alone, with propensity score adjustment as the primary confounding control. The Nature Medicine design is more demanding — and more defensible.

The AHA’s PREVENT equations, as published in Circulation, represent the first cardiovascular risk model to integrate cardiovascular, kidney, and metabolic health measures simultaneously — a recognition that cardiovascular risk doesn’t operate in a single-organ silo. PREVENT’s 10-year and 30-year risk estimates are designed for primary prevention treatment decisions, which means they’re not just actuarial tools. They’re enrollment instruments. When a sponsor uses PREVENT to define high-risk populations in a statin or GLP-1 trial, the equation’s calibration performance in the specific trial geography becomes a regulatory question, not just a clinical one. The Nature Medicine multinational validation makes that question answerable for the first time at this sample size.

Six-point-four million individuals is not a Phase 3 sample size. It’s a population-level signal — and the question for every cardiovascular trial team reading this paper should be whether their current protocol reflects what that signal is saying.

The Operational Reckoning

What should sponsors actually do with this? The answer is more specific than “validate your risk model in each geography.” Three concrete moves follow from the Nature Medicine methodology.

First, any cardiovascular trial enrolling outside the PREVENT derivation geography — meaning outside the North American cohort populations that anchored the original 25-dataset development pool — should include a pre-specified calibration assessment as part of the statistical analysis plan. Not a post-hoc sensitivity analysis. A pre-specified subgroup analysis that can be defended in a Type B or Type C meeting as evidence that the enrollment tool was validated for the populations being enrolled. This is the kind of addition that should be made at the IND stage, not after enrollment is underway in five countries.

Second, sponsors using SCORE2 for European enrollment should document the equation’s ESC-endorsed calibration for the specific European risk region their sites occupy — the ESC stratifies Europe into low, moderate, high, and very high cardiovascular risk regions, and SCORE2 performs differently across those strata. The 2021 ESC guidelines made this explicit. A trial enrolling in Poland and the Netherlands and treating both populations as equivalently “European” in SCORE2 terms is misapplying the model.

Third, and most structurally significant: the Nature Medicine validation‘s fusion of observational and RCT data is a design template that sponsors should be presenting to FDA as a model for hybrid evidence packages. The FDA’s draft guidance on AI in drug development acknowledges the validity of hybrid evidence architectures in principle. The 6.4 million-person validation gives sponsors a published, peer-reviewed precedent in a high-stakes therapeutic area to anchor that conversation. Use it in your Type B meeting request. Name the study. Point to the methodology.

The regulatory affairs director with the cold coffee is still staring at the paper. She now has three specific protocol amendments to draft, two meeting requests to file, and one very uncomfortable conversation with her biostatistics team about a calibration assumption they locked into the SAP eighteen months ago. The Nature Medicine validation didn’t create her problem — it made the problem visible, at 6.4 million-person resolution, in a way that can no longer be argued away in a review division meeting. Whether FDA reviewers are reading the same paper with the same urgency is a question that every cardiovascular sponsor running a multinational outcomes trial should be asking their agency contacts — directly, and on the record.

References

  1. Nature Medicine — “Multinational validation of the PREVENT and SCORE2 cardiovascular risk equations across 6.4 million individuals”
  2. American College of Cardiology — “PREVENT Equations: Development and Validation, Khan et al.”
  3. PMC / European Heart Journal — “SCORE2 risk prediction algorithms: new models to estimate 10-year risk of cardiovascular disease in Europe” (2021)
  4. PubMed — Systematic review: 112 external validations of Framingham risk models (2019)
  5. PACE CME — “2021 ESC Guidelines and SCORE2 Implementation in Clinical Practice”
  6. American Heart Association — PREVENT Calculator: Official Guidance and Clinical Application
  7. AHA Journals / Circulation — PREVENT Equations, Calibration and Discrimination Metrics
  8. American College of Cardiology — SCORE2 Journal Scan: Derivation and Validation Summary (2021)
Website |  + posts

Moe Alsumidaie is Chief Editor of The Clinical Trial Vanguard. Moe holds decades of experience in the clinical trials industry. Moe also serves as Head of Research at CliniBiz and Chief Data Scientist at Annex Clinical Corporation.