Extending Inferences From A Real-world Clinico-Genomic Database Non-Small Cell Lung Cancer Sample To The Underlying Target Population Using Combined Information From Cancer Registrations

preprint OA: closed
📄 Open PDF View at publisher

Abstract

ABSTRACT Purpose The study aimed to extend inferences on the Average Treatment Effects (ATEs) from a Flatiron Health Clinico-Genomic Database (CGDB) Non-Small Cell Cancer (NSCLC) sample to the target population represented by SEER cancer registrations. The work also demonstrates the potential for the non-random selection to cause bias through a quantitative framework that compares the marginal and joint distributions of effect modifiers through each non-random selection process. Methods ATEs for a binary treatment were estimated within the sample (SATE) and extended to the population (PATE) using combined information from SEER and a weighted estimator to standardize the joint distributions of effect modifiers. To understand potential biases through selection, the marginal and joint distributions of effect modifiers were compared through each stepwise process using two referent populations: SEER registrations and a superset of all NSCLC patients in the Flatiron Health network. Results Within a subset of 1,166 stage III-IV NSCLCs receiving a binary treatment and combined information from 149,056 SEER registrations, the SATE & PATE for differences in survival at month 48 were −3.7 (−8.7, 1.6) & −3.4 (−8.7, 2.8) percentage points. Through each sequential selection, the joint distributions of effect modifiers were not discernibly different among the selected & unselected. Estimates of survival were unbiased by selection. Conclusions Combined information from cancer registrations can be used to extend inferences from a selected sample to the target population. ATEs within a CGDB were an unbiased estimate of the population because the sequential selection did not differentially select effect modifiers causative of survival. Key Points Combined information from cancer registrations can be used to extend inferences from a selected sample to the target population We outline a quantitative framework for determining the potential for non-random selection to cause bias, through comparing the marginal and joint distributions of effect modifiers through each sequential selection process Average Treatment Effects within a highly selected genetic cohort were an unbiased estimate of the population because the sequential selection did not differentially select effect modifiers causative of survival Plain Language Summary In pharmacoepidemiology we oftentimes learn about the effectiveness of treatments within a smaller sample in attempt to understand how they would work in the broader population. When this sampling is non-random — like when we use Real-World Data in the form of Electronic Medical Records or insurance claims — what we learn in this sample may not translate to the broader population. In this study we show how we can publicly available data in the form of cancer registrations to better understand how these treatments work in the population. We show that what we learn about treatments within our highly selected sample that required patients to undergo expensive genetic testing does in fact translate to the population. We also provide a framework showing why this was the case: patients in our selected sample were very similar in their characteristics to the broader population.

My notes (saved in your browser only)

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00