Sampling bias in large healthcare claims databases
preprint
OA: closed
CC-BY-4.0
Abstract
Healthcare claims databases that aggregate claims from multiple commercial insurers are increasingly being used to generate real-world evidence. These databases represent a non-random sample of the underlying population, but often little attention is paid to the inherent sampling bias within the data, and how it might affect results. As an illustrative example, we characterize variation in sampling in Optum's de-identified Clinformatics Data Mart Database (CDM) at the zip-code level in 2018, and identify socioeconomic and demographic factors associated with inclusion.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-06-04T02:00:05.705006+00:00
License: CC-BY-4.0