Pan-cancer detection and typing by mining patterns in large genome-wide cell-free DNA sequencing datasets
preprint
OA: closed
Abstract
Background Cell-free DNA (cfDNA) analysis holds great promise for non-invasive cancer screening, diagnosis and monitoring. We hypothesized that mining the patterns of big datasets of shallow whole genome sequencing cfDNA from cancer patients could improve cancer detection. Methods By applying unsupervised clustering and supervised machine learning on large shallow whole-genome sequencing cfDNA datasets from healthy individuals (n=367), patients with different hematological (n=238) and solid malignancies (n=320), we identify cfDNA signatures that enable cancer detection and typing. Results Unsupervised clustering revealed cancer-type-specific sub-grouping. Classification using supervised machine learning model yielded an overall accuracy of 81.62% in discriminating malignant from control samples. The accuracy of disease type prediction was 85% and 70% for the hematological and solid cancers, respectively. We demonstrate the clinical utility of our approach by classifying benign from invasive and borderline adnexal masses with an AUC of 0.8656 and 0.7388, respectively. Conclusions This approach provides a generic and cost-effective strategy for non-invasive pan-cancer detection.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00