⚙
AI-generated deep summary
by claude@2026-07, 2026-07-05
· read from full text
ⓘ
The paper studies how to improve machine learning for extremely sparse and heavy-tailed cell-free RNA (cfRNA) data in human plasma, particularly when labeled data are scarce. It introduces cfRNA-ICL, an in-context learning model trained on synthetic tasks generated from a biologically grounded structural causal model that incorporates empirically measured cfRNA dropout, overdispersion, tissue-mixture latent factors, compositional variability, and sequencing noise. Across multiple cancer classification benchmarks, cfRNA-ICL outperforms tabular ICL models trained on generic synthetic data, with the largest gains in few-shot settings, and it forms representation-level manifolds aligned with biologically coherent structure without supervised constraints. The main limitation explicitly implied by the approach is that training relies on a synthetic task universe produced by the proposed causal model. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.
Abstract
Cell-free RNA (cfRNA) in human plasma provides a minimally invasive readout of tissue physiology, yet its extreme sparsity, heavy-tailed abundance distributions, and weak but structured correlation patterns create major challenges for machine learning. Conventional tabular foundation models are typically trained on synthetic datasets that assume generic statistical properties, and as a result, they fail to capture the distinctive characteristics of cfRNA. These limitations become even more pronounced in settings where labeled data are scarce. We introduce cfRNA-ICL , a cfRNA-specific in-context learning model trained entirely on tasks generated from a biologically grounded structural causal model (SCM) . The SCM produces realistic cfRNA-like scenarios by incorporating empirical measurements of gene-level dropout, overdispersion, tissue-mixture–driven latent factors, compositional variability, and sequencing noise. This synthetic task universe enables cfRNA-ICL to acquire inductive biases that closely reflect the geometry of real cfRNA data. Across multiple cancer classification benchmarks, cfRNA-ICL demonstrates consistently higher performance than tabular ICL models trained on generic synthetic data. The gains are most sub-stantial in few-shot settings, where the model benefits from its exposure to cfRNA-specific statistical regimes during meta-training. Representation-level analyses further show that cfRNA-ICL organizes samples into biologically coherent manifolds, preserving cancer-type identity without the use of supervised constraints. This finding indicates strong alignment between the synthetic prior and real cfRNA structure. Taken together, these results show that domain-aware generative priors can meaningfully enhance in-context learning for biological tabular data. cfRNA-ICL provides a generalizable framework for cfRNA modeling and establishes a practical path toward foundation-scale models that are intrinsically adapted to the statistical landscape of plasma cfRNA.
Full text
2,183 characters
· extracted from
oa-doi-fallback
· click to expand
Abstract
Cell-free RNA (cfRNA) in human plasma provides a minimally invasive readout of tissue physiology, yet its extreme sparsity, heavy-tailed abundance distributions, and weak but structured correlation patterns create major challenges for machine learning. Conventional tabular foundation models are typically trained on synthetic datasets that assume generic statistical properties, and as a result, they fail to capture the distinctive characteristics of cfRNA. These limitations become even more pronounced in settings where labeled data are scarce.
We introduce cfRNA-ICL, a cfRNA-specific in-context learning model trained entirely on tasks generated from a biologically grounded structural causal model (SCM). The SCM produces realistic cfRNA-like scenarios by incorporating empirical measurements of gene-level dropout, overdispersion, tissue-mixture–driven latent factors, compositional variability, and sequencing noise. This synthetic task universe enables cfRNA-ICL to acquire inductive biases that closely reflect the geometry of real cfRNA data.
Across multiple cancer classification benchmarks, cfRNA-ICL demonstrates consistently higher performance than tabular ICL models trained on generic synthetic data. The gains are most sub-stantial in few-shot settings, where the model benefits from its exposure to cfRNA-specific statistical regimes during meta-training. Representation-level analyses further show that cfRNA-ICL organizes samples into biologically coherent manifolds, preserving cancer-type identity without the use of supervised constraints. This finding indicates strong alignment between the synthetic prior and real cfRNA structure.
Taken together, these results show that domain-aware generative priors can meaningfully enhance in-context learning for biological tabular data. cfRNA-ICL provides a generalizable framework for cfRNA modeling and establishes a practical path toward foundation-scale models that are intrinsically adapted to the statistical landscape of plasma cfRNA.
Competing Interest Statement
Authors are employees or founders of Eigen Bio Inc.
Footnotes
ryan{at}eigenbioai.com
beomsoo{at}eigenbioai.com
hyunjin{at}eigenbioai.com
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.