A Biologically Grounded Structural Causal Model Enables cfRNA Specific In-Context Learning

preprint OA: closed
Full text JSON View at publisher
AI-generated deep summary by claude@2026-07, 2026-07-05 · read from full text

The paper studies how to improve machine learning for extremely sparse and heavy-tailed cell-free RNA (cfRNA) data in human plasma, particularly when labeled data are scarce. It introduces cfRNA-ICL, an in-context learning model trained on synthetic tasks generated from a biologically grounded structural causal model that incorporates empirically measured cfRNA dropout, overdispersion, tissue-mixture latent factors, compositional variability, and sequencing noise. Across multiple cancer classification benchmarks, cfRNA-ICL outperforms tabular ICL models trained on generic synthetic data, with the largest gains in few-shot settings, and it forms representation-level manifolds aligned with biologically coherent structure without supervised constraints. The main limitation explicitly implied by the approach is that training relies on a synthetic task universe produced by the proposed causal model. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Cell-free RNA (cfRNA) in human plasma provides a minimally invasive readout of tissue physiology, yet its extreme sparsity, heavy-tailed abundance distributions, and weak but structured correlation patterns create major challenges for machine learning. Conventional tabular foundation models are typically trained on synthetic datasets that assume generic statistical properties, and as a result, they fail to capture the distinctive characteristics of cfRNA. These limitations become even more pronounced in settings where labeled data are scarce. We introduce cfRNA-ICL , a cfRNA-specific in-context learning model trained entirely on tasks generated from a biologically grounded structural causal model (SCM) . The SCM produces realistic cfRNA-like scenarios by incorporating empirical measurements of gene-level dropout, overdispersion, tissue-mixture–driven latent factors, compositional variability, and sequencing noise. This synthetic task universe enables cfRNA-ICL to acquire inductive biases that closely reflect the geometry of real cfRNA data. Across multiple cancer classification benchmarks, cfRNA-ICL demonstrates consistently higher performance than tabular ICL models trained on generic synthetic data. The gains are most sub-stantial in few-shot settings, where the model benefits from its exposure to cfRNA-specific statistical regimes during meta-training. Representation-level analyses further show that cfRNA-ICL organizes samples into biologically coherent manifolds, preserving cancer-type identity without the use of supervised constraints. This finding indicates strong alignment between the synthetic prior and real cfRNA structure. Taken together, these results show that domain-aware generative priors can meaningfully enhance in-context learning for biological tabular data. cfRNA-ICL provides a generalizable framework for cfRNA modeling and establishes a practical path toward foundation-scale models that are intrinsically adapted to the statistical landscape of plasma cfRNA.
Full text 2,183 characters · extracted from oa-doi-fallback · click to expand
Abstract Cell-free RNA (cfRNA) in human plasma provides a minimally invasive readout of tissue physiology, yet its extreme sparsity, heavy-tailed abundance distributions, and weak but structured correlation patterns create major challenges for machine learning. Conventional tabular foundation models are typically trained on synthetic datasets that assume generic statistical properties, and as a result, they fail to capture the distinctive characteristics of cfRNA. These limitations become even more pronounced in settings where labeled data are scarce. We introduce cfRNA-ICL, a cfRNA-specific in-context learning model trained entirely on tasks generated from a biologically grounded structural causal model (SCM). The SCM produces realistic cfRNA-like scenarios by incorporating empirical measurements of gene-level dropout, overdispersion, tissue-mixture–driven latent factors, compositional variability, and sequencing noise. This synthetic task universe enables cfRNA-ICL to acquire inductive biases that closely reflect the geometry of real cfRNA data. Across multiple cancer classification benchmarks, cfRNA-ICL demonstrates consistently higher performance than tabular ICL models trained on generic synthetic data. The gains are most sub-stantial in few-shot settings, where the model benefits from its exposure to cfRNA-specific statistical regimes during meta-training. Representation-level analyses further show that cfRNA-ICL organizes samples into biologically coherent manifolds, preserving cancer-type identity without the use of supervised constraints. This finding indicates strong alignment between the synthetic prior and real cfRNA structure. Taken together, these results show that domain-aware generative priors can meaningfully enhance in-context learning for biological tabular data. cfRNA-ICL provides a generalizable framework for cfRNA modeling and establishes a practical path toward foundation-scale models that are intrinsically adapted to the statistical landscape of plasma cfRNA. Competing Interest Statement Authors are employees or founders of Eigen Bio Inc. Footnotes ryan{at}eigenbioai.com beomsoo{at}eigenbioai.com hyunjin{at}eigenbioai.com

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00