Spectral neural approximations for models of transcriptional dynamics
preprint
OA: closed
CC-BY-4.0
Abstract
The advent of high-throughput transcriptomics provides an opportunity to advance mechanistic understanding of transcriptional processes and their connections to cellular function at an un-precedented, genome-wide scale. These transcriptional systems, which involve discrete, stochastic events, are naturally modeled using Chemical Master Equations (CMEs), which can be solved for probability distributions to fit biophysical rates that govern system dynamics. While CME models have been used as standards in fluorescence transcriptomics for decades to analyze single species RNA distributions, there are often no closed-form solutions to CMEs that model multiple species, such as nascent and mature RNA transcript counts. This has prevented the application of standard likelihood-based statistical methods for analyzing high-throughput, multi-species transcriptomic datasets using biophysical models. Inspired by recent work in machine learning to learn solutions to complex dynamical systems, we leverage neural networks and statistical understanding of system distributions to produce accurate approximations to a steady-state bivariate distribution for a model of the RNA life-cycle that includes nascent and mature molecules. The steady-state distribution to this simple model has no closed-form solution and requires intensive numerical solving techniques: our approach reduces likelihood evaluation time by several orders of magnitude. We demonstrate two approaches, where solutions are approximated by (1) learning the weights of kernel distributions with constrained parameters, or (2) learning both weights and scaling factors for parameters of kernel distributions. We show that our strategies, denoted by kernel weight regression (KWR) and parameter scaled kernel weight regression (psKWR), respectively, enable broad exploration of parameter space and can be used in existing likelihood frameworks to infer transcriptional burst sizes, RNA splicing rates, and mRNA degradation rates from experimental transcriptomic data. Statement of significance The life-cycles of RNA molecules are governed by a set of stochastic events that result in heterogeneous gene expression patterns in genetically identical cells, resulting in the vast diversity of cellular types, responses, and functions. While stochastic models have been used in the field of fluorescence transcriptomics to understand how cells exploit and regulate this inherent randomness, biophysical models have not been widely applied to high-throughput transcriptomic data, as solutions are often intractable and computationally impractical to scale. Our neural approximations of solutions to a two-species transcriptional system enable efficient inference of rates that drive the dynamics of gene expression, thus providing a scalable route to extracting mechanistic information from increasingly available multi-species single-cell transcriptomics data.
My notes (saved in your browser only)
Citation neighborhood (no data yet)
We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.
Source provenance
- europepmc
- last seen: 2026-05-19T01:45:01.086888+00:00
- unpaywall
- last seen: 2026-05-23T02:00:01.238055+00:00
License: CC-BY-4.0