Intrinsic DNA sequence determinants and tissue-specific regulation of human replication origins

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-16

A transformer model and a mutational asymmetry method identified intrinsic DNA sequences and chromatin accessibility as key determinants of constitutive and tissue-specific DNA replication origins in humans.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-16 · read from full text

The paper investigates what determines where DNA replication initiates in human cells, developing ORIFormer, a transformer model that learns a sequence grammar to predict replication initiation sites from intrinsic DNA features. The authors validate newly identified sequence determinants using genetic variants that alter motifs, showing selection and molecular effects of those changes, and they develop MuSAS to map tissue-specific initiation zones by exploiting mutational strand asymmetries from processes like APOBEC activity and mismatch repair deficiency across somatic tissues. Integrating intrinsic sequence-based predictions with tissue-specific chromatin accessibility, they find that sequence features define a high-potential landscape of constitutive origins while chromatin accessibility acts as a permissive switch for tissue-specific usage. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

The accurate duplication of the genome relies on the spatiotemporal control of DNA replication initiation, yet the determinants specifying mammalian origin locations remain elusive. By developing ORIFormer, a transformer-based neural network, we decode a complex, conserved DNA sequence grammar that accurately predicts initiation sites. This approach uncovers novel determinants, which we validate by demonstrating selection and molecular effects of motif-altering genetic variants. To characterize tissue-specific usage, we developed MuSAS, a statistical genomic method leveraging widespread mutational strand asymmetries, such as from APOBEC and mismatch repair deficiency, to map initiation zones across diverse somatic tissues. Integrating these modalities reveals that while intrinsic DNA sequence features establish a high-potential landscape of constitutive origins, tissue-specific usage is governed by local chromatin accessibility acting as a permissive switch. We suggest that human replication initiation is driven by a deterministic genetic code modulated by the epigenetic landscape, providing a unified framework for understanding genome copying.
Full text 1,249 characters · extracted from oa-doi-fallback · click to expand
Abstract The accurate duplication of the genome relies on the spatiotemporal control of DNA replication initiation, yet the determinants specifying mammalian origin locations remain elusive. By developing ORIFormer, a transformer-based neural network, we decode a complex, conserved DNA sequence grammar that accurately predicts initiation sites. This approach uncovers novel determinants, which we validate by demonstrating selection and molecular effects of motif-altering genetic variants. To characterize tissue-specific usage, we developed MuSAS, a statistical genomic method leveraging widespread mutational strand asymmetries, such as from APOBEC and mismatch repair deficiency, to map initiation zones across diverse somatic tissues. Integrating these modalities reveals that while intrinsic DNA sequence features establish a high-potential landscape of constitutive origins, tissue-specific usage is governed by local chromatin accessibility acting as a permissive switch. We suggest that human replication initiation is driven by a deterministic genetic code modulated by the epigenetic landscape, providing a unified framework for understanding genome copying. Competing Interest Statement The authors have declared no competing interest.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: oa-doi-fallback

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00