Dissecting the Expression Pattern during Embryonic Development and Unveiling Sustained ZGA Genes with Oncogenic Relevance

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Zygotic genome activation (ZGA) represents a pivotal transition in early embryonic development, marking the onset of gene expression following fertilization. Despite its fundamental importance, precisely determining the timing and identifying the key genes involved in ZGA remains a significant challenge. Results Based on time-course patterns from RNA-seq data spanning all developmental stages, we proposed a computational framework to identify ZGA genes and the onset of ZGA across diverse species. In mice, we identified 690 ZGA-associated genes, 119 of which were previously uncharacterized. Furthermore, we defined a pivotal gene subset termed Sustained ZGA (S-ZGA) genes, which are activated from ZGA and maintain sustained expression throughout subsequent development. Epigenetic analyses revealed that promoter accessibility and H3K4me3 enrichment are primary regulatory mechanisms for S-ZGA genes. Notably, these genes exhibit strong enrichment in tumorigenesis and metastatic processes. Conclusion This study establishes a computational framework that operates independently of prior knowledge to identify ZGA genes and precisely determining ZGA onset timing across diverse species, and defined a critical subset termed Sustained ZGA (S-ZGA) genes, which are potentially associated with tumorigenesis and metastatic processes. These findings provide a novel perspective on the molecular mechanisms underlying both normal development and cancer.
Full text 207,305 characters · extracted from preprint-html · click to expand
Dissecting the Expression Pattern during Embryonic Development and Unveiling Sustained ZGA Genes with Oncogenic Relevance | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Dissecting the Expression Pattern during Embryonic Development and Unveiling Sustained ZGA Genes with Oncogenic Relevance Han Xing, Yongjian Zhao This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9174297/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 9 You are reading this latest preprint version Abstract Background Zygotic genome activation (ZGA) represents a pivotal transition in early embryonic development, marking the onset of gene expression following fertilization. Despite its fundamental importance, precisely determining the timing and identifying the key genes involved in ZGA remains a significant challenge. Results Based on time-course patterns from RNA-seq data spanning all developmental stages, we proposed a computational framework to identify ZGA genes and the onset of ZGA across diverse species. In mice, we identified 690 ZGA-associated genes, 119 of which were previously uncharacterized. Furthermore, we defined a pivotal gene subset termed Sustained ZGA (S-ZGA) genes, which are activated from ZGA and maintain sustained expression throughout subsequent development. Epigenetic analyses revealed that promoter accessibility and H3K4me3 enrichment are primary regulatory mechanisms for S-ZGA genes. Notably, these genes exhibit strong enrichment in tumorigenesis and metastatic processes. Conclusion This study establishes a computational framework that operates independently of prior knowledge to identify ZGA genes and precisely determining ZGA onset timing across diverse species, and defined a critical subset termed Sustained ZGA (S-ZGA) genes, which are potentially associated with tumorigenesis and metastatic processes. These findings provide a novel perspective on the molecular mechanisms underlying both normal development and cancer. Multi-omics integration Zygotic genome activation Early embryo development Tumorigenesis and metastasis Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Background During the early stages of mammalian embryonic development, the fertilized egg gives rise to all cell types that form the adult organism through sequential rounds of proliferation and differentiation [ 1 ]. This period is marked by a series of coordinated events including maternal mRNA clearance, zygotic genome activation (ZGA), cell polarization, asymmetric division cycles, and the establishment of gene expression patterns governing the first and second cell fate decisions [ 2 , 3 ]. Central to these events is the maternal-to-zygotic transition (MZT) [ 4 ], marked by the shift from maternal to zygotic transcriptional control through embryo-wide ZGA [ 5 ]. In metazoan embryos, transcription is mostly silent for the initial cell divisions until ZGA is initiated [ 6 ]. ZGA manifests as two distinct transcriptional waves: a minor wave during early cleavage (producing limited coding transcripts and non-productive RNA via promiscuous transcription) and a major wave coinciding with the cell cycle arrest in many species [ 7 ]. While the timing and number of divisions preceding ZGA vary across species, the process is highly conserved and precisely timed within each species [ 8 ]. In mice, minor ZGA initiates at mid-to-late 1-cell stage, with major ZGA activating thousands of genes at the late 2-cell stage [ 9 – 13 ]. In humans, low-level transcription is detected from the zygote to the 4-cell embryo stage, and the major wave occurs at the 8-cell stage [ 12 , 14 , 15 ]. ZGA is essential for orchestrating gastrulation, germ-layer specification, and cell fate determination, with ZGA dysregulation leading to developmental failure [ 5 , 16 ]. For instance, the presence of one of the important ZGA genes, Nanog , is essential to initiate the zygotic developmental program and induces clearance of the maternal program in zebrafish, thereby preventing developmental arrest before gastrulation [ 17 ]. In mice, Nanog encodes a homeoprotein critical for maintaining pluripotency in the inner cell mass (ICM) and embryonic stem (ES) cells [ 18 ]. Combined loss of Nanog and other ZGA regulators results in embryonic lethality before gastrulation [ 17 ]. ZGA is a complicated process involving many important types of biological activation. However, Nanog is not the only regulator; numerous studies over the past few years have identified additional ZGA regulators that function within a complex hierarchical network. Notably, the homologous box protein Dux (and human DUX4) acts as a ZGA “accelerator,” enhancing ZGA during early embryonic development. When Dux is absent, activation of certain ZGA genes is delayed, and dose effects of alternative factors may lead to divergent developmental trajectories in Dux knockout embryos [ 19 ]. Unlike the partial compensation seen in Dux mutants, Obox factors simultaneously regulate both minor and major ZGA waves in mouse embryos, and their depletion leads to severe ZGA impairment [ 20 ]. Mechanistically, OBOX facilitates the “pre-loading” of RNA Polymerase II (Pol II), promoting its migration from initial sites to the promoters and enhancers of ZGA genes. Therefore, systematic identification of genes activated during ZGA is indispensable for unraveling its regulatory mechanisms and deciphering the molecular basis of early embryonic development. Current methodologies for identifying zygotic genome activation (ZGA) genes face critical limitations. Experimental approaches, such as α-Amanitin and aphidicolin treatments [ 9 , 21 – 23 ], enable precise ZGA gene detection but are resource-intensive and low-throughput [ 5 ], Conversely, data-driven methods comparing pre- and post-ZGA transcriptomes [ 24 – 26 ] facilitate high-throughput screening yet suffer from noise-induced false positives [ 27 ]. Crucially, both paradigms depend on manually defined ZGA timing, which varies substantially across species [ 28 ], resulting in inconsistent ZGA gene identification across datasets. This species-dependent timing discrepancy, coupled with poor reproducibility in existing studies, underscores the urgent need for robust, prior knowledge-independent ZGA identification frameworks. Given that ZGA represents a dynamic process rather than a discrete event, populations of ZGA gene transcripts straddle the period of ZGA [ 2 , 29 – 31 ]. Single-stage RNA-seq analyses are inherently prone to technical noise, limiting reliable gene identification. Leveraging time-series RNA-seq data, we reconstructs dynamic developmental trajectories [ 32 ] to identify ZGA genes through temporal expression patterns rather than static pre/post-ZGA comparisons. This approach enhances accuracy and robustness by capturing gene activation dynamics across developmental stages, providing mechanistic insights into embryonic genome activation [ 33 ]. Moreover, we uncovered a key subclass of ZGA genes called S-ZGA genes, defined by their activation at ZGA and sustained expression post-activation. We characterized their epigenetic signatures, revealing their potential link to clinical prognosis in multiple cancers. These findings bridge early embryonic development and oncogenesis, suggesting that S-ZGA genes, critical for developmental programs, may also contribute to tumor progression when dysregulated. Results Dissecting gene expression pattern during embryonic development ZGA onset varies significantly across species, posing challenges for manual determination. To dissect ZGA features, we first investigated the gene expression pattern during mouse embryonic development. We obtained gene expression data from 8 mouse early embryonic periods from a previous study [ 26 ] and 8 mouse different tissues with two replicates for each tissue from the ENCODE project and performed hierarchical clustering. Gene expression exhibits three distinct clusters corresponding to: (1) oocyte to early 2-cell embryos, (2) 2-cell to embryonic stem cell (ESC) and (3) ESC to somatic tissue (Fig. S1 A). The transition from first to second cluster well aligns with known ZGA onset at the 2-cell embryo in mice [ 25 ], and also confirmed by PCA analysis (Fig. S1 B). This phenomenon is not an isolated occurrence in mice, but conserved across different species including human [ 34 ], rhesus monkey [ 35 ] and bovine [ 36 ], as their RNA-seq data showed distinct expression patterns (Fig. 1 A), which are consistent with previous studies reporting that in these species ZGA occurs [ 27 , 34 – 36 ]. These findings collectively demonstrate that species-specific ZGA timing is reliably reflected in developmental RNA-seq expression patterns, establishing a robust foundation for automated ZGA event identification. Inspired by these findings, we analyzed the time-series RNA-seq data (Fig. 1 B) to precisely identify ZGA onset time and key genes. This framework begins with hierarchical clustering of developmental samples based on gene expression profiles-a standard approach for quantifying sample similarity in biological studies, and clustered cells from different developmental stages based on similar expression profiles. Gene expression measured by RNA-seq is often noisy due to technical error and computational bias [ 37 , 38 ] and will affect the accuracy of current ZGA gene identification methods, which are mainly based on the fold change of expression at two developmental stages. To overcome this limitation, we implement a three-layer filtering pipeline that analyzes gene expression across all developmental stages. The framework thus categorizes genes into eight distinct classes based on their expression patterns: class 0 genes show ubiquitous expression in early embryos and adult tissues (likely essential housekeeping genes), while class 7 genes remain largely silent throughout embryonic development (Fig. S1 C). Critically, genes in classes 1–6 exhibit dynamic expression patterns coinciding with ZGA transitions, making them the primary focus for identifying ZGA-specific regulatory genes. In mice, we identified six gene classes (89, 195, 316, 571, 119, and 1383 genes, respectively) with developmental stage-specific expression patterns (Fig. S1 D, supplementary table 1 ). Classes 1 and 2 characterized maternal factors. Class 3 genes exhibited minor ZGA signatures, while classes 4 and 5 defined major ZGA, and class 6 genes were exclusively expressed in differentiated tissues (Fig. 1 C). Within the complex spatiotemporal transcriptional network, embryos employ an evolutionary strategy of gene co-option [ 39 ], utilizing epigenetic mechanisms (e.g., DNA methylation, histone modifications) to differentially manage these maternal factors[ 40 , 41 ]. Specific regulators are degraded and permanently silenced, marking the irreversible exit from oogenic programs, whereas a distinct subset of maternal pioneer factors exhibits a dynamic “degrade-silence-reactivate” trajectory; initially deposited to prime the zygotic genome, they are actively suppressed post-ZGA to prevent ectopic activity, yet their circuits are preserved in a poised epigenetic state for future recruitment in specific lineages. Accordingly, we further subdivided these maternal factors into those permanently silenced after ZGA and those reactivated in mature tissues, distinguishing between “one-time developmental components” and “multifunctional developmental components” [ 42 , 43 ]. We designated minor-ZGA (class 3) and major-ZGA (classes 4 and 5) genes as ZGA genes thereafter. Functional enrichment analysis via DAVID revealed: class 3 (minor-ZGA) were mainly enriched in cell division, chromosome segregation and meiotic cell cycle, class 2 (maternal factors) were principally implicated in regulation of inflammatory response and oocyte maturation, class 4 (major-ZGA) were mainly enriched in protein transport and apoptotic process, and class 6 (tissue-specific genes) were predominantly associated with more complex biological processes, such as the angiogenesis, inflammatory response, and extracellular matrix organization (Fig. 1 D). Identified ZGA genes demonstrate critical repression during α-Amanitin treatment α-Amanitin is a specific inhibitor of RNA polymerase II (RNA Pol II), which mediates transcription of protein-coding genes. In preimplantation embryos, α-Amanitin effectively blocks zygotic transcription without globally disrupting maternal mRNA stability. To assess whether the genes identified by our framework require de novo transcription, we analyzed publicly available RNA-seq data from control and α-amanitin-treated 2-cell embryos.[ 44 ]. ZGA-associated genes (classes 3, 4, and 5) exhibited significant downregulation upon α-Amanitin treatment (Wilcoxon signed-rank test, p < 0.001; Fig. 2 A), whereas non-ZGA genes (classes 1, 2, and 6) showed no such reduction. Next, we collected the scRNA-seq data for samples in 2-cell embryos and 2-cell embryos with α-Amanitin. High inter-replicate expression correlation confirmed data reliability and absence of outliers (Fig. S2A), and further analysis of these scRNA-seq datasets revealed that only major ZGA genes (classes 4 and 5) were substantially downregulated (> 20-fold) by α-Amanitin (Fig. S2B), consistent with bulk RNA-seq results. These results demonstrate that ZGA genes our framework identified are transcriptionally activated during zygotic genome activation. It is worth noting that α-amanitin sensitivity is not absolute for all zygotic transcripts. Certain ZGA components exhibit partial resistance-a minority of genes in classes 4 and 5 retained residual expression upon treatment (Fig. S2B) [ 45 ]. This may reflect transcripts with longer maternal contributions, higher intrinsic stability, early activation prior to complete drug penetration, or the previously described “promiscuous” minor ZGA transcription that is less dependent on canonical RNA Pol II activity. Thus, while our findings confirm that the majority of ZGA genes require active transcription, they also align with the emerging view that ZGA encompasses both highly sensitive “core” transcriptional events and more resilient programs. We next assessed gene expression dynamics in α-Amanitin-treated versus control 2-cell embryos with other ZGA identification methods [ 25 , 46 – 48 ]. In our ZGA genes, major-ZGA genes (classes 4 and 5) exhibited the most pronounced transcriptional suppression (residual log2[FPKM + 1] < 0.5 in treated embryos; Fig. 2 B), with significantly greater downregulation than observed using other approaches (Wilcoxon signed-rank test, p < 0.05). Across all methods, 54 major-ZGA genes were consistently identified (Fig. 2 C, supplementary table 2). For example, Mta2 was a typical major-ZGA gene that initiated transcription at the 2-cell stage (Fig. S2C) and was markedly suppressed by α-Amanitin in 2-cell embryos (Fig. S2D). Targeted mutation of the Mta2 gene results in multiple phenotypes that include partial embryonic and perinatal lethality, female infertility, abnormal T-cell activation, and lupus-like autoimmune disease in mice [ 49 ]. Critically, we identified 119 novel major-ZGA genes absent from other methods (Fig. 2 C, supplementary table 2). These genes activated specifically at the 2-cell stage and maintained elevated expression throughout early embryonic development (Fig. S2E). To illustrate the biological relevance of these findings, we selected representative examples encompassing both canonical markers to validate our method and novel candidates to demonstrate its discovery potential. Some of these genes were supported as major-ZGA genes by literature, such as Xist and Nanog [ 17 , 50 ]. Xist is an imprinted gene expressed in early embryos (beginning at the 2-cell stage) [ 51 ], and its expression remained high at the 8-cell stage (Fig. S2F). Xist triggers gene silencing and drives major structural reorganization of the X chromosome, which allows female mammalian embryos to survive development by establishing and maintaining proper dosage compensation [ 50 ]. Nanog is essential to initiate the zygotic developmental program and induces clearance of the maternal program, thereby preventing developmental arrest before gastrulation [ 17 ], and it showed a similar expression profile trend during ZGA (Fig. S2G). In addition to these classical genes, we quantified their biological relevance through rigorous functional annotation and literature screening of 119 identified S-ZGA-specific genes. Approximately 17% (20/119) of the genes were directly involved in classical early mammalian developmental processes, such as "lamina labyrinthina development" and "cell fate determination during primary germ layer formation." Another 9% (11/119) were frequently annotated as proteins related to non-classical metabolic or stress pathways, playing definitive key roles in early embryogenesis of vertebrate models including zebrafish and African clawed frog. The remaining 74% represented novel candidate factors. Others like Mvb12a , Pdf , Uqcc3 , Ccpg1os , E530001F21Rik , AA543401 , and C330021F23Rik showed ZGA-associated expression patterns (Fig. 2 D) and were significantly downregulated by α-Amanitin (Fig. S2H). The variations in ZGA gene sets between our framework and prior studies stem from these methodological differences, highlighting their complementary value. Previous studies, using pairwise comparisons or fold-change cutoffs, effectively capture genes with sharp, transient bursts typical of canonical ZGA. Indeed, major-ZGA genes identified in prior studies but not by our framework predominantly exhibited transient 2-cell-stage expression without sustained elevation (Fig. S2I); while authentic, these likely represent a distinct regulatory layer, such as non-canonical or transient transcriptional events, compared to the sustained activation program captured by our framework. Together, these complementary perspectives offer a more integrated view of the complex, multi-faceted nature of zygotic genome activation. Gene expression pattern is conserved during embryonic development across diverse datasets We next extended to datasets generated using diverse biological and technical methods. First, we examined gene expression patterns in single-cell RNA-seq (scRNA-seq) data from prior studies [ 11 ]. The expression patterns of the ZGA genes were consistent in bulk and scRNA-seq data (Fig. 3 A), demonstrating our framework produces consistent outputs regardless of sequencing platform. We then analyzed human embryonic development and tissue datasets. The framework successfully resolved the six gene classes with distinct expression patterns in humans (Fig. S3A, supplementary table 3), confirming its robust cross-species adaptability. Given the complexity of human tissue samples, we systematically assessed whether varying sample selection or composition could bias results. By analyzing gene expression datasets containing 8, 16, or 24 tissue samples derived from human embryonic stem cells (hESCs), we found minimal variation in ZGA gene identification across these conditions (Fig. 3 B), underscoring the framework’s resilience to tissue-specific heterogeneity. Comparative analysis of ZGA gene classifications between humans and mice revealed species-specific differences in ZGA timing and gene repertoires (Fig. S3B). This divergence aligns with prior observations of distinct regulatory mechanisms governing ZGA in these species [ 27 ]. Collectively, these results demonstrate that expression patterns of ZGA genes are conserved within the same species across datasets and experimental conditions. Sustained ZGA genes are functionally enriched in housekeeping genes and sustain embryonic competence Overall, by applying our hierarchical filtering framework based on expression patterns, we identified 1,006 genes activated during the ZGA window in early mouse embryos-consisting of 316 minor-ZGA and 690 major-ZGA genes, and mapped their global functional profiles. These genes collectively govern essential early developmental processes, spanning from cell division and chromosome segregation in the minor-ZGA to protein transport and apoptosis in the major-ZGA (Fig. 1 E). While this entire cohort signifies the initiation of zygotic transcription, our analysis revealed marked heterogeneity in their temporal dynamics (Fig. 1 D). Specifically, unlike the transient activation peaks characteristic of most ZGA genes, a distinct subset within the major-ZGA (Class 4) displays a unique trajectory of sustained high expression that persists from the initial activation through later developmental stages and into mature tissues. Such persistent activity implies a specialized function for this subpopulation in preserving embryonic potential and sustaining basic cellular operations throughout the life cycle. Consequently, we further focus on class 4 (Fig. 1 D) as a critical subset of major ZGA genes. These genes are transcriptionally activated at ZGA and maintain sustained expression throughout early embryonic development, leading us to designate them Sustained ZGA (S-ZGA) genes (Fig. S4A). We next investigated expression dynamics of S-ZGA genes during oogenesis [ 52 ], revealing a dramatic decline in S-ZGA expression during the transition from grown oocytes (GO) to fully grown oocytes (FGO) (Fig. 4 A). Notably, S-ZGA genes exhibit strong functional relevance: they are highly enriched among 4,781 housekeeping genes involved in basic cell maintenance and expected to maintain constant expression levels across all cells and conditions(Fisher's exact test, p = 1.2e-27; Fig. 4 B) [ 53 ]. This enrichment aligns with ZGA’s critical role in establishing pluripotency and genome stability [ 54 ]. Rather than relying on individual examples, our conclusion is supported by the global trend that housekeeping genes within S-ZGA exhibit significantly higher expression than non-housekeeping genes across multiple developmental stages (Wilcoxon signed-rank test, p < 0.001; Fig. S4B). These findings establish S-ZGA genes as essential mediators of early embryonic competence through their dual role in ZGA activation and maintenance of core cellular functions. The epigenetic landscape of S-ZGA genes during early embryo development During mammal embryos development, epigenetic reprogramming and chromatin remodeling convert terminally differentiated gametes into totipotent or pluripotent cells. Promoter accessibility strongly correlates with gene expression levels in mouse embryos [ 55 – 59 ]. To elucidate transcriptional regulation of S-ZGA genes, we analyzed chromatin accessibility on mouse promoters using ATAC-seq data from Wu et al [ 26 ]. We observed a significant increase in ATAC-seq signal at S-ZGA gene promoters from the early 2-cell to 2-cell stage, followed by sustained high accessibility through subsequent developmental stages (Fig. 4 C- 4 E). This pattern persisted in adult tissues (Fig. S4C) and S-ZGA gene promoters exhibited a substantially higher proportion within ATAC-seq peaks during ZGA in both mouse and human datasets (Fig. S4D, S4E). Species-specific chromatin remodeling strategies were suggested: mouse embryos exhibited a selective opening pattern, where accessibility increases were restricted to specific ZGA-associated gene clusters (Classes 3–5), likely reflecting a rapid, targeted activation mechanism. In contrast, human embryos displayed a global and gradual increase in promoter accessibility across all gene classes, suggesting a broader, more permissive chromatin relaxation consistent with their slower developmental tempo. These findings directly correlate with the sustained expression profile of S-ZGA genes (Fig. 1 C) and previous studies [ 34 ]. We next characterized key epigenetic features of S-ZGA genes, including histone modification and DNA methylation [ 60 – 62 ]. Histone modifications are fundamental epigenetic regulators that control many key cellular processes [ 44 ]. Histone H3 lysine 4 trimethylation (H3K4me3), a canonical mark for active transcription initiation [ 63 ], showed significant enrichment at S-ZGA gene promoters and increased progressively during early embryonic development (Fig. 4 C, S4F). Crucially, this mark remained significantly elevated in S-ZGA promoters relative to other genes across both early embryonic stages and mature tissues (Fig. S4G, S4H). To distinguish whether these epigenetic signatures drive expression or merely reflect high transcriptional activity, we performed an expression-matched control analysis. For each gene class, we selected control genes from other classes with comparable expression levels (within a 50% difference threshold) across early developmental stages (Fig. S4I; see Methods for details). Comparing ATAC-seq and H3K4me3 signals between the target genes and this matched control set revealed distinct patterns: chromatin accessibility (ATAC-seq signal) remained significantly higher in S-ZGA genes even after controlling for expression, suggesting it is a prerequisite for activation and sustained expression of S-ZGA genes during ZGA (Fig. 4 D, 4 E) [ 64 ]. In contrast, the differential enrichment of H3K4me3 largely diminished in the matched comparison, indicating that H3K4me3 accumulation is more likely to be the result of transcriptional activation (Fig. S4J, S4K) [ 65 , 66 ]. These results refine our understanding of the epigenetic regulation of S-ZGA genes. Furthermore, we investigated whether the enrichment of housekeeping genes (HKGs) in S-ZGA genes was merely a consequence of their high expression levels. By analyzing the proportion of HKGs within the expression-matched target and control sets, we found that Class 4 (S-ZGA) genes consistently exhibited the highest percentage of housekeeping functions compared to other classes and their expression-matched controls across all developmental stages (Fig. S4L). These results demonstrate that the enrichment of housekeeping functions in S-ZGA genes is a distinct biological feature, not merely a byproduct of high transcriptional activity. In addition, H3K27me3, H3K9me3, and DNA methylation at S-ZGA promoters remained consistently low during early embryogenesis, while H3K27ac-a mark of active enhancers-was increased at the 2-cell stage (Fig. 4 C). Quantitatively, < 2% of S-ZGA promoters harbored H3K27me3 at 2-cell (Fig. S5A). At the 2-cell stage, the DNA methylation level of promoters was lowest in S-ZGA genes and the demethylation level was highest in class 3 (Fig. S5B, S5C). Concurrently, H3K9me3 density remained minimal at S-ZGA promoters (Fig. S5D, S5E). Biologically, this coordinated depletion of multiple repressive marks (DNA methylation, H3K27me3, and H3K9me3) suggests that S-ZGA promoters reside in a uniquely “permissive” chromatin state, effectively removing epigenetic barriers to facilitate the rapid binding of pioneer transcription factors (TFs) essential for ZGA initiation. Furthermore, Chromatin architecture analysis using Hi-C data [ 67 ] further revealed that S-ZGA gene promoters were significantly enriched at topologically associating domain (TAD) boundaries in 2-cell embryos compared to other gene classes (Fig. 4 F), a feature critical for transcriptional regulation [ 68 ]. By positioning S-ZGA genes at these boundaries, the embryo likely ensures their isolation from surrounding repressive domains while maximizing chromtine accessibility to the “transcriptional machines”-TFs. These integrated epigenetic and structural features-low repressive marks, upregulated H3K4me3/H3K27ac, and TAD boundary localization-demonstrate that establish a “poised-to-active” landscape. This unique configuration underpins the sustained and synchronous expression of S-ZGA genes, enabling them to serve as the primary drivers of the zygotic genome activation program. S-ZGA genes are associated with tumorigenesis and metastasis We next investigated functional roles of S-ZGA genes across the organismal life cycle. Human S-ZGA genes demonstrated significant enrichment in core developmental processes, including extracellular matrix organization, morphogenesis-linked growth regulation, and cell size control (Fig S6A, supplementary table 4). These genes were also enriched among targets of key transcription factors (TFs), such as JUN, FOS, and SP1 (Fig S6B, supplementary table 5), which orchestrate fundamental cellular activities such as proliferation, differentiation, survival, and death [ 69 – 71 ]. We further examined S-ZGA gene associations with genetic variation, aging, and disease. We collected single nucleotide polymorphism (SNP) recorded from NCBI Allele Frequency Aggregator (ALFA) [ 72 ] (see Methods, supplementary table 5), and found that promoter regions of S-ZGA genes exhibited significantly reduced SNP density (Wilcoxon rank-sum test, p < 0.001; Fig. S6C), suggesting stringent regulatory constraints. Concurrently, S-ZGA genes were markedly enriched in human phenotype-associated genes [ 73 ] and senescence-related genes [ 74 ] (Fisher's exact test, p < 0.001; Fig. S6D, Fig. S6E; supplementary table 5). For example, GBA , which ranks second highest within the S-ZGA set with 230 associated phenotypes (Supplementary Table 5), correlates with a broad spectrum of conditions ranging from developmental disorders (e.g., intrauterine growth retardation and pulmonary hypoplasia) to aging-related neurodegeneration (e.g., parkinsonism and dementia) [ 73 ]. These findings reveal that S-ZGA genes serve dual roles: they drive essential developmental processes while concurrently contributing to aging and disease susceptibility. Moreover, aberrant expression of some ZGA genes has been associated with numerous forms of cancer, since the ability of genome activators to reprogram cells toward pluripotency also makes them potent drivers of tumor development [ 4 , 75 ]. To investigate this connection and its specificity across ZGA classes, we analyzed S-ZGA gene associations with cancer. Human S-ZGA genes (class 4) were the only significantly enriched among cancer driver genes [ 76 – 82 ] (Fisher's exact test, p < 0.05; Fig. 5 A), a finding that mirrors their unique biological functional enrichment in apoptosis regulation (Fig. 1 E). Moreover, human S-ZGA genes were enriched in epithelial to mesenchymal transition (EMT) and related biological processes (Fig. 5 B, supplementary table 4), a program critical for embryogenesis, wound healing and malignant progression [ 83 ]. The co-enrichment of EMT and apoptosis regulation pathways in S-ZGA genes suggests a coordinated mechanism for metastasis: while EMT facilitates cell detachment and invasion, the concurrent suppression of apoptosis (particularly anoikis) This dual functionality ensures that disseminating cells survive the stress of circulation and successfully colonize distant organs distant organs. We observed differential expression of S-ZGA genes (e.g., CCND1 , RGL3 , and ELF3 ) between tumor and normal tissues (Fig. S7A), with elevated expression in metastatic primary tumors compared to non-metastatic counterparts (Fig S7B) [ 84 ]. Survival analyses further revealed that high expression of specific S-ZGA genes was associated with significantly worse prognosis across multiple cancer types. Among these, HIST1H3D , NOTCH3 , CD276 , POLE , CARM1 , PLK2 , ELF3 , and CCND1 emerged as the top eight genes, each showing significant prognostic value in the largest number of cancer types (Fig. S7C). These findings are consistent with prior reports linking these factors to tumor progression [ 85 – 87 ]. Finally, to further validate the association between S-ZGA genes and cancer aggressiveness, we performed an overlap analysis integrating S-ZGA genes with established cancer driver genes and differentially expressed genes (DEGs) reported in GEPIA2. Based on this intersection, we identified CCND1 , ELF3 , and RGL3 as the candidates exhibiting the most significant expression alterations for subsequent experimental validation. We respectively transfected siRNAs targeting S-ZGA genes CCND1, ELF3 , RGL3 , and negative control siRNA in human ovarian cancer cell line SKOV3, with knockdown efficiency verified by RT-qPCR. Then, the proliferation, migration and invasion of SKOV3 cells were detected by CCK8 and transwell assays. Two independent siRNAs were designed for each target gene. Both siRNAs effectively knocked down their respective targets (Fig.S7D). We found the proliferation was significantly inhibited after knockdown of ELF3 , CCND1 and RGL3 (Fig. 5 C), and cell migration and invasion were also obviously inhibited (Fig. 5 D-E and Fig.S7E). These results establish that S-ZGA genes play a functional role in regulating cancer cell proliferative, migratory, and invasive capacities in vitro . In conclusion, S-ZGA genes were activated during ZGA and maintained throughout development. They orchestrate essential developmental processes while, once their epigenetic signals are “hijacked” to drive aberrant expression, they may promote tumorigenesis and metastasis. Discussion ZGA marks the transition from maternal to zygotic transcriptional control in metazoan embryos, initiating the major wave of embryonic gene expression [ 6 ]. In this study, we introduced a computational framework that leverages time-series RNA-seq data to capture the dynamic trajectories of gene expression during early embryogenesis. Unlike prior methods reliant on static pre/post-ZGA comparisons [ 23 , 27 , 28 ]. Our framework enables the prior-free ZGA timing identification and robust cross-dataset gene classification. This framework not only recapitulates known ZGA timing in well-characterized organisms but also provides a foundation for future discovery of ZGA dynamics in unexplored species. In addition, the ZGA gene sets we identified represent a valuable resource for developmental and cancer biology research, as the transcriptional mechanisms governing embryonic ZGA are conserved in cellular reprogramming contexts such as induced pluripotent stem cell (iPSC) generation [ 4 ]. This framework enables the systematic discovery of novel genes critical for developmental processes and cancer-driving functions, thereby advancing our understanding of the molecular basis of embryonic genome activation. Furthermore, one subset stood out for its distinctive expression dynamics: class 4 genes, which we term Sustained ZGA (S-ZGA) genes, are activated at ZGA and maintain persistent expression throughout subsequent development (Fig. S4A). This sustained pattern prompted us to investigate whether these genes possess unique regulatory and functional properties that might underlie their potential to be re-deployed in disease contexts. By integrating epigenomic profiling with expression-matched control analyses-a crucial step to address potential confounding by high transcriptional activity-we demonstrated that S-ZGA genes possess a unique regulatory landscape. Notably, their promoters exhibit significantly higher chromatin accessibility (ATAC-seq) even after controlling for expression level, whereas H3K4me3 enrichment is largely explained by transcriptional output. This suggests that open chromatin is an intrinsic feature of S-ZGA genes, poising them for sustained activity, while H3K4me3 reflects ongoing transcription. These findings refine our understanding of how epigenetic mechanisms establish and maintain distinct gene expression programs during development. In addition, there are many similarities between early embryonic development and tumorigenesis in cell invasion behavior, epigenetic regulation, gene expression, metabolism and other important biological behaviors [ 88 , 89 ]. Epithelial-mesenchymal transition (EMT) pathways-a program essential for embryogenesis and aberrantly reactivated in metastasis [ 90 – 92 ]. Using in vitro assays in an ovarian cancer cell line, we further showed that knockdown of selected S-ZGA genes ( CCND1 , ELF3 , RGL3 ) impairs proliferation, migration, and invasion. our work provides a foundation for investigating how early embryonic drivers are re-deployed in disease, converting them from uncharacterized transcriptional events into mechanistically tractable hypotheses. Several limitations of this study should be acknowledged. First, our current framework is specifically designed to identify a discrete developmental transition-the onset of ZGA. For continuous processes such as cellular differentiation or lineage commitment, the sample-clustering approach used here may not be directly applicable. However, the underlying principle of analyzing full temporal trajectories remains valid. By replacing hierarchical clustering with pseudotime ordering or continuous trend-clustering algorithms, the core concept of our framework could be adapted to systematically identify genes with specific dynamic patterns in gradual biological processes; this represents a promising direction for future work. Second, while our framework is designed to be species-flexible, its application to species with completely uncharacterized ZGA timing remains to be systematically validated. Finally, the cancer-related findings presented here are primarily correlative. Although we demonstrate a strong statistical association between S-ZGA genes and cancer driver genes, and provide in vitro evidence for three significantly differentially expressed S-ZGA genes ( CCND1 , ELF3 , RGL3 ) in the SKOV3 cell line, these results do not establish a direct causal role for S-ZGA genes in driving tumorigenesis or metastasis in vivo . The functional validation definitive evidence would require in vivo models and broader interrogation across diverse cancer types. In summary, this study establishes a dynamic framework for dissecting gene expression programs during ZGA, identifies a previously underappreciated class of sustained ZGA genes, and uncovers their potential dual relevance to development and disease. By providing a robust resource and highlighting the interplay between epigenetic regulation and oncogenic potential, our work opens new avenues for investigating the molecular mechanisms that bridge early embryogenesis and cancer biology. Methods scRNA-seq data processing Sequencing reads were trimmed using Trim Galore (version 0.6.5), which is a Perl wrapper based on two tools: Cutadapt and FastQC. Reads with quality > 20 were aligned to the mouse reference genome (mm10) using STAR (version 2.7.3a) [ 93 ]. Low mapping quality sequencing reads were removed using Samtools version 1.7. The gene expression level was calculated by Cufflinks (version 2.2.1) [ 94 ]. The relative abundance of transcripts was measured as fragments per kilobase of exon per million mapped reads (FPKM). Testing of the differential expression between different groups was performed with the Cuffdiff module, and genes with a q-value < 0.05 were selected as being differentially expressed. Data Integration and Normalization To integrate RNA-seq data from early mouse embryonic development with publicly available ENCODE mature tissue data, we implemented a rigorous two-step normalization pipeline. First, to address the inherent skewness and heteroscedasticity of RNA-seq count-derived data, FPKM values were subjected to a log2-transformation [log2(FPKM + 1)]. This step stabilizes variance and approximates a normal distribution, satisfying the assumptions required for subsequent parametric scaling. Second, we applied a row-wise Z-score transformation to the log-transformed matrix. This normalization aligns the distribution of expression values across datasets by ensuring similar mean and variance per gene, thereby minimizing technical biases arising from different sequencing platforms. This approach enables the analysis to focus on the similarity of relative gene expression patterns rather than absolute abundance. Hierarchical clustering Following normalization, hierarchical clustering was performed based on the Weighted average distance (WPGMA) algorithm (this approach employs a recursive formula to update inter-cluster distances upon merging D (Cκ, Ci ∪Cj) and defines the distance between clusters as the weighted average of all pairwise sample distances), which uses a recursive definition for the distance between two clusters and eliminates the influence of class cluster size on clustering. The Spearman distance was used to measure the distance of samples, which is robust to non-linear relationships and outliers [ 95 ]. Principal component analysis We performed a principal component analysis (PCA) of the gene expression-level matrix to further characterize transcriptome distances. The calculation was performed with a singular-value decomposition of the data matrix normalized by Z-Score transformation, which made sure that the mean expression of each gene across cells is 0, and the variance across cells is 1. Then, we used the top two principal component features to describe the clustering relationships among embryonic samples in two-dimensional space. Key Gene screening for each developmental stage To capture the dynamic transcriptional landscape during early embryogenesis, we employed a three-layer filtering strategy designed to classify genes based on their binary expression states (“Active” vs. “Silent”) across the three sequential developmental clusters identified in Fig. 1 C. We screened genes based on their expression. First, we defined the features representing the expression pattern. Briefly, in early embryonic development, the features were genes whose FPKM ≥ 5 at > 25% of the samples or FPKM 25% of the samples or FPKM > 1 at ≤ 25% sample of the cluster. Next, the genes defined as features were transferred to the next filter, while other genes would be discarded. After gene screening by 3-layer filters, 8 gene lists were obtained (Fig. 1 C, Fig. S1 C). To deal with maternal genes with high expression levels before ZGA, we additionally screened maternal genes (in class 2) whose FPKM > 100 at MII oocyte, zygote and early 2-cell stages and FPKM increased fivefold at early 2-cell stage compared with 4-cell stage (FPKM > 1 at ≤ 50% sample for 2-cell to ESC-stages). Assessment of threshold robustness To evaluate the robustness of our gene classification strategy, we performed a sensitivity analysis by varying the sample proportion cutoff used to define consistent expression within each developmental cluster. Specifically, we tested cutoffs ranging from 20% to 30% (default: 25%) Fig. S8A-8D. The core classification of main ZGA genes remained largely stable, particularly for S-ZGA genes, where over 90% of genes were consistently classified across different parameter settings, indicating that our main conclusions are robust to the exact choice of this parameter. We further evaluated the robustness of class 1 classification by adjusting the post-ZGA silence threshold (default: FPKM > 1 at < 50% of samples from 2-cell to ESC). Varying this threshold did not substantially alter the set of identified genes or their enriched biological processes Fig. S8E. Performing a cross-species ortholog retention analysis To investigate whether the functional and developmental importance of S-ZGA genes is reflected in their evolutionary conservation, we performed an ortholog analysis between mouse and human. We identified one-to-one orthologs using the Ensembl Compara database. We then calculated the proportion of human S-ZGA genes (Class 4 in human) that retain a clear ortholog in mouse, and compared this to the ortholog retention rate of non-S-ZGA human genes (all other classes). Validation of Gene Classes Using Public Perturbation Data To validate the functional characteristics of the six identified gene classes, we examined their expression responses to transcriptional inhibition using publicly available RNA-seq datasets from mouse 2-cell embryos treated with α-Amanitin. It is important to note that this analysis was performed as a descriptive comparison of expression patterns rather than a statistical differential expression test. Where biological replicates were available in the source dataset, we calculated the mean FPKM values for each condition to represent the expression profile. We then compared the relative expression levels of genes in each class between the control and treated groups to identify consistent up- or down-regulation trends. Gene set enrichment analysis Gene Ontology (GO) term enrichment analysis in Fig. 1 D was performed for genes in each class using the Database for Annotation, Visualization and Integrated Discovery Bioinformatics Resource (DAVID, https://david.ncifcrf.gov/home.jsp ) [ 96 ]. GO analysis in Fig. 5 B and Fig.S6A was performed by clusterProfiler-enrichGO (V4.0.5) [ 97 ]. Enriched TFs of S-ZGA genes in Fig.S6B were calculated by metascape [ 98 ] and recorded in TRRUST Database[ 99 ]. The Benjamini-Hochberg method was used for correction to control the false discovery rate (FDR). For the Gene Ontology (GO) enrichment analysis, the background gene set was defined as all annotated genes in the genome. Identification of topologically associated domain boundaries We collected Hi-C data for the late 2-cell stage from GEO under accession no. GSE82185 and identified significant interactions using HOMER algorithm with a default parameter ( p = 0.001) and 5-kb resolution. A topologically associated domain (TAD) boundaries were identified at 40-kb bin resolution using the previously described insulation score analysis approach [ 100 ]. Briefly, TAD boundaries were identified by using the insulation score method at 40-kb resolution, with two minor modifications. First, a 200-kb (rather than a 500-kb) genomic region was used to detect a TAD, as a TAD is now considered to be mostly < 200-kb. Second, in delta vector calculations, we used a 120-kb (rather than a 100-kb) window to fit the resolution, which is suitable in detecting TADs with high sensitivity with a low false-positive rate. We collected Hi-C data from GEO under accession. Analyses of SNP within promoters of genes We collected vcf file from NCBI Allele Frequency Aggregator (ALFA, www.ncbi.nlm.nih.gov/snp/docs/gsr/alfa/ , Release Version: 20200227123210) [ 72 ]. We used bcftools to calculate the SNP overlap with each target region (± 1 kb around the transcription start site (TSS)). We then filter the SNP that minor allele frequency (MAF) is 0 (site homozygous for the reference allele and no variant allele detected from the current subject sample size). For the SNP density analysis, the background gene set was defined as all detected protein-coding genes within our dataset, excluding the identified S-ZGA genes. This restriction to protein-coding genes ensured comparability in gene structure and annotation. Notably, Cluster 0 and Cluster 7 were included in the background set; all non-S-ZGA protein-coding genes within these clusters served as controls, as they represent the same biological context and cell states relevant to the study. Epigenetic profiles S-ZGA genes We collected bedgraph or bigwig files from GEO as Datasets parts mentioned. We used the EnrichedHeatmap R package to quantify ATAC-seq signals. Specifically, the normalizeToMatrix function was employed to map genomic signals to target regions (± 5 kb around the TSS). Signal values were averaged across biological replicates where available. Finally, for each gene, we calculated the mean signal intensity across its entire target region to represent the overall chromatin accessibility. Expression-matched control selection Specifically, to control for potential confounding due to differences in gene expression levels, we performed a matched control gene selection for each gene class (classes 1–6). Using expression data from the early 2-cell to ICM stages, genes belonging to a given class were designated as the target set, while all genes from the other classes served as the candidate pool. For each target gene, we searched the pool without replacement and selected the gene whose expression value was closest to that of the target, provided the absolute difference did not exceed 50% of the target’s expression value. If no gene in the pool met this criterion, the target received no matched control (recorded as missing). This procedure yielded a list of target-control pairs with comparable expression levels, thereby minimizing expression-level bias in downstream epigenetic analyses. We then compared the ATAC-seq and H3K4me3 signals at promoter regions (± 1 kb around TSS) between classes 1–6 genes and this expression-matched control set. Analyses of differential gene expression and cancer survival We analyzed differential gene expression between tumor and normal tissues using the Gene Expression Profiling Interactive Analysis (GEPIA2) comprehensive online platform ( http://gepia2.cancer-pku.cn/#index ), which provides rapid and customized delivery of functionalities based on the Cancer Genome Atlas and genotype-tissue expression data [ 101 ]. Using GEPIA2, the survival effects of differentially expressed genes in our cancer samples were assessed to determine the prognostic value of the expression of identified key ZGA genes. Cell culture SKOV3 cell line were purchased from the Cell Bank of Peking Union Medical College.SKOV3 cells were cultured in McCoy’s 5A medium (Gibco,16600-082, USA) supplemented with 10% fetal bovine serum (Gibco,10270-106, USA) at 37°C in a humidified 5% CO 2 atmosphere. Gene silencing by siRNA transfection ELF3 , CCND1 and RGL3 genes were knocked down with their specific siRNAs designed and synthesized by GenePharma Co., Ltd. (Suzhou, China). For siRNA transfection, SKOV3 seeded in 6-well plates at 60 ~ 70% were transfected with specific pooled siRNAs and scrambled negative control at a final concentration of 100 nM mixed with jetPRIME transfection reagent (Polyplus,114 − 15, France) according to the manufacturer’s instructions. After that, the cells were harvested for RT-qPCR after 48 h to detect knockdown efficiency. To rule out off-target effects, knockdown experiments were performed using two distinct siRNA sequences targeting genes. For each siRNA, three independent biological replicates were conducted. The sequences of specific siRNAs were as follows: siNC: 5ʹ-UUCUCCGAACGUGUCACGUTT-3ʹ siELF3 #1: 5ʹ-GCUACCAAGUGGAGAAGAATT-3ʹ siELF3 #2: 5ʹ-GCCAUGAGGUACUACUACATT-3ʹ siCCND1 #1: 5ʹ-GUCUGCGAGGAACAGAAGUTT-3ʹ siCCND1 #2: 5ʹ-GCAUGUUCGUGGCCUCUAATT-3ʹ siRGL3 #1: 5ʹ-GGUGGGAGAGUUGGUGUUUTT-3ʹ siRGL3 #2: 5ʹ-CCUGUGACUAUCAGCUCUUTT-3ʹ CCK8 assay Cell viability was detected by cell counting kit-8 (Dojindo, CK04, Japan). SKOV3 cells were seeded in 96-well plates at a density of 5× 10 3 cells per well, 24 h after transfected siRNAs targeting ELF3 , CCND1 , RGL3 or negative control with jetPRIME transfection reagent (Polyplus,114 − 15, France). Cultured to the indicated times, the medium of each well was replaced with 100 µL fresh culture media containing 10% CCK-8 solution and incubated for 1.5 h at 37°C. The absorbance of each well was measured at a wavelength of 450 nm. All experiments were independently performed in triplicate. RNA extraction and RT-qPCR Forty-eight hours after siRNA transfection, total RNA was extracted from SKOV3 cells using Trizol reagent (Invitrogen,15596026, USA) following the manufacturer's instruction. Synthesis of cDNA was performed using HiScript III All-in-one RT SuperMix (Vazyme, R333, China). qPCR was performed using PowerUpTM SYBRTM Green Master Mix (Applied Biosystems, A25742, USA) and analysed on Mx3000P QPCR Systems (Agilent, Mx3000P, USA). The relative mRNA expression was normalized to GAPDH. Each sample was determined with triplicate independent experiments. The primers used for amplification were as follows: ELF3-Forward: 5ʹ-CAGACTCCGGTGGAAGTGAC-3ʹ ELF3-Reverse: 5ʹ-CCCTCGAGACAGTCCCAGTA-3ʹ CCND1-Forward: 5ʹ-GGAGAACAAACAGATCATCCGC-3ʹ CCND1-Reverse: 5ʹ-GTGTGAGGCGGTAGTAGGAC-3ʹ RGL3-Forward: 5ʹ-CTGTGTGCTGGGTTCCGT-3ʹ RGL3-Reverse: 5ʹ-TGAGCCGGTAGATGGGGTTA-3ʹ GAPDH-Forward: 5ʹ-GTCTCCTCTGACTTCAACAGCG-3ʹ GAPDH-Reverse: 5ʹ-ACCACCCTGTTGCTGTAGCCAA-3ʹ Cell migration and invasion assays Migration and invasion assays were performed by 6.5 mm Transwell® with 8.0 µm Pore Polycarbonate Membrane Insert (Corning, 3422, USA) and BioCoatTM Matrigel Invasion Chamber (Corning, 354480, USA), respectively. The Transwell insert and Matrigel Chamber were respectively seeded with SKOV3 cells in 1.5×10 4 cells/insert or 8×104 cells/chamber in serum-free McCoy’s 5A medium, 24 h after transfected siRNAs targeting ELF3 , CCND1 , RGL3 or negative control. The 24-well plates were added with culture medium containing 10% or 20% fetal bovine serum as chemoattractant, and then the cells were incubated for 24 h or 48 h in migration and invasion assay, respectively. After incubation, the non-migrated or non-invaded cells were removed from the upper surface of the membrane by a cotton swab. The cells on the lower surface of the membrane were fixed and stained with Diff-Quik Stain Kit (Solarbio, G1540, China), then photographed and counted the nine random fields from triplicate membranes of each group by inverted microscope (100×, Nikon, Japan). All experiments were independently performed in triplicate. Statistical Analysis Statistical analysis was performed using GraphPad Prism (GraphPad Software) and R software. Survival analysis based on GEPIA2 uses log-rank test for the hypothesis evaluation. In the verification part of the experiment, paired two-tailed Student’s t tests were employed to conduct statistical analysis, and nominal p values are specified. All values represent the means of at least three independent experiments ± SD. p < 0.05 was considered statistically significant. Specifically, the Wilcoxon signed-rank test was applied for data in Fig. 2 A/B and Fig. S4B, and the Wilcoxon rank-sum test (also called Mann-Whitney U test) was applied for data in Fig. S6C/D; Fisher’s exact test was used for Fig. 4 A and Fig. 5 A; and an unpaired t-test was employed for Fig. 5 C/D. Unless otherwise specified, all experiments were performed in at least three independent replicates, and data are presented as mean ± standard deviation (SD). A p-value < 0.05 was considered statistically significant. Hypothesis Testing: The null hypothesis ( H0 ) for all tests posited no difference between groups (mean/median difference = 0). A two-tailed alternative hypothesis ( H1 ) was assumed unless otherwise specified. Multiple Testing Correction: For high-throughput screening data involving thousands of genes, p-values were adjusted for multiple testing using the Benjamini-Hochberg procedure to control the False Discovery Rate (FDR). For comparisons like the Wilcoxon signed-rank test or unpaired two-tailed Student’s t tests, which were not part of a large-scale multiple testing procedure, we have reported the raw p values as is standard practice. Abbreviations ZGA Zygotic genome activation S-ZGA Sustained ZGA ESC Embryonic stem cell ICM Inner cell mass FGO Fully grown oocytes TAD Topologically associating domain SNP Single nucleotide polymorphism TF Transcription factor iPSC Induced pluripotent stem cell Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Data Availability Mouse RNA-seq data covering developmental stages from MII oocytes and PN5 zygotes to early 2-cell, 2-cell, 4-cell, and 8-cell embryos, as well as ICM and ESCs, were obtained from the GEO database (Accession: GSE66390) [26]. Mouse scRNA-seq data spanning stages from MII oocytes and PN5 zygotes to early 2-cell, 2-cell, 4-cell, and 8-cell embryos, as well as ICM and ESCs, were retrieved from the GEO database (Accession: GSE45719) [11]. Mouse RNA-seq data of 2-cell embryos, including both untreated controls and those treated with α-Amanitin, were sourced from the GEO database (Accession: GSE71434) [44]. Mouse scRNA-seq datasets comparing untreated 2-cell embryos with α-Amanitin-treated counterparts were acquired from the GEO database (Accession: GSE53386) [47]. Mouse RNA-seq data for grown oocytes and fully grown oocytes (FGO) were compiled from the GEO database (Accession: GSE112835) [52]. Mouse ATAC-seq data covering early 2-cell, 2-cell, 4-cell, and 8-cell embryos as well as ICM were obtained from the GEO database (Accession: GSE66390) [26]. Mouse ChIP-seq data for H3K4me3 across sperm, MII oocytes, zygotes, early 2-cell, 2-cell, 4-cell, and 8-cell embryos, and the ICM were retrieved from the GEO database (Accession: GSE71434) [44]. Mouse ChIP-seq data for H3K27me3 across MII oocytes, zygotes, 2-cell embryos, 8-cell embryos, and the ICM were sourced from the GEO database (Accession: GSE76687) [102]. Mouse ChIP-seq data for H3K9me3 across sperm, MII oocytes, zygotes, 2-cell, 4-cell, and 8-cell embryos, and the ICM were acquired from the GEO database (Accession: GSE98149) [103]. Mouse ChIP-seq data for H3K27ac in oocytes, 2-cell embryos, and 8-cell embryos were compiled from the GEO database (Accession: GSE72784) [104]. Mouse DNA methylome data covering sperm, MII oocytes, 2-cell and 4-cell embryos, and the ICM stage were obtained from the GEO database (Accession: GSE56697) [105]. Mouse Hi-C data for late 2-cell embryos were retrieved from the GEO database (Accession: GSE82185) [67]. Mouse RNA-seq data for multiple tissues, including heart, kidney, liver, stomach, duodenum, lung, spleen, and thymus, were sourced from the ENCODE project [106]. Mouse ATAC-seq data for heart and liver tissues were acquired from the ENCODE project [106]. Mouse ChIP-seq data for H3K4me3 in heart and liver tissues were compiled from the ENCODE project [106]. Human RNA-seq data spanning GV oocytes, MII oocytes, 2-cell, 4-cell, and 8-cell embryos, as well as the ICM and ESCs, were obtained from the GEO database (Accession: GSE101571) [34]. Human ATAC-seq data for the 4-cell and 8-cell stages were retrieved from the GEO database (Accession: GSE101571) [34]. Human RNA-seq data comprising 24 samples were sourced from the ENCODE project [106]. Bovine RNA-seq data covering MII oocytes, 2-cell, 4-cell, and 8-cell embryos, as well as morula and blastocyst stages, were acquired from the GEO database (Accession: GSE59186) [107]. Macaque RNA-seq data spanning GV, MI, and MII oocytes, PN zygotes, 2-cell, 4-cell, and 8-cell embryos, through morula and blastocyst stages, were compiled from the GEO database (Accession: GSE103313) [35]. Mouse promoter coordinates were mapped to the mm10 genome assembly. All gene expression analyses utilized pre-computed matrices downloaded from public repositories; no raw data re-processing or re-alignment was performed. Housekeeping genes in mouse were compiled from previous studies [108] Cancer driver genes [76-82] were compiled from the Candidate Cancer Gene Database (CCGD; http://ccgd-starrlab.oit.umn.edu/), incorporating core sets identified by Lawrence et al. (260 genes), Vogelstein et al. (125 genes), Schroeder et al. (285 genes), Zehir et al. (545 genes), Sondka et al. (723 genes), and Martínez-Jiménez et al. (568 genes). Human phenotype-related genes curated in the Human Phenotype Ontology (HPO) Database (https://hpo.jax.org) [73] Human senescence related genes sourced from the Cell Senescence Gene Database (http://csgene.bioinfo-minzhao.org/) [74] Differential expression profiles between tumor and normal tissues derived from GEPIA (http://gepia2.cancer-pku.cn/#about) [101] Gene expression levels in metastasis samples obtained from the Human Cancer Metastasis Database (HCMDB) (http://hcmdb.i-sanger.com/index) [84] Competing interests The authors declare that they have no competing interests. Funding The authors declare that this research received no specific grant from any funding. Author contributions Y.Z. contributed to the conception and design of the study. H.X. performed the experiments and wrote the manuscript. Acknowledgements The authors would like to thank Prof. Hebing Chen for his helpful in study design. References Pijuan-Sala B, Guibentif C, Göttgens B. Single-cell transcriptional profiling: a window into embryonic cell-type specification. Nat Rev Mol Cell Biol. 2018;19(6):399–412. Zernicka-Goetz M, Morris SA, Bruce AW. Making a firm decision: multifaceted regulation of cell fate in the early mouse embryo. Nat Rev Genet. 2009;10(7):467–77. Sha QQ, Zhu YZ, Li S, Jiang Y, Chen L, Sun XH, Shen L, Ou XH, Fan HY. Characterization of zygotic genome activation-dependent maternal mRNA clearance in mouse. Nucleic Acids Res. 2020;48(2):879–94. Schulz KN, Harrison MM. Mechanisms regulating zygotic genome activation. Nat Rev Genet. 2019;20(4):221–34. Chen H, Einstein LC, Little SC, Good MC. Spatiotemporal Patterning of Zygotic Genome Activation in a Model Vertebrate Embryo. Dev Cell. 2019;49(6):852–e866857. De Iaco A, Planet E, Coluccio A, Verp S, Duc J, Trono D. DUX-family transcription factors regulate zygotic genome activation in placental mammals. Nat Genet. 2017;49(6):941–5. Tadros W, Lipshitz HD. The maternal-to-zygotic transition: a play in two acts. Development. 2009;136(18):3033–42. Heyn P, Kircher M, Dahl A, Kelso J, Tomancak P, Kalinka AT, Neugebauer KM. The earliest transcribed zygotic genes are short, newly evolved, and different across species. Cell Rep. 2014;6(2):285–92. Liu B, Xu Q, Wang Q, Feng S, Lai F, Wang P, Zheng F, Xiang Y, Wu J, Nie J et al. The landscape of RNA Pol II binding reveals a stepwise transition during ZGA. Nature 2020. Abe K, Yamamoto R, Franke V, Cao M, Suzuki Y, Suzuki MG, Vlahovicek K, Svoboda P, Schultz RM, Aoki F. The first murine zygotic transcription is promiscuous and uncoupled from splicing and 3' processing. Embo j. 2015;34(11):1523–37. Deng Q, Ramsköld D, Reinius B, Sandberg R. Single-cell RNA-seq reveals dynamic, random monoallelic gene expression in mammalian cells. Sci (New York NY). 2014;343(6167):193–6. Xue Z, Huang K, Cai C, Cai L, Jiang CY, Feng Y, Liu Z, Zeng Q, Cheng L, Sun YE, et al. Genetic programs in human and mouse early embryos revealed by single-cell RNA sequencing. Nature. 2013;500(7464):593–7. Park SJ, Komata M, Inoue F, Yamada K, Nakai K, Ohsugi M, Shirahige K. Inferring the choreography of parental genomes during fertilization from ultralarge-scale whole-transcriptome analysis. Genes Dev. 2013;27(24):2736–48. Yan L, Yang M, Guo H, Yang L, Wu J, Li R, Liu P, Lian Y, Zheng X, Yan J, et al. Single-cell RNA-Seq profiling of human preimplantation embryos and embryonic stem cells. Nat Struct Mol Biol. 2013;20(9):1131–9. Vassena R, Boué S, González-Roca E, Aran B, Auer H, Veiga A, Izpisua Belmonte JC. Waves of early transcriptional activation and pluripotency program initiation during human preimplantation development. Development. 2011;138(17):3699–709. Lee MT, Bonneau AR, Giraldez AJ. Zygotic genome activation during the maternal-to-zygotic transition. Annu Rev Cell Dev Biol. 2014;30:581–613. Lee MT, Bonneau AR, Takacs CM, Bazzini AA, DiVito KR, Fleming ES, Giraldez AJ. Nanog, Pou5f1 and SoxB1 activate zygotic gene expression during the maternal-to-zygotic transition. Nature. 2013;503(7476):360–4. Mitsui K, Tokuzawa Y, Itoh H, Segawa K, Murakami M, Takahashi K, Maruyama M, Maeda M, Yamanaka S. The homeoprotein Nanog is required for maintenance of pluripotency in mouse epiblast and ES cells. Cell. 2003;113(5):631–42. Mingyue G, Yanping Z, Jianfeng Z, Yan B, Junqin X, Ce X, Xiaochen K, Yanhong Z, Yanhe L, Zhifen T et al. Precise temporal regulation of Dux is important for embryo development. Cell Res 2019, 29(11). Shuyan J, Fengling C, Paula S, Jiacheng W, Ziming Z, Lijuan W, Qing Z, Zili L, Bofeng L, Kai X et al. OBOX regulates mouse zygotic genome activation and early development. Nature 2023, 620(7976). Zhang W, Chen F, Chen R, Xie D, Yang J, Zhao X, Guo R, Zhang Y, Shen Y, Göke J, et al. Zscan4c activates endogenous retrovirus MERVL and cleavage embryo genes. Nucleic Acids Res. 2019;47(16):8485–501. Inoue K, Oikawa M, Kamimura S, Ogonuki N, Nakamura T, Nakano T, Abe K, Ogura A. Trichostatin A specifically improves the aberrant expression of transcription factor genes in embryos produced by somatic cell nuclear transfer. Sci Rep. 2015;5:10127. Sonehara H, Nagata M, Aoki F. Roles of the first and second round of DNA replication in the regulation of zygotic gene activation in mice. J Reprod Dev. 2008;54(5):381–4. Li H, Song M, Yang W, Cao P, Zheng L, Zuo Y. A Comparative Analysis of Single-Cell Transcriptome Identifies Reprogramming Driver Factors for Efficiency Improvement. Mol therapy Nucleic acids. 2020;19:1053–64. Abe KI, Funaya S, Tsukioka D, Kawamura M, Suzuki Y, Suzuki MG, Schultz RM, Aoki F. Minor zygotic gene activation is essential for mouse preimplantation development. Proc Natl Acad Sci USA. 2018;115(29):E6780–8. Wu J, Huang B, Chen H, Yin Q, Liu Y, Xiang Y, Zhang B, Liu B, Wang Q, Xia W, et al. The landscape of accessible chromatin in mammalian preimplantation embryos. Nature. 2016;534(7609):652–7. Jukam D, Shariati SAM, Skotheim JM. Zygotic Genome Activation in Vertebrates. Dev Cell. 2017;42(4):316–32. Li L, Lu X, Dean J. The maternal to zygotic transition in mammals. Mol Aspects Med. 2013;34(5):919–38. Collart C, Owens ND, Bhaw-Rosun L, Cooper B, De Domenico E, Patrushev I, Sesay AK, Smith JN, Smith JC, Gilchrist MJ. High-resolution analysis of gene activity during the Xenopus mid-blastula transition. Development. 2014;141(9):1927–39. Harvey SA, Sealy I, Kettleborough R, Fenyes F, White R, Stemple D, Smith JC. Identification of the zebrafish maternal and paternal transcriptomes. Development. 2013;140(13):2703–10. Lott SE, Villalta JE, Schroth GP, Luo S, Tonkin LA, Eisen MB. Noncanonical compensation of zygotic X transcription in early Drosophila melanogaster development revealed through single-embryo RNA-seq. PLoS Biol. 2011;9(2):e1000590. Ding J, Aronow BJ, Kaminski N, Kitzmiller J, Whitsett JA, Bar-Joseph Z. Reconstructing differentiation networks and their regulation from time series single-cell expression data. Genome Res. 2018;28(3):383–95. Østrup O, Andersen IS, Collas P. Chromatin-linked determinants of zygotic genome activation. Cell Mol Life Sci. 2013;70(8):1425–37. Wu J, Xu J, Liu B, Yao G, Wang P, Lin Z, Huang B, Wang X, Li T, Shi S, et al. Chromatin analysis in human early development reveals epigenetic transition during ZGA. Nature. 2018;557(7704):256–60. Chitwood JL, Burruel VR, Halstead MM, Meyers SA, Ross PJ. Transcriptome profiling of individual rhesus macaque oocytes and preimplantation embryos. Biol Reprod. 2017;97(3):353–64. Graf A, Krebs S, Zakhartchenko V, Schwalb B, Blum H, Wolf E. Fine mapping of genome activation in bovine embryos by RNA sequencing. Proc Natl Acad Sci USA. 2014;111(11):4139–44. Clark MB, Mercer TR, Bussotti G, Leonardi T, Haynes KR, Crawford J, Brunck ME, Cao KA, Thomas GP, Chen WY, et al. Quantitative gene profiling of long noncoding RNAs with targeted RNA sequencing. Nat Methods. 2015;12(4):339–42. Liu S, Lin L, Jiang P, Wang D, Xing Y. A comparison of RNA-Seq and high-density exon array for detecting differential gene expression between closely related species. Nucleic Acids Res. 2011;39(2):578–88. John RT, Sean BC. Gene co-option in physiological and morphological evolution. Annu Rev Cell Dev Biol 2002, 18(0). Weikun X, Jiawei X, Guang Y, Guidong Y, Kai X, Xueshan M, Nan Z, Bofeng L, Tong L, Zili L et al. Resetting histone modifications during human parental-to-zygotic transition. Science (New York, NY) 2019, 365(6451). Mina LK, Caroline H, Antonio JG. The maternal-to-zygotic transition: reprogramming of the cytoplasm and nucleus. Nat Rev Genet 2024, 26(4). Qian-Qian S, Yun-Wen W, Meng-Yan J, Yu J, Xing-Xing D, Long-Wen Z, Heng-Yu F. Maternal mRNA clearance is associated with the prevention of precocious transcription and genome instability in mouse early embryos. Sci Bull (Beijing) 2025, 71(5). Qiuyan W, Chuanxin Z, Yanna D, Jiaqi S, Zhuoning Z, Cheng L, Shuiying M, Zongyu L, Hui L, Xiaonan M et al. Maternal factor OTX2 regulates human embryonic genome activation and early development. Nat Genet 2025, 57(11). Zhang B, Zheng H, Huang B, Li W, Xiang Y, Peng X, Ming J, Wu X, Zhang Y, Xu Q, et al. Allelic reprogramming of the histone modification H3K4me3 in early mammalian development. Nature. 2016;537(7621):553–7. K E LR. L: Strain-specific progression of alpha-amanitin-treated mouse embryos beyond the two-cell stage. Mol Reprod Dev 1995, 41(1). Lu F, Liu Y, Inoue A, Suzuki T, Zhao K, Zhang Y. Establishing Chromatin Regulatory Landscape during Mouse Preimplantation Development. Cell. 2016;165(6):1375–88. Fan X, Zhang X, Wu X, Guo H, Hu Y, Tang F, Huang Y. Single-cell RNA-seq transcriptome analysis of linear and circular RNAs in mouse preimplantation embryos. Genome Biol. 2015;16(1):148. Park SJ, Shirahige K, Ohsugi M, Nakai K. DBTMEE: a database of transcriptome in mouse early embryos. Nucleic Acids Res. 2015;43(Database issue):D771–776. Lu X, Kovalev GI, Chang H, Kallin E, Knudsen G, Xia L, Mishra N, Ruiz P, Li E, Su L, et al. Inactivation of NuRD component Mta2 causes abnormal T cell activation and lupus-like autoimmune disease in mice. J Biol Chem. 2008;283(20):13825–33. Loda A, Heard E. Xist RNA in action: Past, present, and future. PLoS Genet. 2019;15(9):e1008333. Wells KD. More with less Xist. Proc Natl Acad Sci USA. 2012;109(2):349–50. Xu Q, Xiang Y, Wang Q, Wang L, Brind'Amour J, Bogutz AB, Zhang Y, Zhang B, Yu G, Xia W, et al. SETD2 regulates the maternal epigenome, genomic imprinting and embryonic development. Nat Genet. 2019;51(5):844–56. Eisenberg E, Levanon EY. Human housekeeping genes, revisited. Trends Genet. 2013;29(10):569–74. Ogawa S, Yamada M, Nakamura A, Sugawara T, Nakamura A, Miyajima S, Harada Y, Ooka R, Okawa R, Miyauchi J, et al. Zscan5b Deficiency Impairs DNA Damage Response and Causes Chromosomal Aberrations during Mitosis. Stem Cell Rep. 2019;12(6):1366–79. Guo F, Li L, Li J, Wu X, Hu B, Zhu P, Wen L, Tang F. Single-cell multi-omics sequencing of mouse early embryos and embryonic stem cells. Cell Res. 2017;27(8):967–88. Li L, Guo F, Gao Y, Ren Y, Yuan P, Yan L, Li R, Lian Y, Li J, Hu B, et al. Single-cell multi-omics sequencing of human early embryos. Nat Cell Biol. 2018;20(7):847–58. Xu Q, Xie W. Epigenome in Early Mammalian Development: Inheritance, Reprogramming and Establishment. Trends Cell Biol. 2018;28(3):237–53. Zhou LQ, Dean J. Reprogramming the genome to totipotency in mouse embryos. Trends Cell Biol. 2015;25(2):82–91. Hontelez S, van Kruijsbergen I, Georgiou G, van Heeringen SJ, Bogdanovic O, Lister R, Veenstra GJC. Embryonic transcription is controlled by maternally defined chromatin state. Nat Commun. 2015;6:10148. Guo H, Zhu P, Yan L, Li R, Hu B, Lian Y, Yan J, Ren X, Lin S, Li J, et al. The DNA methylation landscape of human early embryos. Nature. 2014;511(7511):606–10. Messerschmidt DM, Knowles BB, Solter D. DNA methylation dynamics during epigenetic reprogramming in the germline and preimplantation embryos. Genes Dev. 2014;28(8):812–28. Lindeman LC, Andersen IS, Reiner AH, Li N, Aanes H, Østrup O, Winata C, Mathavan S, Müller F, Aleström P, et al. Prepatterning of developmental gene expression by modified histones before zygotic genome activation. Dev Cell. 2011;21(6):993–1004. Ruthenburg AJ, Allis CD, Wysocka J. Methylation of lysine 4 on histone H3: intricacy of writing and reading a single epigenetic mark. Mol Cell. 2007;25(1):15–30. Ping G, Ji-Han X, Csilla S, Xiao-Ming D, Qin Z, Yuehong Y, Peng Z, Sara Pereira C, Kai Z, Jing Z et al. Biology and Clinical Implications of the 19q13 Aggressive Prostate Cancer Susceptibility Locus. Cell 2018, 174(3). I K. N S, D S, A P, H F, S L: Oxygen utilization during isovolumic pressure-volume loading: effects of prolonged extracorporeal circulation and cardioplegic arrest. Ann Thorac Surg 1986, 41(4). Bingjie Z, Hui Z, Bo H, Wenzhi L, Yunlong X, Xu P, Jia M, Xiaotong W, Yu Z, Qianhua X et al. Allelic reprogramming of the histone modification H3K4me3 in early mammalian development. Nature 2016, 537(7621). Du Z, Zheng H, Huang B, Ma R, Wu J, Zhang X, He J, Xiang Y, Wang Q, Li Y, et al. Allelic reprogramming of 3D chromatin architecture during early mammalian development. Nature. 2017;547(7662):232–5. Hug CB, Grimaldi AG, Kruse K, Vaquerizas JM. Chromatin Architecture Emerges during Zygotic Genome Activation Independent of Transcription. Cell. 2017;169(2):216–e228219. Shaulian E, Karin M. AP-1 as a regulator of cell life and death. Nat Cell Biol. 2002;4(5):E131–136. Wagner EF. Functions of AP1 (Fos/Jun) in bone development. Ann Rheum Dis. 2002;61(Suppl 2):ii40–42. Kaczynski J, Cook T, Urrutia R. Sp1- and Krüppel-like transcription factors. Genome Biol. 2003;4(2):206. ALFA. Allele Frequency Aggregator. [ Köhler S, Gargano M, Matentzoglu N, Carmody LC, Lewis-Smith D, Vasilevsky NA, Danis D, Balagura G, Baynam G, Brower AM, et al. The Human Phenotype Ontology in 2021. Nucleic Acids Res. 2021;49(D1):D1207–17. Zhao M, Chen L, Qu H. CSGene: a literature-based database for cell senescence genes and its application to identify critical cell aging pathways and associated diseases. Cell Death Dis. 2016;7(1):e2053. Liu A, Yu X, Liu S. Pluripotency transcription factors and cancer stem cells: small genes make a big difference. Chin J cancer. 2013;32(9):483–7. Abbott KL, Nyre ET, Abrahante J, Ho YY, Isaksson Vogel R, Starr TK. The Candidate Cancer Gene Database: a database of cancer driver genes from forward genetic screens in mice. Nucleic Acids Res. 2015;43(Database issue):D844–848. Lawrence MS, Stojanov P, Mermel CH, Robinson JT, Garraway LA, Golub TR, Meyerson M, Gabriel SB, Lander ES, Getz G. Discovery and saturation analysis of cancer genes across 21 tumour types. Nature. 2014;505(7484):495–501. Vogelstein B, Papadopoulos N, Velculescu VE, Zhou S, Diaz LA Jr., Kinzler KW. Cancer genome landscapes. Sci (New York NY). 2013;339(6127):1546–58. Schroeder MP, Rubio-Perez C, Tamborero D, Gonzalez-Perez A, Lopez-Bigas N. OncodriveROLE classifies cancer driver genes in loss of function and activating mode of action. Bioinf (Oxford England). 2014;30(17):i549–555. Zehir A, Benayed R, Shah RH, Syed A, Middha S, Kim HR, Srinivasan P, Gao J, Chakravarty D, Devlin SM, et al. Mutational landscape of metastatic cancer revealed from prospective clinical sequencing of 10,000 patients. Nat Med. 2017;23(6):703–13. Sondka Z, Bamford S, Cole CG, Ward SA, Dunham I, Forbes SA. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat Rev Cancer. 2018;18(11):696–705. Martínez-Jiménez F, Muiños F, Sentís I, Deu-Pons J, Reyes-Salazar I, Arnedo-Pac C, Mularoni L, Pich O, Bonet J, Kranas H et al. A compendium of mutational cancer driver genes. Nat Rev Cancer 2020. Dongre A, Weinberg RA. New insights into the mechanisms of epithelial-mesenchymal transition and implications for cancer. Nat Rev Mol Cell Biol. 2019;20(2):69–84. Zheng G, Ma Y, Zou Y, Yin A, Li W, Dong D. HCMDB: the human cancer metastasis database. Nucleic Acids Res. 2018;46(D1):D950–5. Picarda E, Ohaegbulam KC, Zang X. Molecular Pathways: Targeting B7-H3 (CD276) for Human Cancer Immunotherapy. Clin cancer research: official J Am Association Cancer Res. 2016;22(14):3425–31. Boelens MC, Wu TJ, Nabet BY, Xu B, Qiu Y, Yoon T, Azzam DJ, Twyman-Saint Victor C, Wiemann BZ, Ishwaran H, et al. Exosome transfer from stromal to breast cancer cells regulates therapy resistance pathways. Cell. 2014;159(3):499–513. Iwaya T, Fukagawa T, Suzuki Y, Takahashi Y, Sawada G, Ishibashi M, Kurashige J, Sudo T, Tanaka F, Shibata K, et al. Contrasting expression patterns of histone mRNA and microRNA 760 in patients with gastric cancer. Clin cancer research: official J Am Association Cancer Res. 2013;19(23):6438–49. Ma Y, Zhang P, Wang F, Yang J, Yang Z, Qin H. The relationship between early embryo development and tumourigenesis. J Cell Mol Med. 2010;14(12):2697–701. Smith DG, Sturmey RG. Parallels between embryo and cancer cell metabolism. Biochem Soc Trans. 2013;41(2):664–9. Bernardi Y, Strobl-Mazzulla PH. What we can learn from embryos to understand the mesenchymal-to-epithelial transition in tumor progression. Biochem J. 2021;478(9):1809–25. Wieland E, Rodriguez-Vita J, Liebler SS, Mogler C, Moll I, Herberich SE, Espinet E, Herpel E, Menuchin A, Chang-Claude J, et al. Endothelial Notch1 Activity Facilitates Metastasis. Cancer Cell. 2017;31(3):355–67. Zhang S, Chitu V, Stanley ER, Elliott BE, Greer PA. Fes tyrosine kinase expression in the tumor niche correlates with enhanced tumor growth, angiogenesis, circulating tumor cells, metastasis, and infiltrating macrophages. Cancer Res. 2011;71(4):1465–73. Dobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, Gingeras TR. STAR: ultrafast universal RNA-seq aligner. Bioinf (Oxford England). 2013;29(1):15–21. Trapnell C, Roberts A, Goff L, Pertea G, Kim D, Kelley DR, Pimentel H, Salzberg SL, Rinn JL, Pachter L. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat Protoc. 2012;7(3):562–78. Lazar C, Meganck S, Taminau J, Steenhoff D, Coletta A, Molter C, Weiss-Solís DY, Duque R, Bersini H, Nowé A. Batch effect removal methods for microarray gene expression data integration: a survey. Brief Bioinform. 2013;14(4):469–90. Dennis G Jr., Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. DAVID: Database for Annotation, Visualization, and Integrated Discovery. Genome Biol. 2003;4(5):P3. Wu T, Hu E, Xu S, Chen M, Guo P, Dai Z, Feng T, Zhou L, Tang W, Zhan L, et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innov (New York NY). 2021;2(3):100141. Zhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, Benner C, Chanda SK. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10(1):1523. Han H, Cho JW, Lee S, Yun A, Kim H, Bae D, Yang S, Kim CY, Lee M, Kim E, et al. TRRUST v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic Acids Res. 2018;46(D1):D380–6. Schmitt A, Hu M, Jung I, Xu Z, Qiu Y, Tan C, Li Y, Lin S, Lin Y, Barr C. A Compendium of Chromatin Contact Maps Reveals Spatially Active Regions in the Human Genome. Cell Rep. 2016;17(8):2042–59. Tang Z, Kang B, Li C, Chen T, Zhang Z. GEPIA2: an enhanced web server for large-scale expression profiling and interactive analysis. Nucleic Acids Res. 2019;47(W1):W556–60. Zheng H, Huang B, Zhang B, Xiang Y, Du Z, Xu Q, Li Y, Wang Q, Ma J, Peng X, et al. Resetting Epigenetic Memory by Reprogramming of Histone Modifications in Mammals. Mol Cell. 2016;63(6):1066–79. Wang C, Liu X, Gao Y, Yang L, Li C, Liu W, Chen C, Kou X, Zhao Y, Chen J, et al. Reprogramming of H3K9me3-dependent heterochromatin during mammalian embryo development. Nat Cell Biol. 2018;20(5):620–31. Dahl JA, Jung I, Aanes H, Greggains GD, Manaf A, Lerdrup M, Li G, Kuan S, Li B, Lee AY, et al. Broad histone H3K4me3 domains in mouse oocytes modulate maternal-to-zygotic transition. Nature. 2016;537(7621):548–52. Wang L, Zhang J, Duan J, Gao X, Zhu W, Lu X, Yang L, Zhang J, Li G, Ci W, et al. Programming and inheritance of parental DNA methylomes in mammals. Cell. 2014;157(4):979–91. Luo Y, Hitz BC, Gabdank I, Hilton JA, Kagda MS, Lam B, Myers Z, Sud P, Jou J, Lin K, et al. New developments on the Encyclopedia of DNA Elements (ENCODE) data portal. Nucleic Acids Res. 2020;48(D1):D882–9. Duan J, Zhu L, Dong H, Zheng X, Jiang Z, Chen J, Tian XC. Analysis of mRNA abundance for histone variants, histone- and DNA-modifiers in bovine in vivo and in vitro oocytes and embryos. Sci Rep. 2019;9(1):1217. Li B, Qing T, Zhu J, Wen Z, Yu Y, Fukumura R, Zheng Y, Gondo Y, Shi L. A Comprehensive Mouse Transcriptomic BodyMap across 17 Tissues by RNA-seq. Sci Rep. 2017;7(1):4200. Additional Declarations No competing interests reported. Supplementary Files supplementarytables.rar Supplementary tables. S1 to S5 SupplementaryFiguresRevised.pptx Figs. S1 to S8 Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 08 May, 2026 Reviews received at journal 17 Apr, 2026 Reviewers agreed at journal 07 Apr, 2026 Reviews received at journal 03 Apr, 2026 Reviewers agreed at journal 29 Mar, 2026 Reviewers invited by journal 29 Mar, 2026 Editor assigned by journal 24 Mar, 2026 Submission checks completed at journal 23 Mar, 2026 First submitted to journal 23 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9174297","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":615707587,"identity":"ba3dda90-962e-4613-bcb9-87c2befd3308","order_by":0,"name":"Han Xing","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA1ElEQVRIiWNgGAWjYDACZjACggPMBw58+EGaFrbEgzN7iLcIpIXH+DAHGxHKDY7zHn5dUHPHnu9GzofDDDwM8vxiB/BrkWzmS7OecexZ4swbuRsOF1gwGM6cnYBfCz8zj5kxD9vhBAOQlhk8DAkGtwloYQNr+XfY3uBGzoPDPGxEaAHaYvyYt+0w44YbOQzEaZFs5jFj5u07nDjzzDMDYCBLEPaLwfkzxp95vh225zue/PjDhx828vzSBLSAvCOBxJHAqQwZMH8gStkoGAWjYBSMXAAAGFtGGF40LLwAAAAASUVORK5CYII=","orcid":"","institution":"Shandong University","correspondingAuthor":true,"prefix":"","firstName":"Han","middleName":"","lastName":"Xing","suffix":""},{"id":615707589,"identity":"dc89ae9a-e144-4e9e-8b58-49141b050c05","order_by":1,"name":"Yongjian Zhao","email":"","orcid":"","institution":"Shandong University","correspondingAuthor":false,"prefix":"","firstName":"Yongjian","middleName":"","lastName":"Zhao","suffix":""}],"badges":[],"createdAt":"2026-03-20 03:23:28","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9174297/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9174297/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106095235,"identity":"2d0e920c-eef5-4954-afd1-1b11f7a5db04","added_by":"auto","created_at":"2026-04-03 11:46:37","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":2281193,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDissecting gene expression pattern during embryonic development and identify ZGA genes. \u003c/strong\u003e(A) Hierarchical clustering for samples of mouse (GEO: GSE66390 and ENCODE), human (GEO:GSE101571 and ENCODE), rhesus monkey (GEO: GSE103313) and bovine (GEO: GSE59186) early embryos. The red line represents the timing of ZGA. hES: human embryonic stem, GV: germinal vesicle, MI oocyte: metaphase I oocyte, 2PN: 2 pronuclei, 3PN: tripronuclear. (B) Schematic of the framework. Formulas depicting the recursive distance update rule (left) and the pairwise average distance metric (right) employed in the WPGMA hierarchical clustering. (C) Expression patterns (FPKM) of 6 classes identified using expression profiles from bulk RNA-seq on mouse early developmental stages and tissues. (D) Expression patterns (FPKM) of the 6 classes on mouse early developmental stages and tissues. (E) Bar plot of GO term enrichment for genes in 6 classes.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/5d7de50d2878ac0d745ac8ae.png"},{"id":106095186,"identity":"3bc08588-7297-4561-be89-d4e17155a1dc","added_by":"auto","created_at":"2026-04-03 11:46:31","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":692959,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eComparison of identified mouse ZGA genes.\u003c/strong\u003e (A) Expression level in normal 2-cell and 2-cell embryos treated with α-Amanitin in 6 classes identified from bulk RNA-seq data (GEO: GSE71434). Expression level is log2 (FPKM + 1). (B) ZGA gene expression fold changes between 2-cell embryos treated with α-Amanitin and normal late 2-cell embryos. Expression level is log2 (FPKM + 1). (C) Overlap of major-ZGA genes identified (class 4 and class 5) using our framework and previous studies. (D) Expression levels (FPKM) of ZGA genes exclusively identified using our framework in early developmental stages.\u003c/p\u003e","description":"","filename":"floatimage22.png","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/866f316b0a3360e7d0679ea9.png"},{"id":106095602,"identity":"ff7bcbd3-43f1-44a2-a073-495593c83719","added_by":"auto","created_at":"2026-04-03 11:50:02","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":459280,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGene expression pattern is conserved during embryonic development across diverse datasets.\u003c/strong\u003e (A) Expression patterns of the 6 classes (from Fig. 1D) in different developmental stages in single-cell resolution on mouse (GEO: GSE45719). (B) Overlap of the genes in 6 classes identified with 8, 16, and 24 tissue samples (ENCODE) on human.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/c3f065e79d180ac38c328209.png"},{"id":106095697,"identity":"41b08f28-eb6f-4bfd-b1fe-eb2abc7395ec","added_by":"auto","created_at":"2026-04-03 11:50:34","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":127915,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe features and epigenetic properties of mouse ZGA genes. \u003c/strong\u003e(A) Expression patterns (FPKM) of genes in 6 classes from grown oocytes (GO) to fully grown oocytes (FGO) stage (GEO: GSE112835). Each row is a gene and each column is a stage. Each value is the FPKM of the gene. (B) Percentages of housekeeping genes in the 6 classes. (C) Epigenetic profiles of mouse S-ZGA genes (n = 571, resouces: GSE66390, GSE71434, GSE76687, GSE98149, GSE56697 and GSE72784). The average epigenetic modification signal levels of ±5 kb around TSS regions of S-ZGA genes was shown at each stage. (D-E) Boxplot of ATAC-seq signal in ±1 kb around TSS regions of genes in 6 classes from early 2-cell to ICM stage. E2C: early 2-cell, 2C: 2-cell, 4C: 4-cell, 8C: 8-cell, ICM: inner cell mass. (D) Signals from genes originally assigned to Class 1–6 based on unsupervised clustering of transcriptomic profiles (see Methods). (E) Signals from a control set of genes matched for expression level (FPKM) at each stage but randomly selected outside the original class definitions. (F) Percentages of promoters (±2 kb around TSS) in the TAD boundary at 2-cell stage. The numbers in the bar plot are the number of mouse promoters in TAD boundary (GSE82185\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/71710d288d1e1f18cd92fcf5.png"},{"id":106401885,"identity":"53558e31-3e60-4754-ace0-4366ad1d5d3f","added_by":"auto","created_at":"2026-04-08 09:10:11","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":6114438,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFunctional properties of human ZGA genes.\u003c/strong\u003e (A) Percentages of cancer driver genes in the 6 classes that we identified on human (Fig. 3B). The x-axis represents the six identified gene classes (Class 1–6), and the y-axis indicates the sources of cancer driver gene sets reported in previous literature. Dot color reflects the significance of enrichment, while dot size corresponds to the number of genes. (B) GO term enrichment for genes and related genes in human genes. (C) Effect of \u003cem\u003eCCND1\u003c/em\u003e(left), \u003cem\u003eELF3\u003c/em\u003e (middle), and\u003cem\u003e RGL3\u003c/em\u003e (right) knockdown on proliferation of SKOV3 cells detected by CCK8 assays. (D) Representative images of SKOV3 cells treated with siRNA targeting\u003cem\u003e ELF3\u003c/em\u003e or negative control in migration and invasion assays(Left), Quantification(Right) represent mean ± SD. Unpaired t test, *p \u0026lt; 0.05,**p \u0026lt; 0.01, ***p \u0026lt; 0.001, Scale bar, 100μm. (E) Representative images of SKOV3 cells treated with siRNA targeting RGL3 or negative control in migration and invasion assays(Left), Quantification(Right) represent mean ± SD. Unpaired t test, *p \u0026lt; 0.05,**p \u0026lt; 0.01, ***p \u0026lt; 0.001, Scale bar, 100μm.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/3d334adc20709026d5bfbe51.png"},{"id":106405525,"identity":"fbf10ccf-962d-430e-8b20-6532ad85829b","added_by":"auto","created_at":"2026-04-08 09:27:04","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":12556161,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/dd7fb666-ac98-4767-a0bc-8aaea32d3b78.pdf"},{"id":106095185,"identity":"9dbdf228-d72f-44d3-a18f-557870cbc46e","added_by":"auto","created_at":"2026-04-03 11:46:30","extension":"rar","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2230843,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary tables. S1 to S5\u003c/p\u003e","description":"","filename":"supplementarytables.rar","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/562c89b016412035a346eeb1.rar"},{"id":106095244,"identity":"f4ee1fda-9bb0-4faa-ac01-578bc9583154","added_by":"auto","created_at":"2026-04-03 11:46:48","extension":"pptx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":4797503,"visible":true,"origin":"","legend":"\u003cp\u003eFigs. S1 to S8\u003c/p\u003e","description":"","filename":"SupplementaryFiguresRevised.pptx","url":"https://assets-eu.researchsquare.com/files/rs-9174297/v1/460de94427645c9682c97343.pptx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Dissecting the Expression Pattern during Embryonic Development and Unveiling Sustained ZGA Genes with Oncogenic Relevance","fulltext":[{"header":"Background","content":"\u003cp\u003eDuring the early stages of mammalian embryonic development, the fertilized egg gives rise to all cell types that form the adult organism through sequential rounds of proliferation and differentiation [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. This period is marked by a series of coordinated events including maternal mRNA clearance, zygotic genome activation (ZGA), cell polarization, asymmetric division cycles, and the establishment of gene expression patterns governing the first and second cell fate decisions [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Central to these events is the maternal-to-zygotic transition (MZT) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e], marked by the shift from maternal to zygotic transcriptional control through embryo-wide ZGA [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. In metazoan embryos, transcription is mostly silent for the initial cell divisions until ZGA is initiated [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. ZGA manifests as two distinct transcriptional waves: a minor wave during early cleavage (producing limited coding transcripts and non-productive RNA via promiscuous transcription) and a major wave coinciding with the cell cycle arrest in many species [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. While the timing and number of divisions preceding ZGA vary across species, the process is highly conserved and precisely timed within each species [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. In mice, minor ZGA initiates at mid-to-late 1-cell stage, with major ZGA activating thousands of genes at the late 2-cell stage [\u003cspan additionalcitationids=\"CR10 CR11 CR12\" citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. In humans, low-level transcription is detected from the zygote to the 4-cell embryo stage, and the major wave occurs at the 8-cell stage [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eZGA is essential for orchestrating gastrulation, germ-layer specification, and cell fate determination, with ZGA dysregulation leading to developmental failure [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. For instance, the presence of one of the important ZGA genes, \u003cem\u003eNanog\u003c/em\u003e, is essential to initiate the zygotic developmental program and induces clearance of the maternal program in zebrafish, thereby preventing developmental arrest before gastrulation [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. In mice, Nanog encodes a homeoprotein critical for maintaining pluripotency in the inner cell mass (ICM) and embryonic stem (ES) cells [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Combined loss of \u003cem\u003eNanog\u003c/em\u003e and other ZGA regulators results in embryonic lethality before gastrulation [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. ZGA is a complicated process involving many important types of biological activation. However, Nanog is not the only regulator; numerous studies over the past few years have identified additional ZGA regulators that function within a complex hierarchical network. Notably, the homologous box protein Dux (and human DUX4) acts as a ZGA \u0026ldquo;accelerator,\u0026rdquo; enhancing ZGA during early embryonic development. When \u003cem\u003eDux\u003c/em\u003e is absent, activation of certain ZGA genes is delayed, and dose effects of alternative factors may lead to divergent developmental trajectories in \u003cem\u003eDux\u003c/em\u003e knockout embryos [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Unlike the partial compensation seen in \u003cem\u003eDux\u003c/em\u003e mutants, Obox factors simultaneously regulate both minor and major ZGA waves in mouse embryos, and their depletion leads to severe ZGA impairment [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Mechanistically, OBOX facilitates the \u0026ldquo;pre-loading\u0026rdquo; of RNA Polymerase II (Pol II), promoting its migration from initial sites to the promoters and enhancers of ZGA genes. Therefore, systematic identification of genes activated during ZGA is indispensable for unraveling its regulatory mechanisms and deciphering the molecular basis of early embryonic development.\u003c/p\u003e \u003cp\u003eCurrent methodologies for identifying zygotic genome activation (ZGA) genes face critical limitations. Experimental approaches, such as α-Amanitin and aphidicolin treatments [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan additionalcitationids=\"CR22\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], enable precise ZGA gene detection but are resource-intensive and low-throughput [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], Conversely, data-driven methods comparing pre- and post-ZGA transcriptomes [\u003cspan additionalcitationids=\"CR25\" citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] facilitate high-throughput screening yet suffer from noise-induced false positives [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Crucially, both paradigms depend on manually defined ZGA timing, which varies substantially across species [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], resulting in inconsistent ZGA gene identification across datasets. This species-dependent timing discrepancy, coupled with poor reproducibility in existing studies, underscores the urgent need for robust, prior knowledge-independent ZGA identification frameworks.\u003c/p\u003e \u003cp\u003eGiven that ZGA represents a dynamic process rather than a discrete event, populations of ZGA gene transcripts straddle the period of ZGA [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e, \u003cspan additionalcitationids=\"CR30\" citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Single-stage RNA-seq analyses are inherently prone to technical noise, limiting reliable gene identification. Leveraging time-series RNA-seq data, we reconstructs dynamic developmental trajectories [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e] to identify ZGA genes through temporal expression patterns rather than static pre/post-ZGA comparisons. This approach enhances accuracy and robustness by capturing gene activation dynamics across developmental stages, providing mechanistic insights into embryonic genome activation [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Moreover, we uncovered a key subclass of ZGA genes called S-ZGA genes, defined by their activation at ZGA and sustained expression post-activation. We characterized their epigenetic signatures, revealing their potential link to clinical prognosis in multiple cancers. These findings bridge early embryonic development and oncogenesis, suggesting that S-ZGA genes, critical for developmental programs, may also contribute to tumor progression when dysregulated.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDissecting gene expression pattern during embryonic development\u003c/h2\u003e \u003cp\u003eZGA onset varies significantly across species, posing challenges for manual determination. To dissect ZGA features, we first investigated the gene expression pattern during mouse embryonic development. We obtained gene expression data from 8 mouse early embryonic periods from a previous study [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e] and 8 mouse different tissues with two replicates for each tissue from the ENCODE project and performed hierarchical clustering. Gene expression exhibits three distinct clusters corresponding to: (1) oocyte to early 2-cell embryos, (2) 2-cell to embryonic stem cell (ESC) and (3) ESC to somatic tissue (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eA). The transition from first to second cluster well aligns with known ZGA onset at the 2-cell embryo in mice [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e], and also confirmed by PCA analysis (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eB). This phenomenon is not an isolated occurrence in mice, but conserved across different species including human [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e], rhesus monkey [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] and bovine [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e], as their RNA-seq data showed distinct expression patterns (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA), which are consistent with previous studies reporting that in these species ZGA occurs [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan additionalcitationids=\"CR35\" citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. These findings collectively demonstrate that species-specific ZGA timing is reliably reflected in developmental RNA-seq expression patterns, establishing a robust foundation for automated ZGA event identification.\u003c/p\u003e \u003cp\u003eInspired by these findings, we analyzed the time-series RNA-seq data (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB) to precisely identify ZGA onset time and key genes. This framework begins with hierarchical clustering of developmental samples based on gene expression profiles-a standard approach for quantifying sample similarity in biological studies, and clustered cells from different developmental stages based on similar expression profiles. Gene expression measured by RNA-seq is often noisy due to technical error and computational bias [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] and will affect the accuracy of current ZGA gene identification methods, which are mainly based on the fold change of expression at two developmental stages. To overcome this limitation, we implement a three-layer filtering pipeline that analyzes gene expression across all developmental stages. The framework thus categorizes genes into eight distinct classes based on their expression patterns: class 0 genes show ubiquitous expression in early embryos and adult tissues (likely essential housekeeping genes), while class 7 genes remain largely silent throughout embryonic development (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eC). Critically, genes in classes 1\u0026ndash;6 exhibit dynamic expression patterns coinciding with ZGA transitions, making them the primary focus for identifying ZGA-specific regulatory genes.\u003c/p\u003e \u003cp\u003eIn mice, we identified six gene classes (89, 195, 316, 571, 119, and 1383 genes, respectively) with developmental stage-specific expression patterns (Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eD, supplementary table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Classes 1 and 2 characterized maternal factors. Class 3 genes exhibited minor ZGA signatures, while classes 4 and 5 defined major ZGA, and class 6 genes were exclusively expressed in differentiated tissues (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC). Within the complex spatiotemporal transcriptional network, embryos employ an evolutionary strategy of gene co-option [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], utilizing epigenetic mechanisms (e.g., DNA methylation, histone modifications) to differentially manage these maternal factors[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Specific regulators are degraded and permanently silenced, marking the irreversible exit from oogenic programs, whereas a distinct subset of maternal pioneer factors exhibits a dynamic \u0026ldquo;degrade-silence-reactivate\u0026rdquo; trajectory; initially deposited to prime the zygotic genome, they are actively suppressed post-ZGA to prevent ectopic activity, yet their circuits are preserved in a poised epigenetic state for future recruitment in specific lineages. Accordingly, we further subdivided these maternal factors into those permanently silenced after ZGA and those reactivated in mature tissues, distinguishing between \u0026ldquo;one-time developmental components\u0026rdquo; and \u0026ldquo;multifunctional developmental components\u0026rdquo; [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. We designated minor-ZGA (class 3) and major-ZGA (classes 4 and 5) genes as ZGA genes thereafter. Functional enrichment analysis via DAVID revealed: class 3 (minor-ZGA) were mainly enriched in cell division, chromosome segregation and meiotic cell cycle, class 2 (maternal factors) were principally implicated in regulation of inflammatory response and oocyte maturation, class 4 (major-ZGA) were mainly enriched in protein transport and apoptotic process, and class 6 (tissue-specific genes) were predominantly associated with more complex biological processes, such as the angiogenesis, inflammatory response, and extracellular matrix organization (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eIdentified ZGA genes demonstrate critical repression during α-Amanitin treatment\u003c/h3\u003e\n\u003cp\u003eα-Amanitin is a specific inhibitor of RNA polymerase II (RNA Pol II), which mediates transcription of protein-coding genes. In preimplantation embryos, α-Amanitin effectively blocks zygotic transcription without globally disrupting maternal mRNA stability. To assess whether the genes identified by our framework require de novo transcription, we analyzed publicly available RNA-seq data from control and α-amanitin-treated 2-cell embryos.[\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. ZGA-associated genes (classes 3, 4, and 5) exhibited significant downregulation upon α-Amanitin treatment (Wilcoxon signed-rank test, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA), whereas non-ZGA genes (classes 1, 2, and 6) showed no such reduction. Next, we collected the scRNA-seq data for samples in 2-cell embryos and 2-cell embryos with α-Amanitin. High inter-replicate expression correlation confirmed data reliability and absence of outliers (Fig. S2A), and further analysis of these scRNA-seq datasets revealed that only major ZGA genes (classes 4 and 5) were substantially downregulated (\u0026gt;\u0026thinsp;20-fold) by α-Amanitin (Fig. S2B), consistent with bulk RNA-seq results. These results demonstrate that ZGA genes our framework identified are transcriptionally activated during zygotic genome activation.\u003c/p\u003e \u003cp\u003eIt is worth noting that α-amanitin sensitivity is not absolute for all zygotic transcripts. Certain ZGA components exhibit partial resistance-a minority of genes in classes 4 and 5 retained residual expression upon treatment (Fig. S2B) [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. This may reflect transcripts with longer maternal contributions, higher intrinsic stability, early activation prior to complete drug penetration, or the previously described \u0026ldquo;promiscuous\u0026rdquo; minor ZGA transcription that is less dependent on canonical RNA Pol II activity. Thus, while our findings confirm that the majority of ZGA genes require active transcription, they also align with the emerging view that ZGA encompasses both highly sensitive \u0026ldquo;core\u0026rdquo; transcriptional events and more resilient programs.\u003c/p\u003e \u003cp\u003eWe next assessed gene expression dynamics in α-Amanitin-treated versus control 2-cell embryos with other ZGA identification methods [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan additionalcitationids=\"CR47\" citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e]. In our ZGA genes, major-ZGA genes (classes 4 and 5) exhibited the most pronounced transcriptional suppression (residual log2[FPKM\u0026thinsp;+\u0026thinsp;1]\u0026thinsp;\u0026lt;\u0026thinsp;0.5 in treated embryos; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB), with significantly greater downregulation than observed using other approaches (Wilcoxon signed-rank test, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05). Across all methods, 54 major-ZGA genes were consistently identified (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC, supplementary table 2). For example, \u003cem\u003eMta2\u003c/em\u003e was a typical major-ZGA gene that initiated transcription at the 2-cell stage (Fig. S2C) and was markedly suppressed by α-Amanitin in 2-cell embryos (Fig. S2D). Targeted mutation of the \u003cem\u003eMta2\u003c/em\u003e gene results in multiple phenotypes that include partial embryonic and perinatal lethality, female infertility, abnormal T-cell activation, and lupus-like autoimmune disease in mice [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e]. Critically, we identified 119 novel major-ZGA genes absent from other methods (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC, supplementary table 2). These genes activated specifically at the 2-cell stage and maintained elevated expression throughout early embryonic development (Fig. S2E). To illustrate the biological relevance of these findings, we selected representative examples encompassing both canonical markers to validate our method and novel candidates to demonstrate its discovery potential. Some of these genes were supported as major-ZGA genes by literature, such as \u003cem\u003eXist\u003c/em\u003e and \u003cem\u003eNanog\u003c/em\u003e [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. \u003cem\u003eXist\u003c/em\u003e is an imprinted gene expressed in early embryos (beginning at the 2-cell stage) [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e], and its expression remained high at the 8-cell stage (Fig. S2F). \u003cem\u003eXist\u003c/em\u003e triggers gene silencing and drives major structural reorganization of the X chromosome, which allows female mammalian embryos to survive development by establishing and maintaining proper dosage compensation [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e]. Nanog is essential to initiate the zygotic developmental program and induces clearance of the maternal program, thereby preventing developmental arrest before gastrulation [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], and it showed a similar expression profile trend during ZGA (Fig. S2G). In addition to these classical genes, we quantified their biological relevance through rigorous functional annotation and literature screening of 119 identified S-ZGA-specific genes. Approximately 17% (20/119) of the genes were directly involved in classical early mammalian developmental processes, such as \"lamina labyrinthina development\" and \"cell fate determination during primary germ layer formation.\" Another 9% (11/119) were frequently annotated as proteins related to non-classical metabolic or stress pathways, playing definitive key roles in early embryogenesis of vertebrate models including zebrafish and African clawed frog. The remaining 74% represented novel candidate factors. Others like \u003cem\u003eMvb12a\u003c/em\u003e, \u003cem\u003ePdf\u003c/em\u003e, \u003cem\u003eUqcc3\u003c/em\u003e, \u003cem\u003eCcpg1os\u003c/em\u003e, \u003cem\u003eE530001F21Rik\u003c/em\u003e, \u003cem\u003eAA543401\u003c/em\u003e, and \u003cem\u003eC330021F23Rik\u003c/em\u003e showed ZGA-associated expression patterns (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD) and were significantly downregulated by α-Amanitin (Fig. S2H).\u003c/p\u003e \u003cp\u003eThe variations in ZGA gene sets between our framework and prior studies stem from these methodological differences, highlighting their complementary value. Previous studies, using pairwise comparisons or fold-change cutoffs, effectively capture genes with sharp, transient bursts typical of canonical ZGA. Indeed, major-ZGA genes identified in prior studies but not by our framework predominantly exhibited transient 2-cell-stage expression without sustained elevation (Fig. S2I); while authentic, these likely represent a distinct regulatory layer, such as non-canonical or transient transcriptional events, compared to the sustained activation program captured by our framework. Together, these complementary perspectives offer a more integrated view of the complex, multi-faceted nature of zygotic genome activation.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eGene expression pattern is conserved during embryonic development across diverse datasets\u003c/h3\u003e\n\u003cp\u003eWe next extended to datasets generated using diverse biological and technical methods. First, we examined gene expression patterns in single-cell RNA-seq (scRNA-seq) data from prior studies [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e]. The expression patterns of the ZGA genes were consistent in bulk and scRNA-seq data (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA), demonstrating our framework produces consistent outputs regardless of sequencing platform. We then analyzed human embryonic development and tissue datasets. The framework successfully resolved the six gene classes with distinct expression patterns in humans (Fig. S3A, supplementary table 3), confirming its robust cross-species adaptability.\u003c/p\u003e \u003cp\u003eGiven the complexity of human tissue samples, we systematically assessed whether varying sample selection or composition could bias results. By analyzing gene expression datasets containing 8, 16, or 24 tissue samples derived from human embryonic stem cells (hESCs), we found minimal variation in ZGA gene identification across these conditions (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB), underscoring the framework\u0026rsquo;s resilience to tissue-specific heterogeneity.\u003c/p\u003e \u003cp\u003eComparative analysis of ZGA gene classifications between humans and mice revealed species-specific differences in ZGA timing and gene repertoires (Fig. S3B). This divergence aligns with prior observations of distinct regulatory mechanisms governing ZGA in these species [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Collectively, these results demonstrate that expression patterns of ZGA genes are conserved within the same species across datasets and experimental conditions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eSustained ZGA genes are functionally enriched in housekeeping genes and sustain embryonic competence\u003c/h3\u003e\n\u003cp\u003eOverall, by applying our hierarchical filtering framework based on expression patterns, we identified 1,006 genes activated during the ZGA window in early mouse embryos-consisting of 316 minor-ZGA and 690 major-ZGA genes, and mapped their global functional profiles. These genes collectively govern essential early developmental processes, spanning from cell division and chromosome segregation in the minor-ZGA to protein transport and apoptosis in the major-ZGA (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE). While this entire cohort signifies the initiation of zygotic transcription, our analysis revealed marked heterogeneity in their temporal dynamics (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD). Specifically, unlike the transient activation peaks characteristic of most ZGA genes, a distinct subset within the major-ZGA (Class 4) displays a unique trajectory of sustained high expression that persists from the initial activation through later developmental stages and into mature tissues. Such persistent activity implies a specialized function for this subpopulation in preserving embryonic potential and sustaining basic cellular operations throughout the life cycle. Consequently, we further focus on class 4 (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD) as a critical subset of major ZGA genes. These genes are transcriptionally activated at ZGA and maintain sustained expression throughout early embryonic development, leading us to designate them Sustained ZGA (S-ZGA) genes (Fig. S4A).\u003c/p\u003e \u003cp\u003eWe next investigated expression dynamics of S-ZGA genes during oogenesis [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e], revealing a dramatic decline in S-ZGA expression during the transition from grown oocytes (GO) to fully grown oocytes (FGO) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). Notably, S-ZGA genes exhibit strong functional relevance: they are highly enriched among 4,781 housekeeping genes involved in basic cell maintenance and expected to maintain constant expression levels across all cells and conditions(Fisher's exact test, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1.2e-27; Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB) [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e]. This enrichment aligns with ZGA\u0026rsquo;s critical role in establishing pluripotency and genome stability [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. Rather than relying on individual examples, our conclusion is supported by the global trend that housekeeping genes within S-ZGA exhibit significantly higher expression than non-housekeeping genes across multiple developmental stages (Wilcoxon signed-rank test, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Fig. S4B). These findings establish S-ZGA genes as essential mediators of early embryonic competence through their dual role in ZGA activation and maintenance of core cellular functions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eThe epigenetic landscape of S-ZGA genes during early embryo development\u003c/h3\u003e\n\u003cp\u003eDuring mammal embryos development, epigenetic reprogramming and chromatin remodeling convert terminally differentiated gametes into totipotent or pluripotent cells. Promoter accessibility strongly correlates with gene expression levels in mouse embryos [\u003cspan additionalcitationids=\"CR56 CR57 CR58\" citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e]. To elucidate transcriptional regulation of S-ZGA genes, we analyzed chromatin accessibility on mouse promoters using ATAC-seq data from Wu et al [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. We observed a significant increase in ATAC-seq signal at S-ZGA gene promoters from the early 2-cell to 2-cell stage, followed by sustained high accessibility through subsequent developmental stages (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC-\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE). This pattern persisted in adult tissues (Fig. S4C) and S-ZGA gene promoters exhibited a substantially higher proportion within ATAC-seq peaks during ZGA in both mouse and human datasets (Fig. S4D, S4E). Species-specific chromatin remodeling strategies were suggested: mouse embryos exhibited a selective opening pattern, where accessibility increases were restricted to specific ZGA-associated gene clusters (Classes 3\u0026ndash;5), likely reflecting a rapid, targeted activation mechanism. In contrast, human embryos displayed a global and gradual increase in promoter accessibility across all gene classes, suggesting a broader, more permissive chromatin relaxation consistent with their slower developmental tempo. These findings directly correlate with the sustained expression profile of S-ZGA genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC) and previous studies [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe next characterized key epigenetic features of S-ZGA genes, including histone modification and DNA methylation [\u003cspan additionalcitationids=\"CR61\" citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e]. Histone modifications are fundamental epigenetic regulators that control many key cellular processes [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. Histone H3 lysine 4 trimethylation (H3K4me3), a canonical mark for active transcription initiation [\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e], showed significant enrichment at S-ZGA gene promoters and increased progressively during early embryonic development (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC, S4F). Crucially, this mark remained significantly elevated in S-ZGA promoters relative to other genes across both early embryonic stages and mature tissues (Fig. S4G, S4H). To distinguish whether these epigenetic signatures drive expression or merely reflect high transcriptional activity, we performed an expression-matched control analysis. For each gene class, we selected control genes from other classes with comparable expression levels (within a 50% difference threshold) across early developmental stages (Fig. S4I; see Methods for details). Comparing ATAC-seq and H3K4me3 signals between the target genes and this matched control set revealed distinct patterns: chromatin accessibility (ATAC-seq signal) remained significantly higher in S-ZGA genes even after controlling for expression, suggesting it is a prerequisite for activation and sustained expression of S-ZGA genes during ZGA (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eD, \u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eE) [\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e]. In contrast, the differential enrichment of H3K4me3 largely diminished in the matched comparison, indicating that H3K4me3 accumulation is more likely to be the result of transcriptional activation (Fig. S4J, S4K) [\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e, \u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e]. These results refine our understanding of the epigenetic regulation of S-ZGA genes. Furthermore, we investigated whether the enrichment of housekeeping genes (HKGs) in S-ZGA genes was merely a consequence of their high expression levels. By analyzing the proportion of HKGs within the expression-matched target and control sets, we found that Class 4 (S-ZGA) genes consistently exhibited the highest percentage of housekeeping functions compared to other classes and their expression-matched controls across all developmental stages (Fig. S4L). These results demonstrate that the enrichment of housekeeping functions in S-ZGA genes is a distinct biological feature, not merely a byproduct of high transcriptional activity.\u003c/p\u003e \u003cp\u003eIn addition, H3K27me3, H3K9me3, and DNA methylation at S-ZGA promoters remained consistently low during early embryogenesis, while H3K27ac-a mark of active enhancers-was increased at the 2-cell stage (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eC). Quantitatively, \u0026lt;\u0026thinsp;2% of S-ZGA promoters harbored H3K27me3 at 2-cell (Fig. S5A). At the 2-cell stage, the DNA methylation level of promoters was lowest in S-ZGA genes and the demethylation level was highest in class 3 (Fig. S5B, S5C). Concurrently, H3K9me3 density remained minimal at S-ZGA promoters (Fig. S5D, S5E). Biologically, this coordinated depletion of multiple repressive marks (DNA methylation, H3K27me3, and H3K9me3) suggests that S-ZGA promoters reside in a uniquely \u0026ldquo;permissive\u0026rdquo; chromatin state, effectively removing epigenetic barriers to facilitate the rapid binding of pioneer transcription factors (TFs) essential for ZGA initiation. Furthermore, Chromatin architecture analysis using Hi-C data [\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e] further revealed that S-ZGA gene promoters were significantly enriched at topologically associating domain (TAD) boundaries in 2-cell embryos compared to other gene classes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eF), a feature critical for transcriptional regulation [\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e]. By positioning S-ZGA genes at these boundaries, the embryo likely ensures their isolation from surrounding repressive domains while maximizing chromtine accessibility to the \u0026ldquo;transcriptional machines\u0026rdquo;-TFs.\u003c/p\u003e \u003cp\u003eThese integrated epigenetic and structural features-low repressive marks, upregulated H3K4me3/H3K27ac, and TAD boundary localization-demonstrate that establish a \u0026ldquo;poised-to-active\u0026rdquo; landscape. This unique configuration underpins the sustained and synchronous expression of S-ZGA genes, enabling them to serve as the primary drivers of the zygotic genome activation program.\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eS-ZGA genes are associated with tumorigenesis and metastasis\u003c/h2\u003e \u003cp\u003eWe next investigated functional roles of S-ZGA genes across the organismal life cycle. Human S-ZGA genes demonstrated significant enrichment in core developmental processes, including extracellular matrix organization, morphogenesis-linked growth regulation, and cell size control (Fig S6A, supplementary table 4). These genes were also enriched among targets of key transcription factors (TFs), such as JUN, FOS, and SP1 (Fig S6B, supplementary table 5), which orchestrate fundamental cellular activities such as proliferation, differentiation, survival, and death [\u003cspan additionalcitationids=\"CR70\" citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e]. We further examined S-ZGA gene associations with genetic variation, aging, and disease. We collected single nucleotide polymorphism (SNP) recorded from NCBI Allele Frequency Aggregator (ALFA) [\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e] (see Methods, supplementary table 5), and found that promoter regions of S-ZGA genes exhibited significantly reduced SNP density (Wilcoxon rank-sum test, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Fig. S6C), suggesting stringent regulatory constraints. Concurrently, S-ZGA genes were markedly enriched in human phenotype-associated genes [\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e] and senescence-related genes [\u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e] (Fisher's exact test, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.001; Fig. S6D, Fig. S6E; supplementary table 5). For example, \u003cem\u003eGBA\u003c/em\u003e, which ranks second highest within the S-ZGA set with 230 associated phenotypes (Supplementary Table\u0026nbsp;5), correlates with a broad spectrum of conditions ranging from developmental disorders (e.g., intrauterine growth retardation and pulmonary hypoplasia) to aging-related neurodegeneration (e.g., parkinsonism and dementia) [\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e]. These findings reveal that S-ZGA genes serve dual roles: they drive essential developmental processes while concurrently contributing to aging and disease susceptibility.\u003c/p\u003e \u003cp\u003eMoreover, aberrant expression of some ZGA genes has been associated with numerous forms of cancer, since the ability of genome activators to reprogram cells toward pluripotency also makes them potent drivers of tumor development [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e]. To investigate this connection and its specificity across ZGA classes, we analyzed S-ZGA gene associations with cancer. Human S-ZGA genes (class 4) were the only significantly enriched among cancer driver genes [\u003cspan additionalcitationids=\"CR77 CR78 CR79 CR80 CR81\" citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e82\u003c/span\u003e] (Fisher's exact test, p\u0026thinsp;\u0026lt;\u0026thinsp;0.05; Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA), a finding that mirrors their unique biological functional enrichment in apoptosis regulation (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE). Moreover, human S-ZGA genes were enriched in epithelial to mesenchymal transition (EMT) and related biological processes (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB, supplementary table 4), a program critical for embryogenesis, wound healing and malignant progression [\u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e83\u003c/span\u003e]. The co-enrichment of EMT and apoptosis regulation pathways in S-ZGA genes suggests a coordinated mechanism for metastasis: while EMT facilitates cell detachment and invasion, the concurrent suppression of apoptosis (particularly anoikis) This dual functionality ensures that disseminating cells survive the stress of circulation and successfully colonize distant organs distant organs. We observed differential expression of S-ZGA genes (e.g., \u003cem\u003eCCND1\u003c/em\u003e, \u003cem\u003eRGL3\u003c/em\u003e, and \u003cem\u003eELF3\u003c/em\u003e) between tumor and normal tissues (Fig. S7A), with elevated expression in metastatic primary tumors compared to non-metastatic counterparts (Fig S7B) [\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e]. Survival analyses further revealed that high expression of specific S-ZGA genes was associated with significantly worse prognosis across multiple cancer types. Among these, \u003cem\u003eHIST1H3D\u003c/em\u003e, \u003cem\u003eNOTCH3\u003c/em\u003e, \u003cem\u003eCD276\u003c/em\u003e, \u003cem\u003ePOLE\u003c/em\u003e, \u003cem\u003eCARM1\u003c/em\u003e, \u003cem\u003ePLK2\u003c/em\u003e, \u003cem\u003eELF3\u003c/em\u003e, and \u003cem\u003eCCND1\u003c/em\u003e emerged as the top eight genes, each showing significant prognostic value in the largest number of cancer types (Fig. S7C). These findings are consistent with prior reports linking these factors to tumor progression [\u003cspan additionalcitationids=\"CR86\" citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e87\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFinally, to further validate the association between S-ZGA genes and cancer aggressiveness, we performed an overlap analysis integrating S-ZGA genes with established cancer driver genes and differentially expressed genes (DEGs) reported in GEPIA2. Based on this intersection, we identified \u003cem\u003eCCND1\u003c/em\u003e, \u003cem\u003eELF3\u003c/em\u003e, and \u003cem\u003eRGL3\u003c/em\u003e as the candidates exhibiting the most significant expression alterations for subsequent experimental validation. We respectively transfected siRNAs targeting S-ZGA genes \u003cem\u003eCCND1, ELF3\u003c/em\u003e, \u003cem\u003eRGL3\u003c/em\u003e, and negative control siRNA in human ovarian cancer cell line SKOV3, with knockdown efficiency verified by RT-qPCR. Then, the proliferation, migration and invasion of SKOV3 cells were detected by CCK8 and transwell assays. Two independent siRNAs were designed for each target gene. Both siRNAs effectively knocked down their respective targets (Fig.S7D). We found the proliferation was significantly inhibited after knockdown of \u003cem\u003eELF3\u003c/em\u003e, \u003cem\u003eCCND1\u003c/em\u003e and \u003cem\u003eRGL3\u003c/em\u003e (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eC), and cell migration and invasion were also obviously inhibited (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eD-E and Fig.S7E). These results establish that S-ZGA genes play a functional role in regulating cancer cell proliferative, migratory, and invasive capacities \u003cem\u003ein vitro\u003c/em\u003e. In conclusion, S-ZGA genes were activated during ZGA and maintained throughout development. They orchestrate essential developmental processes while, once their epigenetic signals are \u0026ldquo;hijacked\u0026rdquo; to drive aberrant expression, they may promote tumorigenesis and metastasis.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eZGA marks the transition from maternal to zygotic transcriptional control in metazoan embryos, initiating the major wave of embryonic gene expression [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. In this study, we introduced a computational framework that leverages time-series RNA-seq data to capture the dynamic trajectories of gene expression during early embryogenesis. Unlike prior methods reliant on static pre/post-ZGA comparisons [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Our framework enables the prior-free ZGA timing identification and robust cross-dataset gene classification. This framework not only recapitulates known ZGA timing in well-characterized organisms but also provides a foundation for future discovery of ZGA dynamics in unexplored species. In addition, the ZGA gene sets we identified represent a valuable resource for developmental and cancer biology research, as the transcriptional mechanisms governing embryonic ZGA are conserved in cellular reprogramming contexts such as induced pluripotent stem cell (iPSC) generation [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. This framework enables the systematic discovery of novel genes critical for developmental processes and cancer-driving functions, thereby advancing our understanding of the molecular basis of embryonic genome activation.\u003c/p\u003e \u003cp\u003eFurthermore, one subset stood out for its distinctive expression dynamics: class 4 genes, which we term Sustained ZGA (S-ZGA) genes, are activated at ZGA and maintain persistent expression throughout subsequent development (Fig. S4A). This sustained pattern prompted us to investigate whether these genes possess unique regulatory and functional properties that might underlie their potential to be re-deployed in disease contexts. By integrating epigenomic profiling with expression-matched control analyses-a crucial step to address potential confounding by high transcriptional activity-we demonstrated that S-ZGA genes possess a unique regulatory landscape. Notably, their promoters exhibit significantly higher chromatin accessibility (ATAC-seq) even after controlling for expression level, whereas H3K4me3 enrichment is largely explained by transcriptional output. This suggests that open chromatin is an intrinsic feature of S-ZGA genes, poising them for sustained activity, while H3K4me3 reflects ongoing transcription. These findings refine our understanding of how epigenetic mechanisms establish and maintain distinct gene expression programs during development. In addition, there are many similarities between early embryonic development and tumorigenesis in cell invasion behavior, epigenetic regulation, gene expression, metabolism and other important biological behaviors [\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e, \u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e89\u003c/span\u003e]. Epithelial-mesenchymal transition (EMT) pathways-a program essential for embryogenesis and aberrantly reactivated in metastasis [\u003cspan additionalcitationids=\"CR91\" citationid=\"CR90\" class=\"CitationRef\"\u003e90\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR92\" class=\"CitationRef\"\u003e92\u003c/span\u003e]. Using \u003cem\u003ein vitro\u003c/em\u003e assays in an ovarian cancer cell line, we further showed that knockdown of selected S-ZGA genes (\u003cem\u003eCCND1\u003c/em\u003e, \u003cem\u003eELF3\u003c/em\u003e, \u003cem\u003eRGL3\u003c/em\u003e) impairs proliferation, migration, and invasion. our work provides a foundation for investigating how early embryonic drivers are re-deployed in disease, converting them from uncharacterized transcriptional events into mechanistically tractable hypotheses.\u003c/p\u003e \u003cp\u003eSeveral limitations of this study should be acknowledged. First, our current framework is specifically designed to identify a discrete developmental transition-the onset of ZGA. For continuous processes such as cellular differentiation or lineage commitment, the sample-clustering approach used here may not be directly applicable. However, the underlying principle of analyzing full temporal trajectories remains valid. By replacing hierarchical clustering with pseudotime ordering or continuous trend-clustering algorithms, the core concept of our framework could be adapted to systematically identify genes with specific dynamic patterns in gradual biological processes; this represents a promising direction for future work. Second, while our framework is designed to be species-flexible, its application to species with completely uncharacterized ZGA timing remains to be systematically validated. Finally, the cancer-related findings presented here are primarily correlative. Although we demonstrate a strong statistical association between S-ZGA genes and cancer driver genes, and provide in vitro evidence for three significantly differentially expressed S-ZGA genes (\u003cem\u003eCCND1\u003c/em\u003e, \u003cem\u003eELF3\u003c/em\u003e, \u003cem\u003eRGL3\u003c/em\u003e) in the SKOV3 cell line, these results do not establish a direct causal role for S-ZGA genes in driving tumorigenesis or metastasis \u003cem\u003ein vivo\u003c/em\u003e. The functional validation definitive evidence would require \u003cem\u003ein vivo\u003c/em\u003e models and broader interrogation across diverse cancer types.\u003c/p\u003e \u003cp\u003eIn summary, this study establishes a dynamic framework for dissecting gene expression programs during ZGA, identifies a previously underappreciated class of sustained ZGA genes, and uncovers their potential dual relevance to development and disease. By providing a robust resource and highlighting the interplay between epigenetic regulation and oncogenic potential, our work opens new avenues for investigating the molecular mechanisms that bridge early embryogenesis and cancer biology.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003escRNA-seq data processing\u003c/h2\u003e \u003cp\u003eSequencing reads were trimmed using Trim Galore (version 0.6.5), which is a Perl wrapper based on two tools: Cutadapt and FastQC. Reads with quality\u0026thinsp;\u0026gt;\u0026thinsp;20 were aligned to the mouse reference genome (mm10) using STAR (version 2.7.3a) [\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e]. Low mapping quality sequencing reads were removed using Samtools version 1.7. The gene expression level was calculated by Cufflinks (version 2.2.1) [\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e]. The relative abundance of transcripts was measured as fragments per kilobase of exon per million mapped reads (FPKM). Testing of the differential expression between different groups was performed with the Cuffdiff module, and genes with a q-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were selected as being differentially expressed.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eData Integration and Normalization\u003c/h2\u003e \u003cp\u003eTo integrate RNA-seq data from early mouse embryonic development with publicly available ENCODE mature tissue data, we implemented a rigorous two-step normalization pipeline. First, to address the inherent skewness and heteroscedasticity of RNA-seq count-derived data, FPKM values were subjected to a log2-transformation [log2(FPKM\u0026thinsp;+\u0026thinsp;1)]. This step stabilizes variance and approximates a normal distribution, satisfying the assumptions required for subsequent parametric scaling. Second, we applied a row-wise Z-score transformation to the log-transformed matrix. This normalization aligns the distribution of expression values across datasets by ensuring similar mean and variance per gene, thereby minimizing technical biases arising from different sequencing platforms. This approach enables the analysis to focus on the similarity of relative gene expression patterns rather than absolute abundance.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eHierarchical clustering\u003c/h2\u003e \u003cp\u003eFollowing normalization, hierarchical clustering was performed based on the Weighted average distance (WPGMA) algorithm (this approach employs a recursive formula to update inter-cluster distances upon merging D (Cκ, Ci \u0026cup;Cj) and defines the distance between clusters as the weighted average of all pairwise sample distances), which uses a recursive definition for the distance between two clusters and eliminates the influence of class cluster size on clustering. The Spearman distance was used to measure the distance of samples, which is robust to non-linear relationships and outliers [\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003ePrincipal component analysis\u003c/h2\u003e \u003cp\u003eWe performed a principal component analysis (PCA) of the gene expression-level matrix to further characterize transcriptome distances. The calculation was performed with a singular-value decomposition of the data matrix normalized by Z-Score transformation, which made sure that the mean expression of each gene across cells is 0, and the variance across cells is 1. Then, we used the top two principal component features to describe the clustering relationships among embryonic samples in two-dimensional space.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eKey Gene screening for each developmental stage\u003c/h2\u003e \u003cp\u003eTo capture the dynamic transcriptional landscape during early embryogenesis, we employed a three-layer filtering strategy designed to classify genes based on their binary expression states (\u0026ldquo;Active\u0026rdquo; vs. \u0026ldquo;Silent\u0026rdquo;) across the three sequential developmental clusters identified in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC. We screened genes based on their expression. First, we defined the features representing the expression pattern. Briefly, in early embryonic development, the features were genes whose FPKM\u0026thinsp;\u0026ge;\u0026thinsp;5 at \u0026gt;\u0026thinsp;25% of the samples or FPKM\u0026thinsp;\u0026lt;\u0026thinsp;1 in all samples of the cluster. In tissues, the features were genes whose FPKM\u0026thinsp;\u0026ge;\u0026thinsp;5 at \u0026gt;\u0026thinsp;25% of the samples or FPKM\u0026thinsp;\u0026gt;\u0026thinsp;1 at \u0026le;\u0026thinsp;25% sample of the cluster. Next, the genes defined as features were transferred to the next filter, while other genes would be discarded. After gene screening by 3-layer filters, 8 gene lists were obtained (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC, Fig. \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eC). To deal with maternal genes with high expression levels before ZGA, we additionally screened maternal genes (in class 2) whose FPKM\u0026thinsp;\u0026gt;\u0026thinsp;100 at MII oocyte, zygote and early 2-cell stages and FPKM increased fivefold at early 2-cell stage compared with 4-cell stage (FPKM\u0026thinsp;\u0026gt;\u0026thinsp;1 at \u0026le;\u0026thinsp;50% sample for 2-cell to ESC-stages).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eAssessment of threshold robustness\u003c/h2\u003e \u003cp\u003eTo evaluate the robustness of our gene classification strategy, we performed a sensitivity analysis by varying the sample proportion cutoff used to define consistent expression within each developmental cluster. Specifically, we tested cutoffs ranging from 20% to 30% (default: 25%) Fig. S8A-8D. The core classification of main ZGA genes remained largely stable, particularly for S-ZGA genes, where over 90% of genes were consistently classified across different parameter settings, indicating that our main conclusions are robust to the exact choice of this parameter.\u003c/p\u003e \u003cp\u003eWe further evaluated the robustness of class 1 classification by adjusting the post-ZGA silence threshold (default: FPKM\u0026thinsp;\u0026gt;\u0026thinsp;1 at \u0026lt;\u0026thinsp;50% of samples from 2-cell to ESC). Varying this threshold did not substantially alter the set of identified genes or their enriched biological processes Fig. S8E.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003ePerforming a cross-species ortholog retention analysis\u003c/h2\u003e \u003cp\u003eTo investigate whether the functional and developmental importance of S-ZGA genes is reflected in their evolutionary conservation, we performed an ortholog analysis between mouse and human. We identified one-to-one orthologs using the Ensembl Compara database. We then calculated the proportion of human S-ZGA genes (Class 4 in human) that retain a clear ortholog in mouse, and compared this to the ortholog retention rate of non-S-ZGA human genes (all other classes).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eValidation of Gene Classes Using Public Perturbation Data\u003c/h2\u003e \u003cp\u003eTo validate the functional characteristics of the six identified gene classes, we examined their expression responses to transcriptional inhibition using publicly available RNA-seq datasets from mouse 2-cell embryos treated with α-Amanitin. It is important to note that this analysis was performed as a descriptive comparison of expression patterns rather than a statistical differential expression test. Where biological replicates were available in the source dataset, we calculated the mean FPKM values for each condition to represent the expression profile. We then compared the relative expression levels of genes in each class between the control and treated groups to identify consistent up- or down-regulation trends.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eGene set enrichment analysis\u003c/h2\u003e \u003cp\u003eGene Ontology (GO) term enrichment analysis in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD was performed for genes in each class using the Database for Annotation, Visualization and Integrated Discovery Bioinformatics Resource (DAVID, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://david.ncifcrf.gov/home.jsp\u003c/span\u003e\u003cspan address=\"https://david.ncifcrf.gov/home.jsp\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) [\u003cspan citationid=\"CR96\" class=\"CitationRef\"\u003e96\u003c/span\u003e]. GO analysis in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB and Fig.S6A was performed by clusterProfiler-enrichGO (V4.0.5) [\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e]. Enriched TFs of S-ZGA genes in Fig.S6B were calculated by metascape [\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e] and recorded in TRRUST Database[\u003cspan citationid=\"CR99\" class=\"CitationRef\"\u003e99\u003c/span\u003e]. The Benjamini-Hochberg method was used for correction to control the false discovery rate (FDR). For the Gene Ontology (GO) enrichment analysis, the background gene set was defined as all annotated genes in the genome.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eIdentification of topologically associated domain boundaries\u003c/h2\u003e \u003cp\u003eWe collected Hi-C data for the late 2-cell stage from GEO under accession no. GSE82185 and identified significant interactions using HOMER algorithm with a default parameter (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.001) and 5-kb resolution. A topologically associated domain (TAD) boundaries were identified at 40-kb bin resolution using the previously described insulation score analysis approach [\u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e100\u003c/span\u003e]. Briefly, TAD boundaries were identified by using the insulation score method at 40-kb resolution, with two minor modifications. First, a 200-kb (rather than a 500-kb) genomic region was used to detect a TAD, as a TAD is now considered to be mostly\u0026thinsp;\u0026lt;\u0026thinsp;200-kb. Second, in delta vector calculations, we used a 120-kb (rather than a 100-kb) window to fit the resolution, which is suitable in detecting TADs with high sensitivity with a low false-positive rate. We collected Hi-C data from GEO under accession.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eAnalyses of SNP within promoters of genes\u003c/h2\u003e \u003cp\u003eWe collected vcf file from NCBI Allele Frequency Aggregator (ALFA, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ewww.ncbi.nlm.nih.gov/snp/docs/gsr/alfa/\u003c/a\u003e\u003c/span\u003e\u003cspan address=\"http://www.ncbi.nlm.nih.gov/snp/docs/gsr/alfa/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, Release Version: 20200227123210) [\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e]. We used bcftools to calculate the SNP overlap with each target region (\u0026plusmn;\u0026thinsp;1 kb around the transcription start site (TSS)). We then filter the SNP that minor allele frequency (MAF) is 0 (site homozygous for the reference allele and no variant allele detected from the current subject sample size). For the SNP density analysis, the background gene set was defined as all detected protein-coding genes within our dataset, excluding the identified S-ZGA genes. This restriction to protein-coding genes ensured comparability in gene structure and annotation. Notably, Cluster 0 and Cluster 7 were included in the background set; all non-S-ZGA protein-coding genes within these clusters served as controls, as they represent the same biological context and cell states relevant to the study.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eEpigenetic profiles S-ZGA genes\u003c/h2\u003e \u003cp\u003eWe collected bedgraph or bigwig files from GEO as Datasets parts mentioned. We used the EnrichedHeatmap R package to quantify ATAC-seq signals. Specifically, the normalizeToMatrix function was employed to map genomic signals to target regions (\u0026plusmn;\u0026thinsp;5 kb around the TSS). Signal values were averaged across biological replicates where available. Finally, for each gene, we calculated the mean signal intensity across its entire target region to represent the overall chromatin accessibility.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eExpression-matched control selection\u003c/h2\u003e \u003cp\u003eSpecifically, to control for potential confounding due to differences in gene expression levels, we performed a matched control gene selection for each gene class (classes 1\u0026ndash;6). Using expression data from the early 2-cell to ICM stages, genes belonging to a given class were designated as the target set, while all genes from the other classes served as the candidate pool. For each target gene, we searched the pool without replacement and selected the gene whose expression value was closest to that of the target, provided the absolute difference did not exceed 50% of the target\u0026rsquo;s expression value. If no gene in the pool met this criterion, the target received no matched control (recorded as missing). This procedure yielded a list of target-control pairs with comparable expression levels, thereby minimizing expression-level bias in downstream epigenetic analyses. We then compared the ATAC-seq and H3K4me3 signals at promoter regions (\u0026plusmn;\u0026thinsp;1 kb around TSS) between classes 1\u0026ndash;6 genes and this expression-matched control set.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eAnalyses of differential gene expression and cancer survival\u003c/h2\u003e \u003cp\u003eWe analyzed differential gene expression between tumor and normal tissues using the Gene Expression Profiling Interactive Analysis (GEPIA2) comprehensive online platform (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://gepia2.cancer-pku.cn/#index\u003c/span\u003e\u003cspan address=\"http://gepia2.cancer-pku.cn/#index\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), which provides rapid and customized delivery of functionalities based on the Cancer Genome Atlas and genotype-tissue expression data [\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e]. Using GEPIA2, the survival effects of differentially expressed genes in our cancer samples were assessed to determine the prognostic value of the expression of identified key ZGA genes.\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003eCell culture\u003c/h2\u003e \u003cp\u003eSKOV3 cell line were purchased from the Cell Bank of Peking Union Medical College.SKOV3 cells were cultured in McCoy\u0026rsquo;s 5A medium (Gibco,16600-082, USA) supplemented with 10% fetal bovine serum (Gibco,10270-106, USA) at 37\u0026deg;C in a humidified 5% CO\u003csup\u003e2\u003c/sup\u003e atmosphere.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003eGene silencing by siRNA transfection\u003c/h2\u003e \u003cp\u003e \u003cem\u003eELF3\u003c/em\u003e, \u003cem\u003eCCND1\u003c/em\u003e and \u003cem\u003eRGL3\u003c/em\u003e genes were knocked down with their specific siRNAs designed and synthesized by GenePharma Co., Ltd. (Suzhou, China). For siRNA transfection, SKOV3 seeded in 6-well plates at 60\u0026thinsp;~\u0026thinsp;70% were transfected with specific pooled siRNAs and scrambled negative control at a final concentration of 100 nM mixed with jetPRIME transfection reagent (Polyplus,114\u0026thinsp;\u0026minus;\u0026thinsp;15, France) according to the manufacturer\u0026rsquo;s instructions. After that, the cells were harvested for RT-qPCR after 48 h to detect knockdown efficiency. To rule out off-target effects, knockdown experiments were performed using two distinct siRNA sequences targeting genes. For each siRNA, three independent biological replicates were conducted.\u003c/p\u003e \u003cp\u003eThe sequences of specific siRNAs were as follows:\u003c/p\u003e \u003cp\u003esiNC: 5ʹ-UUCUCCGAACGUGUCACGUTT-3ʹ\u003c/p\u003e \u003cp\u003esiELF3 #1: 5ʹ-GCUACCAAGUGGAGAAGAATT-3ʹ\u003c/p\u003e \u003cp\u003esiELF3 #2: 5ʹ-GCCAUGAGGUACUACUACATT-3ʹ\u003c/p\u003e \u003cp\u003esiCCND1 #1: 5ʹ-GUCUGCGAGGAACAGAAGUTT-3ʹ\u003c/p\u003e \u003cp\u003esiCCND1 #2: 5ʹ-GCAUGUUCGUGGCCUCUAATT-3ʹ\u003c/p\u003e \u003cp\u003esiRGL3 #1: 5ʹ-GGUGGGAGAGUUGGUGUUUTT-3ʹ\u003c/p\u003e \u003cp\u003esiRGL3 #2: 5ʹ-CCUGUGACUAUCAGCUCUUTT-3ʹ\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec27\" class=\"Section3\"\u003e \u003ch2\u003eCCK8 assay\u003c/h2\u003e \u003cp\u003eCell viability was detected by cell counting kit-8 (Dojindo, CK04, Japan). SKOV3 cells were seeded in 96-well plates at a density of 5\u0026times; 10\u003csup\u003e3\u003c/sup\u003e cells per well, 24 h after transfected siRNAs targeting \u003cem\u003eELF3\u003c/em\u003e, \u003cem\u003eCCND1\u003c/em\u003e, \u003cem\u003eRGL3\u003c/em\u003e or negative control with jetPRIME transfection reagent (Polyplus,114\u0026thinsp;\u0026minus;\u0026thinsp;15, France). Cultured to the indicated times, the medium of each well was replaced with 100 \u0026micro;L fresh culture media containing 10% CCK-8 solution and incubated for 1.5 h at 37\u0026deg;C. The absorbance of each well was measured at a wavelength of 450 nm. All experiments were independently performed in triplicate.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec28\" class=\"Section2\"\u003e \u003ch2\u003eRNA extraction and RT-qPCR\u003c/h2\u003e \u003cp\u003eForty-eight hours after siRNA transfection, total RNA was extracted from SKOV3 cells using Trizol reagent (Invitrogen,15596026, USA) following the manufacturer's instruction. Synthesis of cDNA was performed using HiScript III All-in-one RT SuperMix (Vazyme, R333, China). qPCR was performed using PowerUpTM SYBRTM Green Master Mix (Applied Biosystems, A25742, USA) and analysed on Mx3000P QPCR Systems (Agilent, Mx3000P, USA). The relative mRNA expression was normalized to GAPDH. Each sample was determined with triplicate independent experiments. The primers used for amplification were as follows:\u003c/p\u003e \u003cp\u003eELF3-Forward: 5ʹ-CAGACTCCGGTGGAAGTGAC-3ʹ\u003c/p\u003e \u003cp\u003eELF3-Reverse: 5ʹ-CCCTCGAGACAGTCCCAGTA-3ʹ\u003c/p\u003e \u003cp\u003eCCND1-Forward: 5ʹ-GGAGAACAAACAGATCATCCGC-3ʹ\u003c/p\u003e \u003cp\u003eCCND1-Reverse: 5ʹ-GTGTGAGGCGGTAGTAGGAC-3ʹ\u003c/p\u003e \u003cp\u003eRGL3-Forward: 5ʹ-CTGTGTGCTGGGTTCCGT-3ʹ\u003c/p\u003e \u003cp\u003eRGL3-Reverse: 5ʹ-TGAGCCGGTAGATGGGGTTA-3ʹ\u003c/p\u003e \u003cp\u003eGAPDH-Forward: 5ʹ-GTCTCCTCTGACTTCAACAGCG-3ʹ\u003c/p\u003e \u003cp\u003eGAPDH-Reverse: 5ʹ-ACCACCCTGTTGCTGTAGCCAA-3ʹ\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec29\" class=\"Section2\"\u003e \u003ch2\u003eCell migration and invasion assays\u003c/h2\u003e \u003cp\u003eMigration and invasion assays were performed by 6.5 mm Transwell\u0026reg; with 8.0 \u0026micro;m Pore Polycarbonate Membrane Insert (Corning, 3422, USA) and BioCoatTM Matrigel Invasion Chamber (Corning, 354480, USA), respectively. The Transwell insert and Matrigel Chamber were respectively seeded with SKOV3 cells in 1.5\u0026times;10\u003csup\u003e4\u003c/sup\u003e cells/insert or 8\u0026times;104 cells/chamber in serum-free McCoy\u0026rsquo;s 5A medium, 24 h after transfected siRNAs targeting \u003cem\u003eELF3\u003c/em\u003e, \u003cem\u003eCCND1\u003c/em\u003e, \u003cem\u003eRGL3\u003c/em\u003e or negative control. The 24-well plates were added with culture medium containing 10% or 20% fetal bovine serum as chemoattractant, and then the cells were incubated for 24 h or 48 h in migration and invasion assay, respectively. After incubation, the non-migrated or non-invaded cells were removed from the upper surface of the membrane by a cotton swab. The cells on the lower surface of the membrane were fixed and stained with Diff-Quik Stain Kit (Solarbio, G1540, China), then photographed and counted the nine random fields from triplicate membranes of each group by inverted microscope (100\u0026times;, Nikon, Japan). All experiments were independently performed in triplicate.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec30\" class=\"Section2\"\u003e \u003ch2\u003eStatistical Analysis\u003c/h2\u003e \u003cp\u003eStatistical analysis was performed using GraphPad Prism (GraphPad Software) and R software. Survival analysis based on GEPIA2 uses log-rank test for the hypothesis evaluation. In the verification part of the experiment, paired two-tailed Student\u0026rsquo;s t tests were employed to conduct statistical analysis, and nominal p values are specified. All values represent the means of at least three independent experiments\u0026thinsp;\u0026plusmn;\u0026thinsp;SD. p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 was considered statistically significant. Specifically, the Wilcoxon signed-rank test was applied for data in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA/B and Fig. S4B, and the Wilcoxon rank-sum test (also called Mann-Whitney U test) was applied for data in Fig. S6C/D; Fisher\u0026rsquo;s exact test was used for Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA and Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA; and an unpaired t-test was employed for Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eC/D. Unless otherwise specified, all experiments were performed in at least three independent replicates, and data are presented as mean\u0026thinsp;\u0026plusmn;\u0026thinsp;standard deviation (SD). A p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05 was considered statistically significant.\u003c/p\u003e \u003cp\u003eHypothesis Testing: The null hypothesis (\u003cem\u003eH0\u003c/em\u003e) for all tests posited no difference between groups (mean/median difference\u0026thinsp;=\u0026thinsp;0). A two-tailed alternative hypothesis (\u003cem\u003eH1\u003c/em\u003e) was assumed unless otherwise specified.\u003c/p\u003e \u003cp\u003eMultiple Testing Correction: For high-throughput screening data involving thousands of genes, p-values were adjusted for multiple testing using the Benjamini-Hochberg procedure to control the False Discovery Rate (FDR). For comparisons like the Wilcoxon signed-rank test or unpaired two-tailed Student\u0026rsquo;s t tests, which were not part of a large-scale multiple testing procedure, we have reported the raw p values as is standard practice.\u003c/p\u003e \u003c/div\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eZGA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eZygotic genome activation\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eS-ZGA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSustained ZGA\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eESC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eEmbryonic stem cell\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eICM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInner cell mass\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eFGO\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eFully grown oocytes\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTAD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTopologically associating domain\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSNP\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSingle nucleotide polymorphism\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eTF\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eTranscription factor\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eiPSC\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eInduced pluripotent stem cell\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eEthics approval and consent to participate\u003c/h2\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003ch2\u003eConsent for publication\u003c/h2\u003e\n\u003cp\u003e\u0026nbsp;Not applicable.\u003c/p\u003e\n\u003ch2\u003eData Availability\u003c/h2\u003e\n\u003cp\u003eMouse RNA-seq data covering developmental stages from MII oocytes and PN5 zygotes to early 2-cell, 2-cell, 4-cell, and 8-cell embryos, as well as ICM and ESCs, were obtained from the GEO database (Accession: GSE66390) [26].\u003c/p\u003e\n\u003cp\u003eMouse scRNA-seq data spanning stages from MII oocytes and PN5 zygotes to early 2-cell, 2-cell, 4-cell, and 8-cell embryos, as well as ICM and ESCs, were retrieved from the GEO database (Accession: GSE45719) [11].\u003c/p\u003e\n\u003cp\u003eMouse RNA-seq data of 2-cell embryos, including both untreated controls and those treated with \u0026alpha;-Amanitin, were sourced from the GEO database (Accession: GSE71434) [44].\u003c/p\u003e\n\u003cp\u003eMouse scRNA-seq datasets comparing untreated 2-cell embryos with \u0026alpha;-Amanitin-treated counterparts were acquired from the GEO database (Accession: GSE53386) [47].\u003c/p\u003e\n\u003cp\u003eMouse RNA-seq data for grown oocytes and fully grown oocytes (FGO) were compiled from the GEO database (Accession: GSE112835) [52].\u003c/p\u003e\n\u003cp\u003eMouse ATAC-seq data covering early 2-cell, 2-cell, 4-cell, and 8-cell embryos as well as ICM were obtained from the GEO database (Accession: GSE66390) [26].\u003c/p\u003e\n\u003cp\u003eMouse ChIP-seq data for H3K4me3 across sperm, MII oocytes, zygotes, early 2-cell, 2-cell, 4-cell, and 8-cell embryos, and the ICM were retrieved from the GEO database (Accession: GSE71434) [44].\u003c/p\u003e\n\u003cp\u003eMouse ChIP-seq data for H3K27me3 across MII oocytes, zygotes, 2-cell embryos, 8-cell embryos, and the ICM were sourced from the GEO database (Accession: GSE76687) [102].\u003c/p\u003e\n\u003cp\u003eMouse ChIP-seq data for H3K9me3 across sperm, MII oocytes, zygotes, 2-cell, 4-cell, and 8-cell embryos, and the ICM were acquired from the GEO database (Accession: GSE98149) [103].\u003c/p\u003e\n\u003cp\u003eMouse ChIP-seq data for H3K27ac in oocytes, 2-cell embryos, and 8-cell embryos were compiled from the GEO database (Accession: GSE72784) [104].\u003c/p\u003e\n\u003cp\u003eMouse DNA methylome data covering sperm, MII oocytes, 2-cell and 4-cell embryos, and the ICM stage were obtained from the GEO database (Accession: GSE56697) [105].\u003c/p\u003e\n\u003cp\u003eMouse Hi-C data for late 2-cell embryos were retrieved from the GEO database (Accession: GSE82185) [67].\u003c/p\u003e\n\u003cp\u003eMouse RNA-seq data for multiple tissues, including heart, kidney, liver, stomach, duodenum, lung, spleen, and thymus, were sourced from the ENCODE project [106].\u003c/p\u003e\n\u003cp\u003eMouse ATAC-seq data for heart and liver tissues were acquired from the ENCODE project [106].\u003c/p\u003e\n\u003cp\u003eMouse ChIP-seq data for H3K4me3 in heart and liver tissues were compiled from the ENCODE project [106].\u003c/p\u003e\n\u003cp\u003eHuman RNA-seq data spanning GV oocytes, MII oocytes, 2-cell, 4-cell, and 8-cell embryos, as well as the ICM and ESCs, were obtained from the GEO database (Accession: GSE101571) [34].\u003c/p\u003e\n\u003cp\u003eHuman ATAC-seq data for the 4-cell and 8-cell stages were retrieved from the GEO database (Accession: GSE101571) [34].\u003c/p\u003e\n\u003cp\u003eHuman RNA-seq data comprising 24 samples were sourced from the ENCODE project [106].\u003c/p\u003e\n\u003cp\u003eBovine RNA-seq data covering MII oocytes, 2-cell, 4-cell, and 8-cell embryos, as well as morula and blastocyst stages, were acquired from the GEO database (Accession: GSE59186) [107].\u003c/p\u003e\n\u003cp\u003eMacaque RNA-seq data spanning GV, MI, and MII oocytes, PN zygotes, 2-cell, 4-cell, and 8-cell embryos, through morula and blastocyst stages, were compiled from the GEO database (Accession: GSE103313) [35].\u003c/p\u003e\n\u003cp\u003eMouse promoter coordinates were mapped to the mm10 genome assembly. All gene expression analyses utilized pre-computed matrices downloaded from public repositories; no raw data re-processing or re-alignment was performed.\u003c/p\u003e\n\u003cp\u003eHousekeeping genes in mouse were compiled from previous studies [108]\u003c/p\u003e\n\u003cp\u003eCancer driver genes [76-82] were compiled from the Candidate Cancer Gene Database (CCGD; http://ccgd-starrlab.oit.umn.edu/), incorporating core sets identified by Lawrence et al. (260 genes), Vogelstein et al. (125 genes), Schroeder et al. (285 genes), Zehir et al. (545 genes), Sondka et al. (723 genes), and Mart\u0026iacute;nez-Jim\u0026eacute;nez et al. (568 genes).\u003c/p\u003e\n\u003cp\u003eHuman phenotype-related genes curated in the Human Phenotype Ontology (HPO) Database (https://hpo.jax.org) [73]\u003c/p\u003e\n\u003cp\u003eHuman senescence related genes sourced from the Cell Senescence Gene Database (http://csgene.bioinfo-minzhao.org/) [74]\u003c/p\u003e\n\u003cp\u003eDifferential expression profiles between tumor and normal tissues derived from GEPIA (http://gepia2.cancer-pku.cn/#about) [101]\u003c/p\u003e\n\u003cp\u003eGene expression levels in metastasis samples obtained from the Human Cancer Metastasis Database (HCMDB) (http://hcmdb.i-sanger.com/index) [84]\u003c/p\u003e\n\u003ch2\u003eCompeting interests\u003c/h2\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThe authors declare that this research received no specific grant from any funding.\u003c/p\u003e\n\u003ch2\u003eAuthor contributions\u003c/h2\u003e\n\u003cp\u003eY.Z. contributed to the conception and design of the study. H.X. performed the experiments and wrote the manuscript.\u003c/p\u003e\n\u003ch2\u003eAcknowledgements\u003c/h2\u003e\n\u003cp\u003eThe authors would like to thank Prof. Hebing Chen for his helpful in study design.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003ePijuan-Sala B, Guibentif C, G\u0026ouml;ttgens B. Single-cell transcriptional profiling: a window into embryonic cell-type specification. Nat Rev Mol Cell Biol. 2018;19(6):399\u0026ndash;412.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZernicka-Goetz M, Morris SA, Bruce AW. Making a firm decision: multifaceted regulation of cell fate in the early mouse embryo. Nat Rev Genet. 2009;10(7):467\u0026ndash;77.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSha QQ, Zhu YZ, Li S, Jiang Y, Chen L, Sun XH, Shen L, Ou XH, Fan HY. Characterization of zygotic genome activation-dependent maternal mRNA clearance in mouse. Nucleic Acids Res. 2020;48(2):879\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchulz KN, Harrison MM. Mechanisms regulating zygotic genome activation. Nat Rev Genet. 2019;20(4):221\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen H, Einstein LC, Little SC, Good MC. Spatiotemporal Patterning of Zygotic Genome Activation in a Model Vertebrate Embryo. Dev Cell. 2019;49(6):852\u0026ndash;e866857.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Iaco A, Planet E, Coluccio A, Verp S, Duc J, Trono D. DUX-family transcription factors regulate zygotic genome activation in placental mammals. Nat Genet. 2017;49(6):941\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTadros W, Lipshitz HD. The maternal-to-zygotic transition: a play in two acts. Development. 2009;136(18):3033\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeyn P, Kircher M, Dahl A, Kelso J, Tomancak P, Kalinka AT, Neugebauer KM. The earliest transcribed zygotic genes are short, newly evolved, and different across species. Cell Rep. 2014;6(2):285\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu B, Xu Q, Wang Q, Feng S, Lai F, Wang P, Zheng F, Xiang Y, Wu J, Nie J et al. The landscape of RNA Pol II binding reveals a stepwise transition during ZGA. Nature 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbe K, Yamamoto R, Franke V, Cao M, Suzuki Y, Suzuki MG, Vlahovicek K, Svoboda P, Schultz RM, Aoki F. The first murine zygotic transcription is promiscuous and uncoupled from splicing and 3' processing. Embo j. 2015;34(11):1523\u0026ndash;37.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeng Q, Ramsk\u0026ouml;ld D, Reinius B, Sandberg R. Single-cell RNA-seq reveals dynamic, random monoallelic gene expression in mammalian cells. Sci (New York NY). 2014;343(6167):193\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXue Z, Huang K, Cai C, Cai L, Jiang CY, Feng Y, Liu Z, Zeng Q, Cheng L, Sun YE, et al. Genetic programs in human and mouse early embryos revealed by single-cell RNA sequencing. Nature. 2013;500(7464):593\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark SJ, Komata M, Inoue F, Yamada K, Nakai K, Ohsugi M, Shirahige K. Inferring the choreography of parental genomes during fertilization from ultralarge-scale whole-transcriptome analysis. Genes Dev. 2013;27(24):2736\u0026ndash;48.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan L, Yang M, Guo H, Yang L, Wu J, Li R, Liu P, Lian Y, Zheng X, Yan J, et al. Single-cell RNA-Seq profiling of human preimplantation embryos and embryonic stem cells. Nat Struct Mol Biol. 2013;20(9):1131\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVassena R, Bou\u0026eacute; S, Gonz\u0026aacute;lez-Roca E, Aran B, Auer H, Veiga A, Izpisua Belmonte JC. Waves of early transcriptional activation and pluripotency program initiation during human preimplantation development. Development. 2011;138(17):3699\u0026ndash;709.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee MT, Bonneau AR, Giraldez AJ. Zygotic genome activation during the maternal-to-zygotic transition. Annu Rev Cell Dev Biol. 2014;30:581\u0026ndash;613.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee MT, Bonneau AR, Takacs CM, Bazzini AA, DiVito KR, Fleming ES, Giraldez AJ. Nanog, Pou5f1 and SoxB1 activate zygotic gene expression during the maternal-to-zygotic transition. Nature. 2013;503(7476):360\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMitsui K, Tokuzawa Y, Itoh H, Segawa K, Murakami M, Takahashi K, Maruyama M, Maeda M, Yamanaka S. The homeoprotein Nanog is required for maintenance of pluripotency in mouse epiblast and ES cells. Cell. 2003;113(5):631\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMingyue G, Yanping Z, Jianfeng Z, Yan B, Junqin X, Ce X, Xiaochen K, Yanhong Z, Yanhe L, Zhifen T et al. Precise temporal regulation of Dux is important for embryo development. Cell Res 2019, 29(11).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShuyan J, Fengling C, Paula S, Jiacheng W, Ziming Z, Lijuan W, Qing Z, Zili L, Bofeng L, Kai X et al. OBOX regulates mouse zygotic genome activation and early development. \u003cem\u003eNature\u003c/em\u003e 2023, 620(7976).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang W, Chen F, Chen R, Xie D, Yang J, Zhao X, Guo R, Zhang Y, Shen Y, G\u0026ouml;ke J, et al. Zscan4c activates endogenous retrovirus MERVL and cleavage embryo genes. Nucleic Acids Res. 2019;47(16):8485\u0026ndash;501.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInoue K, Oikawa M, Kamimura S, Ogonuki N, Nakamura T, Nakano T, Abe K, Ogura A. Trichostatin A specifically improves the aberrant expression of transcription factor genes in embryos produced by somatic cell nuclear transfer. Sci Rep. 2015;5:10127.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSonehara H, Nagata M, Aoki F. Roles of the first and second round of DNA replication in the regulation of zygotic gene activation in mice. J Reprod Dev. 2008;54(5):381\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Song M, Yang W, Cao P, Zheng L, Zuo Y. A Comparative Analysis of Single-Cell Transcriptome Identifies Reprogramming Driver Factors for Efficiency Improvement. Mol therapy Nucleic acids. 2020;19:1053\u0026ndash;64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbe KI, Funaya S, Tsukioka D, Kawamura M, Suzuki Y, Suzuki MG, Schultz RM, Aoki F. Minor zygotic gene activation is essential for mouse preimplantation development. Proc Natl Acad Sci USA. 2018;115(29):E6780\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu J, Huang B, Chen H, Yin Q, Liu Y, Xiang Y, Zhang B, Liu B, Wang Q, Xia W, et al. The landscape of accessible chromatin in mammalian preimplantation embryos. Nature. 2016;534(7609):652\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJukam D, Shariati SAM, Skotheim JM. Zygotic Genome Activation in Vertebrates. Dev Cell. 2017;42(4):316\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi L, Lu X, Dean J. The maternal to zygotic transition in mammals. Mol Aspects Med. 2013;34(5):919\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollart C, Owens ND, Bhaw-Rosun L, Cooper B, De Domenico E, Patrushev I, Sesay AK, Smith JN, Smith JC, Gilchrist MJ. High-resolution analysis of gene activity during the Xenopus mid-blastula transition. Development. 2014;141(9):1927\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarvey SA, Sealy I, Kettleborough R, Fenyes F, White R, Stemple D, Smith JC. Identification of the zebrafish maternal and paternal transcriptomes. Development. 2013;140(13):2703\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLott SE, Villalta JE, Schroth GP, Luo S, Tonkin LA, Eisen MB. Noncanonical compensation of zygotic X transcription in early Drosophila melanogaster development revealed through single-embryo RNA-seq. PLoS Biol. 2011;9(2):e1000590.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing J, Aronow BJ, Kaminski N, Kitzmiller J, Whitsett JA, Bar-Joseph Z. Reconstructing differentiation networks and their regulation from time series single-cell expression data. Genome Res. 2018;28(3):383\u0026ndash;95.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e\u0026Oslash;strup O, Andersen IS, Collas P. Chromatin-linked determinants of zygotic genome activation. Cell Mol Life Sci. 2013;70(8):1425\u0026ndash;37.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu J, Xu J, Liu B, Yao G, Wang P, Lin Z, Huang B, Wang X, Li T, Shi S, et al. Chromatin analysis in human early development reveals epigenetic transition during ZGA. Nature. 2018;557(7704):256\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChitwood JL, Burruel VR, Halstead MM, Meyers SA, Ross PJ. Transcriptome profiling of individual rhesus macaque oocytes and preimplantation embryos. Biol Reprod. 2017;97(3):353\u0026ndash;64.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGraf A, Krebs S, Zakhartchenko V, Schwalb B, Blum H, Wolf E. Fine mapping of genome activation in bovine embryos by RNA sequencing. Proc Natl Acad Sci USA. 2014;111(11):4139\u0026ndash;44.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eClark MB, Mercer TR, Bussotti G, Leonardi T, Haynes KR, Crawford J, Brunck ME, Cao KA, Thomas GP, Chen WY, et al. Quantitative gene profiling of long noncoding RNAs with targeted RNA sequencing. Nat Methods. 2015;12(4):339\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu S, Lin L, Jiang P, Wang D, Xing Y. A comparison of RNA-Seq and high-density exon array for detecting differential gene expression between closely related species. Nucleic Acids Res. 2011;39(2):578\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJohn RT, Sean BC. Gene co-option in physiological and morphological evolution. Annu Rev Cell Dev Biol 2002, 18(0).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeikun X, Jiawei X, Guang Y, Guidong Y, Kai X, Xueshan M, Nan Z, Bofeng L, Tong L, Zili L et al. Resetting histone modifications during human parental-to-zygotic transition. \u003cem\u003eScience (New York, NY)\u003c/em\u003e 2019, 365(6451).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMina LK, Caroline H, Antonio JG. The maternal-to-zygotic transition: reprogramming of the cytoplasm and nucleus. Nat Rev Genet 2024, 26(4).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQian-Qian S, Yun-Wen W, Meng-Yan J, Yu J, Xing-Xing D, Long-Wen Z, Heng-Yu F. Maternal mRNA clearance is associated with the prevention of precocious transcription and genome instability in mouse early embryos. Sci Bull (Beijing) 2025, 71(5).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQiuyan W, Chuanxin Z, Yanna D, Jiaqi S, Zhuoning Z, Cheng L, Shuiying M, Zongyu L, Hui L, Xiaonan M et al. Maternal factor OTX2 regulates human embryonic genome activation and early development. Nat Genet 2025, 57(11).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang B, Zheng H, Huang B, Li W, Xiang Y, Peng X, Ming J, Wu X, Zhang Y, Xu Q, et al. Allelic reprogramming of the histone modification H3K4me3 in early mammalian development. Nature. 2016;537(7621):553\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK E LR. L: Strain-specific progression of alpha-amanitin-treated mouse embryos beyond the two-cell stage. Mol Reprod Dev 1995, 41(1).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu F, Liu Y, Inoue A, Suzuki T, Zhao K, Zhang Y. Establishing Chromatin Regulatory Landscape during Mouse Preimplantation Development. Cell. 2016;165(6):1375\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFan X, Zhang X, Wu X, Guo H, Hu Y, Tang F, Huang Y. Single-cell RNA-seq transcriptome analysis of linear and circular RNAs in mouse preimplantation embryos. Genome Biol. 2015;16(1):148.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark SJ, Shirahige K, Ohsugi M, Nakai K. DBTMEE: a database of transcriptome in mouse early embryos. Nucleic Acids Res. 2015;43(Database issue):D771\u0026ndash;776.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu X, Kovalev GI, Chang H, Kallin E, Knudsen G, Xia L, Mishra N, Ruiz P, Li E, Su L, et al. Inactivation of NuRD component Mta2 causes abnormal T cell activation and lupus-like autoimmune disease in mice. J Biol Chem. 2008;283(20):13825\u0026ndash;33.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLoda A, Heard E. Xist RNA in action: Past, present, and future. PLoS Genet. 2019;15(9):e1008333.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWells KD. More with less Xist. Proc Natl Acad Sci USA. 2012;109(2):349\u0026ndash;50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu Q, Xiang Y, Wang Q, Wang L, Brind'Amour J, Bogutz AB, Zhang Y, Zhang B, Yu G, Xia W, et al. SETD2 regulates the maternal epigenome, genomic imprinting and embryonic development. Nat Genet. 2019;51(5):844\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEisenberg E, Levanon EY. Human housekeeping genes, revisited. Trends Genet. 2013;29(10):569\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOgawa S, Yamada M, Nakamura A, Sugawara T, Nakamura A, Miyajima S, Harada Y, Ooka R, Okawa R, Miyauchi J, et al. Zscan5b Deficiency Impairs DNA Damage Response and Causes Chromosomal Aberrations during Mitosis. Stem Cell Rep. 2019;12(6):1366\u0026ndash;79.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo F, Li L, Li J, Wu X, Hu B, Zhu P, Wen L, Tang F. Single-cell multi-omics sequencing of mouse early embryos and embryonic stem cells. Cell Res. 2017;27(8):967\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi L, Guo F, Gao Y, Ren Y, Yuan P, Yan L, Li R, Lian Y, Li J, Hu B, et al. Single-cell multi-omics sequencing of human early embryos. Nat Cell Biol. 2018;20(7):847\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu Q, Xie W. Epigenome in Early Mammalian Development: Inheritance, Reprogramming and Establishment. Trends Cell Biol. 2018;28(3):237\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou LQ, Dean J. Reprogramming the genome to totipotency in mouse embryos. Trends Cell Biol. 2015;25(2):82\u0026ndash;91.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHontelez S, van Kruijsbergen I, Georgiou G, van Heeringen SJ, Bogdanovic O, Lister R, Veenstra GJC. Embryonic transcription is controlled by maternally defined chromatin state. Nat Commun. 2015;6:10148.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo H, Zhu P, Yan L, Li R, Hu B, Lian Y, Yan J, Ren X, Lin S, Li J, et al. The DNA methylation landscape of human early embryos. Nature. 2014;511(7511):606\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMesserschmidt DM, Knowles BB, Solter D. DNA methylation dynamics during epigenetic reprogramming in the germline and preimplantation embryos. Genes Dev. 2014;28(8):812\u0026ndash;28.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLindeman LC, Andersen IS, Reiner AH, Li N, Aanes H, \u0026Oslash;strup O, Winata C, Mathavan S, M\u0026uuml;ller F, Alestr\u0026ouml;m P, et al. Prepatterning of developmental gene expression by modified histones before zygotic genome activation. Dev Cell. 2011;21(6):993\u0026ndash;1004.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRuthenburg AJ, Allis CD, Wysocka J. Methylation of lysine 4 on histone H3: intricacy of writing and reading a single epigenetic mark. Mol Cell. 2007;25(1):15\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePing G, Ji-Han X, Csilla S, Xiao-Ming D, Qin Z, Yuehong Y, Peng Z, Sara Pereira C, Kai Z, Jing Z et al. Biology and Clinical Implications of the 19q13 Aggressive Prostate Cancer Susceptibility Locus. Cell 2018, 174(3).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eI K. N S, D S, A P, H F, S L: Oxygen utilization during isovolumic pressure-volume loading: effects of prolonged extracorporeal circulation and cardioplegic arrest. Ann Thorac Surg 1986, 41(4).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBingjie Z, Hui Z, Bo H, Wenzhi L, Yunlong X, Xu P, Jia M, Xiaotong W, Yu Z, Qianhua X et al. Allelic reprogramming of the histone modification H3K4me3 in early mammalian development. \u003cem\u003eNature\u003c/em\u003e 2016, 537(7621).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu Z, Zheng H, Huang B, Ma R, Wu J, Zhang X, He J, Xiang Y, Wang Q, Li Y, et al. Allelic reprogramming of 3D chromatin architecture during early mammalian development. Nature. 2017;547(7662):232\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHug CB, Grimaldi AG, Kruse K, Vaquerizas JM. Chromatin Architecture Emerges during Zygotic Genome Activation Independent of Transcription. Cell. 2017;169(2):216\u0026ndash;e228219.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShaulian E, Karin M. AP-1 as a regulator of cell life and death. Nat Cell Biol. 2002;4(5):E131\u0026ndash;136.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWagner EF. Functions of AP1 (Fos/Jun) in bone development. Ann Rheum Dis. 2002;61(Suppl 2):ii40\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaczynski J, Cook T, Urrutia R. Sp1- and Kr\u0026uuml;ppel-like transcription factors. Genome Biol. 2003;4(2):206.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eALFA. Allele Frequency Aggregator. [\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003c/span\u003e\u003cspan address=\"http://www.ncbi.nlm.nih.gov/snp/docs/gsr/alfa/]\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK\u0026ouml;hler S, Gargano M, Matentzoglu N, Carmody LC, Lewis-Smith D, Vasilevsky NA, Danis D, Balagura G, Baynam G, Brower AM, et al. The Human Phenotype Ontology in 2021. Nucleic Acids Res. 2021;49(D1):D1207\u0026ndash;17.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao M, Chen L, Qu H. CSGene: a literature-based database for cell senescence genes and its application to identify critical cell aging pathways and associated diseases. Cell Death Dis. 2016;7(1):e2053.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu A, Yu X, Liu S. Pluripotency transcription factors and cancer stem cells: small genes make a big difference. Chin J cancer. 2013;32(9):483\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbbott KL, Nyre ET, Abrahante J, Ho YY, Isaksson Vogel R, Starr TK. The Candidate Cancer Gene Database: a database of cancer driver genes from forward genetic screens in mice. Nucleic Acids Res. 2015;43(Database issue):D844\u0026ndash;848.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLawrence MS, Stojanov P, Mermel CH, Robinson JT, Garraway LA, Golub TR, Meyerson M, Gabriel SB, Lander ES, Getz G. Discovery and saturation analysis of cancer genes across 21 tumour types. Nature. 2014;505(7484):495\u0026ndash;501.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVogelstein B, Papadopoulos N, Velculescu VE, Zhou S, Diaz LA Jr., Kinzler KW. Cancer genome landscapes. Sci (New York NY). 2013;339(6127):1546\u0026ndash;58.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchroeder MP, Rubio-Perez C, Tamborero D, Gonzalez-Perez A, Lopez-Bigas N. OncodriveROLE classifies cancer driver genes in loss of function and activating mode of action. Bioinf (Oxford England). 2014;30(17):i549\u0026ndash;555.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZehir A, Benayed R, Shah RH, Syed A, Middha S, Kim HR, Srinivasan P, Gao J, Chakravarty D, Devlin SM, et al. Mutational landscape of metastatic cancer revealed from prospective clinical sequencing of 10,000 patients. Nat Med. 2017;23(6):703\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSondka Z, Bamford S, Cole CG, Ward SA, Dunham I, Forbes SA. The COSMIC Cancer Gene Census: describing genetic dysfunction across all human cancers. Nat Rev Cancer. 2018;18(11):696\u0026ndash;705.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMart\u0026iacute;nez-Jim\u0026eacute;nez F, Mui\u0026ntilde;os F, Sent\u0026iacute;s I, Deu-Pons J, Reyes-Salazar I, Arnedo-Pac C, Mularoni L, Pich O, Bonet J, Kranas H et al. A compendium of mutational cancer driver genes. Nat Rev Cancer 2020.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDongre A, Weinberg RA. New insights into the mechanisms of epithelial-mesenchymal transition and implications for cancer. Nat Rev Mol Cell Biol. 2019;20(2):69\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng G, Ma Y, Zou Y, Yin A, Li W, Dong D. HCMDB: the human cancer metastasis database. Nucleic Acids Res. 2018;46(D1):D950\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePicarda E, Ohaegbulam KC, Zang X. Molecular Pathways: Targeting B7-H3 (CD276) for Human Cancer Immunotherapy. Clin cancer research: official J Am Association Cancer Res. 2016;22(14):3425\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoelens MC, Wu TJ, Nabet BY, Xu B, Qiu Y, Yoon T, Azzam DJ, Twyman-Saint Victor C, Wiemann BZ, Ishwaran H, et al. Exosome transfer from stromal to breast cancer cells regulates therapy resistance pathways. Cell. 2014;159(3):499\u0026ndash;513.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eIwaya T, Fukagawa T, Suzuki Y, Takahashi Y, Sawada G, Ishibashi M, Kurashige J, Sudo T, Tanaka F, Shibata K, et al. Contrasting expression patterns of histone mRNA and microRNA 760 in patients with gastric cancer. Clin cancer research: official J Am Association Cancer Res. 2013;19(23):6438\u0026ndash;49.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMa Y, Zhang P, Wang F, Yang J, Yang Z, Qin H. The relationship between early embryo development and tumourigenesis. J Cell Mol Med. 2010;14(12):2697\u0026ndash;701.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmith DG, Sturmey RG. Parallels between embryo and cancer cell metabolism. Biochem Soc Trans. 2013;41(2):664\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBernardi Y, Strobl-Mazzulla PH. What we can learn from embryos to understand the mesenchymal-to-epithelial transition in tumor progression. Biochem J. 2021;478(9):1809\u0026ndash;25.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWieland E, Rodriguez-Vita J, Liebler SS, Mogler C, Moll I, Herberich SE, Espinet E, Herpel E, Menuchin A, Chang-Claude J, et al. Endothelial Notch1 Activity Facilitates Metastasis. Cancer Cell. 2017;31(3):355\u0026ndash;67.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang S, Chitu V, Stanley ER, Elliott BE, Greer PA. Fes tyrosine kinase expression in the tumor niche correlates with enhanced tumor growth, angiogenesis, circulating tumor cells, metastasis, and infiltrating macrophages. Cancer Res. 2011;71(4):1465\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDobin A, Davis CA, Schlesinger F, Drenkow J, Zaleski C, Jha S, Batut P, Chaisson M, Gingeras TR. STAR: ultrafast universal RNA-seq aligner. Bioinf (Oxford England). 2013;29(1):15\u0026ndash;21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTrapnell C, Roberts A, Goff L, Pertea G, Kim D, Kelley DR, Pimentel H, Salzberg SL, Rinn JL, Pachter L. Differential gene and transcript expression analysis of RNA-seq experiments with TopHat and Cufflinks. Nat Protoc. 2012;7(3):562\u0026ndash;78.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLazar C, Meganck S, Taminau J, Steenhoff D, Coletta A, Molter C, Weiss-Sol\u0026iacute;s DY, Duque R, Bersini H, Now\u0026eacute; A. Batch effect removal methods for microarray gene expression data integration: a survey. Brief Bioinform. 2013;14(4):469\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDennis G Jr., Sherman BT, Hosack DA, Yang J, Gao W, Lane HC, Lempicki RA. DAVID: Database for Annotation, Visualization, and Integrated Discovery. Genome Biol. 2003;4(5):P3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu T, Hu E, Xu S, Chen M, Guo P, Dai Z, Feng T, Zhou L, Tang W, Zhan L, et al. clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innov (New York NY). 2021;2(3):100141.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Y, Zhou B, Pache L, Chang M, Khodabakhshi AH, Tanaseichuk O, Benner C, Chanda SK. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10(1):1523.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHan H, Cho JW, Lee S, Yun A, Kim H, Bae D, Yang S, Kim CY, Lee M, Kim E, et al. TRRUST v2: an expanded reference database of human and mouse transcriptional regulatory interactions. Nucleic Acids Res. 2018;46(D1):D380\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchmitt A, Hu M, Jung I, Xu Z, Qiu Y, Tan C, Li Y, Lin S, Lin Y, Barr C. A Compendium of Chromatin Contact Maps Reveals Spatially Active Regions in the Human Genome. Cell Rep. 2016;17(8):2042\u0026ndash;59.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang Z, Kang B, Li C, Chen T, Zhang Z. GEPIA2: an enhanced web server for large-scale expression profiling and interactive analysis. Nucleic Acids Res. 2019;47(W1):W556\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZheng H, Huang B, Zhang B, Xiang Y, Du Z, Xu Q, Li Y, Wang Q, Ma J, Peng X, et al. Resetting Epigenetic Memory by Reprogramming of Histone Modifications in Mammals. Mol Cell. 2016;63(6):1066\u0026ndash;79.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang C, Liu X, Gao Y, Yang L, Li C, Liu W, Chen C, Kou X, Zhao Y, Chen J, et al. Reprogramming of H3K9me3-dependent heterochromatin during mammalian embryo development. Nat Cell Biol. 2018;20(5):620\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDahl JA, Jung I, Aanes H, Greggains GD, Manaf A, Lerdrup M, Li G, Kuan S, Li B, Lee AY, et al. Broad histone H3K4me3 domains in mouse oocytes modulate maternal-to-zygotic transition. Nature. 2016;537(7621):548\u0026ndash;52.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang L, Zhang J, Duan J, Gao X, Zhu W, Lu X, Yang L, Zhang J, Li G, Ci W, et al. Programming and inheritance of parental DNA methylomes in mammals. Cell. 2014;157(4):979\u0026ndash;91.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo Y, Hitz BC, Gabdank I, Hilton JA, Kagda MS, Lam B, Myers Z, Sud P, Jou J, Lin K, et al. New developments on the Encyclopedia of DNA Elements (ENCODE) data portal. Nucleic Acids Res. 2020;48(D1):D882\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDuan J, Zhu L, Dong H, Zheng X, Jiang Z, Chen J, Tian XC. Analysis of mRNA abundance for histone variants, histone- and DNA-modifiers in bovine in vivo and in vitro oocytes and embryos. Sci Rep. 2019;9(1):1217.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi B, Qing T, Zhu J, Wen Z, Yu Y, Fukumura R, Zheng Y, Gondo Y, Shi L. A Comprehensive Mouse Transcriptomic BodyMap across 17 Tissues by RNA-seq. Sci Rep. 2017;7(1):4200.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Multi-omics integration, Zygotic genome activation, Early embryo development, Tumorigenesis and metastasis","lastPublishedDoi":"10.21203/rs.3.rs-9174297/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9174297/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e Zygotic genome activation (ZGA) represents a pivotal transition in early embryonic development, marking the onset of gene expression following fertilization. Despite its fundamental importance, precisely determining the timing and identifying the key genes involved in ZGA remains a significant challenge.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e Based on time-course patterns from RNA-seq data spanning all developmental stages, we proposed a computational framework to identify ZGA genes and the onset of ZGA across diverse species. In mice, we identified 690 ZGA-associated genes, 119 of which were previously uncharacterized. Furthermore, we defined a pivotal gene subset termed Sustained ZGA (S-ZGA) genes, which are activated from ZGA and maintain sustained expression throughout subsequent development. Epigenetic analyses revealed that promoter accessibility and H3K4me3 enrichment are primary regulatory mechanisms for S-ZGA genes. Notably, these genes exhibit strong enrichment in tumorigenesis and metastatic processes.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e This study establishes a computational framework that operates independently of prior knowledge to identify ZGA genes and precisely determining ZGA onset timing across diverse species, and defined a critical subset termed Sustained ZGA (S-ZGA) genes, which are potentially associated with tumorigenesis and metastatic processes. These findings provide a novel perspective on the molecular mechanisms underlying both normal development and cancer.\u003c/p\u003e","manuscriptTitle":"Dissecting the Expression Pattern during Embryonic Development and Unveiling Sustained ZGA Genes with Oncogenic Relevance","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-02 17:40:35","doi":"10.21203/rs.3.rs-9174297/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-05-08T18:37:59+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-17T04:14:17+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"211536195067044134473845347960043996808","date":"2026-04-07T13:08:15+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-04-04T00:30:34+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"207554661198147586750372712954831196590","date":"2026-03-29T17:42:54+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-29T15:14:53+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-24T20:46:15+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-24T02:00:51+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Genomics","date":"2026-03-24T01:54:29+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"31046640-a0a3-4839-a86f-42039490fc72","owner":[],"postedDate":"April 2nd, 2026","published":true,"recentEditorialEvents":[{"type":"decision","content":"Revision requested","date":"2026-05-08T18:37:59+00:00","index":"","fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-05-18T01:38:18+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-02 17:40:35","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9174297","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9174297","identity":"rs-9174297","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00