The Functional Map of Ultraconserved Regions in Humans, Mice and Rats

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract BACKGROUND: Ultraconserved regions (UCRs) encompass 481 DNA segments exceeding 200 base pairs (bp), displaying 100% sequence identity across humans, mice, and rats, indicating profound conservation across taxa and pivotal functional roles in human health and disease. Despite two decades since their discovery, many UCRs remain to be explored owing to incomplete annotation, particularly of newly identified long non-coding RNAs (lncRNAs), and limited data aggregation in large-scale databases. This study offers a comprehensive functional map of 481 UCRs, investigating their genomic and transcriptomic implications: (i) enriching UCR annotation data, including ancestral genomes; (ii) exploring lncRNAs containing T-UCRs across pan-cancers; (iii) elucidating UCR involvement in regulatory elements; and (iv) analyzing population single-nucleotide variations linked to motifs, expression patterns, and diseases. RESULTS: Our results indicate that, although a high number of protein-coding transcripts with UCRs (1,945 from 2,303), 1,775 contained UCRs outside CDS regions. Focusing on non-coding transcripts, 355 are mapped in 85 lncRNA genes, with 35 of them differentially expressed in at least one TCGA cancer type, seven lncRNAs strongly associated with survival time, and 23 differentially expressed according to single-cell cancer analysis. Additionally, we identified regulatory elements in 373 UCRs (77.5%), and found 353 SNP-UCRs (with at least 1% frequency) with potential regulatory effects, such as motif changes, eQTL potential, and associations with disease/traits. Finally, we identified 4 novel UCRs that had not been previously described. CONCLUSION: This report compiles and organizes all the above information, providing new insights into the functional mechanisms of UCRs and their potential diagnostic applications.
Full text 195,194 characters · extracted from preprint-html · click to expand
The Functional Map of Ultraconserved Regions in Humans, Mice and Rats | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article The Functional Map of Ultraconserved Regions in Humans, Mice and Rats Bruno Thiago de Lima Nichio, Liliane Santana Oliveira, Ana Carolina Rodrigues, and 11 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4837600/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract BACKGROUND: Ultraconserved regions (UCRs) encompass 481 DNA segments exceeding 200 base pairs (bp), displaying 100% sequence identity across humans, mice, and rats, indicating profound conservation across taxa and pivotal functional roles in human health and disease. Despite two decades since their discovery, many UCRs remain to be explored owing to incomplete annotation, particularly of newly identified long non-coding RNAs (lncRNAs), and limited data aggregation in large-scale databases. This study offers a comprehensive functional map of 481 UCRs, investigating their genomic and transcriptomic implications: (i) enriching UCR annotation data, including ancestral genomes; (ii) exploring lncRNAs containing T-UCRs across pan-cancers; (iii) elucidating UCR involvement in regulatory elements; and (iv) analyzing population single-nucleotide variations linked to motifs, expression patterns, and diseases. RESULTS: Our results indicate that, although a high number of protein-coding transcripts with UCRs (1,945 from 2,303), 1,775 contained UCRs outside CDS regions. Focusing on non-coding transcripts, 355 are mapped in 85 lncRNA genes, with 35 of them differentially expressed in at least one TCGA cancer type, seven lncRNAs strongly associated with survival time, and 23 differentially expressed according to single-cell cancer analysis. Additionally, we identified regulatory elements in 373 UCRs (77.5%), and found 353 SNP-UCRs (with at least 1% frequency) with potential regulatory effects, such as motif changes, eQTL potential, and associations with disease/traits. Finally, we identified 4 novel UCRs that had not been previously described. CONCLUSION: This report compiles and organizes all the above information, providing new insights into the functional mechanisms of UCRs and their potential diagnostic applications. ultraconserved regions non-coding RNAs lncRNAs ancestral genomes pan-cancer variation Figures Figure 1 Figure 2 Figure 3 Figure 4 1. INTRODUCTION Ultraconserved regions (UCRs) are defined as 481 DNA segments longer than 200 bp, with 100% sequence identity in their alignment with humans, mice, and rats [ 1 ]. These regions are also highly conserved across disparate taxa, including vertebrates, insects, worms, yeast, and other species, and are essential DNA markers for phylogenomic analyses [ 1 – 3 ]. Approximately 93% of the UCRs were transcribed in at least one human tissue tested, with nearly 79% showing ubiquitous transcription across a wide range of normal human tissues. Because of this characteristic, these regions are called transcribed UCRs (T-UCRs) [ 4 ]. The first UCR annotation proposed by Bejerano et al. [ 1 ] focused on their overlap with protein-coding genomic regions, with surprisingly only 23.0% of the UCRs overlapping with known mature mRNA sequences [ 1 ]. They identified 256 sequences with no evidence of transcription from any matching expressed sequence tag or mRNA, and 111 overlapped the mRNAs of known human protein-coding genes. Additionally, 114 sequences were "inconclusive," where the evidence for transcription was inconclusive. In 2010, Mestdagh and colleagues reorganized all UCR sequences into six categories based on protein-coding information, classifying those regions as intergenic (38.7%), intronic (42.6%), exonic (4.2%), partly exonic (5.0%), and exon-containing (5.6%). Additionally, 3.9% were classified as “multiple” when the genomic annotation varied because of the host gene splice variants [ 5 ]. Recently, in 2022, Snetkova and collaborators observed that 371 (77.1%) were non-coding UCRs and 110 (22.9%) were UCRs that overlapped the exons of protein-coding genes. In this study they reviewed the biological importance of these UCRs, mainly in regulating gene expression during embryogenesis [ 6 ]. Coding sequences are typically highly conserved and are subject to more stringent negative selection pressures [ 7 ]. Nevertheless, more than 75% of the UCRs were in intergenic or intronic, indicating that negative selection pressure can also act on these non-coding sequences[ 8 ]. Additionally, among the intronic transcribed UCRs (T-UCRs), almost 58% were detected in the antisense orientation compared with the host gene [ 4 ]. These results suggest that many UCRs may be transcribed as long non-coding RNAs (lncRNAs). However, a general annotation of the 481 ultraconserved regions with overlapping lncRNAs was readily unavailable. In addition to their potential roles in lncRNA function, UCRs are crucial enhancers during development. For example, Visel and colleagues reported that half of the non-exonic UCRs evaluated drove reproducible reporter gene expression in various tissues of developing mouse embryos [ 9 ], and this function was reinforced by others [ 6 , 10 ]. Additionally, highlighting the role of some T-UCRs, deregulated expression has been associated with several pathological conditions, including tumor types and Cancer disease [ 4 , 11 – 13 ]. Although the UCRs play an essential role in human health and disease; there is limited information regarding specific T-UCRs, the investigation of annotated lncRNA-T-UCRs largely improves our knowledge of the potential roles of these molecules. Therefore, even two decades after their initial identification, many UCRs remain to be explored. To address this, we constructed a comprehensive functional map of 481 UCRs. however, In this study, we investigated four significant contributions: (i) we analysed enrichment mining annotation data, including comparisons of the modern genome version and the uncharted ancestral genomes in 15 Neanderthal and ancient humans from different eras; (ii) examination of pan-cancer expression with the support of RNA-Seq and single-cell data analysis, focusing on the T-UCRs in lncRNA genes; (iii) evidence of UCRs in the regulatory elements, including enhancers and microRNA/transcription factor binding sites (TFBS); and (iv) population single-variation analysis, shedding light on their evolutionary dynamics and implications for changes in binding sites, expression, and disease susceptibility. Our research explored the omics landscape of UCRs, shedding light on non-coding regions and deepening our insights into UCR functionality, gene regulation, and disease mechanisms, underscoring their pivotal roles in genomic stability and disease pathogenesis. 2. RESULTS 2.1. UCR Annotation Enrichment Mining Insights We analyzed genomic and transcriptomic annotation data using the currently human reference and we compared this to previous UCR annotations, performing the same analysis for recent mouse and rat genomes (Fig. 1 , Supplementary Tables S1-S5). Adopting the genomic classification of Bejerano et al. [ 1 ], we observed that 126 (26.2%) and 355 (73.8%) UCRs were classified as coding and non-coding, respectively, in the human genome (Fig. 1 A). In the mouse and rat genomes, 133 (27.6%) and 348 (72.4%) sequences were classified as coding, respectively, while 82 (17.0%) and 399 (83.0%) were classified as non-coding, respectively (Fig. 1 B). Using the classification of Mestdag et al. [ 5 ], for human, we identified 261 (54.5%) UCRs as intronic, four as exonic (0.8%), four as exon-containing (0.8%), five as partly exonic (1.0%), 98 as intergenic (20.4%), and 109 as multiple (22.7%) (Fig. 1 C; Supplementary Table S2 ). In mouse, 223 (44.9%) UCRs were mapped to intronic, 13 (2.7%) to exonic, 10 (2.1%) to exon-containing, 21 (4.3%) to partly exonic, 93 (19.3%) to multiple, and 121 (25.3%) to intergenic regions (Fig. 1 C and Supplementary Tables S2 and S4). With respect to the rat genome, our analysis revealed 161 (33.5%) UCRs to an intron, 12 (2.5%) exonic, 14 (2.9%) exon-containing, 16 (3.3%) partly exonic, 28 (5.8%) multiple, and 250 (51.9%) intergenic regions (Fig. 1 C; Supplementary Table S2 ). To map UCR evidence at transcriptomic levels, we investigate the UCR sequences onto human, mouse and rat transcript annotations. In human, 384 UCRs overlapped with 2,303 transcripts (GENCODE v46). Among these, 1,945 transcripts were mapped to protein-coding T-UCRs (205 protein-coding genes overlapped with 315 T-UCRs), 355 to lncRNA-T-UCR (85 lncRNA genes overlapped with 100 T-UCRs), and three other transcripts were mapped to a pseudogene and two small non-coding RNAs (one miRNA and one miscellaneous RNA (abbrev. miscRNA)), each containing one T-UCR (Fig. 1 D). Among the 1,945 protein-coding transcripts annotated, 22 contained UCRs in the coding sequence (CDS) region, 148 overlapped with UCRs in the CDS, and 1,775 contained UCRs outside CDS (UCRs were identified in intronic or exonic non-CDS regions) (Fig. 1 E; Supplementary Table S3). In mouse (GENCODE vM35), 360 UCRs overlapped with 1,125 transcripts, of which 1,022 were protein-coding transcripts (205 protein-coding genes and four protein-coding genes to be experimentally confirmed), 101 were long non-coding transcripts (associated with 46 lncRNA genes), and two were other non-coding RNA (one miRNA and another miscRNA) (Fig. 1 D). Of the 1,022 protein-coding transcripts, 14 contained UCRs into the CDS region, 117 overlapped with UCRs in the CDS, and 891 contained UCRs outside the CDS (Fig. 1 E; Supplementary Table S4). In rat (ENSEMBLE 111), we identified 224 T-UCRs overlapping 935 transcripts, with 826 protein-coding transcripts (118 protein-coding genes), 50 lncRNA transcripts (associated with 17 lncRNA genes), 58 miscRNA transcripts (58 genes), and one miRNA (Fig. 1 D and Supplementary Table S5). Among the protein-coding transcripts, two contained UCRs in the CDS, 54 overlapped with UCRs in the CDS, and 770 contained UCRs outside the CDS (Fig. 1 E). This unique miRNA gene is associated with the uc.420/miR-3064 gene in the three species. 2.2. Conservation of UCRs in Ancient Human Genomes We also investigated the presence of UCRs in five Homo neanderthalensis genomes (one male and four females) and ten Homo sapiens genomes (nine males and one female) from various periods, including the Palaeolithic era to more recent times, and present-day tribal populations, such as those from the Khoisan Bantu people from Kalahari desert (Fig. 2 A) (Supplementary Table S6). We obtained all of 481 UCRs with 100% coverage and identity in these genomas, although some had regions with sequences containing few bases represented as ‘N’ without the exact nucleotide defined (Supplementary Table S7). Furthermore, both modern and ancient human genomes contained an additional set of novel UCRs with high conservation levels (sequences of at least 200 base pairs, minimum alignment coverage of 75%, and identity threshold of 80% for the UCR set) in distinct chromosomal regions, except in the Mezmaiskaya genome (Fig. 2 B and Supplementary Table S8). We found four novel UCRs align perfectly in the human, mouse, and rat genomes and follow Bejerano’s rules for UCR definition (Fig. 2 C; Table 1 and Supplementary Table S9), and we denominated them ‘extra’ UCRs (exUCRs as abbreviations and “euc.” for the code sequence). In particular, we observed the presence of euc.397 (chr18) in the analyzed genomes. We identified euc.129 (chr13) and euc.425 (chr16) in all genomes, except for ZKU and Bantu, respectively. Finally, euc.175 (chr10) was present in these genomes, except in the ZKU and Scandinavian genomes(Fig. 2 C). Table 1 ExUCR vs UCR coordinates, and overlapping gennomic annotations for human, mouse and rat. genome exUCR exUCR coordinates exUCR gene overlapping exUCR gene overlapping class UCR UCR coordinates UCR gene overlapping UCR gene overlapping class human euc.129 chr13:97356685–97356888 MBNL2 (exonic) protein-coding uc.129 chr3:152446598–152446809 MBNL1 (exonic) protein-coding mouse euc.129 chr14:120633013–120633216 MBNL2 (exonic) protein-coding uc.129 chr3:60521998–60522209 MBNL1 (exonic) protein-coding rat euc.129 chr15:97508498–97508701 EFCAB12-ARSJ (intergenic) NA (intergenic) uc.129 chr2:144798866–144799077 MBNL1 (exonic) protein-coding human euc.175 chr10:129893328–129893555 EBF3 (intronic) protein-coding uc.175 chr5:158914830–158915079 EBF1 (intronic) protein-coding mouse euc.175 chr7:136852289 136852516 EBF3 (intronic) protein-coding uc.175 chr11:44691125–44691374 EBF1 (intronic) protein-coding rat euc.175 chr1:192052056 192052283 EBF3 (intronic) protein-coding uc.175 chr10:29284049–29284298 EBF1 (intronic) protein-coding human euc.397 chr18:25285204–25285497 ZNF521 (intronic) protein-coding uc.397 chr16:49701937–49702247 ZNF423 (intronic) protein-coding mouse euc.397 chr18:14029595–14029888 ZFP521 (intronic) protein-coding uc.397 chr8:88563296–88563607 ZFP423 (intronic) protein-coding rat euc.397 chr18:4997387–4997680 ZFP521 (intronic) protein-coding uc.397 chr19:19231570–19231880 ZFP423 (intronic) protein-coding human euc.425 chr16:49701982–49702237 ZNF423 (intronic) protein-coding uc.425 chr18:25285243–25285567 ZNF521 (intronic) protein-coding mouse euc.425 chr8:88563341–88563596 ZNF423 (intronic) protein-coding uc.425 chr18:14029634–14029958 ZNF521 (intronic) protein-coding rat euc.425 chr19:19231580–19231835 SYNPOL2L-MYO5B (intergenic) NA (intergenic) uc.425 chr18:4997426–4997750 ZNF521 (intronic) protein-coding Owing to the lack of comprehensive annotations for the Neanderthal and ancient human genomes, we utilized the human GENCODE v46 reference annotation to describe the exUCRs and compared these annotations with those from mouse and rat. All exUCRs overlapped with protein-coding genes in human and mouse annotations, whereas euc.129 and euc.425 were intergenic in rat (Table 1 and Supplementary Tables S10 and S11). Euc.129 is located between EFCAB12 gene (upstream) and ARSJ gene (downstream), whereas euc.425 is located between the SYNPOL2L gene (upstream) and the MYO5B gene (downstream). Interestingly, all extra UCRs were located in protein-coding genes from the same family from the related UCR, even maintaining overlap within the region (exonic or intronic) of the gene (Table 1 ). This finding suggests that the ultraconserved sequence is essential for the structure of these proteins or for transcript regulation. 2.3. The T-UCR in LncRNAs: Insights in Pan-Cancer Analyses There is limited information regarding T-UCRs (Supplementary Table S12). The investigation of annotated lncRNA-T-UCRs improves our knowledge of the potential functional roles in specific pathologic conditions, such as Cancer disease. We performed differential expression analysis of lncRNA-T-UCRs via RNA-seq data (bulk and single-cell) from public cancer databases, including The Cancer Genome Atlas Program (TCGA) and the Curated Cancer Cell Atlas (3CA). Differential expression of lncRNA-T-UCRs in cancer We selected datasets from the TCGA database using sample size as the only criterion (at least 30 samples per condition, tumor, and non-tumoral). Ten cancer datasets were obtained: breast invasive carcinoma (BRCA), colon adenocarcinoma (COAD), head and neck squamous cell carcinoma (HNSC), kidney renal clear cell carcinoma (KIRC), kidney renal papillary cell carcinoma (KIRP), lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), prostate adenocarcinoma (PRAD), thyroid carcinoma (THCA), and uterine corpus endometrial carcinoma (UCEC). The number of non-tumoral and tumoral samples for each cancer is detailed in Supplementary Table S13. This analysis aimed to detect lncRNAs containing at least one T-UCR transcript isoform. Thus, to identify differentially expressed lncRNA genes, we used the 85 lncRNAs identified in our transcript mapping analysis (i.e., 85 lncRNAs mapped to 100 T-UCRs). However, the RNA-Seq data did not include data for two lncRNAs (ENSG00000288692 and ENSG00000289413). Therefore, we performed tumor and single-cell differential expression analyses of 83 lncRNAs. Using the Wilcoxon rank-sum test to calculate differentially expressed genes and filtering the results by p-values ( ≤ 0.05), we detected differentially expressed (DE) lncRNAs in each cancer type (Supplementary Figure S1 and Supplementary Table S14). Considering the threshold of + 2 log2foldChange, the number of genes upregulated in cancers ranges from three (HNSC and PRAD) to 11 (LUSC). With respect to the downregulated lncRNAs (log 2 foldchange ≤ -2), we identified only one lncRNA-T-UCR in LUAD, THCA, and UCEC cancer types and four in KIRC. BRCA and PRAD cancers did not present downregulated lncRNA-T-UCR (Supplementary Table S14). When we compared the expression of lncRNA genes across all tumor types, we identified 35 lncRNAs that were differentially expressed in at least one cancer type (Supplementary Table S15). Using TPM values, we constructed a heatmap using these 35 DE lncRNA-T-UCRs for each cancer type (Supplementary Figure S2 ). None of the lncRNA-T-UCRs were differentially expressed in any cancer type. However, we detected two genes, FEZF1-AS1 /uc.232 and DLX6-AS1 /uc.220/uc.221, which were upregulated in the seven cancer types (Fig. 3 A). The lncRNA FEZF1-AS1 /uc.232 (ENSG00000230316.7) was upregulated in BRCA, COAD, HSNC, LUAD, LUSC, PRAD, and UCEC (fold-change ranging from 2.3 in UCEC to 8,82 in COAD); whereas DLX6-AS1 /uc.220/uc.221 was upregulated in BRCA, COAD, HSNC, KIRC, LUAD, LUSC, and UCEC (fold-change ranging from 1.07 in PRAD to 6,7 in LUSC). We did not identify any downregulated lncRNA-T-UCRs across all cancer types. However, we observed downregulated three lncRNAs in three cancers: MECOM-AS1 /uc.136 (HNSC, KIRC, KIRP), MEF2C-AS1 /uc.167 (COAD, LUSC, and UCEC), and ENSG00000257277.2/uc.70, an lncRNA without a gene symbol defined and antisense to the coding ARHGAP15 gene (differentially expressed in COAD, LUAD, LUSC). Furthermore, four lncRNAs were up-regulated in one or two cancer types and downregulated in another ( SAMMSON /uc.116, HOXB-AS3 /uc.415/uc.416/uc.417, LINC01505 /uc.266, and ERICH2-DT /uc.95). Finally, regarding lncRNAs expressed exclusively in one type of cancer, 10 were upregulated and seven were downregulated (Fig. 3 A). The impact of lncRNA-T-UCRs on patient prognosis To analyze the impact of lncRNA-T-UCRs on patient prognosis via TCGA survival data, we investigated the expression of all 83 human lncRNA-T-UCRs in the 10 TCGA cancer types selected, and 19 lncRNAs were associated with survival in the Cox proportional-hazard analysis, seven of which also exhibited significance in the Kaplan-Meier model (Supplementary Figure S3 and Supplementary Table S16). Focusing on lncRNA-T-UCRs that were significantly associated with survival in both analyses, we identified LINC00583/uc.250, FEZF1-AS1/uc.233, FOXG1-AS1/uc.361, DPH6-DT/uc.382, PKN2-AS1/uc.31, AQP4-AS1/uc.427, and ENSG00000270087.5/uc.286. Interestingly, the lncRNA-UCR analysis revealed a cancer-specific clinical impact, since none were shared among more than one cancer type. This result emphasizes the tissue specificity of these molecules, which is a biologically prominent characteristic of these molecules, highlighting their use as prognostic markers. lncRNA-T-UCR single-cell RNA-seq (scRNA-seq) data analysis To identify the 83 lncRNA-T-UCRs expressed within the tumor microenvironment, we obtained publicly available scRNA-seq data from the Curated Cancer Cell Atlas database. Among the lncRNA-T-UCRs, 35 and 23 were differentially expressed according to single-cell analysis (Fig. 3 B). We discovered that NR2F1-AS1/uc.169 was the most highly expressed lncRNA in the fibroblasts of all the tumor types analyzed. NR2F1-AS1/uc.169 was also expressed in pericytes from BRCA and COAD patients and in malignant cells from LUAD and COAD patients (Fig. 3 B). Additionally, IQCH-AS1/uc.389 was highly expressed in epithelial cells from HNSC and mast cells from BRCA and LUAD. MEF2C-AS1/uc.167 was highly expressed in B cells from HNSC, BRCA, and COAD. Specifically, in LUAD, FEZF1-AS1/uc.232 was highly expressed in malignant cells, HOXB-AS3/uc.415 to uc.417 was highly expressed in fibroblasts, and AQP4-AS1/uc.427 was highly expressed in epithelial cells. LINC01117/uc.109 was highly expressed in the stromal and malignant cells of BRCA and COAD. LINC01505/uc.266 was highly expressed in NK cells of KIRC, whereas SAMMSON/uc.116 was expressed in the B cells of KIRC. SATB1-AS1/uc.113 was highly expressed in immune cells, and PKN2-AS1/uc.31 was highly expressed in malignant and epithelial cells in HNSC (Fig. 3 B). 2.4. UCRs in Regulatory Elements We performed enrichment mining for the presence of regulatory elements within UCRs via the ORegAnno tool [ 14 ]. We identified regulatory elements in 373 UCRs (77.5%), comprising 234 (62.7%) instances as regulatory regions (RR) (including promoters, enhancers, etc.), 73 (19.6%) as transcription factor binding sites (TFBS), 64 (17.1%) as combinations of TFBS and RR, one as a TFBS and miRNA binding site (SMARCA4 and miR-429, respectively), and one as a miRNA binding site (miR-214-3p) in the human genome (Supplementary Table S17). Additionally, 308 UCRs were 100% inserted into RRs and TFBS. Among the 45 distinct TFBSs, SMARCA4 was the binding site most representative of 62 UCRs, followed by EGR1, FOXA1, CTCF, and CEBPB, which contained 20, 13, 12, and 11 UCRs, respectively. Despite SMARCA4, 51 UCRs were completely inserted into the TFBS sequence, and another 11 UCRs partially overlapped. Among the UCRs related to EGR1, FOXA1, CTFC, and CEBPB TFs, 18, 5, 5, and 4 UCRs, respectively, were contained in these TFBSs. The 19 UCRs exhibited agglomerative numbers of TFBS: uc.213 and uc.246 related to six TFBS; uc.88, uc.329, uc.374, and uc.392 with five TFs; uc.95, uc.162, uc.393, and uc.417 with four TFBS; and uc.74, uc.189, uc.232, uc.277, uc.398, uc.409, uc.412, uc.416, and uc.456 with three TFBS. Finally, in 26 UCRs, we identified two TFBS associated with the same UCR, and a single TFBS in the remaining 93 UCRs. According to the VISTA Enhancer database, 295 of the 373 (79.1%) regulatory regions within UCRs were identified as human enhancers [ 15 ]. For four UCRs (uc.29, uc.347, uc.350 and uc.355) we report the presence of two simultaneous enhancers, totaling 299 enhancers. Among these enhancers, almost 80% (240/299) were experimentally validated and linked with the respective UCR, according to previous studies [ 9 , 10 ]. In addition, 152 and 147 enhancers were classified as positively and negatively regulating transcription, respectively (Supplementary Table S17). 2.5. Personalized Genomics: Pan-SNP Data Analysis SNP-UCR compendium identified in human databases and pangenomes The availability of complete human genomes in various databases allowed us to examine the presence of single nucleotide polymorphisms (SNPs) in UCR sequences and investigate their potential functional effects. In our study, we identified 353 SNPs (with at least 1% frequency) that mapped to 223 UCRs in the Genome Aggregation Database (gnomAD) [ 16 ], 1000 Genomes Project (1Kgenomes) and Archive of Brazilian Mutations (ABraOM) [ 17 ] (see Supplementary Table S18). We found that 17 (4.8%) of these SNPs are shared across all three databases (gnomAD, 1kgenomes, and ABraOM), 301 (85.3%) were shared between gnomAD and 1k genomes but were not present in ABraOM, 33 (9.3%) are unique to gnomAD, and two (0.6%) are unique to 1kgenomes (Fig. 4 A). The predominant SNP types were single nucleotide variants (SNVs) (332/353), followed by insertion or deletion markers (INDELs) (17/353), deletions (3/353), and insertions (1/353). Most of the variants were located in introns of the coding gene, including 146 (41%) SNPs. Additionally, 13 SNPs were found in exonic-coding regions (of them, six were synonymous and seven were nonsynonymous SNVs), seven in the 3’UTR, three in the 5’UTR of protein-coding genes, and six in exonic and 58 in intronic regions of lncRNA genes. A total of 120 SNPs were identified in the intergenic regions. Concerning the ethnicity identified in these databases, 36 SNPs were present in African, European, American, East, and South Asian populations, whereas 63, 48, 3, 19, and 20 SNPs were unique to African, European, American, East, and South Asian populations, respectively. In addition, 164 SNPs were distributed among at least two ethnicities, and 17 SNPs were shared with the Brazilian population (Fig. 4 A). UCR-SNPs in regulatory regions associated with diseases To understand the effects of some SNPs on the regulation of nearby genes and those associated with diseases, we assessed whether UCR-SNPs mapped to a sequence of a transcription factor binding motif (called a motif SNP), which SNPs affects gene expression (called eQTL SNPs), are associated with a complex disease or trait through a Genome-Wide Association Study (GWAS SNPs) [ 18 ]. Using the Super Enhancer database (SEdb v2.0), we identified 48 SNPs with potential regulatory effects in 43 UCRs (Supplementary Tables S19 and S20). 21 UCRs containing exclusively motif SNPs; 12 UCRs with eQTL and motif SNPs; 11 with eQTL SNPs; one with motif, GWAS, and eQTL SNPs; one containing eQTL and GWAS SNPs; and two UCRs containing motif and GWAS SNPs (Supplementary Table S19 and S20). For UCRs with GWAS SNPs, we identified that rs12981/uc.268 ( RC3H2 gene, 3’UTR variant) was associated with coronary artery calcification [ 19 ], rs17291131/uc.211 ( SKAP2 gene, intron variant) was associated with type 2 diabetes [ 20 ], rs2056117/uc.140 (intergenic variant) was associated with systemic lupus erythematosus [ 21 ], and rs11190870/uc.302 (intergenic variant) was associated with scoliosis [ 22 – 24 ]. In addition, we identified eight UCRs with SNPs in more than 1% of the global population that were correlated with disease/trait according to the GWAS catalogue [ 25 ] (Supplementary Table S21). rs1861100/uc.53 contains an intron variant in LINC01122 and was related to body mass index (obesity) [ 26 – 28 ]). Additionally, rs142252570/uc.380, an intergenic variant related to the LINC02304/LINC02325 genes, was associated with menarche (age at onset) [ 29 ]. The SNP rs142252570/uc.380, an intergenic variant related with the RSRC1 gene, was associated with body mass index and drinks per week [ 30 , 31 ]. The SNP rs73174306/uc.136, an intron variant in MECOM and MECOM-AS1 genes, was related to blood glucose levels and pancreatic hormone levels [ 32 – 36 ]. The SNP rs1538101/uc.252, an intron variant in the BNC2 gene, was related to velopharyngeal dysfunction [ 37 ]. The SNP rs11190870/uc.302, an intergenic variant reported with the LINC01514/LBX1 genes, was related to scoliosis [ 22 , 24 , 38 , 39 ]. The SNP rs182041872/uc.315, an intergenic variant reported with the ACADSB/HMX3 genes, was related to antidepressant treatment resistance [ 33 ]. Investigation of SNPs in UCR sequences for ancient humans, Neanderthals, and human pangenomes (Fig. 4 B and Supplementary Table S22) revealed six altered SNPs with a prevalence of > 50%. Among these SNPs, five were located in intronic regions correlated with GWAS/SEdb with impact disease. Rs2682406/uc.133 (Fig. 4 B in the upper triangle) and rs1861100/uc.53 (Fig. 4 B in the lower triangle) were related to breast cancer and obesity, respectively [ 40 , 41 ]. Rs1538101/uc.252, mapped to the protein-coding gene BNC2 , was associated with velopharyngeal dysfunction and rs11896224/uc.82, an intron variant of the lncRNA LINC01876 was also associated with velopharyngeal dysfunction. Finally, two SNVs were found to be intergenic: rs11190870/uc.302 (associated with scoliosis) and rs7092999/uc.295 with no disease association. 3. DISCUSSION Some regions related to the 481 UCRs initially described by Bejerano and colleagues in 2004 [ 1 ] have been described as important in cellular processes and are associated with diseases [ 12 ]. However, despite two decades since the first description, the role of most regions has yet to be discovered. Previous UCR annotations have focused predominantly on DNA coding to provide functional evidence. In this study, we innovatively enriched the annotation of UCR data at both the genomic and transcriptomic levels by incorporating the latest human, mouse, and rat genome versions. Our analysis reinforces the presence of most UCRs in the non-coding and outside CDS regions. Despite this, the transcripts associated with these UCRs are mostly annotated in the intronic regions of protein-coding transcripts, with a quarter annotated as lncRNA genes on the basis of biotype annotation. To the best of our knowledge, this is the first study to focus on UCRs mapped to lncRNAs, providing accessible and organized information such as gene symbols and IDs. Comparing human, mouse, and rat annotations is an interesting way to observe human and mouse similarities. For example, the percentages of intergenic UCRs in human and mouse is approximately 20 and 25%, respectively; in contrast, 52% of UCRs were found in intergenic positions in rat. This difference was also verified by examining the transcripts that included UCR. The observed differences may result from the limited information available for the rat genome because the rat transcript data are significantly less complete than those available for the mouse, particularly in terms of transcribed and exon regions [ 42 ]. With respect to transcript analysis, a unique miRNA gene, MIR3064 , was wholly contained in the uc.420 genomic region of the three organisms. MiR-3064 is related to diverse diseases [ 43 – 47 ]. In cancer, miR-3064 is targeted by several lncRNAs/circRNAs associated with important pathways such as PI3K/Akt [ 48 ], hTERT [ 49 ], glycolysis [ 50 ], and Src, and plays crucial tumor suppressive roles [ 48 , 51 – 59 ]. Given the extreme conservation of UCRs in mammals, they are expected to be preserved in the ancestral hominids. However, this has not been previously confirmed. Considering the analysis of the 481 UCRs within ancestral genomes, including Neanderthal and ancient human genomes from different eras, as expected, UCR sequences in hominid genomes were also 100% conserved. Notably, we identified four novel UCRs in humans (modern and ancestral genomes), mouse, and rat that had not been previously described. Interestingly, these four novel UCRs (exUCRs) are protein-coding genes from the same family as the UCR related genes, suggesting a potential motif associated with the conserved sequence in these molecules. euc.129 is a novel UCR related to uc.129, which maps to the MBNL1 gene. In contrast, euc.129 is located in the MBNL2 gene, which plays a role in MBNL1. Similarly, the new UCR exuc.175 overlaps the EBF3 gene while uc.175 the EBF1 ; exuc.397 overlaps the ZNF521 gene while the uc.397 the ZNF423 ; and the exuc.425 overlaps the ZNF423 gene, while uc.425 the ZNF521 . MBNL and ZNF are both zinc finger proteins, and EBF are DNA-binding proteins, all of which act as transcription factors [ 60 ]. In human, our approach enabled us to associate 85 lncRNA-T-UCRs with previously unrelated functions, demonstrating its potential for analysis in public databases through Ensemble codes and gene symbols. This is the first study to analyze the lncRNA-T-UCR panel in a pancancer context. These lncRNAs play essential roles in cancer [ 61 , 62 ]; however, they are not previously recognized as containing T-UCRs. For example, the uc.232 sequence is almost entirely in exon 2 of FEZF1-AS1-202 and in the opposite strand in exon 1 of FEZF1-203 (with no protein associated with the retained intron). A unique study focused on uc.232 showed that the T-UCR was differentially expressed in TCGA gastric cancer samples compared with their non-tumoral counterparts, but no expression of this T-UCR was observed when a different cohort of GC samples was utilized [ 63 ]. Here, we reinforced the importance of FEZF1-AS1 in multiple tumor types and highlighted the relationship between this lncRNA and uc.232. In the survival analysis, we identified seven lncRNA-T-UCRs in cancer cohorts: FEZF1-AS1/uc.233, LINC00583/uc.250, FOXG1-AS1/uc.361, DPH6-DT/uc.382, PKN2-AS1/uc.31, AQP4-AS1/uc.427, and ENSG00000270087.5/uc.286. Two of these genes were previously associated with prognosis in cancer patients, including the widespread deregulation of FEZF1-AS1, which is associated with worse patient prognosis in 13 cancer types [ 64 ]. Additionally, PKN2-AS1 was previously identified as being related to the prognosis of patients with bladder cancer [ 65 ]. The expression of the other five lncRNAs is associated with survival time for the first time herein. To the best of our knowledge, this is the first study to analyze T-UCR expression via single-cell data. Data from 35 lncRNA-T-UCRs were obtained, and NR2F1-AS1/uc.169 was the most highly expressed lncRNA in the fibroblasts of all the tumor types analyzed. NR2F1-AS1/uc.169 was also expressed in pericytes from BRCA and COAD patients and in malignant cells from LUAD and COAD patients. The uc.169 sequence overlapped with the coding gene NR2F1 and its antisense transcript NR2F1-AS1. NR2F1 encodes a nuclear hormone receptor that functions as a transcription factor. This protein is critical in the developmental process of the brain[ 66 ]. In contrast, antisense transcripts have mostly been studied in the context of cancer [ 67 ]. NR2F1-AS1 plays an oncogenic role in several tumor types and is associated with proliferation, migration, and drug resistance [ 68 ]. In colorectal and cervical cancers, this lncRNA is downregulated and associated with a tumor-suppressive role[ 67 – 69 ]. Single-cell analysis of NR2F1-AS1 expression in different cell types is underexplored. This lncRNA is usually highly expressed in cancer samples, but in our study, this lncRNA was explicitly found in fibroblast cells for all tumor types analyzed. Cancer-associated fibroblasts (CAFs) are part of the tumor stroma. Although not malignant, fibroblasts may play a role in cancer progression, such as supporting tumor growth, epithelial-mesenchymal transition/metastasis, and therapy resistance [ 70 ]. Fibroblasts may also be associated with angiogenesis, as the interaction between tumors and stromal cells can increase vascular endothelial growth factor (VEGF) [ 71 ]. Furthermore, NR2F1-AS1/uc.169 was expressed in pericytes in BRCA and COAD. Pericytes are fibroblast-like cells that wrap around endothelial cells in arterioles, capillaries, and venules and are strongly associated with angiogenesis [ 72 ]. Our findings underscore the importance of studying lncRNA-T-UCRs in the context of single cells, suggesting more specific functions than previously known. By organizing data on regulatory elements within the UCRs of the human genome, we found evidence of regulatory elements in 373 UCRs (77.5% of all UCRs). These regulatory elements include enhancers, transcription factors, and miRNA binding sites. The two most abundant TFs in the UCR sequences are SMARCA4 and EGR1. Both TFs are essential for the regulation of critical networks under both physiological and pathological conditions [ 73 , 74 ], are deregulated during cancer development, and may deregulate the regulation of transcripts related to UCRs. Finally, we investigated frequent SNPs in UCR sequences, including thousands (~ 128.700) of human whole-genome sequences. Previously, a broad analysis and utilizing ultraconserved regions that include comparisons with other species revealed that some polymorphisms have also been identified, however, the vast majority of these polymorphisms are rare events [ 75 ]. In the present study, we focused on 481 UCRs and in the most frequent SNPs (> 1%), identifying 353 SNPs 25 of which were associated with eQTL effects. For example, rs34384113/uc.109 is in the intron portion of LINC01117 with an eQTL effect in HOXD10 and LINC01117 expression; rs117486161/uc.220, in the intron of the DLX6-AS1 gene has an eQTL effect on DLX6-AS1 expression; and rs56805315/uc.417 mapped at the intron of HOXB-AS3 and 5’UTR of the HOXB6 gene has an eQTL effect on AC091133.1, CDK5RAP3, and HOXB6 expression. Interestingly, the lncRNAs LINC01117/uc.109, DLX6-AS1/uc.220, and HOXB-AS3/uc.417 were dysregulated in cancer, as highlighted in our expression analysis. Our in-depth and comprehensive functional map revealed that almost all UCRs provide tips for understanding their functions. These complete maps of UCRs provide new insights into these regions, accelerating the understanding of functional mechanisms and their potential use as biomarkers of these related regions. 4. METHODS Data sources UCR data acquisition The human UCR coordinates were obtained from Bejerano et al. (2004). The sequences of the UCRs were extracted from the human genome via coordinates (obtained from GENCODE release 45 to GRCh38 (hg38).p14)). We used bedtools[ 76 ] version 2.31.0 with the getfasta algorithm. General command: bedtools getfasta [OPTIONS] -fi -bed The genomic annotation data The human genomic annotation (GFF/GTF formats) was obtained from GENCODE[ 77 ] release 45 to the GRCh38(hg38).p14 genome (available at https://www.gencodegenes.org/human/ ). For mouse genomic annotation, we used the GENCODE release M34 for the GRcm39/mm39 genome (available at https://www.gencodegenes.org/mouse ). Finally, the rat genomic annotation was obtained from RefSeq 106 (ENSEMBLE release 111) for the mRatBN7.2 rn7.2 genome (available at https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_015227675.2 ). UCR genome remapping For human, we remapped the 481 UCRs for NCBI37(hg16) genomic coordinates (BED format) via the liftOver tool ( https://genome.ucsc.edu/cgi-bin/hgLiftOver ) to human GRCh37(hg19), mouse NCBI34(mm6) and rat Baylor 3.1(rn3). The coordinates were converted from GRCh37(hg19) to GRCh38(hg38).p14 for reference to the human genome, NCBI34(mm6) to GRCm39/mm39 for mouse and Baylor 3.1(rn3) to mRatBN7.2 rn7 for rat via the NCBI-remapping tool ( https://www.ncbi.nlm.nih.gov/genome/tools/remap ). UCR overlapping functions The UCR coordinates (BED format) were overlapped with respective genomic annotation releases in GFF/GTF (GENCODE and RefSeq annotations) or BED formats (mirRNA, regulatory, and SNP/SNVs annotations) and extracted via the R (v4.3.2) script built with the foverlaps function from package data.table (v1.14.2) available at https://www.rdocumentation.org/packages/data.table/versions/1.14.2 . Identification of UCRs in the miRNA annotation data To detect which UCRs overlapped with miRNAs and mirtrons, we obtained these features from miRBase release 22.1 [ 78 ], MirGeneDB 2.1 [ 79 ], and MirtronDB 1.1 [ 80 ]. Neanderthal and ancient human genomes Neanderthal genomes were obtained from direct genome projects in BAM format: Altai ( https://www.eva.mpg.de/genetics/genome-projects/neandertal/ ) [ 81 ], Chagyrskaya 8 [ 82 ] ( https://www.eva.mpg.de/genetics/genome-projects/chagyrskaya-neandertal/ ) , Mezmaskaya 3 (PRJNA765125), Vindija 19.33 [ 83 ] ( http://cdna.eva.mpg.de/neandertal/Vindija/bam/Pruefer_etal_2017/Vindija33.19/ ) , and Denisovan 8 [ 84 ] ( https://www.eva.mpg.de/genetics/genome-projects/denisova/ ). The ancient humans in the ENA database: Ust'-Ishim (PRJEB6622), Yana (PRJEB29700), AHUR_2064 (PRJEB29074), Sunghir Burial 3 (SIII) (PRJEB22592), Zlatý kůň (ZKU) (PRJEB39040), Sumidouro 5 (PRJEB29074) Scandinavian hunter-gatherers (PRJEB32786), Ötzi (Tyrolean Iceman) (PRJEB56570), and Ayayema A460 (PRJEB29074). The Kalahari genome was obtained from RAW data (FASTq, paired-end sequences) from the work of Schuster et al. 2010 work [ 85 ]. More information about the samples is provided in Supplementary Table S6. Kalahari genome assembly All genomic FASTAs were obtained in the BAM format via the consensus algorithm ( http://www.htslib.org/doc/samtools-consensus.html ) from SAMtools [ 86 ] ( http://www.htslib.org/ ). The Kalahari genome was assembled chromosome by chromosome and aligned with GRCh38(hg38) as a reference assembly. The BAM format was also converted to FASTA via the aforementioned method. Human pangenome project data From the human pangenome project [ 87 ] ( https://humanpangenome.org/ ) we selected the 47 genome haplotypes (primary assemblies) data from the NCBI bioproject (PRJNA730822) in genomic FASTA format, which included the following: 04 Southern Han Chinese (SHC), 07 Puerto Rican in Puerto Rico (PUR), 03 Colombian in Medellin, Colombia (CLM), 02 Esan in Nigeria (ESN), 07 African ancestries from Barbados (Caribbean) (ACB), 04 Peruvian in Lima, Peru (PEL), 01 Vietnamese Kinh In Ho Chi Minh City, Vietnam (KHV), 08 Mandinka in Gambia - Western Division (MWG), 04 Mende in Sierra Leone (MSL), 01 Punjabi in Lahore, Pakistan (PLP), 02 Yoruban in Ibadan, Nigeria (YIN), 02 Kinyawa, Kenya (KWK), 02 Ashkenazim Son HG002 (AIS), 01 Han Chinese Son/HG005 (HCS). More information about the samples is available in Supplementary Table S6. Identification of ‘extra’ UCRs (exUCRs) For this analysis, we used the Neanderthal and human ancient genomes previously described and the pangenome genome and human genome references GRCh38(hg38).p14 and T2T-CHM13(hs1) in genomic FASTA format (T2T-CHM13(hs1) reference has complete the sequence of the Y chromosome[ 88 ] and was included to mapping the UCRs). First, each genome was converted into the BLAST database format via makeblastdb -in genome.fasta -dbtype nucl -out genomedb.DB . We subsequently performed similarity searches via the BLASTn algorithm of the 481 UCRs against the genomes via the following command line: blastn -query UCRs.fasta -db genome.DB -out ucr-genomeDB.txt -evalue 0.000001 -max_target_seqs 200 -outfmt 6 qseqid sseqid pident length mismatch gapopen qstart qend sstart send evalue bitscore qlen slen sstrand qcovs qcovhsp . We then filtered the results on the basis of size of the aligned region (at least 200 bp), percent identity (at least 75%), and query coverage (at least 80%). Finally, we consider those results located in different coordinates in relation to the reference sequence. That is, a region was considered an “extra UCR” if its coordinates were different from the original UCR in the reference genome. Literature review of UCRs and UCR overlapped genes The electronic search was performed up to June 2024 in the PubMed and Google Scholar databases. We have included original studies, reviews, and chapters written in English. In the PubMed database, the terms “uc.x” or “gene name” were searched, and the title and abstracts were screened. In Google Scholar, the terms “ultraconserved region” + “uc.x” were used. Duplicates, theses, preprints, and unrelated texts were excluded from the study. Cancer Expression Data Sources from TCGA For gene expression analysis, we obtained RNA-Seq datasets for different cancer types from the TCGA database. We selected cancer datasets containing at least 30 samples per condition (tumor and normal tissue). To obtain the gene read counts and TPM values for each cancer type, we used the TCGAbiolinks package from R [ 89 ]. For the differential expression analyses, we selected genes annotated as lncRNAs that contained at least one UCR in some of its isoforms. Differentialexpression (DE) analysis and visualization of results For differential expression analysis, we used the Wilcoxon rank-sum test, as described in [ 90 ]. For this analysis, we used the gene read counts of the 83 lncRNA-T-UCRs previously selected. As a first step, we filtered genes that had very low counts via the filterByExpr function from the egdeR package. After that, we normalized the gene counts via the trimmed mean of M values (TMM) method. We subsequently calculated each gene’s counts-per-million (CPM) values and used it as input in the wilcox.test function in R to calculate the p-value. Finally, we set a p-value cut-off on the basis of an FDR threshold using the Benjamini & Hochberg method. To visualize the differentially expressed lncRNA genes across the 10 cancer types, we constructed Vulcan plots for each of them via the ggplot2 package in R. In addition, we selected genes that were differentially expressed in at least one type of cancer. Using this list, we constructed heatmaps for each cancer type using the TPM values obtained from TCGA database. The heatmaps were constructed via the function Heatmap of the ComplexHeatmap package from R. To analyze the differentially expressed lncRNA-T-UCRs among the 10 cancer types, we constructed a bar plot via package of ggplot2 from R [ 91 ]. Cancer survival analysis from TCGA impact measurement To analyze the impact of lncRNA-T-UCRs on patients prognosis via TCGA survival data, we employed both the Cox proportional-hazards model and the Kaplan-Meier (KM) method via the survival R package [ 92 ]. TCGA survival data were downloaded via XenaBrowser ( https://xenabrowser.net/datapages/ ) and only patients used for differential expression analysis were included. For the Cox model, Bonferroni correction was applied post-hoc for multiple testing adjustments. When a significant result was found with the Cox model, the same lncRNA-T-UCR was further analyzed via the KM model, which groups patients on the basis of median expression value (high and low). Graphical visualizations were generated via the survminer R package ( https://cran.r-hub.io/web/packages/survminer/index.html ). Single-cell RNA-sequencing (scRNA-seq) analysis We analyzed publicly available scRNA-seq data to determine the transcriptional profiles of 83 lncRNA-T-UCRs mapped to UTR and annotated in TCGA. BRCA (GSE148673, GSE161529, and E-MTAB-8107), KIRC (EGAS00001002325, GSE159115, phs002065.v1.p1, nc9bc8dn4m.1), LUAD (GSE131907, GSE123904, HRA000154, and E-MTAB-6149), PRAD (GSE141445), HNSC (nasopharyngeal cancer; GSE150430), and COAD (EGAS00001003779) scRNA-seq data were downloaded from the Curated Cancer Cell Atlas [ 93 ] ( https://www.weizmann.ac.il/sites/3CA/ ) database. We selected 162 untreated primary samples belonging to the 10X platform whose count data were available. The selected datasets and samples are described in Supplementary Table S23. THCA and UCEC were excluded from this analysis because of the lack of data availability and our inclusion criteria. The count data were processed via the Seurat (v4.0.6) package in R (v4.2.2) [ 94 ]. KIRC, BRCA, and LUAD had more than one available dataset and were integrated via the IntegrateData function [ 95 ]. The identities of the cell types were based on the expression of canonical markers. In total, the RNA sequencing data from 596,400 single cells were analyzed. The heatmaps were generated and analyzed via the Morpheus online tool ( https://software.broadinstitute.org/morpheus/ ) Regulatory element annotation and enrichment sources To identify possible regulatory evidence in UCRs, we obtained transcription factors (TFs) and enhancers from the human genome (hg38) from the Open Regulatory Annotation database (ORegAnnoDB) version 3.0 [ 14 ]. These datasets are also accessible through the UCSC Genome Browser https://genome.ucsc.edu/cgi-bin/hgTrackUi?hgsid=686342163_2it3aVMQVoXWn0wuCjkNOVX39wxy&c=chr1&g=oreganno ). The Enhancers dataset is enriched with the VISTA Enhancer [ 15 ], Visel et al. (2008) and Snetkova et al. (2021) publications [ 9 , 10 ]. SNP/SNV data acquisition SNP identification and annotation were performed via dbSNP version 156 for GRCh38 VCF as a reference [ 96 ]. Tabix (HTSlib v1.13) [ 97 ] extracted variants in UCRs, resulting in a filtered VCF file. To add variant frequency information, the VCF file was annotated via Annovar (v2020-06-08) [ 98 ] with the following formats and versions provided by the tool: ABraOM (v20181204) [ 99 ], 1,000 Genomes (v20150824) [ 100 ], and gnomAD 4.0 (v20231127) [ 16 ]. Variants were filtered to retain those with a frequency equal to or greater than 1% in at least one subpopulation from the analyzed databases (ABraOM; 1,000 Genomes: African, Admixed American, East Asian, European, South Asian; gnomAD: African, Amish, Latino/Admixed American, Ashkenazi Jewish, East Asian, Finnish, Non-Finnish European, Middle Eastern, South Asian, Other, XX, XY). Additionally, to find SNPs in UCRs in ancient humans and pangenome data we aligned all the UCRs with the MUSCLE algorithm [ 101 ]. Here, we selected the common SNPs from the Super Enhancer database (SEdb) release v2.0 [ 18 ] (available at https://bio.liclab.net/sedb/download.php ) to search for SNP motifs, SNP GWASs, and SNP eQTLs. Declarations Author Contribution B.T.L.N. wrote the main manuscript text and prepared all figures and tables. L.S.O. contributed to drafting the manuscript and participated in the data analysis, data curation and validation of RNA-seq analyses. A.C.R. and J.C.O. contributed to the literature review and manuscript editing. C.M. assisted in data analysis and interpretation of survival patient prognostic. D.F.G. provided critical revisions and contributed to the drafting of the manuscript. A.H.U. helped with methodology development and data curation of SNVs and SNPs. F.P. provided oversight throughout the research project. V.L.S.C., S.S.C. and A.P.S. participated in the data analysis, data curation and validation of scRNA-seq analyses. G.A.C., R.F.C., A.R.P. and J.C.O. provided expert advice and critical feedback on the study's findings. A.R.P. and J.C.O. supervised the project, reviewed the manuscript and contributed to designing the study. All authors reviewed and approved the final manuscript. Acknowledgement The results published here are in whole or in part based upon data generated by the TCGA Research Network: https://www.cancer.gov/tcga. This study is supported by Araucaria Foundation - Fundação Araucária in NAPI Bioinformática (# PDI 66/2021) and CNPq (# 440412/2022-6). References Bejerano G, Pheasant M, Makunin I, Stephen S, Kent WJ, Mattick JS, Haussler D. Ultraconserved elements in the human genome. Science. 2004;304:1321–5. Stephen S, Pheasant M, Makunin IV, Mattick JS. Large-scale appearance of ultraconserved elements in tetrapod genomes and slowdown of the molecular clock. Mol Biol Evol. 2008;25:402–8. Carter JK, Kimball RT, Funk ER, Kane NC, Schield DR, Spellman GM, Safran RJ. Estimating phylogenies from genomes: A beginners review of commonly used genomic data in vertebrate phylogenomics. J Hered. 2023;114:1–13. Calin GA, Liu C-G, Ferracin M, et al. Ultraconserved regions encoding ncRNAs are altered in human leukemias and carcinomas. Cancer Cell. 2007;12:215–29. Mestdagh P, Fredlund E, Pattyn F, et al. An integrative genomics screen uncovers ncRNA T-UCR functions in neuroblastoma tumours. Oncogene. 2010;29:3583–92. Snetkova V, Pennacchio LA, Visel A, Dickel DE. Perfect and imperfect views of ultraconserved sequences. Nat Rev Genet. 2022;23:182–94. Prabh N, Rödelsperger C. Are orphan genes protein-coding, prediction artifacts, or non-coding RNAs? BMC Bioinformatics. 2016;17:226. de Oliveira JC. Transcribed Ultraconserved Regions: New regulators in cancer signaling and potential biomarkers. Genet Mol Biol. 2023;46:e20220125. Visel A, Prabhakar S, Akiyama JA, Shoukry M, Lewis KD, Holt A, Plajzer-Frick I, Afzal V, Rubin EM, Pennacchio LA. Ultraconservation identifies a small subset of extremely constrained developmental enhancers. Nat Genet. 2008;40:158–60. Snetkova V, Ypsilanti AR, Akiyama JA, et al. Ultraconserved enhancer function does not require perfect sequence conservation. Nat Genet. 2021;53:521–8. Fabris L, Calin GA. (2017) Chapter Four - Understanding the Genomic Ultraconservations: T-UCRs and Cancer. In: Galluzzi L, Vitale I, editors International Review of Cell and Molecular Biology. Academic Press, pp 159–172. Pereira Zambalde E, Mathias C, Rodrigues AC, de Souza Fonseca Ribeiro EM, Fiori Gradia D, Calin GA, de Carvalho J. Highlighting transcribed ultraconserved regions in human diseases. Wiley Interdiscip Rev RNA. 2020;11:e1567. Gibert MK Jr, Sarkar A, Chagari B, Cells et al. https://doi.org/10.3390/cells11101684 Lesurf R, Cotto KC, Wang G, Griffith M, Kasaian K, Jones SJM, Montgomery SB, Griffith OL, Open Regulatory Annotation Consortium. ORegAnno 3.0: a community-driven resource for curated regulatory annotation. Nucleic Acids Res. 2016;44:D126–32. Visel A, Minovitsky S, Dubchak I, Pennacchio LA. VISTA Enhancer Browser–a database of tissue-specific human enhancers. Nucleic Acids Res. 2007;35:D88–92. Chen S, Francioli LC, Goodrich JK, et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature. 2024;625:92–100. Naslavsky MS, Scliar MO, Yamamoto GL, et al. Whole-genome sequencing of 1,171 elderly admixed individuals from São Paulo, Brazil. Nat Commun. 2022;13:1004. Wang Y, Song C, Zhao J, et al. SEdb 2.0: a comprehensive super-enhancer database of human and mouse. Nucleic Acids Res. 2023;51:D280–90. Ferguson JF, Matthews GJ, Townsend RR, et al. Candidate gene association study of coronary artery calcification in chronic kidney disease: findings from the CRIC study (Chronic Renal Insufficiency Cohort). J Am Coll Cardiol. 2013;62:789–98. Diabetes Genetics Initiative of Broad Institute of Harvard and MIT, Lund University, and Novartis Institutes of BioMedical Research, Saxena R, Voight BF et al. (2007) Genome-wide association analysis identifies loci for type 2 diabetes and triglyceride levels. Science 316:1331–1336. Chung SA, Brown EE, Williams AH, et al. Lupus nephritis susceptibility loci in women with systemic lupus erythematosus. J Am Soc Nephrol. 2014;25:2859–70. Takahashi Y, Kou I, Takahashi A, et al. A genome-wide association study identifies common variants near LBX1 associated with adolescent idiopathic scoliosis. Nat Genet. 2011;43:1237–40. Miyake A, Kou I, Takahashi Y, et al. Identification of a susceptibility locus for severe adolescent idiopathic scoliosis on chromosome 17q24.3. PLoS ONE. 2013;8:e72802. Ogura Y, Kou I, Miura S, et al. A Functional SNP in BNC2 Is Associated with Adolescent Idiopathic Scoliosis. Am J Hum Genet. 2015;97:337–42. Sollis E, Mosaku A, Abid A, et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51:D977–85. Zhu Z, Guo Y, Shi H, et al. Shared genetic and experimental links between obesity-related traits and asthma subtypes in UK Biobank. J Allergy Clin Immunol. 2020;145:537–49. Koskeridis F, Evangelou E, Said S, Boyle JJ, Elliott P, Dehghan A, Tzoulaki I. Pleiotropic genetic architecture and novel loci for C-reactive protein levels. Nat Commun. 2022;13:6939. Huang J, Huffman JE, Huang Y, et al. Genomics and phenomics of body mass index reveals a complex disease network. Nat Commun. 2022;13:7973. Horikoshi M, Day FR, Akiyama M, et al. Elucidating the genetic architecture of reproductive ageing in the Japanese population. Nat Commun. 2018;9:1977. Pulit SL, Stoneman C, Morris AP, et al. Meta-analysis of genome-wide association studies for body fat distribution in 694 649 individuals of European ancestry. Hum Mol Genet. 2019;28:166–74. Saunders GRB, Wang X, Chen F, et al. Genetic diversity fuels gene discovery for tobacco and alcohol use. Nature. 2022;612:720–4. Sinnott-Armstrong N, Tanigawa Y, Amar D, et al. Genetics of 35 blood and urine biomarkers in the UK Biobank. Nat Genet. 2021;53:185–94. Sakaue S, Kanai M, Tanigawa Y, et al. A cross-population atlas of genetic associations for 220 human phenotypes. Nat Genet. 2021;53:1415–24. Sun BB, Maranville JC, Peters JE, et al. Genomic atlas of the human plasma proteome. Nature. 2018;558:73–9. Lagou V, Jiang L, Ulrich A, et al. GWAS of random glucose in 476,326 individuals provide insights into diabetes pathophysiology, complications and treatment stratification. Nat Genet. 2023;55:1448–61. Pietzner M, Wheeler E, Carrasco-Zanini J, et al. Mapping the proteo-genomic convergence of human diseases. Science. 2021;374:eabj1541. Chernus J, Roosenboom J, Ford M, et al. GWAS reveals loci associated with velopharyngeal dysfunction. Sci Rep. 2018;8:8470. Khanshour AM, Kou I, Fan Y, et al. Genome-wide meta-analysis and replication studies in multiple ethnicities identify novel adolescent idiopathic scoliosis susceptibility loci. Hum Mol Genet. 2018;27:3986–98. Kou I, Otomo N, Takeda K, et al. Genome-wide association study identifies 14 previously unreported susceptibility loci for adolescent idiopathic scoliosis in Japanese. Nat Commun. 2019;10:3685. Shen H, Lu C, Jiang Y, et al. Genetic variants in ultraconserved elements and risk of breast cancer in Chinese population. Breast Cancer Res Treat. 2011;128:855–61. Yang R, Frank B, Hemminki K, et al. SNPs in ultraconserved elements and familial breast cancer risk. Carcinogenesis. 2008;29:351–5. Ji X, Li P, Fuscoe JC, et al. A comprehensive rat transcriptome built from large scale RNA-seq-based annotation. Nucleic Acids Res. 2020;48:8320–31. Xu C, Wang Z, Liu YJ, Duan K, Guan J. Harnessing GMNP-loaded BMSC-derived EVs to target miR-3064-5p via MEG3 overexpression: Implications for diabetic osteoporosis therapy in rats. Cell Signal. 2024;118:111055. Yang W, Tu H, Tang K, Huang H, Ou S, Wu J. MiR-3064 in Epicardial Adipose-Derived Exosomes Targets Neuronatin to Regulate Adipogenic Differentiation of Epicardial Adipose Stem Cells. Front Cardiovasc Med. 2021;8:709079. Huang M, Li X, Li G. Mesenchyme homeobox 1 mediated-promotion of osteoblastic differentiation is negatively regulated by mir-3064-5p. Differentiation. 2021;120:19–27. Grosu Ș-A, Dobre M, Milanesi E, Hinescu ME. (2023) Blood-Based MicroRNAs in Psychotic Disorders-A Systematic Review. Biomedicines. https://doi.org/10.3390/biomedicines11092536 Patel RB, Bajpai AK, Thirumurugan K. Differential Expression of MicroRNAs and Predicted Drug Target in Amyotrophic Lateral Sclerosis. J Mol Neurosci. 2023;73:375–90. Luo Z, Hao S, Yuan J, Zhu K, Liu S, Zhang J, Yao L. Long non-coding RNA LINC00958 promotes colorectal cancer progression by enhancing the expression of LEM domain containing 1 via microRNA miR-3064-5p. Bioengineered. 2021;12:8100–15. Bai L, Wang H, Wang A-H, Zhang L-Y, Bai J. MicroRNA-532 and microRNA-3064 inhibit cell proliferation and invasion by acting as direct regulators of human telomerase reverse transcriptase in ovarian cancer. PLoS ONE. 2017;12:e0173912. He R, Zhang FH, Shen N. LncRNA FEZF1-AS1 enhances epithelial-mesenchymal transition (EMT) through suppressing E-cadherin and regulating WNT pathway in non-small cell lung cancer (NSCLC). Biomed Pharmacother. 2017;95:331–8. Zhang P, Ha M, Li L, Huang X, Liu C. MicroRNA-3064-5p sponged by MALAT1 suppresses angiogenesis in human hepatocellular carcinoma by targeting the FOXA1/CD24/Src pathway. FASEB J. 2020;34:66–81. Xiao E, Zhang D, Zhan W, Yin H, Ma L, Wei J, Kang Y, Mao Z. circNFIX facilitates hepatocellular carcinoma progression by targeting miR-3064-5p/HMGA2 to enhance glutaminolysis. Am J Transl Res. 2021;13:8697–710. Wei M, Chen Y, Du W. LncRNA LINC00858 enhances cervical cancer cell growth through miR-3064-5p/ VMA21 axis. Cancer Biomark. 2021;32:479–89. Shih C-H, Chuang L-L, Tsai M-H, Chen L-H, Chuang EY, Lu T-P, Lai L-C. Hypoxia-Induced MALAT1 Promotes the Proliferation and Migration of Breast Cancer Cells by Sponging MiR-3064-5p. Front Oncol. 2021;11:658151. Wang S, Ping M, Song B, Guo Y, Li Y, Jia J. Exosomal CircPRRX1 Enhances Doxorubicin Resistance in Gastric Cancer by Regulating MiR-3064-5p/PTPN14 Signaling. Yonsei Med J. 2020;61:750–61. Yan J, Jia Y, Chen H, Chen W, Zhou X. Long non-coding RNA PXN-AS1 suppresses pancreatic cancer progression by acting as a competing endogenous RNA of miR-3064 to upregulate PIP4K2B expression. J Exp Clin Cancer Res. 2019;38:390. Khalilian S, Mohajer Z, Khazeei Tabari MA, Ghobadinezhad F, Ghafouri-Fard S. circGFRA1: A circular RNA with important roles in human carcinogenesis. Pathol Res Pract. 2023;248:154588. Meng M, Wu Y-C. (2022) LMX1B Activated Circular RNA GFRA1 Modulates the Tumorigenic Properties and Immune Escape of Prostate Cancer. J Immunol Res 2022:7375879. Ji X, Lv C, Huang J, Dong W, Sun W, Zhang H. ALKBH5-induced circular RNA NRIP1 promotes glycolysis in thyroid cancer cells by targeting PKM2. Cancer Sci. 2023;114:2318–34. Brown GR, Hem V, Katz KS, et al. Gene: a gene-centered information resource at NCBI. Nucleic Acids Res. 2015;43:D36–42. Shi C, Sun L, Song Y. FEZF1-AS1: a novel vital oncogenic lncRNA in multiple human malignancies. Biosci Rep. 2019. https://doi.org/10.1042/BSR20191202 . Hu C, Liu K, Wang B, Xu W, Lin Y, Yuan C. DLX6-AS1: An Indispensable Cancer-related Long Non-coding RNA. Curr Pharm Des. 2021;27:1211–8. Khalafiyan A, Emadi-Baygi M, Wolfien M, Salehzadeh-Yazdi A, Nikpour P. Construction of a three-component regulatory network of transcribed ultraconserved regions for the identification of prognostic biomarkers in gastric cancer. J Cell Biochem. 2023;124:396–408. Zhou Y, Xu S, Xia H, Gao Z, Huang R, Tang E, Jiang X. Long noncoding RNA FEZF1-AS1 in human cancers. Clin Chim Acta. 2019;497:20–6. Yang FL, Hong K, Zhao GJ, Liu C, Song YM, Ma LL. [Construction of prognostic model and identification of prognostic biomarkers based on the expression of long non-coding RNA in bladder cancer via bioinformatics]. Beijing Da Xue Xue Bao. 2019;51:615–22. Tocco C, Bertacchi M, Studer M. Structural and Functional Aspects of the Neurodevelopmental Gene NR2F1: From Animal Models to Human Pathology. Front Mol Neurosci. 2021;14:767965. Hu J, Peng F, Qiu X, Yang J, Li J, Shen C, Yuan C. NR2F1-AS1: A Functional Long Noncoding RNA in Tumorigenesis. Curr Med Chem. 2023;30:4266–76. Ghafouri-Fard S, Khoshbakht T, Hussen BM, Baniahmad A, Taheri M, Samsami M. A review on the role of NR2F1-AS1 in the development of cancer. Pathol Res Pract. 2022;240:154210. Luo D, Liu Y, Yuan S, Bi X, Yang Y, Zhu H, Li Z, Ji L, Yu X. The emerging role of NR2F1-AS1 in the tumorigenesis and progression of human cancer. Pathol Res Pract. 2022;235:153938. Tao L, Huang G, Song H, Chen Y, Chen L. Cancer associated fibroblasts: An essential role in the tumor microenvironment. Oncol Lett. 2017;14:2611–20. Gomes FG, Nedel F, Alves AM, Nör JE, Tarquinio SBC. Tumor angiogenesis and lymphangiogenesis: tumor/endothelial crosstalk and cellular/microenvironmental signaling mechanisms. Life Sci. 2013;92:101–7. Jiang Z, Zhou J, Li L, Liao S, He J, Zhou S, Zhou Y. Pericytes in the tumor microenvironment. Cancer Lett. 2023;556:216074. Dreier MR, Walia J, de la Serna IL. (2024) Targeting SWI/SNF Complexes in Cancer: Pharmacological Approaches and Implications. Epigenomes. https://doi.org/10.3390/epigenomes8010007 Wang B, Guo H, Yu H, Chen Y, Xu H, Zhao G. The Role of the Transcription Factor EGR1 in Cancer. Front Oncol. 2021;11:642547. Habic A, Mattick JS, Calin GA, Krese R, Konc J, Kunej T. Genetic Variations of Ultraconserved Elements in the Human Genome. OMICS. 2019;23:549–59. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841–2. Frankish A, Carbonell-Sala S, Diekhans M, et al. GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res. 2023;51:D942–9. Kozomara A, Birgaoanu M, Griffiths-Jones S. miRBase: from microRNA sequences to function. Nucleic Acids Res. 2019;47:D155–62. Fromm B, Høye E, Domanska D, et al. MirGeneDB 2.1: toward a complete sampling of all major animal phyla. Nucleic Acids Res. 2022;50:D204–10. Da Fonseca BHR, Domingues DS, Paschoal AR. mirtronDB: a mirtron knowledge base. Bioinformatics. 2019;35:3873–4. Prüfer K, Racimo F, Patterson N, et al. The complete genome sequence of a Neanderthal from the Altai Mountains. Nature. 2014;505:43–9. Mafessoni F, Grote S, de Filippo C, et al. A high-coverage Neandertal genome from Chagyrskaya Cave. Proc Natl Acad Sci U S A. 2020;117:15132–6. Prüfer K, de Filippo C, Grote S, et al. A high-coverage Neandertal genome from Vindija Cave in Croatia. Science. 2017;358:655–8. Meyer M, Kircher M, Gansauge M-T, et al. A high-coverage genome sequence from an archaic Denisovan individual. Science. 2012;338:222–6. Schuster SC, Miller W, Ratan A, et al. Complete Khoisan and Bantu genomes from southern Africa. Nature. 2010;463:943–7. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, 1000 Genome Project Data Processing Subgroup. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078–9. Liao W-W, Asri M, Ebler J, et al. A draft human pangenome reference. Nature. 2023;617:312–24. Rhie A, Nurk S, Cechova M, et al. The complete sequence of a human Y chromosome. Nature. 2023;621:344–54. Colaprico A, Silva TC, Olsen C, et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016;44:e71. Li Y, Ge X, Peng F, Li W, Li JJ. Exaggerated false positives by popular differential expression methods when analyzing human population samples. Genome Biol. 2022;23:79. Wickham H. ggplot2: Elegant Graphics for Data Analysis. Springer; 2016. Borgan Ø, Patricia M, grambsch. Springer-Verlag, New York, 2000. No. Of pages: Xiii + 350. Price: $ 69.95. ISBN 0‐387‐98784‐3. Stat Med 20:2053–2054. Gavish A, Tyler M, Greenwald AC, et al. Hallmarks of transcriptional intratumour heterogeneity across a thousand tumours. Nature. 2023;618:598–606. Satija R, Farrell JA, Gennert D, Schier AF, Regev A. Spatial reconstruction of single-cell gene expression data. Nat Biotechnol. 2015;33:495–502. Stuart T, Butler A, Hoffman P, Hafemeister C, Papalexi E, Mauck WM 3rd, Hao Y, Stoeckius M, Smibert P, Satija R. Comprehensive Integration of Single-Cell Data. Cell. 2019;177:1888–e190221. Sherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 2001;29:308–11. Bonfield JK, Marshall J, Danecek P, Li H, Ohan V, Whitwham A, Keane T, Davies RM. (2021) HTSlib: C library for reading/writing high-throughput sequencing data. Gigascience. https://doi.org/10.1093/gigascience/giab007 Wang K, Li M, Hakonarson H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 2010;38:e164. Naslavsky MS, Yamamoto GL, de Almeida TF, et al. Exomic variants of an elderly cohort of Brazilians in the ABraOM database. Hum Mutat. 2017;38:751–63. 1000 Genomes Project Consortium, Auton A, Brooks LD, et al. A global reference for human genetic variation. Nature. 2015;526:68–74. Edgar RC. MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics. 2004;5:113. Additional Declarations No competing interests reported. Supplementary Files SupplementaryFigsS1S2S3.pdf SupplementaryTablesS1S23.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4837600","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":345280951,"identity":"b342a980-bc2c-4505-abea-f2693dc3e4af","order_by":0,"name":"Bruno Thiago de Lima Nichio","email":"","orcid":"","institution":"Federal University of Technology - Paraná - UTFPR","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Bruno","middleName":"Thiago de Lima","lastName":"Nichio","suffix":""},{"id":345280952,"identity":"aca8ba5b-9da4-47b7-ad4d-3f23df864ae6","order_by":1,"name":"Liliane Santana Oliveira","email":"","orcid":"","institution":"Federal University of Technology - Paraná - UTFPR","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Liliane","middleName":"Santana","lastName":"Oliveira","suffix":""},{"id":345280953,"identity":"f9cc9a1c-5e7c-4afa-9ac9-ce8820b6e0b0","order_by":2,"name":"Ana Carolina Rodrigues","email":"","orcid":"","institution":"Federal University of Parana","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ana","middleName":"Carolina","lastName":"Rodrigues","suffix":""},{"id":345280954,"identity":"abed9da5-bd82-458d-9acc-a59bb701f120","order_by":3,"name":"Carolina Mathias","email":"","orcid":"","institution":"Federal University of Parana","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Carolina","middleName":"","lastName":"Mathias","suffix":""},{"id":345280955,"identity":"cf70d65e-7e7e-4505-ba11-86c15557fd33","order_by":4,"name":"Daniela Fiori Gradia","email":"","orcid":"","institution":"Federal University of Parana","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Daniela","middleName":"Fiori","lastName":"Gradia","suffix":""},{"id":345280956,"identity":"b621b9fd-a6ff-4353-b180-0009e80c71c7","order_by":5,"name":"Alysson Henrique Urbanski","email":"","orcid":"","institution":"Carlos Chagas Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Alysson","middleName":"Henrique","lastName":"Urbanski","suffix":""},{"id":345280957,"identity":"106eb202-e416-47c1-b2ed-deaabbaa35d3","order_by":6,"name":"Fabio Passetti","email":"","orcid":"","institution":"Carlos Chagas Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Fabio","middleName":"","lastName":"Passetti","suffix":""},{"id":345280958,"identity":"24e8f1eb-25ba-4066-897e-4a6c1668ef50","order_by":7,"name":"Victória Larissa Schimidt Camargo","email":"","orcid":"","institution":"São Paulo State University (UNESP)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Victória","middleName":"Larissa Schimidt","lastName":"Camargo","suffix":""},{"id":345280959,"identity":"71517e42-3037-4895-b7ea-423a96441c51","order_by":8,"name":"Sarah Santiloni Cury","email":"","orcid":"","institution":"São Paulo State University (UNESP)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Sarah","middleName":"Santiloni","lastName":"Cury","suffix":""},{"id":345280960,"identity":"ba9d7efc-d0b2-4769-be82-6b6e59ec70a2","order_by":9,"name":"Amanda Piveta Schnepper","email":"","orcid":"","institution":"São Paulo State University (UNESP)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Amanda","middleName":"Piveta","lastName":"Schnepper","suffix":""},{"id":345280961,"identity":"f6e6b643-16b6-4034-844f-ae8cf9da73be","order_by":10,"name":"Robson Francisco Carvalho","email":"","orcid":"","institution":"São Paulo State University (UNESP)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Robson","middleName":"Francisco","lastName":"Carvalho","suffix":""},{"id":345280962,"identity":"65837859-747f-47e1-8129-478230559813","order_by":11,"name":"George A. Calin","email":"","orcid":"","institution":"The University of Texas MD Anderson Cancer Center","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"George","middleName":"A.","lastName":"Calin","suffix":""},{"id":345280963,"identity":"9eee52e6-10f3-495e-bfa1-73559659f517","order_by":12,"name":"Jaqueline Carvalho Oliveira","email":"","orcid":"","institution":"Federal University of Parana","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jaqueline","middleName":"Carvalho","lastName":"Oliveira","suffix":""},{"id":345280964,"identity":"90f5898a-828c-4219-a450-fc6a887879ec","order_by":13,"name":"Alexandre Rossi Paschoal","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5ElEQVRIiWNgGAWjYJACZiCW4QcSH8DcAwxsRGnhkWxgZpxBmhaDA8RqMWc//vBzQYUdj/GN/IPNhXsY5PhuJLA9rsCjxbInx1h6xplkHrMbyYzNM54xGEveSGA3PINHi8GBHDZm3jZmkBb2xzwHGBI3AG2RbMCn5fzzZ0At9TzGM4C2ALXUE9ZyI8EMqOUwj4EEREuCASEtljPeGEvznDnOI3HmsWHzjAMShjPPPGw3xKfFnD/94Weeimo5/vbEh80FB2zk+Y4nH3uI12HIHGAESQApRnwaMLWMglEwCkbBKMAEAISISnbCCwlqAAAAAElFTkSuQmCC","orcid":"","institution":"The Rosalind Franklin Institute","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Alexandre","middleName":"Rossi","lastName":"Paschoal","suffix":""}],"badges":[],"createdAt":"2024-07-31 18:48:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4837600/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4837600/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":63470058,"identity":"b8edb077-25e7-4886-9dac-51a7c0de994b","added_by":"auto","created_at":"2024-08-28 13:02:55","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1842290,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenomic and transcript annotation of UCRs and the classification of the lncRNA transcripts that overlap with UCRs\u003c/strong\u003e. (\u003cstrong\u003eA.\u003c/strong\u003e) Pie chart showing the genomic annotation of ultraconserved sequences in the human genome proposed by Bejerano et al. (2004), Snetkova et al. (2022) and our work (GENCODE v46), which classifies the position of UCRs in humans (\u003cstrong\u003eB.\u003c/strong\u003e) Revision of 481 UCR genomic annotations for mouse (GRCmm39.mm39 from GENCODE vM35) and rat (mRatBN7.2 rn7 from RefSeq 106). (\u003cstrong\u003eC\u003c/strong\u003e.) The bar chart shows the distribution of UCR genomic classification according to Mestdagh et al. (2010) mapped to the human genome from GENCODE v46, mouse GENCODE vM35 and rat RefSeq rn7. (\u003cstrong\u003eD.\u003c/strong\u003e) Transcripts annotation revealed that 383 T-UCRs overlapped with 2,303 transcripts in human (GENCODE v46), 360 T-UCRs overlapped with 1,125 transcripts in mouse (GENCODE vM35), and 224 T-UCRs overlapped with 935 transcripts in rat (RefSeq rn7). (\u003cstrong\u003eE.\u003c/strong\u003e) From 1,945 protein-coding transcripts annotated in human (GENCODE v46), 22 contained UCRs in the coding sequence (CDS) region, 148 overlapped UCR in the CDS, and 1,775 contained UCRs outside the CDS (UCRs were identified in intronic or exonic non-CDS regions). For mouse annotation (GENCODE vM35), from 1,022 protein-coding transcripts, 14 contained UCRs in the CDS region, 117 overlapped UCRs in CDS, and 891 contained UCRs outside the CDS. For rat annotation (RefSeq rn7), from 826 protein-coding transcripts, two contained UCRs in CDS, 54 overlapped with UCRs in the CDS, and 770 contained UCRs outside the CDS.\u003c/p\u003e","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/1cf72199bc90d61216731227.png"},{"id":63470634,"identity":"ce5eeb20-ea3a-4b37-b81a-23f52389762e","added_by":"auto","created_at":"2024-08-28 13:10:55","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":891235,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eHomo sapiens \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003eand \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eHomo neanderthalensis\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003efossil lineages presented in this work and the ‘extra’ UCRs (exUCRs) identified and annotated on extra chromosome regions\u003c/strong\u003e. (\u003cstrong\u003eA\u003c/strong\u003e). The ten ancient humans (in the gray chart), the five Neanderthal fossils (in the orange chart), and their coexistence in kilo annus (ka) linneage. (\u003cstrong\u003eB\u003c/strong\u003e). Scatter plot of exUCRs identified in ancient human and Neanderthal genomes according to UCR coverage and UCR alignment length. (\u003cstrong\u003eC\u003c/strong\u003e). The UpSet plot of the four exUCR (euc.129, euc.175, euc.397, and euc.425) intersections in the \u003cem\u003eHomo sapiens \u003c/em\u003eand \u003cem\u003eHomo neanderthalensis \u003c/em\u003egroups and the genomic annotation via GENCODE v46 for human genome. The circlize diagram expresses these four exUCR chromosome regions and the UCR, which are derived.\u003c/p\u003e","description":"","filename":"Fig2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/f9609012a7a9707d5aaa3237.jpg"},{"id":63470060,"identity":"effb0927-aba9-4f3c-a0ce-1d80f78c1b6e","added_by":"auto","created_at":"2024-08-28 13:02:56","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1207100,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ePancancer of lncRNA-T-UCRs differentially expressed with the support of RNA-seq and single-cell data analysis. \u003c/strong\u003e(A) LncRNA-T-UCR was differentially expressed in ten TCGA cancer types. The graph shows that 35 lncRNA-T-UCRs were differentially expressed in at least one of the cancer types. The bars above the x-axis represent upregulated lncRNAs, and the bars below the x-axis represent downregulated lncRNA-T-UCRs. (B) Single-cell expression landscape of lncRNAs across the seven tumor types. Heatmap of the average expression of 35 lncRNAs in seven tumor types. Both rows (genes) and columns (cell types) were clustered using Euclidean distance. Heatmaps were generated and analyzed via the Morpheus online tool (hMps://soPware.broadins0tute.org/morpheus/). NK: natural killer cells; DC: dendritic cells; BRCA: breast cancer; COAD: colorectal cancer; HNSCC: head and neck cancer (nasopharyngeal); KIRC: kidney cancer; LUAD: lung adenocarcinoma; PRAD: prostate cancer.\u003c/p\u003e","description":"","filename":"Fig3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/8f5e8857f8ac9b0442e196c2.jpg"},{"id":63470061,"identity":"f13c28e1-c0ef-4236-aec2-8ad44f41cc37","added_by":"auto","created_at":"2024-08-28 13:02:56","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":5940499,"visible":true,"origin":"","legend":"\u003cp\u003e(\u003cstrong\u003eA\u003c/strong\u003e) \u003cstrong\u003ePanorama of theSNPs identified in the UCR sequences\u003c/strong\u003e.\u003cstrong\u003e \u003c/strong\u003eThe SNPs in UCRs with a frequency exceeding 1% in the populations identified in the gnomAD (v.3.1), 1Kgenomes, and ABraOM databases. These SNPs were observed across diverse ethnicities, including African, American, European, South, East Asian, and Brazilian populations. (B) Representation of SNPs in ultraconserved sequences in protein-coding and non-coding regions (lncRNA and intergenic regions) for ancient humans, Neanderthal, and human pangenomes. The frequencies are shown on the world map: the upper triangle was represented the rs2682406 identified in uc.133 (rs2682406/uc.133), a \u003cem\u003eRSRC1 \u003c/em\u003eprotein-coding gene intron variant and the lower triangle represents the rs1861100 identified in uc.53 (rs1861100/uc.53), a \u003cem\u003eLINC1122 \u003c/em\u003elong non-coding gene intron variant.\u003c/p\u003e","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/6c5dca4207709c42712442a3.png"},{"id":63472105,"identity":"21bdea7c-7c23-4cd6-a6f3-fc91e84b4703","added_by":"auto","created_at":"2024-08-28 13:27:04","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9885079,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/5ebd5a7c-92f4-4781-a29f-170011218ea3.pdf"},{"id":63470064,"identity":"ea5d69e0-9952-4f2b-b065-8b18f76dc0fe","added_by":"auto","created_at":"2024-08-28 13:02:56","extension":"pdf","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":3969088,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigsS1S2S3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/1051eb633bc0793496508c73.pdf"},{"id":63470063,"identity":"3f075131-b4c7-48e4-adcc-f970fab09c83","added_by":"auto","created_at":"2024-08-28 13:02:56","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":2264518,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTablesS1S23.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4837600/v1/b0162713b048387c50cc19e1.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"The Functional Map of Ultraconserved Regions in Humans, Mice and Rats","fulltext":[{"header":"1. INTRODUCTION","content":"\u003cp\u003eUltraconserved regions (UCRs) are defined as 481 DNA segments longer than 200 bp, with 100% sequence identity in their alignment with humans, mice, and rats [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. These regions are also highly conserved across disparate taxa, including vertebrates, insects, worms, yeast, and other species, and are essential DNA markers for phylogenomic analyses [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. Approximately 93% of the UCRs were transcribed in at least one human tissue tested, with nearly 79% showing ubiquitous transcription across a wide range of normal human tissues. Because of this characteristic, these regions are called transcribed UCRs (T-UCRs) [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe first UCR annotation proposed by Bejerano et al. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] focused on their overlap with protein-coding genomic regions, with surprisingly only 23.0% of the UCRs overlapping with known mature mRNA sequences [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. They identified 256 sequences with no evidence of transcription from any matching expressed sequence tag or mRNA, and 111 overlapped the mRNAs of known human protein-coding genes. Additionally, 114 sequences were \"inconclusive,\" where the evidence for transcription was inconclusive. In 2010, Mestdagh and colleagues reorganized all UCR sequences into six categories based on protein-coding information, classifying those regions as intergenic (38.7%), intronic (42.6%), exonic (4.2%), partly exonic (5.0%), and exon-containing (5.6%). Additionally, 3.9% were classified as \u0026ldquo;multiple\u0026rdquo; when the genomic annotation varied because of the host gene splice variants [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Recently, in 2022, Snetkova and collaborators observed that 371 (77.1%) were non-coding UCRs and 110 (22.9%) were UCRs that overlapped the exons of protein-coding genes. In this study they reviewed the biological importance of these UCRs, mainly in regulating gene expression during embryogenesis [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eCoding sequences are typically highly conserved and are subject to more stringent negative selection pressures [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Nevertheless, more than 75% of the UCRs were in intergenic or intronic, indicating that negative selection pressure can also act on these non-coding sequences[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Additionally, among the intronic transcribed UCRs (T-UCRs), almost 58% were detected in the antisense orientation compared with the host gene [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. These results suggest that many UCRs may be transcribed as long non-coding RNAs (lncRNAs). However, a general annotation of the 481 ultraconserved regions with overlapping lncRNAs was readily unavailable.\u003c/p\u003e \u003cp\u003eIn addition to their potential roles in lncRNA function, UCRs are crucial enhancers during development. For example, Visel and colleagues reported that half of the non-exonic UCRs evaluated drove reproducible reporter gene expression in various tissues of developing mouse embryos [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], and this function was reinforced by others [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Additionally, highlighting the role of some T-UCRs, deregulated expression has been associated with several pathological conditions, including tumor types and Cancer disease [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan additionalcitationids=\"CR12\" citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. Although the UCRs play an essential role in human health and disease; there is limited information regarding specific T-UCRs, the investigation of annotated lncRNA-T-UCRs largely improves our knowledge of the potential roles of these molecules. Therefore, even two decades after their initial identification, many UCRs remain to be explored. To address this, we constructed a comprehensive functional map of 481 UCRs. however,\u003c/p\u003e \u003cp\u003eIn this study, we investigated four significant contributions: (i) we analysed enrichment mining annotation data, including comparisons of the modern genome version and the uncharted ancestral genomes in 15 Neanderthal and ancient humans from different eras; (ii) examination of pan-cancer expression with the support of RNA-Seq and single-cell data analysis, focusing on the T-UCRs in lncRNA genes; (iii) evidence of UCRs in the regulatory elements, including enhancers and microRNA/transcription factor binding sites (TFBS); and (iv) population single-variation analysis, shedding light on their evolutionary dynamics and implications for changes in binding sites, expression, and disease susceptibility. Our research explored the omics landscape of UCRs, shedding light on non-coding regions and deepening our insights into UCR functionality, gene regulation, and disease mechanisms, underscoring their pivotal roles in genomic stability and disease pathogenesis.\u003c/p\u003e"},{"header":"2. RESULTS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. UCR Annotation Enrichment Mining Insights\u003c/h2\u003e \u003cp\u003eWe analyzed genomic and transcriptomic annotation data using the currently human reference and we compared this to previous UCR annotations, performing the same analysis for recent mouse and rat genomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, Supplementary Tables S1-S5). Adopting the genomic classification of Bejerano et al. [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e], we observed that 126 (26.2%) and 355 (73.8%) UCRs were classified as coding and non-coding, respectively, in the human genome (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). In the mouse and rat genomes, 133 (27.6%) and 348 (72.4%) sequences were classified as coding, respectively, while 82 (17.0%) and 399 (83.0%) were classified as non-coding, respectively (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eUsing the classification of Mestdag et al. [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], for human, we identified 261 (54.5%) UCRs as intronic, four as exonic (0.8%), four as exon-containing (0.8%), five as partly exonic (1.0%), 98 as intergenic (20.4%), and 109 as multiple (22.7%) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC; Supplementary Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). In mouse, 223 (44.9%) UCRs were mapped to intronic, 13 (2.7%) to exonic, 10 (2.1%) to exon-containing, 21 (4.3%) to partly exonic, 93 (19.3%) to multiple, and 121 (25.3%) to intergenic regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC and Supplementary Tables S2 and S4). With respect to the rat genome, our analysis revealed 161 (33.5%) UCRs to an intron, 12 (2.5%) exonic, 14 (2.9%) exon-containing, 16 (3.3%) partly exonic, 28 (5.8%) multiple, and 250 (51.9%) intergenic regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC; Supplementary Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo map UCR evidence at transcriptomic levels, we investigate the UCR sequences onto human, mouse and rat transcript annotations. In human, 384 UCRs overlapped with 2,303 transcripts (GENCODE v46). Among these, 1,945 transcripts were mapped to protein-coding T-UCRs (205 protein-coding genes overlapped with 315 T-UCRs), 355 to lncRNA-T-UCR (85 lncRNA genes overlapped with 100 T-UCRs), and three other transcripts were mapped to a pseudogene and two small non-coding RNAs (one miRNA and one miscellaneous RNA (abbrev. miscRNA)), each containing one T-UCR (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD). Among the 1,945 protein-coding transcripts annotated, 22 contained UCRs in the coding sequence (CDS) region, 148 overlapped with UCRs in the CDS, and 1,775 contained UCRs outside CDS (UCRs were identified in intronic or exonic non-CDS regions) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE; Supplementary Table S3).\u003c/p\u003e \u003cp\u003eIn mouse (GENCODE vM35), 360 UCRs overlapped with 1,125 transcripts, of which 1,022 were protein-coding transcripts (205 protein-coding genes and four protein-coding genes to be experimentally confirmed), 101 were long non-coding transcripts (associated with 46 lncRNA genes), and two were other non-coding RNA (one miRNA and another miscRNA) (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD). Of the 1,022 protein-coding transcripts, 14 contained UCRs into the CDS region, 117 overlapped with UCRs in the CDS, and 891 contained UCRs outside the CDS (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE; Supplementary Table S4).\u003c/p\u003e \u003cp\u003eIn rat (ENSEMBLE 111), we identified 224 T-UCRs overlapping 935 transcripts, with 826 protein-coding transcripts (118 protein-coding genes), 50 lncRNA transcripts (associated with 17 lncRNA genes), 58 miscRNA transcripts (58 genes), and one miRNA (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD and Supplementary Table S5). Among the protein-coding transcripts, two contained UCRs in the CDS, 54 overlapped with UCRs in the CDS, and 770 contained UCRs outside the CDS (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE). This unique miRNA gene is associated with the uc.420/miR-3064 gene in the three species.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Conservation of UCRs in Ancient Human Genomes\u003c/h2\u003e \u003cp\u003eWe also investigated the presence of UCRs in five \u003cem\u003eHomo neanderthalensis\u003c/em\u003e genomes (one male and four females) and ten \u003cem\u003eHomo sapiens\u003c/em\u003e genomes (nine males and one female) from various periods, including the Palaeolithic era to more recent times, and present-day tribal populations, such as those from the Khoisan Bantu people from Kalahari desert (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA) (Supplementary Table S6).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe obtained all of 481 UCRs with 100% coverage and identity in these genomas, although some had regions with sequences containing few bases represented as \u0026lsquo;N\u0026rsquo; without the exact nucleotide defined (Supplementary Table S7). Furthermore, both modern and ancient human genomes contained an additional set of novel UCRs with high conservation levels (sequences of at least 200 base pairs, minimum alignment coverage of 75%, and identity threshold of 80% for the UCR set) in distinct chromosomal regions, except in the Mezmaiskaya genome (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB and Supplementary Table S8). We found four novel UCRs align perfectly in the human, mouse, and rat genomes and follow Bejerano\u0026rsquo;s rules for UCR definition (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC; Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Table S9), and we denominated them \u0026lsquo;extra\u0026rsquo; UCRs (exUCRs as abbreviations and \u0026ldquo;euc.\u0026rdquo; for the code sequence). In particular, we observed the presence of euc.397 (chr18) in the analyzed genomes. We identified euc.129 (chr13) and euc.425 (chr16) in all genomes, except for ZKU and Bantu, respectively. Finally, euc.175 (chr10) was present in these genomes, except in the ZKU and Scandinavian genomes(Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC).\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eExUCR vs UCR coordinates, and overlapping gennomic annotations for human, mouse and rat.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003egenome\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eexUCR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eexUCR coordinates\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eexUCR gene overlapping\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eexUCR gene overlapping class\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eUCR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eUCR coordinates\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eUCR gene overlapping\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eUCR gene overlapping class\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ehuman\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr13:97356685\u0026ndash;97356888\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eMBNL2\u003c/em\u003e (exonic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr3:152446598\u0026ndash;152446809\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eMBNL1\u003c/em\u003e (exonic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003emouse\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr14:120633013\u0026ndash;120633216\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eMBNL2\u003c/em\u003e (exonic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr3:60521998\u0026ndash;60522209\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eMBNL1\u003c/em\u003e (exonic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003erat\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr15:97508498\u0026ndash;97508701\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEFCAB12-ARSJ (intergenic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNA (intergenic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.129\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr2:144798866\u0026ndash;144799077\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eMBNL1\u003c/em\u003e (exonic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ehuman\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.175\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr10:129893328\u0026ndash;129893555\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eEBF3\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.175\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr5:158914830\u0026ndash;158915079\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eEBF1\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003emouse\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.175\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr7:136852289 136852516\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eEBF3\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.175\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr11:44691125\u0026ndash;44691374\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eEBF1\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003erat\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.175\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr1:192052056 192052283\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eEBF3\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.175\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr10:29284049\u0026ndash;29284298\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eEBF1\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ehuman\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.397\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr18:25285204\u0026ndash;25285497\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eZNF521\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.397\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr16:49701937\u0026ndash;49702247\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eZNF423\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003emouse\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.397\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr18:14029595\u0026ndash;14029888\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eZFP521\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.397\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr8:88563296\u0026ndash;88563607\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eZFP423\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003erat\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.397\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr18:4997387\u0026ndash;4997680\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eZFP521\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.397\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr19:19231570\u0026ndash;19231880\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eZFP423\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ehuman\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.425\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr16:49701982\u0026ndash;49702237\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eZNF423\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.425\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr18:25285243\u0026ndash;25285567\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eZNF521\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003emouse\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.425\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr8:88563341\u0026ndash;88563596\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eZNF423\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.425\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr18:14029634\u0026ndash;14029958\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eZNF521\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003erat\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeuc.425\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003echr19:19231580\u0026ndash;19231835\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cem\u003eSYNPOL2L-MYO5B\u003c/em\u003e (intergenic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eNA (intergenic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003euc.425\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003echr18:4997426\u0026ndash;4997750\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e\u003cem\u003eZNF521\u003c/em\u003e (intronic)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003eprotein-coding\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eOwing to the lack of comprehensive annotations for the Neanderthal and ancient human genomes, we utilized the human GENCODE v46 reference annotation to describe the exUCRs and compared these annotations with those from mouse and rat. All exUCRs overlapped with protein-coding genes in human and mouse annotations, whereas euc.129 and euc.425 were intergenic in rat (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Supplementary Tables S10 and S11). Euc.129 is located between \u003cem\u003eEFCAB12\u003c/em\u003e gene (upstream) and \u003cem\u003eARSJ\u003c/em\u003e gene (downstream), whereas euc.425 is located between the \u003cem\u003eSYNPOL2L\u003c/em\u003e gene (upstream) and the \u003cem\u003eMYO5B\u003c/em\u003e gene (downstream). Interestingly, all extra UCRs were located in protein-coding genes from the same family from the related UCR, even maintaining overlap within the region (exonic or intronic) of the gene (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). This finding suggests that the ultraconserved sequence is essential for the structure of these proteins or for transcript regulation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. The T-UCR in LncRNAs: Insights in Pan-Cancer Analyses\u003c/h2\u003e \u003cp\u003eThere is limited information regarding T-UCRs (Supplementary Table S12). The investigation of annotated lncRNA-T-UCRs improves our knowledge of the potential functional roles in specific pathologic conditions, such as Cancer disease. We performed differential expression analysis of lncRNA-T-UCRs via RNA-seq data (bulk and single-cell) from public cancer databases, including The Cancer Genome Atlas Program (TCGA) and the Curated Cancer Cell Atlas (3CA).\u003c/p\u003e \u003cp\u003e \u003cb\u003eDifferential expression of lncRNA-T-UCRs in cancer\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe selected datasets from the TCGA database using sample size as the only criterion (at least 30 samples per condition, tumor, and non-tumoral). Ten cancer datasets were obtained: breast invasive carcinoma (BRCA), colon adenocarcinoma (COAD), head and neck squamous cell carcinoma (HNSC), kidney renal clear cell carcinoma (KIRC), kidney renal papillary cell carcinoma (KIRP), lung adenocarcinoma (LUAD), lung squamous cell carcinoma (LUSC), prostate adenocarcinoma (PRAD), thyroid carcinoma (THCA), and uterine corpus endometrial carcinoma (UCEC). The number of non-tumoral and tumoral samples for each cancer is detailed in Supplementary Table S13.\u003c/p\u003e \u003cp\u003eThis analysis aimed to detect lncRNAs containing at least one T-UCR transcript isoform. Thus, to identify differentially expressed lncRNA genes, we used the 85 lncRNAs identified in our transcript mapping analysis (i.e., 85 lncRNAs mapped to 100 T-UCRs). However, the RNA-Seq data did not include data for two lncRNAs (ENSG00000288692 and ENSG00000289413). Therefore, we performed tumor and single-cell differential expression analyses of 83 lncRNAs.\u003c/p\u003e \u003cp\u003eUsing the Wilcoxon rank-sum test to calculate differentially expressed genes and filtering the results by p-values (\u003cb\u003e\u0026le;\u003c/b\u003e\u0026thinsp;0.05), we detected differentially expressed (DE) lncRNAs in each cancer type (Supplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e and Supplementary Table S14). Considering the threshold of +\u0026thinsp;2 log2foldChange, the number of genes upregulated in cancers ranges from three (HNSC and PRAD) to 11 (LUSC). With respect to the downregulated lncRNAs (log\u003csub\u003e2\u003c/sub\u003efoldchange \u0026le; -2), we identified only one lncRNA-T-UCR in LUAD, THCA, and UCEC cancer types and four in KIRC. BRCA and PRAD cancers did not present downregulated lncRNA-T-UCR (Supplementary Table S14).\u003c/p\u003e \u003cp\u003eWhen we compared the expression of lncRNA genes across all tumor types, we identified 35 lncRNAs that were differentially expressed in at least one cancer type (Supplementary Table S15). Using TPM values, we constructed a heatmap using these 35 DE lncRNA-T-UCRs for each cancer type (Supplementary Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). None of the lncRNA-T-UCRs were differentially expressed in any cancer type. However, we detected two genes, \u003cem\u003eFEZF1-AS1\u003c/em\u003e/uc.232 and \u003cem\u003eDLX6-AS1\u003c/em\u003e/uc.220/uc.221, which were upregulated in the seven cancer types (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA). The lncRNA \u003cem\u003eFEZF1-AS1\u003c/em\u003e/uc.232 (ENSG00000230316.7) was upregulated in BRCA, COAD, HSNC, LUAD, LUSC, PRAD, and UCEC (fold-change ranging from 2.3 in UCEC to 8,82 in COAD); whereas \u003cem\u003eDLX6-AS1\u003c/em\u003e/uc.220/uc.221 was upregulated in BRCA, COAD, HSNC, KIRC, LUAD, LUSC, and UCEC (fold-change ranging from 1.07 in PRAD to 6,7 in LUSC).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe did not identify any downregulated lncRNA-T-UCRs across all cancer types. However, we observed downregulated three lncRNAs in three cancers: \u003cem\u003eMECOM-AS1\u003c/em\u003e/uc.136 (HNSC, KIRC, KIRP), \u003cem\u003eMEF2C-AS1\u003c/em\u003e/uc.167 (COAD, LUSC, and UCEC), and ENSG00000257277.2/uc.70, an lncRNA without a gene symbol defined and antisense to the coding \u003cem\u003eARHGAP15\u003c/em\u003e gene (differentially expressed in COAD, LUAD, LUSC). Furthermore, four lncRNAs were up-regulated in one or two cancer types and downregulated in another (\u003cem\u003eSAMMSON\u003c/em\u003e/uc.116, \u003cem\u003eHOXB-AS3\u003c/em\u003e/uc.415/uc.416/uc.417, \u003cem\u003eLINC01505\u003c/em\u003e/uc.266, and \u003cem\u003eERICH2-DT\u003c/em\u003e/uc.95). Finally, regarding lncRNAs expressed exclusively in one type of cancer, 10 were upregulated and seven were downregulated (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003cb\u003eThe impact of lncRNA-T-UCRs on patient prognosis\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo analyze the impact of lncRNA-T-UCRs on patient prognosis via TCGA survival data, we investigated the expression of all 83 human lncRNA-T-UCRs in the 10 TCGA cancer types selected, and 19 lncRNAs were associated with survival in the Cox proportional-hazard analysis, seven of which also exhibited significance in the Kaplan-Meier model (Supplementary Figure S3 and Supplementary Table S16).\u003c/p\u003e \u003cp\u003eFocusing on lncRNA-T-UCRs that were significantly associated with survival in both analyses, we identified LINC00583/uc.250, FEZF1-AS1/uc.233, FOXG1-AS1/uc.361, DPH6-DT/uc.382, PKN2-AS1/uc.31, AQP4-AS1/uc.427, and ENSG00000270087.5/uc.286. Interestingly, the lncRNA-UCR analysis revealed a cancer-specific clinical impact, since none were shared among more than one cancer type. This result emphasizes the tissue specificity of these molecules, which is a biologically prominent characteristic of these molecules, highlighting their use as prognostic markers.\u003c/p\u003e \u003cp\u003e \u003cb\u003elncRNA-T-UCR single-cell RNA-seq (scRNA-seq) data analysis\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo identify the 83 lncRNA-T-UCRs expressed within the tumor microenvironment, we obtained publicly available scRNA-seq data from the Curated Cancer Cell Atlas database. Among the lncRNA-T-UCRs, 35 and 23 were differentially expressed according to single-cell analysis (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB). We discovered that NR2F1-AS1/uc.169 was the most highly expressed lncRNA in the fibroblasts of all the tumor types analyzed. NR2F1-AS1/uc.169 was also expressed in pericytes from BRCA and COAD patients and in malignant cells from LUAD and COAD patients (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB).\u003c/p\u003e \u003cp\u003eAdditionally, IQCH-AS1/uc.389 was highly expressed in epithelial cells from HNSC and mast cells from BRCA and LUAD. MEF2C-AS1/uc.167 was highly expressed in B cells from HNSC, BRCA, and COAD. Specifically, in LUAD, FEZF1-AS1/uc.232 was highly expressed in malignant cells, HOXB-AS3/uc.415 to uc.417 was highly expressed in fibroblasts, and AQP4-AS1/uc.427 was highly expressed in epithelial cells. LINC01117/uc.109 was highly expressed in the stromal and malignant cells of BRCA and COAD. LINC01505/uc.266 was highly expressed in NK cells of KIRC, whereas SAMMSON/uc.116 was expressed in the B cells of KIRC. SATB1-AS1/uc.113 was highly expressed in immune cells, and PKN2-AS1/uc.31 was highly expressed in malignant and epithelial cells in HNSC (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. UCRs in Regulatory Elements\u003c/h2\u003e \u003cp\u003eWe performed enrichment mining for the presence of regulatory elements within UCRs via the ORegAnno tool [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. We identified regulatory elements in 373 UCRs (77.5%), comprising 234 (62.7%) instances as regulatory regions (RR) (including promoters, enhancers, etc.), 73 (19.6%) as transcription factor binding sites (TFBS), 64 (17.1%) as combinations of TFBS and RR, one as a TFBS and miRNA binding site (SMARCA4 and miR-429, respectively), and one as a miRNA binding site (miR-214-3p) in the human genome (Supplementary Table S17). Additionally, 308 UCRs were 100% inserted into RRs and TFBS. Among the 45 distinct TFBSs, SMARCA4 was the binding site most representative of 62 UCRs, followed by EGR1, FOXA1, CTCF, and CEBPB, which contained 20, 13, 12, and 11 UCRs, respectively. Despite SMARCA4, 51 UCRs were completely inserted into the TFBS sequence, and another 11 UCRs partially overlapped. Among the UCRs related to EGR1, FOXA1, CTFC, and CEBPB TFs, 18, 5, 5, and 4 UCRs, respectively, were contained in these TFBSs.\u003c/p\u003e \u003cp\u003eThe 19 UCRs exhibited agglomerative numbers of TFBS: uc.213 and uc.246 related to six TFBS; uc.88, uc.329, uc.374, and uc.392 with five TFs; uc.95, uc.162, uc.393, and uc.417 with four TFBS; and uc.74, uc.189, uc.232, uc.277, uc.398, uc.409, uc.412, uc.416, and uc.456 with three TFBS. Finally, in 26 UCRs, we identified two TFBS associated with the same UCR, and a single TFBS in the remaining 93 UCRs.\u003c/p\u003e \u003cp\u003eAccording to the VISTA Enhancer database, 295 of the 373 (79.1%) regulatory regions within UCRs were identified as human enhancers [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. For four UCRs (uc.29, uc.347, uc.350 and uc.355) we report the presence of two simultaneous enhancers, totaling 299 enhancers. Among these enhancers, almost 80% (240/299) were experimentally validated and linked with the respective UCR, according to previous studies [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. In addition, 152 and 147 enhancers were classified as positively and negatively regulating transcription, respectively (Supplementary Table S17).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5. Personalized Genomics: Pan-SNP Data Analysis\u003c/h2\u003e \u003cp\u003e \u003cb\u003eSNP-UCR compendium identified in human databases and pangenomes\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe availability of complete human genomes in various databases allowed us to examine the presence of single nucleotide polymorphisms (SNPs) in UCR sequences and investigate their potential functional effects. In our study, we identified 353 SNPs (with at least 1% frequency) that mapped to 223 UCRs in the Genome Aggregation Database (gnomAD) [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], 1000 Genomes Project (1Kgenomes) and Archive of Brazilian Mutations (ABraOM) [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e] (see Supplementary Table S18).\u003c/p\u003e \u003cp\u003eWe found that 17 (4.8%) of these SNPs are shared across all three databases (gnomAD, 1kgenomes, and ABraOM), 301 (85.3%) were shared between gnomAD and 1k genomes but were not present in ABraOM, 33 (9.3%) are unique to gnomAD, and two (0.6%) are unique to 1kgenomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA). The predominant SNP types were single nucleotide variants (SNVs) (332/353), followed by insertion or deletion markers (INDELs) (17/353), deletions (3/353), and insertions (1/353). Most of the variants were located in introns of the coding gene, including 146 (41%) SNPs. Additionally, 13 SNPs were found in exonic-coding regions (of them, six were synonymous and seven were nonsynonymous SNVs), seven in the 3\u0026rsquo;UTR, three in the 5\u0026rsquo;UTR of protein-coding genes, and six in exonic and 58 in intronic regions of lncRNA genes. A total of 120 SNPs were identified in the intergenic regions.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eConcerning the ethnicity identified in these databases, 36 SNPs were present in African, European, American, East, and South Asian populations, whereas 63, 48, 3, 19, and 20 SNPs were unique to African, European, American, East, and South Asian populations, respectively. In addition, 164 SNPs were distributed among at least two ethnicities, and 17 SNPs were shared with the Brazilian population (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eA).\u003c/p\u003e \u003cp\u003e \u003cb\u003eUCR-SNPs in regulatory regions associated with diseases\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo understand the effects of some SNPs on the regulation of nearby genes and those associated with diseases, we assessed whether UCR-SNPs mapped to a sequence of a transcription factor binding motif (called a motif SNP), which SNPs affects gene expression (called eQTL SNPs), are associated with a complex disease or trait through a Genome-Wide Association Study (GWAS SNPs) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Using the Super Enhancer database (SEdb v2.0), we identified 48 SNPs with potential regulatory effects in 43 UCRs (Supplementary Tables S19 and S20). 21 UCRs containing exclusively motif SNPs; 12 UCRs with eQTL and motif SNPs; 11 with eQTL SNPs; one with motif, GWAS, and eQTL SNPs; one containing eQTL and GWAS SNPs; and two UCRs containing motif and GWAS SNPs (Supplementary Table S19 and S20). For UCRs with GWAS SNPs, we identified that rs12981/uc.268 (\u003cem\u003eRC3H2\u003c/em\u003e gene, 3\u0026rsquo;UTR variant) was associated with coronary artery calcification [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], rs17291131/uc.211 (\u003cem\u003eSKAP2\u003c/em\u003e gene, intron variant) was associated with type 2 diabetes [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e], rs2056117/uc.140 (intergenic variant) was associated with systemic lupus erythematosus [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e], and rs11190870/uc.302 (intergenic variant) was associated with scoliosis [\u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn addition, we identified eight UCRs with SNPs in more than 1% of the global population that were correlated with disease/trait according to the GWAS catalogue [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] (Supplementary Table S21). rs1861100/uc.53 contains an intron variant in LINC01122 and was related to body mass index (obesity) [\u003cspan additionalcitationids=\"CR27\" citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]). Additionally, rs142252570/uc.380, an intergenic variant related to the LINC02304/LINC02325 genes, was associated with menarche (age at onset) [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. The SNP rs142252570/uc.380, an intergenic variant related with the RSRC1 gene, was associated with body mass index and drinks per week [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. The SNP rs73174306/uc.136, an intron variant in MECOM and MECOM-AS1 genes, was related to blood glucose levels and pancreatic hormone levels [\u003cspan additionalcitationids=\"CR33 CR34 CR35\" citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. The SNP rs1538101/uc.252, an intron variant in the BNC2 gene, was related to velopharyngeal dysfunction [\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. The SNP rs11190870/uc.302, an intergenic variant reported with the LINC01514/LBX1 genes, was related to scoliosis [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e, \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The SNP rs182041872/uc.315, an intergenic variant reported with the ACADSB/HMX3 genes, was related to antidepressant treatment resistance [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eInvestigation of SNPs in UCR sequences for ancient humans, Neanderthals, and human pangenomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB and Supplementary Table S22) revealed six altered SNPs with a prevalence of \u0026gt;\u0026thinsp;50%. Among these SNPs, five were located in intronic regions correlated with GWAS/SEdb with impact disease. Rs2682406/uc.133 (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB in the upper triangle) and rs1861100/uc.53 (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eB in the lower triangle) were related to breast cancer and obesity, respectively [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e, \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Rs1538101/uc.252, mapped to the protein-coding gene \u003cem\u003eBNC2\u003c/em\u003e, was associated with velopharyngeal dysfunction and rs11896224/uc.82, an intron variant of the lncRNA LINC01876 was also associated with velopharyngeal dysfunction. Finally, two SNVs were found to be intergenic: rs11190870/uc.302 (associated with scoliosis) and rs7092999/uc.295 with no disease association.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. DISCUSSION","content":"\u003cp\u003eSome regions related to the 481 UCRs initially described by Bejerano and colleagues in 2004 [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e] have been described as important in cellular processes and are associated with diseases [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. However, despite two decades since the first description, the role of most regions has yet to be discovered. Previous UCR annotations have focused predominantly on DNA coding to provide functional evidence. In this study, we innovatively enriched the annotation of UCR data at both the genomic and transcriptomic levels by incorporating the latest human, mouse, and rat genome versions. Our analysis reinforces the presence of most UCRs in the non-coding and outside CDS regions. Despite this, the transcripts associated with these UCRs are mostly annotated in the intronic regions of protein-coding transcripts, with a quarter annotated as lncRNA genes on the basis of biotype annotation. To the best of our knowledge, this is the first study to focus on UCRs mapped to lncRNAs, providing accessible and organized information such as gene symbols and IDs.\u003c/p\u003e \u003cp\u003eComparing human, mouse, and rat annotations is an interesting way to observe human and mouse similarities. For example, the percentages of intergenic UCRs in human and mouse is approximately 20 and 25%, respectively; in contrast, 52% of UCRs were found in intergenic positions in rat. This difference was also verified by examining the transcripts that included UCR. The observed differences may result from the limited information available for the rat genome because the rat transcript data are significantly less complete than those available for the mouse, particularly in terms of transcribed and exon regions [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWith respect to transcript analysis, a unique miRNA gene, \u003cem\u003eMIR3064\u003c/em\u003e, was wholly contained in the uc.420 genomic region of the three organisms. MiR-3064 is related to diverse diseases [\u003cspan additionalcitationids=\"CR44 CR45 CR46\" citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e]. In cancer, miR-3064 is targeted by several lncRNAs/circRNAs associated with important pathways such as PI3K/Akt [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e], hTERT [\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e], glycolysis [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e], and Src, and plays crucial tumor suppressive roles [\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e, \u003cspan additionalcitationids=\"CR52 CR53 CR54 CR55 CR56 CR57 CR58\" citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eGiven the extreme conservation of UCRs in mammals, they are expected to be preserved in the ancestral hominids. However, this has not been previously confirmed. Considering the analysis of the 481 UCRs within ancestral genomes, including Neanderthal and ancient human genomes from different eras, as expected, UCR sequences in hominid genomes were also 100% conserved. Notably, we identified four novel UCRs in humans (modern and ancestral genomes), mouse, and rat that had not been previously described. Interestingly, these four novel UCRs (exUCRs) are protein-coding genes from the same family as the UCR related genes, suggesting a potential motif associated with the conserved sequence in these molecules. euc.129 is a novel UCR related to uc.129, which maps to the MBNL1 gene. In contrast, euc.129 is located in the MBNL2 gene, which plays a role in MBNL1. Similarly, the new UCR exuc.175 overlaps the \u003cem\u003eEBF3\u003c/em\u003e gene while uc.175 the \u003cem\u003eEBF1\u003c/em\u003e; exuc.397 overlaps the \u003cem\u003eZNF521\u003c/em\u003e gene while the uc.397 the \u003cem\u003eZNF423\u003c/em\u003e; and the exuc.425 overlaps the \u003cem\u003eZNF423\u003c/em\u003e gene, while uc.425 the \u003cem\u003eZNF521\u003c/em\u003e. MBNL and ZNF are both zinc finger proteins, and EBF are DNA-binding proteins, all of which act as transcription factors [\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn human, our approach enabled us to associate 85 lncRNA-T-UCRs with previously unrelated functions, demonstrating its potential for analysis in public databases through Ensemble codes and gene symbols. This is the first study to analyze the lncRNA-T-UCR panel in a pancancer context. These lncRNAs play essential roles in cancer [\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e, \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e]; however, they are not previously recognized as containing T-UCRs. For example, the uc.232 sequence is almost entirely in exon 2 of FEZF1-AS1-202 and in the opposite strand in exon 1 of FEZF1-203 (with no protein associated with the retained intron). A unique study focused on uc.232 showed that the T-UCR was differentially expressed in TCGA gastric cancer samples compared with their non-tumoral counterparts, but no expression of this T-UCR was observed when a different cohort of GC samples was utilized [\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e]. Here, we reinforced the importance of FEZF1-AS1 in multiple tumor types and highlighted the relationship between this lncRNA and uc.232.\u003c/p\u003e \u003cp\u003eIn the survival analysis, we identified seven lncRNA-T-UCRs in cancer cohorts: FEZF1-AS1/uc.233, LINC00583/uc.250, FOXG1-AS1/uc.361, DPH6-DT/uc.382, PKN2-AS1/uc.31, AQP4-AS1/uc.427, and ENSG00000270087.5/uc.286. Two of these genes were previously associated with prognosis in cancer patients, including the widespread deregulation of FEZF1-AS1, which is associated with worse patient prognosis in 13 cancer types [\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e]. Additionally, PKN2-AS1 was previously identified as being related to the prognosis of patients with bladder cancer [\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e]. The expression of the other five lncRNAs is associated with survival time for the first time herein.\u003c/p\u003e \u003cp\u003eTo the best of our knowledge, this is the first study to analyze T-UCR expression via single-cell data. Data from 35 lncRNA-T-UCRs were obtained, and NR2F1-AS1/uc.169 was the most highly expressed lncRNA in the fibroblasts of all the tumor types analyzed. NR2F1-AS1/uc.169 was also expressed in pericytes from BRCA and COAD patients and in malignant cells from LUAD and COAD patients. The uc.169 sequence overlapped with the coding gene \u003cem\u003eNR2F1\u003c/em\u003e and its antisense transcript NR2F1-AS1. \u003cem\u003eNR2F1\u003c/em\u003e encodes a nuclear hormone receptor that functions as a transcription factor. This protein is critical in the developmental process of the brain[\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e]. In contrast, antisense transcripts have mostly been studied in the context of cancer [\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e]. NR2F1-AS1 plays an oncogenic role in several tumor types and is associated with proliferation, migration, and drug resistance [\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e]. In colorectal and cervical cancers, this lncRNA is downregulated and associated with a tumor-suppressive role[\u003cspan additionalcitationids=\"CR68\" citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSingle-cell analysis of NR2F1-AS1 expression in different cell types is underexplored. This lncRNA is usually highly expressed in cancer samples, but in our study, this lncRNA was explicitly found in fibroblast cells for all tumor types analyzed. Cancer-associated fibroblasts (CAFs) are part of the tumor stroma. Although not malignant, fibroblasts may play a role in cancer progression, such as supporting tumor growth, epithelial-mesenchymal transition/metastasis, and therapy resistance [\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e]. Fibroblasts may also be associated with angiogenesis, as the interaction between tumors and stromal cells can increase vascular endothelial growth factor (VEGF) [\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e]. Furthermore, NR2F1-AS1/uc.169 was expressed in pericytes in BRCA and COAD. Pericytes are fibroblast-like cells that wrap around endothelial cells in arterioles, capillaries, and venules and are strongly associated with angiogenesis [\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e]. Our findings underscore the importance of studying lncRNA-T-UCRs in the context of single cells, suggesting more specific functions than previously known.\u003c/p\u003e \u003cp\u003eBy organizing data on regulatory elements within the UCRs of the human genome, we found evidence of regulatory elements in 373 UCRs (77.5% of all UCRs). These regulatory elements include enhancers, transcription factors, and miRNA binding sites. The two most abundant TFs in the UCR sequences are SMARCA4 and EGR1. Both TFs are essential for the regulation of critical networks under both physiological and pathological conditions [\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e, \u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e74\u003c/span\u003e], are deregulated during cancer development, and may deregulate the regulation of transcripts related to UCRs.\u003c/p\u003e \u003cp\u003eFinally, we investigated frequent SNPs in UCR sequences, including thousands (~\u0026thinsp;128.700) of human whole-genome sequences. Previously, a broad analysis and utilizing ultraconserved regions that include comparisons with other species revealed that some polymorphisms have also been identified, however, the vast majority of these polymorphisms are rare events [\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e]. In the present study, we focused on 481 UCRs and in the most frequent SNPs (\u0026gt;\u0026thinsp;1%), identifying 353 SNPs 25 of which were associated with eQTL effects. For example, rs34384113/uc.109 is in the intron portion of LINC01117 with an eQTL effect in HOXD10 and LINC01117 expression; rs117486161/uc.220, in the intron of the \u003cem\u003eDLX6-AS1\u003c/em\u003e gene has an eQTL effect on DLX6-AS1 expression; and rs56805315/uc.417 mapped at the intron of \u003cem\u003eHOXB-AS3\u003c/em\u003e and 5\u0026rsquo;UTR of the HOXB6 gene has an eQTL effect on AC091133.1, CDK5RAP3, and HOXB6 expression. Interestingly, the lncRNAs LINC01117/uc.109, DLX6-AS1/uc.220, and HOXB-AS3/uc.417 were dysregulated in cancer, as highlighted in our expression analysis.\u003c/p\u003e \u003cp\u003eOur in-depth and comprehensive functional map revealed that almost all UCRs provide tips for understanding their functions. These complete maps of UCRs provide new insights into these regions, accelerating the understanding of functional mechanisms and their potential use as biomarkers of these related regions.\u003c/p\u003e"},{"header":"4. METHODS","content":"\u003cp\u003e \u003cb\u003eData sources\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eUCR data acquisition\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe human UCR coordinates were obtained from Bejerano et al. (2004). The sequences of the UCRs were extracted from the human genome via coordinates (obtained from GENCODE release 45 to GRCh38 (hg38).p14)). We used bedtools[\u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e] version 2.31.0 with the getfasta algorithm. General command: \u003cem\u003ebedtools getfasta [OPTIONS] -fi\u0026thinsp;\u0026lt;\u0026thinsp;input FASTA\u0026gt; -bed\u0026thinsp;\u0026lt;\u0026thinsp;BED/GFF/VCF\u0026gt;\u003c/em\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eThe genomic annotation data\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe human genomic annotation (GFF/GTF formats) was obtained from GENCODE[\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e] release 45 to the GRCh38(hg38).p14 genome (available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gencodegenes.org/human/\u003c/span\u003e\u003cspan address=\"https://www.gencodegenes.org/human/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). For mouse genomic annotation, we used the GENCODE release M34 for the GRcm39/mm39 genome (available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gencodegenes.org/mouse\u003c/span\u003e\u003cspan address=\"https://www.gencodegenes.org/mouse\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Finally, the rat genomic annotation was obtained from RefSeq 106 (ENSEMBLE release 111) for the mRatBN7.2 rn7.2 genome (available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/datasets/genome/GCF_015227675.2\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/datasets/genome/GCF_015227675.2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eUCR genome remapping\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFor human, we remapped the 481 UCRs for NCBI37(hg16) genomic coordinates (BED format) via the liftOver tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://genome.ucsc.edu/cgi-bin/hgLiftOver\u003c/span\u003e\u003cspan address=\"https://genome.ucsc.edu/cgi-bin/hgLiftOver\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e to human GRCh37(hg19), mouse NCBI34(mm6) and rat Baylor 3.1(rn3). The coordinates were converted from GRCh37(hg19) to GRCh38(hg38).p14 for reference to the human genome, NCBI34(mm6) to GRCm39/mm39 for mouse and Baylor 3.1(rn3) to mRatBN7.2 rn7 for rat via the NCBI-remapping tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/genome/tools/remap\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/genome/tools/remap\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eUCR overlapping functions\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe UCR coordinates (BED format) were overlapped with respective genomic annotation releases in GFF/GTF (GENCODE and RefSeq annotations) or BED formats (mirRNA, regulatory, and SNP/SNVs annotations) and extracted via the R (v4.3.2) script built with the foverlaps function from package data.table (v1.14.2) available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.rdocumentation.org/packages/data.table/versions/1.14.2\u003c/span\u003e\u003cspan address=\"https://www.rdocumentation.org/packages/data.table/versions/1.14.2\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cb\u003eIdentification of UCRs in the miRNA annotation data\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo detect which UCRs overlapped with miRNAs and mirtrons, we obtained these features from miRBase release 22.1 [\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e], MirGeneDB 2.1 [\u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e79\u003c/span\u003e], and MirtronDB 1.1 [\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e80\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cb\u003eNeanderthal and ancient human genomes\u003c/b\u003e \u003c/p\u003e \u003cp\u003eNeanderthal genomes were obtained from direct genome projects in BAM format: Altai (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.eva.mpg.de/genetics/genome-projects/neandertal/\u003c/span\u003e\u003cspan address=\"https://www.eva.mpg.de/genetics/genome-projects/neandertal/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e [\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e81\u003c/span\u003e], Chagyrskaya 8 [\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e82\u003c/span\u003e] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.eva.mpg.de/genetics/genome-projects/chagyrskaya-neandertal/\u003c/span\u003e\u003cspan address=\"https://www.eva.mpg.de/genetics/genome-projects/chagyrskaya-neandertal/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e, Mezmaskaya 3 (PRJNA765125), Vindija 19.33 [\u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e83\u003c/span\u003e] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://cdna.eva.mpg.de/neandertal/Vindija/bam/Pruefer_etal_2017/Vindija33.19/\u003c/span\u003e\u003cspan address=\"http://cdna.eva.mpg.de/neandertal/Vindija/bam/Pruefer_etal_2017/Vindija33.19/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e, and Denisovan 8 [\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e84\u003c/span\u003e] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.eva.mpg.de/genetics/genome-projects/denisova/\u003c/span\u003e\u003cspan address=\"https://www.eva.mpg.de/genetics/genome-projects/denisova/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e The ancient humans in the ENA database: Ust'-Ishim (PRJEB6622), Yana (PRJEB29700), AHUR_2064 (PRJEB29074), Sunghir Burial 3 (SIII) (PRJEB22592), Zlat\u0026yacute; kůň (ZKU) (PRJEB39040), Sumidouro 5 (PRJEB29074) Scandinavian hunter-gatherers (PRJEB32786), \u0026Ouml;tzi (Tyrolean Iceman) (PRJEB56570), and Ayayema A460 (PRJEB29074). The Kalahari genome was obtained from RAW data (FASTq, paired-end sequences) from the work of Schuster et al. 2010 work [\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e85\u003c/span\u003e]. More information about the samples is provided in Supplementary Table S6.\u003c/p\u003e \u003cp\u003e \u003cb\u003eKalahari genome assembly\u003c/b\u003e \u003c/p\u003e \u003cp\u003eAll genomic FASTAs were obtained in the BAM format via the consensus algorithm (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.htslib.org/doc/samtools-consensus.html\u003c/span\u003e\u003cspan address=\"http://www.htslib.org/doc/samtools-consensus.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e from SAMtools [\u003cspan citationid=\"CR86\" class=\"CitationRef\"\u003e86\u003c/span\u003e] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.htslib.org/\u003c/span\u003e\u003cspan address=\"http://www.htslib.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e The Kalahari genome was assembled chromosome by chromosome and aligned with GRCh38(hg38) as a reference assembly. The BAM format was also converted to FASTA via the aforementioned method.\u003c/p\u003e \u003cp\u003e \u003cb\u003eHuman pangenome project data\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFrom the human pangenome project [\u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e87\u003c/span\u003e] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://humanpangenome.org/\u003c/span\u003e\u003cspan address=\"https://humanpangenome.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e we selected the 47 genome haplotypes (primary assemblies) data from the NCBI bioproject (PRJNA730822) in genomic FASTA format, which included the following: 04 Southern Han Chinese (SHC), 07 Puerto Rican in Puerto Rico (PUR), 03 Colombian in Medellin, Colombia (CLM), 02 Esan in Nigeria (ESN), 07 African ancestries from Barbados (Caribbean) (ACB), 04 Peruvian in Lima, Peru (PEL), 01 Vietnamese Kinh In Ho Chi Minh City, Vietnam (KHV), 08 Mandinka in Gambia - Western Division (MWG), 04 Mende in Sierra Leone (MSL), 01 Punjabi in Lahore, Pakistan (PLP), 02 Yoruban in Ibadan, Nigeria (YIN), 02 Kinyawa, Kenya (KWK), 02 Ashkenazim Son HG002 (AIS), 01 Han Chinese Son/HG005 (HCS). More information about the samples is available in Supplementary Table S6.\u003c/p\u003e \u003cp\u003e \u003cb\u003eIdentification of \u0026lsquo;extra\u0026rsquo; UCRs (exUCRs)\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFor this analysis, we used the Neanderthal and human ancient genomes previously described and the pangenome genome and human genome references GRCh38(hg38).p14 and T2T-CHM13(hs1) in genomic FASTA format (T2T-CHM13(hs1) reference has complete the sequence of the Y chromosome[\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e88\u003c/span\u003e] and was included to mapping the UCRs). First, each genome was converted into the BLAST database format via \u003cem\u003emakeblastdb -in genome.fasta -dbtype nucl -out genomedb.DB\u003c/em\u003e. We subsequently performed similarity searches via the BLASTn algorithm of the 481 UCRs against the genomes via the following command line: \u003cem\u003eblastn -query UCRs.fasta -db genome.DB -out ucr-genomeDB.txt -evalue 0.000001 -max_target_seqs 200 -outfmt 6 qseqid sseqid pident length mismatch gapopen qstart qend sstart send evalue bitscore qlen slen sstrand qcovs qcovhsp\u003c/em\u003e.\u003c/p\u003e \u003cp\u003eWe then filtered the results on the basis of size of the aligned region (at least 200 bp), percent identity (at least 75%), and query coverage (at least 80%). Finally, we consider those results located in different coordinates in relation to the reference sequence. That is, a region was considered an \u0026ldquo;extra UCR\u0026rdquo; if its coordinates were different from the original UCR in the reference genome.\u003c/p\u003e \u003cp\u003e \u003cb\u003eLiterature review of UCRs and UCR overlapped genes\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe electronic search was performed up to June 2024 in the PubMed and Google Scholar databases. We have included original studies, reviews, and chapters written in English. In the PubMed database, the terms \u0026ldquo;uc.x\u0026rdquo; or \u0026ldquo;gene name\u0026rdquo; were searched, and the title and abstracts were screened. In Google Scholar, the terms \u0026ldquo;ultraconserved region\u0026rdquo; + \u0026ldquo;uc.x\u0026rdquo; were used. Duplicates, theses, preprints, and unrelated texts were excluded from the study.\u003c/p\u003e \u003cp\u003e \u003cb\u003eCancer Expression Data Sources from TCGA\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFor gene expression analysis, we obtained RNA-Seq datasets for different cancer types from the TCGA database. We selected cancer datasets containing at least 30 samples per condition (tumor and normal tissue). To obtain the gene read counts and TPM values for each cancer type, we used the TCGAbiolinks package from R [\u003cspan citationid=\"CR89\" class=\"CitationRef\"\u003e89\u003c/span\u003e]. For the differential expression analyses, we selected genes annotated as lncRNAs that contained at least one UCR in some of its isoforms.\u003c/p\u003e \u003cp\u003e \u003cb\u003eDifferentialexpression (DE) analysis and visualization of results\u003c/b\u003e \u003c/p\u003e \u003cp\u003eFor differential expression analysis, we used the Wilcoxon rank-sum test, as described in [\u003cspan citationid=\"CR90\" class=\"CitationRef\"\u003e90\u003c/span\u003e]. For this analysis, we used the gene read counts of the 83 lncRNA-T-UCRs previously selected. As a first step, we filtered genes that had very low counts via the filterByExpr function from the egdeR package. After that, we normalized the gene counts via the trimmed mean of M values (TMM) method. We subsequently calculated each gene\u0026rsquo;s counts-per-million (CPM) values and used it as input in the wilcox.test function in R to calculate the p-value. Finally, we set a p-value cut-off on the basis of an FDR threshold using the Benjamini \u0026amp; Hochberg method. To visualize the differentially expressed lncRNA genes across the 10 cancer types, we constructed Vulcan plots for each of them via the ggplot2 package in R. In addition, we selected genes that were differentially expressed in at least one type of cancer. Using this list, we constructed heatmaps for each cancer type using the TPM values obtained from TCGA database. The heatmaps were constructed via the function Heatmap of the ComplexHeatmap package from R. To analyze the differentially expressed lncRNA-T-UCRs among the 10 cancer types, we constructed a bar plot via package of ggplot2 from R [\u003cspan citationid=\"CR91\" class=\"CitationRef\"\u003e91\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cb\u003eCancer survival analysis from TCGA impact measurement\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo analyze the impact of lncRNA-T-UCRs on patients prognosis via TCGA survival data, we employed both the Cox proportional-hazards model and the Kaplan-Meier (KM) method via the survival R package [\u003cspan citationid=\"CR92\" class=\"CitationRef\"\u003e92\u003c/span\u003e]. TCGA survival data were downloaded via XenaBrowser (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://xenabrowser.net/datapages/\u003c/span\u003e\u003cspan address=\"https://xenabrowser.net/datapages/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and only patients used for differential expression analysis were included. For the Cox model, Bonferroni correction was applied post-hoc for multiple testing adjustments. When a significant result was found with the Cox model, the same lncRNA-T-UCR was further analyzed via the KM model, which groups patients on the basis of median expression value (high and low). Graphical visualizations were generated via the survminer R package (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://cran.r-hub.io/web/packages/survminer/index.html\u003c/span\u003e\u003cspan address=\"https://cran.r-hub.io/web/packages/survminer/index.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003cb\u003eSingle-cell RNA-sequencing (scRNA-seq) analysis\u003c/b\u003e \u003c/p\u003e \u003cp\u003eWe analyzed publicly available scRNA-seq data to determine the transcriptional profiles of 83 lncRNA-T-UCRs mapped to UTR and annotated in TCGA. BRCA (GSE148673, GSE161529, and E-MTAB-8107), KIRC (EGAS00001002325, GSE159115, phs002065.v1.p1, nc9bc8dn4m.1), LUAD (GSE131907, GSE123904, HRA000154, and E-MTAB-6149), PRAD (GSE141445), HNSC (nasopharyngeal cancer; GSE150430), and COAD (EGAS00001003779) scRNA-seq data were downloaded from the Curated Cancer Cell Atlas [\u003cspan citationid=\"CR93\" class=\"CitationRef\"\u003e93\u003c/span\u003e] (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.weizmann.ac.il/sites/3CA/\u003c/span\u003e\u003cspan address=\"https://www.weizmann.ac.il/sites/3CA/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e database. We selected 162 untreated primary samples belonging to the 10X platform whose count data were available. The selected datasets and samples are described in Supplementary Table S23. THCA and UCEC were excluded from this analysis because of the lack of data availability and our inclusion criteria. The count data were processed via the Seurat (v4.0.6) package in R (v4.2.2) [\u003cspan citationid=\"CR94\" class=\"CitationRef\"\u003e94\u003c/span\u003e]. KIRC, BRCA, and LUAD had more than one available dataset and were integrated via the \u003cem\u003eIntegrateData\u003c/em\u003e function [\u003cspan citationid=\"CR95\" class=\"CitationRef\"\u003e95\u003c/span\u003e]. The identities of the cell types were based on the expression of canonical markers. In total, the RNA sequencing data from 596,400 single cells were analyzed. The heatmaps were generated and analyzed via the Morpheus online tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://software.broadinstitute.org/morpheus/\u003c/span\u003e\u003cspan address=\"https://software.broadinstitute.org/morpheus/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e)\u003c/span\u003e\u003c/p\u003e \u003cp\u003e \u003cb\u003eRegulatory element annotation and enrichment sources\u003c/b\u003e \u003c/p\u003e \u003cp\u003eTo identify possible regulatory evidence in UCRs, we obtained transcription factors (TFs) and enhancers from the human genome (hg38) from the Open Regulatory Annotation database (ORegAnnoDB) version 3.0 [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. These datasets are also accessible through the UCSC Genome Browser \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://genome.ucsc.edu/cgi-bin/hgTrackUi?hgsid=686342163_2it3aVMQVoXWn0wuCjkNOVX39wxy\u0026amp;c=chr1\u0026amp;g=oreganno\u003c/span\u003e\u003cspan address=\"https://genome.ucsc.edu/cgi-bin/hgTrackUi?hgsid=686342163_2it3aVMQVoXWn0wuCjkNOVX39wxy\u0026amp;c=chr1\u0026amp;g=oreganno\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003cspan type=\"Underline\" class=\"Underline\" name=\"Emphasis\"\u003e).\u003c/span\u003e The Enhancers dataset is enriched with the VISTA Enhancer [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], Visel et al. (2008) and Snetkova et al. (2021) publications [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003cb\u003eSNP/SNV data acquisition\u003c/b\u003e \u003c/p\u003e \u003cp\u003eSNP identification and annotation were performed via dbSNP version 156 for GRCh38 VCF as a reference [\u003cspan citationid=\"CR96\" class=\"CitationRef\"\u003e96\u003c/span\u003e]. Tabix (HTSlib v1.13) [\u003cspan citationid=\"CR97\" class=\"CitationRef\"\u003e97\u003c/span\u003e] extracted variants in UCRs, resulting in a filtered VCF file. To add variant frequency information, the VCF file was annotated via Annovar (v2020-06-08) [\u003cspan citationid=\"CR98\" class=\"CitationRef\"\u003e98\u003c/span\u003e] with the following formats and versions provided by the tool: ABraOM (v20181204) [\u003cspan citationid=\"CR99\" class=\"CitationRef\"\u003e99\u003c/span\u003e], 1,000 Genomes (v20150824) [\u003cspan citationid=\"CR100\" class=\"CitationRef\"\u003e100\u003c/span\u003e], and gnomAD 4.0 (v20231127) [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Variants were filtered to retain those with a frequency equal to or greater than 1% in at least one subpopulation from the analyzed databases (ABraOM; 1,000 Genomes: African, Admixed American, East Asian, European, South Asian; gnomAD: African, Amish, Latino/Admixed American, Ashkenazi Jewish, East Asian, Finnish, Non-Finnish European, Middle Eastern, South Asian, Other, XX, XY). Additionally, to find SNPs in UCRs in ancient humans and pangenome data we aligned all the UCRs with the MUSCLE algorithm [\u003cspan citationid=\"CR101\" class=\"CitationRef\"\u003e101\u003c/span\u003e]. Here, we selected the common SNPs from the Super Enhancer database (SEdb) release v2.0 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e] (available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://bio.liclab.net/sedb/download.php\u003c/span\u003e\u003cspan address=\"https://bio.liclab.net/sedb/download.php\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) to search for SNP motifs, SNP GWASs, and SNP eQTLs.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eB.T.L.N. wrote the main manuscript text and prepared all figures and tables. L.S.O. contributed to drafting the manuscript and participated in the data analysis, data curation and validation of RNA-seq analyses. A.C.R. and J.C.O. contributed to the literature review and manuscript editing. C.M. assisted in data analysis and interpretation of survival patient prognostic. D.F.G. provided critical revisions and contributed to the drafting of the manuscript. A.H.U. helped with methodology development and data curation of SNVs and SNPs. F.P. provided oversight throughout the research project. V.L.S.C., S.S.C. and A.P.S. participated in the data analysis, data curation and validation of scRNA-seq analyses. G.A.C., R.F.C., A.R.P. and J.C.O. provided expert advice and critical feedback on the study's findings. A.R.P. and J.C.O. supervised the project, reviewed the manuscript and contributed to designing the study. All authors reviewed and approved the final manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThe results published here are in whole or in part based upon data generated by the TCGA Research Network: https://www.cancer.gov/tcga. This study is supported by Araucaria Foundation - Funda\u0026ccedil;\u0026atilde;o Arauc\u0026aacute;ria in NAPI Bioinform\u0026aacute;tica (# PDI 66/2021) and CNPq (# 440412/2022-6).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eBejerano G, Pheasant M, Makunin I, Stephen S, Kent WJ, Mattick JS, Haussler D. Ultraconserved elements in the human genome. Science. 2004;304:1321\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStephen S, Pheasant M, Makunin IV, Mattick JS. Large-scale appearance of ultraconserved elements in tetrapod genomes and slowdown of the molecular clock. Mol Biol Evol. 2008;25:402\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCarter JK, Kimball RT, Funk ER, Kane NC, Schield DR, Spellman GM, Safran RJ. Estimating phylogenies from genomes: A beginners review of commonly used genomic data in vertebrate phylogenomics. J Hered. 2023;114:1\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCalin GA, Liu C-G, Ferracin M, et al. Ultraconserved regions encoding ncRNAs are altered in human leukemias and carcinomas. Cancer Cell. 2007;12:215\u0026ndash;29.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMestdagh P, Fredlund E, Pattyn F, et al. An integrative genomics screen uncovers ncRNA T-UCR functions in neuroblastoma tumours. Oncogene. 2010;29:3583\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSnetkova V, Pennacchio LA, Visel A, Dickel DE. Perfect and imperfect views of ultraconserved sequences. Nat Rev Genet. 2022;23:182\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrabh N, R\u0026ouml;delsperger C. Are orphan genes protein-coding, prediction artifacts, or non-coding RNAs? BMC Bioinformatics. 2016;17:226.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ede Oliveira JC. Transcribed Ultraconserved Regions: New regulators in cancer signaling and potential biomarkers. Genet Mol Biol. 2023;46:e20220125.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVisel A, Prabhakar S, Akiyama JA, Shoukry M, Lewis KD, Holt A, Plajzer-Frick I, Afzal V, Rubin EM, Pennacchio LA. Ultraconservation identifies a small subset of extremely constrained developmental enhancers. Nat Genet. 2008;40:158\u0026ndash;60.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSnetkova V, Ypsilanti AR, Akiyama JA, et al. Ultraconserved enhancer function does not require perfect sequence conservation. Nat Genet. 2021;53:521\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFabris L, Calin GA. (2017) Chapter Four - Understanding the Genomic Ultraconservations: T-UCRs and Cancer. In: Galluzzi L, Vitale I, editors International Review of Cell and Molecular Biology. Academic Press, pp 159\u0026ndash;172.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePereira Zambalde E, Mathias C, Rodrigues AC, de Souza Fonseca Ribeiro EM, Fiori Gradia D, Calin GA, de Carvalho J. Highlighting transcribed ultraconserved regions in human diseases. Wiley Interdiscip Rev RNA. 2020;11:e1567.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGibert MK Jr, Sarkar A, Chagari B, Cells et al. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/cells11101684\u003c/span\u003e\u003cspan address=\"10.3390/cells11101684\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLesurf R, Cotto KC, Wang G, Griffith M, Kasaian K, Jones SJM, Montgomery SB, Griffith OL, Open Regulatory Annotation Consortium. ORegAnno 3.0: a community-driven resource for curated regulatory annotation. Nucleic Acids Res. 2016;44:D126\u0026ndash;32.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVisel A, Minovitsky S, Dubchak I, Pennacchio LA. VISTA Enhancer Browser\u0026ndash;a database of tissue-specific human enhancers. Nucleic Acids Res. 2007;35:D88\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S, Francioli LC, Goodrich JK, et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature. 2024;625:92\u0026ndash;100.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaslavsky MS, Scliar MO, Yamamoto GL, et al. Whole-genome sequencing of 1,171 elderly admixed individuals from S\u0026atilde;o Paulo, Brazil. Nat Commun. 2022;13:1004.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang Y, Song C, Zhao J, et al. SEdb 2.0: a comprehensive super-enhancer database of human and mouse. Nucleic Acids Res. 2023;51:D280\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFerguson JF, Matthews GJ, Townsend RR, et al. Candidate gene association study of coronary artery calcification in chronic kidney disease: findings from the CRIC study (Chronic Renal Insufficiency Cohort). J Am Coll Cardiol. 2013;62:789\u0026ndash;98.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDiabetes Genetics Initiative of Broad Institute of Harvard and MIT, Lund University, and Novartis Institutes of BioMedical Research, Saxena R, Voight BF et al. (2007) Genome-wide association analysis identifies loci for type 2 diabetes and triglyceride levels. Science 316:1331\u0026ndash;1336.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChung SA, Brown EE, Williams AH, et al. Lupus nephritis susceptibility loci in women with systemic lupus erythematosus. J Am Soc Nephrol. 2014;25:2859\u0026ndash;70.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTakahashi Y, Kou I, Takahashi A, et al. A genome-wide association study identifies common variants near LBX1 associated with adolescent idiopathic scoliosis. Nat Genet. 2011;43:1237\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiyake A, Kou I, Takahashi Y, et al. Identification of a susceptibility locus for severe adolescent idiopathic scoliosis on chromosome 17q24.3. PLoS ONE. 2013;8:e72802.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOgura Y, Kou I, Miura S, et al. A Functional SNP in BNC2 Is Associated with Adolescent Idiopathic Scoliosis. Am J Hum Genet. 2015;97:337\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSollis E, Mosaku A, Abid A, et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51:D977\u0026ndash;85.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu Z, Guo Y, Shi H, et al. Shared genetic and experimental links between obesity-related traits and asthma subtypes in UK Biobank. J Allergy Clin Immunol. 2020;145:537\u0026ndash;49.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKoskeridis F, Evangelou E, Said S, Boyle JJ, Elliott P, Dehghan A, Tzoulaki I. Pleiotropic genetic architecture and novel loci for C-reactive protein levels. Nat Commun. 2022;13:6939.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang J, Huffman JE, Huang Y, et al. Genomics and phenomics of body mass index reveals a complex disease network. Nat Commun. 2022;13:7973.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHorikoshi M, Day FR, Akiyama M, et al. Elucidating the genetic architecture of reproductive ageing in the Japanese population. Nat Commun. 2018;9:1977.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePulit SL, Stoneman C, Morris AP, et al. Meta-analysis of genome-wide association studies for body fat distribution in 694 649 individuals of European ancestry. Hum Mol Genet. 2019;28:166\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSaunders GRB, Wang X, Chen F, et al. Genetic diversity fuels gene discovery for tobacco and alcohol use. Nature. 2022;612:720\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSinnott-Armstrong N, Tanigawa Y, Amar D, et al. Genetics of 35 blood and urine biomarkers in the UK Biobank. Nat Genet. 2021;53:185\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSakaue S, Kanai M, Tanigawa Y, et al. A cross-population atlas of genetic associations for 220 human phenotypes. Nat Genet. 2021;53:1415\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSun BB, Maranville JC, Peters JE, et al. Genomic atlas of the human plasma proteome. Nature. 2018;558:73\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLagou V, Jiang L, Ulrich A, et al. GWAS of random glucose in 476,326 individuals provide insights into diabetes pathophysiology, complications and treatment stratification. Nat Genet. 2023;55:1448\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePietzner M, Wheeler E, Carrasco-Zanini J, et al. Mapping the proteo-genomic convergence of human diseases. Science. 2021;374:eabj1541.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChernus J, Roosenboom J, Ford M, et al. GWAS reveals loci associated with velopharyngeal dysfunction. Sci Rep. 2018;8:8470.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhanshour AM, Kou I, Fan Y, et al. Genome-wide meta-analysis and replication studies in multiple ethnicities identify novel adolescent idiopathic scoliosis susceptibility loci. Hum Mol Genet. 2018;27:3986\u0026ndash;98.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKou I, Otomo N, Takeda K, et al. Genome-wide association study identifies 14 previously unreported susceptibility loci for adolescent idiopathic scoliosis in Japanese. Nat Commun. 2019;10:3685.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShen H, Lu C, Jiang Y, et al. Genetic variants in ultraconserved elements and risk of breast cancer in Chinese population. Breast Cancer Res Treat. 2011;128:855\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang R, Frank B, Hemminki K, et al. SNPs in ultraconserved elements and familial breast cancer risk. Carcinogenesis. 2008;29:351\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi X, Li P, Fuscoe JC, et al. A comprehensive rat transcriptome built from large scale RNA-seq-based annotation. Nucleic Acids Res. 2020;48:8320\u0026ndash;31.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu C, Wang Z, Liu YJ, Duan K, Guan J. Harnessing GMNP-loaded BMSC-derived EVs to target miR-3064-5p via MEG3 overexpression: Implications for diabetic osteoporosis therapy in rats. Cell Signal. 2024;118:111055.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang W, Tu H, Tang K, Huang H, Ou S, Wu J. MiR-3064 in Epicardial Adipose-Derived Exosomes Targets Neuronatin to Regulate Adipogenic Differentiation of Epicardial Adipose Stem Cells. Front Cardiovasc Med. 2021;8:709079.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuang M, Li X, Li G. Mesenchyme homeobox 1 mediated-promotion of osteoblastic differentiation is negatively regulated by mir-3064-5p. Differentiation. 2021;120:19\u0026ndash;27.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGrosu Ș-A, Dobre M, Milanesi E, Hinescu ME. (2023) Blood-Based MicroRNAs in Psychotic Disorders-A Systematic Review. Biomedicines. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/biomedicines11092536\u003c/span\u003e\u003cspan address=\"10.3390/biomedicines11092536\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePatel RB, Bajpai AK, Thirumurugan K. Differential Expression of MicroRNAs and Predicted Drug Target in Amyotrophic Lateral Sclerosis. J Mol Neurosci. 2023;73:375\u0026ndash;90.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo Z, Hao S, Yuan J, Zhu K, Liu S, Zhang J, Yao L. Long non-coding RNA LINC00958 promotes colorectal cancer progression by enhancing the expression of LEM domain containing 1 via microRNA miR-3064-5p. Bioengineered. 2021;12:8100\u0026ndash;15.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBai L, Wang H, Wang A-H, Zhang L-Y, Bai J. MicroRNA-532 and microRNA-3064 inhibit cell proliferation and invasion by acting as direct regulators of human telomerase reverse transcriptase in ovarian cancer. PLoS ONE. 2017;12:e0173912.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHe R, Zhang FH, Shen N. LncRNA FEZF1-AS1 enhances epithelial-mesenchymal transition (EMT) through suppressing E-cadherin and regulating WNT pathway in non-small cell lung cancer (NSCLC). Biomed Pharmacother. 2017;95:331\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang P, Ha M, Li L, Huang X, Liu C. MicroRNA-3064-5p sponged by MALAT1 suppresses angiogenesis in human hepatocellular carcinoma by targeting the FOXA1/CD24/Src pathway. FASEB J. 2020;34:66\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXiao E, Zhang D, Zhan W, Yin H, Ma L, Wei J, Kang Y, Mao Z. circNFIX facilitates hepatocellular carcinoma progression by targeting miR-3064-5p/HMGA2 to enhance glutaminolysis. Am J Transl Res. 2021;13:8697\u0026ndash;710.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWei M, Chen Y, Du W. LncRNA LINC00858 enhances cervical cancer cell growth through miR-3064-5p/ VMA21 axis. Cancer Biomark. 2021;32:479\u0026ndash;89.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShih C-H, Chuang L-L, Tsai M-H, Chen L-H, Chuang EY, Lu T-P, Lai L-C. Hypoxia-Induced MALAT1 Promotes the Proliferation and Migration of Breast Cancer Cells by Sponging MiR-3064-5p. Front Oncol. 2021;11:658151.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang S, Ping M, Song B, Guo Y, Li Y, Jia J. Exosomal CircPRRX1 Enhances Doxorubicin Resistance in Gastric Cancer by Regulating MiR-3064-5p/PTPN14 Signaling. Yonsei Med J. 2020;61:750\u0026ndash;61.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan J, Jia Y, Chen H, Chen W, Zhou X. Long non-coding RNA PXN-AS1 suppresses pancreatic cancer progression by acting as a competing endogenous RNA of miR-3064 to upregulate PIP4K2B expression. J Exp Clin Cancer Res. 2019;38:390.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhalilian S, Mohajer Z, Khazeei Tabari MA, Ghobadinezhad F, Ghafouri-Fard S. circGFRA1: A circular RNA with important roles in human carcinogenesis. Pathol Res Pract. 2023;248:154588.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeng M, Wu Y-C. (2022) LMX1B Activated Circular RNA GFRA1 Modulates the Tumorigenic Properties and Immune Escape of Prostate Cancer. J Immunol Res 2022:7375879.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJi X, Lv C, Huang J, Dong W, Sun W, Zhang H. ALKBH5-induced circular RNA NRIP1 promotes glycolysis in thyroid cancer cells by targeting PKM2. Cancer Sci. 2023;114:2318\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBrown GR, Hem V, Katz KS, et al. Gene: a gene-centered information resource at NCBI. Nucleic Acids Res. 2015;43:D36\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShi C, Sun L, Song Y. FEZF1-AS1: a novel vital oncogenic lncRNA in multiple human malignancies. Biosci Rep. 2019. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1042/BSR20191202\u003c/span\u003e\u003cspan address=\"10.1042/BSR20191202\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu C, Liu K, Wang B, Xu W, Lin Y, Yuan C. DLX6-AS1: An Indispensable Cancer-related Long Non-coding RNA. Curr Pharm Des. 2021;27:1211\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhalafiyan A, Emadi-Baygi M, Wolfien M, Salehzadeh-Yazdi A, Nikpour P. Construction of a three-component regulatory network of transcribed ultraconserved regions for the identification of prognostic biomarkers in gastric cancer. J Cell Biochem. 2023;124:396\u0026ndash;408.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Y, Xu S, Xia H, Gao Z, Huang R, Tang E, Jiang X. Long noncoding RNA FEZF1-AS1 in human cancers. Clin Chim Acta. 2019;497:20\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang FL, Hong K, Zhao GJ, Liu C, Song YM, Ma LL. [Construction of prognostic model and identification of prognostic biomarkers based on the expression of long non-coding RNA in bladder cancer via bioinformatics]. Beijing Da Xue Xue Bao. 2019;51:615\u0026ndash;22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTocco C, Bertacchi M, Studer M. Structural and Functional Aspects of the Neurodevelopmental Gene NR2F1: From Animal Models to Human Pathology. Front Mol Neurosci. 2021;14:767965.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHu J, Peng F, Qiu X, Yang J, Li J, Shen C, Yuan C. NR2F1-AS1: A Functional Long Noncoding RNA in Tumorigenesis. Curr Med Chem. 2023;30:4266\u0026ndash;76.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGhafouri-Fard S, Khoshbakht T, Hussen BM, Baniahmad A, Taheri M, Samsami M. A review on the role of NR2F1-AS1 in the development of cancer. Pathol Res Pract. 2022;240:154210.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuo D, Liu Y, Yuan S, Bi X, Yang Y, Zhu H, Li Z, Ji L, Yu X. The emerging role of NR2F1-AS1 in the tumorigenesis and progression of human cancer. Pathol Res Pract. 2022;235:153938.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTao L, Huang G, Song H, Chen Y, Chen L. Cancer associated fibroblasts: An essential role in the tumor microenvironment. Oncol Lett. 2017;14:2611\u0026ndash;20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGomes FG, Nedel F, Alves AM, N\u0026ouml;r JE, Tarquinio SBC. Tumor angiogenesis and lymphangiogenesis: tumor/endothelial crosstalk and cellular/microenvironmental signaling mechanisms. Life Sci. 2013;92:101\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJiang Z, Zhou J, Li L, Liao S, He J, Zhou S, Zhou Y. Pericytes in the tumor microenvironment. Cancer Lett. 2023;556:216074.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDreier MR, Walia J, de la Serna IL. (2024) Targeting SWI/SNF Complexes in Cancer: Pharmacological Approaches and Implications. Epigenomes. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/epigenomes8010007\u003c/span\u003e\u003cspan address=\"10.3390/epigenomes8010007\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang B, Guo H, Yu H, Chen Y, Xu H, Zhao G. The Role of the Transcription Factor EGR1 in Cancer. Front Oncol. 2021;11:642547.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHabic A, Mattick JS, Calin GA, Krese R, Konc J, Kunej T. Genetic Variations of Ultraconserved Elements in the Human Genome. OMICS. 2019;23:549\u0026ndash;59.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQuinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics. 2010;26:841\u0026ndash;2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrankish A, Carbonell-Sala S, Diekhans M, et al. GENCODE: reference annotation for the human and mouse genomes in 2023. Nucleic Acids Res. 2023;51:D942\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKozomara A, Birgaoanu M, Griffiths-Jones S. miRBase: from microRNA sequences to function. Nucleic Acids Res. 2019;47:D155\u0026ndash;62.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFromm B, H\u0026oslash;ye E, Domanska D, et al. MirGeneDB 2.1: toward a complete sampling of all major animal phyla. Nucleic Acids Res. 2022;50:D204\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDa Fonseca BHR, Domingues DS, Paschoal AR. mirtronDB: a mirtron knowledge base. Bioinformatics. 2019;35:3873\u0026ndash;4.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePr\u0026uuml;fer K, Racimo F, Patterson N, et al. The complete genome sequence of a Neanderthal from the Altai Mountains. Nature. 2014;505:43\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMafessoni F, Grote S, de Filippo C, et al. A high-coverage Neandertal genome from Chagyrskaya Cave. Proc Natl Acad Sci U S A. 2020;117:15132\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePr\u0026uuml;fer K, de Filippo C, Grote S, et al. A high-coverage Neandertal genome from Vindija Cave in Croatia. Science. 2017;358:655\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeyer M, Kircher M, Gansauge M-T, et al. A high-coverage genome sequence from an archaic Denisovan individual. Science. 2012;338:222\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSchuster SC, Miller W, Ratan A, et al. Complete Khoisan and Bantu genomes from southern Africa. Nature. 2010;463:943\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, Marth G, Abecasis G, Durbin R, 1000 Genome Project Data Processing Subgroup. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25:2078\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao W-W, Asri M, Ebler J, et al. A draft human pangenome reference. Nature. 2023;617:312\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRhie A, Nurk S, Cechova M, et al. The complete sequence of a human Y chromosome. Nature. 2023;621:344\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eColaprico A, Silva TC, Olsen C, et al. TCGAbiolinks: an R/Bioconductor package for integrative analysis of TCGA data. Nucleic Acids Res. 2016;44:e71.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Ge X, Peng F, Li W, Li JJ. Exaggerated false positives by popular differential expression methods when analyzing human population samples. Genome Biol. 2022;23:79.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWickham H. ggplot2: Elegant Graphics for Data Analysis. Springer; 2016.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorgan \u0026Oslash;, Patricia M, grambsch. Springer-Verlag, New York, 2000. No. Of pages: Xiii\u0026thinsp;+\u0026thinsp;350. Price: \u003cspan\u003e$\u003c/span\u003e69.95. ISBN 0‐387‐98784‐3. Stat Med 20:2053\u0026ndash;2054.\u003c/span\u003e \u003c/li\u003e \u003cli\u003e\u003cspan\u003eGavish A, Tyler M, Greenwald AC, et al. Hallmarks of transcriptional intratumour heterogeneity across a thousand tumours. Nature. 2023;618:598\u0026ndash;606.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSatija R, Farrell JA, Gennert D, Schier AF, Regev A. Spatial reconstruction of single-cell gene expression data. Nat Biotechnol. 2015;33:495\u0026ndash;502.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStuart T, Butler A, Hoffman P, Hafemeister C, Papalexi E, Mauck WM 3rd, Hao Y, Stoeckius M, Smibert P, Satija R. Comprehensive Integration of Single-Cell Data. Cell. 2019;177:1888\u0026ndash;e190221.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSherry ST, Ward MH, Kholodov M, Baker J, Phan L, Smigielski EM, Sirotkin K. dbSNP: the NCBI database of genetic variation. Nucleic Acids Res. 2001;29:308\u0026ndash;11.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBonfield JK, Marshall J, Danecek P, Li H, Ohan V, Whitwham A, Keane T, Davies RM. (2021) HTSlib: C library for reading/writing high-throughput sequencing data. Gigascience. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/gigascience/giab007\u003c/span\u003e\u003cspan address=\"10.1093/gigascience/giab007\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang K, Li M, Hakonarson H. ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res. 2010;38:e164.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaslavsky MS, Yamamoto GL, de Almeida TF, et al. Exomic variants of an elderly cohort of Brazilians in the ABraOM database. Hum Mutat. 2017;38:751\u0026ndash;63.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e1000 Genomes Project Consortium, Auton A, Brooks LD, et al. A global reference for human genetic variation. Nature. 2015;526:68\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEdgar RC. MUSCLE: a multiple sequence alignment method with reduced time and space complexity. BMC Bioinformatics. 2004;5:113.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"ultraconserved regions, non-coding RNAs, lncRNAs, ancestral genomes, pan-cancer, variation","lastPublishedDoi":"10.21203/rs.3.rs-4837600/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4837600/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBACKGROUND: Ultraconserved regions (UCRs) encompass 481 DNA segments exceeding 200 base pairs (bp), displaying 100% sequence identity across humans, mice, and rats, indicating profound conservation across taxa and pivotal functional roles in human health and disease. Despite two decades since their discovery, many UCRs remain to be explored owing to incomplete annotation, particularly of newly identified long non-coding RNAs (lncRNAs), and limited data aggregation in large-scale databases. This study offers a comprehensive functional map of 481 UCRs, investigating their genomic and transcriptomic implications: (i) enriching UCR annotation data, including ancestral genomes; (ii) exploring lncRNAs containing T-UCRs across pan-cancers; (iii) elucidating UCR involvement in regulatory elements; and (iv) analyzing population single-nucleotide variations linked to motifs, expression patterns, and diseases.\u003c/p\u003e\n\u003cp\u003eRESULTS: Our results indicate that, although a high number of protein-coding transcripts with UCRs (1,945 from 2,303), 1,775 contained UCRs outside CDS regions. Focusing on non-coding transcripts, 355 are mapped in 85 lncRNA genes, with 35 of them differentially expressed in at least one TCGA cancer type, seven lncRNAs strongly associated with survival time, and 23 differentially expressed according to single-cell cancer analysis. Additionally, we identified regulatory elements in 373 UCRs (77.5%), and found 353 SNP-UCRs (with at least 1% frequency) with potential regulatory effects, such as motif changes, eQTL potential, and associations with disease/traits. Finally, we identified 4 novel UCRs that had not been previously described.\u003c/p\u003e\n\u003cp\u003eCONCLUSION: This report compiles and organizes all the above information, providing new insights into the functional mechanisms of UCRs and their potential diagnostic applications.\u003c/p\u003e","manuscriptTitle":"The Functional Map of Ultraconserved Regions in Humans, Mice and Rats","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-28 13:02:51","doi":"10.21203/rs.3.rs-4837600/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"adcec2de-122f-4326-944e-2c89429b9ee5","owner":[],"postedDate":"August 28th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-09-03T08:38:26+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-28 13:02:51","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4837600","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4837600","identity":"rs-4837600","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00