Deciphering Complex Regions within the Human Genome and Unraveling Their Critical Biomedical Regulatory Functions | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Deciphering Complex Regions within the Human Genome and Unraveling Their Critical Biomedical Regulatory Functions Duo Du, Ming Li, Chaoyu Zhu, Xuan Li, Feizhen Wu, Wen Zhang, Xunjia Cheng, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4800291/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background: Nuclear genomic DNA plays a crucial role in individual development and phenotype determination. The genetic landscape within populations exhibits significant heterogeneity, contributing to diverse human traits. Current studies of human genome heterogeneity often focus on specific segments of high-frequency phenotype-associated sequences or structurally complex regions. Therefore, to overcome the limitations of previous studies and more directly explore population heterogeneity, it is essential to study the entire genome rather than focusing only on known phenotype-associated regions. Results: Using set theory, we have clearly defined Complex Regions (Complex_Region) by integrating pan-genome datasets, covering about 8.1% of the human genome. These regions exhibit high sequence diversity and nonrandom long continuous fragments (≥450kb), thus reflecting population genetic complexity. Our enrichment analysis revealed that genes within Complex_Region are primarily involved in immunity and metabolism, indicating chromosome-specific functional enrichment. Notably, immune genes are mainly located on chromosomes 6 and 19, which are closely associated with disease occurrence. Moreover, these regions are enriched for human phenotype-related signals and tumor somatic mutations, providing novel insights for large-scale cohort studies. We also detected ancient viral sequences, particularly ~9.47 kb human endogenous retroviruses (HERV) insertion sequence NC_022518, which is diverse in humans but remains conserved across primates, to be implicated in regulating bodily functions and various diseases. Conclusions: Our study highlights the biomedical importance of Complex_Region by revealing associations among genotypes, environment, and phenotypes. This enhances our understanding of life regulation and phenotype shaping, highlighting the role of these regions in immunity, metabolism, and disease association. Complex Regions Genotypes Phenotypes Immunity and Metabolism Pan-genome HERV Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Background Nuclear genomic DNA, which is pivotal to biological processes, orchestrates complex interactions involving gene expression regulation, epigenetic modifications, microbiome interactions, and environmental influences on organismal phenotype development. A deeper understanding of the genotype-environment-phenotype relationship can enhance our comprehension of biological diversity and ultimately provide new perspectives and strategies for disease diagnosis, treatment and prevention. Genome-wide association studies (GWAS) have revealed that over 10,000 genetic variants are associated with various phenotypes and diseases 1 . By integrating diverse human genotype-phenotype data, the DisGeNET database now includes more than 24,000 diseases and traits, and 17,000 genes with 117,000 variants 2 . In addition, researchers have identified approximately 8–9 SNPs per 10 kb in the human genome, revealing a significant nonrandom mutation pattern across populations 3 , 4 . Notably, regions such as the KIR immune gene family on chromosome 19 are highly variable within the genome 5 . Moreover, the precise resolution of large-scale genome sequences remains constrained by cost and technological limitations at this stage, with only a few complete genomes like CHM13 and HG002 available 6 , 7 . The above potential needs highlight the importance of further exploration of the complex characteristics of the human genome and population diversity. Current studies exploring the heterogeneity of the human genome tend to focus on specific similar segments of high-frequency phenotype-associated sequences in populations or structurally complex regions. On the basis of linkage disequilibrium and allele frequency, researchers have identified regions associated with complex traits, shedding light on various evolutionary patterns influencing these trait loci 8 . Furthermore, the integration of 3D structure data with GWAS signals has highlighted genes in close proximity to known phenotypic sites, facilitating the identification of potentially schizophrenia-associated genomic regions 9 . Additionally, sequences like tandem or proximal repeats, which are intricate regions of eukaryotic genomes but pose challenges for molecular identification because of their high similarity, play crucial roles in understanding human diseases, developmental processes, and environmental adaptability 10 . By analyzing the genomes of humans and eight nonhuman primates, scientists initially introduced the concept of structurally variable regions (SDRs) in primate genomes, identified several SDRs closely associated with human diseases such as Joubert syndrome 11 . To overcome the limitations of previous research and more directly explore population heterogeneity rather than focusing solely on known phenotype-associated regions. Inspired by set theory, this study has explicitly defined the Complex_Region explicitly by integrating graphical pan-genomic data, revealing the primary functions of genes in these regions associated with immune and metabolic processes. It subsequently conducted a comprehensive exploration of the sequence characteristics, potential phenotypes, and regulatory mechanisms of these regions. From another perspective, our defined Complex_Region, which reflects the "degree of sequence difference" in bioinformatic analysis, and it focuses on high-frequency, long segments, and nonrandom genomic regions within populations. However, owing to limitations in current graph pan-genome construction methods regarding repetitive sequences, our Complex_Region includes more polymorphic regions of the human genome rather than highly repetitive regions. Briefly, the Complex_Region, characterized by high sequence diversity entropy and long nonrandom contiguous segment features of at least 450kb, which can be used to characterize the complexity of genetic information at the population level, facilitating the advancement of precision medicine, with great significant potential for diverse biomedical applications. Results Defining Complex Regions within the human genome Genetic diversity within populations is a key driver of phenotypic variation, yet comprehensive large-scale genomic analysis often incurs substantial costs. Consequently, this study employs the Complex_Region as an approximate representation of the entire human genome, showing the genetic diversity of the human population. Assuming that the genome contains numerous elements, this study treats different genomic regions as distinct set elements, highlighting their varying degrees of similarity and differences among samples. The Complex_Region, defined in this study contains those elements with the most significant differences across the population, thus these regions reflect inter-individual diversity and can be used as an initial population characterization (Fig. 1 a). Considering the presence of similar or divergent sequences across individuals, the cross-overlap between circles indicates the high frequency region sequences in the population (enclosed by the light purple dotted circle). As the analysis relies on a reference-based approach, it is necessary to exclude highly similar sequences (within the light blue dotted circle) to define the final Complex_Region. This region, demarcated by the red arrow, represents the final outcome of the screening process. To better define these regions, we utilized more data from the human pan-genome 12 , 13 (HPRC and CPC) as well as the 1000 Genomes Project at different levels, using hg38 coordinates to identify and integrate genomic regions where variant density or average sequence depth was greater than the overall mean plus two standard deviations ( mean + 2* SD ) and with continuous segment lengths of at least 450 kb. To further illustrate the above definition, we visualized the Complex_Region regions on chromosome 19, and these highlighted regions can be used to roughly outline the major heterogeneity of genetic information in human populations represented by any dataset of HPRC, CPC or 1000 genomes, confirming that our selected regions meet the previous definition and that the datasets used are sufficiently representative of the population (Fig. 1 b). The Complex_Region sequences may reflect the heterogeneity within the human population, prompting us to further confirm them and mine their sequence characteristics for subsequent functional exploration. Therefore, we compared them with non-Complex_Region sequences from the hg38 genome by selecting random sequences with similar length distributions, and performed PCA dimensionality reduction after vectorisation using 6-mer tokenisation. The analysis revealed that Complex_Region sequences are slightly different from the random non-Complex_Region genomic sequences due to their inherent human origin and display greater diversity (Supplementary Fig. 1a). To minimize potential noise and accurately identify Complex_Region sequences, we constructed Complex_Region dataset in which random genomic regions were specifically used as classification control sequences. Moreover, we constructed the genomic Random dataset using a similar strategy. The final binary and multi-classification results demonstrate that the fine-tuning pre-trained DNA_bert_6 and human_gpt2-v1 models can effectively identify the Complex_Region sequences, and the classification accuracies of the Random dataset range from 0.2864 to 0.5205, whereas those of the Complex_Region dataset are within the range of 0.7942 ~ 0.8700 (Supplementary Table 1). Compared with the Random datasets, these results highlight the distinctiveness of Complex_Region sequences with random genomic regions, reflecting their high internal diversity and sequence uniqueness, as indicated by the four classification results of the Complex_Region dataset, which show the differentiation of coding sequences from non-coding regions. Sequence characterization of Complex Regions In this study, we totally identified approximately 248.69 Mb of Complex_Region, accounting for approximately 8.1% of the human genome, containing 239 distinct sequence segments, with the largest contiguous region spanning 5.433 Mb (Supplementary Table 2). These regions are distributed both in clustered and scattered across the hg38 reference genome, varying in density and location (Fig. 1 c). Notably, these regions exhibit a variable chromosomal distribution pattern, particularly within chromosomes 16, 6, 2, and 1, which contain more cumulative sequences, and similarly the clinical medical genes 14 are more abundant on chromosomes 16 and 19 (Supplementary Fig. 1b, Supplementary Table 2). This observation highlights the fragmented characterization and nonuniform distribution of the Complex_Region across chromosomes, which is consistent with previously reported nonrandom genomic mutations 15 . This further implies that these patterns may result from chromosome-specific evolutionary processes as well as being influenced by environmental factors. To determine whether the Complex_Region is adequately representative of population diversity, we further analyzed its population characteristics, inter-sequence similarity, and sequence-specific types within these regions to explore its function. The Complex_Region of the HPRC pan-genome exhibited high sequence diversity entropy, with some regions also showing sequence repetitiveness (Fig. 1 d). Compared with the whole hg38 genome, we then assessed whether there was an enrichment of specific sequence types within these regions. The results revealed a general decrease in overall repeat sequences, LINE repeat sequences, and DNase I hypersensitive site regions (DNase_Clusters), but a significant enrichment for SV genetic variants, CpG islands, high GC content, immunogenetic regions, chromosomal topological domains and pan-cancer variants in the Complex_Region. Small variants were significantly more abundant than in non-repetitive genomic regions, and slightly exceeded the overall genomic level, which is consistent with the presence of long region fragments and increased SNVs in genomic segmental duplications 16 . Additionally, other features such as partial repeat sequences (LINE-Alu; SINE; LTR-ERVK), gene numbers, regulatory elements, the human virus interaction database, HERV coding sequences and the GWAS locus were also proportionally greater than those at the genome-wide level, whereas other characteristics remained generally consistent (Supplementary Fig. 1c). The high genetic variation is similar to the sequence diversity of the HPRC pan-genome, and the low percentage of repetitive sequences supports the analysis. The low percentage of DNase_Clusters also reflects the sequence complexity within the Complex_Region 17 , 18 . The remaining features also imply the potential functional role of these regions, as evidenced by the enriched signals for somatic mutations in pan-cancer and GWAS locus, which directly indicate their potential biomedical value. Genes functional characteristics of Complex Regions The Complex_Region, which reflects the greatest genetic heterogeneity among human populations, studies of the functions of related genes help to elucidate the biomedical functions of sequences within this region and the influence of regional environments in shaping population phenotypes. Despite the insignificant enrichment of related genes in the Complex_Region compared with the genome-wide level (Supplementary Fig. 1c), genes within these regions still play a crucial role in regulating differential gene expression regulation across populations. Functional enrichment analysis revealed that these related genes are associated primarily with immunity and metabolism, development, growth, and pathways for specific infections and autoimmune diseases (Fig. 2 a, Supplementary Fig. 2a). These genes exhibit tissue-specific expression in the spleen, bone marrow, and blood, particularly in CD34 + and BDCA4 + cells (Supplementary Fig. 2b, Supplementary Table 3). Notably, compared with other regions of the autosomes, the Complex_Region shows greater significant enrichment in molecular functions such as immune system regulation, including cytokine signaling, antigen processing and presentation, immunoregulatory interactions and other processes. Meanwhile, processes such as pathogen response and metabolism, such as herpes simplex virus 1 infection, beta-defensins and fatty acid metabolism, were significantly enriched in the Complex_Region and other regions within the autosomes. However, functions related to human development, sensory perception, and complex genetics were not significantly enriched in Complex_Region, which may be potentially due to the influence of potentially repetitive sequences and complex organismal regulatory mechanisms (Supplementary Table 4). In addition, the intragenomic disease-specific enrichment revealed that genes in these regions are significantly associated with infection-induced diseases, genetic and developmental diseases, autoimmune diseases, neuromuscular diseases, and other diseases, showing greater enrichment than those in non-Complex_Region regions (Fig. 2 b, Supplementary Table 5). Given the uneven distribution of Complex_Region sequences and genes across chromosomes, we aimed to further explore their gene enrichment and their specific functional roles at different chromosomal levels. Initially, we quantified the proportion of Complex_Region genes across different chromosomes (both all and protein-coding genes), and reported that chromosomes 1, 2, 6, 11, 16, and 19 presented a notable enrichment trend, whereas chromosomes 3, 4, 8, and 10 exhibited a significant reduction in genes (Fig. 2 c). Subsequent functional enrichment analysis of overlapping genes by chromosome (chr1, chr2, chr6, chr11, chr16, chr19, and others) identified specific functional enrichments across different chromosomes (Fig. 2 d, Supplementary Table 6). For instance, chromosome 1 is associated with immunity, metabolism and development, and chromosome 2 focuses on metabolism, locomotion, and cytoarchitecture, chromosome 6 is involved in antigen processing, immunomodulation and cellular stress, chromosome 11 is closely associated with olfactory perception, antigen processing and cancer, chromosome 16 is particularly related to ethanol catabolism, membrane protein insertion and inflammation. These findings of chromosome-specific enrichments in the Complex_Region, which further enhances our understanding of population heterogeneity and local sequence function within the genome. To explore the implications of chromosome-specific immune and metabolic genes in disease mechanisms, we selected an inflammatory gene set on chromosome 19 consisting of the KIR gene cluster 5 ( KIR2DL3 , KIR2DS4 , KIR3DL1 , KIR3DL2 ), members of the immunoglobulin superfamily 19 ( LILRB1 , LILRB2 ) and genes involved in complement with inflammation regulation 20 , 21 ( C3 , NLRP2 ) for further analysis (Supplementary Fig. 3a). Moreover, the selected metabolic gene set included ATP energy metabolism genes 22 ( ATP2A1 , ATP2C2 , ATP6V0D1 ) located on chromosome 16, carbohydrate metabolism genes 23 ( AMY1A , AMY1B , AMY1C ) on chromosome 1, as well as detoxification and drug metabolism genes 24 ( GSTM1 , GSTM2 , GSTM3 ) together (Supplementary Fig. 3b). We found that these immune and metabolic gene sets exhibited strong associations with autoimmune disorders, infections, cardiovascular diseases, and malignancies, highlighting the need for deeper exploration of sequence function in the Complex_Region. Further studies we will include refining variant-level phenotypic associations and exploring the factors that regulate gene function. Disease associated signals enriched in Complex Regions Recently, large-scale GWAS studies have identified numerous genetic phenotype signals 25 . This study utilized a hypergeometric distribution analysis to further investigate the enrichment of these signals within the Complex_Region, revealing a substantial enrichment mainly in immunity-related regions (CD64 on CD14 + CD16 + monocytes; etc.) and metabolism (Triacylglycerol (56:6) [M + NH4] + levels; etc.) (Fig. 3 a, Supplementary Fig. 4a, Supplementary Table 7). This raises the question of whether the Complex_Region defined in this study can approximate genome-wide level results. To explore this further, we demonstrated the chromosomal distribution of the major genetic susceptibility loci of HIV-associated variants in infectious immune diseases and RA-associated variants in chronic immune diseases, and the results suggest that the redefined Complex_Region can cover the major association loci associated with these two diseases. Consequently, using the Complex_Region for the above enriched phenotypes in large-scale cohort studies may provide a cost-effective strategy without compromising the integrity of the research results (Fig. 3 b, Supplementary Fig. 4b). Tumors have attracted much attention as a major medical challenge. This study explores the distribution of strong somatic mutation association signals within the Complex_Region and their underlying genetic implications. A hypergeometric test revealed substantial enrichment of tumor-causing mutations in these regions, particularly in gastric adenocarcinoma (TCGA-STAD) and kidney clear cell carcinoma (TCGA-KIRC) (Fig. 3 c, Supplementary Table 8). Further analysis of genes within the Complex_Region associated with these malignancies revealed that the top 10 mutated genes in gastric adenocarcinoma predominantly harbored missense mutations, with CSMD1 gene expression strongly correlated with patient prognosis. Similarly, in kidney clear cell carcinoma, the top 10 mutated genes contained a high frequency of missense mutations and frameshift deletions, and the expression level of the PBRM1 gene, which presented the greatest number of mutations, was also strongly associated with patient survival (Fig. 3 d-e). Therefore, for tumors that are significantly enriched in somatic mutations within the Complex_Region, researchers can directly analyze the overlapping genes to detect the pathogenic genes associated with this tumor. Furthermore, prioritizing this region in large-scale tumor cohort studies may reduce experimental and computational costs without significantly compromising effectiveness. Characteristics of putative viral sequences in Complex Regions The identified Complex_Region exhibits unique sequence polymorphisms in human populations and their potential biomedical significance. Non-random regions exceeding 450 kb in length suggest that these sequences may increase population adaptability, genetic complexity, and phenotypic plasticity. Which factors contribute to this phenomenon? Environmental viral microbes, as important external driving forces in human evolution, may partially explain the emergence of this phenomenon. Therefore, we investigated putative viral sequences across the entire human genome, focusing particularly on those within the Complex_Region. Initially, the reference sequence hg38 and its endogenous viral regions were aligned separately with the Virus-Host Database 26 , revealing that the largest cumulative alignment lengths were primarily for retroviridae, herpesviridae and poxviridae, with several chromosomal variations (Supplementary Table 9). We subsequently assessed the enrichment of potential viral types in the human endogenous retrovirus, the human-virus interaction regions (HVIDB), and the defined Complex_Region, using the potential viral sequences accumulated on the different chromosomes of hg38 as a reference (Fig. 4 a). The results show that these regions also contain many potential viruses, such as retroviridae, herpesviridae, poxviridae, and baculoviridae, but clearly exhibit a clear chromosome-specific enrichment. Within the Complex_Region, the most pronounced retroviridae enrichment is on chromosomes 6 and 19, with the parvoviridae and papillomaviridae also showing chromosome-specific enrichment of potential viral sequences (Supplementary Table 10). Compared with the HERV and HVIDB regions, the Complex_Region exhibits a greater diversity of chromosome-specific viral enrichment, suggesting that external viruses contribute to the formation of these sequences, which have been integrated into human DNA through evolutionary selection, indicating that more remnants of ancient viruses may still remain in the human genome. We found that the sequences within the Complex_Region contained an ~ 2-fold enrichment of viral sequences compared with the whole hg38 genome. The majority of viral sequence remnants were short, predominantly from families like retroviridae, herpesviridae, and poxviridae, with only the Proteus phage VB_PmiS-Isfahan and the human endogenous retrovirus K113 (NC_022518) having remnants exceeding 500 bp and over 95% similarity (Fig. 4 b). Given the genome-wide similarity among many sequence fragments within the human genome, this phenomenon is likely to be related to genome sequence complexity and its dynamic regulation during evolution 27 . As a prominent manifestation of population heterogeneity, does the Complex_Region and its contained putative viral sequences exhibit a similar trend? Further exploration of sequence collinearity across the entire hg38 genome revealed that most of the highly similar segments between chromosomes are smaller than 50 kb, with only a few larger segments, reflecting the genomic similarity in the Complex_Region. The putative viral sequences within the Complex_Region also show inter-chromosomal sequence similarity, but they exhibit chromosomal preferences, such as the absence of this phenomenon on chromosomes 13 and 14, whereas NC_022518 (~ 9.47 kb) is widely present across multiple chromosomes (Supplementary Fig. 5a-b, Fig. 4 c). To further explore the organismic activities associated with the NC_022518 sequence, a functional investigation of human protein-encoding genes carrying potential NC_022518 sequences was conducted (Supplementary Table 11), which revealed that these 10 protein-encoding genes are crucial for fundamental biological processes and closely related to various diseases. The study revealed that the NC_022518 sequence predominantly resides in the intronic regions, with ENSG00000283809 present in exons 5 and 6, and these 10 genes can be clustered into two major groups (Supplementary Fig. 6a). To further investigate the characterization of sequence traits associated with potential NC_022518 sequences across primates, diverse populations, and individual genomes, and we assessed the similarity scores of these sequences, followed by PCA for dimensionality reduction (Fig. 4 d, Supplementary Fig. 6b-c). The analysis revealed no significant differences among the NC_022518-like sequences in primates, whereas the human sequences presented greater diversity. Multiple coding gene sequences highly similar to those of NC_022518 in humans clustered with primate and viral sequences, whereas the longest non-coding sequence displayed distinct characteristics. Notably, non-coding sequences closely related to NC_022518 exhibited greater disparities than coding regions did, which was particularly evident in the differences between the two haplotype segments within East Asian populations, indicating individual-level discrepancies in potential NC_022518 sequences rather than the population, and this phenomenon may be related to the adaptability of populations. The evolutionary tree constructed by the potentially similar sequences of NC_022518 also showed the same trend, in which these sequences had no obvious evolutionary distance among non-human primates, but had obvious differentiation in the population (Supplementary Fig. 6d). Furthermore, alignment results from non-redundant nucleic acid databases revealed the presence of the NC_022518-like sequences in various primate species (Supplementary Table 12), and the results of the corresponding species tree indicated that the collected sequences clustered into three primary branches, one closely aligned with HERVK-related sequences of the NC_022518 sequence, another predominantly composed of HERVK (I) sequences, and a third represented by HERVK HML-2 sequences in primates (Supplementary Fig. 7). In addition, recent research findings have highlighted HERVK as a virus recently integrated into the genome, playing a significant role in regulating processes such as aging and neurological diseases 28 , 29 , which further highlights the crucial role of the NC_022518 sequence in organismic regulation and disease. The pathogenicity and regulation of putative viral sequences in Complex Regions The immune system, a vital component of the body’s response to the environment, plays a critical role in defending against diseases and eliminating pathogens 30 , 31 . It has been found that ancient viral sequences, integrated into the human genome via reverse transcription mechanisms, persist in regulating gene expression. These sequences trigger innate immune responses by activating viral defense pathways, which are closely associated with cancer and neurological disorders 32 . To further explore the functions of these putative viral sequences in the Complex_Region, this study assessed the chromosomal distribution and related gene regulatory networks of long terminal repeat retrotransposons (LTR_Repeat), putative viral regions (Virus), immunogenetic regions (Immunogenetic), and human-virus protein interaction zones (HVIDB). The findings revealed that sequences from LTR_Repeat, Virus, and HVIDB predominantly extended across chromosomes 1, 2, 6, and 16, whereas the immunogenetic sequences were located mainly on chromosomes 6, 8, 14, and 19 (Fig. 5 a). Building on the premise that sequence structure determines function, this study investigated the common characteristics among Virus, HVIDB, and Immunogenetic sequences, revealing that these regions collectively encompass approximately 275.471 kb and include 86 genes (Fig. 5 b). The disease enrichment analysis for these genes revealed that they are associated primarily with immune and infection-related diseases, tumors and specific neurological disorders. Moreover, the average pathogenic prediction scores for gene sets associated with specific diseases are similar to the enrichment results, with conditions like Agnosia for Pain achieving a high pathogenicity score of 0.7, whereas those such as HIV-1 are around 0.4. (Supplementary Fig. 8a). Moreover, the results of the gene-protein interaction network analysis revealed that these overlapping genes are organized into three clusters: MODULE_1 involves graft-versus-host disease, antigen processing, and the presentation of exogenous peptide antigens; MODULE_2 is related to transmembrane receptor protein tyrosine kinase signaling pathways and enzyme-linked receptor-associated signaling pathways, etc.; MODULE_3 is linked to graft-versus-host disease, antigen processing and presentation, and natural killer cell-mediated cytotoxicity (Fig. 5 c). Consequently, specific regions within the Complex_Region containing putative viral sequences continue to interact with external viruses and contribute to immune regulation, signal transduction, and graft rejection processes. Notably, given that many putative viral sequences are located in non-coding regions, we used the susceptibility prediction scores of the SNPs to assess their average pathogenicity to the corresponding regions within Complex_Region. Compared with sequences within the entire Complex_Region, putative viral sequences on various chromosomes presented differing pathogenicity scores, with Chromosomes 1, 3, 4, 6, 11, 16, and 19 showing greater pathogenicity for Virus and HVIDB sequences, whereas the immunogenetic sequences on chromosomes 4, 6, 11, and 19 presented increased pathogenicity (Supplementary Fig. 8b, Supplementary Table 11). These pathogenicity assessments will continue to guide the interpretation of their functions, facilitating ongoing research. To assess the impact of potential putative viral variants in specific regions on the disease, this study further demonstrated their chromosomal distribution and associated phenotypes. The results of the cumulative small variants analysis revealed that chromosomes 1, 2, 6, 8, 12, 16, and 19 presented a relatively high density of Virus and HVIDB-related variants, with chromosomes 6, 14, and 19 exhibiting more immunogenetic variants. At the structural variants level, chromosomes 1, 5, 9, 15, 16 and 22 presented greater distributions of virus and HVIDB variants, with chromosomes 1 and 9 also harboring significant immunogenetic mutations (Supplementary Fig. 9a-b). Utilizing the results of GWAS analyses, we identified the complex phenotypes associated with these viral variants, mainly in regions related to immunity and its related diseases, metabolic and cardiovascular indicators, and general health markers (Fig. 5 d). Additionally, these GWAS-associated potential viruses show strong signals of recent positive selection (XP-EHH, etc.) and local adaptation (Fst), highlighting their active role in recent regulatory selection of the organism (Supplementary Table 14–15). Further analysis of the pathogenicity scores, functional annotations, and population frequencies of these putative viral variants revealed that only a few had high pathogenicity scores, resulting in a limited number of mutations within different score regions (Supplementary Fig. 9c, Supplementary Table 16). High pathogenic mutations often occur at enhancers and transcription factor binding sites, displaying varying frequencies across populations, such as rs3748805, which has a high frequency in all populations, rs116399863 with a higher frequency in African populations, and rs2230209 with a higher frequency in South Asian populations (Supplementary Fig. 9d). Meanwhile, these highly pathogenic mutations associated with the 3D genome involve genes regulating the immune system ( CD70 ), development and differentiation 33 ( UNCX , TBX10 ), signal transduction 34 ( G6B , DPCR1 ), cell apoptosis and growth 35 ( LGALS7 , LZRS ), and gene expression regulation 36 ( CRMP1/MIR137BD1 , DOCK11/DDX11-AS1 ), and other functional processes. Further dynamic regulatory action analysis of single putative viral mutations revealed that the high-frequency rs3748805 variant plays a complex role in gene regulation, and is located in the exon region of the THEM4 gene (Supplementary Fig. 10, Supplementary Table 17). This mutation is closely associated with the inhibition of protein kinase B (Akt), cell apoptosis, insulin signaling, cancer, eye movement dysfunctions in schizophrenia, etc 35 , 36 , 37 ., and is positively associated with AMR and SAS populations, and may be closely associated with the regional environmental viruses. Meanwhile, the rs10484554 variant in the PSORS1 gene has been found to significantly increase psoriasis risk, particularly in early-onset cases, with males being more susceptible to severe disease, possibly due to hormonal imbalances, immune response variations, or sex-specific genetic background 38 , 39 . Discussion Understanding the basis for nucleotide-level differences among us from the perspective of nuclear genomic DNA, and how major regional sequences interact with environmental factors to shape individual phenotypes, is crucial for understanding the essence of life phenomena, which will further to benefit biomedical research. This study starts from this idea, inspired by mathematical set theory, and has redefined the Complex_Region by collecting as much genome sequence data as possible from different sources and levels, accounting for about 8.1% of the genome, comprising approximately 248.69 Mb of nonrandom continuous long regions, the longest of which is about 5.43Mb. These regions have a significant distribution of functional sequences on chromosomes 1, 2, 6, and 16, and most fragments within these regions exhibit high genetic diversity and low overall repetition, and there is a significant enrichment of signals from Pan-Cancer and GWAS studies. To further explore the potential biomedical value of these sequences revealed that genes within this region are involved in immunological and metabolic processes, indicating chromosome-specific functional enrichment. Notably, many phenotypic genetic signals and somatic mutations in tumors are significantly enriched in these region s , which undoubtedly provides new perspectives for subsequent studies. The analysis of putative viral sequences in the human genome revealed that this region contains about two folds the enrichment, including sequences from retroviridae, poxviridae, and herpesviridae viruses, with the longest being the HERV K113 sequence (NC_022518; ~9.47 kb), which is located on several protein-coding genes involved in disease. In exploring the pathogenicity and regulatory roles of these putative viral sequences, we discovered that genes within the putative viral regions are involved in immune regulation, signal transduction and rejection processes, and still interact with external viruses. These interactions are likely to be involved in physiological characteristics and complex disease phenotype regulation, potentially reflecting recent environmental adaptation selection events. The Complex_Region represents a comprehensive understanding of "sequence difference degree" in the bioinformatic analysis, overcoming the limitations of previous studies that focused only on specific phenotypic regions. Despite our efforts to investigate their origins, there remain shortcomings such as incomplete high-frequency repetitive sequences, insufficient understanding of sequence causation, limited knowledge of sequence regulatory networks, and a lack of genotype-phenotype associations remain. Owing to current limitations with pan-genome data, we cannot completely cover regions related to high-frequency repetitive sequences such as centromeres, which can be directly integrated into our defined intervals in the future 16 , 37 . To deepen our understanding of sequence causality, we will continue to collaborate with evolutionary biologists based on the interpretation of environmental viruses. and adopt computational strategies such as pre-trained genome models, to deepen our understanding of their origins using a "human in the loop" approach. The complex regulatory networks of living organisms, such as gene regulatory networks and the influence of non-coding sequences on gene expression, which cannot be studied in humans by methods like genetic perturbation, so the Complex_Region should be extended to experimental animal models in the future. Currently, public databases collecting both genotypes and complex phenotypes, particularly for non-European-American populations, are still scarce and more data need to be collected and pre-processed. Therefore, our future work will focus on understanding this phenomenon from more methods and perspectives, expanding this definition to other mammals and building model animals. We will continue to collect and integrate more "genotype-complex phenotype" data to validate and explore, and apply the Complex_Region to specific biomedical research, such as the use of CRISPR-Cas9 technology. In summary, in our study within the Complex_Region, we attempt to understand the associations between genotype, the external environment, and external phenotypes, fundamentally exploring potential regulatory factors behind population differences in phenotypes. The results can be summarized in the following aspects: 1) The definition and characteristics of the Complex_Region, it was defined by various data sources and levels, exhibiting high sequence diversity entropy, nonrandom distribution, and long continuous segments (≥ 450 kb), which can roughly characterize the complexity of population genetic information. 2) The biomedical significance and functions of the Complex_Region, where gene function enrichment is focused mainly on immunity and metabolism, and there is chromosome-specific functional enrichment. Many genetically related phenotypes and somatic mutations in tumors are significantly enriched in these regions, thus allowing for inexpensive and comprehensive preliminary exploration in cohort studies of these diseases. 3) With respect to the origin and regulation of putative viral sequences in the Complex_Region, there is significant potential enrichment of exogenous viral sequence enrichment, particularly the longest fragment NC_022518-like sequence, which is classified in the HERVK family. This sequence is particularly present in several disease-regulatory coding genes, reflecting high population diversity and conservation across primates. This study further explored the pathogenicity and regulatory roles of these putative viral sequences and their impact on life processes and phenotype shaping. In essence, the Complex_Region contributes to a deeper understanding of human genetic complexity at the population level, further advances precision medicine, and provides a guide for subsequent sequence function studies, but it still needs to be further explored and interpreted. Methods Complex Regions selection by integrating multiple data types A range of data sources and types were used to ensure that the complex regions identified were sufficiently representative. The selection process included GFA format screening based on human graphic pangenomes (HPRC and CPC), VCF file screening based on human graphic pangenomes HPRC constructed by Minigraph-Cactus and PGGB respectively 38 , 39 , and 1000 Genomes VCF file screening. For the screening of GFA files based on human pangenomes, the HPRC and CPC graphic pangenomes constructed via the Minigraph-Cactus process were initially converted to the GFA format using vg convert, and then together with the HPRC graphic pangenomes constructed by the PGGB method, the GFA files were further transformed using odgi build. This was followed by a preliminary screening for complex regions, where we calculated the average mapping depth within 500 bp windows using odgi depth 40 and bedtools across the three pan-genome datasets. After sorting the depth data in ascending order, we removed the top and bottom 20% to avoid outliers, using mean + 2* SD as our filtering threshold for depths within 500 bp windows, and merged the windows within 15 kb using bedtools merge, with the cumulative lengths for each graphical pan-genome separately. When screening the VCF files from the graphical pan-Genome constructed by Minigraph-Cactus and PGGB, we retained loci where the alternative allele frequencies were not less than 0.3, and then counted the corresponding SNPs within 500 bp windows. We then used a similar filtering approach to further obtain the potential regions. For screenings based on the 1000 Genomes VCF files, the filters used are as described above. We then combined the potential complex regions obtained by filtering these varied data sources, sorted them by position and coordinates, and merged intervals within 15 kb to form preliminary complex regions. To emphasize longer segments and ensure the broad applicability of identified complex regions in biomedical research, this study retained regions exceeding 450 kb as the final complex regions. To increase their representativeness, the longest sequences from the Y and M chromosomes were also included, creating approximately 248.468 Mb of complex regions. To demonstrate the screening process of the above process, we visualized the average depth (HPRC and CPC) and variant count (1000 Genomes) within 500 bp windows of these complex regions on chromosome 19. This visualization enhanced our understanding of the characteristics of these genomic regions. Uniqueness sequence characterization within Complex Regions To assess the sequence specificity of the identified complex regions, we first obtained potential non-complex regions within the human genome using bedtools subtract, and finally used bedtools shuffle to finally generate non-complex regions similar to the distribution of complex regions. The corresponding DNA sequences of these regions from the hg38 reference were extracted by coordinates using bedtools getfasta by coordinates, and transformed them into numerical 6-mer feature vectors using CountVectorizer. We standardized the resulting sparse matrix using StandardScaler without subtracting the mean of each element during the process to maintain sparsity, and further used SparsePCA for dimensionality reduction and presentation using the above results. Next, following the strategy described in a recent publication 41 , we created the Complex_Region dataset and the Random dataset. We then fine-tuned the DNA_bert_6 and human_gpt2-v1 models for binary and multiple classification of DNA sequences using the hugging face interface. In the binary classification task of sequences, the Complex_Region dataset aims to distinguish between complex regions and non-complex regions sequences; the Random dataset is used to classify genomic random region sequences that are similarly distributed to the Complex_Region dataset. In the quaternary classification of sequences, the Complex_Region dataset further distinguishes whether a sequence is located in the coding region or not using the above labels, whereas the Random dataset is used to randomly distinguish genomic sequences. Sequence characterization of identified Complex Regions After obtaining the complex regions, we proceeded to quantify the longest segment and cumulative length of each chromosome, and used the CMRG_v1.00_HG002 clinical gene set region as a reference for visual display and statistics. To further delineate the diversity and repetitiveness of sequences within the complex region, we extracted the potential sequences of each interval from the HPRC pan-genome using odgi extract and transformed them into FASTA sequences using odgi paths after optimization and sequencing by odgi sort, respectively. Then the sequences exceeding 10kb were retained using seqkit seq, and the pgr-tk’s pgr-pbundle-decomp 42 was used to obtain the primary components, leading to turn constructed the GFA graph files of the corresponding regions. Additionally, we performed a simple diffusion model for each extracted complex region sequence using the pgrtk.compute_graph_diffusion_entropy function to obtain the sequence diversity entropy and node weights, and used the last 32 elements mean in the weight list as an assessment of sequence repeatability. Finally, we visualized the relationship between the sequence diversity entropy value and its repeatability for all complex regions. Furthermore, by integrating data from UCSC, GENCODE, and pertinent literature, we calculated and displayed the proportions of cumulative sequence lengths in complex regions relative to the entire genome, and displayed them after enrichment rate calculation. Enrichment of overlapping gene functions in Complex Regions Using the hg38 reference genome annotations, we utilized bedtools intersect to extract annotations for genes overlapping with complex regions (≥ 150 bp) for further analysis. We then employed the Metascape 43 to enrich and perform disease-related analysis through DisGeNET using the selected protein-coding gene set. Concurrently, we evaluated the functional enrichment and disease characteristics of genes in complex regions and non-complex region from each chromosome, to further verify their functional specificity via Metascape in complex regions. Subsequently, the genes within the complex regions were categorized into protein-coding (Protein_coding) and all genes (All), with proportions calculated based on the basis of chromosomal distribution against the hg38 reference. Chromosomes were labeled as (chr1, chr16, chr11, chr6, chr19, chr2, others). Functional enrichment and DisGeNET-related analyses were conducted for the protein-coding gene sets on each chromosome using Metascape, with the results displayed according to the new chromosome labels. Furthermore, select immune and metabolic genes within the complex regions were analyzed for gene-disease associations via DisGeNET. Genetic phenotypic signals and somatic mutation enrichment in Complex Regions We initially downloaded the GWAS summary files from the UCSC website, identifying 10,376 phenotypes with 403,362 signals across the entire genome, of which 3,975 phenotypes comprising 38,859 signals were located within complex regions. To quantify the prevalence of trait enrichment in complex regions versus the whole genome, we first counted the cumulative Trait_num for all the signal sites associated with the 3,975 phenotypes. Subsequently, we normalized this count by the length of the corresponding region (Len), resulting in Trait_num/Len. For each of the 3,975 phenotypes, we calculated the enrichment ratio (ER) by comparing the Trait_num/Len of the complex regions to that at the genome-wide level. To determine which phenotypes were enriched in the complex region, we performed a hypergeometric distribution test for each phenotype with the following formula: phyper( N -1, M , Total_num - M , Total_case , lower.tail = F), where N represents the number of signals in the complex regions (Trait_num1), M represents those at the genome level (Trait_num2), Total_num is the total 403,362 association signals across the genome, and Total_case is the 38,859 signals within the complex regions. We applied the Benjamini-Hochberg method for P -value correction, classifying significant phenotypes into three major categories: immune system and cellular biology, metabolism and biochemistry, and health status and other biometrics. Phenotypes with an ER greater than five and an adjusted P -value less than 0.05 were visualized using EnhancedVolcano plots. Furthermore, GWAS signal loci for diseases like the rheumatoid arthritis (RA) and human immunodeficiency virus (HIV) infection were identified at both the genome-wide and complex region levels to demonstrate the adequately representativeness of complex regions. Using a similar computational approach, we analyzed the downloaded pan-cancer data for 33 tumor types, calculating the Trait_num and ER of signals within genome-wide and complex regions. We applied the hypergeometric distribution test phyper( N -1, M , Total_num - M , Total_case , lower.tail = F), where N is the tumor signal Trait_num1 in the complex regions, M being the tumour signal Trait_num2 at the genome level, and Total_num being 3,038,556 tumor association signals at the genome-wide level, and Total_case is 325,537 tumor association signals in the complex region. Next, we used EnhancedVolcano for volcano plot visualization of the enrichment results with the pan-cancer type as the label, and selected gastric adenocarcinoma (STAD) and renal clear cell carcinoma (KIRC) as the most significantly enriched cancers. In addition, the genes with the most mutated types in STAD and KIRC were displayed after statistical analysis by maftools 44 , and the most mutated genes were selected for survival analysis using the RNA sequencing, respectively. Enrichment of putative viral sequences within Complex Regions To comprehensively analyze the viral sequences within complex regions, we adopted a dual alignment strategy at both the endogenous retrovirus (HERV) and whole-genome levels. Initially, at the HERV level, we constructed a virus sequence database using makeblastdb, and then extracted the HERV sequences from the reference hg38 using bedtools getfasta, and utilized bedtools intersect to isolate structural variant (SV) sequences within the HERV regions from the 1000 Genomes cohort. Subsequently, we aligned these HERV sequences from hg38 and 1000 Genomes to the viral database using blastn respectively. At the whole-genome level, SVs from the 1000 Genomes data were converted into sequences and merged with the reference hg38, and a corresponding database was constructed using makeblastdb, followed by comparing Virushostdb sequences to this database. For the above comparison results, we filtered and retained sequences longer than 30 bp with over 95% similarity, categorizing them based on the viral family information in the Virushostdb database to obtain a potential viral dataset, including both HERV and other aligned viral sequences. Using this filtered viral dataset, we calculated the non-redundant cumulative alignment lengths for specific regions (hg38, Complex_Region, HERV_Region and HVIDB_Region), and demonstrated the distribution of potential viral alignment lengths on different chromosomes. We also assessed the enrichment ratios of potential aligned viral sequences within a specific region (Complex_Region, HERV_Region, HVIDB_ Region), based on the proportion of viral sequences on different chromosomes of hg38. Finally, we showcased all potential viral alignment sequences within the Complex_Region. Characteristics and species distribution of the NC_022518 Sequence To comprehensively analyze the NC_022518 sequence within the complex regions, we initiated our study by utilizing nucmer to align against each sequence within the reference hg38, and then the corresponding region alignment relationships were shown by show-coords and used as controls 45 . Subsequent extractions of sequences from both the Complex_Region and the Virus regions (including NC_022518), were aligned back to hg38 to obtain chromosomal sequence similarity. The alignment results were then categorized into four major categories: 10–50 kb, 50–100 kb, 100–500 kb, and ≥ 500 kb, and further to visualize based on this classification. To explore the characteristics of NC_022518-like sequences in humans, we identified ten protein-coding genes containing this region. We then assessed their functions, pathways, and cell or tissue specificity using the GenDoma database, and explored their corresponding population and primate sequence features using UCSC. To elucidate the relationships between similar sequences among these genes, sequences that highly similar to NC_022518 were extracted with bedtools getfasta, constructed an evolutionary tree using iqtree after mafft multiple sequence comparison 46 , 47 , and finally visualized it. To explain the divergence of NC_022518 in primates and humans, we identified its longest protein-coding TRPC6 region (chr11:101695064–101704528) and a similar fragment within the LncRNA gene ENSG00000286016 (chr12:127153658–127168021) on the genome hg38, and continue to analyze the sequences within these regions in depth. Firstly, we extracted the sequence fragments in primate, human pan-genome and genome hg38 according to the above coordinates in turn. For sequence extraction in primates, we used minimap2 to align the fragments within ENSG00000286016 and chromosome 12 to the primate genome respectively, and then extracted the corresponding sequences from the aligned regions. For human pan-genome sequence extraction, the above two region sequences in HPRC and CPC were extracted by odgi extract separately, and the corresponding sequences were obtained after odgi sort and odgi paths conversion. For the sequence fragment of genome hg38, we used bedtools getfasta to retain the potential protein coding sequence and the sequence within ENSG00000286016. Immediately following the above steps, we combined the potential sequences together to filter out the sequences with more than 80% of non-A/G/C/T bases, and the filtered sequences were compared by mafft to obtain their multiple sequence alignment results. In order to further portray the degree of similarity between sequences, we used pgrtk.get_shmmr_pairs_from_seq to extract the key information from the multi-sequence alignment results, and combined it with the sample annotation information to perform PCA downscaling and visualization. Meanwhile, we also constructed evolutionary trees using iqtree and ggtree to demonstrate the inter-sequence relationships. Furthermore, we aligned the NC_022518 sequences to the non-redundant nucleic acid database of NCBI, further statistically analyzed the alignment results and constructed the corresponding species sequence trees. Chromosomal distribution, regulatory networks, and pathogenicity prediction of putative viral regions Using bedtools intersect, we extracted sequences from specified regions (LTR_Repeat, Virus, Immunogenetic, HVIDB) within the complex regions and displayed them after counting them separately by chromosome. To assess the immune characteristics of these sequences, we first used bedtools intersect to count overlapping sequences between endogenous and other aligned viral sequences (Virus), human-virus interaction region sequences (HVIDB), and human immune region sequences (Immunogenetic). We selected the gene overlap among these three groups and used the Metascape website to analyze the protein interaction networks and visualize the core gene modules. For the pathogenicity assessment of specific regions, we first extracted single-base pathogenicity scores predicted by PrimateAI and AlphaMissense 48 , 49 , and then used a bedtools map to calculate the average pathogenicity scores of the desired region (Complex_Region, Virus, Immunogenetic, HVIDB), and presented them at the chromosome level. Moreover, we employed a similar strategy to calculate the pathogenicity scores of overlapping genes and demonstrated the similarity between the pathogenicity scores and the functional enrichment results of these genes. Pathogenicity and intrabody regulation of putative viral variants To analyze putative viral mutations within specified genomic regions (Virus, HVIDB, Immunogenetic), we initially employed bedtools intersect to extract small variants and SVs, followed by a chromosomal distribution analysis of these variants. Subsequently, we extracted GWAS association results to delineate the characteristics of virus-related variants. Utilizing the GSEL toolkit 50 , we identified several evolutionary selection indicators in GWAS association signals, including signals for natural selection and adaptivity differentiation (xPEHH; iES; GERP), as well as measures for genetic diversity and population differentiation (Fst) and conservation assessments (PhastCons; phyloP100). Furthermore, we calculated the pathogenicity scores of all identified viral variants and classified them into different pathogenicity score ranges, which included (0,0.3], (0.3,0.5], (0.5,0.7], and (0.7,1]), and then counted and displayed the number of variants in the different ranges. To ensure the reliability of our findings, we repeatedly verified that these variants were located in the HERV region, followed by selecting the variants with scores above 0.7 for conversion to the reference hg19 using liftover 51 . We employed the 3DSNP database to perform regulatory analysis on these high-pathogenicity variants. Finally, we describe and discuss the functional regions (Enhancer, Promoter, TFBS, Motif) and intra-population frequencies of these variants, and finely explore the in vivo regulation of some high-frequency variants and their associated diseases. Abbreviations HERV Human endogenous retrovirus GWAS Whole-genome association studies DNase_Clusters DNase I hypersensitive site regions SDRs structurally variable regions LTR_Repeat long terminal repeat retrotransposons HVIDB human-virus PPI database zones SV structural variant Declarations Generative AI in scientific writing AI-assisted technologies have been employed to increase readability and improve language proficiency. Nevertheless, the final results were exclusively produced by the authors, who meticulously edited the language to conform to domain terminology. Consequently, we are wholly responsible and accountable for the content of this study. Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Competing Interest None declared. Funding This work was supported by the Peak Disciplines (Type IV) of Institutions of Higher Learning in Shanghai. Author Contribution Conceptualization: DD, FZ, LL; Formal Analysis: DD; Funding Acquisition: LL; Investigation: DD, ML; Methodology: DD, ML; Project Administration: FZ, LL; Resources: LL; Supervision: FZ, LL; Useful Suggestions: WZ, FZW, XC; Visualization: DD, CYZ; Writing – Original Draft: DD, XL; Writing – Review & Editing: DD, FZ, LL. Acknowledgement This work was supported by the Peak Disciplines (Type IV) of Institutions of Higher Learning in Shanghai. We thank all the authors for their hard work. This work was supported by the Medical Research Data Center of Fudan University. Data Availability All relevant raw data supporting the key findings of this study are available within the article and its Supplementary Information page. All custom analytical pipeline codes used in this work are available at https://github.com/ GeorgeBGM/Complex-genome_analysis.git. References Momozawa Y, Mizukami K. Unique roles of rare variants in the genetics of complex diseases in humans. J Hum Genet. 2021;66:11–23. Pinero J, et al. The DisGeNET knowledge platform for disease genomics: 2019 update. Nucleic Acids Res. 2020;48:D845–55. Zhao Z, Fu YX, Hewett-Emmett D, Boerwinkle E. Investigating single nucleotide polymorphism (SNP) density in the human genome and its implications for molecular evolution. Gene. 2003;312:207–13. Campbell CD, Eichler EE. Properties and rates of germline mutations in humans. Trends Genet. 2013;29:575–84. Middleton D, Gonzelez F. The extensive polymorphism of KIR genes. Immunology. 2010;129:8–19. Nurk S, et al. The complete sequence of a human genome. Science. 2022;376:44–53. Guo Y, Feng X, Li H. Evaluation of haplotype-aware long-read error correction with hifieval. Bioinformatics 39, (2023). Abraham A, LaBella AL, Capra JA, Rokas A. Mosaic patterns of selection in genomic regions associated with diverse human traits. PLoS Genet. 2022;18:e1010494. Buxton DS, Batten DJ, Crofts JJ, Chuzhanova N. Predicting novel genomic regions linked to genetic disorders using GWAS and chromosome conformation data - a case study of schizophrenia. Sci Rep. 2019;9:17940. Ranz J, Clifton B. Characterization and evolutionary dynamics of complex regions in eukaryotic genomes. Sci China Life Sci. 2019;62:467–88. Mao Y, et al. Structurally divergent and recurrently mutated regions of primate genomes. Cell. 2024;187:1547–e15621513. Liao WW, et al. A draft human pangenome reference. Nature. 2023;617:312–24. Gao Y, et al. A pangenome reference of 36 Chinese populations. Nature. 2023;619:112–21. Wagner J, et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat Biotechnol. 2022;40:672–80. Monroe JG, et al. Mutation bias reflects natural selection in Arabidopsis thaliana. Nature. 2022;602:101–5. Vollger MR, et al. Increased mutation and gene conversion within human segmental duplications. Nature. 2023;617:325–34. Thurman RE, et al. The accessible chromatin landscape of the human genome. Nature. 2012;489:75–82. Chen A, Chen D, Chen Y. Advances of DNase-seq for mapping active gene regulatory elements across the genome in animals. Gene. 2018;667:83–94. Zeller T, et al. Dual checkpoint blockade of CD47 and LILRB1 enhances CD20 antibody-dependent phagocytosis of lymphoma cells by macrophages. Front Immunol. 2022;13:929339. Zarantonello A, Revel M, Grunenwald A, Roumenina LT. C3-dependent effector functions of complement. Immunol Rev. 2023;313:120–38. Zhang T, et al. NLRP2 in health and disease. Immunology. 2024;171:170–80. Fujikura Y, et al. Ketogenic diet containing medium-chain triglyceride ameliorates transcriptome disruption in skeletal muscles of rat models of duchenne muscular dystrophy. Biochem Biophys Rep. 2022;32:101378. Hariharan R, Mousa A, de Courten B. Influence of AMY1A copy number variations on obesity and other cardiometabolic risk factors: A review of the evidence. Obes Rev. 2021;22:e13205. Zhang J, et al. Comprehensive analysis of the glutathione S-transferase Mu (GSTM) gene family in ovarian cancer identifies prognostic and expression significance. Front Oncol. 2022;12:968547. Chang M, He L, Cai L. An Overview of Genome-Wide Association Studies. Methods Mol Biol. 2018;1754:97–108. Mihara T, et al. Linking Virus Genomes with Host Taxonomy. Viruses. 2016;8:66. Guarracino A, et al. Recombination between heterologous human acrocentric chromosomes. Nature. 2023;617:335–43. Jern P, Sperber GO, Blomberg J. Use of endogenous retroviral sequences (ERVs) and structural markers for retroviral phylogenetic inference and taxonomy. Retrovirology. 2005;2:50. Wang J, Lu X, Zhang W, Liu GH. Endogenous retroviruses in development and health. Trends Microbiol. 2024;32:342–54. Jo EK. Interplay between host and pathogen: immune defense and beyond. Exp Mol Med. 2019;51:1–3. Kang SH, Sun YD, Atallah OO, Huguet-Tapia JC, Noble JD, Folimonova SY. A Long Non-Coding RNA of Citrus tristeza virus: Role in the Virus Interplay with the Host Immunity. Viruses 11, (2019). Jakobsson J, Vincendeau M, SnapShot. Human endogenous retroviruses. Cell. 2022;185:400–e400401. Yonezawa Y, et al. Identification of a Functional Susceptibility Variant for Adolescent Idiopathic Scoliosis that Upregulates Early Growth Response 1 (EGR1)-Mediated UNCX Expression. J Bone Min Res. 2023;38:144–53. Radoux-Mergault A, Oberhauser L, Aureli S, Gervasio FL, Stoeber M. Subcellular location defines GPCR signal transduction. Sci Adv. 2023;9:eadf6059. Sewgobind NV, Albers S, Pieters RJ. Functions and Inhibition of Galectin-7, an Emerging Target in Cellular Pathophysiology. Biomolecules 11, (2021). Chen Y, et al. The regulation of DOCK family proteins on T and B cells. J Leukoc Biol. 2021;109:383–94. Dolzhenko E et al. Characterization and visualization of tandem repeats at genome scale. Nat Biotechnol, (2024). Hickey G, et al. Pangenome graph construction from genome alignments with Minigraph-Cactus. Nat Biotechnol. 2024;42:663–73. Garrison E et al. Building pangenome graphs. bioRxiv, (2023). Guarracino A, Heumos S, Nahnsen S, Prins P, Garrison E. ODGI: understanding pangenome graphs. Bioinformatics. 2022;38:3319–26. Du D, Zhong F, Liu L. Enhancing Recognition and Interpretation of Functional Phenotypic Sequences through Fine-Tuning Pre-Trained Genomic Models. bioRxiv, 2023.2012.2005.570173 (2023). Chin CS, et al. Multiscale analysis of pangenomes enables improved representation of genomic diversity for repetitive and clinically relevant genes. Nat Methods. 2023;20:1213–21. Zhou Y, et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10:1523. Mayakonda A, Lin DC, Assenov Y, Plass C, Koeffler HP. Maftools: efficient and comprehensive analysis of somatic variants in cancer. Genome Res. 2018;28:1747–56. Riva G, Mauri M. MuMMER: How Robotics Can Reboot Social Interaction and Customer Engagement in Shops and Malls. Cyberpsychol Behav Soc Netw. 2021;24:210–1. Nguyen LT, Schmidt HA, von Haeseler A, Minh BQ. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol Biol Evol. 2015;32:268–74. Katoh K, Rozewicki J, Yamada KD. MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization. Brief Bioinform. 2019;20:1160–6. Gao H, et al. The landscape of tolerated genetic variation in humans and primates. Science. 2023;380:eabn8153. Cheng J, et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science. 2023;381:eadg7492. Abraham A, Labella AL, Benton ML, Rokas A, Capra JA. GSEL: a fast, flexible python package for detecting signatures of diverse evolutionary forces on genomic regions. Bioinformatics 39, (2023). Park KJ, Yoon YA, Park JH. Evaluation of Liftover Tools for the Conversion of Genome Reference Consortium Human Build 37 to Build 38 Using ClinVar Variants. Genes (Basel) 14, (2023). Additional Declarations No competing interests reported. Supplementary Files 2024724ManuscriptTables.xlsx 2024724SupplementaryAppendix.pdf Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4800291","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":336437056,"identity":"9d3445bf-4df1-41e2-a823-d35ad3aa7117","order_by":0,"name":"Duo Du","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"prefix":"","firstName":"Duo","middleName":"","lastName":"Du","suffix":""},{"id":336437057,"identity":"58580adf-082e-4c21-bcc3-bb38f8b70061","order_by":1,"name":"Ming Li","email":"","orcid":"","institution":"University of Konstanz","correspondingAuthor":false,"prefix":"","firstName":"Ming","middleName":"","lastName":"Li","suffix":""},{"id":336437058,"identity":"31258bcd-c43a-4e2b-a404-904fcdae7b52","order_by":2,"name":"Chaoyu Zhu","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"prefix":"","firstName":"Chaoyu","middleName":"","lastName":"Zhu","suffix":""},{"id":336437059,"identity":"1f0077c7-b332-4bbe-b0f4-f148b6f2199a","order_by":3,"name":"Xuan Li","email":"","orcid":"","institution":"CAS","correspondingAuthor":false,"prefix":"","firstName":"Xuan","middleName":"","lastName":"Li","suffix":""},{"id":336437060,"identity":"cd20276b-58a5-4e70-9869-87910bf5f0ea","order_by":4,"name":"Feizhen Wu","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"prefix":"","firstName":"Feizhen","middleName":"","lastName":"Wu","suffix":""},{"id":336437061,"identity":"cda12b52-691e-4e6b-9a3f-f0ea4e0c2e77","order_by":5,"name":"Wen Zhang","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"prefix":"","firstName":"Wen","middleName":"","lastName":"Zhang","suffix":""},{"id":336437062,"identity":"0b340665-43b1-46c5-94b8-f2b0c5e63121","order_by":6,"name":"Xunjia Cheng","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"prefix":"","firstName":"Xunjia","middleName":"","lastName":"Cheng","suffix":""},{"id":336437063,"identity":"daae7650-7e9a-482b-a475-8fc75b433546","order_by":7,"name":"Fan Zhong","email":"","orcid":"","institution":"Fudan University","correspondingAuthor":false,"prefix":"","firstName":"Fan","middleName":"","lastName":"Zhong","suffix":""},{"id":336437064,"identity":"2c56626d-ad73-4c9c-a9c5-bd0348d4181d","order_by":8,"name":"Lei Liu","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA0ElEQVRIie3QrQ7CMBSG4W9pcjANsyUQuIUmJLO7lVZhIJmcQCDIJgBPwk2AQ0KWTBWPHAYNDkHCT/iRtBLRR1WcN+k5gOf9oTpDsIEESCCoVDq0J8SAT8JkZUqHBM/kSYAahzFzSGpcbk4J4nozP6Z6RAjzibJ8jMvtXEJnLRPt9boFYXZLa1JwCUWiT3ttCFIMHJKrREyid0x0xhyTx8WCTKgIjgkl2+lrl65QpuTWXcKwWFWXK+LOIj+cL+mwHeaz38nb7fviLuOe53mexR10fz7o+Y8QTgAAAABJRU5ErkJggg==","orcid":"","institution":"Fudan University","correspondingAuthor":true,"prefix":"","firstName":"Lei","middleName":"","lastName":"Liu","suffix":""}],"badges":[],"createdAt":"2024-07-25 08:36:34","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4800291/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4800291/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":61955639,"identity":"89630214-d00e-4585-844e-f7e5a7cbc596","added_by":"auto","created_at":"2024-08-07 13:31:38","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1297160,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDefinition and sequences characteristics in Complex Regions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Representation of the collection and definition of Complex_Region. The diagram uses circles (light blue; light green; light orange) to represent the nuclear genomic DNA sequences from different individuals, with black dots pinpointing distinct regional sequences within the whole genome. The potentially high frequency occurrence of sequences in the population (within the light purple dotted circle) and the highly similar sequences (within the light blue dotted circle) are highlighted. The red arrows indicate the region that will eventually form the Complex_Region, which emerges between the light purple and light blue dotted circles. b. Schematic representation of the complex region screened on chromosome 19. To ensure the representativeness of the redefined Complex_Region, we included data from the human pan-genome project (HPRC; CPC) and the 1000 Genomes. The x-axis shows the coordinates of human chromosome 19, while the y-axis measures the variation density and sequence depth from the three datasets (HPRC in red, CPC in blue and the 1000 Genomes in green), with highlight regions indicating selected sequences from chromosome 19 that show dense clusters of population heterogeneity variants. The middle part of the x-axis is the centromeric region of the chromosome. c. The chromosome distribution of selected regions within Complex_Region. The Complex_Region is shown in red. d. The HPRC pan-genomic sequence diversity and repeatability across 239 segments within the Complex_ Region. The x-axis may reflect sequence diversity entropy with colors representing different chromosomes, while the y-axis represents sequence repeatability.\u003c/p\u003e","description":"","filename":"Fig.1N.png","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/40076e506dbf1a64567f274a.png"},{"id":61955641,"identity":"7caf8565-8816-4b52-a093-22bcad8d861d","added_by":"auto","created_at":"2024-08-07 13:31:38","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":656425,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eFunctional analysis of genes within Complex Regions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. GO enrichment analysis of all protein-coding genes within Complex_Region. This figure presents the functional enrichment analysis for all genes within Complex_Region, with the primary functions of these genes are highlighted in red, and the x-axis indicates the \u003cem\u003ep\u003c/em\u003e-value of the corresponding terms. b. The results of disease enrichment analysis for all protein-coding genes within Complex_Region. c. The distribution of different types of overlapping gene ratios at the chromosome level. Using hg38 as a reference, the figure shows the percentage change of all genes and protein-coding genes across different chromosomes within the Complex_Region. Light red indicates all genes on hg38, light green indicates protein-coding genes on hg38, light blue indicates all genes within the Complex_Region, and light purple indicates protein-coding genes within Complex_Region. d. The functional enrichment analysis of specific protein coding genes within Complex_Region. Functional enrichment analysis of protein coding genes within Complex _Region was performed by chromosome, and the enrichment results are categorized by chromosomes 1, 2, 6, 11, 16, 19, and others based on the gene proportions in c. The statistics for enrichment, including the count of enriched genes and their \u003cem\u003ep\u003c/em\u003e-values, are detailed in the legend on the right.\u003c/p\u003e","description":"","filename":"Fig.2.png","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/414d786144eab03cbe463da0.png"},{"id":61956288,"identity":"e8342d65-6d44-4994-8167-2ade77a4f802","added_by":"auto","created_at":"2024-08-07 13:39:38","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1947556,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGenetic phenotypic signals and tumor somatic mutations enriched within Complex Regions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Enrichment of human genetic phenotypes within Complex_Region. The x-axis represents the enrichment ratio (\u003cem\u003eER\u003c/em\u003e), and the \u003cem\u003ey\u003c/em\u003e-axis shows the adjusted \u003cem\u003ep\u003c/em\u003e-value. Phenotypic signals are categorized based on \u003cem\u003ep \u003c/em\u003e≤ 0.05 and \u003cem\u003eER\u003c/em\u003e ≥2 into: neither NS (grey) for signals that meet neither criterion; \u003cem\u003eER\u003c/em\u003e for those meeting only the enrichment criterion; Adjusted \u003cem\u003ep\u003c/em\u003e-value (blue) for those meeting only the \u003cem\u003ep\u003c/em\u003e-value criterion; Adjusted \u003cem\u003ep\u003c/em\u003e-value and \u003cem\u003eER\u003c/em\u003e (yellow) for signals meeting both criteria. A black dashed line at \u003cem\u003ey\u003c/em\u003e=0.05 demarcates significance. Each phenotypic point is plotted according to the above criteria and labelled some of the actual names. b. Genetic susceptibility loci for the HIV infectious disease covered by the complex region. The upper section of the figure presents GWAS signal points for HIV infection, with a cutoff indicated by red dashed lines at \u003cem\u003ep\u003c/em\u003e=5x10^\u003csup\u003e-8\u003c/sup\u003e. The lower section shows different signal sites within Complex_Region by chromosomes. c. The distribution of human tumor somatic mutation enrichment in Complex_Region. We used a similar strategy as in a to display and annotate the signals of tumors specifically identified within the Complex_Region. d. The most mutated genes and their prognosis in Complex_Region of stomach cancer (STAD). Genetic variants in STAD are analyzed in this figure, with the classification of these variants shown on the left, the most highly mutated genes in STAD selected in the middle, and survival analysis of the CSMD1 gene on the right. e. The most mutated genes and their prognosis in Complex_Region of kidney clear cell carcinoma (KIRC). We analyzed KIRC using the strategy in d, containing the classification of variants, the display of the most mutated genes, and the survival analysis of specific genes.\u003c/p\u003e","description":"","filename":"Fig.3.png","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/1efdbf548fe9c2e9681410ad.png"},{"id":61955648,"identity":"353e6db7-2c66-4ad3-88d2-8f167850d8d5","added_by":"auto","created_at":"2024-08-07 13:31:38","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":3157422,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCharacteristics of putative viral sequences in Complex Regions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. The distribution and characteristics of putative viral sequences. As shown in the figure, we firstly demonstrated the classification of potential viral-like sequences (AlignedLength) within the reference hg38 by known virus types. Then the potential viral enrichment ratios (EnrichmentRate) in the Complex Region, HERV and the human-virus interaction database (HVIDB) were evaluated relative to hg38. The chromosome order is consistent across the four plots, as shown by the corresponding numbering on the diagrams. b. The alignment of potential virus sequences within Complex_Region. We show the aligned sequences of potential viruses with over 90% similarity within this region, the horizontal axis shows the length of the aligned sequences, and the different symbols indicate the potential virus types (~150 viruses on hg38; ~32 in the Complex_Region and it is ~8.1% of hg38). c. The similarity between the putative viral sequences within Complex_Region and hg38. We classified and visualized the putative viral sequences within this region into different types based on their aligned lengths, and demonstrated the specificity of NC_022518's potential sequences within hg38. d. The sequence similarity relationships between Inter-primates, in-population and intra-genomic for the NC_022518 sequence. This figure presents a principal component analysis of sequence similarity for NC_022518, illustrating principal component 1 (PC1) and principal component 2 (PC2). Virus represents the NC_022518 sequence, Non_Hominidae represents the non-hominid primate family, EUR, EAS, SAS, AMR and AFR represent the European, East Asian, South Asian, white American and African populations respectively.\u003c/p\u003e","description":"","filename":"Fig.4.png","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/83bdde3a40c16cef7a5afcc9.png"},{"id":61955649,"identity":"67d62ed0-db16-4b89-8140-cf6f26b0cca1","added_by":"auto","created_at":"2024-08-07 13:31:39","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1411493,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRegulatory processes and pathogenicity of specific sequences in Complex Regions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ea. Chromosomal distribution of different types of genomic sequences. This figure illustrates the cumulative length distribution of LTR_Repeat, Virus, Immunogenic, and HVIDB sequences across the human genome. Light red is LTR_Repeat sequences, light green indicates Virus sequences, light blue represents Immunogenic sequences and light purple is HVIDB sequences. b. The intersection of virus sequences, HVIDB data and immunogenic regions in human body. The left side shows the intersection between the three sequences (Seqs), light red represents the virus sequences in human, light green represents the HVIDB data sequences and light blue is the Immunogenic sequences. c. The core modules of the gene interaction network in the Complex_ Region. MCODE_1 is marked in red, MCODE_2 in blue, and MCODE_3 in green. d. The associated phenotypes of variation in the putative viral regions. The diagram presents selected GWAS signal sites within the putative viral regions, classified by their potential biological functions types, and displays their \u003cem\u003ep\u003c/em\u003e-values and odds ratios (\u003cem\u003eOR\u003c/em\u003e) of these loci.\u003c/p\u003e","description":"","filename":"Fig.5N.png","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/be9f3b78967c6b440b59c503.png"},{"id":61956921,"identity":"cb65c8f8-bb63-4275-bc6c-3b7a8fbcd11f","added_by":"auto","created_at":"2024-08-07 13:47:45","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":10119743,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/3b97e2ac-a32b-4de9-8e8c-ec2e17c2bd90.pdf"},{"id":61956287,"identity":"e51f9e44-6495-4f8b-b289-d835a1f6ab50","added_by":"auto","created_at":"2024-08-07 13:39:38","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":737419,"visible":true,"origin":"","legend":"","description":"","filename":"2024724ManuscriptTables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/39c6d7672c790f123332e43e.xlsx"},{"id":61955645,"identity":"a15adde7-25b6-431c-8668-ca0f00ff5d88","added_by":"auto","created_at":"2024-08-07 13:31:38","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":2843632,"visible":true,"origin":"","legend":"","description":"","filename":"2024724SupplementaryAppendix.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4800291/v1/4ff7bbce0eff2717b455df53.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Deciphering Complex Regions within the Human Genome and Unraveling Their Critical Biomedical Regulatory Functions","fulltext":[{"header":"Background","content":"\u003cp\u003eNuclear genomic DNA, which is pivotal to biological processes, orchestrates complex interactions involving gene expression regulation, epigenetic modifications, microbiome interactions, and environmental influences on organismal phenotype development. A deeper understanding of the genotype-environment-phenotype relationship can enhance our comprehension of biological diversity and ultimately provide new perspectives and strategies for disease diagnosis, treatment and prevention. Genome-wide association studies (GWAS) have revealed that over 10,000 genetic variants are associated with various phenotypes and diseases\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e. By integrating diverse human genotype-phenotype data, the DisGeNET database now includes more than 24,000 diseases and traits, and 17,000 genes with 117,000 variants\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. In addition, researchers have identified approximately 8\u0026ndash;9 SNPs per 10 kb in the human genome, revealing a significant nonrandom mutation pattern across populations\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Notably, regions such as the KIR immune gene family on chromosome 19 are highly variable within the genome\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Moreover, the precise resolution of large-scale genome sequences remains constrained by cost and technological limitations at this stage, with only a few complete genomes like CHM13 and HG002 available\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. The above potential needs highlight the importance of further exploration of the complex characteristics of the human genome and population diversity.\u003c/p\u003e \u003cp\u003eCurrent studies exploring the heterogeneity of the human genome tend to focus on specific similar segments of high-frequency phenotype-associated sequences in populations or structurally complex regions. On the basis of linkage disequilibrium and allele frequency, researchers have identified regions associated with complex traits, shedding light on various evolutionary patterns influencing these trait loci\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Furthermore, the integration of 3D structure data with GWAS signals has highlighted genes in close proximity to known phenotypic sites, facilitating the identification of potentially schizophrenia-associated genomic regions\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. Additionally, sequences like tandem or proximal repeats, which are intricate regions of eukaryotic genomes but pose challenges for molecular identification because of their high similarity, play crucial roles in understanding human diseases, developmental processes, and environmental adaptability\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e. By analyzing the genomes of humans and eight nonhuman primates, scientists initially introduced the concept of structurally variable regions (SDRs) in primate genomes, identified several SDRs closely associated with human diseases such as Joubert syndrome\u003csup\u003e\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTo overcome the limitations of previous research and more directly explore population heterogeneity rather than focusing solely on known phenotype-associated regions. Inspired by set theory, this study has explicitly defined the Complex_Region explicitly by integrating graphical pan-genomic data, revealing the primary functions of genes in these regions associated with immune and metabolic processes. It subsequently conducted a comprehensive exploration of the sequence characteristics, potential phenotypes, and regulatory mechanisms of these regions. From another perspective, our defined Complex_Region, which reflects the \"degree of sequence difference\" in bioinformatic analysis, and it focuses on high-frequency, long segments, and nonrandom genomic regions within populations. However, owing to limitations in current graph pan-genome construction methods regarding repetitive sequences, our Complex_Region includes more polymorphic regions of the human genome rather than highly repetitive regions. Briefly, the Complex_Region, characterized by high sequence diversity entropy and long nonrandom contiguous segment features of at least 450kb, which can be used to characterize the complexity of genetic information at the population level, facilitating the advancement of precision medicine, with great significant potential for diverse biomedical applications.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eDefining Complex Regions within the human genome\u003c/h2\u003e \u003cp\u003eGenetic diversity within populations is a key driver of phenotypic variation, yet comprehensive large-scale genomic analysis often incurs substantial costs. Consequently, this study employs the Complex_Region as an approximate representation of the entire human genome, showing the genetic diversity of the human population. Assuming that the genome contains numerous elements, this study treats different genomic regions as distinct set elements, highlighting their varying degrees of similarity and differences among samples. The Complex_Region, defined in this study contains those elements with the most significant differences across the population, thus these regions reflect inter-individual diversity and can be used as an initial population characterization (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). Considering the presence of similar or divergent sequences across individuals, the cross-overlap between circles indicates the high frequency region sequences in the population (enclosed by the light purple dotted circle). As the analysis relies on a reference-based approach, it is necessary to exclude highly similar sequences (within the light blue dotted circle) to define the final Complex_Region. This region, demarcated by the red arrow, represents the final outcome of the screening process. To better define these regions, we utilized more data from the human pan-genome\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e (HPRC and CPC) as well as the 1000 Genomes Project at different levels, using hg38 coordinates to identify and integrate genomic regions where variant density or average sequence depth was greater than the overall mean plus two standard deviations (\u003cem\u003emean\u003c/em\u003e\u0026thinsp;+\u0026thinsp;2*\u003cem\u003eSD\u003c/em\u003e) and with continuous segment lengths of at least 450 kb. To further illustrate the above definition, we visualized the Complex_Region regions on chromosome 19, and these highlighted regions can be used to roughly outline the major heterogeneity of genetic information in human populations represented by any dataset of HPRC, CPC or 1000 genomes, confirming that our selected regions meet the previous definition and that the datasets used are sufficiently representative of the population (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe Complex_Region sequences may reflect the heterogeneity within the human population, prompting us to further confirm them and mine their sequence characteristics for subsequent functional exploration. Therefore, we compared them with non-Complex_Region sequences from the hg38 genome by selecting random sequences with similar length distributions, and performed PCA dimensionality reduction after vectorisation using 6-mer tokenisation. The analysis revealed that Complex_Region sequences are slightly different from the random non-Complex_Region genomic sequences due to their inherent human origin and display greater diversity (Supplementary Fig.\u0026nbsp;1a). To minimize potential noise and accurately identify Complex_Region sequences, we constructed Complex_Region dataset in which random genomic regions were specifically used as classification control sequences. Moreover, we constructed the genomic Random dataset using a similar strategy. The final binary and multi-classification results demonstrate that the fine-tuning pre-trained DNA_bert_6 and human_gpt2-v1 models can effectively identify the Complex_Region sequences, and the classification accuracies of the Random dataset range from 0.2864 to 0.5205, whereas those of the Complex_Region dataset are within the range of 0.7942\u0026thinsp;~\u0026thinsp;0.8700 (Supplementary Table\u0026nbsp;1). Compared with the Random datasets, these results highlight the distinctiveness of Complex_Region sequences with random genomic regions, reflecting their high internal diversity and sequence uniqueness, as indicated by the four classification results of the Complex_Region dataset, which show the differentiation of coding sequences from non-coding regions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003eSequence characterization of Complex Regions\u003c/h2\u003e \u003cp\u003eIn this study, we totally identified approximately 248.69 Mb of Complex_Region, accounting for approximately 8.1% of the human genome, containing 239 distinct sequence segments, with the largest contiguous region spanning 5.433 Mb (Supplementary Table\u0026nbsp;2). These regions are distributed both in clustered and scattered across the hg38 reference genome, varying in density and location (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec). Notably, these regions exhibit a variable chromosomal distribution pattern, particularly within chromosomes 16, 6, 2, and 1, which contain more cumulative sequences, and similarly the clinical medical genes\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e are more abundant on chromosomes 16 and 19 (Supplementary Fig.\u0026nbsp;1b, Supplementary Table\u0026nbsp;2). This observation highlights the fragmented characterization and nonuniform distribution of the Complex_Region across chromosomes, which is consistent with previously reported nonrandom genomic mutations\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. This further implies that these patterns may result from chromosome-specific evolutionary processes as well as being influenced by environmental factors.\u003c/p\u003e \u003cp\u003eTo determine whether the Complex_Region is adequately representative of population diversity, we further analyzed its population characteristics, inter-sequence similarity, and sequence-specific types within these regions to explore its function. The Complex_Region of the HPRC pan-genome exhibited high sequence diversity entropy, with some regions also showing sequence repetitiveness (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ed). Compared with the whole hg38 genome, we then assessed whether there was an enrichment of specific sequence types within these regions. The results revealed a general decrease in overall repeat sequences, LINE repeat sequences, and DNase I hypersensitive site regions (DNase_Clusters), but a significant enrichment for SV genetic variants, CpG islands, high GC content, immunogenetic regions, chromosomal topological domains and pan-cancer variants in the Complex_Region. Small variants were significantly more abundant than in non-repetitive genomic regions, and slightly exceeded the overall genomic level, which is consistent with the presence of long region fragments and increased SNVs in genomic segmental duplications\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e. Additionally, other features such as partial repeat sequences (LINE-Alu; SINE; LTR-ERVK), gene numbers, regulatory elements, the human virus interaction database, HERV coding sequences and the GWAS locus were also proportionally greater than those at the genome-wide level, whereas other characteristics remained generally consistent (Supplementary Fig.\u0026nbsp;1c). The high genetic variation is similar to the sequence diversity of the HPRC pan-genome, and the low percentage of repetitive sequences supports the analysis. The low percentage of DNase_Clusters also reflects the sequence complexity within the Complex_Region\u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e. The remaining features also imply the potential functional role of these regions, as evidenced by the enriched signals for somatic mutations in pan-cancer and GWAS locus, which directly indicate their potential biomedical value.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003eGenes functional characteristics of Complex Regions\u003c/h2\u003e \u003cp\u003eThe Complex_Region, which reflects the greatest genetic heterogeneity among human populations, studies of the functions of related genes help to elucidate the biomedical functions of sequences within this region and the influence of regional environments in shaping population phenotypes. Despite the insignificant enrichment of related genes in the Complex_Region compared with the genome-wide level (Supplementary Fig.\u0026nbsp;1c), genes within these regions still play a crucial role in regulating differential gene expression regulation across populations. Functional enrichment analysis revealed that these related genes are associated primarily with immunity and metabolism, development, growth, and pathways for specific infections and autoimmune diseases (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea, Supplementary Fig.\u0026nbsp;2a). These genes exhibit tissue-specific expression in the spleen, bone marrow, and blood, particularly in CD34\u003csup\u003e+\u003c/sup\u003e and BDCA4\u003csup\u003e+\u003c/sup\u003e cells (Supplementary Fig.\u0026nbsp;2b, Supplementary Table\u0026nbsp;3). Notably, compared with other regions of the autosomes, the Complex_Region shows greater significant enrichment in molecular functions such as immune system regulation, including cytokine signaling, antigen processing and presentation, immunoregulatory interactions and other processes. Meanwhile, processes such as pathogen response and metabolism, such as herpes simplex virus 1 infection, beta-defensins and fatty acid metabolism, were significantly enriched in the Complex_Region and other regions within the autosomes. However, functions related to human development, sensory perception, and complex genetics were not significantly enriched in Complex_Region, which may be potentially due to the influence of potentially repetitive sequences and complex organismal regulatory mechanisms (Supplementary Table\u0026nbsp;4). In addition, the intragenomic disease-specific enrichment revealed that genes in these regions are significantly associated with infection-induced diseases, genetic and developmental diseases, autoimmune diseases, neuromuscular diseases, and other diseases, showing greater enrichment than those in non-Complex_Region regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb, Supplementary Table\u0026nbsp;5).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGiven the uneven distribution of Complex_Region sequences and genes across chromosomes, we aimed to further explore their gene enrichment and their specific functional roles at different chromosomal levels. Initially, we quantified the proportion of Complex_Region genes across different chromosomes (both all and protein-coding genes), and reported that chromosomes 1, 2, 6, 11, 16, and 19 presented a notable enrichment trend, whereas chromosomes 3, 4, 8, and 10 exhibited a significant reduction in genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). Subsequent functional enrichment analysis of overlapping genes by chromosome (chr1, chr2, chr6, chr11, chr16, chr19, and others) identified specific functional enrichments across different chromosomes (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed, Supplementary Table\u0026nbsp;6). For instance, chromosome 1 is associated with immunity, metabolism and development, and chromosome 2 focuses on metabolism, locomotion, and cytoarchitecture, chromosome 6 is involved in antigen processing, immunomodulation and cellular stress, chromosome 11 is closely associated with olfactory perception, antigen processing and cancer, chromosome 16 is particularly related to ethanol catabolism, membrane protein insertion and inflammation. These findings of chromosome-specific enrichments in the Complex_Region, which further enhances our understanding of population heterogeneity and local sequence function within the genome.\u003c/p\u003e \u003cp\u003eTo explore the implications of chromosome-specific immune and metabolic genes in disease mechanisms, we selected an inflammatory gene set on chromosome 19 consisting of the KIR gene cluster\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eKIR2DL3\u003c/em\u003e, \u003cem\u003eKIR2DS4\u003c/em\u003e, \u003cem\u003eKIR3DL1\u003c/em\u003e, \u003cem\u003eKIR3DL2\u003c/em\u003e), members of the immunoglobulin superfamily\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eLILRB1\u003c/em\u003e, \u003cem\u003eLILRB2\u003c/em\u003e) and genes involved in complement with inflammation regulation\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eC3\u003c/em\u003e, \u003cem\u003eNLRP2\u003c/em\u003e) for further analysis (Supplementary Fig.\u0026nbsp;3a). Moreover, the selected metabolic gene set included ATP energy metabolism genes\u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eATP2A1\u003c/em\u003e, \u003cem\u003eATP2C2\u003c/em\u003e, \u003cem\u003eATP6V0D1\u003c/em\u003e) located on chromosome 16, carbohydrate metabolism genes\u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eAMY1A\u003c/em\u003e, \u003cem\u003eAMY1B\u003c/em\u003e, \u003cem\u003eAMY1C\u003c/em\u003e) on chromosome 1, as well as detoxification and drug metabolism genes\u003csup\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eGSTM1\u003c/em\u003e, \u003cem\u003eGSTM2\u003c/em\u003e, \u003cem\u003eGSTM3\u003c/em\u003e) together (Supplementary Fig.\u0026nbsp;3b). We found that these immune and metabolic gene sets exhibited strong associations with autoimmune disorders, infections, cardiovascular diseases, and malignancies, highlighting the need for deeper exploration of sequence function in the Complex_Region. Further studies we will include refining variant-level phenotypic associations and exploring the factors that regulate gene function.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003eDisease associated signals enriched in Complex Regions\u003c/h2\u003e \u003cp\u003eRecently, large-scale GWAS studies have identified numerous genetic phenotype signals\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/sup\u003e. This study utilized a hypergeometric distribution analysis to further investigate the enrichment of these signals within the Complex_Region, revealing a substantial enrichment mainly in immunity-related regions (CD64 on CD14\u003csup\u003e+\u003c/sup\u003e CD16\u003csup\u003e+\u003c/sup\u003e monocytes; etc.) and metabolism (Triacylglycerol (56:6) [M\u0026thinsp;+\u0026thinsp;NH4]\u003csup\u003e+\u003c/sup\u003e levels; etc.) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea, Supplementary Fig.\u0026nbsp;4a, Supplementary Table\u0026nbsp;7). This raises the question of whether the Complex_Region defined in this study can approximate genome-wide level results. To explore this further, we demonstrated the chromosomal distribution of the major genetic susceptibility loci of HIV-associated variants in infectious immune diseases and RA-associated variants in chronic immune diseases, and the results suggest that the redefined Complex_Region can cover the major association loci associated with these two diseases. Consequently, using the Complex_Region for the above enriched phenotypes in large-scale cohort studies may provide a cost-effective strategy without compromising the integrity of the research results (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb, Supplementary Fig.\u0026nbsp;4b).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTumors have attracted much attention as a major medical challenge. This study explores the distribution of strong somatic mutation association signals within the Complex_Region and their underlying genetic implications. A hypergeometric test revealed substantial enrichment of tumor-causing mutations in these regions, particularly in gastric adenocarcinoma (TCGA-STAD) and kidney clear cell carcinoma (TCGA-KIRC) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec, Supplementary Table\u0026nbsp;8). Further analysis of genes within the Complex_Region associated with these malignancies revealed that the top 10 mutated genes in gastric adenocarcinoma predominantly harbored missense mutations, with \u003cem\u003eCSMD1\u003c/em\u003e gene expression strongly correlated with patient prognosis. Similarly, in kidney clear cell carcinoma, the top 10 mutated genes contained a high frequency of missense mutations and frameshift deletions, and the expression level of the \u003cem\u003ePBRM1\u003c/em\u003e gene, which presented the greatest number of mutations, was also strongly associated with patient survival (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ed-e). Therefore, for tumors that are significantly enriched in somatic mutations within the Complex_Region, researchers can directly analyze the overlapping genes to detect the pathogenic genes associated with this tumor. Furthermore, prioritizing this region in large-scale tumor cohort studies may reduce experimental and computational costs without significantly compromising effectiveness.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003eCharacteristics of putative viral sequences in Complex Regions\u003c/h2\u003e \u003cp\u003eThe identified Complex_Region exhibits unique sequence polymorphisms in human populations and their potential biomedical significance. Non-random regions exceeding 450 kb in length suggest that these sequences may increase population adaptability, genetic complexity, and phenotypic plasticity. Which factors contribute to this phenomenon? Environmental viral microbes, as important external driving forces in human evolution, may partially explain the emergence of this phenomenon. Therefore, we investigated putative viral sequences across the entire human genome, focusing particularly on those within the Complex_Region. Initially, the reference sequence hg38 and its endogenous viral regions were aligned separately with the Virus-Host Database\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e, revealing that the largest cumulative alignment lengths were primarily for retroviridae, herpesviridae and poxviridae, with several chromosomal variations (Supplementary Table\u0026nbsp;9). We subsequently assessed the enrichment of potential viral types in the human endogenous retrovirus, the human-virus interaction regions (HVIDB), and the defined Complex_Region, using the potential viral sequences accumulated on the different chromosomes of hg38 as a reference (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ea). The results show that these regions also contain many potential viruses, such as retroviridae, herpesviridae, poxviridae, and baculoviridae, but clearly exhibit a clear chromosome-specific enrichment. Within the Complex_Region, the most pronounced retroviridae enrichment is on chromosomes 6 and 19, with the parvoviridae and papillomaviridae also showing chromosome-specific enrichment of potential viral sequences (Supplementary Table\u0026nbsp;10). Compared with the HERV and HVIDB regions, the Complex_Region exhibits a greater diversity of chromosome-specific viral enrichment, suggesting that external viruses contribute to the formation of these sequences, which have been integrated into human DNA through evolutionary selection, indicating that more remnants of ancient viruses may still remain in the human genome.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe found that the sequences within the Complex_Region contained an ~\u0026thinsp;2-fold enrichment of viral sequences compared with the whole hg38 genome. The majority of viral sequence remnants were short, predominantly from families like retroviridae, herpesviridae, and poxviridae, with only the Proteus phage VB_PmiS-Isfahan and the human endogenous retrovirus K113 (NC_022518) having remnants exceeding 500 bp and over 95% similarity (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003eb). Given the genome-wide similarity among many sequence fragments within the human genome, this phenomenon is likely to be related to genome sequence complexity and its dynamic regulation during evolution\u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e. As a prominent manifestation of population heterogeneity, does the Complex_Region and its contained putative viral sequences exhibit a similar trend? Further exploration of sequence collinearity across the entire hg38 genome revealed that most of the highly similar segments between chromosomes are smaller than 50 kb, with only a few larger segments, reflecting the genomic similarity in the Complex_Region. The putative viral sequences within the Complex_Region also show inter-chromosomal sequence similarity, but they exhibit chromosomal preferences, such as the absence of this phenomenon on chromosomes 13 and 14, whereas NC_022518 (~\u0026thinsp;9.47 kb) is widely present across multiple chromosomes (Supplementary Fig.\u0026nbsp;5a-b, Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ec).\u003c/p\u003e \u003cp\u003eTo further explore the organismic activities associated with the NC_022518 sequence, a functional investigation of human protein-encoding genes carrying potential NC_022518 sequences was conducted (Supplementary Table\u0026nbsp;11), which revealed that these 10 protein-encoding genes are crucial for fundamental biological processes and closely related to various diseases. The study revealed that the NC_022518 sequence predominantly resides in the intronic regions, with ENSG00000283809 present in exons 5 and 6, and these 10 genes can be clustered into two major groups (Supplementary Fig.\u0026nbsp;6a). To further investigate the characterization of sequence traits associated with potential NC_022518 sequences across primates, diverse populations, and individual genomes, and we assessed the similarity scores of these sequences, followed by PCA for dimensionality reduction (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003ed, Supplementary Fig.\u0026nbsp;6b-c). The analysis revealed no significant differences among the NC_022518-like sequences in primates, whereas the human sequences presented greater diversity. Multiple coding gene sequences highly similar to those of NC_022518 in humans clustered with primate and viral sequences, whereas the longest non-coding sequence displayed distinct characteristics. Notably, non-coding sequences closely related to NC_022518 exhibited greater disparities than coding regions did, which was particularly evident in the differences between the two haplotype segments within East Asian populations, indicating individual-level discrepancies in potential NC_022518 sequences rather than the population, and this phenomenon may be related to the adaptability of populations. The evolutionary tree constructed by the potentially similar sequences of NC_022518 also showed the same trend, in which these sequences had no obvious evolutionary distance among non-human primates, but had obvious differentiation in the population (Supplementary Fig.\u0026nbsp;6d). Furthermore, alignment results from non-redundant nucleic acid databases revealed the presence of the NC_022518-like sequences in various primate species (Supplementary Table\u0026nbsp;12), and the results of the corresponding species tree indicated that the collected sequences clustered into three primary branches, one closely aligned with HERVK-related sequences of the NC_022518 sequence, another predominantly composed of HERVK (I) sequences, and a third represented by HERVK HML-2 sequences in primates (Supplementary Fig.\u0026nbsp;7). In addition, recent research findings have highlighted HERVK as a virus recently integrated into the genome, playing a significant role in regulating processes such as aging and neurological diseases\u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, which further highlights the crucial role of the NC_022518 sequence in organismic regulation and disease.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eThe pathogenicity and regulation of putative viral sequences in Complex Regions\u003c/h2\u003e \u003cp\u003eThe immune system, a vital component of the body\u0026rsquo;s response to the environment, plays a critical role in defending against diseases and eliminating pathogens\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. It has been found that ancient viral sequences, integrated into the human genome via reverse transcription mechanisms, persist in regulating gene expression. These sequences trigger innate immune responses by activating viral defense pathways, which are closely associated with cancer and neurological disorders\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. To further explore the functions of these putative viral sequences in the Complex_Region, this study assessed the chromosomal distribution and related gene regulatory networks of long terminal repeat retrotransposons (LTR_Repeat), putative viral regions (Virus), immunogenetic regions (Immunogenetic), and human-virus protein interaction zones (HVIDB). The findings revealed that sequences from LTR_Repeat, Virus, and HVIDB predominantly extended across chromosomes 1, 2, 6, and 16, whereas the immunogenetic sequences were located mainly on chromosomes 6, 8, 14, and 19 (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ea).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eBuilding on the premise that sequence structure determines function, this study investigated the common characteristics among Virus, HVIDB, and Immunogenetic sequences, revealing that these regions collectively encompass approximately 275.471 kb and include 86 genes (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eb). The disease enrichment analysis for these genes revealed that they are associated primarily with immune and infection-related diseases, tumors and specific neurological disorders. Moreover, the average pathogenic prediction scores for gene sets associated with specific diseases are similar to the enrichment results, with conditions like Agnosia for Pain achieving a high pathogenicity score of 0.7, whereas those such as HIV-1 are around 0.4. (Supplementary Fig.\u0026nbsp;8a). Moreover, the results of the gene-protein interaction network analysis revealed that these overlapping genes are organized into three clusters: MODULE_1 involves graft-versus-host disease, antigen processing, and the presentation of exogenous peptide antigens; MODULE_2 is related to transmembrane receptor protein tyrosine kinase signaling pathways and enzyme-linked receptor-associated signaling pathways, etc.; MODULE_3 is linked to graft-versus-host disease, antigen processing and presentation, and natural killer cell-mediated cytotoxicity (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ec). Consequently, specific regions within the Complex_Region containing putative viral sequences continue to interact with external viruses and contribute to immune regulation, signal transduction, and graft rejection processes. Notably, given that many putative viral sequences are located in non-coding regions, we used the susceptibility prediction scores of the SNPs to assess their average pathogenicity to the corresponding regions within Complex_Region. Compared with sequences within the entire Complex_Region, putative viral sequences on various chromosomes presented differing pathogenicity scores, with Chromosomes 1, 3, 4, 6, 11, 16, and 19 showing greater pathogenicity for Virus and HVIDB sequences, whereas the immunogenetic sequences on chromosomes 4, 6, 11, and 19 presented increased pathogenicity (Supplementary Fig.\u0026nbsp;8b, Supplementary Table\u0026nbsp;11). These pathogenicity assessments will continue to guide the interpretation of their functions, facilitating ongoing research.\u003c/p\u003e \u003cp\u003eTo assess the impact of potential putative viral variants in specific regions on the disease, this study further demonstrated their chromosomal distribution and associated phenotypes. The results of the cumulative small variants analysis revealed that chromosomes 1, 2, 6, 8, 12, 16, and 19 presented a relatively high density of Virus and HVIDB-related variants, with chromosomes 6, 14, and 19 exhibiting more immunogenetic variants. At the structural variants level, chromosomes 1, 5, 9, 15, 16 and 22 presented greater distributions of virus and HVIDB variants, with chromosomes 1 and 9 also harboring significant immunogenetic mutations (Supplementary Fig.\u0026nbsp;9a-b). Utilizing the results of GWAS analyses, we identified the complex phenotypes associated with these viral variants, mainly in regions related to immunity and its related diseases, metabolic and cardiovascular indicators, and general health markers (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003ed). Additionally, these GWAS-associated potential viruses show strong signals of recent positive selection (XP-EHH, etc.) and local adaptation (Fst), highlighting their active role in recent regulatory selection of the organism (Supplementary Table\u0026nbsp;14\u0026ndash;15). Further analysis of the pathogenicity scores, functional annotations, and population frequencies of these putative viral variants revealed that only a few had high pathogenicity scores, resulting in a limited number of mutations within different score regions (Supplementary Fig.\u0026nbsp;9c, Supplementary Table\u0026nbsp;16). High pathogenic mutations often occur at enhancers and transcription factor binding sites, displaying varying frequencies across populations, such as rs3748805, which has a high frequency in all populations, rs116399863 with a higher frequency in African populations, and rs2230209 with a higher frequency in South Asian populations (Supplementary Fig.\u0026nbsp;9d). Meanwhile, these highly pathogenic mutations associated with the 3D genome involve genes regulating the immune system (\u003cem\u003eCD70\u003c/em\u003e), development and differentiation\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eUNCX\u003c/em\u003e, \u003cem\u003eTBX10\u003c/em\u003e), signal transduction\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eG6B\u003c/em\u003e, \u003cem\u003eDPCR1\u003c/em\u003e), cell apoptosis and growth\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eLGALS7\u003c/em\u003e, \u003cem\u003eLZRS\u003c/em\u003e), and gene expression regulation\u003csup\u003e\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e (\u003cem\u003eCRMP1/MIR137BD1\u003c/em\u003e, \u003cem\u003eDOCK11/DDX11-AS1\u003c/em\u003e), and other functional processes. Further dynamic regulatory action analysis of single putative viral mutations revealed that the high-frequency rs3748805 variant plays a complex role in gene regulation, and is located in the exon region of the THEM4 gene (Supplementary Fig.\u0026nbsp;10, Supplementary Table\u0026nbsp;17). This mutation is closely associated with the inhibition of protein kinase B (Akt), cell apoptosis, insulin signaling, cancer, eye movement dysfunctions in schizophrenia, etc\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e,\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e,\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e., and is positively associated with AMR and SAS populations, and may be closely associated with the regional environmental viruses. Meanwhile, the rs10484554 variant in the \u003cem\u003ePSORS1\u003c/em\u003e gene has been found to significantly increase psoriasis risk, particularly in early-onset cases, with males being more susceptible to severe disease, possibly due to hormonal imbalances, immune response variations, or sex-specific genetic background\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eUnderstanding the basis for nucleotide-level differences among us from the perspective of nuclear genomic DNA, and how major regional sequences interact with environmental factors to shape individual phenotypes, is crucial for understanding the essence of life phenomena, which will further to benefit biomedical research. This study starts from this idea, inspired by mathematical set theory, and has redefined the Complex_Region by collecting as much genome sequence data as possible from different sources and levels, accounting for about 8.1% of the genome, comprising approximately 248.69 Mb of nonrandom continuous long regions, the longest of which is about 5.43Mb. These regions have a significant distribution of functional sequences on chromosomes 1, 2, 6, and 16, and most fragments within these regions exhibit high genetic diversity and low overall repetition, and there is a significant enrichment of signals from Pan-Cancer and GWAS studies. To further explore the potential biomedical value of these sequences revealed that genes within this region are involved in immunological and metabolic processes, indicating chromosome-specific functional enrichment. Notably, many phenotypic genetic signals and somatic mutations in tumors are significantly enriched in these region\u003cb\u003es\u003c/b\u003e, which undoubtedly provides new perspectives for subsequent studies. The analysis of putative viral sequences in the human genome revealed that this region contains about two folds the enrichment, including sequences from retroviridae, poxviridae, and herpesviridae viruses, with the longest being the HERV K113 sequence (NC_022518; ~9.47 kb), which is located on several protein-coding genes involved in disease. In exploring the pathogenicity and regulatory roles of these putative viral sequences, we discovered that genes within the putative viral regions are involved in immune regulation, signal transduction and rejection processes, and still interact with external viruses. These interactions are likely to be involved in physiological characteristics and complex disease phenotype regulation, potentially reflecting recent environmental adaptation selection events.\u003c/p\u003e \u003cp\u003eThe Complex_Region represents a comprehensive understanding of \"sequence difference degree\" in the bioinformatic analysis, overcoming the limitations of previous studies that focused only on specific phenotypic regions. Despite our efforts to investigate their origins, there remain shortcomings such as incomplete high-frequency repetitive sequences, insufficient understanding of sequence causation, limited knowledge of sequence regulatory networks, and a lack of genotype-phenotype associations remain. Owing to current limitations with pan-genome data, we cannot completely cover regions related to high-frequency repetitive sequences such as centromeres, which can be directly integrated into our defined intervals in the future\u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. To deepen our understanding of sequence causality, we will continue to collaborate with evolutionary biologists based on the interpretation of environmental viruses. and adopt computational strategies such as pre-trained genome models, to deepen our understanding of their origins using a \"human in the loop\" approach. The complex regulatory networks of living organisms, such as gene regulatory networks and the influence of non-coding sequences on gene expression, which cannot be studied in humans by methods like genetic perturbation, so the Complex_Region should be extended to experimental animal models in the future. Currently, public databases collecting both genotypes and complex phenotypes, particularly for non-European-American populations, are still scarce and more data need to be collected and pre-processed. Therefore, our future work will focus on understanding this phenomenon from more methods and perspectives, expanding this definition to other mammals and building model animals. We will continue to collect and integrate more \"genotype-complex phenotype\" data to validate and explore, and apply the Complex_Region to specific biomedical research, such as the use of CRISPR-Cas9 technology.\u003c/p\u003e \u003cp\u003eIn summary, in our study within the Complex_Region, we attempt to understand the associations between genotype, the external environment, and external phenotypes, fundamentally exploring potential regulatory factors behind population differences in phenotypes. The results can be summarized in the following aspects: 1) The definition and characteristics of the Complex_Region, it was defined by various data sources and levels, exhibiting high sequence diversity entropy, nonrandom distribution, and long continuous segments (\u0026ge;\u0026thinsp;450 kb), which can roughly characterize the complexity of population genetic information. 2) The biomedical significance and functions of the Complex_Region, where gene function enrichment is focused mainly on immunity and metabolism, and there is chromosome-specific functional enrichment. Many genetically related phenotypes and somatic mutations in tumors are significantly enriched in these regions, thus allowing for inexpensive and comprehensive preliminary exploration in cohort studies of these diseases. 3) With respect to the origin and regulation of putative viral sequences in the Complex_Region, there is significant potential enrichment of exogenous viral sequence enrichment, particularly the longest fragment NC_022518-like sequence, which is classified in the HERVK family. This sequence is particularly present in several disease-regulatory coding genes, reflecting high population diversity and conservation across primates. This study further explored the pathogenicity and regulatory roles of these putative viral sequences and their impact on life processes and phenotype shaping. In essence, the Complex_Region contributes to a deeper understanding of human genetic complexity at the population level, further advances precision medicine, and provides a guide for subsequent sequence function studies, but it still needs to be further explored and interpreted.\u003c/p\u003e"},{"header":"Methods","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003eComplex Regions selection by integrating multiple data types\u003c/h2\u003e \u003cp\u003eA range of data sources and types were used to ensure that the complex regions identified were sufficiently representative. The selection process included GFA format screening based on human graphic pangenomes (HPRC and CPC), VCF file screening based on human graphic pangenomes HPRC constructed by Minigraph-Cactus and PGGB respectively\u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e, \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u003c/sup\u003e, and 1000 Genomes VCF file screening. For the screening of GFA files based on human pangenomes, the HPRC and CPC graphic pangenomes constructed via the Minigraph-Cactus process were initially converted to the GFA format using vg convert, and then together with the HPRC graphic pangenomes constructed by the PGGB method, the GFA files were further transformed using odgi build. This was followed by a preliminary screening for complex regions, where we calculated the average mapping depth within 500 bp windows using odgi depth\u003csup\u003e\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e and bedtools across the three pan-genome datasets. After sorting the depth data in ascending order, we removed the top and bottom 20% to avoid outliers, using \u003cem\u003emean\u003c/em\u003e\u0026thinsp;+\u0026thinsp;2*\u003cem\u003eSD\u003c/em\u003e as our filtering threshold for depths within 500 bp windows, and merged the windows within 15 kb using bedtools merge, with the cumulative lengths for each graphical pan-genome separately. When screening the VCF files from the graphical pan-Genome constructed by Minigraph-Cactus and PGGB, we retained loci where the alternative allele frequencies were not less than 0.3, and then counted the corresponding SNPs within 500 bp windows. We then used a similar filtering approach to further obtain the potential regions. For screenings based on the 1000 Genomes VCF files, the filters used are as described above.\u003c/p\u003e \u003cp\u003eWe then combined the potential complex regions obtained by filtering these varied data sources, sorted them by position and coordinates, and merged intervals within 15 kb to form preliminary complex regions. To emphasize longer segments and ensure the broad applicability of identified complex regions in biomedical research, this study retained regions exceeding 450 kb as the final complex regions. To increase their representativeness, the longest sequences from the Y and M chromosomes were also included, creating approximately 248.468 Mb of complex regions. To demonstrate the screening process of the above process, we visualized the average depth (HPRC and CPC) and variant count (1000 Genomes) within 500 bp windows of these complex regions on chromosome 19. This visualization enhanced our understanding of the characteristics of these genomic regions.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eUniqueness sequence characterization within Complex Regions\u003c/h2\u003e \u003cp\u003eTo assess the sequence specificity of the identified complex regions, we first obtained potential non-complex regions within the human genome using bedtools subtract, and finally used bedtools shuffle to finally generate non-complex regions similar to the distribution of complex regions. The corresponding DNA sequences of these regions from the hg38 reference were extracted by coordinates using bedtools getfasta by coordinates, and transformed them into numerical 6-mer feature vectors using CountVectorizer. We standardized the resulting sparse matrix using StandardScaler without subtracting the mean of each element during the process to maintain sparsity, and further used SparsePCA for dimensionality reduction and presentation using the above results.\u003c/p\u003e \u003cp\u003eNext, following the strategy described in a recent publication\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e, we created the Complex_Region dataset and the Random dataset. We then fine-tuned the DNA_bert_6 and human_gpt2-v1 models for binary and multiple classification of DNA sequences using the hugging face interface. In the binary classification task of sequences, the Complex_Region dataset aims to distinguish between complex regions and non-complex regions sequences; the Random dataset is used to classify genomic random region sequences that are similarly distributed to the Complex_Region dataset. In the quaternary classification of sequences, the Complex_Region dataset further distinguishes whether a sequence is located in the coding region or not using the above labels, whereas the Random dataset is used to randomly distinguish genomic sequences.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eSequence characterization of identified Complex Regions\u003c/h2\u003e \u003cp\u003eAfter obtaining the complex regions, we proceeded to quantify the longest segment and cumulative length of each chromosome, and used the CMRG_v1.00_HG002 clinical gene set region as a reference for visual display and statistics. To further delineate the diversity and repetitiveness of sequences within the complex region, we extracted the potential sequences of each interval from the HPRC pan-genome using odgi extract and transformed them into FASTA sequences using odgi paths after optimization and sequencing by odgi sort, respectively. Then the sequences exceeding 10kb were retained using seqkit seq, and the pgr-tk\u0026rsquo;s pgr-pbundle-decomp\u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e was used to obtain the primary components, leading to turn constructed the GFA graph files of the corresponding regions. Additionally, we performed a simple diffusion model for each extracted complex region sequence using the pgrtk.compute_graph_diffusion_entropy function to obtain the sequence diversity entropy and node weights, and used the last 32 elements mean in the weight list as an assessment of sequence repeatability. Finally, we visualized the relationship between the sequence diversity entropy value and its repeatability for all complex regions. Furthermore, by integrating data from UCSC, GENCODE, and pertinent literature, we calculated and displayed the proportions of cumulative sequence lengths in complex regions relative to the entire genome, and displayed them after enrichment rate calculation.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eEnrichment of overlapping gene functions in Complex Regions\u003c/h2\u003e \u003cp\u003eUsing the hg38 reference genome annotations, we utilized bedtools intersect to extract annotations for genes overlapping with complex regions (\u0026ge;\u0026thinsp;150 bp) for further analysis. We then employed the Metascape\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e to enrich and perform disease-related analysis through DisGeNET using the selected protein-coding gene set. Concurrently, we evaluated the functional enrichment and disease characteristics of genes in complex regions and non-complex region from each chromosome, to further verify their functional specificity via Metascape in complex regions.\u003c/p\u003e \u003cp\u003eSubsequently, the genes within the complex regions were categorized into protein-coding (Protein_coding) and all genes (All), with proportions calculated based on the basis of chromosomal distribution against the hg38 reference. Chromosomes were labeled as (chr1, chr16, chr11, chr6, chr19, chr2, others). Functional enrichment and DisGeNET-related analyses were conducted for the protein-coding gene sets on each chromosome using Metascape, with the results displayed according to the new chromosome labels. Furthermore, select immune and metabolic genes within the complex regions were analyzed for gene-disease associations via DisGeNET.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eGenetic phenotypic signals and somatic mutation enrichment in Complex Regions\u003c/h2\u003e \u003cp\u003eWe initially downloaded the GWAS summary files from the UCSC website, identifying 10,376 phenotypes with 403,362 signals across the entire genome, of which 3,975 phenotypes comprising 38,859 signals were located within complex regions. To quantify the prevalence of trait enrichment in complex regions versus the whole genome, we first counted the cumulative Trait_num for all the signal sites associated with the 3,975 phenotypes. Subsequently, we normalized this count by the length of the corresponding region (Len), resulting in Trait_num/Len. For each of the 3,975 phenotypes, we calculated the enrichment ratio (ER) by comparing the Trait_num/Len of the complex regions to that at the genome-wide level. To determine which phenotypes were enriched in the complex region, we performed a hypergeometric distribution test for each phenotype with the following formula: phyper(\u003cem\u003eN\u003c/em\u003e-1, \u003cem\u003eM\u003c/em\u003e, \u003cem\u003eTotal_num\u003c/em\u003e-\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eTotal_case\u003c/em\u003e, lower.tail\u0026thinsp;=\u0026thinsp;F), where \u003cem\u003eN\u003c/em\u003e represents the number of signals in the complex regions (Trait_num1), \u003cem\u003eM\u003c/em\u003e represents those at the genome level (Trait_num2), Total_num is the total 403,362 association signals across the genome, and Total_case is the 38,859 signals within the complex regions. We applied the Benjamini-Hochberg method for \u003cem\u003eP\u003c/em\u003e-value correction, classifying significant phenotypes into three major categories: immune system and cellular biology, metabolism and biochemistry, and health status and other biometrics. Phenotypes with an ER greater than five and an adjusted \u003cem\u003eP\u003c/em\u003e-value less than 0.05 were visualized using EnhancedVolcano plots. Furthermore, GWAS signal loci for diseases like the rheumatoid arthritis (RA) and human immunodeficiency virus (HIV) infection were identified at both the genome-wide and complex region levels to demonstrate the adequately representativeness of complex regions.\u003c/p\u003e \u003cp\u003eUsing a similar computational approach, we analyzed the downloaded pan-cancer data for 33 tumor types, calculating the Trait_num and ER of signals within genome-wide and complex regions. We applied the hypergeometric distribution test phyper(\u003cem\u003eN\u003c/em\u003e-1, \u003cem\u003eM\u003c/em\u003e, \u003cem\u003eTotal_num\u003c/em\u003e-\u003cem\u003eM\u003c/em\u003e, \u003cem\u003eTotal_case\u003c/em\u003e, lower.tail\u0026thinsp;=\u0026thinsp;F), where \u003cem\u003eN\u003c/em\u003e is the tumor signal Trait_num1 in the complex regions, \u003cem\u003eM\u003c/em\u003e being the tumour signal Trait_num2 at the genome level, and Total_num being 3,038,556 tumor association signals at the genome-wide level, and Total_case is 325,537 tumor association signals in the complex region. Next, we used EnhancedVolcano for volcano plot visualization of the enrichment results with the pan-cancer type as the label, and selected gastric adenocarcinoma (STAD) and renal clear cell carcinoma (KIRC) as the most significantly enriched cancers. In addition, the genes with the most mutated types in STAD and KIRC were displayed after statistical analysis by maftools\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e, and the most mutated genes were selected for survival analysis using the RNA sequencing, respectively.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eEnrichment of putative viral sequences within Complex Regions\u003c/h2\u003e \u003cp\u003eTo comprehensively analyze the viral sequences within complex regions, we adopted a dual alignment strategy at both the endogenous retrovirus (HERV) and whole-genome levels. Initially, at the HERV level, we constructed a virus sequence database using makeblastdb, and then extracted the HERV sequences from the reference hg38 using bedtools getfasta, and utilized bedtools intersect to isolate structural variant (SV) sequences within the HERV regions from the 1000 Genomes cohort. Subsequently, we aligned these HERV sequences from hg38 and 1000 Genomes to the viral database using blastn respectively. At the whole-genome level, SVs from the 1000 Genomes data were converted into sequences and merged with the reference hg38, and a corresponding database was constructed using makeblastdb, followed by comparing Virushostdb sequences to this database. For the above comparison results, we filtered and retained sequences longer than 30 bp with over 95% similarity, categorizing them based on the viral family information in the Virushostdb database to obtain a potential viral dataset, including both HERV and other aligned viral sequences. Using this filtered viral dataset, we calculated the non-redundant cumulative alignment lengths for specific regions (hg38, Complex_Region, HERV_Region and HVIDB_Region), and demonstrated the distribution of potential viral alignment lengths on different chromosomes. We also assessed the enrichment ratios of potential aligned viral sequences within a specific region (Complex_Region, HERV_Region, HVIDB_ Region), based on the proportion of viral sequences on different chromosomes of hg38. Finally, we showcased all potential viral alignment sequences within the Complex_Region.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eCharacteristics and species distribution of the NC_022518 Sequence\u003c/h2\u003e \u003cp\u003eTo comprehensively analyze the NC_022518 sequence within the complex regions, we initiated our study by utilizing nucmer to align against each sequence within the reference hg38, and then the corresponding region alignment relationships were shown by show-coords and used as controls\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e. Subsequent extractions of sequences from both the Complex_Region and the Virus regions (including NC_022518), were aligned back to hg38 to obtain chromosomal sequence similarity. The alignment results were then categorized into four major categories: 10\u0026ndash;50 kb, 50\u0026ndash;100 kb, 100\u0026ndash;500 kb, and \u0026ge;\u0026thinsp;500 kb, and further to visualize based on this classification. To explore the characteristics of NC_022518-like sequences in humans, we identified ten protein-coding genes containing this region. We then assessed their functions, pathways, and cell or tissue specificity using the GenDoma database, and explored their corresponding population and primate sequence features using UCSC. To elucidate the relationships between similar sequences among these genes, sequences that highly similar to NC_022518 were extracted with bedtools getfasta, constructed an evolutionary tree using iqtree after mafft multiple sequence comparison\u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e, and finally visualized it.\u003c/p\u003e \u003cp\u003eTo explain the divergence of NC_022518 in primates and humans, we identified its longest protein-coding \u003cem\u003eTRPC6\u003c/em\u003e region (chr11:101695064\u0026ndash;101704528) and a similar fragment within the LncRNA gene ENSG00000286016 (chr12:127153658\u0026ndash;127168021) on the genome hg38, and continue to analyze the sequences within these regions in depth. Firstly, we extracted the sequence fragments in primate, human pan-genome and genome hg38 according to the above coordinates in turn. For sequence extraction in primates, we used minimap2 to align the fragments within ENSG00000286016 and chromosome 12 to the primate genome respectively, and then extracted the corresponding sequences from the aligned regions. For human pan-genome sequence extraction, the above two region sequences in HPRC and CPC were extracted by odgi extract separately, and the corresponding sequences were obtained after odgi sort and odgi paths conversion. For the sequence fragment of genome hg38, we used bedtools getfasta to retain the potential protein coding sequence and the sequence within ENSG00000286016. Immediately following the above steps, we combined the potential sequences together to filter out the sequences with more than 80% of non-A/G/C/T bases, and the filtered sequences were compared by mafft to obtain their multiple sequence alignment results. In order to further portray the degree of similarity between sequences, we used pgrtk.get_shmmr_pairs_from_seq to extract the key information from the multi-sequence alignment results, and combined it with the sample annotation information to perform PCA downscaling and visualization. Meanwhile, we also constructed evolutionary trees using iqtree and ggtree to demonstrate the inter-sequence relationships. Furthermore, we aligned the NC_022518 sequences to the non-redundant nucleic acid database of NCBI, further statistically analyzed the alignment results and constructed the corresponding species sequence trees.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eChromosomal distribution, regulatory networks, and pathogenicity prediction of putative viral regions\u003c/h2\u003e \u003cp\u003eUsing bedtools intersect, we extracted sequences from specified regions (LTR_Repeat, Virus, Immunogenetic, HVIDB) within the complex regions and displayed them after counting them separately by chromosome. To assess the immune characteristics of these sequences, we first used bedtools intersect to count overlapping sequences between endogenous and other aligned viral sequences (Virus), human-virus interaction region sequences (HVIDB), and human immune region sequences (Immunogenetic). We selected the gene overlap among these three groups and used the Metascape website to analyze the protein interaction networks and visualize the core gene modules. For the pathogenicity assessment of specific regions, we first extracted single-base pathogenicity scores predicted by PrimateAI and AlphaMissense\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e, \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e, and then used a bedtools map to calculate the average pathogenicity scores of the desired region (Complex_Region, Virus, Immunogenetic, HVIDB), and presented them at the chromosome level. Moreover, we employed a similar strategy to calculate the pathogenicity scores of overlapping genes and demonstrated the similarity between the pathogenicity scores and the functional enrichment results of these genes.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003ePathogenicity and intrabody regulation of putative viral variants\u003c/h2\u003e \u003cp\u003eTo analyze putative viral mutations within specified genomic regions (Virus, HVIDB, Immunogenetic), we initially employed bedtools intersect to extract small variants and SVs, followed by a chromosomal distribution analysis of these variants. Subsequently, we extracted GWAS association results to delineate the characteristics of virus-related variants. Utilizing the GSEL toolkit\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e, we identified several evolutionary selection indicators in GWAS association signals, including signals for natural selection and adaptivity differentiation (xPEHH; iES; GERP), as well as measures for genetic diversity and population differentiation (Fst) and conservation assessments (PhastCons; phyloP100). Furthermore, we calculated the pathogenicity scores of all identified viral variants and classified them into different pathogenicity score ranges, which included (0,0.3], (0.3,0.5], (0.5,0.7], and (0.7,1]), and then counted and displayed the number of variants in the different ranges. To ensure the reliability of our findings, we repeatedly verified that these variants were located in the HERV region, followed by selecting the variants with scores above 0.7 for conversion to the reference hg19 using liftover\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. We employed the 3DSNP database to perform regulatory analysis on these high-pathogenicity variants. Finally, we describe and discuss the functional regions (Enhancer, Promoter, TFBS, Motif) and intra-population frequencies of these variants, and finely explore the in vivo regulation of some high-frequency variants and their associated diseases.\u003c/p\u003e \u003c/div\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eHERV\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eHuman endogenous retrovirus\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eGWAS\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eWhole-genome association studies\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eDNase_Clusters\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDNase I hypersensitive site regions\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eSDRs\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003estructurally variable regions\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eLTR_Repeat\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003elong terminal repeat retrotransposons\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eHVIDB\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ehuman-virus PPI database zones\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003cb\u003eSV\u003c/b\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003estructural variant\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eGenerative AI in scientific writing\u003c/h2\u003e \u003cp\u003eAI-assisted technologies have been employed to increase readability and improve language proficiency. Nevertheless, the final results were exclusively produced by the authors, who meticulously edited the language to conform to domain terminology. Consequently, we are wholly responsible and accountable for the content of this study.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eDeclarations\u003c/h2\u003e \u003cp\u003e \u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eConsent for publication\u003c/strong\u003e \u003cp\u003eNot applicable.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eCompeting Interest\u003c/strong\u003e \u003cp\u003eNone declared.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e \u003cp\u003eThis work was supported by the Peak Disciplines (Type IV) of Institutions of Higher Learning in Shanghai.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eConceptualization: DD, FZ, LL; Formal Analysis: DD; Funding Acquisition: LL; Investigation: DD, ML; Methodology: DD, ML; Project Administration: FZ, LL; Resources: LL; Supervision: FZ, LL; Useful Suggestions: WZ, FZW, XC; Visualization: DD, CYZ; Writing \u0026ndash; Original Draft: DD, XL; Writing \u0026ndash; Review \u0026amp; Editing: DD, FZ, LL.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eThis work was supported by the Peak Disciplines (Type IV) of Institutions of Higher Learning in Shanghai. We thank all the authors for their hard work. This work was supported by the Medical Research Data Center of Fudan University.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eAll relevant raw data supporting the key findings of this study are available within the article and its Supplementary Information page. All custom analytical pipeline codes used in this work are available at https://github.com/ GeorgeBGM/Complex-genome_analysis.git.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eMomozawa Y, Mizukami K. Unique roles of rare variants in the genetics of complex diseases in humans. J Hum Genet. 2021;66:11\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePinero J, et al. The DisGeNET knowledge platform for disease genomics: 2019 update. Nucleic Acids Res. 2020;48:D845\u0026ndash;55.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao Z, Fu YX, Hewett-Emmett D, Boerwinkle E. Investigating single nucleotide polymorphism (SNP) density in the human genome and its implications for molecular evolution. Gene. 2003;312:207\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCampbell CD, Eichler EE. Properties and rates of germline mutations in humans. Trends Genet. 2013;29:575\u0026ndash;84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMiddleton D, Gonzelez F. The extensive polymorphism of KIR genes. Immunology. 2010;129:8\u0026ndash;19.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNurk S, et al. The complete sequence of a human genome. Science. 2022;376:44\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo Y, Feng X, Li H. Evaluation of haplotype-aware long-read error correction with hifieval. Bioinformatics 39, (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbraham A, LaBella AL, Capra JA, Rokas A. Mosaic patterns of selection in genomic regions associated with diverse human traits. PLoS Genet. 2022;18:e1010494.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBuxton DS, Batten DJ, Crofts JJ, Chuzhanova N. Predicting novel genomic regions linked to genetic disorders using GWAS and chromosome conformation data - a case study of schizophrenia. Sci Rep. 2019;9:17940.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRanz J, Clifton B. Characterization and evolutionary dynamics of complex regions in eukaryotic genomes. Sci China Life Sci. 2019;62:467\u0026ndash;88.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMao Y, et al. Structurally divergent and recurrently mutated regions of primate genomes. Cell. 2024;187:1547\u0026ndash;e15621513.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiao WW, et al. A draft human pangenome reference. Nature. 2023;617:312\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao Y, et al. A pangenome reference of 36 Chinese populations. Nature. 2023;619:112\u0026ndash;21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWagner J, et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat Biotechnol. 2022;40:672\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMonroe JG, et al. Mutation bias reflects natural selection in Arabidopsis thaliana. Nature. 2022;602:101\u0026ndash;5.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVollger MR, et al. Increased mutation and gene conversion within human segmental duplications. Nature. 2023;617:325\u0026ndash;34.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThurman RE, et al. The accessible chromatin landscape of the human genome. Nature. 2012;489:75\u0026ndash;82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen A, Chen D, Chen Y. Advances of DNase-seq for mapping active gene regulatory elements across the genome in animals. Gene. 2018;667:83\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZeller T, et al. Dual checkpoint blockade of CD47 and LILRB1 enhances CD20 antibody-dependent phagocytosis of lymphoma cells by macrophages. Front Immunol. 2022;13:929339.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZarantonello A, Revel M, Grunenwald A, Roumenina LT. C3-dependent effector functions of complement. Immunol Rev. 2023;313:120\u0026ndash;38.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang T, et al. NLRP2 in health and disease. Immunology. 2024;171:170\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFujikura Y, et al. Ketogenic diet containing medium-chain triglyceride ameliorates transcriptome disruption in skeletal muscles of rat models of duchenne muscular dystrophy. Biochem Biophys Rep. 2022;32:101378.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHariharan R, Mousa A, de Courten B. Influence of AMY1A copy number variations on obesity and other cardiometabolic risk factors: A review of the evidence. Obes Rev. 2021;22:e13205.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang J, et al. Comprehensive analysis of the glutathione S-transferase Mu (GSTM) gene family in ovarian cancer identifies prognostic and expression significance. Front Oncol. 2022;12:968547.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChang M, He L, Cai L. An Overview of Genome-Wide Association Studies. Methods Mol Biol. 2018;1754:97\u0026ndash;108.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMihara T, et al. Linking Virus Genomes with Host Taxonomy. Viruses. 2016;8:66.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuarracino A, et al. Recombination between heterologous human acrocentric chromosomes. Nature. 2023;617:335\u0026ndash;43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJern P, Sperber GO, Blomberg J. Use of endogenous retroviral sequences (ERVs) and structural markers for retroviral phylogenetic inference and taxonomy. Retrovirology. 2005;2:50.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang J, Lu X, Zhang W, Liu GH. Endogenous retroviruses in development and health. Trends Microbiol. 2024;32:342\u0026ndash;54.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJo EK. Interplay between host and pathogen: immune defense and beyond. Exp Mol Med. 2019;51:1\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKang SH, Sun YD, Atallah OO, Huguet-Tapia JC, Noble JD, Folimonova SY. A Long Non-Coding RNA of Citrus tristeza virus: Role in the Virus Interplay with the Host Immunity. \u003cem\u003eViruses\u003c/em\u003e 11, (2019).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJakobsson J, Vincendeau M, SnapShot. Human endogenous retroviruses. Cell. 2022;185:400\u0026ndash;e400401.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYonezawa Y, et al. Identification of a Functional Susceptibility Variant for Adolescent Idiopathic Scoliosis that Upregulates Early Growth Response 1 (EGR1)-Mediated UNCX Expression. J Bone Min Res. 2023;38:144\u0026ndash;53.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRadoux-Mergault A, Oberhauser L, Aureli S, Gervasio FL, Stoeber M. Subcellular location defines GPCR signal transduction. Sci Adv. 2023;9:eadf6059.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSewgobind NV, Albers S, Pieters RJ. Functions and Inhibition of Galectin-7, an Emerging Target in Cellular Pathophysiology. \u003cem\u003eBiomolecules\u003c/em\u003e 11, (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen Y, et al. The regulation of DOCK family proteins on T and B cells. J Leukoc Biol. 2021;109:383\u0026ndash;94.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDolzhenko E et al. Characterization and visualization of tandem repeats at genome scale. Nat Biotechnol, (2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHickey G, et al. Pangenome graph construction from genome alignments with Minigraph-Cactus. Nat Biotechnol. 2024;42:663\u0026ndash;73.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarrison E et al. Building pangenome graphs. bioRxiv, (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuarracino A, Heumos S, Nahnsen S, Prins P, Garrison E. ODGI: understanding pangenome graphs. Bioinformatics. 2022;38:3319\u0026ndash;26.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDu D, Zhong F, Liu L. Enhancing Recognition and Interpretation of Functional Phenotypic Sequences through Fine-Tuning Pre-Trained Genomic Models. bioRxiv, 2023.2012.2005.570173 (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChin CS, et al. Multiscale analysis of pangenomes enables improved representation of genomic diversity for repetitive and clinically relevant genes. Nat Methods. 2023;20:1213\u0026ndash;21.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhou Y, et al. Metascape provides a biologist-oriented resource for the analysis of systems-level datasets. Nat Commun. 2019;10:1523.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMayakonda A, Lin DC, Assenov Y, Plass C, Koeffler HP. Maftools: efficient and comprehensive analysis of somatic variants in cancer. Genome Res. 2018;28:1747\u0026ndash;56.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRiva G, Mauri M. MuMMER: How Robotics Can Reboot Social Interaction and Customer Engagement in Shops and Malls. Cyberpsychol Behav Soc Netw. 2021;24:210\u0026ndash;1.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen LT, Schmidt HA, von Haeseler A, Minh BQ. IQ-TREE: a fast and effective stochastic algorithm for estimating maximum-likelihood phylogenies. Mol Biol Evol. 2015;32:268\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKatoh K, Rozewicki J, Yamada KD. MAFFT online service: multiple sequence alignment, interactive sequence choice and visualization. Brief Bioinform. 2019;20:1160\u0026ndash;6.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao H, et al. The landscape of tolerated genetic variation in humans and primates. Science. 2023;380:eabn8153.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng J, et al. Accurate proteome-wide missense variant effect prediction with AlphaMissense. Science. 2023;381:eadg7492.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbraham A, Labella AL, Benton ML, Rokas A, Capra JA. GSEL: a fast, flexible python package for detecting signatures of diverse evolutionary forces on genomic regions. Bioinformatics 39, (2023).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePark KJ, Yoon YA, Park JH. Evaluation of Liftover Tools for the Conversion of Genome Reference Consortium Human Build 37 to Build 38 Using ClinVar Variants. Genes (Basel) 14, (2023).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Complex Regions, Genotypes, Phenotypes, Immunity and Metabolism, Pan-genome, HERV","lastPublishedDoi":"10.21203/rs.3.rs-4800291/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4800291/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground: \u003c/strong\u003eNuclear genomic DNA plays a crucial role in individual development and phenotype determination. The genetic landscape within populations exhibits significant heterogeneity, contributing to diverse human traits. Current studies of human genome heterogeneity often focus on specific segments of high-frequency phenotype-associated sequences or structurally complex regions. Therefore, to overcome the limitations of previous studies and more directly explore population heterogeneity, it is essential to study the entire genome rather than focusing only on known phenotype-associated regions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e Using set theory, we have clearly defined Complex Regions (Complex_Region) by integrating pan-genome datasets, covering about 8.1% of the human genome. These regions exhibit high sequence diversity and nonrandom long continuous fragments (≥450kb), thus reflecting population genetic complexity. Our enrichment analysis revealed that genes within Complex_Region are primarily involved in immunity and metabolism, indicating chromosome-specific functional enrichment. Notably, immune genes are mainly located on chromosomes 6 and 19, which are closely associated with disease occurrence. Moreover, these regions are enriched for human phenotype-related signals and tumor somatic mutations, providing novel insights for large-scale cohort studies. We also detected ancient viral sequences, particularly ~9.47 kb human endogenous retroviruses (HERV) insertion sequence NC_022518, which is diverse in humans but remains conserved across primates, to be implicated in regulating bodily functions and various diseases.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions:\u003c/strong\u003e Our study highlights the biomedical importance of Complex_Region by revealing associations among genotypes, environment, and phenotypes. This enhances our understanding of life regulation and phenotype shaping, highlighting the role of these regions in immunity, metabolism, and disease association.\u003c/p\u003e","manuscriptTitle":"Deciphering Complex Regions within the Human Genome and Unraveling Their Critical Biomedical Regulatory Functions","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-08-07 13:31:33","doi":"10.21203/rs.3.rs-4800291/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"80c08058-2d7c-4f4d-b408-e5387e8f58f9","owner":[],"postedDate":"August 7th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-08-07T13:31:34+00:00","versionOfRecord":[],"versionCreatedAt":"2024-08-07 13:31:33","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4800291","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4800291","identity":"rs-4800291","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.