{"paper_id":"aed5c60e-86a6-4959-bc84-4c12dd50f093","body_text":"Sequencing the gaps: dark genomic regions persist in CHM13 despite long-read \nadvances \nMark E. Wadsworth1,2,3, Madeline L. Page1,2,3, Bernardo Aguzzoli Heberle1,2, Justin B. Miller1,3,4,5, Cody \nSteely1,3, Mark T. W. Ebbert1,2,3* \n \n1Sanders-Brown Center on Aging, University of Kentucky, Lexington, KY; 2Department of Neuroscience, \nCollege of Medicine, University of Kentucky, Lexington, KY; 3Division of Biomedical Informatics, Internal \nMedicine, College of Medicine, University of Kentucky, Lexington, KY; 4Department of Pathology and \nLaboratory Medicine, University of Kentucky, Lexington, KY, USA; 5Microbiology, Immunology and Molecular \nGenetics, College of Medicine, University of Kentucky, Lexington, KY , USA \n \n \n*To whom correspondence should be addressed: Mark T. W. Ebbert (mark.ebbert@uky.edu)   \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nAbstract \n \nComprehensive genomic analysis is essential for advancing our understanding of human genetics and \ndisease. However, short-read sequencing technologies are inherently limited in their ability to resolve highly \nrepetitive, structurally complex, and low-mappability genomic regions, previously coined as \"dark\" regions. \nLong-read sequencing technologies, such as PacBio and Oxford Nanopore Technologies (ONT), offer \nimproved resolution of these regions, yet they are not perfect. With the advent of the new Telomere-to-\nTelomere (T2T) CHM13 reference genome, exploring its effect on dark regions is prudent. In this study, we \nsystematically analyze dark regions across four human genome references—HG19, HG38 (with and without \nalternate contigs), and CHM13—using both short- and long-read sequencing data. We found that dark regions \nincrease as the reference becomes more complete, especially dark-by-MAPQ regions, but that long-read \nsequencing significantly reduces the number of dark regions in the genome, particularly within gene bodies. \nHowever, we identify potential alignment challenges in long-read data, such as centromeric regions. These \nfindings highlight the importance of both reference genome selection and sequencing technology choice in \nachieving a truly comprehensive genomic analysis.\n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nIntroduction \n \nThe ultimate goal of human genomics research is to improve disease diagnostics and treatments. To this end, \nperforming a comprehensive and personalized analysis on an individual’s genome, transcriptome, and \nepigenome, in combination with other factors (e.g., environmental factors) is essential. Specifically, a \n“complete” analysis at the DNA level should, at minimum, capture all DNA variants (small and structural) that \nthe individual carries, and interpret and predict the downstream implications of these variants. Unfortunately, \nmany gaps remain, but major technological and scientific advances in the past decade have provided \nsignificant gains toward that goal. Specifically, short-read sequencing became widely accessible in the early \n2010’s, propelling our understanding of the human genome and transcriptome to a level that was difficult to \nimagine only a decade before. While short-read sequencing has been a major boon to genomics research, \nmany genomic regions remained inaccessible (i.e., dark or camouflaged) and structural variants could not be \naccurately resolved because of the inherent limitations of short read lengths [1–3].  \n \nIn 2019, we systematically characterized the long-known, but poorly appreciated issue surrounding the “dark” \ngenome—regions of the genome that cannot be accurately resolved and are thus entirely overlooked [1]. \nSpecifically, using short-read sequencing we characterized two basic forms of “dark” regions, including: (1) \nregions that are dark-by-depth, where few or no sequencing reads are present in the alignment; and (2) dark-\nby-mapping-quality (dark-by-MAPQ) where the region contains aligned reads, but the reads do not align \nuniquely to that region. In both cases, any variants within the region are completely overlooked using standard \nshort-read sequencing and downstream analyses because of ambiguous alignments [1–3]. We further \ndemonstrated the breadth of this issue and how it affects genes already known to be involved in human \ndisease. We termed genes containing dark-by-MAPQ regions because of either full or partial genomic \nduplications as “camouflaged” genes [1]. In all, we identified >6000 gene bodies where some portion of the \ngene’s sequence is “dark” using standard short-read sequencing approaches, and 2128 were ≥ 5% dark [1]. \nDark and camouflaged genes make it difficult (if not impossible) to perform a complete analysis of an \nindividual’s genome, including for both small and large DNA variants [1]. We also showed that long-read \nsequencing resolves most dark and camouflaged regions overlooked by short-read sequencing [1]. More \nrecent work has shown that long reads also characterize and quantify individual RNA isoform expression [4–7], \nultimately bringing us closer to the reality of a comprehensive and personalized genomic analysis. \n \nA complete and accurate representation of the human genome is essential to understanding its complexity, but \nthe human genome has only recently been completed. Leveraging long-read sequencing, the Telomere-2-\nTelomere (T2T) consortium assembled the first complete human genome sequence from a completely \nhomozygous cell line (CHM13) in 2022 [8, 9]. The new T2T CHM13 reference genome added ~200 \nmegabases to the previously most complete reference genome (GRCh38), predominantly in telomeric, \ncentromeric, and acrocentric chromosomal regions [8]. \n \nThough tempting to assume the human genome is finally “complete”, as the efforts and latest results from the \nHuman Pangenome Reference Consortium demonstrate, the human genome is still far from “complete” \nbecause no single reference can accurately represent all of humanity [10]; individuals have their own unique \ncombination of not only single-nucleotide variants, but also larger structural DNA variations [11–14]. Thus, \nbeing able to perform a complete analysis of an individual’s genome (whether by de novo assembly or \nalignment to a reference genome) is essential to understanding their genetic predisposition for various \nphenotypes, including those involved in human health and disease [11–13, 15]. Unfortunately, we are still a \nlong way from achieving the goal to perform a truly complete analysis of an individual’s genome, especially \nsince most clinical and research still rely on standard short-read sequencing approaches and pipelines that are \nknown to overlook critical regions of an individual’s genome [1].  \n \nTo per\nform a comprehensive analysis, we first need a “comprehensive” reference genome to serve as a \nbaseline, for which CHM13 and the pangenome are the beginning. Counterintuitively, however, the more \ncomplete the human reference genome becomes, the more challenging it becomes to properly analyze and \ninterpret an individual’s unique combination of DNA variants because of genomic duplications and ambiguous \nalignments. In this challenging area of research, we address three important knowledge gaps, herein: (1) how \nthe interaction between sequencing platform choice and reference genome, especially CHM13, affects our \nability to assess the “dark” and “camouflaged” regions of the genome; (2) how dark and camouflaged regions \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\naffect related short-read sequencing assays beyond standard DNA sequencing (e.g., ChIP-Seq, bisulfite \nsequencing, etc.) and further prevent a complete genomic analysis using short-read data; and (3) the \nimportance of supplementary alignments in long-read datasets. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nResults \n \nTo better understand how the intersection of sequencing length/platform and reference genome affects our \nability to accurately resolve dark and camouflaged regions of the genome, we compared dark regions from \nlong-read (PacBio and Oxford Nanopore Technologies) and short-read sequencing data (Illumina 100bp reads \nand 250bp reads; Fig. 1a) across four different human genome references. We compared reference genomes \nrepresenting a continuum of genome completeness: (1) HG19; (2) HG38 (excluding alternate contigs); (3) \nHG38 (including alternate contigs); and (4) the Telomere-to-Telomere (T2T) CHM13 v2.0 (Fig. 1b). Using the \nsame 20 short-read samples from our original paper (ten 100bp Illumina and ten 250bp Illumina) [1], and ten \nsamples sequenced on both PacBio and ONT (Fig. 1a,b), we aligned each sample to the four reference \ngenomes (excluding secondary and supplementary alignments, per our 2019 analysis) and identified dark and \ncamouflaged regions using the Dark Region Finder (Fig. 1c) [1]. For clarity, secondary reads are those that \nmap equally well in multiple places in the genome while supplementary reads are chimeric reads, mapping \ndifferent regions of the same read independently. We then quantified dark-by-depth and dark-by-MAPQ regions \nacross platforms and references genome-wide. Building upon our previous work, we performed a comparison \nof dark regions between HG38 and T2T’s CHM13 (Fig. 1di) [8, 16, 17]. As part of our analyses, we also \nassessed the effect of dark and camouflage regions in other short-read based sequencing assays, including \nepigenetic assays like ChIP-Seq and bisulfite sequencing (Fig. 1dii). \n \nIn our 2019 work, we excluded secondary and supplementary alignments for simplicity. Here, we demonstrate \nthat excluding supplementary reads falsely inflates estimates for dark-by-depth regions in long-read data, and \nthus demonstrates the importance of supplementary reads. We discovered this because of a region that was \ndark-by-depth in long reads but resolved in short reads when excluding supplementary reads. Upon deeper \ninvestigation, we found that including supplementary reads resolved this issue.  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nFigure 1: Dark region analytical workflow. (a) We obtained 20 samples sequenced using Illumina short-reads (100bp and 250bp) and 10 \nsamples sequenced using PacBio and ONT. (b) Using bwa and minimap2 for short and long reads, respectively, we aligned all samples to \nHG19, HG38 (with and without alternate contigs), and CHM13. (c) We then identified dark and camouflaged regions using our Dark Region \nFinder. (d) i. Finally, we characterized and quantified dark regions genome wide, along with their differences between references and \nsequencing platforms. ii. We further assessed how these dark and camouflaged regions affect other short-read based DNA assays (i.e., \nepigenetic assays). iii. We identified important long-read alignment challenges that need to be addressed. Created with BioRender.com. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nNumber of dark regions increase in CHM13 \n \nIn our previous work, we characterized and quantified dark and camouflaged regions across the genome using \nshort-read sequencing technologies and assessed how well long reads resolved these regions, when using \nprimary alignments only [1]. Here, we build on that work by including CHM13. After aligning the samples to \neach of the four reference genomes, we classified dark regions into three main classes: (1) dark-by-depth \nregions, which have <5x coverage; (2) dark-by-MAPQ regions, in which 90% of the reads covering the region \nhave a mapping quality (MAPQ) less than 10; and (3) camouflaged regions, which are a subset of dark-by-\nMAPQ regions that are also 98% identical to another genomic region, as calculated by BLAT [18]. \n \nWe observed several important patterns when comparing genome-wide dark-by-MAPQ bases for each \nplatform across the four reference genomes. As the reference genome became more complete, between HG19 \nand CHM13, the number of dark-by-MAPQ bases increased. Specifically, there was >6x increase in dark-by-\nMAPQ bases in CHM13 compared to HG19 for Illumuna100 reads (Fig. 2a; Table 1). In CHM13, there are \n3.3x more dark-by-MAPQ bases in short-read data compared to long-reads (Fig. 2a; Table 1) [23,24,35]. \nSpecifically, Illumina100 reads had 206,689,498 dark-by-MAPQ bases when aligned to CHM13, while PacBio \nand ONT had 78,273,227 (resolved 62.1%) and 62,833,689 (resolved 69.6%), respectively. Upon initial \ninspection, the number of dark-by-depth bases appears to be enriched in long-read technologies (Fig. 2b), but \nwe show later that this enrichment is artificially inflated in ONT because we omitted supplementary reads in our \n2019 work and in most analyses in this work. When comparing within gene bodies (i.e. not genome wide), \nPacBio and ONT had 6,071,156 and 1,802,133 dark-by-MAPQ bases, respectively, showing that within gene \nbodies, ONT resolved 3.4 times more bases than PacBio (Fig. 2c,d).  \n \n \n \nTable 1: Dark region statistics reveal long-read dark regions are fewer in number but span larger regions. For all four sequencing \ntechnologies and all four reference genomes, we compared the number of dark bases, regions, average bases per region, and median number of \nbases per region for dark-by-depth, dark-by-MAPQ, and the combination of those two. We found that short-read sequencing data has more, smaller \ndark regions, while long-read sequencing has fewer, larger dark regions. \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n  \n \nFigure 2: Dark-by-MAPQ dark bases predominate. (a) As the genome becomes more complete, short-read data exhibits increasing dark-by-\nMAPQ nucleotides while long reads plateau. Green and pink arrows show general trends for Illumina100 and PacBio, respectively. (b) While millions \nof bases are obscured by dark-by-depth regions, they pale in comparison to dark-by-MAPQ (c) ONT has the least number of dark bases within \ngenes in all types of dark regions (i), dark-by-depth (ii), dark-by-MAPQ (iii), and camouflaged regions (iv) across all references. HG38 with \nalternates has the highest amount of dark bases. (d) ONT has the least number of dark bases within protein coding genes in all types of dark regions \n(i), dark-by-depth (ii), dark-by-MAPQ (iii), and camouflaged regions (iv) across all references. HG38 with alternates has the worst amount of dark \nbases.  \n \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nONT outperforms each technology for resolving dark regions within gene bodies, including in the \nCHM13 reference genome in primary alignments \n \nAfter comparing the dark bases genome-wide, we specifically compared dark bases within gene bodies, which \nwas a major focus of our previous work using primary alignments, only [1]. Here, we found that HG38 with \nalternate contigs had 2-3x more total gene body dark bases (either type) for each technology, compared to \nCHM13, except ONT which was almost even (Illumina100: 1.97x; Illumina250: 2.25x; PacBio: 3.27x; ONT: \n1.15x; Fig. 2ci). PacBio also had between 2.4x (CHM13) and 6.7x (HG38 Alt) more total gene body dark bases \nthan ONT across all four reference genomes. We additionally observed that HG38 with alternate contigs has \ngreater total gene-body dark-by-MAPQ bases in PacBio data (20.4 Mb) compared to HG38 without alternate \ncontigs (2.8 Mb) and CHM13 (6.1 Mb), respectively, demonstrating the alignment challenges that the alternate \ncontigs cause (Fig. 2c). One reason this may affect PacBio more than ONT is that PacBio has shorter read \nlengths than ONT (PacBio median read length, averaged across samples: 15,577bp; standard deviation: \n1,495; ONT median read length, averaged across samples: 24,798bp; standard deviation: 9,769; \nSupplementary Fig. S1e,f). This trend also held when subsetting the genes to only protein coding genes (Fig. \n2d), showing that CHM13 provides both a complete genome and avoids over representing gene regions. \nSignificant work is still needed for CHM13 to be ready for widespread use, however. Specifically, the \nbackground work already done on previous genome assemblies, including allele frequencies and gene \nannotations, is not easily converted to a new reference genome (accurately). \n \nAfter comparing the total gene body dark bases, we assessed the number of genes containing any dark region \nwith at least 20 contiguous nucleotides. Exactly 8,077 genes contain dark regions (of either type) in short-read \ndata aligned to CHM13, whereas ONT only had 286 (Illumina100: 8,077; Illumina250: 6,377; PacBio: 990; \nONT: 286; Fig. 3a). Of the 8,077 genes containing dark regions, 6,059 contained dark-by-MAPQ (Illumina100: \n6,059; Illumina250: 4,255; PacBio: 824; ONT: 180; Fig. 3b) and 3,709 genes contained dark-by-depth regions \nin CHM13 (Illumina100: 3,709; Illumina250: 3,372; PacBio: 319; ONT: 107; Fig. 3c). For clarity, the sum of \ndark-by-depth and dark-by-MAPQ gene counts add up to 9,768 for the 100-bp Illumina data (not 8,077) \nbecause 1,354 genes contain both dark-by-depth and dark-by-MAPQ regions within the gene body \n(Illumina100: 1,354; Illumina250: 907; PacBio: 57; ONT: 4; Fig. 3d).  \n \nE\nxamples of genes that suffer from both forms of dark regions include CTAG1A, CTAG1B, and IL3RA. \nCTAG1A and CTAG1B (both originally known as NY-ESO-1 before the duplication was discovered) are located \nin a duplicated region of chromosome X causing two approximately 35kb dark-by-MAPQ regions (Fig. 3e.i) \nand are primarily expressed in specific cell types within the testes and in several types of cancer (e.g. breast \ncancer, leukemia, etc.) making it a potential therapeutic target [19–23]. Problematically, CTAG1A and CTAG1B \nare 97.9% dark (dark-by-depth: 23.5%, dark-by-MAPQ: 74.4%; Fig. 3e.ii) and 99.6% dark (dark-by-depth: \n26.9%, dark-by-MAPQ: 72.8%; Fig. 3e.iii) in Illumina100 sequencing aligned to CHM13, respectively. CTAG1A \ncontains a dark-by-depth region in the first intron for both Illumina100 and Illumina250. CTAG1B also contains \na dark-by-depth region in the intron which breaks up the dark-by-MAPQ regions that span the rest of the gene \nin Illumina100 and Illumina250 samples. PacBio sequencing resolved most of this large, duplicated region on \nchromosome X, while ONT resolved the entire region (Fig. 3e.i). While only ONT was able to fully resolve the \nentire region, PacBio had better coverage in CTAG1A and always has better single molecule per-base \naccuracy. CTAG1A and CTAG1B are perfect examples of genes and regions that prevent researchers from \nperforming a truly complete genomic analysis, because short reads cannot be accurately disentangled across \nthe region. Long-read data properly illuminate these dark regions. \n \nIL3RA (a.k.a. CD123) is a subunit of a cytokine receptor with two copies, with one located on the X and Y \nchromosomes, each. It is a top therapeutic target for acute myeloid leukemia and blastic plasmacytoid dendritic \ncell neoplasm because it is known to have higher expression in leukemic cells [24–27]. IL3RA has a \nfascinating and complicated short-read alignment structure, oscillating from non-dark to dark-by-MAPQ or \ndark-by-depth (Fig. 3f). In total, it is 41.5% dark with a combined 25.8% of the gene being dark-by-MAPQ and \n15.7% of the gene being dark-by-depth. The sporadic dark regions across IL3RA using short-read data prevent \na complete analysis of this gene, which is believed to be key to treating leukemic cells. PacBio has decreased \ncoverage at the ends of IL3RA but better single molecule per-base accuracy, whereas ONT had more \nconsistent, deep coverage across the gene. By leveraging long-read data, a complete analysis of this gene is \npossible and could provide additional clues for treating disease.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nLooking at gene bodies that are at least 5% dark (of either type), we observed 3,874 when aligning Illumina100 \nreads to CHM13, which is nearly double the number of gene bodies that were 5% dark when aligned to HG38 \nwithout alternate contigs (2140), and nearly four and fourteen times more than PacBio (976) and ONT (283) \nwhen aligned to CHM13, respectively (Fig. 3g). Exactly 3,754 (96.9%) of the dark genes contained dark-by-\nMAPQ regions in CHM13 (Fig. 3h). ONT resolved 95.2% (1 −\n180\n3754) of the Illumina100 dark-by-MAPQ genes, \nwhile Illumina250 only resolves 14.4% and PacBio resolves 78.3% in CHM13 (Fig. 3h). When stratifying by \ndark-by-depth (Fig. 3i) or gene bodies that contain both types (Fig. 3j), HG38 with alternate contigs stands out \nwith the greatest number of gene bodies that are at least 5% dark in short-read data, as expected; we \nexpected HG38 with alternate contigs to have more dark gene regions because the alternate contigs introduce \nadditional gene copies. The trends remain similar when comparing genes that are 100% camouflaged (Fig. \nS2a-d). Long-read sequencing technologies outperform the short-reads in all categories, and ONT outperforms \nPacBio. \n \nTo understand the difference between references, we assessed the number of camouflaged genes that are at \nleast 5% camouflaged. We also assessed only those that are 100% camouflaged, specifically. For reference, \ncamouflaged genes are dark-by-MAPQ regions that fall within a gene body and have at least 98% similarity to \nanother genomic region (i.e., a subset of dark-by-MAPQ). Exactly 2,813 genes have more than 5% of the gene \nbody camouflaged in Illumina100 data when aligned to CHM13 (Illumina100: 2,813; Illumina250: 2,330; \nPacBio: 509; ONT: 119; Fig. 3k) and 512 gene bodies were 100% camouflaged (Illumina100: 512; Illumina250: \n455; PacBio: 201; ONT: 100; Fig. 3l). The 512 gene bodies that were 100% camouflaged with Illumina100 \nreads aligned to CHM13 is nearly double than when aligned to HG38 without alternate contigs (275). These \ngene bodies were predominantly classified pseudogene, protein coding, and lincRNA biotypes in those that are \ngreater than 5% and 100% camouflaged (Fig. S2e,f).  \n \nONT can further elucidate the role of large variants in dark reads where even PacBio struggles. An example is \nthe intron of TMEM88B, which contains a dark-by-depth region flanked by dark-by-MAPQ regions in the \nIllumina100 data (49.5% dark-by-depth; Fig. S2g). The long-read data decreases the size of the dark region, \nbut only ONT unequivocally resolved the region, showing that this subject likely has a heterozygous deletion \nwithin the TMEM88B intron. Little is known about TMEM88B, though it appears to be a negative regulator of \nWNT signaling, particularly in the heart [28, 29]. Thus, the effect of this intragenic deletion is undetermined, but \nworthy of future research. Many more ONT reads properly map to the general region (46; Fig. S2g) compared \nto PacBio (10) despite similar overall sequencing depth, and many more ONT reads spanned the entire region. \nThe heterozygous deletion is indicated by approximately half of the ONT reads mapping through the entire \nregion containing a deletion, while the other half map through the region without a deletion. Thus, long reads \nsuggest that the short-read alignment challenges in the TMEM88B intron are likely due to a large intronic \ndeletion. TMEM88B highlights the utility of long-read data for identifying relatively large deletions and therefore \nfurther the quest for a truly complete analysis by shedding light on the dark regions of the genome. \n \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nFigure 3. The T2T-CHM13 genome contained more dark and camouflaged genes than HG38; ONT outperformed other platforms. (a) 8,077 \ngene bodies contained dark regions in CHM13 for Illumina100 while ONT samples only had 286 dark gene bodies. (b) 6,059 gene bodies \ncontained dark-by-MAPQ regions in CHM13 for Illumina100 while ONT samples only had 180 dark-by-MAPQ gene bodies. (c) 3,709 gene bodies \ncontained dark-by-depth regions in CHM13 for Illumina100 while ONT samples only had 107 dark-by-MAPQ gene bodies. (d) 1,354 gene bodies \ncontained both dark-by-depth and dark-by-MAPQ regions in CHM13 for Illumina100 while ONT samples only had 4 gene bodies containing both \ndark-by-depth and dark-by-MAPQ regions. (e) CTAG1A/B contained both dark-by-depth and dark-by-MAPQ within a larger dark region. (f) IL3RA is \n41.5% dark (in Illumina100) made up of both types of dark regions. (g) CHM13 had almost double the number of at least 5% dark genes compared \nto HG38 without alternate contigs. (h) While CHM13 had far more genes with at least 5% dark-by-MAPQ, ONT resolved 95%. (i) Genes that were \nat least 5% dark-by-depth were an order of magnitude less of an issue (see Y-axis) than dark-by-MAPQ, but ONT still performed the best across \nthe board. (j) HG38 with alternate contigs had more gene bodies with at least 5% dark by both types. (k) CHM13 generally had the most genes \nwith at least 5% camouflaged, the vast majority of which were resolved with long reads. (l) CHM13 generally had the most 100% camouflaged \ngenes.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nDark regions from short reads obscure short-read DNA epigenetic assay results \n \nAfter quantifying genome-wide dark and camouflaged regions, we anticipated these dark regions would likely \nbe consistent across the spectrum of genomic DNA assays, including those related to epigenetics. Specifically, \nwe looked at bisulfite sequencing, chromatin immunoprecipitation (ChIP) sequencing, and high-throughput \nchromosome conformation capture (HiC) sequencing. Whole-genome bisulfite sequencing is a common assay \nto assess the CpG methylation status genome wide by converting unmethylated cytosines to uracils (Fig. 4a). \nChIP-Seq, on the other hand, is an assay designed to identify DNA binding locations for proteins to a specific \nplace in the genome (Fig. 4b). HiC data is an epigenetic assay that identifies regions of the genome that \nphysically interact and is used to identify boundary regions within which an enhancer or other distal regulatory \nelement could interact with a given promoter called Topologically Associated Domains (TADs; Fig. 4c) [30]. All \nthree assays historically rely on short-read sequencing and are thus likely susceptible to issues related to dark \nand camouflaged regions. To test the effect dark regions have on epigenetic assay results, we identified dark \nbases and camouflaged genes in each assay and found that epigenetic assays are equally susceptible to \nchallenges related to dark regions. \n \nWe compared the difference in total dark bases, dark-by-depth bases, and dark-by-MAPQ bases with whole-\ngenome bisulfite sequencing, ChIP-seq, and the previously analyzed Illumina100 dark data. ChIP-seq assays \nwill only include DNA bound to the queried proteins (e.g., H3K4me3 histone marks), thus we expect most of \nthe genome to be dark-by-depth, by nature of the assay. We leveraged frontal cortex (BA9) H3K4me3 \npromoter associated histone mark data [21, 31] for two individuals we obtained from the Encode Project. \nH3K4me3 dysfunction is linked to several neurological conditions including Alzheimer’s disease [32], Autism \n[33, 34], and Schizophrenia [35]. As expected, we found the ChIP data had far more dark bases than both \nIllumina100 and whole-genome bisulfite sequencing. For those regions where coverage is expected, however, \nhigh-quality alignments are essential. Of the approximately 35 megabases (35,180,776 bp) where ChIP-seq \nsequencing data were obtained, 1.2 megabases (3.4%) of ChIP data were dark-by-MAPQ (Fig. 4d). Thus, \neven those regions where sequencing data are expected, short-read ChIP-seq results leave a large proportion \ninaccessible because they are dark-by-MAPQ. \n \nTo assess the effect of dark regions on whole-genome bisulfite sequencing data, we leveraged publicly \navailable data from the Encode Project from male adrenal gland tissue, where the sequencing read length was \n150 bp. Even knowing that bisulfite treatment is extremely harsh on DNA, we observed a surprisingly large \nnumber of dark-by-depth bases, resulting in an 8.7x increase (98,331,392 bp) compared to Illumina100 whole-\ngenome sequencing (Fig. 4d). Such a large number of bases without sufficient coverage precludes the ability \nto assess DNA methylation. While the increase in dark-by-depth is concerning, it is not necessarily surprising \nbecause harsh bisulfite treatment shears the DNA, and biases are well known [36–38]. Finally, about 202 \nmegabases (202,350,000 bp) of the whole-genome bisulfite sequencing data are dark-by-MAPQ. We conclude \nthat short-read based epigenetic assays are dramatically affected by dark regions, predominantly dark-by-\ndepth bases (Fig. 4d).  \n \nH\nere, we highlight an example camouflaged gene for each of the epigenetic assays. First, we investigated the \nwhole-genome bisulfite sequencing camouflaged genes. We plotted the per base genomic total coverage and \nhigh-quality coverage (MAPQ ≥ 10) for the AMY1C gene, which is 49.6% camouflaged in Illumina100 data and \nwhose copy number relates to glucose absorption [39, 40] (Fig. 4e). The promoter proximal region of AMY1C \nhas sufficient total depth to assess CpG methylation (Fig. 4e top); however, except for one small region of the \ngene, the region is dark-by-MAPQ (Fig. 4e bottom), and therefore is unrepresented in the processed data. \nUsing short-read bisulfite sequencing data, it is not possible to assess the epigenetic regulation of the AMY1C \ngene by DNA methylation using standard analysis pipelines. An awareness of dark-by-MAPQ regions in genes \nis prudent when assessing CpG methylation. In our original paper [1], we developed a method to rescue and \nanalyze reads from dark-by-MAPQ regions, but we maintain that this approach is simply a stopgap and long \nreads are the ultimate solution—especially since long-read technologies can directly measure genome-wide \nDNA methylation during standard sequencing without special DNA preparation. \n \nSecondly, we highlighted several genes in ChIP-Seq data that are camouflaged by each other, including the \nhistone genes H2AC18 and H2AC19 (100% and 99.8% camouflaged respectively in Illumina100; Fig. 4f), and \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nthe heat shock proteins HSPA1A and HSPA1B (41.4% and 39.5% camouflaged respectively in Illumina100; \nFig. S3a). For both sets of genes, we saw large peaks in the total coverage in the promoter proximal region of \nthese genes; however, the peaks in the promoter proximal region are truncated or completely obscured by \nlarge dark-by-MAPQ regions. While little is known about the differential expression of H2AC18 and H2AC19, \nthey are differentially expressed upon exposure to a Human Papillomavirus oncoprotein [41] therefore an \nunderstanding of how these genes are epigenetically regulated has potential medical importance. HSPA1A and \nHSPA1B are two subunits of the HSP70 complex with potential therapeutic relevance to ALS [42–44].  When \ncamouflaged regions obscure epigenetic results for potentially medically relevant genes such as H2AC18/19 \nand HSPA1A/B, a complete understanding of gene regulation and a comprehensive analysis is inhibited. \n \nFinally, HiC data is also plagued by gaps due to dark regions because it relies on short-read sequencing. For \nexample, the Alzheimer’s disease-associated gene, CR1, has a tandem domain duplication that camouflages a \nlarge segment of the gene [1]. This region is almost completely dark in HG38 HiC data (Fig. 4g). Meng et al. \nidentified an overlap of Alzheimer’s disease associated GWAS SNPs with HiC interactions and eQTLs [45], but \ntheir results would likely overlook any association within CR1 because the region is camouflaged. The use of \nTADs through the HiC assay in Alzheimer’s disease would leave the analysis blind to any physical interactions \nor GWAS SNPs with this region of CR1 and it would increase the complexity of TAD boundary identification.  \n \nUltimately, a complete analysis requires the integration of multiple assays to fully understand genomic \nprocesses. The loss of valuable information from these epigenetic assays because of short-read sequencing \ninhibits a comprehensive genomic analysis. The inability to quantify DNA methylation, histone modifications, \nand genomic looping causes a gap in our understanding of gene regulation and overall genomic structure of \nthese regions. In addition to largely resolving camouflaged regions of the genome through long-read \nsequencing, PacBio and ONT have the added benefit of being able to identify DNA methylation simultaneously.  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nFigure 4: Dark regions obscure epigenetic assay results. (a) Whole genome bisulfite sequencing is a common genome-wide short-read \nsequencing assay that identifies CpG methylation sites by converting unmethylated cytosines to uracils. (b) ChIP-Seq is an assay that pulls down \nDNA bound to a target protein. (c) HiC data identifies the location of chromatin loops in nuclear DNA. (d) Compared to 100 bp Illumina data, the \nwhole-genome bisulfite sequencing samples have more dark-by-depth and less dark-by-MAPQ. Most of the ChIP data, as expected, are dark-by-\ndepth; however, 1.2 Mb (3.4%) are dark-by-MAPQ. (e) Due to close paralogs few, if any, CpG sites in AMY1C can be quantified using short-read \nsequencing because AMY1C is dark-by-MAPQ and camouflaged by AMY1A and AMY1B. (f) H2AC18/19 camouflage each other and, as a result, \nthe promoter associated histone mark H3K4me3 peaks are obscured and left unanalyzed. (g) The CR1 tandem domain repeats completely obscure \nDNA looping data in HiC data. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nCHM13 results in decreased dark bases in CR1 resulting in better identification of CR1 major allele.  \n \nInsertions and deletions can modify copy number which can influence disease development for a range of \ndiseases, including Alzheimer’s disease (APP duplication) [46], Parkinson’s disease (SNCA duplication) [47, \n48], and various heart diseases [49, 50]. Previous work has even suggested that copy number of the CR1 \nC3B/C4B binding domain is associated with Alzheimer’s disease risk [51]. The CR1 C3B/C4B binding domain \nis known to be variable in the population [52]. Specifically, CR1 is known to have at least four primary \nhaplotypes known as CR1-A, CR1-B, CR1-C, and CR1-D [51], where all four haplotypes are believed to \noriginate from varying combinations of three highly similar low-copy repeats (LCR) known as LCR1, LCR1’, \nand LCR2 that make up the C3B/C4B binding domain and plays an important role in the complement cascade.  \n \nAt the protein level, all three LCRs are identical except for a single amino acid difference in LCR1 (Ala405Thr). \nThus, treating these three repeats as the same fundamental domain, HG38 contains three total copies of the \nbinding domain [1], whereas CHM13 contains only two [52] (Fig. 5a,b). Yang et al. recently demonstrated that \nthe CR1 allele represented in CHM13 (two copies) is the major allele found in 79 of their 94 samples (84%), \nwhile HG38 contains the minor allele (three copies) [52]. The ramifications of the reference genome used when \nanalyzing sequencing data for this region are important. Because the binding domain is repeated, it is also \ndark-by-MAPQ (and camouflaged) in short-read sequencing data, as we demonstrated previously [1]. When \nanalyzing Illumina100 data for this region using HG38, CR1 is 22.7% camouflaged, whereas it is only 3.1% \ncamouflaged when using CHM13. This dramatic difference is not only because one of the copies is missing, \nbut because the CR1-A allele contains the two most distinct LCRs (LCR1 and LCR2) [53–55]. Analyzing short-\nread data in CHM13 makes it possible to more easily identify variants that would be obscured by HG38 but will \nalso result in more false variants being reported for individuals carrying the CR1-B allele because of true \ndifferences between the three LCR domains at the nucleotide level. This shows that the reference genome you \nuse matters when assessing regions with variable copy numbers in the population in short read data. Notably, \nboth PacBio and ONT maintain coverage through this region regardless of the reference genome used. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nFigure 5: CR1 tandem C3B/C4B domain repeat number differs between HG38 and CHM13 with important ramifications for short-read \nalignments while long reads perform well on both. CR1 haplotypes are made up of varying numbers of low-complexity repeat regions (LCRs). (a) \nHG38 contains the CR1-B haplotype with three copies of the C3B/C4B binding site domain representing the minor allele where we see widespread \ndark-by-MAPQ (camouflaged) regions with short-read data. CR1 is 22.49% camouflaged in HG38 (including introns). Both PacBio and ONT resolve \nthis region well. (b) CHM13 contains the CR1-A haplotype with two copies of the C3B/C4B binding domain. The amount of dark-by-MAPQ \n(camouflaged) bases is dramatically fewer, where CR1 is 3.07% camouflaged in CHM13.  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nCentromeric satellite variability on chromosome 10 drives poor alignments for both short- and long-\nread data in CHM13. \n \nWe were curious about the nature of genomic regions where short-read data exhibit both dark-by-MAPQ and \ndark-by-depth behavior. We identified a large region in CHM13 near the chromosome 10 centromere that \nexhibited this behavior (Fig. 6a). We found that the region is highly repetitive with multiple repeat types, \nincluding human satellite II (HSATII) and Beta satellite sequences (Fig. 6a). Most of this region is dark-by-\nMAPQ, which is easily explained because of ambiguous alignments from sequence duplication. We anticipate \nthat the dark-by-depth regions could only occur for two reasons: (1) there simply are no reads from that region \nfor the individual (e.g., genuine genomic deletions, sequencing artifacts, etc.); or (2) because of an alignment \nissue. \n \nIn addition to the co-occurring dark-by-MAPQ and dark-by-depth regions within short-read data, we noticed \nthat the long-read alignments were also problematic. Part of what makes this example particularly striking is \nthe clear contrast between the PacBio and ONT alignments. For PacBio, a large portion of this region is well \nrepresented with quality alignments but reaches a sudden dark-by-depth region that is not observed in ONT \ndata (Fig. 6a). On the other hand, while ONT maintains better coverage throughout the region, there is a \nsudden drop in alignment quality not observed in the PacBio data that coincides with a known boundary \nbetween a Beta satellite and HSATII (Fig. 6a). The drop in ONT alignment quality is shown by the sudden \nincrease in mismatches in the histogram above the ONT reads (Fig. 6a).  \n \nTo better understand this phenomenon, we collected all ONT reads that aligned to this region \n(chr10:42,530,255-42,588,899) for this sample (HG00096) and analyzed them individually using \nRepeatMasker after converting all the reads to the same forward direction. Based on the RepeatMasker \nresults, we were surprised to find three different haplotypes in this region: (1) the first consisted solely of \nHSATII sequence, matching the reference genome (57/80 reads; 71.25%); (2) the second contained the \nannotated HSATII followed by a combination of beta satellites, LSAU, and composite beta/LSAU repeats, and \ntandem repeats (HSATII/BSat/LSAU/TR; 11/80; 13.75%); and (3) the third was an inverted version of the \nsecond (12/80; 15%; Fig. 6b). The composite beta/LSAU repeats were identified during the T2T studies [16], \nand this region appears to include families 1,4, and 10. Observing three haplotypes is entirely unexpected and \nsuggests there may be mosaicism occurring in this region. This finding deserves to be followed up in future \nstudies. Notably, Altemose et al. recently reported that satellites make up 6.2% of the CHM13 reference \ngenome, and specifically discussed the chromosome 10 centromere as being highly variable because of \nstructural variation [9]. We also noticed that one of the two known copies of the DUX4 gene is located in this \nregion. A repeat retraction in both the chromosome 4 [56] and 10 [57] copies are known to cause \nfacioscapulohumeral muscular dystrophy (FSHD). Looking at alignments within a broader region, we see a \ndecrease in general coverage over this centromeric region and an increase in dark regions and insertions and \ndeletions in long-read sequencing pointing to incomplete alignments likely due to the satellite rearrangements \n(Fig. 6c). In all, these results suggest that centromeric satellite variability drives alignment challenges for both \nshort- and long-read data—perhaps because using a static reference genome imposes the reference \ngenome’s structure on the individuals. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nFigure 6: Centromeric/pericentromeric satellites cause problems in alignments. (a) The pericentromeric region exhibits major alignment \nchallenges with varying degrees of dark-by-MAPQ and dark-by-depth regions. (b) Using Repeat Masker, we identified multiple types of repetitive \nelements rather than just the annotated HSATII. (c) When zooming out we see that ONT has more deletions and insertions, combined with a large \nsection with nearly complete sequence variability that still aligns (indicated by high mismatch rate in histogram above the ONT reads). On the other \nhand, PacBio has large dark-by-depth regions. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nSupplementary alignments resolve some dark regions in long-read data.  \n \nDuring our analyses, we identified dark regions in long-read data aligned to CHM13 overlapping known \nstructural variants that were properly resolved with short-read data. One example is on the q-arm of \nchromosome 8 (chr8:112,191,464-112,192,351), corresponding to a documented inversion in CHM13 \ncompared to HG38 with the Cactus [58] alignment annotations found on the UCSC Genome Browser [59] (Fig. \n7a). We identified this region because of a stark, consistent dark-by-depth region in all long-read samples \naligned to CHM13 (using only primary alignments). This region of 887 bases is completely dark and ends \nabruptly (Fig. 7b). In this situation, the reads were long enough to span such a short region (887 bases), thus it \nis unlikely that the long-read data itself was the problem. Short-read sequencing data clearly identified the \ninversion where paired-end reads are oriented in the same direction (highlighted reads; Fig. 7b) with a break in \nthe coverage at the breakpoints of the inverted region. The reads for all ONT and PacBio samples were soft \nclipped to exactly match the same boundaries. Thousands of bases were soft-clipped per read in PacBio, and \ntens of thousands of bases were soft-clipped per read in ONT. Notably, this gap was not present in HG38 \nalignments (Fig. 7c). \n \nIn our original paper [1], for simplicity we excluded all secondary and supplementary reads when assessing \ndark regions, therefore including only primary reads. In these analyses, we found that the break in alignments \nstems from our exclusion of supplementary alignments. Minimap2 uses supplementary alignments to span the \nbreaks in the case of inversions. To test this, we re-analyzed the samples including the supplementary \nalignments and, in this case, including supplementary reads resolved this dark-by-depth region (Fig. 7d). To \nassess the overall effect of using supplementary reads, we determined to compare the differences between the \nprimary only results and the primary and supplementary results.  \n \n  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n  \nFigure 7: Inversions cause decreased coverage in long read alignments. (a) Chromosome 8 contains a region of 887 bp that is consistently \ndark-by-depth in all PacBio and ONT samples where minimap2 soft-clipped the reads and Cactus alignments show an inversion between HG38 and \nCHM13. (b) In CHM13, short-read alignments clearly identify the inversion, while primary long-read alignments (removing supplementary alignments) \nsoft-clip large segments of reads. (c) In contrast to the CHM13 alignments, this region is not inverted in HG38 and is completely covered in all \nalignments. (d) Supplementary reads resolve dark region caused by only considering primary reads in the inverted region. \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nWhen including supplementary alignments, we found an expected increase in dark-by-MAPQ bases in short-\nread sequencing data. (Fig. S4a-f; Table 2; Fig. 8a). The number of dark-by-MAPQ bases across the genome \nonly slightly increased in HG38 with and without alternates for every sequencing platform except ONT which \nexhibited a decrease of almost two megabases when including supplementary reads in the analysis (1,685,814 \nbases in HG38 no alternate contigs; 1,695,547 bases in HG38 with alternate contigs). In ONT specifically, \nsupplementary alignments consistently resulted in fewer dark-by-MAPQ bases across all references, though \nsurprisingly, PacBio had a significant reduction in CHM13, compared to ONT, with 4 and 1 Mb decreases, \nrespectively (Fig. 8a; Table 2). Contrastingly, Illumina100 had almost a 2 Mb (1,911,880 base) increase in \ndark-by-MAPQ bases in CHM13 (Fig 8a). Including supplementary reads consistently reduced dark-by-depth \nregions for all platforms, especially for ONT aligned to CHM13 which decreased by 15.5 Mb (Fig. 8b; Table 2).  \n \nThese differences in dark regions resulted in all sequencing platforms, except Illumina100, resolving dark \nregions within genes (Fig. 8c). We found that, across all platforms, dark-by-MAPQ had increased or no real \nchange in the number of genes (Fig. 8d), while all platforms had a decrease in dark-by-depth genes when \nincluding supplementary reads (Fig. 8e). When looking at all dark genes, we observed an increase in \nIllumina100, while the others decreased. (Fig. 8c). All references across all the sequencing platforms exhibit a \ndecrease in genes that are at least 5% dark-by-depth (Fig. 8e). Finally, only Illumina100 has an increase in \ngenes that are at least 5% camouflaged when you add in supplementary reads (Fig. 8f). We plotted the \nlocation of the dark-by-depth regions in ONT sequencing that are resolved when including supplementary \nreads, revealing the biggest groupings are on the q-arms of chromosomes 1 and 9 (Fig. 8g).  \n \nThese findings suggest that including supplementary reads results in marginal differences in short-read data \nwhile large, potentially significant differences, especially in structural variant identification, are found in long-\nread data, particularly ONT. This is likely because the increased read lengths often allow for spanning \nbreakpoints that create chimeric reads.  \n \n \nTable 2: Dark bases and regions with supplementary read inclusion. When comparing this table to Table 1, dark regions and bases are mostly \ndecreased across the board.  \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n  \nFigure 8: Adding supplementary reads to the process mainly effects long-read sequencing analyses. For each of the figures a-f a negative \nvalue means that adding in the supplementary reads resolved the dark regions, while a positive value means the dark regions are exacerbated. For \nfigures c-f all genes were filtered based on having at least 5% of the gene being dark or camouflaged. (a) When comparing primary only alignments \nto primary with supplementary alignments included, ONT, in HG38 and CHM13, and PacBio, in CHM13, exhibit a large number of bases that are no \nlonger dark-by-MAPQ. Illumina100 increased dark-by-MAPQ bases by almost 2 Mb. (b) ONT resolved the most dark-by-depth bases compared to all \nthe other platforms across all references. (c) Including supplementary reads results in an increase in genes that are at least 5% dark in Illumina100 \nsequencing, and less for all other sequencing types. (d) Illumina100 has issues with increased numbers of at least 5% dark -by-MAPQ genes, while \nthe other platforms change very little. (e) All platforms across all references have a decrease of genes that are at least 5% dark-by-depth. (f) Only \nIllumina100 saw an increase in genes that are at least 5% camouflaged with including supplementary reads. (g) We plotted the location of all of the \nONT dark-by-depth regions resolved by including supplementary reads in the analysis.  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n \nDiscussion: \n \nBuilding off our previous work, in this study, we found that dark-by-MAPQ bases are the predominant dark \nregions across the spectrum of references—especially CHM13. Long-read platforms outperform  short-read \nplatforms in resolving dark regions, where ONT remains the most effective, including in gene bodies. \nAdditionally, all short-read assay results are dramatically affected by dark regions, including epigenetic assays. \nStill, much work remains to achieve a comprehensive analysis, because even long-read technologies struggle \nto properly resolve highly variable regions when constrained to the structure of a static reference genome, \nparticularly the highly variable centromeric satellites on chr10.  \n \nWe also demonstrate the importance of including supplementary alignments with long-read data to properly \nresolve structural variants; the longer reads become the more necessary aligning reads chimerically becomes \nwhen aligning to a static reference genome. Finally, we also demonstrate that the reference genome can have \nimportant implications using CR1 as a case example. From our analysis, we conclude that in many cases \nCHM13 and long-reads are likely the best choice, but in cases like addressing the minor allele of CR1, CHM13 \nmay not be the best choice. \n \nNo single reference genome will be perfect. Based on certain statistics, HG19 may appear ideal since here we \nshow that it has overall fewer dark regions than the other reference genomes—but it has fewer dark regions \nbecause it is incomplete. The more complete the human reference genome becomes, the more challenging it \nis to properly analyze and interpret an individual’s unique combination of DNA variants. These challenges are \noften due to genomic duplications and ambiguous alignments leading to dark and camouflaged regions, among \nother complications. Understanding what contributes to these challenges and how we can overcome them is \nimperative.  \n \nOur study further demonstrates the importance of long reads to achieve a comprehensive analysis but also \nshows that more work remains. The biggest limitation and future direction to this analysis is that we did not \ninclude the human pangenome project. The pangenome is the first reference genome that begins to account \nfor population variability rather than attempting to represent all of humanity using a single, static reference \ngenome. Therefore, the next logical step is to assess how dark regions impact the pangenome. \n \nUltimately, many challenges remain to realizing a comprehensive and personalized genomic analysis, including \nhaving a deeper understanding of the human genome itself, which is limited by imperfect sequencing and \ndownstream analyses. Even a relatively basic comparison between the various reference genomes and \nrespective sequencing technologies becomes challenging. We demonstrated a range of differences between \nthe issues related to short- and long-read sequencing, but even important differences between long-read \ntechnologies and their tendencies. There remain significant gaps in our understanding of the human genome \nthat limit our ability to study it. The problem becomes significantly harder when trying to account for genomic \ndiversity across the population. Having the first truly complete human reference genome is an important step \nforward, but significant work remains.  \n \n  \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nMethods \n \nSample Selection \nFor our dark region analyses, we selected 10 samples for each type of data. We used the same short-read \nsamples as we used in our 2019 analysis. For 100 bp read length Illumina data, we selected 10 male hispanic \nnon-related samples from the Alzheimer’s Disease Sequencing Project (ADSP) whole-genome sequencing \n(WGS) dataset: A-CUHS-CU000208-BL-COL-56227BL1, A-CUHS-CU000406-BL-COL-52870BL1, A-CUHS-\nCU000779-BL-COL-31428BL1, A-CUHS-CU001010-BL-COL-52679BL1, A-CUHS-CU002031-BL-COL-\n25771BL1, A-CUHS-CU002707-BL-COL-40848BL1, A-CUHS-CU002997-BL-COL-47280BL1, A-CUHS-\nCU003023-BL-COL-47464BL1, A-CUHS-CU003090-BL-COL-47998BL1, and A-CUHS-CU003128-BL-COL-\n49696BL1. For Illumina250 samples, we used 10 male samples from the Thousand Genomes Project: 4 SAS \n(HG01583, HG03006, HG03742, NA20845), 3 AMR (HG01112, HG01051, HG01565), 2 EUR (HG00096, \nHG01500), 1 AFR (HG01879). Finally, we selected our long-read (PacBio and ONT) data from the 1000 \nGenomes Project, HGSVC3 for the following 10 male samples: 4 AFR (HG01890, HG02666, NA19317, \nNA19347), 3 EAS (HG01596, NA18534, NA18989), 3 EUR (HG00096, HG00268, HG00358). We specifically \nselected all male samples to ensure that we include the Y chromosome. Additionally, we specifically chose \ndiverse population backgrounds to minimize bias to any specific population. We aligned the short-read data \nusing bwa-mem [60] and the long-read with minimap2 [61, 62]. All samples were originally aligned to GRCh38 \nand then aligned to the four different references and have at least 30x average coverage (Supplementary Fig. \nS1a-d).  \n \nReference Genomes \nWe used the NCBI version of HG19 with no alternate contigs \n(\nhttps://ftp.ncbi.nlm.nih.gov/genomes/archive/old_genbank/Eukaryotes/vertebrates_mammals/Homo_sapiens/\nGRCh37.p13/seqs_for_alignment_pipelines/GCA_000001405.14_GRCh37.p13_no_alt_analysis_set.fna.gz) \nand HG38 without alternate contigs  \n(\nhttps://ftp.ncbi.nlm.nih.gov/genomes/all/GCA/000/001/405/GCA_000001405.15_GRCh38/seqs_for_alignment\n_pipelines.ucsc_ids/GCA_000001405.15_GRCh38_no_alt_analysis_set.fna.gz). We used the 1000 Genomes \nversion from 2015 of HG38 with alternate contigs \n(https://github.com/igsr/1000Genomes_data_indexes/blob/master/data_collections/1000_genomes_project/RE\nADME.1000genomes.GRCh38DH.alignment). Finally, we used version 2.0 of CHM13 (https://s3-us-west-\n2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/analysis_set/chm13v2.0.fa.gz). \n \nEpigenetic Data \nWe pulled WGBS data from (https://www.encodeproject.org/experiments/ENCSR042LOG/) Encode for adrenal \ngland tissue – (ENCFF384LDT, ENCFF288SYU). We then trimmed reads with TrimGalore (trim_galore --\npaired $R1 $R2) [63]  and aligned them to CHM13 with Bismark (bismark $GENOME -1 $R1 -2 $R2 -\n-parallel 2 --ambig_bam  --unmapped --ambiguous) [64]. We used --ambiguous because \nBismark removes low MAPQ reads by default, but this flag forces it to write out the ambiguous reads. We then \nmerged the ambiguous reads with the confident reads. We aligned ChIP-Seq data for AD and MCI patients \n(ENCLB308YWZ, ENCLB142GGP H3K4me3 89 AD; ENCLB305RHN, ENCLB555ZFE H3K4me3 90 MCI) from \nEncode using bowtie2 to HG38 (bowtie2 -k 10 -q -x $Ref -p 20 -U $Trimmed_fq -S \n$Sample_sam) [65] after trimming the reads with TrimGalore (trim_galore $file) [63]. Both assays were \nfiltered using samtools view for the annotated genes of interest, then we calculated the coverage using \nbedtools genomecov [66]. We then processed the output using our Epigenetics Rmarkdown script \n(https://github.com/UK-SBCoA-\nEbbertLab/DarkRegionCamoPaperFigures/blob/main/EpigeneticsFigs/EpigeneticCamoRegions.Rmd).  \n \nThe HiC data was plotted and assessed using the UCSC Genome Browser (HESC HiC \nhttps://genome.ucsc.edu/cgi-\nbin/hgTables?db=hg38&hgta_group=regulation&hgta_track=hicAndMicroC&hgta_table=h1hescInsitu&hgta_do\nSchema=describe+table+schema). The Genome Browser only had the data aligned and analyzed in HG38, \nwhich is why we used that reference rather than CHM13 for the figure. \n \nAll software for this section was run within a singularity container hosted by Sylabs that can be pulled with the \nfollowing command: singularity pull --arch amd64 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nlibrary://mewadsworth/dark_region_followup/epigenetics_software_2023_02_22.sif:sh\na256.d9ccbfbb1905e73e09a9bc3237fb5cc70cbe1b1be782e4e52573c0446f40c508. \n \nDarkRegionFinder (DRF) \nOur original paper used a series of bash scripts to run the Dark Region Finder. We converted and updated the \nscripts to a nextflow pipeline [67] that can be found on Github (https://github.com/UK-SBCoA-\nEbbertLab/Dark_and_Camouflaged_Genes_Pipeline/tree/master). This runs in conjunction with a singularity \ncontainer hosted by Sylabs that can be pulled with the following command: singularity pull --arch amd64 \nlibrary://mewadsworth/dark_region_followup/rescue_camo_variants_2024_10_26.sif:sha256.e9d459cc13d8afa\n980900a2981ff84c3072939fabeba38207b6cef034dc91a62. We converted our bash scripts to a self-contained \nautomated Nextflow pipeline. We also provide a web application that creates basic plots such as those in \nFigures 2, 3, and S2. It also allows the user to filter based on the types of genes and percentage camouflaged.  \n \nThe Dark and Camouflaged Region identifier is made up of six steps as described in our 2019 work [1]. Briefly, \nwe realigned the samples (for both short- and long-read data) to each respective reference. For this work, we \nincluded HG19, HG38 with and without alternate contigs, and CHM13. We used bwa-mem (version 0.7.17-\nr1188) to align short-read data and minimap2 (version 2.26) for the long-read data. In the original pipeline, we \nused minimap2’s map-pb option for both ONT and PacBio data because it performed better [61, 62]. The \ncurrent pipeline uses an updated version of minimap2 and we used the map-ont for ONT data. The second \nstep of the pipline is to run the Dark Region Finder (DRF) (https://github.com/mebbert/DarkRegionFinder\n) to \nidentify dark-by-depth regions (regions of less than 5x coverage) and dark-by-MapQ regions (90% of reads \nhave a mapping quality of less than 10). Step three combines the output from the individual samples. Step four \nprepares the gene annotation bed from the gene annotation GTF file. The fifth step creates the final output bed \nfiles for the different types of regions. These bed files are publicly available (https://github.com/UK-SBCoA-\nEbbertLab/DRF_PaperApp_V2/tree/main/data) and leveraged for the aforementioned WebApp. Finally, the \nsixth step, which creates a masked genome, is used to rescue variants from camouflaged genomic regions but \nwas not used in this study. \n \nIncluding supplementary alignments \nTo analyze the dark regions including supplementary alignments, we removed the -M option used in BWA. The \n-M option marks all supplementary reads determined as small by the algorithm as secondary alignments. \nSince it would skew our results if we didn’t include all supplementary alignments, we removed that option to \nkeep those reads as supplementary. We additionally added an option to exclude only secondary alignments in \nhtslib used in DRF. \n \nDue to massive increases in depth in some areas when including supplementary reads in our dark region \nanalyses, we limited the number of reads loaded at a time to 10,000 reads in HTSJDK \n(setMaxReadsToAccumulatePerLocus(10000)). This allowed us to run the DRF with less than 500 gigabytes of \nmemory. Additionally, we removed the intervals for the run with supplementary reads, because if a \nsupplementary read was in a different interval from its primary alignment, HTSJDK failed. The edited pipeline is \nlocated on the following branch: \nhttps://github.com/UK-SBCoA-\nEbbertLab/Dark_and_Camouflaged_Genes_Pipeline/tree/nextflow-supplementary-pipeline.  \n \nWeb APP & Rscripts \nWe developed our web application using plotly’s dash app written in R \n(https://ebbertlab.com/dark_region_comparison.html). The app is split into three tabs: (a) Primary Alignments \nOnly, (b) Primary + Supplementary Alignments, and (c) Comparison. The Primary Alignments Only tab contains \nplots and tables calculated from the output from the Dark Region Finder app run excluding secondary and \nsupplementary reads. The Primary + Supplementary Alignments tab contains plots and tables calculated from \nthe output from the Dark Region Finder app run excluding only secondary and including primary and \nsupplementary reads. The Comparison tab contains plots and tables created with data from the previous two \ntabs to compare the results. Most of the plots in this paper were created in and taken from the app. The \npackages and versions used to deploy the app are listed at the bottom of the page. \n \nOur web app is hosted on Heroku. It uses the heroku-24 stack, an R buildpack \n(https://github.com/virtualstaticvoid/heroku-buildpack-r\n) running version 4.4.2, and a basic dyno. Additional \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\ninstallation information can be found in the init.R file on the app's github page (https://github.com/UK-SBCoA-\nEbbertLab/DRF_PaperApp/blob/main/init.R). We use the Github deployment method which allows us to deploy \nthe app directly from github.  \n \nIGV screenshots \nThe different IGV figures were generated using IGV version 2.11.9 [68]. We chose a representative sample \nfrom each platform type: Illumina100 - A-CUHS-CU000208-BL-COL-56227BL1, Illumina250 - HG00096, ONT - \nHG00096, and PacBio – HG00096. Of note we had three sequencing platforms for the 1KG sample HG00096 \n(Illumin250, PacBio, and ONT). Therefore, when looking at the IGV screen shots the bottom three alignments \nare for the same sample. We limit the number of insertions and deletions seen by only allowing IGV to show \ninsertions and deletions that are at least 1000 base pairs.  \n \nCR1 LCR placement  \nFor figure 5, we used data from Brouwers et al. to place the CR1 LCRs using comparison between the two \nisoforms with and without LCR1’ [53]. We corroborated the location based on the original papers that identified \nthe repeat in the late 1980’s [54, 69] as well as the structure of the C3B binding domain [55].  \n \nRepeat Analysis \n \nWe extracted all the reads that were aligned to the area of the beta-satellite and the HSATII using samtools. \nWe then put all of them on the same strand by extracting all reads aligned to the negative strand (samtools \nview -f 16) and performed a reverse complement using python and then appended them to the forward strand \nreads (samtools view -F 20). We used RepeatMasker (version 4.1.5) [70] to analyze the repeat structure of the \nhighly variable region on Chromosome 10. First, we analyzed the reads that were completely contained (i.e., \ndid not exceed the boundaries) within the repetitive region (chr10:42,530,255-42,588,899) using the “-s” flag in \nRepeatMasker [70] to perform a more sensitive search within the curated Dfam library (version 3.7). These \nresults were manually analyzed to identify the repeats in each read within the region. This process was then \nrepeated with all the reads that overlapped with this region to determine if these reads also supported the \nrepeat structure that we identified. RepeatMasker [70] identified the HSATII repeat as alternating “HSATII” and \n“(CATTC)n simple repeats”, which has been noted previously (Altemose, et al. supplementary materials) [9].   \n \nAnnotations \nWe used the CHM13 repetitive element annotation bed file and intersected it with our dark region bed files \nusing bedtools intersect (\nhttps://s3-us-west-2.amazonaws.com/human-\npangenomics/T2T/CHM13/assemblies/annotation/chm13v2.0_RepeatMasker_4.1.2p1.2022Apr14.bed). \nAdditionally, for our centromeric satellite annotation comparison from CenSat for CHM13 (https://s3-us-west-\n2.amazonaws.com/human-pangenomics/T2T/CHM13/assemblies/annotation/chm13v2.0_censat_v2.0.bed ). \nTo compare our dark regions against regions that are annotated as unique in CHM13 we leveraged the UCSC \nGenome Browser CHM13 Unique annotation (https://genome.ucsc.edu/cgi-\nbin/hgTrackUi?hgsid=1725076402_LoDEV1XMCzGoxZTAnVOGyqEAkw4U&db=hub_3671779_hs1&c=chr5&\ng=hub_3671779_hgUnique ). Finally, to compare the differences between HG38 and CHM13 we compared the \nCactus alignment from UCSC Genome Browser (https://genome.ucsc.edu/cgi-\nbin/hgTables?db=hub_3671779_hs1&hgta_group=compGeno&hgta_track=hub_3671779_cactus&hgta_table=\nhub_3671779_snakeHg38&hgta_doSchema=describe+table+schema). \n \nGene Annotations Used \nWe used the following gene annotations for each reference. \n• HG19 - \nhttps://ftp.ensembl.org/pub/grch37/release-\n107/gff3/homo_sapiens/Homo_sapiens.GRCh37.87.chr.gff3.gz \n• HG38 with or without alternates - https://ftp.ensembl.org/pub/release-\n107/gff3/homo_sapiens/Homo_sapiens.GRCh38.107.chr.gff3.gz  \n• CHM13 - https://s3-us-west-2.amazonaws.com/human-\npangenomics/T2T/CHM13/assemblies/annotation/chm13.draft_v2.0.gene_annotation.gff3 \n \n \n \nContributions \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nMW and ME developed and designed the study. MW, ME, and MP wrote the manuscript. MW designed and \ndeveloped the website. MW and MP developed the nextflow pipeline with assistance from BAH. MW, MP, and \nCS performed all analyses. CS analyzed the chr10 satellite region and advised on the writing of that section. \nMW developed the DashBio app, and MP embedded the DashBio app into ebbertlab.com. JM provided \nimportant intellectual contributions. \n \nCompeting interests \nThe authors report no competing interests. \n \nData availability \nResults from our analyses can be found at the following link: https://github.com/UK-SBCoA-\nEbbertLab/DRF_PaperApp_V2/tree/main. Main results data are available in the data directory in the Github \nrepository. The subfolders labelled “Updated_output_01_17_2025” contain the primary + supplementary \nalignments. The others contain the primary only alignments. \n \nFunding \nThis work was supported by the National Institutes of Health [R35GM138636, R01AG068331 to M.E.], the \nBrightFocus Foundation [A2020161S to M.E.], Alzheimer’s Association [2019-AARG- 644082 to M.E.], PhRMA \nFoundation [RSGTMT17 to M.E. and predoctoral fellowship to B.A.H]. \n \nAcknowledgments \nWe appreciate the contributions of the Sanders-Brown Center on Aging at the University of Kentucky. We \nwould like to thank the University of Kentucky Center for Computational Sciences and Information Technology \nServices Research Computing for their support and use of the Morgan Compute Cluster and associated \nresearch computing resources.  \n \nWe further appreciate and acknowledge data used by both the Alzheimer’s Disease Sequencing Project \n(ADSP) and the 1000 Genomes Project, and especially the participants of these studies who made this \nresearch possible. \n \nThe Alzheimer’s Disease Sequencing Project (ADSP) is comprised of two Alzheimer’s Disease (AD) genetics \nconsortia and three National Human Genome Research Institute (NHGRI) funded Large Scale Sequencing and \nAnalysis Centers (LSAC). The two AD genetics consortia are the Alzheimer’s Disease Genetics Consortium \n(ADGC) funded by NIA (U01 AG032984), and the Cohorts for Heart and Aging Research in Genomic \nEpidemiology (CHARGE) funded by NIA (R01 AG033193), the National Heart, Lung, and Blood Institute \n(NHLBI), other National Institute of Health (NIH) institutes and other foreign governmental and non-\ngovernmental organizations. The Discovery Phase analysis of sequence data is supported through \nUF1AG047133 (to Drs. Schellenberg, Farrer, Pericak-Vance, Mayeux, and Haines); U01AG049505 to Dr. \nSeshadri; U01AG049506 to Dr. Boerwinkle; U01AG049507 to Dr. Wijsman; and U01AG049508 to Dr. Goate \nand the Discovery Extension Phase analysis is supported through U01AG052411 to Dr. Goate, U01AG052410 \nto Dr. Pericak-Vance and U01 AG052409 to Drs. Seshadri and Fornage. \n \nSequencing for the Follow Up Study (FUS) is supported through U01AG057659 (to Drs. PericakVance, \nMayeux, and Vardarajan) and U01AG062943 (to Drs. Pericak-Vance and Mayeux). Data generation and \nharmonization in the Follow-up Phase is supported by U54AG052427 (to Drs. Schellenberg and Wang). The \nFUS Phase analysis of sequence data is supported through U01AG058589 (to Drs. Destefano, Boerwinkle, De \nJager, Fornage, Seshadri, and Wijsman), U01AG058654 (to Drs. Haines, Bush, Farrer, Martin, and Pericak-\nVance), U01AG058635 (to Dr. Goate), RF1AG058066 (to Drs. Haines, Pericak-Vance, and Scott), \nRF1AG057519 (to Drs. Farrer and Jun), R01AG048927 (to Dr. Farrer), and RF1AG054074 (to Drs. Pericak-\nVance and Beecham). \nThe ADGC cohorts include: Adult Changes in Thought (ACT) (U01 AG006781, U19 AG066567), the \nAlzheimer’s Disease Research Centers (ADRC) (P30 AG062429, P30 AG066468, P30 AG062421, P30 \nAG066509, P30 AG066514, P30 AG066530, P30 AG066507, P30 AG066444, P30 AG066518, P30 \nAG066512, P30 AG066462, P30 AG072979, P30 AG072972, P30 AG072976, P30 AG072975, P30 \nAG072978, P30 AG072977, P30 AG066519, P30 AG062677, P30 AG079280, P30 AG062422, P30 AG066511, \nP30 AG072946, P30 AG062715, P30 AG072973, P30 AG066506, P30 AG066508, P30 AG066515, P30 \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nAG072947, P30 AG072931, P30 AG066546, P20 AG068024, P20 AG068053, P20 AG068077, P20 \nAG068082, P30 AG072958, P30 AG072959), the Chicago Health and Aging Project (CHAP) (R01 AG11101, \nRC4 AG039085, K23 AG030944), Indiana Memory and Aging Study (IMAS) (R01 AG019771), Indianapolis \nIbadan (R01 AG009956, P30 AG010133), the Memory and Aging Project (MAP) ( R01 AG17917), Mayo Clinic \n(MAYO) (R01 AG032990, U01 AG046139, R01 NS080820, RF1 AG051504, P50 AG016574), Mayo \nParkinson’s Disease controls (NS039764, NS071674, 5RC2HG005605), University of Miami (R01 AG027944, \nR01 AG028786, R01 AG019085, IIRG09133827, A2011048), the Multi-Institutional Research in Alzheimer’s \nGenetic Epidemiology Study (MIRAGE) (R01 AG09029, R01 AG025259), the National Centralized Repository \nfor Alzheimer’s Disease and Related Dementias (NCRAD) (U24 AG021886), the National Institute on Aging \nLate Onset Alzheimer’s Disease Family Study (NIA- LOAD) (U24 AG056270), the Religious Orders Study \n(ROS) (P30 AG10161, R01 AG15819), the Texas Alzheimer’s Research and Care Consortium (TARCC) \n(funded by the Darrell K Royal Texas Alzheimer’s Initiative), Vanderbilt University/Case Western Reserve \nUniversity (VAN/CWRU) (R01 AG019757, R01 AG021547, R01 AG027944, R01 AG028786, P01 NS026630, \nand Alzheimer’s Association), the Washington Heights-Inwood Columbia Aging Project (WHICAP) (RF1 \nAG054023), the University of Washington Families (VA Research Merit Grant, NIA: P50AG005136, \nR01AG041797, NINDS: R01NS069719), the Columbia University Hispanic Estudio Familiar de Influencia \nGenetica de Alzheimer (EFIGA) (RF1 AG015473), the University of Toronto (UT) (funded by Wellcome Trust, \nMedical Research Council, Canadian Institutes of Health Research), and Genetic Differences (GD) (R01 \nAG007584). The CHARGE cohorts are supported in part by National Heart, Lung, and Blood Institute (NHLBI) \ninfrastructure grant HL105756 (Psaty), RC2HL102419 (Boerwinkle) and the neurology working group is \nsupported by the National Institute on Aging (NIA) R01 grant AG033193. \n \nThe CHARGE cohorts participating in the ADSP include the following: Austrian Stroke Prevention Study \n(ASPS), ASPS-Family study, and the Prospective Dementia Registry-Austria (ASPS/PRODEM-Aus), the \nAtherosclerosis Risk in Communities (ARIC) Study, the Cardiovascular Health Study (CHS), the Erasmus \nRucphen Family Study (ERF), the Framingham Heart Study (FHS), and the Rotterdam Study (RS). ASPS is \nfunded by the Austrian Science Fond (FWF) grant number P20545-P05 and P13180 and the Medical \nUniversity of Graz. The ASPS-Fam is funded by the Austrian Science Fund (FWF) project I904), the EU Joint \nProgramme – Neurodegenerative Disease Research (JPND) in frame of the BRIDGET project (Austria, \nMinistry of Science) and the Medical University of Graz and the Steiermärkische Krankenanstalten \nGesellschaft. PRODEM-Austria is supported by the Austrian Research Promotion agency (FFG) (Project No. \n827462) and by the Austrian National Bank (Anniversary Fund, project 15435. ARIC research is carried out as \na collaborative study supported by NHLBI contracts (HHSN268201100005C, HHSN268201100006C, \nHHSN268201100007C, HHSN268201100008C, HHSN268201100009C, HHSN268201100010C, \nHHSN268201100011C, and HHSN268201100012C). Neurocognitive data in ARIC is collected by U01 \n2U01HL096812, 2U01HL096814, 2U01HL096899, 2U01HL096902, 2U01HL096917 from the NIH (NHLBI, \nNINDS, NIA and NIDCD), and with previous brain MRI examinations funded by R01-HL70825 from the NHLBI. \nCHS research was supported by contracts HHSN268201200036C, HHSN268200800007C, N01HC55222, \nN01HC85079, N01HC85080, N01HC85081, N01HC85082, N01HC85083, N01HC85086, and grants \nU01HL080295 and U01HL130114 from the NHLBI with additional contribution from the National Institute of \nNeurological Disorders and Stroke (NINDS). Additional support was provided by R01AG023629, R01AG15928, \nand R01AG20098 from the NIA. FHS research is supported by NHLBI contracts N01-HC-25195 and \nHHSN268201500001I. This study was also supported by additional grants from the NIA (R01s AG054076, \nAG049607 and AG033040 and NINDS (R01 NS017950). The ERF study as a part of EUROSPAN (European \nSpecial Populations Research Network) was supported by European Commission FP6 STRP grant number \n018947 (LSHG-CT-2006-01947) and also received funding from the European Community’s Seventh \nFramework Programme (FP7/2007-2013)/grant agreement HEALTH-F4- 2007-201413 by the European \nCommission under the programme “Quality of Life and Management of the Living Resources” of 5th \nFramework Programme (no. QLG2-CT-2002- 01254). High-throughput analysis of the ERF data was supported \nby a joint grant from the Netherlands Organization for Scientific Research and the Russian Foundation for \nBasic Research (NWO-RFBR 047.017.043). The Rotterdam Study is funded by Erasmus Medical Center and \nErasmus University, Rotterdam, the Netherlands Organization for Health Research and Development \n(ZonMw), the Research Institute for Diseases in the Elderly (RIDE), the Ministry of Education, Culture and \nScience, the Ministry for Health, Welfare and Sports, the European Commission (DG XII), and the municipality \nof Rotterdam. Genetic data sets are also supported by the Netherlands Organization of Scientific Research \nNWO Investments (175.010.2005.011, 911-03-012), the Genetic Laboratory of the Department of Internal \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nMedicine, Erasmus MC, the Research Institute for Diseases in the Elderly (014-93-015; RIDE2), and the \nNetherlands Genomics Initiative (NGI)/Netherlands Organization for Scientific Research (NWO) Netherlands \nConsortium for Healthy Aging (NCHA), project 050-060-810. All studies are grateful to their participants, faculty \nand staff. The content of these manuscripts is solely the responsibility of the authors and does not necessarily \nrepresent the official views of the National Institutes of Health or the U.S. Department of Health and Human \nServices. \n \nThe FUS cohorts include: the Alzheimer’s Disease Research Centers (ADRC) (P30 AG062429, P30 \nAG066468, P30 AG062421, P30 AG066509, P30 AG066514, P30 AG066530, P30 AG066507, P30 \nAG066444, P30 AG066518, P30 AG066512, P30 AG066462, P30 AG072979, P30 AG072972, P30 \nAG072976, P30 AG072975, P30 AG072978, P30 AG072977, P30 AG066519, P30 AG062677, P30 \nAG079280, P30 AG062422, P30 AG066511, P30 AG072946, P30 AG062715, P30 AG072973, P30 AG066506, \nP30 AG066508, P30 AG066515, P30 AG072947, P30 AG072931, P30 AG066546, P20 AG068024, P20 \nAG068053, P20 AG068077, P20 AG068082, P30 AG072958, P30 AG072959), Alzheimer’s Disease \nNeuroimaging Initiative (ADNI) (U19AG024904), Amish Protective Variant Study (RF1AG058066), Cache \nCounty Study (R01AG11380, R01AG031272, R01AG21136, RF1AG054052), Case Western Reserve \nUniversity Brain Bank (CWRUBB) (P50AG008012), Case Western Reserve University Rapid Decline \n(CWRURD) (RF1AG058267, NU38CK000480), CubanAmerican Alzheimer’s Disease Initiative (CuAADI) \n(3U01AG052410), Estudio Familiar de Influencia Genetica en Alzheimer (EFIGA) (5R37AG015473, \nRF1AG015473, R56AG051876), Genetic and Environmental Risk Factors for Alzheimer Disease Among \nAfrican Americans Study (GenerAAtions) (2R01AG09029, R01AG025259, 2R01AG048927), Gwangju \nAlzheimer and Related Dementias Study (GARD) (U01AG062602), Hillblom Aging Network (2014-A-004-NET, \nR01AG032289, R01AG048234), Hussman Institute for Human Genomics Brain Bank (HIHGBB) \n(R01AG027944, Alzheimer’s Association “Identification of Rare Variants in Alzheimer Disease”), Ibadan Study \nof Aging (IBADAN) (5R01AG009956), Longevity Genes Project (LGP) and LonGenity (R01AG042188, \nR01AG044829, R01AG046949, R01AG057909, R01AG061155, P30AG038072), Mexican Health and Aging \nStudy (MHAS) (R01AG018016), Multi-Institutional Research in Alzheimer’s Genetic Epidemiology (MIRAGE) \n(2R01AG09029, R01AG025259, 2R01AG048927), Northern Manhattan Study (NOMAS) (R01NS29993), Peru \nAlzheimer’s Disease Initiative (PeADI) (RF1AG054074), Puerto Rican 1066 (PR1066) (Wellcome Trust \n(GR066133/GR080002), European Research Council (340755)), Puerto Rican Alzheimer Disease Initiative \n(PRADI) (RF1AG054074), Reasons for Geographic and Racial Differences in Stroke (REGARDS) \n(U01NS041588), Research in African American Alzheimer Disease Initiative (REAAADI) (U01AG052410), the \nReligious Orders Study (ROS) (P30 AG10161, P30 AG72975, R01 AG15819, R01 AG42210), the RUSH \nMemory and Aging Project (MAP) (R01 AG017917, R01 AG42210Stanford Extreme Phenotypes in AD \n(R01AG060747), University of Miami Brain Endowment Bank (MBB), University of Miami/Case Western/North \nCarolina A&T African American (UM/CASE/NCAT) (U01AG052410, R01AG028786), and Wisconsin Registry \nfor Alzheimer’s Prevention (WRAP) (R01AG027161 and R01AG054047). \n \nThe four LSACs are: the Human Genome Sequencing Center at the Baylor College of Medicine (U54 \nHG003273), the Broad Institute Genome Center (U54HG003067), The American Genome Center at the \nUniformed Services University of the Health Sciences (U01AG057659), and the Washington University \nGenome Institute (U54HG003079). Genotyping and sequencing for the ADSP FUS is also conducted at John \nP . Hussman Institute for Human Genomics (HIHG) Center for Genome Technology (CGT). \n \nBiological samples and associated phenotypic data used in primary data analyses were stored at Study \nInvestigators institutions, and at the National Centralized Repository for Alzheimer’s Disease and Related \nDementias (NCRAD, U24AG021886) at Indiana University funded by NIA. Associated Phenotypic Data used in \nprimary and secondary data analyses were provided by Study Investigators, the NIA funded Alzheimer’s \nDisease Centers (ADCs), and the National Alzheimer’s Coordinating Center (NACC, U24AG072122) and the \nNational Institute on Aging Genetics of Alzheimer’s Disease Data Storage Site (NIAGADS, U24AG041689) at \nthe University of Pennsylvania, funded by NIA. Harmonized phenotypes were provided by the ADSP \nPhenotype Harmonization Consortium (ADSP-PHC), funded by NIA (U24 AG074855, U01 AG068057 and R01 \nAG059716) and Ultrascale Machine Learning to Empower Discovery in Alzheimer’s Disease Biobanks (AI4AD, \nU01 AG068057). This research was supported in part by the Intramural Research Program of the National \nInstitutes of health, National Library of Medicine. Contributors to the Genetic Analysis Data included Study \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nInvestigators on projects that were individually funded by NIA, and other NIH institutes, and by private U.S. \norganizations, or foreign governmental or nongovernmental organizations. \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\nReferences: \n1. Ebbert MTW, Jensen TD, Jansen-West K, Sens JP , Reddy JS, Ridge PG, et al. Systema�c analysis of dark and \ncamouﬂaged genes reveals disease-relevant genes hiding in plain sight. Genome Biology 2019 20:1. 2019;20:1–\n23. \n2. Karimzadeh M, Ernst C, Kundaje A, Hoﬀman MM. Umap and Bismap: quan�fying genome and methylome \nmappability. Nucleic Acids Research. 2018;46:e120–e120. \n3. Lee H, Schatz MC. Genomic dark mater: the reliability of short read mapping illustrated by the genome \nmappability score. Bioinforma�cs. 2012;28:2097–105. \n4. Glinos DA, Garborcauskas G, Hoﬀman P , Ehsan N, Jiang L, Gokden A, et al. Transcriptome varia�on in human \n�ssues revealed by long-read sequencing. Nature 2022. 2022;:1–8. \n5. Leung SK, Jeﬀries AR, Castanho I, Jordan BT , Moore K, Davies JP , et al. Full-length transcript sequencing of \nhuman and mouse cerebral cortex iden�ﬁes widespread isoform diversity and alterna�ve splicing. Cell Rep. \n2021;37:110022. \n6. Aguzzoli Heberle B, Brandon JA, Page ML, Na�ons KA, Dikobe KI, White BJ, et al. Mapping medically relevant \nRNA isoform diversity in the aged human frontal cortex with deep long-read RNA-seq. Nat Biotechnol. 2024;:1–\n12. \n7. Page ML, Aguzzoli Heberle B, Brandon JA, Wadsworth M, Gordon LA, Na�ons KA, et al. Surveying the \nlandscape of RNA isoform diversity and expression across 9 GTEx �ssues using long-read sequencing data. \nbioRxiv. 2024;:2024–02. \n8. Nurk S, Koren S, Rhie A, Rau�ainen M, Bzikadze AV, Mikheenko A, et al. The complete sequence of a human \ngenome. Science. 2022;376:44–53. \n9. Altemose N, Logsdon GA, Bzikadze AV, Sidhwani P , Langley SA, Caldas GV, et al. Complete genomic and \nepigene�c maps of human centromeres. Science. 2022;376. \n10. Wang T, Antonacci-Fulton L, Howe K, Lawson HA, Lucas JK, Phillippy AM, et al. The Human Pangenome \nProject: a global resource to map genomic diversity. Nature 2022 604:7906. 2022;604:437–46. \n11. Hehir-Kwa JY , Marschall T , Kloosterman WP , Francioli LC, Baaijens JA, Dijkstra LJ, et al. A high-quality human \nreference panel reveals the complexity and distribu�on of genomic structural variants. Nature \nCommunica�ons. 2016;7. \n12. Chiang C, Scot AJ, Davis JR, Tsang EK, Li X, Kim Y , et al. The impact of structural varia�on on human gene \nexpression. Nature gene�cs. 2017;49:692. \n13. Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, et al. A structural varia�on reference for \nmedical and popula�on gene�cs. Nature 2020 581:7809. 2020;581:444–51. \n14. Jonsson H, Magnusdo�r E, Eggertsson HP , Stefansson OA, Arnado�r GA, Eiriksson O, et al. Diﬀerences \nbetween germline genomes of monozygo�c twins. Nature Gene�cs 2021 53:1. 2021;53:27–34. \n15. Porubsky D, Vollger MR, Harvey WT , Rozanski AN, Ebert P , Hickey G, et al. Gaps and complex structurally \nvariant loci in phased genome assemblies. Genome Research. 2023;33:496–510. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n16. Hoyt SJ, Storer JM, Hartley GA, Grady PGS, Gershman A, de Lima LG, et al. From telomere to telomere: The \ntranscrip�onal and epigene�c state of human repeat elements. Science. 2022;376. \n17. Aganezov S, Yan SM, Soto DC, Kirsche M, Zarate S, Avdeyev P , et al. A complete reference genome improves \nanalysis of human gene�c varia�on. Science. 2022;376. \n18. Kent WJ. BLAT—The BLAST-Like Alignment Tool. Genome Res. 2002;12:656–64. \n19. Lee HJ, Kim JY , Song IH, Park IA, Yu JH, Gong G. Expression of NY-ESO-1 in Triple-Nega�ve Breast Cancer Is \nAssociated with Tumor-Inﬁltra�ng Lymphocytes and a Good Prognosis. Oncology. 2015;89:337–44. \n20. IBUKIĆ A, RAMIĆ S, ZOVAK M, BILIĆ Z, TOMAS D, DEMIROVIĆ A. Expression and Prognos�c Signiﬁcance of \nPD-L1 and NY-ESO1 in Gallbladder Carcinoma. In Vivo. 2023;37:1828–37. \n21. Meng X, Sun X, Liu Z, He Y . A novel era of cancer/tes�s an�gen in cancer immunotherapy. Interna�onal \nImmunopharmacology. 2021;98:107889. \n22. Hashimoto K, Nishimura S, Shinyashiki Y , Ito T, Kakinoki R, Akagi M. Involvement of NY-ESO-1 and MAGE-A4 \nin the pathogenesis of desmoid tumors. Medicine (Bal�more). 2023;102:e33908. \n23. Endo M, de Graaﬀ MA, Ingram DR, Lim S, Lev DC, Briaire-de Bruijn IH, et al. NY-ESO-1 (CTAG1B) expression \nin mesenchymal tumors. Mod Pathol. 2015;28:587–95. \n24. Jordan CT, Upchurch D, Szilvassy SJ, Guzman ML, Howard DS, Pe�grew AL, et al. The interleukin-3 receptor \nalpha chain is a unique marker for human acute myelogenous leukemia stem cells. Leukemia. 2000;14:1777–\n84. \n25. Cuglievan B, Connors J, He J, Khazal S, Yedururi S, Dai J, et al. Blas�c plasmacytoid dendri�c cell neoplasm: a \ncomprehensive review in pediatrics, adolescents, and young adults (AYA) and an update of novel therapies. \nLeukemia. 2023;37:1767–78. \n26. Pemmaraju N, Wilson NR, Senapa� J, Economides MP , Guzman ML, Neelapu SS, et al. CD123-directed \nallogeneic chimeric-an�gen receptor T-cell therapy (CAR-T) in blas�c plasmacytoid dendri�c cell neoplasm \n(BPDCN): Clinicopathological insights. Leukemia Research. 2022;121:106928. \n27. Qazi S, Uckun FM. Augmented Expression of the IL3RA/CD123 Gene in MLL/KMT2A-Rearranged Pediatric \nAML and Infant ALL. Onco. 2022;2:245–63. \n28. Palpant NJ, Pabon L, Rabinowitz JS, Hadland BK, Stoick-Cooper CL, Paige SL, et al. Transmembrane protein \n88: a Wnt regulatory protein that speciﬁes cardiomyocyte development. Development. 2013;140:3799–808. \n29. Zhang M, Liu J, Mao A, Ning G, Cao Y , Zhang W, et al. Tmem88 conﬁnes ectodermal Wnt2bb signaling in \npharyngeal arch artery progenitors for balancing cell cycle progression and cell fate decision. Nat Cardiovasc \nRes. 2023;2:234–50. \n30. Lieberman-Aiden E, van Berkum NL, Williams L, Imakaev M, Ragoczy T , Telling A, et al. Comprehensive \nmapping of long range interac�ons reveals folding principles of the human genome. Science. 2009;326:289–\n93. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n31. Dincer A, Gavin DP , Xu K, Zhang B, Dudley JT , Schadt EE, et al. Deciphering H3K4me3 broad domains \nassociated with gene-regulatory networks and conserved epigenomic landscapes in the human brain. Transl \nPsychiatry. 2015;5:e679–e679. \n32. Persico G, Casciaro F, Amatori S, Rusin M, Cantatore F, Perna A, et al. Histone H3 Lysine 4 and 27 \nTrimethyla�on Landscape of Human Alzheimer’s Disease. Cells. 2022;11:734. \n33. De Rubeis S, He X, Goldberg AP , Poultney CS, Samocha K, Ercument Cicek A, et al. Synap�c, transcrip�onal \nand chroma�n genes disrupted in au�sm. Nature. 2014;515:209–15. \n34. Shulha HP , Cheung I, Whitle C, Wang J, Virgil D, Lin CL, et al. Epigene�c Signatures of Au�sm: Trimethylated \nH3K4 Landscapes in Prefrontal Neurons. Archives of General Psychiatry. 2012;69:314–24. \n35. Cheung I, Shulha HP , Jiang Y , Matevossian A, Wang J, Weng Z, et al. Developmental regula�on and \nindividual diﬀerences of neuronal H3K4me3 epigenomes in the prefrontal cortex. Proceedings of the Na�onal \nAcademy of Sciences. 2010;107:8824–9. \n36. Vaisvila R, Ponnaluri VKC, Sun Z, Langhorst BW, Saleh L, Guan S, et al. Enzyma�c methyl sequencing detects \nDNA methyla�on at single-base resolu�on from picograms of DNA. Genome Res. 2021;31:1280–9. \n37. Tanaka K, Okamoto A. Degrada�on of DNA by bisulﬁte treatment. Bioorganic & Medicinal Chemistry \nLeters. 2007;17:1912–5. \n38. Olova N, Krueger F, Andrews S, Oxley D, Berrens RV, Branco MR, et al. Comparison of whole-genome \nbisulﬁte sequencing library prepara�on strategies iden�ﬁes sources of biases aﬀec�ng DNA methyla�on data. \nGenome Biol. 2018;19:33. \n39. Valsesia A, Kulkarni SS, Marquis J, Leone P , Mironova P , Walter O, et al. Salivary α-amylase copy number is \nnot associated with weight trajectories and glycemic improvements following clinical weight loss: results from \na 2-phase dietary interven�on study. The American Journal of Clinical Nutri�on. 2019;109:1029–37. \n40. Barber TM, Bha� AA, Elder PJD, Ball SP , Calvez R, Ramsden DB, et al. AMY1 Gene Copy Number Correlates \nWith Glucose Absorp�on and Visceral Fat Volume, but Not with Insulin Resistance. The Journal of Clinical \nEndocrinology & Metabolism. 2020;105:e3586–96. \n41. Chen X, Li M, Tang Y , Liang Q, Hua C, He H, et al. Gene Expression Proﬁle Analysis of Human Epidermal \nKera�nocytes Expressing Human Papillomavirus Type 8 E7. Pathol Oncol Res. 2022;28:1610176. \n42. François-Moutal L, Scot DD, Ambrose AJ, Zerio CJ, Rodriguez-Sanchez M, Dissanayake K, et al. Heat shock \nprotein Grp78/BiP/HspA5 binds directly to TDP-43 and mi�gates toxicity associated with disease pathology. Sci \nRep. 2022;12:8140. \n43. Kalmar B, Lu C-H, Greensmith L. The role of heat shock proteins in Amyotrophic Lateral Sclerosis: The \ntherapeu�c poten�al of Arimoclomol. Pharmacology & Therapeu�cs. 2014;141:40–54. \n44. Karlin S, Brocchieri L. Heat shock protein 60 sequence comparisons: Duplica�ons, lateral transfer, and \nmitochondrial evolu�on. Proceedings of the Na�onal Academy of Sciences of the United States of America. \n2000;97:11348–53. \n45. Meng G, Xu H, Lu D, Li S, Zhao Z, Li H, et al. Three-dimensional chroma�n architecture datasets for aging \nand Alzheimer’s disease. Sci Data. 2023;10:51. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n46. Sleegers K, Brouwers N, Gijselinck I, Theuns J, Goossens D, Wauters J, et al. APP duplica�on is suﬃcient to \ncause early onset Alzheimer’s demen�a with cerebral amyloid angiopathy. Brain. 2006;129:2977–83. \n47. Konno T , Ross OA, Puschmann A, Dickson DW, Wszolek ZK. Autosomal dominant Parkinson’s disease caused \nby SNCA duplica�ons. Parkinsonism & related disorders. 2016;22 Suppl 1:S1. \n48. Singleton AB, Farrer M, Johnson J, Singleton A, Hague S, Kachergus J, et al. α-Synuclein Locus Triplica�on \nCauses Parkinson’s Disease. Science. 2003;302:841. \n49. Butensky A, de Rinaldis CP , Patel S, Edman S, Bailey A, McGinn DE, et al. Cardiac evalua�on of pa�ents with \n22q11.2 duplica�on syndrome. American Journal of Medical Gene�cs Part A. 2021;185:753–8. \n50. Hasten E, McDonald-McGinn DM, Crowley TB, Zackai E, Emanuel BS, Morrow BE, et al. Dysregula�on of \nTBX1 dosage in the anterior heart ﬁeld results in congenital heart disease resembling the 22q11.2 duplica�on \nsyndrome. Human Molecular Gene�cs. 2018;27:1847–57. \n51. Kucukkilic E, Brookes K, Barber I, Gueta-Baranes T , Morgan K, Hollox EJ, et al. Complement receptor 1 gene \n(CR1) intragenic duplica�on and risk of Alzheimer’s disease. Hum Genet. 2018;137:305–14. \n52. Yang X, Wang X, Zou Y , Zhang S, Xia M, Fu L, et al. Characteriza�on of large-scale genomic diﬀerences in the \nﬁrst complete human genome. Genome Biology. 2023;24:1–25. \n53. Brouwers N, Van Cauwenberghe C, Engelborghs S, Lambert J-C, Betens K, Le Bastard N, et al. Alzheimer \nrisk associated with a copy number varia�on in the complement receptor 1 increasing C3b/C4b binding sites. \nMol Psychiatry. 2012;17:223–33. \n54. Wong WW, Cahill JM, Rosen MD, Kennedy CA, Bonaccio ET, Morris MJ, et al. Structure of the human CR1 \ngene. Molecular basis of the structural and quan�ta�ve polymorphisms and iden�ﬁca�on of a new CR1-like \nallele. Journal of Experimental Medicine. 1989;169:847–63. \n55. Smith BO, Mallin RL, Krych-Goldberg M, Wang X, Hauhart RE, Bromek K, et al. Structure of the C3b Binding \nSite of CR1 (CD35), the Immune Adherence Receptor. Cell. 2002;108:769–80. \n56. Wijmenga C, Hewit JE, Sandkuijl LA, Clark LN, Wright TJ, Dauwerse HG, et al. Chromosome 4q DNA \nrearrangements associated with facioscapulohumeral muscular dystrophy. Nat Genet. 1992;2:26–30. \n57. van Deutekom JC, Wijmenga C, van Tienhoven EA, Gruter AM, Hewit JE, Padberg GW, et al. FSHD \nassociated DNA rearrangements are due to dele�ons of integral copies of a 3.2 kb tandemly repeated unit. \nHum Mol Genet. 1993;2:2037–42. \n58. Armstrong J, Hickey G, Diekhans M, Fiddes IT, Novak AM, Deran A, et al. Progressive Cactus is a mul�ple-\ngenome aligner for the thousand-genome era. Nature. 2020;587:246. \n59. Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, et al. The Human Genome Browser at \nUCSC. Genome Research. 2002;12:996. \n60. Li H. Aligning sequence reads, clone sequences and assembly con�gs with BWA-MEM. 2013. \n61. Li H. Minimap2: pairwise alignment for nucleo�de sequences. Bioinforma�cs. 2018;34:3094–100. \n62. Li H. New strategies to improve minimap2 alignment accuracy. Bioinforma�cs. 2021;37:4572–4. \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint \n\n63. FelixKrueger/TrimGalore: A wrapper around Cutadapt and FastQC to consistently apply adapter and quality \ntrimming to FastQ ﬁles, with extra func�onality for RRBS data. htps://github.com/FelixKrueger/TrimGalore. \nAccessed 12 Jan 2024. \n64. Krueger F, Andrews SR. Bismark: a ﬂexible aligner and methyla�on caller for Bisulﬁte-Seq applica�ons. \nBioinforma�cs. 2011;27:1571. \n65. Langmead B, Salzberg SL. Fast gapped-read alignment with Bow�e 2. Nature Methods. 2012;9:357–9. \n66. Quinlan AR, Hall IM. BEDTools: a ﬂexible suite of u�li�es for comparing genomic features. Bioinforma�cs. \n2010;26:841. \n67. Tommaso PD, Chatzou M, Floden EW, Barja PP , Palumbo E, Notredame C. Nex�low enables reproducible \ncomputa�onal workﬂows. Nature Biotechnology 2017 35:4. 2017;35:316–9. \n68. Thorvaldsdó�r H, Robinson JT, Mesirov JP . Integra�ve Genomics Viewer (IGV): high-performance genomics \ndata visualiza�on and explora�on. Brieﬁngs in Bioinforma�cs. 2013;14:178–92. \n69. Klickstein LB, Bartow TJ, Mile�c V, Rabson LD, Smith JA, Fearon DT. Iden�ﬁca�on of dis�nct C3b and C4b \nrecogni�on sites in the human C3b/C4b receptor (CR1, CD35) by dele�on mutagenesis. Journal of \nExperimental Medicine. 1988;168:1699–717. \n70. Smit A, Hubley R, Green P . RepeatMasker Open-4.0. htp://www.repeatmasker.org. 2015. \n \n \n.CC-BY 4.0 International licenseavailable under a \nwas not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprint (whichthis version posted May 28, 2025. ; https://doi.org/10.1101/2025.05.23.655776doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}