The Impact of Structural Variation on Alzheimer’s Disease in the Alzheimer’s Disease Sequencing Project

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Introduction : Structural variants (SV), genomic alterations spanning more than 50 base pairs, can significantly impact gene expression and protein function. However, their contribution to Alzheimer’s Disease (AD) remains poorly understood. Leveraging a novel SV calling pipeline, we identified SVs with high accuracy in a diverse sample of the Alzheimer's Disease Sequencing Project (ADSP) and investigated the role of SVs in AD. Results : We analyzed SVs in 16,841 individuals from ADSP whole genome sequencing data using BioGraph, a semi-assembly-based method that employs graph-based representation for accurate SV detection. We identified 456,644 high-quality SVs, 65% of which were novel. Of these, 272,728 SVs directly impact genes, including 86 AD-related genes. Association analyses were performed within three ancestry groups, including 3,371 African (AFR), 6,327 European (EUR), and 2,126 Latin (LAT). Multiple deletions and insertions were observed in moderate to high linkage disequilibrium with known AD loci, including TPCN1 and TMEM106B . In EUR, genome-wide association analysis identified two significant low-frequency deletions associated with AD, located in introns of CCDC12 and CCDC88B , both encoding coiled-coil domain-containing proteins. Gene-based analyses further identified rare pathogenic SVs in several known AD genes, including PSEN1 in LAT and ABCA7 in AFR. Conclusions : Using a novel graph-based SV calling pipeline, we identified high-quality SVs across a large and ancestrally diverse cohort. Our analyses revealed both common and rare SVs associated with AD. These findings provide valuable insights into the genetic architecture of AD, emphasizing the value of including diverse populations in AD genomic studies.
Full text 195,241 characters · extracted from preprint-html · click to expand
The Impact of Structural Variation on Alzheimer’s Disease in the Alzheimer’s Disease Sequencing Project | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article The Impact of Structural Variation on Alzheimer’s Disease in the Alzheimer’s Disease Sequencing Project Songmi Lee, Adam C English, Gina M Peloso, Joshua C Bis, Eric Boerwinkle, and 8 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8562759/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 9 You are reading this latest preprint version Abstract Introduction : Structural variants (SV), genomic alterations spanning more than 50 base pairs, can significantly impact gene expression and protein function. However, their contribution to Alzheimer’s Disease (AD) remains poorly understood. Leveraging a novel SV calling pipeline, we identified SVs with high accuracy in a diverse sample of the Alzheimer's Disease Sequencing Project (ADSP) and investigated the role of SVs in AD. Results : We analyzed SVs in 16,841 individuals from ADSP whole genome sequencing data using BioGraph, a semi-assembly-based method that employs graph-based representation for accurate SV detection. We identified 456,644 high-quality SVs, 65% of which were novel. Of these, 272,728 SVs directly impact genes, including 86 AD-related genes. Association analyses were performed within three ancestry groups, including 3,371 African (AFR), 6,327 European (EUR), and 2,126 Latin (LAT). Multiple deletions and insertions were observed in moderate to high linkage disequilibrium with known AD loci, including TPCN1 and TMEM106B . In EUR, genome-wide association analysis identified two significant low-frequency deletions associated with AD, located in introns of CCDC12 and CCDC88B , both encoding coiled-coil domain-containing proteins. Gene-based analyses further identified rare pathogenic SVs in several known AD genes, including PSEN1 in LAT and ABCA7 in AFR. Conclusions : Using a novel graph-based SV calling pipeline, we identified high-quality SVs across a large and ancestrally diverse cohort. Our analyses revealed both common and rare SVs associated with AD. These findings provide valuable insights into the genetic architecture of AD, emphasizing the value of including diverse populations in AD genomic studies. Structural variation Alzheimer’s Disease genetics association study whole genome sequence Figures Figure 1 Figure 2 1. Background Alzheimer’s Disease (AD), characterized by progressive memory loss and declining cognitive function, is the most common form of dementia among older adults. It is estimated that almost 7 million Americans aged 65 and older are currently living with AD and this number is projected to double by 2060 [ 1 , 2 ]. To date, there are few effective disease-modifying treatments or prevention strategies for AD, underscoring a critical need to better understand its etiology. Genetic factors play a substantial role in AD, with a disease heritability estimated between 60–80% [ 3 ]. To date, genetic approaches have yielded important novel insights into AD etiology [ 4 ] and promise further advances in prevention, diagnosis, and treatment. Very rare highly-penetrant mutations have been identified in the Amyloid Beta Precursor Protein ( APP ), Presenilin 1 ( PSEN1 ) and Presenilin 2 ( PSEN2 ) genes that cause Mendelian forms of AD, typically with early onset [ 5 ]. More common alleles identified in large genome-wide association studies (GWAS) of sporadic late-onset AD have uncovered genes involved in cholesterol metabolism, endocytosis/phagocytosis, amyloid plaque and neurofibrillary tangle formation, and the innate immune system [ 6 , 7 ]. Despite significant progress in understanding the genetic basis of AD, a substantial proportion of AD genetic architecture remains unknown [ 8 ]. Addressing this gap in knowledge requires a comprehensive characterization of all forms of genetic variation, beyond single nucleotide variants (SNVs), and including structural variation. Structural variants (SVs) are typically defined as genomic alterations comprising 50 or more base pairs [ 9 , 10 ]. These variants can be classified into five different types: insertions, deletions, inversions, duplications, and translocations. Compared to SNVs, SVs are numerically fewer but are larger in size, and therefore have a greater impact on DNA sequence and, consequently, on gene expression and function [ 11 , 12 ]. SVs often occur in highly repetitive and polymorphic regions of the genome [ 13 ], making them challenging to detect with short-read DNA sequencing technology [ 10 , 14 ]. Insertions are especially problematic and their role across diseases remains understudied [ 10 ]. Over the past decade, technological and methodological developments have improved SV detection from short-read whole genome sequence (WGS) data [ 10 , 15 ], providing an opportunity to more comprehensively and accurately evaluate their impact on complex disease etiology. In this study, we have implemented a novel SV calling method, BioGraph [ 16 ], on the WGS data from 16,841 subjects of the Alzheimer’s Disease Sequencing Project (ADSP, Release 3) and examined the association of the detected SVs with AD. Our study investigated the role of SVs in known AD loci, providing insights into the genetic architecture of AD. We also examined associations between common SVs across the genome and AD, and performed gene-based association testing to analyze rare SVs in AD. 2. Results 2.1. Identification of high-quality SVs using BioGraph in the ADSP We implemented our novel method BioGraph [ 16 ] to generate a highly accurate SV call set from 16,841 WGS in the ADSP (Fig. 1 ). BioGraph is a unique approach to SV detection and genotyping that leverages reference guided assembly of short reads to improve the detection of SV. Additionally, BioGraph uses machine learning techniques to assign useful quality scores to the identified candidate SVs. Comparison of BioGraph performance in detecting SV with that of other SV calling tools, including Manta [ 17 ], Parliament2 [ 18 ], and Smoove [ 19 ], using benchmark data from the Genome In A Bottle HG002 Challenging Medically Relevant Regions [ 20 ] and Truvari [ 21 ] are described in detail in the Supplementary Material . Our initial raw set of SV calls across the ADSP’s 16,841 sample set generated 1,019,035 SVs. We deployed novel sample-based filtering approaches to further ensure high quality and accuracy of SV calls across the samples. We first leveraged technical replicates present in the ADSP data. In total, there are 601 replicate samples derived from 283 unique individuals providing 428 replicate pairs. Using Truvari, we compared the SV calls between each replicate pair. First, we compared the quality score distributions of calls which were found consistently between replicates to those which were inconsistent ( Supplementary Fig. 1 ). We determined a quality score threshold of 50 best segregates the calls by their consistency. These data and additional data examining the quality score’s relationship to False Positive and True Positive measurements in benchmarking experiments ( Supplementary Material ) suggest a minimum quality score of 50 should be applied for high quality calls. In total, 562,391 SVs were removed using this filter. We also analyzed the similarity of the consistent calls between replicate pairs. Over 97% of consistent calls have a sequence/size similarity of 95% or greater, suggesting that SVs with ≥ 95% sequence and size similarity should be considered the same when performing inter-sample merging. Our rigorous QC procedure enabled us to assign pass and fail values across all SVs. Moreover, we collapsed and filtered 55.2% of the initially inferred SV and thus avoided many potential false positive or redundant alleles. After QC, we identified 456,644 SVs, including 254,716 deletions (55.78%), 190,786 insertions (41.78%) and 11,142 inversions (2.44%). The distribution of SV size and type across the ADSP sample is shown in Fig. 2 A. Most SVs identified were less than 5 kbp in length (94.74%) with a majority ranging from 50–100 bp (42.18%). We observed the expected ALU peak (~ 300 bp) for deletions but even a more prominent peak for insertions. The latter is due to our merging strategy from Truvari where the SV are only merged if their sequence similarity exceeds 95% in addition to type and length constraints [ 21 ]. The number of deletions and insertions were similar across most size categories with some exceptions. We identified fewer insertions in size categories larger than 1 kbp. This contrasts with other short read SV calling approaches where the number of insertions declines rapidly starting at 500 bp [ 10 ]. We also observed a size bias across inversions with only 210 inversions over 1 kbp detected. Figure 2 B shows the site frequency spectrum (log scale). Overall, we observed that the insertions and deletions occur at similar frequencies and the majority of them (87.3%) have an allele frequency < 1%. Most of the inversions are singletons (70.7%) compared to deletions and insertions for which the proportion of singletons is 37.2% and 33.6%, respectively. These findings reflect the challenge to correctly infer inversions from short-read sequence data [ 14 ]. The singleton rate appears to vary with SV size, being lowest in smaller SVs (50–100 bp: 27.4% singletons) compared to midsize SVs (1 kbp-2.5 kbp: 51.1%) and large SVs (> 5 kbp: 60.7% singletons). We speculate that this is likely due to a combination of larger events accumulating mutations over time as well as larger events being less consistently discovered. 2.2. A total of 297,034 SVs are novel in the ADSP sample We evaluated the overlap of our high-quality SV call set with previously reported SVs from major reference databases, including the Center for Common Disease Genomics (CCDG), Trans-omics for Precision Medicine (TOPMed), 1000Genomes Project (1KGP), and the Genome Aggregation Database (gnomAD). Only 34.9% of the SV identified in the ADSP overlapped with the reported SVs. The greatest overlap was found with gnomAD data (23.3% of ADSP SV calls), followed by 1KGP (19.4%). When investigating overlap by SV type, 18.9% of the overlapping SV were deletions, with the greatest overlap with gnomAD (12.33%); and 16% were insertions, with the greatest overlap with 1KGP (27.6%) rather than gnomaAD (26.3%). This may be due to tandem duplications in gnomAD reported as insertions in the 1KGP and our call sets, a common SV type swap [ 13 ]. We investigated the correlation of allele frequency among common overlapping SVs (MAF > 1%), and observed a strong overall concordance (r 2 = 0.73, P < 0.01), underscoring the accuracy and reliability of the SV call set. As expected, the correlation was higher for deletions (r 2 = 0.89, P < 0.01) compared to insertions (r 2 = 0.63, P < 0.01). Interestingly, these correlations improved further when comparing the ADSP SV call set with external reference datasets. For example, the correlation of overlapping SVs in Biograph ADSP SV call sets and in gnomAD call sets was 0.94 for deletions and 0.82 for insertions. Notably, many of the novel SVs identified in our study were within size ranges that were more effectively captured by BioGraph than methods relying solely on paired-end read distances. For example, the average deletion size in gnomAD was 7.4 kbp compared to 2.1 kbp in the Biograph ADSP data. For insertions, these numbers were 895 bp vs. 184 bp, respectively. These findings further highlight the strength of our dataset including SVs that may have been under-called by previous short-read studies. 2.3. Annotation and overlap of SVs with genes We identified SVs that directly overlapped or were in close proximity (within 5 kbp) to gene sequences using SVAfotate [ 22 ]. All 456,644 SVs were annotated, of which 143,397 SVs (31.4%) mapped to intergenic regions, 131,988 SVs (28.9%) were reported in proximity of genes but not overlapping them directly, and 272,728 SVs (59.7%) directly impacted a gene. These 272,728 SVs mapped to a total of 30,333 (62.8%) genes, suggesting that they are impacting the same gene more than once across different individuals. The majority of these gene-impacting SVs were deletions (53.3%) followed by insertions (36.6%) and inversions (10.1%). The lower number of gene-impacting insertions compared to deletions is likely because insertions are measured as affecting only the direct base pair at which they are reported, while deletions span multiple base pairs on the reference. The majority (96.9%) of gene-impacting SVs were located within introns whereas 3.0% mapped to the 3’ untranslated region (UTR) and only 1.6% mapped to the 5’ UTR. The SVs mapped to UTR have a higher chance to impact regulatory function itself. Only 2.3% of SVs were directly overlapped coding sequences. Among 86 AD genes reported by the ADSP Gene Verification Committee, 82 genes intersected 1,223 SVs. Most genes (71) had SVs within ± 5kbp as well as overlapping the gene body. Filtering to only common (AF ≥ 1%) SVs hitting non-intronic gene bodies left 69 variants over 30 genes ( Supplementary Table 1 ). SVs of note included a 6,137bp deletion on PRDM7 (AF = 2.4%), 4 tandem repeat expansions between 51bp and 98bp of a 12bp VNTR in RBCK1 , and a 322bp deletion on TMEM106B with a frequency of 49.7%. 2.4. PCA and ancestry inference We derived principal components from our SV data (N = 12,908) to account for possible population structure in the data. PC2 was associated with read length and demonstrated complete separation of samples as shown in Supplementary Fig. 2 . To minimize confounding by batch effects in the association analyses, study participants were further restricted to those with a read-length of 150 (N = 11,890; 5,585 cases and 6,305 controls), which represents more than 90% of the sample. PCA analyses revealed similarities in results between PCs derived from SVs and those derived from SNVs ( Supplementary Fig. 3 ). Based on the results of GrafPop ( Supplementary Fig. 4 ), our study included 3,371 individuals of African (AFR) ancestry, 6,327 of European (EUR) ancestry, 2,126 of Latin (LAT) ancestry, and 66 participants that did not cluster with those three ancestry groups and were therefore excluded from subsequent analyses ( Supplementary Table 2 ). 2.5. SVs in linkage disequilibrium (LD) with Alzheimer’s Disease known loci Ancestry-specific LD analyses identified 9 SVs in EUR, 5 SVs in AFR, and 9 SVs in LAT that were in moderate or high LD (r 2 = 0.43–0.99) with at least one of the SNPs previously identified in AD GWAS ( Supplementary Tables 3–5 ). These SNPs did not exhibit strong associations with AD in our dataset, due to limited statistical power compared to the GWAS sample size in which they were discovered. Among the identified SVs, 4 SVs in EUR, 1 SV in AFR, and 1 SV in LAT showed suggestive evidence of association with AD (SV P-value < 0.1). The strongest SV association with AD was observed in AFR, involving a 122-bp deletion in moderate LD (r 2 = 0.46) with rs2633682 tagging the ALCAM locus ( Supplementary Table 4 ). This SNP, previously associated with AD specifically in an African American population [ 23 ], showed a suggestive association with AD in our dataset (SNP P-value = 0.009). However, conditional analyses indicated that neither the SVs nor the SNP remained significant after adjusting for each other, suggesting non-independence of the signals at the ALCAM locus. In EUR, a similar trend was observed at the ALCAM locus, although the associations were weaker. A 319-bp deletion in an intron of TPCN1 was observed in all three ancestry groups and was in high LD (r 2 = 0.97 in all groups) with the tagging SNP. In EUR and LAT, two deletions, including the 319-bp deletion, and one insertion were in moderate or high LD with a SNP tagging at the TPCN1 locus ( Supplementary Tables 3 and 5 ). Haplotype estimation analysis suggested that all detected SVs lie on the same haplotype as the AD risk allele at this locus ( Supplementary Fig. 5 ). A 68-bp deletion located in an intron of SLC8B1 were detected in EUR and LAT, with suggestive association with AD observed in EUR (SV P-value = 0.03). While the intronic SNP (rs6489896) tagging TPCN1 has been previously associated with AD at genome-wide significance, SLC8B1 has not. Additionally, a 322-bp Alu deletion in exon 8 and 3’ untranslated region of TMEM106B exhibited strong LDs with two tagging SNPs in EUR and LAT, and moderate LD in AFR. Conditional analyses adjusting for the corresponding SNPs revealed that none of the SV associations remained (adjusted SV P-value > 0.1), indicating that the observed suggestive SV associations were not independent of SNPs in LD at those loci. However, in LAT ( Supplementary Table 5 ), a SNP tagging WNT3 / MATP locus remained significant (adjusted SNP P-value = 0.006) after conditioning on a 314-bp deletion in moderate LD (r 2 = 0.68), suggesting that the SNP is independently associated with AD at this locus. 2.6. Genome-wide association of common or low-frequency SVs with Alzheimer’s Disease For each ancestry group, we performed single variant association analyses of high-quality SVs with MAF > 0.5%, and in HWE (as defined in section 4.8). In total, we analyzed 28,942 SVs in AFR, 14,656 SVs in EUR, and 30,394 SVs in LAT ( Supplementary Table 6 ). No SVs were significantly associated with AD in the AFR or LAT analyses, or in the meta-analyses. In EUR, two deletions were significantly associated with AD at the Bonferroni-corrected threshold and were observed exclusively in this ancestry group (Table 1 ; Supplementary Fig. 6 ). Both deletions mapped to introns of genes encoding coiled-coil domain containing proteins. Table 1 Bonferroni-significant SVs associated with AD identified by genome-wide single variant association analyses Group SV type SV locus SV Breakpoints Length QUAL Gene(s) Location Allele Frequencies Single variant association tests AFR EUR LAT OR [95% CI] P-value EUR DEL 3p21.31 46978504–46978622 119 89 CCDC12 intron 0 0.013 0 3.19 [2.1–4.9] 7.66E-08 EUR DEL 11q13.1 64341844–64341923 80 94 CCDC88B intron 0 0.011 0 2.89 [1.9–4.4] 2.09E-06 AFR, African ancestry group; LAT, Latin ancestry group; EUR, European ancestry group; DEL, deletion; QUAL, quality score; OR, odds ratio; CI, confidence interval Table 2 Bonferroni-significant or suggestive genes associated with AD in gene-based association analyses of coding variants (A) and non-coding variants (B) A. Coding SNVs/INDELs + SVs Group Gene Locus Category n variants cMAC P-value LAT PSEN1 14q24.2 pLof + disruptive missense 11 61 1.96E-07 LAT PSEN1 14q24.2 missense 14 74 8.69E-07 LAT SMOC1 14q24.2 synonymous 14 88 1.50E-06 LAT ACOT4 14q24.3 disruptive missense 3 35 2.72E-06 LAT ACOT4 14q24.3 missense 9 61 5.44E-06 LAT ACOT4 14q24.3 synonymous 10 54 5.70E-06 B. Non-coding SNVs/INDELs + SVs Group Gene Locus Category n variants cMAC P-value LAT ELMSAN1 14q24.3 ncRNA 3 23 1.86E-09 LAT AC005225.2 14q24.3 ncRNA 6 44 6.61E-08 LAT LOC100506476 14q24.3 Promoter (CAGE) 21 93 7.17E-08 LAT LOC100506476 14q24.3 Enhancer (CAGE) 26 103 1.03E-07 LAT ACOT6 14q24.3 Promoter (DHS) 21 104 3.61E-07 LAT AL390763.1 10q26.2 ncRNA 5 10 2.24E-06 LAT ACOT4 14q24.3 Enhancer, Promoter (CAGE) 17 59 2.53E-06 LAT AC005225.2 14q24.3 Promoter (DHS) 43 263 5.09E-06 LAT ACOT4 14q24.3 Enhancer (DHS) 17 93 5.25E-06 LAT LINC01500 14q24.1 Promoter (CAGE) 6 69 6.25E-06 LAT TIRAP 11q24.2 UTR 6 16 7.43E-06 LAT PROX2 14q24.3 Enhancer (DHS) 27 130 7.49E-06 LAT ACOT4 14q24.3 Promoter (DHS) 19 103 9.08E-06 LAT AL163974.1 14q32.2 Upstream 16 90 9.88E-06 LAT ACOT6 14q24.3 Enhancer (DHS) 23 107 1.29E-07 SNV, single nucleotide variants; INDELs, insertion and deletions; pLof, putative loss of function; cMAC, cumulative minor allele count; ncRNA, non-coding RNA; CAGE, cap analysis of gene expression; DHS, DNase I hypersensitive site; UTR, untranslated region 2.7. Gene-based association of rare SVs and SNVs with Alzheimer’s Disease We performed gene-based analyses to test association between aggregated rare SVs and AD. In the primary gene-based association analyses, two analyses were carried out based on the SV types: coding SVs ( Supplementary Fig. 7 ) and noncoding SVs ( Supplementary Fig. 8 ). No genes reached genome-wide significant or suggestive significance thresholds in either analysis. In the secondary analyses, we conducted gene-based analyses to assess association between AD and aggregates of rare SVs and rare SNVs/INDELs. Coding variant analyses were performed using five categories of SNVs/INDELs, combined with pathogenic coding SVs ( Supplementary Fig. 9 ), while non-coding variant analyses included eight categories of SNVs/INDELs in combination with pathogenic non-coding SVs ( Supplementary Fig. 10 ). In the secondary coding variant analyses, we identified PSEN1 as being suggestively associated with AD in LAT when aggregating SVs and SNVs/INDELs classified as pLoF and disruptive missense (P = 1.9E-07), disruptive missense (P = 1.9E-07), or missense (P = 8.6E-07) variants (Table 2 A). For this region, we observed two deletions and two insertions, all more frequent or exclusively observed in AD controls ( Supplementary Table 7 ). Conditional analyses on the aggregates of SNVs/INDELs in PSEN1 indicated that the four pathogenic coding SVs were independently associated with AD (adjusted P-values < 0.05). No SVs were detected for the SMOC1 and ACOT4 regions. In the secondary non-coding variant analyses, multiple genes showed significant or suggestive associations with AD in LAT (Table 2 B); however, no SVs were detected for those genes, and thus the observed associations were entirely driven by SNVs/INDELs. Not surprisingly, several of the associated genes overlapped with those identified in a previous study based solely on SNVs and INDELs, including ELMSAN1 , ACOT6 , and ACOT4 [ 24 ]. In the candidate gene analyses, we examined 15 previously reported AD-associated genes to determine whether rare SVs contributed to their associations signals. No genes reached statistical significance in the primary analyses. In the secondary coding variant analyses, PSEN1 , TREM2 , and ABCA7 showed evidence of association with AD (FDR Q < 0.05) ( Supplementary Table 8A ). Conditional analyses on the aggregates of SNVs/INDELs within each gene suggested independent effects of coding SVs for PSEN1 and ABCA7 (adjusted P-value < 0.05). Notably, ABCA7 showed evidence of an independent SV association with AD, driven by a 605 bp exonic deletion (chr19:1050368–1050972) observed exclusively in four AD cases within AFR. This association remained significant after conditioning on disruptive missense SNVs/INDELs, indicating an independent effect of the deletion on AD risk. In the secondary non-coding variant analyses, TREM2 and ABI3 showed evidence of association with AD, however, no evidence of independent SV association with AD was observed ( Supplementary Table 8B ). 3. Discussion In this study, we analyzed SVs in 16,841 individuals from ADSP using BioGraph, a novel SV calling pipeline. We identified 456,644 high-quality SVs, approximately 65% of which were novel. Notably, the vast majority of novel SVs were insertions, which may have been under-detected in previous studies. Among common or low-frequency SVs within each ancestry group, several SVs were found to be in moderate or high LD with known AD loci, offering additional insights into the genetic architecture of AD. Genome-wide association analyses identified two low-frequency deletions associated with AD in individuals of European ancestry, both located within genes encoding coiled-coil domain-containing proteins. Gene-based analyses further revealed that PSEN1 and ABCA7 harbor rare pathogenic SVs associated with AD. Our use of BioGraph, a semi-assembly-based SV calling method, enabled the identification of many insertions not previously reported. Long-read sequencing and genome assembly studies have shown that insertions are the most prevalent SV class, often representing tandem repeat expansions or transposable element integrations that are not in the reference genome. This fact makes insertions challenging to identify, but also biologically intriguing as they have been reported to affect splicing or induce mosaic variants in proximity (e.g. ALUY insertions) [ 25 ]. Our SV call set demonstrated high accuracy and precision for both insertions and deletions, supported by rigorous benchmarking on replicate samples within ADSP and assessment across control samples. We have further introduced detection of inversions from BioGraph results that yielded multiple inversions candidates. This is noteworthy as the correct identification of inversions remains highly challenging [ 26 ]. We comprehensively assessed SVs and their potential impact on AD. Notably, many genes previously highlighted by SNV-based GWAS exhibited SVs either within the gene itself or within 5 kb. Overall, 95.3% of the postulated genes showed SV overlap. To further explore the role of SVs in established AD loci while accounting for ancestral differences, we examined SVs in known AD loci within each inferred ancestry group. We identified several SVs in moderate or high LD with known AD loci across different ancestry groups. At the TPCN1 locus, a 319-bp intronic deletion was observed across the three ancestry groups in high LD with the tagging SNP. This deletion fully overlaps with a previously reported 309-bp deletion associated with Lewy Body dementia, which was validated using long-read sequencing [ 27 ]. TPCN1 , which is highly expressed in the brain, encodes the two-pore calcium channel protein 1 located on endolysosomal membranes. Beyond its association with AD identified in previous GWAS [ 6 ], the function of TPCN1 has been demonstrated in knockout mice, which exhibit impairments in spatial learning and memory [ 28 ]. In both EUR and LAT, we additionally discovered a deletion and an insertion at the TPCN1 locus that showed evidence of association with AD in EUR. The deletion was located in SLC8B1 , which encodes a mitochondrial Na+/Ca2 + exchanger. Notably, a recent study demonstrated that a deletion of the SLC8B1 region in knockout mice, spanning the region of our four deletions, is sufficient to induce AD-like pathology, including age-related cognitive decline [ 29 ]. Our haplotype estimation analysis suggested that all detected SVs lie on the same haplotype as the AD risk allele at the TPCN1 locus. These findings suggest that multiple genes within the TPCN1 locus may influence AD risk and underscore the need for further investigation into the role of SVs at the TPCN1 locus across ancestrally diverse populations and in other neurodegenerative diseases. TMEM106B encodes a transmembrane glycoprotein that localizes to late lysosome and endosome [ 30 ]. At the TMEM106 locus, we detected a 322-bp Alu deletion in exon 8 or 3’ untranslated region, which has previously been reported as a likely causal variant and validated using long-read sequencing data [ 31 ]. This deletion has been associated with not only with AD, but also with frontotemporal lobar dementia with TDP-43 inclusions (FTLD-TDP) [ 31 ], neurodegeneration [ 32 ], and several AD-related phenotypes, including tangles density, TDP-43, and cognitive resilience [ 33 ]. From genome-wide analyses of common or low frequency SVs, we identified two significant deletions associated with AD among EUR. Both deletions are located in genes encoding CCDC proteins. Members of this family are characterized by an N-terminal potential microtubule binding domain, a central coiled-coiled and a C-terminal Hook-related domain. An 80-bp deletion on chromosome 11 is located in intron 7 of CCDC88B and encompasses MIR7155 (chr11:64,341,849 − 64,341,904). CCDC88B has been shown to act as a positive regulator of T-cell maturation and inflammatory function [ 34 ]. The low frequency deletion on chromosome 3 is located in intron 1 of CCDC12 and 2 Kb upstream of neurobeachin like 2 ( NBEAL2 ). The functions of CCDC12 remain unclear but it is predicted to be part of the spliceosomal complex. NBEAL2 is thought to play a role in megakaryocyte alpha-granule biogenesis. In public databases, these two deletions are annotated as indels with rsIDs rs1553653356 (chr3) and rs1591274862 (chr11), respectively. In the gnomAD database (v4.1.0) [ 35 ], rs1553653356 shows a low frequency (AF = 0.002), consistent with our findings. In contrast, rs1591274862 (chr11) shows a notable discrepancy between the exome data (AF = 0.2) and genome data (AF = 0.002), although this variant failed quality control in both datasets. This discrepancy highlights the need for further investigation. Nonetheless, at the gene-level, a pQTL for CCDC88B and an eQTL for CCDC12 have been previously associated with AD [ 36 , 37 ], suggesting potential causal links between these genes and AD. Functional validation of the two deletions is warranted. We identified multiple genes associated with AD in gene-based analyses of rare SVs and SNVs/INDELs, particularly in LAT. The Latino population is genetically admixed, with varying proportions of European, African, and Amerindigenous genetic backgrounds, which adds genetic complexity [ 38 ]. According to GrafPop, the Latin American 1 population primarily represents individuals with European and African ancestry components, whereas the Latin American 2 population mainly represents individuals with European and Amerindigenous components [ 39 ]. Notably, the Latin American populations exhibit unique LD patterns and haplotype structures derived from admixture [ 40 ], which may enhance the detection of rare variants. For PSEN1 , we identified four rare, coding SVs with evidence of association with AD. All four SVs were annotated as highly pathogenic by indirectly altering PSEN1 regulatory elements. Indeed, all four SVs are located in regions of neighboring genes, not within PSEN1 itself, highlighting the impact of SVs through long range regulatory mechanisms [ 41 ]. For ABCA7 , we identified a rare 605-bp deletion that partially overlaps intron 18 and exon 19, showing evidence of association with AD risk in AFR. This observation aligns with findings from a recent study [ 42 ], despite their use of a different SV caller and statistical model. This deletion is located approximately 3 kb downstream of a previously reported 44-bp deletion (rs142076058) associated with AD risk in African American individuals [ 43 ]. Additionally, it overlaps with a well-characterized SNP (rs115550680) previously associated with late-onset AD in African Americans populations [ 44 ]. In our dataset, all four individuals carrying the rare 605-bp deletion had AD and did not carry the previously reported deletion or SNPs associated with AD among African populations, whereas two common SNPs (rs3764650 and rs3752246) previously identified in European AD GWAS were observed [ 45 ]. Our findings provide new insights into the genetic architecture of the ABCA7 locus in African ancestry populations. SPARC-related modular calcium-binding protein 1 (SMOC1 ) has consistently been reported as a biomarker for early AD in proteomics studies [ 36 , 46 – 49 ], although the underlying genetic basis remains unclear. While we did not observe any rare SVs for SMOC1 , the observed aggregate of synonymous SNVs/INDELs suggestively associated with AD in this gene may partially explain the genetic contribution to the increased levels of SMOC1 in AD [ 50 ]. We acknowledge several limitations in this study. Despite the relatively large sample of ascertained AD cases, statistical power remains limited, particularly for rare noncoding SV analyses within ancestry subgroups. Analyses using pooled populations did not yield additional associations for common SVs and may have introduced potential false positives in the gene-based analyses, possibly due to data structure complexities inherent to rare SVs and incomplete adjustment for population stratification using PCs. Another limitation is the lack of replication for our findings. While we identified two significant deletions associated with AD, their low frequency poses challenges for replication. Nonetheless, validation in larger and independent cohorts will be essential. 4. Conclusions In conclusion, we identified high-quality SVs in ADSP samples using a novel SV calling method. Our analysis revealed ancestry-specific SVs at known AD loci, as well as both common and rare SVs associated with AD. These findings provide new insights into genetic architecture of AD. Future studies are warranted to validate our results and investigate the functional impact of these SVs. 5. Methods 5.1. Study samples The Alzheimer’s Disease Sequencing Project (ADSP) was initiated in 2012 to elucidate the genetic architecture of AD, with major goals to identify genes and gene variants that confer risk for or protection against AD, to provide insight as to the biological impact of these genes and variants, and to identify potential therapeutic targets [ 51 ]. The WGS data release used in the present study (Release 3) includes data from 16,841 diverse individuals with and without AD from 24 cohorts. Raw data were obtained from the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS). After removing duplicates (Section 5.4 ), outliers subjects (Section 5.5 ), and subjects with missing phenotypic information, 12,908 samples (6,604 controls, 6,304 cases) remained for analysis. 5.2. SV detection BioGraph version v6.0.4 was run per-sample using GRCh38 as the reference [ 16 ]. Variant Call Format (VCF) files were filtered to variant sites at least 50 bp long and with a PASS filter. Inversions were identified from all VCF entries with at least a 50 bp reference and 50 bp alternate allele reported where the sequence similarity of the reference and the reverse complement of the alternate allele was at least 80%. SVs were merged using bcftools v1.15 [ 52 ] and SVs with over 95% sequence and size similarity within 1000 bp were consolidated using Truvari collapse v3.1 with parameter `--keep maxqual` [ 21 ]. When necessary, SVs were cross-referenced to intersecting tandem repeat regions from the adotto TR catalog [ 53 ]. 5.3. SV benchmarking using challenging, medically-relevant genes (CMRG) We used WGS data from HG002, a sample with broad consent for open genomic data sharing through the Personal Genome Project [ 54 ]. SVs were called using BioGraph [ 16 ], Manta v1.6.0 [ 17 ], Parliament2 [ 18 ], and smoove 0.2.6 [ 19 ]. Truvari v3.5 [ 21 ] was used to compare the resulting SV calls against the CMRG benchmark [ 20 ]. Default truvari parameters were used for BioGraph and Manta. The parameters `--dup-to-ins` and `--pctsim 0` were used for Parliament2 and smoove as neither tool produces sequence resolved calls. 5.4. Quality control of SVs using replicates analysis Truvari v3.1 was run between 428 replicate pairs (ADSP participants with more than one sample sequenced). SV calls with over 70% sequence and size similarity between the replicates were classified as being consistent and the remainder classified as being inconsistent. Truvari annotations of PctSeqSimilarity and PctSizeSimilarity between consistent SV pairs were also analyzed to identify SVs that are the same across samples. 5.5. Quality control of samples using One-Class Support Vector Machine (SVM) Passing SV counts by type were collected for each sample. Classification of the 2% of outlier samples by counts was performed using scikit-learn v1.1.3 and their OneClass SVM with hyper-parameters kernel=’poly’ and nu = 0.02. Intersection of samples with the study which provided them showed a concentration of outlier samples from 3 of the 24 studies comprising the ADSP study sample. All samples included in these three studies (N = 421) were dropped. 5.6. Intersection with known SVs and annotation to genes SVAfotate version 0.0.1 was used to intersect the discovered SVs with known SVs [ 22 ]. This program comprises an annotated file containing boundaries of SVs from 1000G [ 9 ], CCDG [ 55 ], and gnomAD [ 56 ]. TopMed SVs freeze 1.1 [ 57 ] were also collected and consolidated into the annotated file. SVs were annotated to genes with VEP using VEP-ensembl version ​​107.0 [ 58 ]. Next, we examined whether any of the discovered SVs mapped to genes reported by the ADSP Gene Verification Committee [ 59 ]. 5.7. Global ancestry inference Global ancestry inference of the study samples was performed using GrafPop [ 39 , 60 ], a distance-based method that uses a reference composed of nearly 100,000 fingerprint SNPs extracted from dbGaP [ 61 ]. Grafpop estimates ancestry by calculating genetic distances between each individual and the reference populations, and subjects are clustered using genetic distances based on their genetic similarity. This tool considers that individuals’ genomes are admixed from three ancestries: European (E), African (F), and Asian (A) and estimates ancestral proportions P e , P f , and P a based on genetic distances score using barycentric coordinates. In GrafPop, the cutoff thresholds were empirically defined to facilitate the grouping of dbGaP subjects. Due to the incompatibility of GrafPop with SV data, we used ADSP WGS data on single nucleotide variants (SNV) to perform global ancestry inference. Using the cutoff standard established by GrafPop, ADSP participants with WGS data were clustered into nine groups defined by study-reported populations within dbGAP. These groups include European, African, East Asian, African American, Latin American 1, Asian-Pacific Islander, South Asian, Latin American 2, and Other, based on their ancestral proportions and genetic distance [ 39 ]. We grouped African and African American populations as African ancestry group (AFR) and Latin American 1 and Latin American 2 populations as Latin ancestry group (LAT), and European population as European ancestry group (EUR) [ 24 ]. Participants identified as East Asian, Asian-Pacific Islander, South Asian, and other populations were grouped as others and were excluded from subgroup association analyses due to limited sample size. 5.8. Principal Component Analysis Principal component analysis (PCA) was performed using PC-AiR [ 62 ] in the GENetic EStimation and Inference in Structured samples (GENESIS) package [ 63 ]. We calculated PCs for all individuals (N = 12,908) in the study sample using high quality deletions, insertions, and inversions with minor allele frequency (MAF) greater than 1%, and with Hardy Weinberg Equilibrium (HWE) P-value greater than the Bonferroni-corrected threshold based on the total number of SVs (P = 5.2E-07). SVs in linkage disequilibrium (LD) were excluded using a r 2 threshold greater than 0.1. For comparison with SNVs, we performed PCA on the same sample using biallelic SNPs with MAF > 1%, HWE P-value > 1E-06, and call rate > 95%. For ancestry-specific SV association analyses, we calculated PCs using SVs with MAF > 1% and HWE P-values exceeding the Bonferroni-corrected threshold based on the number of SVs with MAF > 1% in each ancestry group (AFR: P = 4.7E-07; EUR: P = 6.5E-07; LAT: P = 5.4E-07). For ancestry-specific SNVs association analyses, we focused on previously reported AD GWAS SNPs that were in LD with our SV calls. For these analyses, PCs were calculated using WGS data filtered for MAF > 1%, HWE P > 1E-06, and call rate > 95% within each ancestry group. 5.9. SVs tagging known AD GWAS SNPs To investigate the role of SVs in known AD GWAS loci, we examined 147 SNPs tagging AD loci identified in previous GWAS [ 6 , 7 , 23 , 64 , 65 ]. These variants were extracted from our WGS data, and we performed pairwise LD analysis between the AD-associated SNPs and our SVs within each ancestry group (Fig. 1 A). Among common or low-frequency SVs with MAF > 0.5% and in HWE (see above), we specifically focused on SVs that were in LD (r 2 > 0.4) with at least one of the AD-associated SNPs. LD calculation was carried out using PLINK v1.9 with parameters --ld-window-r2 0.4 and --r2. Haplotype estimation was performed using PLINK v1.9 based on pairwise LD patterns. 5.10. Association analyses 5.10.1. Models and Covariates Within each ancestry group inferred based on genetic similarity, association analyses were conducted using a mixed effects logistic regression model. Detailed models and software for common and rare SV analyses are provided in the corresponding sections below. Covariates included sex, SV-derived PC 1–5 of each ancestry, relatedness via a genetic relatedness matrix (GRM), and technical covariates including sequencing center and whether the sample preparation was PCR-free. For the analysis evaluating SVs in LD with SNPs tagging AD loci, we applied a conditional model that further included the corresponding SNP dosage. We also performed association analyses of the identified SNPs in LD with SVs, replacing SV PCs with SNP PCs, and including the corresponding SV dosage as a covariate in the conditional model. 5.10.2. Single variant analysis of common or low-frequency SVs Association analyses of common and low frequency SVs (MAF > 0.5% and passing the HWE criterion) were conducted using a mixed-effects logistic regression model implemented in the GENESIS R-package [ 39 ], with covariates and models as described above. All association analyses were performed within each ancestry group. For genome-wide association analysis, a meta-analysis was additionally performed across all ancestry subgroups using METAL software, implementing Stouffer method to weight results by sample size. To identify AD-associated SVs, we considered several P-value thresholds: For evaluating SVs in LD with known AD loci (Fig. 1 A), a suggestive significance threshold (P < 0.1) was used. For genome-wide analysis of all common or low-frequency SVs (Fig. 1 B), two significance thresholds were considered: a Bonferroni-corrected threshold (AFR: P < 1.7E-06; EUR: P < 3.4E-06; LAT: P < 1.6E-06) based on the total number of SVs analyzed and the conventional genome-wide significance threshold (P < 5E-08). 5.10.3. Gene-based analysis of aggregates of rare SVs Gene-based association analyses of aggregated rare SVs with AD were conducted as the primary analyses. We included deletions, insertion, and inversion with MAF < 1% in each ancestry group and estimated their pathogenicity using PhenoSV [ 66 ]. PhenoSV is a machine learning based method that predicts the functional consequences of both coding and non-coding SVs that may directly or indirectly influence genes. SVs were classified as coding if they overlapped at least 1bp with exons of protein-coding gene based on GENCODE v40 annotations [ 67 ], considering only high-confidence representative transcript, otherwise they were classified as noncoding. Independently, SVs were evaluated for their potential to affect genes directly or indirectly. Non-coding SVs were tested for their indirect effect on genes within 1Mb upstream and downstream, as defined by default. Gene-level pathogenicity scores ranged from 0 to 1 and were used to classify SVs into pathogenic (≥ 0.5) and benign (< 0.5) groups [ 66 ]. Only rare pathogenic SVs were included in the analyses, which were performed separately for non-coding and coding variants (Fig. 1 C, 1 D). In secondary analyses, we integrated SV data with SNVs/INDELs derived from ADSP 17K WGS data [ 24 ] to increase statistical power. The WGS data had been previously processed and quality-controlled according to the Genome Center for Alzheimer’s Disease (GCAD)/ADSP QC pipeline [ 68 ]. Using the WGS data annotated with FAVOR, we classified SNVs/INDELs as coding or non-coding based on the STAAR pipeline [ 69 ]. Variants with MAF < 1% within each ancestry group were included in the analysis. The coding SNVs/INDELs were categorized into five functional groups: putative loss of function (pLof), missense, disruptive missense, pLof + disruptive missense, or synonymous. The non-coding SNVs/INDELs were grouped into eight categories: promoter or enhancer overlaid with cap analysis of gene expression (CAGE) or DNase I hypersensitive site (DHS) sites, untranslated region (UTR), upstream, downstream, and noncoding RNA genes. Gene-based analyses were then performed within each category of coding and noncoding variants, combining rare pathogenic SVs of the corresponding type (Fig. 1 E, 1 F). All gene-based association analyses were performed using the variant-set mixed model association test (SMMAT), implemented in the GMMAT R package [ 70 ]. We used a hybrid test (SMMAT-E), which combines burden and SKAT tests and has been shown to offer greater statistical power than either test alone. MAF was used as a weight by default. Genes with a cumulative minor allele count (cMAC) ≥ 10 were included in the analysis. All association analyses were stratified by ancestry group, and meta-analyses combining results across ancestry subgroups were conducted using the metap R package [ 71 ] with the Stouffer method to account for sample size differences. We applied two significance thresholds for gene-based tests: Bonferroni-corrected threshold (P < 1E-07) to account for approximately 20,000 genes tested, and a suggestive threshold (P < 1E-05). For genes with significance in the secondary analyses, we performed conditional analyses on the aggregates of SNVs/INDELs within each gene to assess whether the association signal was driven by SVs. Finally, we conducted candidate gene association analyses focusing on 15 AD genes previously reported to harbor rare variant associations [ 72 , 73 ], evaluating them in both primary and secondary gene-based analyses. We computed the false discovery rate (FDR) Q value within each analysis group using the Benjamini-Hochberg procedure to assess statistical significance. For genes showing evidence of association with AD (FDR Q < 0.05) in the secondary analyses, we conducted additional conditional analyses on the aggregated SNVs/INDELs to evaluate the contribution of SVs to the observed signal. Declarations Author Contribution S.L., A.C.E., M.F., and F.J.S. contributed to the conception and design of the work, data interpretation, and manuscript drafting. A.C.E. developed new software used in the study. S.L., R.X., and A.C.E. performed the data analysis. G.M.P., S.H.C., and A.L.D. contributed to data acquisition. S.L. and A.C.E. contributed equally to this work. M.F. and F.J.S. contributed equally to this work. All authors read, reviewed, and approved the final manuscript. Data availability ADSP whole genome sequencing data (NG00067) are available through the National Institute on Aging Genetics of Alzheimer’s Disease Data Storage Site (NIAGADS) (https://www.niagads.org). Acknowledgments We thank the participants and their families for making this research possible. Data used in this study were generated via the Alzheimer Disease Sequencing Project (ADSP). Full ADSP acknowledgements can be found here: https://adsp.niagads.org/acknowledgment/ Data used in this study were obtained from the Alzheimer's Disease Neuroimaging Initiative (ADNI) via the ADSP. As such, the investigators within the ADNI contributed to the design and implementation of ADNI and/or provided data but did not participate in the analysis or writing of this report. A complete listing of ADNI investigators can be found at: http://adni.loni.usc.edu/wp-content/uploads/how_to_apply/ADNI_Acknowledgement_List.pdf Sources of Funding This work was primarily supported by grants U01AG058589, U01AG052409, and U01AG070112 from the National Institute on Aging. Additional support was provided by grant U01AG068221. Data for this study were prepared, archived, and distributed by the National Institute on Aging Alzheimer’s Disease Data Storage Site (NIAGADS) at the University of Pennsylvania (U24-AG041689), funded by the National Institute on Aging. A complete description of the funding support for the ADSP is provided at https://adsp.niagads.org/acknowledgment/. Consent Statement This study constitutes secondary research, utilizing de-identified data obtained from primary data repositories. In accordance with NIH policy, this research does not qualify as human subject research, and therefore, obtaining consent from individual participants is not required. All contributing studies included in this work received ethical oversight from their respective institutions. References Rajan KB, Weuve J, Barnes LL, McAninch EA, Wilson RS, Evans DA. Population estimate of people with clinical Alzheimer's disease and mild cognitive impairment in the United States (2020–2060). Alzheimers Dement. 2021;17:1966–75. 2025 Alzheimer’s disease facts and figures. Alzheimer's & Dementia 2025. Gatz M, Reynolds CA, Fratiglioni L, Johansson B, Mortimer JA, Berg S, Fiske A, Pedersen NL. Role of genes and environments for explaining Alzheimer disease. Arch Gen Psychiatry. 2006;63:168–74. Sims R, Hill M, Williams J. The multiplex model of the genetics of Alzheimer's disease. Nat Neurosci. 2020;23:311–22. Tanzi RE. The genetics of Alzheimer disease. Cold Spring Harb Perspect Med 2012, 2. Bellenguez C, Kucukali F, Jansen IE, Kleineidam L, Moreno-Grau S, Amin N, Naj AC, Campos-Martin R, Grenier-Boley B, Andrade V, et al. New insights into the genetic etiology of Alzheimer's disease and related dementias. Nat Genet. 2022;54:412–36. Wightman DP, Jansen IE, Savage JE, Shadrin AA, Bahrami S, Holland D, Rongve A, Borte S, Winsvold BS, Drange OK, et al. A genome-wide association study with 1,126,563 individuals identifies new risk loci for Alzheimer's disease. Nat Genet. 2021;53:1276–82. Andrews SJ, Renton AE, Fulton-Howard B, Podlesny-Drabiniok A, Marcora E, Goate AM. The complex genetic architecture of Alzheimer's disease: novel insights and future directions. EBioMedicine. 2023;90:104511. Sudmant PH, Rausch T, Gardner EJ, Handsaker RE, Abyzov A, Huddleston J, Zhang Y, Ye K, Jun G, Fritz MH, et al. An integrated map of structural variation in 2,504 human genomes. Nature. 2015;526:75–81. Mahmoud M, Gobet N, Cruz-Davalos DI, Mounier N, Dessimoz C, Sedlazeck FJ. Structural variant calling: the long and the short of it. Genome Biol. 2019;20:246. Scott AJ, Chiang C, Hall IM. Structural variants are a major source of gene expression differences in humans and often affect multiple nearby genes. Genome Res. 2021;31:2249–57. Chiang C, Scott AJ, Davis JR, Tsang EK, Li X, Kim Y, Hadzic T, Damani FN, Ganel L, Consortium GT, et al. The impact of structural variation on human gene expression. Nat Genet. 2017;49:692–9. Audano PA, Sulovari A, Graves-Lindsay TA, Cantsilieris S, Sorensen M, Welch AE, Dougherty ML, Nelson BJ, Shah A, Dutcher SK, et al. Characterizing the Major Structural Variant Alleles of the Human Genome. Cell. 2019;176:663–e675619. Sedlazeck FJ, Rescheneder P, Smolka M, Fang H, Nattestad M, von Haeseler A, Schatz MC. Accurate detection of complex structural variations using single-molecule sequencing. Nat Methods. 2018;15:461–8. De Coster W, Weissensteiner MH, Sedlazeck FJ. Towards population-scale long-read sequencing. Nat Rev Genet. 2021;22:572–87. English AC, McCarthy N, Flickenger R, Maheshwari S, Meed L, Mangubat A, Shekar SN. Leveraging a WGS compression and indexing format with dynamic graph references to call structural variants. bioRxiv 2020:2020.2004.2024.060202.. Chen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Kallberg M, Cox AJ, Kruglyak S, Saunders CT. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics. 2016;32:1220–2. Zarate S, Carroll A, Mahmoud M, Krasheninina O, Jun G, Salerno WJ, Schatz MC, Boerwinkle E, Gibbs RA, Sedlazeck FJ. Parliament2: Accurate structural variant calling at scale. Gigascience 2020, 9. Layer RM, Chiang C, Quinlan AR, Hall IM. LUMPY: a probabilistic framework for structural variant discovery. Genome Biol. 2014;15:R84. Wagner J, Olson ND, Harris L, McDaniel J, Cheng H, Fungtammasan A, Hwang YC, Gupta R, Wenger AM, Rowell WJ, et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat Biotechnol. 2022;40:672–80. English AC, Menon VK, Gibbs RA, Metcalf GA, Sedlazeck FJ. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol. 2022;23:271. Nicholas TJ, Cormier MJ, Quinlan AR. Annotation of structural variants with reported allele frequencies and related metrics from multiple datasets using SVAFotate. BMC Bioinformatics. 2022;23:490. Kunkle BW, Schmidt M, Klein HU, Naj AC, Hamilton-Nelson KL, Larson EB, Evans DA, De Jager PL, Crane PK, Buxbaum JD, et al. Novel Alzheimer Disease Risk Loci and Pathways in African American Individuals Using the African Genome Resources Panel: A Meta-analysis. JAMA Neurol. 2021;78:102–13. Lee WP, Choi SH, Shea MG, Cheng PL, Dombroski BA, Pitsillides AN, Heard-Costa NL, Wang H, Bulekova K, Kuzma AB, et al. Association of common and rare variants with Alzheimer's disease in more than 13,000 diverse individuals with whole-genome sequencing from the Alzheimer's Disease Sequencing Project. Alzheimers Dement. 2024;20:8470–83. Smolka M, Paulin LF, Grochowski CM, Horner DW, Mahmoud M, Behera S, Kalef-Ezra E, Gandhi M, Hong K, Pehlivan D, et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat Biotechnol. 2024;42:1571–80. Sanders AD, Hills M, Porubsky D, Guryev V, Falconer E, Lansdorp PM. Characterizing polymorphic inversions in human genomes by single-cell sequencing. Genome Res. 2016;26:1575–87. Kaivola K, Chia R, Ding J, Rasheed M, Fujita M, Menon V, Walton RL, Collins RL, Billingsley K, Brand H, et al. Genome-wide structural variant analysis identifies risk loci for non-Alzheimer's dementias. Cell Genom. 2023;3:100316. Mallmann RT, Klugbauer N. Genetic Inactivation of Two-Pore Channel 1 Impairs Spatial Learning and Memory. Behav Genet. 2020;50:401–10. Jadiya P, Cohen HM, Kolmetzky DW, Kadam AA, Tomar D, Elrod JW. Neuronal loss of NCLX-dependent mitochondrial calcium efflux mediates age-associated cognitive decline. iScience. 2023;26:106296. Stagi M, Klein ZA, Gould TJ, Bewersdorf J, Strittmatter SM. Lysosome size, motility and stress response regulated by fronto-temporal dementia modifier TMEM106B. Mol Cell Neurosci. 2014;61:226–40. Chemparathy A, Le Guen Y, Zeng Y, Gorzynski J, Jensen TD, Yang C, Kasireddy N, Talozzi L, Belloy M, Stewart I, et al. A 3'UTR Insertion Is a Candidate Causal Variant at the TMEM106B Locus Associated With Increased Risk for FTLD-TDP. Neurol Genet. 2024;10:e200124. Salazar A, Tesi N, Knoop L, Pijnenburg Y, van der Lee S, Wijesekera S, Krizova J, Hiltunen M, Damme M, Petrucelli L et al. An AluYb8 retrotransposon characterises a risk haplotype of TMEM106B associated in neurodegeneration. medRxiv 2023:2023.2007.2016.23292721.. Vialle RA, de Paiva Lopes K, Li Y, Ng B, Schneider JA, Buchman AS, Wang Y, Farfel JM, Barnes LL, Wingo AP, et al. Structural variants linked to Alzheimer's disease and other common age-related clinical and neuropathologic traits. Genome Med. 2025;17:20. Kennedy JM, Fodil N, Torre S, Bongfen SE, Olivier JF, Leung V, Langlais D, Meunier C, Berghout J, Langat P, et al. CCDC88B is a novel regulator of maturation and effector functions of T cells during pathological inflammation. J Exp Med. 2014;211:2519–35. Chen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alfoldi J, Watts NA, Vittal C, Gauthier LD, et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature. 2024;625:92–100. Ali M, Timsina J, Western D, Liu M, Beric A, Budde J, Do A, Heo G, Wang L, Gentsch J, et al. Multi-cohort cerebrospinal fluid proteomics identifies robust molecular signatures across the Alzheimer disease continuum. Neuron. 2025;113:1363–e13791369. Mathys H, Davila-Velderrain J, Peng Z, Gao F, Mohammadi S, Young JZ, Menon M, He L, Abdurrob F, Jiang X, et al. Single-cell transcriptomic analysis of Alzheimer's disease. Nature. 2019;570:332–7. Mao X, Bigham AW, Mei R, Gutierrez G, Weiss KM, Brutsaert TD, Leon-Velarde F, Moore LG, Vargas E, McKeigue PM, et al. A genomewide admixture mapping panel for Hispanic/Latino populations. Am J Hum Genet. 2007;80:1171–8. Jin Y, Schaffer AA, Feolo M, Holmes JB, Kattman BL. GRAF-pop: A Fast Distance-Based Method To Infer Subject Ancestry from Multiple Genotype Datasets Without Principal Components Analysis. G3 (Bethesda) 2019, 9:2447–2461. da Cruz PRS, Ananina G, Secolin R, Gil-da-Silva-Lopes VL, Lima CSP, de Franca PHC, Donatti A, Lourenco GJ, de Araujo TK, Simioni M et al. Demographic history differences between Hispanics and Brazilians imprint haplotype features. G3 (Bethesda) 2022, 12. Boyling A, Perez-Siles G, Kennerson ML. Structural Variation at a Disease Mutation Hotspot: Strategies to Investigate Gene Regulation and the 3D Genome. Front Genet. 2022;13:842860. Wang H, Dombroski BA, Cheng PL, Tucci A, Si YQ, Farrell JJ, Tzeng JY, Leung YY, Malamon JS et al. Alzheimer's Disease Sequencing P, : Structural variation detection and association analysis of whole-genome-sequence data from 16,543 Alzheimer's disease sequencing project subjects. Alzheimers Dement 2025, 21:e70277. Cukier HN, Kunkle BW, Vardarajan BN, Rolati S, Hamilton-Nelson KL, Kohli MA, Whitehead PL, Dombroski BA, Van Booven D, Lang R, et al. ABCA7 frameshift deletion associated with Alzheimer disease in African Americans. Neurol Genet. 2016;2:e79. Reitz C, Jun G, Naj A, Rajbhandary R, Vardarajan BN, Wang LS, Valladares O, Lin CF, Larson EB, Graff-Radford NR, et al. Variants in the ATP-binding cassette transporter (ABCA7), apolipoprotein E ϵ4,and the risk of late-onset Alzheimer disease in African Americans. JAMA. 2013;309:1483–92. Dib S, Pahnke J, Gosselet F. Role of ABCA7 in Human Health and in Alzheimer's Disease. Int J Mol Sci 2021, 22. Sung YJ, Yang C, Norton J, Johnson M, Fagan A, Bateman RJ, Perrin RJ, Morris JC, Farlow MR, Chhatwal JP, et al. Proteomics of brain, CSF, and plasma identifies molecular signatures for distinguishing sporadic and genetic Alzheimer's disease. Sci Transl Med. 2023;15:eabq5923. Guo Y, Chen SD, You J, Huang SY, Chen YL, Zhang Y, Wang LB, He XY, Deng YT, Zhang YR, et al. Multiplex cerebrospinal fluid proteomics identifies biomarkers for diagnosis and prediction of Alzheimer's disease. Nat Hum Behav. 2024;8:2047–66. Wang H, Dey KK, Chen PC, Li Y, Niu M, Cho JH, Wang X, Bai B, Jiao Y, Chepyala SR, et al. Integrated analysis of ultra-deep proteomes in cortex, cerebrospinal fluid and serum reveals a mitochondrial signature in Alzheimer's disease. Mol Neurodegener. 2020;15:43. Watson CM, Dammer EB, Ping L, Duong DM, Modeste E, Carter EK, Johnson ECB, Levey AI, Lah JJ, Roberts BR, Seyfried NT. Quantitative Mass Spectrometry Analysis of Cerebrospinal Fluid Protein Biomarkers in Alzheimer's Disease. Sci Data. 2023;10:261. Oelschlaeger P. Molecular Mechanisms and the Significance of Synonymous Mutations. Biomolecules 2024, 14. Beecham GW, Bis JC, Martin ER, Choi SH, DeStefano AL, van Duijn CM, Fornage M, Gabriel SB, Koboldt DC, Larson DE, et al. The Alzheimer's Disease Sequencing Project: Study design and sample selection. Neurol Genet. 2017;3:e194. Danecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. Twelve years of SAMtools and BCFtools. Gigascience 2021, 10. English AC, Dolzhenko E, Ziaei Jam H, McKenzie SK, Olson ND, De Coster W, Park J, Gu B, Wagner J, Eberle MA, et al. Analysis and benchmarking of small and large genomic variants across tandem repeats. Nat Biotechnol. 2025;43:431–42. Ball MP, Thakuria JV, Zaranek AW, Clegg T, Rosenbaum AM, Wu X, Angrist M, Bhak J, Bobe J, Callow MJ, et al. A public resource facilitating clinical use of genomes. Proc Natl Acad Sci U S A. 2012;109:11920–7. Abel HJ, Larson DE, Regier AA, Chiang C, Das I, Kanchi KL, Layer RM, Neale BM, Salerno WJ, Reeves C, et al. Mapping and characterization of structural variation in 17,795 human genomes. Nature. 2020;583:83–9. Collins RL, Brand H, Karczewski KJ, Zhao X, Alfoldi J, Francioli LC, Khera AV, Lowther C, Gauthier LD, Wang H, et al. A structural variation reference for medical and population genetics. Nature. 2020;581:444–51. Jun G, English AC, Metcalf GA, Yang J, Chaisson MJ, Pankratz N, Menon VK, Salerno WJ, Krasheninina O, Smith AV et al. Structural variation across 138,134 samples in the TOPMed consortium. bioRxiv 2023. McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GR, Thormann A, Flicek P, Cunningham F. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:122. List of AD Loci and Genes with Genetic Evidence Compiled. by ADSP Gene Verification Committee [ https://adsp.niagads.org/gvc-top-hits-list/] Jin Y, Schaffer AA, Sherry ST, Feolo M. Quickly identifying identical and closely related subjects in large databases using genotype data. PLoS ONE. 2017;12:e0179106. Tryka KA, Hao L, Sturcke A, Jin Y, Wang ZY, Ziyabari L, Lee M, Popova N, Sharopova N, Kimura M, Feolo M. NCBI's Database of Genotypes and Phenotypes: dbGaP. Nucleic Acids Res. 2014;42:D975–979. Conomos MP, Miller MB, Thornton TA. Robust inference of population structure for ancestry prediction and correction of stratification in the presence of relatedness. Genet Epidemiol. 2015;39:276–93. Gogarten SM, Sofer T, Chen H, Yu C, Brody JA, Thornton TA, Rice KM, Conomos MP. Genetic association testing using the GENESIS R/Bioconductor package. Bioinformatics. 2019;35:5346–8. Sherva R, Zhang R, Sahelijo N, Jun G, Anglin T, Chanfreau C, Cho K, Fonda JR, Gaziano JM, Harrington KM, et al. African ancestry GWAS of dementia in a large military cohort identifies significant risk loci. Mol Psychiatry. 2023;28:1293–302. Kunkle BW, Grenier-Boley B, Sims R, Bis JC, Damotte V, Naj AC, Boland A, Vronskaya M, van der Lee SJ, Amlie-Wolf A, et al. Genetic meta-analysis of diagnosed Alzheimer's disease identifies new risk loci and implicates Abeta, tau, immunity and lipid processing. Nat Genet. 2019;51:414–30. Xu Z, Li Q, Marchionni L, Wang K. PhenoSV: interpretable phenotype-aware model for the prioritization of genes affected by structural variants. Nat Commun. 2023;14:7805. Frankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong J, Barnes I, et al. Gencode 2021. Nucleic Acids Res. 2021;49:D916–23. Naj AC, Lin H, Vardarajan BN, White S, Lancour D, Ma Y, Schmidt M, Sun F, Butkiewicz M, Bush WS, et al. Quality control and integration of genotypes from two calling pipelines for whole genome sequence data in the Alzheimer's disease sequencing project. Genomics. 2019;111:808–18. STAARpipeline. an all-in-one rare-variant tool for biobank-scale whole-genome sequencing data. Nat Methods. 2022;19:1532–3. Chen H, Huffman JE, Brody JA, Wang C, Lee S, Li Z, Gogarten SM, Sofer T, Bielak LF, Bis JC, et al. Efficient Variant Set Mixed Model Association Tests for Continuous and Binary Traits in Large-Scale Whole-Genome Sequencing Studies. Am J Hum Genet. 2019;104:260–74. Dewey M. metap: Meta-Analysis of Significance Values. 2025. Wang D, Scalici A, Wang Y, Lin H, Pitsillides A, Heard-Costa N, Cruchaga C, Ziegemeier E, Bis JC, Fornage M, et al. Frequency of variants in Mendelian Alzheimer's disease genes within the Alzheimer's Disease Sequencing Project. J Alzheimers Dis. 2025;104:841–51. De Deyn L, Sleegers K. The impact of rare genetic variants on Alzheimer disease. Nat Rev Neurol. 2025;21:127–39. Additional Declarations Competing interest reported. F.J.S. received research support from Illumina, PacBio, and ONT Supplementary Files SupplementaryMaterial.docx SupplementaryTables.xlsx SupplementaryFigures.docx Cite Share Download PDF Status: Under Review Version 1 posted Editorial decision: Revision requested 05 Feb, 2026 Reviews received at journal 04 Feb, 2026 Reviews received at journal 02 Feb, 2026 Reviewers agreed at journal 19 Jan, 2026 Reviewers agreed at journal 17 Jan, 2026 Reviewers invited by journal 16 Jan, 2026 Editor assigned by journal 16 Jan, 2026 Submission checks completed at journal 12 Jan, 2026 First submitted to journal 09 Jan, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8562759","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":573392259,"identity":"b05a5d5b-e4ce-4aa3-9628-f2912d968f69","order_by":0,"name":"Songmi Lee","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYNACAwYZBgbmAwyMDVABHiK0ANWwJTAcbGCQIFILWA2PAXFa5Nt7D7+6UcDAw8/e80364w6GOt32A4wP3rbhcdKZc2nWOUCHSfac3SZx8AyDhNmZBGbDufi0SOSYGYO0GNzIBWppA2q5wcAmzYtHi/wMqBb7+2+ewbSw/8anheFGjvFjsC0SPGxwW5jxaTE4c8aMOQeoXuJMmrHF2TYJyW1nEpsl55zD47D2HuPPOX9s5PjbDz+8Udlmw292/PDBD2/K8DgMGIXAqJCAcUAMeBrACZg/EFIxCkbBKBgFIxwAAOlaSqr2n00HAAAAAElFTkSuQmCC","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":true,"prefix":"","firstName":"Songmi","middleName":"","lastName":"Lee","suffix":""},{"id":573392260,"identity":"011e5446-476f-40a1-9fd1-a25c01bcb9d8","order_by":1,"name":"Adam C English","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Adam","middleName":"C","lastName":"English","suffix":""},{"id":573392261,"identity":"7ae7be57-98c5-4caa-800f-204d22336a7d","order_by":2,"name":"Gina M Peloso","email":"","orcid":"","institution":"Boston University","correspondingAuthor":false,"prefix":"","firstName":"Gina","middleName":"M","lastName":"Peloso","suffix":""},{"id":573392262,"identity":"7c38dcfd-8c38-4181-b1e9-d4cbac12bd6a","order_by":3,"name":"Joshua C Bis","email":"","orcid":"","institution":"University of Washington","correspondingAuthor":false,"prefix":"","firstName":"Joshua","middleName":"C","lastName":"Bis","suffix":""},{"id":573392264,"identity":"854f4687-776d-48ec-a6ea-849e8ff263fe","order_by":4,"name":"Eric Boerwinkle","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"prefix":"","firstName":"Eric","middleName":"","lastName":"Boerwinkle","suffix":""},{"id":573392265,"identity":"515b223b-9946-4ae4-ad4a-ee723b4c879a","order_by":5,"name":"Seung Hoan Choi","email":"","orcid":"","institution":"Boston University","correspondingAuthor":false,"prefix":"","firstName":"Seung","middleName":"Hoan","lastName":"Choi","suffix":""},{"id":573392267,"identity":"36a74f22-217b-4fef-8a02-7150458a6d8f","order_by":6,"name":"Nancy L Heard-Costa","email":"","orcid":"","institution":"Framingham Heart Study","correspondingAuthor":false,"prefix":"","firstName":"Nancy","middleName":"L","lastName":"Heard-Costa","suffix":""},{"id":573392268,"identity":"a84834bd-e287-4339-8eda-0343f8eee34a","order_by":7,"name":"Honghuang Lin","email":"","orcid":"","institution":"University of Massachusetts Chan Medical School","correspondingAuthor":false,"prefix":"","firstName":"Honghuang","middleName":"","lastName":"Lin","suffix":""},{"id":573392269,"identity":"1954466a-4b53-4f08-8852-8c0cf3dfd6bb","order_by":8,"name":"Rui Xia","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"prefix":"","firstName":"Rui","middleName":"","lastName":"Xia","suffix":""},{"id":573392274,"identity":"327559f8-8a26-415c-95cd-5ff964830cd8","order_by":9,"name":"Sudha Seshadri","email":"","orcid":"","institution":"The University of Texas Health Science Center at San Antonio","correspondingAuthor":false,"prefix":"","firstName":"Sudha","middleName":"","lastName":"Seshadri","suffix":""},{"id":573392276,"identity":"5fbf69b2-2d0d-48c2-80df-d4c776345c59","order_by":10,"name":"Anita L Destefano","email":"","orcid":"","institution":"Boston University","correspondingAuthor":false,"prefix":"","firstName":"Anita","middleName":"L","lastName":"Destefano","suffix":""},{"id":573392279,"identity":"0783491e-e58b-48fb-aa42-220d5dd15cdb","order_by":11,"name":"Myriam Fornage","email":"","orcid":"","institution":"The University of Texas Health Science Center at Houston","correspondingAuthor":false,"prefix":"","firstName":"Myriam","middleName":"","lastName":"Fornage","suffix":""},{"id":573392281,"identity":"98160cca-684d-44d6-ae1f-bf08af657636","order_by":12,"name":"Fritz J Sedlazeck","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Fritz","middleName":"J","lastName":"Sedlazeck","suffix":""}],"badges":[],"createdAt":"2026-01-09 16:08:33","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8562759/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8562759/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":100112502,"identity":"8de5e565-bb48-43e7-b0c3-24c9efbe0261","added_by":"auto","created_at":"2026-01-13 06:59:16","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":5345441,"visible":true,"origin":"","legend":"","description":"","filename":"TheImpactofStructuralVariationonAD.docx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/ca283241c4b3891f2807a9e6.docx"},{"id":100112488,"identity":"96d1ada9-7e18-4a54-8534-974c7d61c33d","added_by":"auto","created_at":"2026-01-13 06:59:15","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":14044,"visible":true,"origin":"","legend":"","description":"","filename":"98254494866b4acfba4d8dec3764483f.json","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/b84d031c09b53b28806f37d8.json"},{"id":100367025,"identity":"a206b4e4-7cd0-40b6-849a-715086dd72da","added_by":"auto","created_at":"2026-01-16 07:56:44","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":2520297,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/c415f081714ce1a35fb6da7c.docx"},{"id":100365755,"identity":"532e4691-e260-4d8d-b46e-8cf45257509d","added_by":"auto","created_at":"2026-01-16 07:55:35","extension":"docx","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":515388,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/727b3b18772aea920cb2de69.docx"},{"id":100112494,"identity":"f5a70717-f975-4806-8ba1-521efce58b33","added_by":"auto","created_at":"2026-01-13 06:59:15","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":54952,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/cd6b1992df48e97d70fd211b.xlsx"},{"id":100112500,"identity":"2489826c-dc09-4523-86f3-1d9693c52b5c","added_by":"auto","created_at":"2026-01-13 06:59:16","extension":"xml","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":213938,"visible":true,"origin":"","legend":"","description":"","filename":"98254494866b4acfba4d8dec3764483f1enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/c849fd6ecc42afdb24c10139.xml"},{"id":100365629,"identity":"2a88fee6-2def-4ab5-bab3-8b56e217520e","added_by":"auto","created_at":"2026-01-16 07:55:26","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":200588,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/1f66767479269cd213c2d925.png"},{"id":100365076,"identity":"fc23bceb-cc19-4e03-a9e5-d69d8fbf0db9","added_by":"auto","created_at":"2026-01-16 07:54:39","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":42820,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/0dfd47c0c8e84ead9d925c83.png"},{"id":100365830,"identity":"691ebaa6-2b5e-47c8-b0a9-ef3864b6732a","added_by":"auto","created_at":"2026-01-16 07:55:40","extension":"jpeg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":245974,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/c8e542d23b92452daea5351f.jpeg"},{"id":100112492,"identity":"cb27ee1b-fd53-4cdb-8df0-f31ac3ef72f1","added_by":"auto","created_at":"2026-01-13 06:59:15","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":58479,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/d22ae64b499ae219e8097617.png"},{"id":100112497,"identity":"49ec220b-d3da-4af9-a92e-361f1a2d06c2","added_by":"auto","created_at":"2026-01-13 06:59:15","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":12733,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/c01ff42ff618a6f925d2f544.png"},{"id":100365089,"identity":"5d826a25-669a-4082-8f76-d4713a32f5fa","added_by":"auto","created_at":"2026-01-16 07:54:39","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":28094,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/af8ce269ce5531eebdc84cce.png"},{"id":100367043,"identity":"8b57f2d4-38ef-4cf4-b955-36baab4b5aac","added_by":"auto","created_at":"2026-01-16 07:56:44","extension":"xml","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":212005,"visible":true,"origin":"","legend":"","description":"","filename":"98254494866b4acfba4d8dec3764483f1structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/267ba4be579c110ad09f0cca.xml"},{"id":100112505,"identity":"2dba7ec2-0862-4701-9cc5-a6c766f649b9","added_by":"auto","created_at":"2026-01-13 06:59:16","extension":"html","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":230818,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/da07665925fc30b4ddb5b1a7.html"},{"id":100112487,"identity":"4a0785ce-7ae0-4588-9353-e32cb4196d32","added_by":"auto","created_at":"2026-01-13 06:59:15","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":227701,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eStudy Overview.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWhole-genome sequencing (WGS) data from 16,841 ADSP participants were analyzed using a novel SV calling pipeline (BioGraph), followed by extensive quality control at both sample and variant levels, resulting in 456,644 high-quality SVs. Samples were stratified into three ancestry groups: African (AFR), European (EUR), and LAT (Latin), based on genetic similarity. Common or low frequency SVs were evaluated using single variant association analyses, including analyses of SVs in known AD loci previously reported in genome-wide association study (GWAS) (A) as well as genome-wide association analyses (B). Rare SVs were assessed using gene-based association analyses as primary analyses, separately evaluating noncoding pathogenic SVs (C) and coding pathogenic SVs (D). Secondary gene-based analyses incorporated single-nucleotide variants (SNVs) and small insertion/deletions (INDELs) in combination with SVs to increase statistical power for noncoding (E) and coding (F) variants.\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/0ab44cefbf7011fa96c17703.png"},{"id":100364902,"identity":"2bfe3acc-83a4-4b3d-adad-524535925f68","added_by":"auto","created_at":"2026-01-16 07:54:28","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":81052,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSize and allele frequency distributions of SVs in the ADSP\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Distribution of SV sizes by variant type across the ADSP participants. SVs are grouped into size bins and colored by type: deletion (blue), insertion (orange), and inversion (green). The y-axis indicates the number of SVs in each size of category. (B) Site frequency spectrum of SVs by variant type, showing the distribution of allele frequencies for deletion (blue), insertion (orange), and inversion (green). The y-axis is shown on a logarithmic scale.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/1125ec2c62ac53d2913c9c9b.png"},{"id":100382220,"identity":"21726a15-0263-45e9-9311-e05541c50672","added_by":"auto","created_at":"2026-01-16 10:41:29","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1828677,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/b92b3a95-b038-403a-8b3e-678cd37dd334.pdf"},{"id":100112489,"identity":"054edbca-7e91-4b8c-b14c-da287514c8e2","added_by":"auto","created_at":"2026-01-13 06:59:15","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":515388,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.docx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/b7a9aa2dc093b0d102511ca3.docx"},{"id":100365523,"identity":"c3e9ed5f-1378-4755-9716-f0687d272575","added_by":"auto","created_at":"2026-01-16 07:55:18","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":54952,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryTables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/67ae8d0086d3b8bef4330b90.xlsx"},{"id":100112503,"identity":"b304d64f-1c82-4d5b-8bd8-d2e693ab8a46","added_by":"auto","created_at":"2026-01-13 06:59:16","extension":"docx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":2520297,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryFigures.docx","url":"https://assets-eu.researchsquare.com/files/rs-8562759/v1/6abfcb60d53ac9f7fd94bed7.docx"}],"financialInterests":"Competing interest reported. F.J.S. received research support from Illumina, PacBio, and ONT","formattedTitle":"The Impact of Structural Variation on Alzheimer’s Disease in the Alzheimer’s Disease Sequencing Project","fulltext":[{"header":"1. Background","content":"\u003cp\u003eAlzheimer\u0026rsquo;s Disease (AD), characterized by progressive memory loss and declining cognitive function, is the most common form of dementia among older adults. It is estimated that almost 7\u0026nbsp;million Americans aged 65 and older are currently living with AD and this number is projected to double by 2060 [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. To date, there are few effective disease-modifying treatments or prevention strategies for AD, underscoring a critical need to better understand its etiology.\u003c/p\u003e \u003cp\u003eGenetic factors play a substantial role in AD, with a disease heritability estimated between 60\u0026ndash;80% [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. To date, genetic approaches have yielded important novel insights into AD etiology [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e] and promise further advances in prevention, diagnosis, and treatment. Very rare highly-penetrant mutations have been identified in the Amyloid Beta Precursor Protein (\u003cem\u003eAPP\u003c/em\u003e), Presenilin 1 (\u003cem\u003ePSEN1\u003c/em\u003e) and Presenilin 2 (\u003cem\u003ePSEN2\u003c/em\u003e) genes that cause Mendelian forms of AD, typically with early onset [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. More common alleles identified in large genome-wide association studies (GWAS) of sporadic late-onset AD have uncovered genes involved in cholesterol metabolism, endocytosis/phagocytosis, amyloid plaque and neurofibrillary tangle formation, and the innate immune system [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Despite significant progress in understanding the genetic basis of AD, a substantial proportion of AD genetic architecture remains unknown [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. Addressing this gap in knowledge requires a comprehensive characterization of all forms of genetic variation, beyond single nucleotide variants (SNVs), and including structural variation.\u003c/p\u003e \u003cp\u003eStructural variants (SVs) are typically defined as genomic alterations comprising 50 or more base pairs [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. These variants can be classified into five different types: insertions, deletions, inversions, duplications, and translocations. Compared to SNVs, SVs are numerically fewer but are larger in size, and therefore have a greater impact on DNA sequence and, consequently, on gene expression and function [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. SVs often occur in highly repetitive and polymorphic regions of the genome [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e], making them challenging to detect with short-read DNA sequencing technology [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Insertions are especially problematic and their role across diseases remains understudied [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. Over the past decade, technological and methodological developments have improved SV detection from short-read whole genome sequence (WGS) data [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e], providing an opportunity to more comprehensively and accurately evaluate their impact on complex disease etiology.\u003c/p\u003e \u003cp\u003eIn this study, we have implemented a novel SV calling method, BioGraph [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], on the WGS data from 16,841 subjects of the Alzheimer\u0026rsquo;s Disease Sequencing Project (ADSP, Release 3) and examined the association of the detected SVs with AD. Our study investigated the role of SVs in known AD loci, providing insights into the genetic architecture of AD. We also examined associations between common SVs across the genome and AD, and performed gene-based association testing to analyze rare SVs in AD.\u003c/p\u003e"},{"header":"2. Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Identification of high-quality SVs using BioGraph in the ADSP\u003c/h2\u003e \u003cp\u003eWe implemented our novel method BioGraph [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] to generate a highly accurate SV call set from 16,841 WGS in the ADSP (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). BioGraph is a unique approach to SV detection and genotyping that leverages reference guided assembly of short reads to improve the detection of SV. Additionally, BioGraph uses machine learning techniques to assign useful quality scores to the identified candidate SVs. Comparison of BioGraph performance in detecting SV with that of other SV calling tools, including Manta [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], Parliament2 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], and Smoove [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], using benchmark data from the Genome In A Bottle HG002 Challenging Medically Relevant Regions [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e] and Truvari [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] are described in detail in the \u003cb\u003eSupplementary Material\u003c/b\u003e.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eOur initial raw set of SV calls across the ADSP\u0026rsquo;s 16,841 sample set generated 1,019,035 SVs. We deployed novel sample-based filtering approaches to further ensure high quality and accuracy of SV calls across the samples. We first leveraged technical replicates present in the ADSP data. In total, there are 601 replicate samples derived from 283 unique individuals providing 428 replicate pairs. Using Truvari, we compared the SV calls between each replicate pair. First, we compared the quality score distributions of calls which were found consistently between replicates to those which were inconsistent (\u003cb\u003eSupplementary Fig.\u0026nbsp;1\u003c/b\u003e). We determined a quality score threshold of 50 best segregates the calls by their consistency. These data and additional data examining the quality score\u0026rsquo;s relationship to False Positive and True Positive measurements in benchmarking experiments (\u003cb\u003eSupplementary Material\u003c/b\u003e) suggest a minimum quality score of 50 should be applied for high quality calls. In total, 562,391 SVs were removed using this filter.\u003c/p\u003e \u003cp\u003eWe also analyzed the similarity of the consistent calls between replicate pairs. Over 97% of consistent calls have a sequence/size similarity of 95% or greater, suggesting that SVs with \u0026ge;\u0026thinsp;95% sequence and size similarity should be considered the same when performing inter-sample merging. Our rigorous QC procedure enabled us to assign pass and fail values across all SVs. Moreover, we collapsed and filtered 55.2% of the initially inferred SV and thus avoided many potential false positive or redundant alleles.\u003c/p\u003e \u003cp\u003eAfter QC, we identified 456,644 SVs, including 254,716 deletions (55.78%), 190,786 insertions (41.78%) and 11,142 inversions (2.44%). The distribution of SV size and type across the ADSP sample is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA. Most SVs identified were less than 5 kbp in length (94.74%) with a majority ranging from 50\u0026ndash;100 bp (42.18%). We observed the expected ALU peak (~\u0026thinsp;300 bp) for deletions but even a more prominent peak for insertions. The latter is due to our merging strategy from Truvari where the SV are only merged if their sequence similarity exceeds 95% in addition to type and length constraints [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e].\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe number of deletions and insertions were similar across most size categories with some exceptions. We identified fewer insertions in size categories larger than 1 kbp. This contrasts with other short read SV calling approaches where the number of insertions declines rapidly starting at 500 bp [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. We also observed a size bias across inversions with only 210 inversions over 1 kbp detected. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB shows the site frequency spectrum (log scale). Overall, we observed that the insertions and deletions occur at similar frequencies and the majority of them (87.3%) have an allele frequency\u0026thinsp;\u0026lt;\u0026thinsp;1%. Most of the inversions are singletons (70.7%) compared to deletions and insertions for which the proportion of singletons is 37.2% and 33.6%, respectively. These findings reflect the challenge to correctly infer inversions from short-read sequence data [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The singleton rate appears to vary with SV size, being lowest in smaller SVs (50\u0026ndash;100 bp: 27.4% singletons) compared to midsize SVs (1 kbp-2.5 kbp: 51.1%) and large SVs (\u0026gt;\u0026thinsp;5 kbp: 60.7% singletons). We speculate that this is likely due to a combination of larger events accumulating mutations over time as well as larger events being less consistently discovered.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. A total of 297,034 SVs are novel in the ADSP sample\u003c/h2\u003e \u003cp\u003eWe evaluated the overlap of our high-quality SV call set with previously reported SVs from major reference databases, including the Center for Common Disease Genomics (CCDG), Trans-omics for Precision Medicine (TOPMed), 1000Genomes Project (1KGP), and the Genome Aggregation Database (gnomAD). Only 34.9% of the SV identified in the ADSP overlapped with the reported SVs. The greatest overlap was found with gnomAD data (23.3% of ADSP SV calls), followed by 1KGP (19.4%). When investigating overlap by SV type, 18.9% of the overlapping SV were deletions, with the greatest overlap with gnomAD (12.33%); and 16% were insertions, with the greatest overlap with 1KGP (27.6%) rather than gnomaAD (26.3%). This may be due to tandem duplications in gnomAD reported as insertions in the 1KGP and our call sets, a common SV type swap [\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe investigated the correlation of allele frequency among common overlapping SVs (MAF\u0026thinsp;\u0026gt;\u0026thinsp;1%), and observed a strong overall concordance (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.73, P\u0026thinsp;\u0026lt;\u0026thinsp;0.01), underscoring the accuracy and reliability of the SV call set. As expected, the correlation was higher for deletions (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.89, P\u0026thinsp;\u0026lt;\u0026thinsp;0.01) compared to insertions (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.63, P\u0026thinsp;\u0026lt;\u0026thinsp;0.01). Interestingly, these correlations improved further when comparing the ADSP SV call set with external reference datasets. For example, the correlation of overlapping SVs in Biograph ADSP SV call sets and in gnomAD call sets was 0.94 for deletions and 0.82 for insertions.\u003c/p\u003e \u003cp\u003eNotably, many of the novel SVs identified in our study were within size ranges that were more effectively captured by BioGraph than methods relying solely on paired-end read distances. For example, the average deletion size in gnomAD was 7.4 kbp compared to 2.1 kbp in the Biograph ADSP data. For insertions, these numbers were 895 bp vs. 184 bp, respectively. These findings further highlight the strength of our dataset including SVs that may have been under-called by previous short-read studies.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Annotation and overlap of SVs with genes\u003c/h2\u003e \u003cp\u003eWe identified SVs that directly overlapped or were in close proximity (within 5 kbp) to gene sequences using SVAfotate [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. All 456,644 SVs were annotated, of which 143,397 SVs (31.4%) mapped to intergenic regions, 131,988 SVs (28.9%) were reported in proximity of genes but not overlapping them directly, and 272,728 SVs (59.7%) directly impacted a gene. These 272,728 SVs mapped to a total of 30,333 (62.8%) genes, suggesting that they are impacting the same gene more than once across different individuals. The majority of these gene-impacting SVs were deletions (53.3%) followed by insertions (36.6%) and inversions (10.1%). The lower number of gene-impacting insertions compared to deletions is likely because insertions are measured as affecting only the direct base pair at which they are reported, while deletions span multiple base pairs on the reference. The majority (96.9%) of gene-impacting SVs were located within introns whereas 3.0% mapped to the 3\u0026rsquo; untranslated region (UTR) and only 1.6% mapped to the 5\u0026rsquo; UTR. The SVs mapped to UTR have a higher chance to impact regulatory function itself. Only 2.3% of SVs were directly overlapped coding sequences.\u003c/p\u003e \u003cp\u003eAmong 86 AD genes reported by the ADSP Gene Verification Committee, 82 genes intersected 1,223 SVs. Most genes (71) had SVs within \u0026plusmn;\u0026thinsp;5kbp as well as overlapping the gene body. Filtering to only common (AF\u0026thinsp;\u0026ge;\u0026thinsp;1%) SVs hitting non-intronic gene bodies left 69 variants over 30 genes (\u003cb\u003eSupplementary Table\u0026nbsp;1\u003c/b\u003e). SVs of note included a 6,137bp deletion on \u003cem\u003ePRDM7\u003c/em\u003e (AF\u0026thinsp;=\u0026thinsp;2.4%), 4 tandem repeat expansions between 51bp and 98bp of a 12bp VNTR in \u003cem\u003eRBCK1\u003c/em\u003e, and a 322bp deletion on \u003cem\u003eTMEM106B\u003c/em\u003e with a frequency of 49.7%.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. PCA and ancestry inference\u003c/h2\u003e \u003cp\u003eWe derived principal components from our SV data (N\u0026thinsp;=\u0026thinsp;12,908) to account for possible population structure in the data. PC2 was associated with read length and demonstrated complete separation of samples as shown in \u003cb\u003eSupplementary Fig.\u0026nbsp;2\u003c/b\u003e. To minimize confounding by batch effects in the association analyses, study participants were further restricted to those with a read-length of 150 (N\u0026thinsp;=\u0026thinsp;11,890; 5,585 cases and 6,305 controls), which represents more than 90% of the sample. PCA analyses revealed similarities in results between PCs derived from SVs and those derived from SNVs (\u003cb\u003eSupplementary Fig.\u0026nbsp;3\u003c/b\u003e). Based on the results of GrafPop (\u003cb\u003eSupplementary Fig.\u0026nbsp;4\u003c/b\u003e), our study included 3,371 individuals of African (AFR) ancestry, 6,327 of European (EUR) ancestry, 2,126 of Latin (LAT) ancestry, and 66 participants that did not cluster with those three ancestry groups and were therefore excluded from subsequent analyses (\u003cb\u003eSupplementary Table\u0026nbsp;2\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5. SVs in linkage disequilibrium (LD) with Alzheimer\u0026rsquo;s Disease known loci\u003c/h2\u003e \u003cp\u003eAncestry-specific LD analyses identified 9 SVs in EUR, 5 SVs in AFR, and 9 SVs in LAT that were in moderate or high LD (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.43\u0026ndash;0.99) with at least one of the SNPs previously identified in AD GWAS (\u003cb\u003eSupplementary Tables\u0026nbsp;3\u0026ndash;5\u003c/b\u003e). These SNPs did not exhibit strong associations with AD in our dataset, due to limited statistical power compared to the GWAS sample size in which they were discovered. Among the identified SVs, 4 SVs in EUR, 1 SV in AFR, and 1 SV in LAT showed suggestive evidence of association with AD (SV P-value\u0026thinsp;\u0026lt;\u0026thinsp;0.1).\u003c/p\u003e \u003cp\u003eThe strongest SV association with AD was observed in AFR, involving a 122-bp deletion in moderate LD (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.46) with rs2633682 tagging the \u003cem\u003eALCAM\u003c/em\u003e locus (\u003cb\u003eSupplementary Table\u0026nbsp;4\u003c/b\u003e). This SNP, previously associated with AD specifically in an African American population [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e], showed a suggestive association with AD in our dataset (SNP P-value\u0026thinsp;=\u0026thinsp;0.009). However, conditional analyses indicated that neither the SVs nor the SNP remained significant after adjusting for each other, suggesting non-independence of the signals at the \u003cem\u003eALCAM\u003c/em\u003e locus. In EUR, a similar trend was observed at the \u003cem\u003eALCAM\u003c/em\u003e locus, although the associations were weaker.\u003c/p\u003e \u003cp\u003eA 319-bp deletion in an intron of \u003cem\u003eTPCN1\u003c/em\u003e was observed in all three ancestry groups and was in high LD (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.97 in all groups) with the tagging SNP. In EUR and LAT, two deletions, including the 319-bp deletion, and one insertion were in moderate or high LD with a SNP tagging at the \u003cem\u003eTPCN1\u003c/em\u003e locus (\u003cb\u003eSupplementary Tables\u0026nbsp;3 and 5\u003c/b\u003e). Haplotype estimation analysis suggested that all detected SVs lie on the same haplotype as the AD risk allele at this locus (\u003cb\u003eSupplementary Fig.\u0026nbsp;5\u003c/b\u003e). A 68-bp deletion located in an intron of \u003cem\u003eSLC8B1\u003c/em\u003e were detected in EUR and LAT, with suggestive association with AD observed in EUR (SV P-value\u0026thinsp;=\u0026thinsp;0.03). While the intronic SNP (rs6489896) tagging \u003cem\u003eTPCN1\u003c/em\u003e has been previously associated with AD at genome-wide significance, \u003cem\u003eSLC8B1\u003c/em\u003e has not. Additionally, a 322-bp Alu deletion in exon 8 and 3\u0026rsquo; untranslated region of \u003cem\u003eTMEM106B\u003c/em\u003e exhibited strong LDs with two tagging SNPs in EUR and LAT, and moderate LD in AFR.\u003c/p\u003e \u003cp\u003eConditional analyses adjusting for the corresponding SNPs revealed that none of the SV associations remained (adjusted SV P-value\u0026thinsp;\u0026gt;\u0026thinsp;0.1), indicating that the observed suggestive SV associations were not independent of SNPs in LD at those loci. However, in LAT (\u003cb\u003eSupplementary Table\u0026nbsp;5\u003c/b\u003e), a SNP tagging \u003cem\u003eWNT3\u003c/em\u003e/\u003cem\u003eMATP\u003c/em\u003e locus remained significant (adjusted SNP P-value\u0026thinsp;=\u0026thinsp;0.006) after conditioning on a 314-bp deletion in moderate LD (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.68), suggesting that the SNP is independently associated with AD at this locus.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6. Genome-wide association of common or low-frequency SVs with Alzheimer\u0026rsquo;s Disease\u003c/h2\u003e \u003cp\u003eFor each ancestry group, we performed single variant association analyses of high-quality SVs with MAF\u0026thinsp;\u0026gt;\u0026thinsp;0.5%, and in HWE (as defined in section 4.8). In total, we analyzed 28,942 SVs in AFR, 14,656 SVs in EUR, and 30,394 SVs in LAT (\u003cb\u003eSupplementary Table\u0026nbsp;6\u003c/b\u003e). No SVs were significantly associated with AD in the AFR or LAT analyses, or in the meta-analyses. In EUR, two deletions were significantly associated with AD at the Bonferroni-corrected threshold and were observed exclusively in this ancestry group (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e; \u003cb\u003eSupplementary Fig.\u0026nbsp;6\u003c/b\u003e). Both deletions mapped to introns of genes encoding coiled-coil domain containing proteins.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBonferroni-significant SVs associated with AD identified by genome-wide single variant association analyses\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"15\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c12\" colnum=\"12\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c13\" colnum=\"13\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c14\" colnum=\"14\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c15\" colnum=\"15\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e Group\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSV type\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSV\u003c/p\u003e \u003cp\u003elocus\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSV Breakpoints\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLength\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eQUAL\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eGene(s)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"5\" nameend=\"c13\" namest=\"c9\"\u003e \u003cp\u003eAllele Frequencies\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c15\" namest=\"c14\"\u003e \u003cp\u003eSingle variant association tests\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eAFR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c11\" namest=\"c10\"\u003e \u003cp\u003eEUR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c13\" namest=\"c12\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c14\"\u003e \u003cp\u003eOR [95% CI]\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c15\"\u003e \u003cp\u003eP-value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEUR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDEL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3p21.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e46978504\u0026ndash;46978622\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e119\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003eCCDC12\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eintron\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c10\" namest=\"c9\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c12\" namest=\"c11\"\u003e \u003cp\u003e0.013\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e3.19 [2.1\u0026ndash;4.9]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c15\"\u003e \u003cp\u003e7.66E-08\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eEUR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDEL\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e11q13.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e64341844\u0026ndash;64341923\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003eCCDC88B\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eintron\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c10\" namest=\"c9\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"2\" nameend=\"c12\" namest=\"c11\"\u003e \u003cp\u003e0.011\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e2.89 [1.9\u0026ndash;4.4]\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c15\"\u003e \u003cp\u003e2.09E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"15\"\u003eAFR, African ancestry group; LAT, Latin ancestry group; EUR, European ancestry group; DEL, deletion; QUAL, quality score; OR, odds ratio; CI, confidence interval\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eBonferroni-significant or suggestive genes associated with AD in gene-based association analyses of coding variants (A) and non-coding variants (B)\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e \u003cp\u003eA. Coding SNVs/INDELs\u0026thinsp;+\u0026thinsp;SVs\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGroup\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGene\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLocus\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eCategory\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003en variants\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ecMAC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eP-value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003ePSEN1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003epLof\u0026thinsp;+\u0026thinsp;disruptive missense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.96E-07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003ePSEN1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003emissense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e8.69E-07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eSMOC1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003esynonymous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.50E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT4\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003edisruptive missense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e35\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.72E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT4\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003emissense\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5.44E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT4\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003esynonymous\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5.70E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colspan=\"7\" nameend=\"c7\" namest=\"c1\"\u003e \u003cp\u003e\u003cb\u003eB. Non-coding SNVs/INDELs\u0026thinsp;+\u0026thinsp;SVs\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eGroup\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eGene\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003eLocus\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eCategory\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003en variants\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003ecMAC\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003eP-value\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eELMSAN1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003encRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e1.86E-09\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eAC005225.2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003encRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e44\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e6.61E-08\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eLOC100506476\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePromoter (CAGE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e7.17E-08\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eLOC100506476\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnhancer (CAGE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e26\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e103\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.03E-07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT6\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePromoter (DHS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e104\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e3.61E-07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eAL390763.1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e10q26.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003encRNA\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.24E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT4\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnhancer, Promoter (CAGE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.53E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eAC005225.2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePromoter (DHS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e43\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e263\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5.09E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT4\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnhancer (DHS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e5.25E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eLINC01500\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePromoter (CAGE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e6.25E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eTIRAP\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e11q24.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUTR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e7.43E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003ePROX2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnhancer (DHS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e27\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e130\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e7.49E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT4\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePromoter (DHS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e103\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.08E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eAL163974.1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q32.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eUpstream\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e90\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e9.88E-06\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLAT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eACOT6\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e14q24.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eEnhancer (DHS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e107\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.29E-07\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003ctfoot\u003e \u003ctr\u003e\u003ctd colspan=\"7\"\u003eSNV, single nucleotide variants; INDELs, insertion and deletions; pLof, putative loss of function; cMAC, cumulative minor allele count; ncRNA, non-coding RNA; CAGE, cap analysis of gene expression; DHS, DNase I hypersensitive site; UTR, untranslated region\u003c/td\u003e\u003c/tr\u003e \u003c/tfoot\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7. Gene-based association of rare SVs and SNVs with Alzheimer\u0026rsquo;s Disease\u003c/h2\u003e \u003cp\u003eWe performed gene-based analyses to test association between aggregated rare SVs and AD. In the primary gene-based association analyses, two analyses were carried out based on the SV types: coding SVs (\u003cb\u003eSupplementary Fig.\u0026nbsp;7\u003c/b\u003e) and noncoding SVs (\u003cb\u003eSupplementary Fig.\u0026nbsp;8\u003c/b\u003e). No genes reached genome-wide significant or suggestive significance thresholds in either analysis.\u003c/p\u003e \u003cp\u003eIn the secondary analyses, we conducted gene-based analyses to assess association between AD and aggregates of rare SVs and rare SNVs/INDELs. Coding variant analyses were performed using five categories of SNVs/INDELs, combined with pathogenic coding SVs (\u003cb\u003eSupplementary Fig.\u0026nbsp;9\u003c/b\u003e), while non-coding variant analyses included eight categories of SNVs/INDELs in combination with pathogenic non-coding SVs (\u003cb\u003eSupplementary Fig.\u0026nbsp;10\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eIn the secondary coding variant analyses, we identified \u003cem\u003ePSEN1\u003c/em\u003e as being suggestively associated with AD in LAT when aggregating SVs and SNVs/INDELs classified as pLoF and disruptive missense (P\u0026thinsp;=\u0026thinsp;1.9E-07), disruptive missense (P\u0026thinsp;=\u0026thinsp;1.9E-07), or missense (P\u0026thinsp;=\u0026thinsp;8.6E-07) variants (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA). For this region, we observed two deletions and two insertions, all more frequent or exclusively observed in AD controls (\u003cb\u003eSupplementary Table\u0026nbsp;7\u003c/b\u003e). Conditional analyses on the aggregates of SNVs/INDELs in \u003cem\u003ePSEN1\u003c/em\u003e indicated that the four pathogenic coding SVs were independently associated with AD (adjusted P-values\u0026thinsp;\u0026lt;\u0026thinsp;0.05). No SVs were detected for the \u003cem\u003eSMOC1\u003c/em\u003e and \u003cem\u003eACOT4\u003c/em\u003e regions. In the secondary non-coding variant analyses, multiple genes showed significant or suggestive associations with AD in LAT (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB); however, no SVs were detected for those genes, and thus the observed associations were entirely driven by SNVs/INDELs. Not surprisingly, several of the associated genes overlapped with those identified in a previous study based solely on SNVs and INDELs, including \u003cem\u003eELMSAN1\u003c/em\u003e, \u003cem\u003eACOT6\u003c/em\u003e, and \u003cem\u003eACOT4\u003c/em\u003e [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn the candidate gene analyses, we examined 15 previously reported AD-associated genes to determine whether rare SVs contributed to their associations signals. No genes reached statistical significance in the primary analyses. In the secondary coding variant analyses, \u003cem\u003ePSEN1\u003c/em\u003e, \u003cem\u003eTREM2\u003c/em\u003e, and \u003cem\u003eABCA7\u003c/em\u003e showed evidence of association with AD (FDR Q\u0026thinsp;\u0026lt;\u0026thinsp;0.05) (\u003cb\u003eSupplementary Table\u0026nbsp;8A\u003c/b\u003e). Conditional analyses on the aggregates of SNVs/INDELs within each gene suggested independent effects of coding SVs for \u003cem\u003ePSEN1\u003c/em\u003e and \u003cem\u003eABCA7\u003c/em\u003e (adjusted P-value\u0026thinsp;\u0026lt;\u0026thinsp;0.05). Notably, \u003cem\u003eABCA7\u003c/em\u003e showed evidence of an independent SV association with AD, driven by a 605 bp exonic deletion (chr19:1050368\u0026ndash;1050972) observed exclusively in four AD cases within AFR. This association remained significant after conditioning on disruptive missense SNVs/INDELs, indicating an independent effect of the deletion on AD risk. In the secondary non-coding variant analyses, \u003cem\u003eTREM2\u003c/em\u003e and \u003cem\u003eABI3\u003c/em\u003e showed evidence of association with AD, however, no evidence of independent SV association with AD was observed (\u003cb\u003eSupplementary Table\u0026nbsp;8B\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Discussion","content":"\u003cp\u003eIn this study, we analyzed SVs in 16,841 individuals from ADSP using BioGraph, a novel SV calling pipeline. We identified 456,644 high-quality SVs, approximately 65% of which were novel. Notably, the vast majority of novel SVs were insertions, which may have been under-detected in previous studies. Among common or low-frequency SVs within each ancestry group, several SVs were found to be in moderate or high LD with known AD loci, offering additional insights into the genetic architecture of AD. Genome-wide association analyses identified two low-frequency deletions associated with AD in individuals of European ancestry, both located within genes encoding coiled-coil domain-containing proteins. Gene-based analyses further revealed that \u003cem\u003ePSEN1\u003c/em\u003e and \u003cem\u003eABCA7\u003c/em\u003e harbor rare pathogenic SVs associated with AD.\u003c/p\u003e \u003cp\u003eOur use of BioGraph, a semi-assembly-based SV calling method, enabled the identification of many insertions not previously reported. Long-read sequencing and genome assembly studies have shown that insertions are the most prevalent SV class, often representing tandem repeat expansions or transposable element integrations that are not in the reference genome. This fact makes insertions challenging to identify, but also biologically intriguing as they have been reported to affect splicing or induce mosaic variants in proximity (e.g. ALUY insertions) [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. Our SV call set demonstrated high accuracy and precision for both insertions and deletions, supported by rigorous benchmarking on replicate samples within ADSP and assessment across control samples. We have further introduced detection of inversions from BioGraph results that yielded multiple inversions candidates. This is noteworthy as the correct identification of inversions remains highly challenging [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe comprehensively assessed SVs and their potential impact on AD. Notably, many genes previously highlighted by SNV-based GWAS exhibited SVs either within the gene itself or within 5 kb. Overall, 95.3% of the postulated genes showed SV overlap. To further explore the role of SVs in established AD loci while accounting for ancestral differences, we examined SVs in known AD loci within each inferred ancestry group. We identified several SVs in moderate or high LD with known AD loci across different ancestry groups. At the \u003cem\u003eTPCN1\u003c/em\u003e locus, a 319-bp intronic deletion was observed across the three ancestry groups in high LD with the tagging SNP. This deletion fully overlaps with a previously reported 309-bp deletion associated with Lewy Body dementia, which was validated using long-read sequencing [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. \u003cem\u003eTPCN1\u003c/em\u003e, which is highly expressed in the brain, encodes the two-pore calcium channel protein 1 located on endolysosomal membranes. Beyond its association with AD identified in previous GWAS [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], the function of \u003cem\u003eTPCN1\u003c/em\u003e has been demonstrated in knockout mice, which exhibit impairments in spatial learning and memory [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. In both EUR and LAT, we additionally discovered a deletion and an insertion at the \u003cem\u003eTPCN1\u003c/em\u003e locus that showed evidence of association with AD in EUR. The deletion was located in \u003cem\u003eSLC8B1\u003c/em\u003e, which encodes a mitochondrial Na+/Ca2\u0026thinsp;+\u0026thinsp;exchanger. Notably, a recent study demonstrated that a deletion of the \u003cem\u003eSLC8B1\u003c/em\u003e region in knockout mice, spanning the region of our four deletions, is sufficient to induce AD-like pathology, including age-related cognitive decline [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. Our haplotype estimation analysis suggested that all detected SVs lie on the same haplotype as the AD risk allele at the \u003cem\u003eTPCN1\u003c/em\u003e locus. These findings suggest that multiple genes within the \u003cem\u003eTPCN1\u003c/em\u003e locus may influence AD risk and underscore the need for further investigation into the role of SVs at the \u003cem\u003eTPCN1\u003c/em\u003e locus across ancestrally diverse populations and in other neurodegenerative diseases. \u003cem\u003eTMEM106B\u003c/em\u003e encodes a transmembrane glycoprotein that localizes to late lysosome and endosome [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. At the \u003cem\u003eTMEM106\u003c/em\u003e locus, we detected a 322-bp Alu deletion in exon 8 or 3\u0026rsquo; untranslated region, which has previously been reported as a likely causal variant and validated using long-read sequencing data [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. This deletion has been associated with not only with AD, but also with frontotemporal lobar dementia with TDP-43 inclusions (FTLD-TDP) [\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e], neurodegeneration [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], and several AD-related phenotypes, including tangles density, TDP-43, and cognitive resilience [\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFrom genome-wide analyses of common or low frequency SVs, we identified two significant deletions associated with AD among EUR. Both deletions are located in genes encoding CCDC proteins. Members of this family are characterized by an N-terminal potential microtubule binding domain, a central coiled-coiled and a C-terminal Hook-related domain. An 80-bp deletion on chromosome 11 is located in intron 7 of \u003cem\u003eCCDC88B\u003c/em\u003e and encompasses \u003cem\u003eMIR7155\u003c/em\u003e (chr11:64,341,849\u0026thinsp;\u0026minus;\u0026thinsp;64,341,904). \u003cem\u003eCCDC88B\u003c/em\u003e has been shown to act as a positive regulator of T-cell maturation and inflammatory function [\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. The low frequency deletion on chromosome 3 is located in intron 1 of \u003cem\u003eCCDC12\u003c/em\u003e and 2 Kb upstream of neurobeachin like 2 (\u003cem\u003eNBEAL2\u003c/em\u003e). The functions of \u003cem\u003eCCDC12\u003c/em\u003e remain unclear but it is predicted to be part of the spliceosomal complex. \u003cem\u003eNBEAL2\u003c/em\u003e is thought to play a role in megakaryocyte alpha-granule biogenesis. In public databases, these two deletions are annotated as indels with rsIDs rs1553653356 (chr3) and rs1591274862 (chr11), respectively. In the gnomAD database (v4.1.0) [\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e], rs1553653356 shows a low frequency (AF\u0026thinsp;=\u0026thinsp;0.002), consistent with our findings. In contrast, rs1591274862 (chr11) shows a notable discrepancy between the exome data (AF\u0026thinsp;=\u0026thinsp;0.2) and genome data (AF\u0026thinsp;=\u0026thinsp;0.002), although this variant failed quality control in both datasets. This discrepancy highlights the need for further investigation. Nonetheless, at the gene-level, a pQTL for \u003cem\u003eCCDC88B\u003c/em\u003e and an eQTL for \u003cem\u003eCCDC12\u003c/em\u003e have been previously associated with AD [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e], suggesting potential causal links between these genes and AD. Functional validation of the two deletions is warranted.\u003c/p\u003e \u003cp\u003eWe identified multiple genes associated with AD in gene-based analyses of rare SVs and SNVs/INDELs, particularly in LAT. The Latino population is genetically admixed, with varying proportions of European, African, and Amerindigenous genetic backgrounds, which adds genetic complexity [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. According to GrafPop, the Latin American 1 population primarily represents individuals with European and African ancestry components, whereas the Latin American 2 population mainly represents individuals with European and Amerindigenous components [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Notably, the Latin American populations exhibit unique LD patterns and haplotype structures derived from admixture [\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e], which may enhance the detection of rare variants. For \u003cem\u003ePSEN1\u003c/em\u003e, we identified four rare, coding SVs with evidence of association with AD. All four SVs were annotated as highly pathogenic by indirectly altering \u003cem\u003ePSEN1\u003c/em\u003e regulatory elements. Indeed, all four SVs are located in regions of neighboring genes, not within \u003cem\u003ePSEN1\u003c/em\u003e itself, highlighting the impact of SVs through long range regulatory mechanisms [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. For \u003cem\u003eABCA7\u003c/em\u003e, we identified a rare 605-bp deletion that partially overlaps intron 18 and exon 19, showing evidence of association with AD risk in AFR. This observation aligns with findings from a recent study [\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e], despite their use of a different SV caller and statistical model. This deletion is located approximately 3 kb downstream of a previously reported 44-bp deletion (rs142076058) associated with AD risk in African American individuals [\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. Additionally, it overlaps with a well-characterized SNP (rs115550680) previously associated with late-onset AD in African Americans populations [\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e]. In our dataset, all four individuals carrying the rare 605-bp deletion had AD and did not carry the previously reported deletion or SNPs associated with AD among African populations, whereas two common SNPs (rs3764650 and rs3752246) previously identified in European AD GWAS were observed [\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. Our findings provide new insights into the genetic architecture of the \u003cem\u003eABCA7\u003c/em\u003e locus in African ancestry populations. SPARC-related modular calcium-binding protein 1 (SMOC1\u003cem\u003e)\u003c/em\u003e has consistently been reported as a biomarker for early AD in proteomics studies [\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan additionalcitationids=\"CR47 CR48\" citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e], although the underlying genetic basis remains unclear. While we did not observe any rare SVs for \u003cem\u003eSMOC1\u003c/em\u003e, the observed aggregate of synonymous SNVs/INDELs suggestively associated with AD in this gene may partially explain the genetic contribution to the increased levels of SMOC1 in AD [\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eWe acknowledge several limitations in this study. Despite the relatively large sample of ascertained AD cases, statistical power remains limited, particularly for rare noncoding SV analyses within ancestry subgroups. Analyses using pooled populations did not yield additional associations for common SVs and may have introduced potential false positives in the gene-based analyses, possibly due to data structure complexities inherent to rare SVs and incomplete adjustment for population stratification using PCs. Another limitation is the lack of replication for our findings. While we identified two significant deletions associated with AD, their low frequency poses challenges for replication. Nonetheless, validation in larger and independent cohorts will be essential.\u003c/p\u003e"},{"header":"4. Conclusions","content":"\u003cp\u003eIn conclusion, we identified high-quality SVs in ADSP samples using a novel SV calling method. Our analysis revealed ancestry-specific SVs at known AD loci, as well as both common and rare SVs associated with AD. These findings provide new insights into genetic architecture of AD. Future studies are warranted to validate our results and investigate the functional impact of these SVs.\u003c/p\u003e"},{"header":"5. Methods","content":"\u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e5.1. Study samples\u003c/h2\u003e \u003cp\u003eThe Alzheimer\u0026rsquo;s Disease Sequencing Project (ADSP) was initiated in 2012 to elucidate the genetic architecture of AD, with major goals to identify genes and gene variants that confer risk for or protection against AD, to provide insight as to the biological impact of these genes and variants, and to identify potential therapeutic targets [\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e]. The WGS data release used in the present study (Release 3) includes data from 16,841 diverse individuals with and without AD from 24 cohorts. Raw data were obtained from the National Institute on Aging Genetics of Alzheimer's Disease Data Storage Site (NIAGADS). After removing duplicates (Section \u003cspan refid=\"Sec16\" class=\"InternalRef\"\u003e5.4\u003c/span\u003e), outliers subjects (Section \u003cspan refid=\"Sec17\" class=\"InternalRef\"\u003e5.5\u003c/span\u003e), and subjects with missing phenotypic information, 12,908 samples (6,604 controls, 6,304 cases) remained for analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e5.2. SV detection\u003c/h2\u003e \u003cp\u003eBioGraph version v6.0.4 was run per-sample using GRCh38 as the reference [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Variant Call Format (VCF) files were filtered to variant sites at least 50 bp long and with a PASS filter. Inversions were identified from all VCF entries with at least a 50 bp reference and 50 bp alternate allele reported where the sequence similarity of the reference and the reverse complement of the alternate allele was at least 80%. SVs were merged using bcftools v1.15 [\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e] and SVs with over 95% sequence and size similarity within 1000 bp were consolidated using Truvari collapse v3.1 with parameter `--keep maxqual` [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. When necessary, SVs were cross-referenced to intersecting tandem repeat regions from the adotto TR catalog [\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003e5.3. SV benchmarking using challenging, medically-relevant genes (CMRG)\u003c/h2\u003e \u003cp\u003eWe used WGS data from HG002, a sample with broad consent for open genomic data sharing through the Personal Genome Project [\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e]. SVs were called using BioGraph [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e], Manta v1.6.0 [\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e], Parliament2 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], and smoove 0.2.6 [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Truvari v3.5 [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e] was used to compare the resulting SV calls against the CMRG benchmark [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Default truvari parameters were used for BioGraph and Manta. The parameters `--dup-to-ins` and `--pctsim 0` were used for Parliament2 and smoove as neither tool produces sequence resolved calls.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e5.4. Quality control of SVs using replicates analysis\u003c/h2\u003e \u003cp\u003eTruvari v3.1 was run between 428 replicate pairs (ADSP participants with more than one sample sequenced). SV calls with over 70% sequence and size similarity between the replicates were classified as being consistent and the remainder classified as being inconsistent. Truvari annotations of PctSeqSimilarity and PctSizeSimilarity between consistent SV pairs were also analyzed to identify SVs that are the same across samples.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e5.5. Quality control of samples using One-Class Support Vector Machine (SVM)\u003c/h2\u003e \u003cp\u003ePassing SV counts by type were collected for each sample. Classification of the 2% of outlier samples by counts was performed using scikit-learn v1.1.3 and their OneClass SVM with hyper-parameters kernel=\u0026rsquo;poly\u0026rsquo; and nu\u0026thinsp;=\u0026thinsp;0.02. Intersection of samples with the study which provided them showed a concentration of outlier samples from 3 of the 24 studies comprising the ADSP study sample. All samples included in these three studies (N\u0026thinsp;=\u0026thinsp;421) were dropped.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e5.6. Intersection with known SVs and annotation to genes\u003c/h2\u003e \u003cp\u003eSVAfotate version 0.0.1 was used to intersect the discovered SVs with known SVs [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. This program comprises an annotated file containing boundaries of SVs from 1000G [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], CCDG [\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e], and gnomAD [\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e]. TopMed SVs freeze 1.1 [\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e] were also collected and consolidated into the annotated file. SVs were annotated to genes with VEP using VEP-ensembl version ​​107.0 [\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e]. Next, we examined whether any of the discovered SVs mapped to genes reported by the ADSP Gene Verification Committee [\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e5.7. Global ancestry inference\u003c/h2\u003e \u003cp\u003eGlobal ancestry inference of the study samples was performed using GrafPop [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e, \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e], a distance-based method that uses a reference composed of nearly 100,000 fingerprint SNPs extracted from dbGaP [\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e]. Grafpop estimates ancestry by calculating genetic distances between each individual and the reference populations, and subjects are clustered using genetic distances based on their genetic similarity. This tool considers that individuals\u0026rsquo; genomes are admixed from three ancestries: European (E), African (F), and Asian (A) and estimates ancestral proportions P\u003csub\u003ee\u003c/sub\u003e, P\u003csub\u003ef\u003c/sub\u003e, and P\u003csub\u003ea\u003c/sub\u003e based on genetic distances score using barycentric coordinates. In GrafPop, the cutoff thresholds were empirically defined to facilitate the grouping of dbGaP subjects. Due to the incompatibility of GrafPop with SV data, we used ADSP WGS data on single nucleotide variants (SNV) to perform global ancestry inference.\u003c/p\u003e \u003cp\u003eUsing the cutoff standard established by GrafPop, ADSP participants with WGS data were clustered into nine groups defined by study-reported populations within dbGAP. These groups include European, African, East Asian, African American, Latin American 1, Asian-Pacific Islander, South Asian, Latin American 2, and Other, based on their ancestral proportions and genetic distance [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. We grouped African and African American populations as African ancestry group (AFR) and Latin American 1 and Latin American 2 populations as Latin ancestry group (LAT), and European population as European ancestry group (EUR) [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Participants identified as East Asian, Asian-Pacific Islander, South Asian, and other populations were grouped as others and were excluded from subgroup association analyses due to limited sample size.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e5.8. Principal Component Analysis\u003c/h2\u003e \u003cp\u003ePrincipal component analysis (PCA) was performed using PC-AiR [\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e] in the GENetic EStimation and Inference in Structured samples (GENESIS) package [\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e]. We calculated PCs for all individuals (N\u0026thinsp;=\u0026thinsp;12,908) in the study sample using high quality deletions, insertions, and inversions with minor allele frequency (MAF) greater than 1%, and with Hardy Weinberg Equilibrium (HWE) P-value greater than the Bonferroni-corrected threshold based on the total number of SVs (P\u0026thinsp;=\u0026thinsp;5.2E-07). SVs in linkage disequilibrium (LD) were excluded using a r\u003csup\u003e2\u003c/sup\u003e threshold greater than 0.1. For comparison with SNVs, we performed PCA on the same sample using biallelic SNPs with MAF\u0026thinsp;\u0026gt;\u0026thinsp;1%, HWE P-value\u0026thinsp;\u0026gt;\u0026thinsp;1E-06, and call rate\u0026thinsp;\u0026gt;\u0026thinsp;95%.\u003c/p\u003e \u003cp\u003eFor ancestry-specific SV association analyses, we calculated PCs using SVs with MAF\u0026thinsp;\u0026gt;\u0026thinsp;1% and HWE P-values exceeding the Bonferroni-corrected threshold based on the number of SVs with MAF\u0026thinsp;\u0026gt;\u0026thinsp;1% in each ancestry group (AFR: P\u0026thinsp;=\u0026thinsp;4.7E-07; EUR: P\u0026thinsp;=\u0026thinsp;6.5E-07; LAT: P\u0026thinsp;=\u0026thinsp;5.4E-07). For ancestry-specific SNVs association analyses, we focused on previously reported AD GWAS SNPs that were in LD with our SV calls. For these analyses, PCs were calculated using WGS data filtered for MAF\u0026thinsp;\u0026gt;\u0026thinsp;1%, HWE P\u0026thinsp;\u0026gt;\u0026thinsp;1E-06, and call rate\u0026thinsp;\u0026gt;\u0026thinsp;95% within each ancestry group.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003e5.9. SVs tagging known AD GWAS SNPs\u003c/h2\u003e \u003cp\u003eTo investigate the role of SVs in known AD GWAS loci, we examined 147 SNPs tagging AD loci identified in previous GWAS [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e, \u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e]. These variants were extracted from our WGS data, and we performed pairwise LD analysis between the AD-associated SNPs and our SVs within each ancestry group (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA). Among common or low-frequency SVs with MAF\u0026thinsp;\u0026gt;\u0026thinsp;0.5% and in HWE (see above), we specifically focused on SVs that were in LD (r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.4) with at least one of the AD-associated SNPs. LD calculation was carried out using PLINK v1.9 with parameters --ld-window-r2 0.4 and --r2. Haplotype estimation was performed using PLINK v1.9 based on pairwise LD patterns.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003e5.10. Association analyses\u003c/h2\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003e5.10.1. Models and Covariates\u003c/h2\u003e \u003cp\u003eWithin each ancestry group inferred based on genetic similarity, association analyses were conducted using a mixed effects logistic regression model. Detailed models and software for common and rare SV analyses are provided in the corresponding sections below.\u003c/p\u003e \u003cp\u003eCovariates included sex, SV-derived PC 1\u0026ndash;5 of each ancestry, relatedness via a genetic relatedness matrix (GRM), and technical covariates including sequencing center and whether the sample preparation was PCR-free.\u003c/p\u003e \u003cp\u003eFor the analysis evaluating SVs in LD with SNPs tagging AD loci, we applied a conditional model that further included the corresponding SNP dosage. We also performed association analyses of the identified SNPs in LD with SVs, replacing SV PCs with SNP PCs, and including the corresponding SV dosage as a covariate in the conditional model.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section3\"\u003e \u003ch2\u003e5.10.2. Single variant analysis of common or low-frequency SVs\u003c/h2\u003e \u003cp\u003eAssociation analyses of common and low frequency SVs (MAF\u0026thinsp;\u0026gt;\u0026thinsp;0.5% and passing the HWE criterion) were conducted using a mixed-effects logistic regression model implemented in the GENESIS R-package [\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e], with covariates and models as described above. All association analyses were performed within each ancestry group. For genome-wide association analysis, a meta-analysis was additionally performed across all ancestry subgroups using METAL software, implementing Stouffer method to weight results by sample size.\u003c/p\u003e \u003cp\u003eTo identify AD-associated SVs, we considered several P-value thresholds: For evaluating SVs in LD with known AD loci (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA), a suggestive significance threshold (P\u0026thinsp;\u0026lt;\u0026thinsp;0.1) was used. For genome-wide analysis of all common or low-frequency SVs (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB), two significance thresholds were considered: a Bonferroni-corrected threshold (AFR: P\u0026thinsp;\u0026lt;\u0026thinsp;1.7E-06; EUR: P\u0026thinsp;\u0026lt;\u0026thinsp;3.4E-06; LAT: P\u0026thinsp;\u0026lt;\u0026thinsp;1.6E-06) based on the total number of SVs analyzed and the conventional genome-wide significance threshold (P\u0026thinsp;\u0026lt;\u0026thinsp;5E-08).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003e5.10.3. Gene-based analysis of aggregates of rare SVs\u003c/h2\u003e \u003cp\u003eGene-based association analyses of aggregated rare SVs with AD were conducted as the primary analyses. We included deletions, insertion, and inversion with MAF\u0026thinsp;\u0026lt;\u0026thinsp;1% in each ancestry group and estimated their pathogenicity using PhenoSV [\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e]. PhenoSV is a machine learning based method that predicts the functional consequences of both coding and non-coding SVs that may directly or indirectly influence genes. SVs were classified as coding if they overlapped at least 1bp with exons of protein-coding gene based on GENCODE v40 annotations [\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e], considering only high-confidence representative transcript, otherwise they were classified as noncoding. Independently, SVs were evaluated for their potential to affect genes directly or indirectly. Non-coding SVs were tested for their indirect effect on genes within 1Mb upstream and downstream, as defined by default. Gene-level pathogenicity scores ranged from 0 to 1 and were used to classify SVs into pathogenic (\u0026ge;\u0026thinsp;0.5) and benign (\u0026lt;\u0026thinsp;0.5) groups [\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e]. Only rare pathogenic SVs were included in the analyses, which were performed separately for non-coding and coding variants (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC, \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD).\u003c/p\u003e \u003cp\u003eIn secondary analyses, we integrated SV data with SNVs/INDELs derived from ADSP 17K WGS data [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e] to increase statistical power. The WGS data had been previously processed and quality-controlled according to the Genome Center for Alzheimer\u0026rsquo;s Disease (GCAD)/ADSP QC pipeline [\u003cspan citationid=\"CR68\" class=\"CitationRef\"\u003e68\u003c/span\u003e]. Using the WGS data annotated with FAVOR, we classified SNVs/INDELs as coding or non-coding based on the STAAR pipeline [\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e]. Variants with MAF\u0026thinsp;\u0026lt;\u0026thinsp;1% within each ancestry group were included in the analysis. The coding SNVs/INDELs were categorized into five functional groups: putative loss of function (pLof), missense, disruptive missense, pLof\u0026thinsp;+\u0026thinsp;disruptive missense, or synonymous. The non-coding SNVs/INDELs were grouped into eight categories: promoter or enhancer overlaid with cap analysis of gene expression (CAGE) or DNase I hypersensitive site (DHS) sites, untranslated region (UTR), upstream, downstream, and noncoding RNA genes. Gene-based analyses were then performed within each category of coding and noncoding variants, combining rare pathogenic SVs of the corresponding type (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eE, \u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eF).\u003c/p\u003e \u003cp\u003eAll gene-based association analyses were performed using the variant-set mixed model association test (SMMAT), implemented in the GMMAT R package [\u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e70\u003c/span\u003e]. We used a hybrid test (SMMAT-E), which combines burden and SKAT tests and has been shown to offer greater statistical power than either test alone. MAF was used as a weight by default. Genes with a cumulative minor allele count (cMAC)\u0026thinsp;\u0026ge;\u0026thinsp;10 were included in the analysis. All association analyses were stratified by ancestry group, and meta-analyses combining results across ancestry subgroups were conducted using the metap R package [\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e] with the Stouffer method to account for sample size differences.\u003c/p\u003e \u003cp\u003eWe applied two significance thresholds for gene-based tests: Bonferroni-corrected threshold (P\u0026thinsp;\u0026lt;\u0026thinsp;1E-07) to account for approximately 20,000 genes tested, and a suggestive threshold (P\u0026thinsp;\u0026lt;\u0026thinsp;1E-05). For genes with significance in the secondary analyses, we performed conditional analyses on the aggregates of SNVs/INDELs within each gene to assess whether the association signal was driven by SVs.\u003c/p\u003e \u003cp\u003eFinally, we conducted candidate gene association analyses focusing on 15 AD genes previously reported to harbor rare variant associations [\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e, \u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e73\u003c/span\u003e], evaluating them in both primary and secondary gene-based analyses. We computed the false discovery rate (FDR) Q value within each analysis group using the Benjamini-Hochberg procedure to assess statistical significance. For genes showing evidence of association with AD (FDR Q\u0026thinsp;\u0026lt;\u0026thinsp;0.05) in the secondary analyses, we conducted additional conditional analyses on the aggregated SNVs/INDELs to evaluate the contribution of SVs to the observed signal.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eS.L., A.C.E., M.F., and F.J.S. contributed to the conception and design of the work, data interpretation, and manuscript drafting. A.C.E. developed new software used in the study. S.L., R.X., and A.C.E. performed the data analysis. G.M.P., S.H.C., and A.L.D. contributed to data acquisition. S.L. and A.C.E. contributed equally to this work. M.F. and F.J.S. contributed equally to this work. All authors read, reviewed, and approved the final manuscript.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eADSP whole genome sequencing data (NG00067) are available through the National Institute on Aging Genetics of Alzheimer\u0026rsquo;s Disease Data Storage Site (NIAGADS) (https://www.niagads.org).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe thank the participants and their families for making this research possible. Data used in this study were generated via the Alzheimer Disease Sequencing Project (ADSP). Full ADSP acknowledgements can be found here: https://adsp.niagads.org/acknowledgment/ Data used in this study were obtained from the Alzheimer\u0026apos;s Disease Neuroimaging Initiative (ADNI) via the ADSP. As such, the investigators within the ADNI contributed to the design and implementation of ADNI and/or provided data but did not participate in the analysis or writing of this report. A complete listing of ADNI investigators can be found at:\u003cbr\u003e \u003cstrong\u003ehttp://adni.loni.usc.edu/wp-content/uploads/how_to_apply/ADNI_Acknowledgement_List.pdf\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSources of Funding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was primarily supported by grants U01AG058589, U01AG052409, and U01AG070112 from the National Institute on Aging. Additional support was provided by grant U01AG068221. Data for this study were prepared, archived, and distributed by the National Institute on Aging Alzheimer\u0026rsquo;s Disease Data Storage Site (NIAGADS) at the University of Pennsylvania (U24-AG041689), funded by the National Institute on Aging. A complete description of the funding support for the ADSP is provided at https://adsp.niagads.org/acknowledgment/.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent Statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study constitutes secondary research, utilizing de-identified data obtained from primary data repositories. In accordance with NIH policy, this research does not qualify as human subject research, and therefore, obtaining consent from individual participants is not required. All contributing studies included in this work received ethical oversight from their respective institutions.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eRajan KB, Weuve J, Barnes LL, McAninch EA, Wilson RS, Evans DA. Population estimate of people with clinical Alzheimer's disease and mild cognitive impairment in the United States (2020\u0026ndash;2060). Alzheimers Dement. 2021;17:1966\u0026ndash;75.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003e2025 Alzheimer\u0026rsquo;s disease facts and figures. \u003cem\u003eAlzheimer's \u0026amp; Dementia\u003c/em\u003e 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGatz M, Reynolds CA, Fratiglioni L, Johansson B, Mortimer JA, Berg S, Fiske A, Pedersen NL. Role of genes and environments for explaining Alzheimer disease. Arch Gen Psychiatry. 2006;63:168\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSims R, Hill M, Williams J. The multiplex model of the genetics of Alzheimer's disease. Nat Neurosci. 2020;23:311\u0026ndash;22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTanzi RE. The genetics of Alzheimer disease. Cold Spring Harb Perspect Med 2012, 2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBellenguez C, Kucukali F, Jansen IE, Kleineidam L, Moreno-Grau S, Amin N, Naj AC, Campos-Martin R, Grenier-Boley B, Andrade V, et al. New insights into the genetic etiology of Alzheimer's disease and related dementias. Nat Genet. 2022;54:412\u0026ndash;36.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWightman DP, Jansen IE, Savage JE, Shadrin AA, Bahrami S, Holland D, Rongve A, Borte S, Winsvold BS, Drange OK, et al. A genome-wide association study with 1,126,563 individuals identifies new risk loci for Alzheimer's disease. Nat Genet. 2021;53:1276\u0026ndash;82.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndrews SJ, Renton AE, Fulton-Howard B, Podlesny-Drabiniok A, Marcora E, Goate AM. The complex genetic architecture of Alzheimer's disease: novel insights and future directions. EBioMedicine. 2023;90:104511.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSudmant PH, Rausch T, Gardner EJ, Handsaker RE, Abyzov A, Huddleston J, Zhang Y, Ye K, Jun G, Fritz MH, et al. An integrated map of structural variation in 2,504 human genomes. Nature. 2015;526:75\u0026ndash;81.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMahmoud M, Gobet N, Cruz-Davalos DI, Mounier N, Dessimoz C, Sedlazeck FJ. Structural variant calling: the long and the short of it. Genome Biol. 2019;20:246.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eScott AJ, Chiang C, Hall IM. Structural variants are a major source of gene expression differences in humans and often affect multiple nearby genes. Genome Res. 2021;31:2249\u0026ndash;57.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChiang C, Scott AJ, Davis JR, Tsang EK, Li X, Kim Y, Hadzic T, Damani FN, Ganel L, Consortium GT, et al. The impact of structural variation on human gene expression. Nat Genet. 2017;49:692\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAudano PA, Sulovari A, Graves-Lindsay TA, Cantsilieris S, Sorensen M, Welch AE, Dougherty ML, Nelson BJ, Shah A, Dutcher SK, et al. Characterizing the Major Structural Variant Alleles of the Human Genome. Cell. 2019;176:663\u0026ndash;e675619.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSedlazeck FJ, Rescheneder P, Smolka M, Fang H, Nattestad M, von Haeseler A, Schatz MC. Accurate detection of complex structural variations using single-molecule sequencing. Nat Methods. 2018;15:461\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Coster W, Weissensteiner MH, Sedlazeck FJ. Towards population-scale long-read sequencing. Nat Rev Genet. 2021;22:572\u0026ndash;87.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEnglish AC, McCarthy N, Flickenger R, Maheshwari S, Meed L, Mangubat A, Shekar SN. Leveraging a WGS compression and indexing format with dynamic graph references to call structural variants. \u003cem\u003ebioRxiv\u003c/em\u003e 2020:2020.2004.2024.060202..\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen X, Schulz-Trieglaff O, Shaw R, Barnes B, Schlesinger F, Kallberg M, Cox AJ, Kruglyak S, Saunders CT. Manta: rapid detection of structural variants and indels for germline and cancer sequencing applications. Bioinformatics. 2016;32:1220\u0026ndash;2.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZarate S, Carroll A, Mahmoud M, Krasheninina O, Jun G, Salerno WJ, Schatz MC, Boerwinkle E, Gibbs RA, Sedlazeck FJ. Parliament2: Accurate structural variant calling at scale. Gigascience 2020, 9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLayer RM, Chiang C, Quinlan AR, Hall IM. LUMPY: a probabilistic framework for structural variant discovery. Genome Biol. 2014;15:R84.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWagner J, Olson ND, Harris L, McDaniel J, Cheng H, Fungtammasan A, Hwang YC, Gupta R, Wenger AM, Rowell WJ, et al. Curated variation benchmarks for challenging medically relevant autosomal genes. Nat Biotechnol. 2022;40:672\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEnglish AC, Menon VK, Gibbs RA, Metcalf GA, Sedlazeck FJ. Truvari: refined structural variant comparison preserves allelic diversity. Genome Biol. 2022;23:271.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNicholas TJ, Cormier MJ, Quinlan AR. Annotation of structural variants with reported allele frequencies and related metrics from multiple datasets using SVAFotate. BMC Bioinformatics. 2022;23:490.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKunkle BW, Schmidt M, Klein HU, Naj AC, Hamilton-Nelson KL, Larson EB, Evans DA, De Jager PL, Crane PK, Buxbaum JD, et al. Novel Alzheimer Disease Risk Loci and Pathways in African American Individuals Using the African Genome Resources Panel: A Meta-analysis. JAMA Neurol. 2021;78:102\u0026ndash;13.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee WP, Choi SH, Shea MG, Cheng PL, Dombroski BA, Pitsillides AN, Heard-Costa NL, Wang H, Bulekova K, Kuzma AB, et al. Association of common and rare variants with Alzheimer's disease in more than 13,000 diverse individuals with whole-genome sequencing from the Alzheimer's Disease Sequencing Project. Alzheimers Dement. 2024;20:8470\u0026ndash;83.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSmolka M, Paulin LF, Grochowski CM, Horner DW, Mahmoud M, Behera S, Kalef-Ezra E, Gandhi M, Hong K, Pehlivan D, et al. Detection of mosaic and population-level structural variants with Sniffles2. Nat Biotechnol. 2024;42:1571\u0026ndash;80.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSanders AD, Hills M, Porubsky D, Guryev V, Falconer E, Lansdorp PM. Characterizing polymorphic inversions in human genomes by single-cell sequencing. Genome Res. 2016;26:1575\u0026ndash;87.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaivola K, Chia R, Ding J, Rasheed M, Fujita M, Menon V, Walton RL, Collins RL, Billingsley K, Brand H, et al. Genome-wide structural variant analysis identifies risk loci for non-Alzheimer's dementias. Cell Genom. 2023;3:100316.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMallmann RT, Klugbauer N. Genetic Inactivation of Two-Pore Channel 1 Impairs Spatial Learning and Memory. Behav Genet. 2020;50:401\u0026ndash;10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJadiya P, Cohen HM, Kolmetzky DW, Kadam AA, Tomar D, Elrod JW. Neuronal loss of NCLX-dependent mitochondrial calcium efflux mediates age-associated cognitive decline. iScience. 2023;26:106296.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStagi M, Klein ZA, Gould TJ, Bewersdorf J, Strittmatter SM. Lysosome size, motility and stress response regulated by fronto-temporal dementia modifier TMEM106B. Mol Cell Neurosci. 2014;61:226\u0026ndash;40.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChemparathy A, Le Guen Y, Zeng Y, Gorzynski J, Jensen TD, Yang C, Kasireddy N, Talozzi L, Belloy M, Stewart I, et al. A 3'UTR Insertion Is a Candidate Causal Variant at the TMEM106B Locus Associated With Increased Risk for FTLD-TDP. Neurol Genet. 2024;10:e200124.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSalazar A, Tesi N, Knoop L, Pijnenburg Y, van der Lee S, Wijesekera S, Krizova J, Hiltunen M, Damme M, Petrucelli L et al. An AluYb8 retrotransposon characterises a risk haplotype of TMEM106B associated in neurodegeneration. \u003cem\u003emedRxiv\u003c/em\u003e 2023:2023.2007.2016.23292721..\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVialle RA, de Paiva Lopes K, Li Y, Ng B, Schneider JA, Buchman AS, Wang Y, Farfel JM, Barnes LL, Wingo AP, et al. Structural variants linked to Alzheimer's disease and other common age-related clinical and neuropathologic traits. Genome Med. 2025;17:20.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKennedy JM, Fodil N, Torre S, Bongfen SE, Olivier JF, Leung V, Langlais D, Meunier C, Berghout J, Langat P, et al. CCDC88B is a novel regulator of maturation and effector functions of T cells during pathological inflammation. J Exp Med. 2014;211:2519\u0026ndash;35.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S, Francioli LC, Goodrich JK, Collins RL, Kanai M, Wang Q, Alfoldi J, Watts NA, Vittal C, Gauthier LD, et al. A genomic mutational constraint map using variation in 76,156 human genomes. Nature. 2024;625:92\u0026ndash;100.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAli M, Timsina J, Western D, Liu M, Beric A, Budde J, Do A, Heo G, Wang L, Gentsch J, et al. Multi-cohort cerebrospinal fluid proteomics identifies robust molecular signatures across the Alzheimer disease continuum. Neuron. 2025;113:1363\u0026ndash;e13791369.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMathys H, Davila-Velderrain J, Peng Z, Gao F, Mohammadi S, Young JZ, Menon M, He L, Abdurrob F, Jiang X, et al. Single-cell transcriptomic analysis of Alzheimer's disease. Nature. 2019;570:332\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMao X, Bigham AW, Mei R, Gutierrez G, Weiss KM, Brutsaert TD, Leon-Velarde F, Moore LG, Vargas E, McKeigue PM, et al. A genomewide admixture mapping panel for Hispanic/Latino populations. Am J Hum Genet. 2007;80:1171\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJin Y, Schaffer AA, Feolo M, Holmes JB, Kattman BL. GRAF-pop: A Fast Distance-Based Method To Infer Subject Ancestry from Multiple Genotype Datasets Without Principal Components Analysis. \u003cem\u003eG3 (Bethesda)\u003c/em\u003e 2019, 9:2447\u0026ndash;2461.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eda Cruz PRS, Ananina G, Secolin R, Gil-da-Silva-Lopes VL, Lima CSP, de Franca PHC, Donatti A, Lourenco GJ, de Araujo TK, Simioni M et al. Demographic history differences between Hispanics and Brazilians imprint haplotype features. \u003cem\u003eG3 (Bethesda)\u003c/em\u003e 2022, 12.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoyling A, Perez-Siles G, Kennerson ML. Structural Variation at a Disease Mutation Hotspot: Strategies to Investigate Gene Regulation and the 3D Genome. Front Genet. 2022;13:842860.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang H, Dombroski BA, Cheng PL, Tucci A, Si YQ, Farrell JJ, Tzeng JY, Leung YY, Malamon JS et al. Alzheimer's Disease Sequencing P, : Structural variation detection and association analysis of whole-genome-sequence data from 16,543 Alzheimer's disease sequencing project subjects. \u003cem\u003eAlzheimers Dement\u003c/em\u003e 2025, 21:e70277.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCukier HN, Kunkle BW, Vardarajan BN, Rolati S, Hamilton-Nelson KL, Kohli MA, Whitehead PL, Dombroski BA, Van Booven D, Lang R, et al. ABCA7 frameshift deletion associated with Alzheimer disease in African Americans. Neurol Genet. 2016;2:e79.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReitz C, Jun G, Naj A, Rajbhandary R, Vardarajan BN, Wang LS, Valladares O, Lin CF, Larson EB, Graff-Radford NR, et al. Variants in the ATP-binding cassette transporter (ABCA7), apolipoprotein E ϵ4,and the risk of late-onset Alzheimer disease in African Americans. JAMA. 2013;309:1483\u0026ndash;92.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDib S, Pahnke J, Gosselet F. Role of ABCA7 in Human Health and in Alzheimer's Disease. Int J Mol Sci 2021, 22.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSung YJ, Yang C, Norton J, Johnson M, Fagan A, Bateman RJ, Perrin RJ, Morris JC, Farlow MR, Chhatwal JP, et al. Proteomics of brain, CSF, and plasma identifies molecular signatures for distinguishing sporadic and genetic Alzheimer's disease. Sci Transl Med. 2023;15:eabq5923.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuo Y, Chen SD, You J, Huang SY, Chen YL, Zhang Y, Wang LB, He XY, Deng YT, Zhang YR, et al. Multiplex cerebrospinal fluid proteomics identifies biomarkers for diagnosis and prediction of Alzheimer's disease. Nat Hum Behav. 2024;8:2047\u0026ndash;66.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang H, Dey KK, Chen PC, Li Y, Niu M, Cho JH, Wang X, Bai B, Jiao Y, Chepyala SR, et al. Integrated analysis of ultra-deep proteomes in cortex, cerebrospinal fluid and serum reveals a mitochondrial signature in Alzheimer's disease. Mol Neurodegener. 2020;15:43.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWatson CM, Dammer EB, Ping L, Duong DM, Modeste E, Carter EK, Johnson ECB, Levey AI, Lah JJ, Roberts BR, Seyfried NT. Quantitative Mass Spectrometry Analysis of Cerebrospinal Fluid Protein Biomarkers in Alzheimer's Disease. Sci Data. 2023;10:261.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOelschlaeger P. Molecular Mechanisms and the Significance of Synonymous Mutations. Biomolecules 2024, 14.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBeecham GW, Bis JC, Martin ER, Choi SH, DeStefano AL, van Duijn CM, Fornage M, Gabriel SB, Koboldt DC, Larson DE, et al. The Alzheimer's Disease Sequencing Project: Study design and sample selection. Neurol Genet. 2017;3:e194.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDanecek P, Bonfield JK, Liddle J, Marshall J, Ohan V, Pollard MO, Whitwham A, Keane T, McCarthy SA, Davies RM, Li H. Twelve years of SAMtools and BCFtools. \u003cem\u003eGigascience\u003c/em\u003e 2021, 10.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEnglish AC, Dolzhenko E, Ziaei Jam H, McKenzie SK, Olson ND, De Coster W, Park J, Gu B, Wagner J, Eberle MA, et al. Analysis and benchmarking of small and large genomic variants across tandem repeats. Nat Biotechnol. 2025;43:431\u0026ndash;42.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBall MP, Thakuria JV, Zaranek AW, Clegg T, Rosenbaum AM, Wu X, Angrist M, Bhak J, Bobe J, Callow MJ, et al. A public resource facilitating clinical use of genomes. Proc Natl Acad Sci U S A. 2012;109:11920\u0026ndash;7.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAbel HJ, Larson DE, Regier AA, Chiang C, Das I, Kanchi KL, Layer RM, Neale BM, Salerno WJ, Reeves C, et al. Mapping and characterization of structural variation in 17,795 human genomes. Nature. 2020;583:83\u0026ndash;9.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollins RL, Brand H, Karczewski KJ, Zhao X, Alfoldi J, Francioli LC, Khera AV, Lowther C, Gauthier LD, Wang H, et al. A structural variation reference for medical and population genetics. Nature. 2020;581:444\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJun G, English AC, Metcalf GA, Yang J, Chaisson MJ, Pankratz N, Menon VK, Salerno WJ, Krasheninina O, Smith AV et al. Structural variation across 138,134 samples in the TOPMed consortium. \u003cem\u003ebioRxiv\u003c/em\u003e 2023.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcLaren W, Gil L, Hunt SE, Riat HS, Ritchie GR, Thormann A, Flicek P, Cunningham F. The Ensembl Variant Effect Predictor. Genome Biol. 2016;17:122.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eList of AD Loci and Genes with Genetic Evidence Compiled. by ADSP Gene Verification Committee [\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://adsp.niagads.org/gvc-top-hits-list/]\u003c/span\u003e\u003cspan address=\"https://adsp.niagads.org/gvc-top-hits-list/]\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJin Y, Schaffer AA, Sherry ST, Feolo M. Quickly identifying identical and closely related subjects in large databases using genotype data. PLoS ONE. 2017;12:e0179106.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTryka KA, Hao L, Sturcke A, Jin Y, Wang ZY, Ziyabari L, Lee M, Popova N, Sharopova N, Kimura M, Feolo M. NCBI's Database of Genotypes and Phenotypes: dbGaP. Nucleic Acids Res. 2014;42:D975\u0026ndash;979.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eConomos MP, Miller MB, Thornton TA. Robust inference of population structure for ancestry prediction and correction of stratification in the presence of relatedness. Genet Epidemiol. 2015;39:276\u0026ndash;93.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGogarten SM, Sofer T, Chen H, Yu C, Brody JA, Thornton TA, Rice KM, Conomos MP. Genetic association testing using the GENESIS R/Bioconductor package. Bioinformatics. 2019;35:5346\u0026ndash;8.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSherva R, Zhang R, Sahelijo N, Jun G, Anglin T, Chanfreau C, Cho K, Fonda JR, Gaziano JM, Harrington KM, et al. African ancestry GWAS of dementia in a large military cohort identifies significant risk loci. Mol Psychiatry. 2023;28:1293\u0026ndash;302.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKunkle BW, Grenier-Boley B, Sims R, Bis JC, Damotte V, Naj AC, Boland A, Vronskaya M, van der Lee SJ, Amlie-Wolf A, et al. Genetic meta-analysis of diagnosed Alzheimer's disease identifies new risk loci and implicates Abeta, tau, immunity and lipid processing. Nat Genet. 2019;51:414\u0026ndash;30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu Z, Li Q, Marchionni L, Wang K. PhenoSV: interpretable phenotype-aware model for the prioritization of genes affected by structural variants. Nat Commun. 2023;14:7805.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFrankish A, Diekhans M, Jungreis I, Lagarde J, Loveland JE, Mudge JM, Sisu C, Wright JC, Armstrong J, Barnes I, et al. Gencode 2021. Nucleic Acids Res. 2021;49:D916\u0026ndash;23.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaj AC, Lin H, Vardarajan BN, White S, Lancour D, Ma Y, Schmidt M, Sun F, Butkiewicz M, Bush WS, et al. Quality control and integration of genotypes from two calling pipelines for whole genome sequence data in the Alzheimer's disease sequencing project. Genomics. 2019;111:808\u0026ndash;18.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSTAARpipeline. an all-in-one rare-variant tool for biobank-scale whole-genome sequencing data. Nat Methods. 2022;19:1532\u0026ndash;3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen H, Huffman JE, Brody JA, Wang C, Lee S, Li Z, Gogarten SM, Sofer T, Bielak LF, Bis JC, et al. Efficient Variant Set Mixed Model Association Tests for Continuous and Binary Traits in Large-Scale Whole-Genome Sequencing Studies. Am J Hum Genet. 2019;104:260\u0026ndash;74.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDewey M. metap: Meta-Analysis of Significance Values. 2025.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang D, Scalici A, Wang Y, Lin H, Pitsillides A, Heard-Costa N, Cruchaga C, Ziegemeier E, Bis JC, Fornage M, et al. Frequency of variants in Mendelian Alzheimer's disease genes within the Alzheimer's Disease Sequencing Project. J Alzheimers Dis. 2025;104:841\u0026ndash;51.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDe Deyn L, Sleegers K. The impact of rare genetic variants on Alzheimer disease. Nat Rev Neurol. 2025;21:127\u0026ndash;39.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"genome-biology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gbio","sideBox":"Learn more about [Genome Biology](https://genomebiology.biomedcentral.com/)","snPcode":"13059","submissionUrl":"https://submission.springernature.com/new-submission/13059/3","title":"Genome Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Structural variation, Alzheimer’s Disease, genetics, association study, whole genome sequence","lastPublishedDoi":"10.21203/rs.3.rs-8562759/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8562759/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eIntroduction\u003c/strong\u003e: Structural variants (SV), genomic alterations spanning more than 50 base pairs, can significantly impact gene expression and protein function. However, their contribution to Alzheimer’s Disease (AD) remains poorly understood. Leveraging a novel SV calling pipeline, we identified SVs with high accuracy in a diverse sample of the Alzheimer's Disease Sequencing Project (ADSP) and investigated the role of SVs in AD.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e: We analyzed SVs in 16,841 individuals from ADSP whole genome sequencing data using BioGraph, a semi-assembly-based method that employs graph-based representation for accurate SV detection. We identified 456,644 high-quality SVs, 65% of which were novel. Of these, 272,728 SVs directly impact genes, including 86 AD-related genes. Association analyses were performed within three ancestry groups, including 3,371 African (AFR), 6,327 European (EUR), and 2,126 Latin (LAT). Multiple deletions and insertions were observed in moderate to high linkage disequilibrium with known AD loci, including \u003cem\u003eTPCN1\u003c/em\u003e and \u003cem\u003eTMEM106B\u003c/em\u003e. In EUR, genome-wide association analysis identified two significant low-frequency deletions associated with AD, located in introns of \u003cem\u003eCCDC12\u003c/em\u003e and \u003cem\u003eCCDC88B\u003c/em\u003e, both encoding coiled-coil domain-containing proteins. Gene-based analyses further identified rare pathogenic SVs in several known AD genes, including \u003cem\u003ePSEN1\u003c/em\u003e in LAT and \u003cem\u003eABCA7 \u003c/em\u003ein AFR.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e: Using a novel graph-based SV calling pipeline, we identified high-quality SVs across a large and ancestrally diverse cohort. Our analyses revealed both common and rare SVs associated with AD. These findings provide valuable insights into the genetic architecture of AD, emphasizing the value of including diverse populations in AD genomic studies.\u003c/p\u003e","manuscriptTitle":"The Impact of Structural Variation on Alzheimer’s Disease in the Alzheimer’s Disease Sequencing Project","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-01-13 06:59:11","doi":"10.21203/rs.3.rs-8562759/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2026-02-05T16:22:16+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-05T00:47:46+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2026-02-03T02:45:25+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"271441013482318441211739063370349189778","date":"2026-01-19T15:41:02+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"4916117584451475878927536216457547877","date":"2026-01-17T15:12:56+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-01-16T20:42:45+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-01-16T09:26:42+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-01-12T15:45:11+00:00","index":"","fulltext":""},{"type":"submitted","content":"Genome Biology","date":"2026-01-09T15:59:26+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"genome-biology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gbio","sideBox":"Learn more about [Genome Biology](https://genomebiology.biomedcentral.com/)","snPcode":"13059","submissionUrl":"https://submission.springernature.com/new-submission/13059/3","title":"Genome Biology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"fe71c2ce-eb0f-4470-b41a-b784fe06b12e","owner":[],"postedDate":"January 13th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-16T14:53:08+00:00","versionOfRecord":[],"versionCreatedAt":"2026-01-13 06:59:11","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8562759","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8562759","identity":"rs-8562759","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00