{"paper_id":"dd13abc4-066b-49c3-a941-5fa23ed0cb58","body_text":"Auwerx et al., 2023 1 \nCopy-number variants as modulators of  1 \ncommon disease susceptibility 2 \n 3 \nChiara Auwerx1,2,3,4,*, Maarja Jõeloo5,6, Marie C. Sadler2,3,4, Nicolò Tesio1, Sven Ojavee2,3, 4 \nCharlie J. Clark1, Reedik Mägi6, Estonian Biobank Research Team6,§, Alexandre 5 \nReymond1,#,* & Zoltán Kutalik2,3,4,#,* 6 \n 7 \n 8 \n1 Center for Integrative Genomics, University of Lausanne, Lausanne, Switzerland 9 \n2 Department of Computational Biology, University of Lausanne, Lausanne, Switzerland 10 \n3 Swiss Institute of Bioinformatics, Lausanne, Switzerland 11 \n4 University Center for Primary Care and Public Health, Lausanne, Switzerland 12 \n5 Institute of Molecular and Cell Biology, University of Tartu, Tartu, Estonia  13 \n6 Estonian Genome Centre, Institute of Genomics, University of Tartu, Tartu, Estonia 14 \n§ Estonian Biobank Research Team: Tõnu Esko, Andres Metspalu, Lili Milani, Reedik Mägi, Mari Nelis 15 \n# These authors jointly supervised this work. 16 \n 17 \n* Correspondence:  18 \nChiara Auwerx: chiara.auwerx@unil.ch; Alexandre Reymond  : alexandre.reymond@unil.ch; Zoltán 19 \nKutalik: zoltan.kutalik@unil.ch. 20 \n \n \n \n \n \n \n \n \n \n \n \n \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \nNOTE: This preprint reports new research that has not been certified by peer review and should not be used to guide clinical practice.\n\nAuwerx et al., 2023 2 \nABSTRACT 21 \nBackground: Copy-number variations (CNVs) have been associated with  rare and 22 \ndebilitating genomic syndromes but their impact on health later in life in the general population 23 \nremains poorly described. 24 \nMethods: Assessing four modes of CNV action, we  performed genome-wide association 25 \nscans (GWASs) between  the copy-number of CNV -proxy probes and 60 curated ICD-10 26 \nbased clinical diagnoses in 331,522 unrelated white UK Biobank participants with replication 27 \nin the Estonian Biobank. 28 \nResults: We identified 73 signals involving 40 diseases, all of which indicating that CNVs  29 \nincreased disease risk and  caused earlier onset. Even after correcting for these signals, a 30 \nhigher CNV burden increased risk for 18 disorders, mainly through the number of deleted 31 \ngenes, suggesting a polygenic CNV architecture. Number and identity of genes disturbed by 32 \nCNVs affected their pathogenicity, with many associations being supported by colocalization 33 \nwith both common and rare single nucleotide variant association signals . Dissection of 34 \nassociation signals provided insights into the epidemiology of known gene-disease pairs (e.g., 35 \ndeletions in BRCA1 and LDLR increased risk for ovarian cancer and ischemic heart disease, 36 \nrespectively), clarified dosage mechanisms of action (e.g., both increased and decreased 37 \ndosage of 17q12 impacts renal health), and identified putative causal genes (e.g., ABCC6 for 38 \nkidney stones). Characterization of the pleiotropic pathological consequences of recurrent 39 \nCNVs at 15q13, 16p13.11, 16p12.2, and 22q11.2 in adulthood indicated variable expressivity 40 \nof these regions and the involvement of multiple genes. 41 \nConclusions: Our results shed light on the prominent role of CNVs in  determining common 42 \ndisease susceptibility within the general population  and provide actionable insights allowing 43 \nto anticipate later-onset comorbidities in carriers of recurrent CNVs. 44 \n 45 \nKEYWORDS 46 \nStructural variation; CNV; GWAS; common diseases; pleiotropy; genomic disorders. 47 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 3 \nBACKGROUND 48 \nCopy-number variants (CNVs) refer to duplicated or deleted DNA fragments (≥ 50bp) and 49 \nrepresent an important source of inter-individual variation [1,2]. As a highly diverse mutational 50 \nclass, they can alter the copy-number of dosage sensitive genes, induce gain- or loss-of-51 \nfunction (LoF) through gene fusion or truncation, unmask recessive alleles, or disrupt 52 \nregulatory sequences, thereby representing potent phenotypic modifiers  [3]. As such, t heir 53 \nrole in human disease has mainly been  studied in clinically ascertained  cohorts often 54 \npresenting with congenital anomalies and/or severe neurological (e.g., developmental delay 55 \nand intellectual disability, epilepsy) or psychiatric (e.g., autism or schizophrenia) symptoms 56 \n[4–7] and today, close to 100 genomic disorders (i.e., disease caused by genomic 57 \nrearrangements) have been described  [8,9]. D espite their deleteriousness, some of these 58 \nCNVs flanked by repeats recurrently appear and remain at a low but stable frequency in the 59 \npopulation [10].  60 \n  61 \nThe emergence of large biobanks coupling genotype to phenotype data has fostered the study 62 \nof CNVs in the general population. Whole genome sequencing represents the best approach 63 \nto characterize the full human CNV landscape  [1,11,12] but current long- and short -read 64 \nsequencing association studies have limited samples size  [13–15]. Alternatively, l arger 65 \nsample sizes are available for exome sequencing data, offering the possibility to assess the 66 \nphenotypic consequence of small CNVs [16,17], while microarray-based CNV calls are better-67 \nsuited for the study of large CNV s and have been successful ly used in association studies  68 \n[8,18–28]. Performing a CNV genome -wide association study (GWAS) on 57  medically 69 \nrelevant continuous traits  in the UK Biobank (UKBB)  [29], we previously identified 131 70 \nindependent associations, including allelic series wherein carriers of CNVs at loci previously 71 \nassociated with rare Mendelian disorders exhibited subtle changes in disease-associated 72 \nphenotypes but lacked the corresponding clinical diagnosis [23]. Paralleling findings for point 73 \nmutations [30–33], this supports a model of variable expressivity, where CNVs can cause a 74 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 4 \nwide spectrum of phenotypic alteration ranging from severe, early -onset diseases to mild 75 \nsubclinical symptoms, opening the question as to whether these loci are also associated with 76 \ncommon diseases.  77 \n  78 \nUnlike continuous traits that can be  objectively measured in all participants , population 79 \ncohorts, such as UKBB, have low numbers of  diseased individuals [34]. Moreover, defining 80 \ncases relies on the arbitrary dichotomization of complex  underlying pathophysiological 81 \nprocesses [35]. Beyond the inherent loss of power associated to usage of binary variables 82 \n[36], cases might be missed because an  individual did not consult a physician, was 83 \nmisdiagnosed due to atypical clinical presentation, or is in a prodromal disease phase. Studies 84 \ninvestigating CNV-disease associations in the general population have either focused on only 85 \nfew diseases  [27,37] or well-established recurrent CNVs [20,38,39].  Alternatively, high-86 \nthroughput studies have assessed  a broad range  of continuous and binary traits 87 \nsimultaneously [17,24,25] without any precautions to accommodate the aforementioned 88 \nchallenges. To date, the largest disease CNV-GWAS meta-analyzed ~1,000,000 individuals 89 \n[8]. While boost ing power through increased sample size, it comes at the cost of extensive 90 \ndata harmonization, resulting in the exclusion of smaller CNVs (≤ 100kb) and broader disease 91 \ncategories (e.g., “immune a bnormality”). Moreover, as th is study includes several clinical 92 \ncohorts, phenotypes are biased towards neuropsychiatric disorders (24/54 phenotypes) for 93 \nwhich the role of CNVs is well-established [4–7].   94 \n 95 \nUsing tailored CNV -GWAS models mimicking four mechanisms of CNV action and time -to-96 \nevent analysis , we  investigate the relation ship between CNVs and 60  carefully defined  97 \ncommon diseases affecting a broad range of physiological systems in 331,522 unrelated white 98 \nUKBB participants. Extensively validating our results , we report associations according to 99 \nconfidence tiers and take advantage of rich individual-level phenotypic data to demonstrate 100 \nthe contribution of CNVs to the common disease burden in the general population. 101 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 5 \nMETHODS 102 \n1. Study material 103 \nDiscovery cohort: UK Biobank  104 \nThe UK Biobank (UKBB) is composed of ~500,000 volunteers (54% females) from the general 105 \nUK population for which microarray -based genotyping and extensive phenotyping data  – 106 \nincluding hospital based International Classification of Diseases, 10th Revision (ICD -10) 107 \ncodes (up to September 2021) and self-reported conditions – are available [29]. Participants 108 \nsigned a broad informed consent form and data were accessed through application #16389. 109 \n 110 \nReplication cohort: Estonian Biobank  111 \nThe Estonian Biobank (EstBB) is a population-based cohort of ~208,000 Estonian individuals 112 \n(65% females; data freeze 2022v01 [12/04/2022]) for which microarray-based genotyping data 113 \nand ICD-10 codes from crosslinking with national and hospital databases (up to end 2021) are 114 \navailable [40]. The activities of the EstBB are regulated by the Human Genes Research Act, 115 \nwhich was adopted in 2000 specifically for the operations of the EstBB. Individual level data 116 \nanalysis in the EstBB was carried out under ethical approval 1.1 -12/624 from the Estoni an 117 \nCommittee on Bioethics and Human Research (Estonian Ministry of Social Affairs), using data 118 \naccording to release application 3-10/GI/34668 [20/12/2022] from the EstBB. All participants 119 \nsigned a broad informed consent form. 120 \n 121 \nOther public resources 122 \nThe PheCode Map 1.2 (beta) (https://phewascatalog.org/phecodes_icd10) was used for ICD-123 \n10 code classification  [41]. Genomic regions were annotated with the NHGRI-EBI GWAS 124 \nCatalog (https://www.ebi.ac.uk/gwas/; 24/11/2022) [42], and the Online Mendelian Inheritance 125 \nin Man (OMIM; https://www.omim.org/; 27/07/2022) [43]. Recurrent CNV coordinates w ere 126 \nretrieved from DECIPHER ( https://www.deciphergenomics.org/) [9]. Unless specified 127 \notherwise, Neale UK Biobank  single-nucleotide polymorphism  (SNP)-GWASs summary 128 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 6 \nstatistics were used ( http://www.nealelab.is/uk-biobank). Allele frequencies and genomic 129 \nconstraint scores (probability of LoF Intolerance (pLI); LoF Observed over Expected Upper 130 \nbound Fraction (LOEUF) ) originate from the Genome Aggregation Database (GnomAD; 131 \nhttps://gnomad.broadinstitute.org/) [44]; pHaplo and pTriplo scores from  [8]. Tissue-specific 132 \ngene expression was assessed in the Genotype -Tissue Expression project (GTEx; 133 \nhttps://gtexportal.org/home/) [45]. 134 \n 135 \nSoftware versions 136 \nCNVs were called with PennCNV v1.0.5  [46] using PennCNV-Affy (27/08/2009) and filtered 137 \nbased on a quality scoring  pipeline [47]. Genetic analyses were conducted with PLINK v1.9 138 \nand v2.0 [48]. ANNOVAR (24/10/2019) was used to map genes to genetic regions [49]. The 139 \nUCSC Genome Browser was used to determine the human genome size (GRCh37/hg19) and 140 \nthe LiftOver tool was used to lift over genomic coordinates  [50]. Statistical analyses were 141 \nperformed with R v3.6.1 and graphs were generated with R v4.1.3. 142 \n 143 \n2. CNV association studies in the UK Biobank 144 \nMicroarray-based CNV calling 145 \nUKBB g enotype microarray data were acquired  from two arrays with 95% probe overlap 146 \n(Applied Biosystems UK Biobank Axiom Array : 438,427 samples; Applied Biosystems UK 147 \nBiLEVE Axiom Array by Affymetrix: 49,950 samples) [29] and used to call CNVs as previously 148 \ndescribed [23]. Briefly, CNVs were called using standard PennCNV settings and samples on 149 \ngenotyping plates with a mean CNV count per sample > 100 and samples with > 200 CNVs 150 \nor single CNV > 10 Mb were excluded. Remaining CNVs were attributed a probabilistic quality 151 \nscore (QS) ranging from -1 (likely deletion) to 1 (likely duplication) [47]. High confidence CNVs, 152 \nstringently defined by |QS| > 0.5, were retained and encoded in chromosome-wide probe-by-153 \nsample matrices (i.e., entries of 1, -1, or 0 indicate probes overlapping high confidence 154 \nduplication, deletion, or no/low quality CNV, respectively) [47], which were converted into three 155 \nPLINK binary file sets to accommodate association analysis according to four modes of CNV 156 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 7 \naction. Details about the CNV encoding  and handling of chromosome X  are provided in 157 \nSupplemental Note 1. Probe-level CNV frequency was calculated [23]. All results in this study 158 \nare based on the human genome reference build GRCh37/hg19. 159 \n 160 \nCase-control definition and age-at-disease onset calculation 161 \nA pool of 331,522 unrelated white British UKBB participants (54% females) was considered 162 \nafter excluding retracted (up to August 2020), as well as related, high missingness and non-163 \nwhite British samples (used.in.pca.calculation = 0 and in.white.British.ancestry.subset = 0 in 164 \nSample-QC v2 file) . CNV outliers (Microarray-based CNV calling) and individuals reporting  165 \nblood malignancies (i.e., possibly harboring somatic CNVs; UKBB field #20001: 10047, 1048, 166 \n1050, 1051, 1052, 1053, 1055, 1056, 1056 ; #41270: ICD-10 codes mapping to  PheCode 167 \nexclusion range “cancer of lymphatic and hematopoietic tissue”), were further excluded. 168 \n 169 \nCases and controls were assigned for 60 ICD-10-based clinical diagnoses using diagnosis – 170 \nICD10 (#41270), cancer code, self -reported (#20001), and non-cancer illness code, self -171 \nreported (#20002) to build exclusion and inclusion lists. For each disease, we first defined all 172 \n331,522 individuals as control s. We excluded  individuals with a self-reported or hospital 173 \ndiagnosis of a broad set of conditions that include the disease of interest, as well as related 174 \ndisorders/ICD-10 codes (e.g., other cancers or radio-/chemotherapy for breast cancer; mood 175 \nor personality disorders for schizophrenia). We then re-introduce as cases individuals having 176 \nreceived a restricted range of ICD-10 codes matching our disease definition. For second level 177 \nICD-10 codes, all subcodes are considered, otherwise only the specified ones. The disease 178 \nburden was calculated as the number of diagnoses (out of the 60 assessed) an individual has 179 \nreceived. For male- (prostate cancer) and female - (menstruation disorders, endometriosis, 180 \nbreast cancer, ovarian cancer ) specific diseases, downstream analyses were conducted 181 \nexcluding individuals from the opposite sex. 182 \n 183 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 8 \nBased on the date at first in -patient diagnosis – ICD10 (#41280) and the individual’s month 184 \n(#52) and year (#34) of birth (birthday assumed on average to be the 15th), the age at diagnosis 185 \nwas calculated by subtracting the earlie st diagnosis date for codes on the inclusion list from 186 \nthe birth date and converting it to years by dividing by 365.25 to account for leap years.  187 \n 188 \nProbe and covariate selection 189 \nRelevant covariates and probes were pre-selected to fit tailored main CNV-GWAS models and 190 \nreduce computation time. For each disease, a logistic regression was fitted to explain disease 191 \nprobability as a function of age (#21003), sex, genotyping array , and the 40 first principal 192 \ncomponents (PCs). Nominally significantly associated covariates (p ≤ 0.05) were retained for 193 \nthe main analysis . CNV-proxy probes with a CNV frequency ≥ 0.01% were pruned at r2 > 194 \n0.9999 in PLINKCNV (–-indep-pairwise 500 250 0.9999 PLINK v2.0) to group probes 195 \nat the core of CNV regions while retaining resolution at breakpoints (BPs), resulting in 18,725 196 \nprobes. For each disease , 2-by-3 genotypic Fisher test s assessed dependence between 197 \ndisease status and probe copy-number (rows: control versus case; columns: deletion versus 198 \ncopy-neutral versus duplication; –-model fisher PLINK v1.9; TEST column GENO). Probes 199 \nwith p ≤ 0.001 and a minimum of two disease cases among CNV, duplication, or deletion 200 \ncarriers were retained for  assessment through  the mirror/U -shaped, duplication -only, or 201 \ndeletion-only model, respectively.  202 \n 203 \nGenome-wide significance threshold 204 \nDue to the recurrent nature of CNVs, the 18,725 probes retained after frequency filter and 205 \npruning remain highly correlated  and are thus not independent. Accounting for all of them 206 \nwould result in an overly strict multiple testing correction. Using an established protocol  207 \n[22,23,51], we estimated the number of effective tests performed to  Neff = 6,633, setting the 208 \ngenome-wide (GW) threshold for significance at p ≤ 0.05/6,633 = 7.5 x 10-6. This threshold is 209 \nof the same order of magnitude as what others have estimated for disease CNV-GWAS [8]. 210 \n 211 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 9 \nMain CNV-GWAS model 212 \nAssociation between  disease risk and  copy-number of CNV-proxy probes was assessed 213 \nthrough logistic regression with Firth fallback (--covar-variance-standardize –-glm 214 \nfirth-fallback omit-ref no-x-sex hide-covar -–ci 0.95 PLINK v2.0), using 215 \ndisease- and model -specific probes and covariates ( Probe and covariate selection ). Four 216 \nassociation models were assessed: the mirror model assesses the additive effect of each 217 \nadditional copy (PLINKCNV); the U-shape model assesses a consistent effect of any deviation 218 \nfrom the copy neutral state  (PLINKCNV, using the hetonly option in glm PLINK v2.0); the 219 \nduplication-only model (PLINKDUP) assesses the impact of a duplication while disregarding 220 \ndeletions; the deletion-only model  (PLINKDEL) assesses the impact of a deletion while 221 \ndisregarding duplications. Odds ratios (OR) and their 95% confidence interval (CI)  were 222 \nharmonized (A1 to “T”; Supplemental Note 1  – Table 1 ), i.e., 𝑂𝑅!\"# =\n$\n%&!\"#\n and 223 \n𝐶𝐼!\"# \t =\t𝑒'() (%&$%&)±$../∗12'() \t(,-!\"#), respectively. GW-significant associations (p ≤ 7.5 x 10-6; 224 \nGenome-wide significance threshold ) were pruned at r2 > 0.8 (--indep-pairwise 3000 225 \n500 0. 8 PLINK v2 .0), giving priority to probes with the strongest association signal by 226 \ninputting a scaled negative logarithm of association p -value as frequency (–-read-freq 227 \nPLINK v2.0). For the U-shape model, pruning was performed using custom code by extracting 228 \nprobes from PLINKCNV and recoding them to match U-shape numerical encoding. Number of 229 \nindependent signals per disease was determined by stepwise conditional analysis. Briefly, for 230 \neach disease and association model, the numerical CNV genotype of the lead probe was 231 \nincluded along selected covariates in the logistic regression model. This process was repeated 232 \nuntil no more GW-significant signal remained. 233 \n 234 \nDue to its continuous nature , the disease burden CNV -GWAS was based on linear 235 \nregressions between copy-number of selected probes and the disease burden (–-glm omit-236 \nref no-x-sex hide-covar allow-covars PLINK v2 .0), correcting for selected 237 \ncovariates. Post-GWAS processing was performed as previously described [23]. 238 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 10 \n 239 \nCNV region definition and annotation  240 \nCNV region (CNVR) boundaries were defined by the most distant probe within ± 3Mb and r2 ≥ 241 \n0.5 of independent lead probes (–-show-tags --tag-kb 3000 --tag-r2 0.5 PLINK 242 \nv1.9; U -shape model: custom code, as described previously for pruning). Signals from 243 \ndifferent models were merged when overlapping (≥ 1bp) and involving the same disease, with 244 \nCNVR boundaries defined as the maximal CNVR. Characteristics of the most significant 245 \nmodel (i.e., “best model”) are reported. The “main model” indicates which CNV type mainly 246 \ndrives the association, i.e., when associations were found through multiple models, priority 247 \nwas given to either the duplication-only or deletion -only models, otherwise to the model 248 \nyielding the lowest p -value. CNVRs were annotated with  hg19 HGNC and ENSEMBL gene 249 \nnames using annotate_variation.pl from ANNOVAR (--geneanno). Number of genes 250 \nmapped to a CNVR was calculated and set to zero for CNVRs with REGION not equaling 251 \n“exonic”. Resulting count was used as a predictor for pleiotropy (i.e., number of associations) 252 \nthrough linear regression. 253 \n 254 \nStatistical confidence tiers  255 \nFollowing primary assessment through logistic regression (Main CNV-GWAS model), three 256 \nstatistical approaches were implemented to gauge robustness of the lead probe’s association 257 \nsignal. First, we assessed post hoc the p-value of 2-by-3 genotypic Fisher tests (Probe and 258 \ncovariate selection). Second, we transformed the binary disease status into a continuous 259 \nvariable by computing the response residuals of the logistic regression of disease status on 260 \ndisease-relevant covariates. This allowed the usage of linear regressions to estimate the effect 261 \nof the CNV genotype (encoded according to all significantly associated models in the primary 262 \nanalysis) on disease risk. The model generating the lowest p-value for the CNV encoding is 263 \nreported. Third, time-to-event analysis was used to assess  whether CNVs influence age -at-264 \ndisease onset. Age at last healthy measurement was calculated as age -at-disease onset for 265 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 11 \ncases and date of last recorded diagnosis (30/09/2021) minus birth date converted  to years 266 \nfor controls (Case-control definition and age-at-disease onset calculation). Cox proportional-267 \nhazards (CoxPH) models were fitted including disease -relevant covariates and numerically 268 \nencoded CNV genotype for either of the four association models as predictors, using coxph() 269 \nfunction from the R survival package [52]. The model with the lowest  CNV genotype p-270 \nvalue is reported. CNV-disease associations were classified in confidence tiers depending on 271 \nwhether they were confirmed by  3 (tier 1) , 2  (tier 2) , or 1  (tier 3)  of the above -described 272 \napproaches at the  arbitrary validation significance threshold of p ≤ 1  x 10-4. Validation 273 \napproaches not being suited for continuous variables – which do not suffer from the same 274 \ncaveats as binary traits – all disease burden associations were classified as tier 1. 275 \n 276 \nLiterature-based supporting evidence 277 \nUsing three literature-based approaches, we examined whether disease-associated CNVRs 278 \nhad previously been linked to relevant phenotypes. First, we investigated the colocalization of 279 \nautosomal CNVRs with SNP-GWAS signals. GRCh38/hg38 lifted CNVR coordinates were 280 \ninputted in the GWAS Catalog and associations (p ≤ 1x10 -7) relevant to the inve stigated 281 \ndisease (i.e., synonym, continuous proxy, or major risk factor) were identified through manual 282 \ncuration. Second, we overlapped OMIM morbid genes (i.e., linked to an OMIM disorder; 283 \nmorbidmap.txt) with disease-associated CNVRs. Through manual curati on, we flagged 284 \nOMIM genes associated to Mendelian disorders  sharing clinical features with the common 285 \ndisease associated through CNV-GWAS. Third, we examined if implicated CNVRs overlapped 286 \nregions at which CNVs were found to modulate continuous traits [23] or disease risk [20,25].  287 \n 288 \n3. Replication in the Estonian Biobank 289 \nCNV calling and sample selection 290 \nAutosomal CNVs were called from Illumina Global Screening Array (GSA) genotype data for 291 \n193,844 individuals that survived general quality control  and had matching genotype -292 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 12 \nphenotype identifiers, matching inferred versus reported sex, a SNP-call rate ≥ 98%, and were 293 \nincluded in the EstBB SNP imputation pipeline. CNV outliers and individuals with a reported 294 \nblood malignancy were excluded, as previously described. High confidence CNV calls (|QS| 295 \n> 0.5) of the 156,254 remaining individuals were encoded into three PLINK binary file sets, 296 \nfollowing the procedure described for the UKBB (CNV association studies in the UK Biobank).  297 \n 298 \nEstBB disease definition 299 \nDisease cases and disease burden were defined similarly than in the UKBB. To account for 300 \ndifferences in recording practices between the countries, Z12 (routine preventive screens for 301 \ncancer), and D22-23 (benign skin lesions) subcodes, were removed from the exclusion list of 302 \ncancer traits as they were much more frequent than in the UKBB and strongly reduced the 303 \nnumber of controls. Due to lack of matching data in the EstBB, no self-reported diseases and 304 \ncancers were used as an exclusion criterion for disease definition.  305 \n 306 \nEstBB replication analysis 307 \nRelated individuals with available CNV calls were pruned (KING kinship coefficient > 0.0884), 308 \nprioritizing individuals whose disease status was least often missing, leaving 90,211 unrelated 309 \nsamples for the replication study. Disease-relevant covariates were selected among sex, year 310 \nof birth, genotyping batch (1 -11), and PC1 -20. For each of the 73 UKBB signals, probes 311 \noverlapping the CNVR and with a n EstBB CNV, duplication, or deletion frequency ≥ 0.01%, 312 \nwere retained, depending on whether the mirror/U -shape, duplication-only, or deletion -only 313 \nwas the best UKBB model , respectively. Association studies were performed on remaining 314 \nprobes using disease-specific covariates and the best UKBB model, following the previously 315 \ndescribed procedure. Forty signals (55%) could not be assessed due to null/low CNV 316 \nfrequency, failure of the regression to converge, absence of at least one case CNV carrier, or 317 \nbecause not mapping on the autosomes. For the remaining 33 signals, summary statistics of 318 \nthe probe showing the strongest association within the CNVR were retained and p-values 319 \nwere adjusted to account for directional concordance with UKBB effect s by rewarding and 320 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 13 \npenalizing signals with matching and non-matching effect size signs, respectively. Specifically, 321 \none-sided p-values were obtained as 𝑝!\"# =\t\n3!\"#\n4  and 𝑝!\"# =\t1 − (\n3!\"#\n4 ) for 24 concordant and 322 \n9 non -concordant signals, respectively.  Accounting for 33 testable signals, the replication 323 \nthreshold for significance was set at  p ≤ 0.05/ 33 = 1.5 x 10-3. One-sided binomial tests 324 \n(binom.test()) were used to assess e nrichment of observed versus expected significant 325 \nassociations at various thresholds ( 𝛼 = 0.1 to 0.005 by steps of 0.005) , with the R function 326 \narguments: 𝑥 the number of observed signals at 𝛼, 𝑛 the number of testable signals (i.e., 33), 327 \nand 𝑝 the expected probability of signals meeting 𝛼 (i.e., 𝛼). 328 \n 329 \n4. CNV region constraint analysis 330 \nEvolutionary c onstraint of genes overlapping disease-associated CNVRs , i.e., “disease 331 \ngenes” (CNV region definition and annotation), was assessed by comparing their pLI, LOEUF, 332 \npHaplo, and pTriplo scores to the ones of “background genes”. The latter were identified by 333 \nannotating ranges of  one or multiple consecutive probes with CNV frequency ≥ 0.01% with 334 \nANNOVAR (hg19 HGNC gene names) and excluding disease genes. For pLi and LOEUF, all 335 \ndisease genes were considered together. For pHaplo and pTriplo,  two disease gene groups 336 \nwere considered: genes overlappin g CNVRs with at least one association through the 337 \nduplication-only model and genes overlapping CNVRs with at least one association through 338 \nthe deletion-only model. As many CNVRs associated through both models, the analysis was 339 \nrepeated considering genes ov erlapping CNVRs with at least one association through the 340 \nduplication-only and none through the deletion -only model and vice-versa. Comparison with 341 \nbackground genes was done through two-sided Wilcoxon rank-sum test.  342 \n 343 \n5. Extended phenotypic assessment 344 \nTo elaborate on specific associations, we made use of the rich phenotypic data available for 345 \nUKBB participants , as detailed  in Supplemental Note 2. For fine -mapping of association 346 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 14 \nsignals, CNV carriers were divided in subgroups based on visual inspection of C NV 347 \nbreakpoints and segmental duplications, as detailed in Supplemental Note 3.  348 \n 349 \nCNV versus copy-neutral comparisons 350 \nComparisons between groups of CNV carriers and copy -neutral individuals always exclude 351 \nlow quality CNV (|QS| ≤ 0.5) carriers altogether. For diseases, prevalence is estimated as 𝑞 =352 \n5\n!  , with 𝑐 and 𝑛 are the number of cases and total number of individuals in a group, and 353 \n𝑆𝐸(𝑞) =\t 4 6∗($76)\n! . Differences in prevalence compared to cop y-neutral individuals were 354 \nassessed by two-sided Fisher test. For continuous traits, comparisons are based on two-sided 355 \nt-tests. 356 \n 357 \n6. CNV burden analyses  358 \nCNV burden association studies 359 \nIn the UKBB, individual -level CNV, duplication, and deletion  burden were calculated as the 360 \nnumber of Mb or genes affected by high -confidence (|QS| > 0.5) autosomal CNVs , 361 \nduplications, and deletions, respectively,  as previously described [23]. Association between 362 \nburden values and the  60 diseases  (logistic regress ion) or the di sease burden  (linear 363 \nregression), was assessed including disease-relevant covariates in the model. Accounting for 364 \nthe 61 evaluated traits, significance was defined at p ≤ 0.05/61 = 8.2 x 10-4. We next corrected 365 \nburden values for CNV-GWAS signals. For each disease, CNVs, duplications, and deletions 366 \noverlapping (≥ 1bp) a CNVR significantly associated with the disease of interest through CNV-367 \nGWAS were omitted from the CNV, duplication, and deletion burden calculations if the CNVR 368 \nhad been found to associate with the disease through the mirror/U-shape, duplication-only, or 369 \ndeletion-only model, respectively. Association studies were repeated using corrected burden 370 \nvalues. Only the most significant burden types are reported in the text. 371 \n 372 \nRelative importance of protein coding regions in mediating the burden’s effect 373 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 15 \nThe average genome -wide gene density ( 𝐺𝐷89) was estimated to 8.4 genes/Mb based on 374 \n26,289 genes (i.e., unique HGNC gene names in hg19 RefSeq, excluding microRNAs but 375 \nincluding genes of uncertain function ( i.e., “LOC”)) and a human genome length of 376 \n3,137,161,264 bp. If CNVs affecting the coding and non-coding DNA have similar effects, we 377 \nexpect that the impact of 1Mb affected by CNVs to be equivalent to 8.4 genes being affected 378 \nby CNVs. Hence, the association effect size  of the CNV burden measured in Mb (𝛽:;) is 379 \nexpected to be 8.4-times larger than the one measured in number of affected genes (𝛽<\"!\"),  380 \ni.e.,  \n=/0\n=1%$%\n=8.4. This hypothesis was tested independently for the deletion and duplication 381 \nburdens for traits with at least one significant uncorrected burden association (i.e., 20 diseases 382 \n+ disease burden). Significant deviations from the expected ratio were assessed by t-statistic: 383 \n𝑡 =\t\n8.4 − \t 𝛽<:;\n𝛽<<\"!\"\n𝑆𝐷=  384 \nwhere 𝛽<:; and 𝛽<<\"!\" are the estimated effects of the burden measured in Mb or number of 385 \ngenes impacted by CNVs, respectively, on the assessed trait. 𝑆𝐷=  is the empirically observed 386 \nstandard deviation of the \n=>/0\n=>1%$%\n ratio, estimated based on 10,000 simulations of 𝛽>:;\t~\t𝑁(𝛽<:;,387 \n𝑉𝑎𝑟E (𝛽<:;)) and 𝛽><\"!\" \t~\t𝑁(𝛽<<\"!\" ,\t\t\t𝑉𝑎𝑟E (𝛽<<\"!\" )). P-values were computed based on a two-sided 388 \none sample t -test and deemed significant at p ≤ 0.05/ 21 = 2.4  x 10-3. This analysis was 389 \nrepeated for the modified burden definition s not accounting for CNVs overlapping disease -390 \nassociated CNVRs. 391 \n 392 \nRESULTS 393 \nThe spectrum of common diseases in the UK Biobank 394 \nTo mitigate issues related to disease definition, we used a three-step approach to designate 395 \ncases and controls in the UKBB (Figure 1A; top). Starting from 331,522 unrelated white British 396 \nindividuals, we defined cases based on a narrow list of hospital -based diagnoses (i.e., ICD-397 \n10 codes) and excluded self-reported cases, as well as self-reported and hospital diagnoses 398 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 16 \nof related conditions. Sixty disorders spanning 12 ICD -10 chapters were selected to cover a 399 \nwide range of physiological systems, favoring conditions with sufficiently large sample size 400 \nand a likely genetic basis ( Figure 1B ; Figure S1A; Table S1 ). Except for systemic lupus 401 \nerythematosus (N = 422) and polycystic kidney disease (N = 454), all diseases had over 500 402 \ncases. Nineteen diseases had a case count > 10,000, with osteoarthrosis (N = 62,175) and 403 \nessential hypertension (N = 97,860) being the most frequent. Seven diseases h ad a median 404 \nage of onset ≤ 60 years , predominantly female reproductive disorders , autoimmune 405 \nconditions, and psychiatric  disorders. Conversely, the nine diseases with a  median age at 406 \nonset ≥ 70 years were mainly degenerative disorders of the brain, eye , and kidney, overall 407 \naligning align with epidemiological knowledge of the respective diseases.408 \n 409 \nFigure 1. Overview of the study 410 \n(A) Schematic representation of the analysis workflow. Trait definition: For each of the 60 investigated 411 \ndiseases, unrelated white British UK Biobank participants were assigned as controls, individuals self -412 \nreporting or diagnosed with the disease of interest or a broader set of related conditions were excluded 413 \nand set as missing, and individuals with a hospital-based ICD-10 diagnosis of the condition of interest 414 \nwere re-introduced as cases. Primary association study: Disease -specific relevant covariates were 415 \nselected. Probes were pre -filtered based on copy -number variant (CNV) frequency , required to 416 \nassociate with the disease, and a minimum of two diseased carriers was required for the probe to be 417 \ncarried forward. Disease- and model-specific covariates and probes were used to generate tailored 418 \nCNV genome-wide association studies (GWASs) based on Firth fallback logistic regression according 419 \nto a mirror, U -shape, duplication -only (i.e., considering only duplications), and deletion -only (i.e., 420 \nconsidering only deletions) models. Independent lead signals were identified through stepwise 421 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 17 \nconditional analysis and CNV regions were defined based on probe correlation and merged across 422 \nmodels. Validation: Statistical validation methods (i.e., Fisher test, residuals regression , and Cox 423 \nproportional hazards model (CoxPH)) were used to rank associations in confid ence tiers. Literature 424 \nvalidation approaches leverage data from independent studies to corroborate that genetic perturbation 425 \n(i.e., single-nucleotide polymorphisms (SNP), rare variants from the OMIM database, and CNVs) in the 426 \nregion are linked to the disease. Independent replication in the Estonian Biobank. (B) Age at onset for 427 \nthe 60 assessed diseases, categorized based on ICD-10 chapters and colored according to case count. 428 \nData are represented as boxplots; outliers are not shown. 429 \n \nCopy-number variant genome-wide association study  430 \nTo assess whether disease susceptibility is modulated by CNVs, we performed CNV genome-431 \nwide association studies (GWASs), i.e., test if the copy-number of selected probes influence 432 \nthe probability to develop a  disease or an individual’s  disease burden (i.e., number of 433 \ndiagnoses among the 60 studied diseases ) (see Methods; Figure 1A ; middle).  Briefly, 434 \nmicroarray-called CNVs for 331,522 unrelated white British individuals were transformed to 435 \nthe probe level after quality -control [23]. As CNVs can act through different gene dosage 436 \nmechanisms, four association models were assessed: mirror and U-shape models consider 437 \ndeletions and duplications simultaneously, assuming that they impact disease risk in opposite 438 \nor identical direction, respectively, while the CNV type-specific duplication- and deletion-only 439 \nmodels assess independently the effect of duplications and deletions, respectively. To reduce 440 \nthe number and complexity of implemented logistic regressions, pre-processing steps  441 \nselected relevant covariates and probes for each disease and model combination , thereby 442 \nlowering computation time and decreasing the multiple testing burden (Figure S2; Table S2). 443 \nAll summary statistics are available (Data Availability). 444 \n 445 \nStepwise conditional analysis narrowed GW significant associations (p ≤ 7.5 x 10 -6; see 446 \nMethods for threshold calculation) to 40, 41, 21, and 38 independent signals for the mirror, U-447 \nshape, duplication-only, and deletion-only models, respectively. These were combined into 70 448 \nrisk-increasing (i.e., no  disease-protecting CNV) associations and 3 disease burden 449 \nassociations (Figure 2; Table S3). Forty-five associations (45/73 = 62%) were supported at 450 \nGW significance by multiple models, the lowest p -value (i.e., “best model”)  being obtained 451 \nthrough the mirror, deletion-only, U-shape, and duplication-only models for 24, 23, 21, and 5 452 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 18 \nof the signals, respectively. No association was detected at GW significance by both the 453 \nduplication-only and deletion-only models, so that each signal was attributed a “main model” 454 \nthat indicates whether the association is primarily driven by duplications or deletion s (see 455 \nMethods; Figure 2). The main model should be interpreted with caution as both deletions and 456 \nduplications might influence disease risk but only one CNV type -specific model might reach 457 \nGW significance (e.g., due to higher frequency) . This is particularly relevant as 73% (33/45) 458 \nof disease-associated CNV regions ( CNVRs) have a hig her duplication than deletion 459 \nfrequency (Figure 2A). Hence, 95% (20/21) of signals mainly driven by duplications were also 460 \nidentified by the mirror/U-shape model(s) and contribution of deletions cannot be excluded.  461 \n \nFigure 2. CNV-disease association map 462 \n(A) Duplication and deletion frequencies ([%]; y-axis; break: //) of the lead probe for each unique and 463 \nnon-overlapping disease-associated CNV region (CNVR), labeled with corresponding cytogenic band 464 \n(x-axis; 16p11.2 is split to  distinguish the distal 220kb breakpoint BP2 -3 and proximal 600kb BP4 -5 465 \nCNVRs; non-overlapping CNVRs on the same cytogenic band are numbered). If signals mapping to 466 \nthe same CNVR have different lead probes, the maximal frequency was plotted. ( B) Associations 467 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 19 \nbetween CNVRs (x-axis) and diseases (y-axis) identified through CNV-GWAS. Color indicates the main 468 \nassociation model. Size and transparency reflect the statistical confidence tier. Black contours indicate 469 \noverlap with OMIM gene causing a disease with shared phenotypic features. Black crosses indicate 470 \noverlap with SNP-GWAS signal for a related trait. Grey shaded lines indicate CNVRs with continuous 471 \ntrait associations [23]. N provides count for various features. 472 \n \nValidation of identified CNV-GWAS signals 473 \nAcross the 45 CNVRs, CNV frequencies were low, ranging between 0.01% ( our frequency 474 \ncutoff) and 0.36%, with 87% ( 39/45) of  CNVRs having a frequency ≤ 0.1% ( Figure 2 A). 475 \nConsequently, associations rely on a low number of diseased CNV carriers  and require 476 \nvalidation (see Methods; Figure 1A; bottom; Figure 2B; Table S3). We used three statistical 477 \napproaches to assess the robustness of CNV-diseases associations: i) Fisher test, ii) residual 478 \nregression, and iii) time -to-event analysis through CoxPH modeling. We replicated 28/70 479 \n(40%), 23/70 (33%), and 70/70 (100%) of the associations with the respective methods at the 480 \narbitrary validation threshold of p ≤ 10 -4. This allowed to stratify associations in confidence 481 \ntiers, with 17  signals replicating with all methods (tier 1), 20 with two (tier 2), and 36 only 482 \nthrough time-to-event analysis (tier 3). Importantly, time-to-event analysis showed that CNVs 483 \nalways contributed to an earlier age of disease onset, in line with the paradigm that diseases 484 \nwith a strong genetic etiology have earlier onset [53]. 485 \n 486 \nIn parallel, we gathered literature evidence linking genetic variation at CNVRs with relevant 487 \nphenotypes (Table S3). Forty-eight signals (48/73 = 64%) mapped to a CNVR harboring a 488 \nleast one OMIM morbid gene and in 15 cases, the gene was linked to a Mendelian disorder 489 \nsharing phenotypic features with the associated common disease . For instance, association 490 \nbetween 4q35 CNVs and corneal conditions (chr4:186,687,554-187,182,384; ORU-shape = 18.2; 491 \n95%-CI [5.2; 63.1]; p = 5.0 x 10-6) encompasses CYP4V2 [MIM: 608614], a gene associated 492 \nwith autosomal recessive Bietti crystalline corneoretinal dystrophy [MIM: 210370], a disorder 493 \nthat impairs vision and progresses to blindness  by age 50-60 years [54]. We next assessed 494 \nwhether SNPs overlapping disease-associated CNVRs were reported to associate with the 495 \nimplicated disease or a biomarker thereof in the GWAS Catalog. This was the case for 28 496 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 20 \n(28/66 = 42%) autosomal signals, a similar proportion (38%) than for continuous trait CNV -497 \nGWAS [23]. For instance,  distal 22q11.2  CNVs increased risk for disorders of mineral 498 \nmetabolism (chr22:21,797,101-22,661,627; ORmirror = 0.02; 95%-CI [0.006; 0.083]; p = 9.9 x 499 \n10-9) and overlap ped heel bone mineral density SNP -GWASs signals, while 3q29 CNVs 500 \nincreased Alzheimer’s disease risk (chr3:196,953,177-197,331,898; ORU-shape = 11.8; 95%-CI 501 \n[4.0; 34.7]; p = 6.6 x 10 -6) and overlap ped with SNP-GWAS signal for PHF -tau levels, and 502 \nsuggestive sig nals (p < 5 x 10 -6) for  frontotemporal dementi a and cognitive decline in 503 \nAlzheimer’s disease. Finally, 37 signals (37/73 = 51%) mapped to nine CNVRs previously 504 \nfound to be associated with complex traits [23].  505 \n 506 \nWe also set out to replicate association signals in 90,211 unrelated EstBB individuals [40], 507 \nusing similarly case definition than for the UKBB  (see Methods; Figure S1B). Requesting at 508 \nleast one diseased CNV carrier, 33 of 73 associations could be evaluated, among which four 509 \nwere strictly replicated (p ≤ 0.05/33 = 1.5 x 10 -3) and five additional ones reached nominal 510 \nsignificance (p ≤ 0.05)  (Table S3). Compared to what would be expected by chance, t his 511 \ncorresponds to a 5.5-fold (pbinomial = 2.5 x 10-5) and 30.3-fold (pbinomial = 6.6 x 10-7) enrichment 512 \nfor replication at p ≤ 0.05 and p ≤ 5 x 10-3, respectively (Figure 3A). Despite low power, these 513 \nresults support validity of the primary UKBB association signals. Signals replicating at nominal 514 \nsignificance are detailed in Figure 3B . Many harbor SNP -GWAS signals for related 515 \nphenotypes (7/9), relevant morbid OMIM genes (4/9), or map to CNVRs previously associated 516 \nwith similar diseases (6/ 7) or biomarkers (4/9). Among them, two are in the lowest UKBB 517 \nconfidence tier. 15q13 duplications were linked to increased risk for acute kidney injury (AKI; 518 \nchr15:30,946,160-31,881,106 | UKBB: ORdup = 4.6; 95%-CI [2.5; 8.4]; p = 7.1 x 10-7 | EstBB: 519 \np = 2.4 x 10-4). Homozygous mutations in FAN1 [MIM: 613534], one of the five genes mapping 520 \nto this CNVR, have been  linked to karyomegalic interstitial nephritis [MIM: 614817], a 521 \nprogressive renal condition that leads to CKD [55]. The second example links CNVs affecting 522 \nexon 2 and intron 2-3 of PRKN ([MIM: 602544]) – a gene causing juvenile autosomal recessive 523 \nParkinson’s disease [MIM: 600116] – to sleep disorders such as insomnia and hypersomnia 524 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 21 \n(chr6:162,705,164-162,873,489 | UKBB: ORmirror = 0.12; 95%-CI [0.05; 0.26]; p = 1.6 x 10 -7 | 525 \nEstBB: p = 0.047). This finding is particularly relevant given the region’s high CNV frequency 526 \n(0.22%) and the fact that sleep disturbances are among the earliest symptoms of Parkinson’s 527 \ndisease [56]. Follow-up studies should determine whether these individuals are more prone 528 \nto develop Parkinson’s disease in the future. 529 \n \nFigure 3. Replication of CNV-disease associations in the Estonian Biobank 530 \n(A) Enrichment for signal replication (y-axis; 95% confidence interval as grey ribbon) at different levels 531 \nof significance (alpha; x-axis) in the Estonian Biobank (EstBB). Color and size indicate the p -value of 532 \nthe enrichment (one -sided binomial test) and the number of observed associations, respectively. 533 \nDashed red line indicates 1x enrichment, i.e., the number of observed associations matches the number 534 \nof expected ones. ( B) Associations replicated at nominal significance in the EstBB , color -stratified 535 \naccording to whether they meet the discovery genome-wide (GW; p ≤ 7.5 x 10-6; dark green), replication 536 \n(p ≤ 1.5 x 10-3; green), or nominal (p ≤ 0.05; light green) significance threshold. Disease (CKD = chronic 537 \nkidney disease; AKI = acute kidney injury; HTN = hypertension; COPD = chronic obstructive lung 538 \ndisease), cytogenic band and coordinates, best model (MIR = mirror; U = U-shape; DUP = duplication-539 \nonly; DEL = deletion-only), odds ratio (OR), p-value (P) and statistical confidence tier are given for the 540 \nUK Biobank (UKBB) discovery analysis. OR, one -sided p-values, and number of cases among CNV 541 \ncarriers are provided for the EstBB replication. Overlap with SNP-GWAS signals for a related trait (✓ = 542 \nyes; ✗ = no) or a relevant OMIM gene (RCAD = renal cyst and diabetes; KIN = karyomegalic interstitial 543 \nnephritis; PXE = pseudoxanthoma elasticum; PD = Parkinson’s disease ) is indicated. Previous 544 \nassociation with diseases[20] (duplication (DUP) or deletion (DEL) was associated with indicated 545 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 22 \ndisease; no association (✗); some CNVRs were not tested) and continuous traits [23] (disease-relevant 546 \nbiomarkers are specified; other traits (*); no association (✗)) are listed. 547 \n \nEvidence provided by statistical, literature-based, or independent replication help prioritizing 548 \nthe most promising associations for follow-up studies and pinpoint plausible candidate genes. 549 \nWe highlight several examples where deviations by one copy-number are linked to common 550 \ndiseases sharing clinical features with rare Mendelian conditions caused by homozygous 551 \nperturbations of the same genetic region. This argues against a dichotomic view on dominant 552 \nversus recessive modes of inheritanc e and analogously to allelic series,  suggest that 553 \nMendelian and common diseases represent different ends of the phenotypic spectrum caused 554 \nby genetic variation at a given locus.   555 \n  556 \nGlobal characterization of disease-associated CNV regions 557 \nWe sought to identify global characteristics that distinguish disease-associated CNVRs (Table 558 \nS4). Number of p rotein-coding genes embedded in disease -associated CNVRs, hereafter 559 \nreferred to as “disease genes”, ranged from 0 to over 30  and generally correlated with the 560 \nnumber of encompassed probes ( rPearson = 0.50; p = 4.2 x 10 -4; Figure S3A). Exceptions 561 \ninclude single-gene CNVRs overlapping known pathogenic genes captured thanks to high 562 \nprobe coverage  (e.g., BRCA1). While only seven CNVRs (16%) associated with multiple 563 \ndiseases, propensity for pleiotropy depended on CNV length (+0.16 association/disease gene; 564 \np = 1.5 x 10-5). Accordingly, CNVRs containing more than five genes were also more likely to 565 \nassociate with continuous traits (ORFisher = 53.2; p = 8.5 x 10-6) [23]. One CNVR that stood out 566 \nis the 600kb 16p11.2 BP4-5 region (Figure 2B). Originally identified as a major risk factor for 567 \nautism, schizophrenia, developmental delay and intellectual disability , macro-/microcephaly, 568 \nepilepsy, and obesity [57–63], we previously found the region to associate with 26 continuous 569 \ncomplex traits  [23]. Here, we show that 16p11.2 BP4 -5 deletions increase the risk of 12 570 \ndiseases – including both new and previously reported associations across multiple organ 571 \nsystems – as well as the disease burden (+3 diseases/deletion; p = 1.2 x 10 -26), while the 572 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 23 \nregion’s duplication drove increased risk for psychiatric conditions (i.e., bipolar disorder, 573 \nschizophrenia, and depression), in line with previous findings [62]. 574 \n 575 \nNext, we assessed whether disease genes were under stronger evolutionary constraint (i.e., 576 \nless tolerant to mutations) than genes affected by CNVs at the same frequency but not 577 \nassociated with any disease (i.e., “background genes”). Compared to background genes, the 578 \n231 disease genes had more constrained pLI (pWilcoxon = 1.3 x 10-4; Figure S3B) and LOEUF 579 \n(pWilcoxon = 1.9 x 10-7; Figure S3C) scores, suggesting stronger intolerance to LoF mutations. 580 \nSplitting CNVRs depending on whether they have at least one association through either the 581 \nduplication-only or deletion-only model, we evaluated whether embedded disease genes were 582 \nsensitive to having less (i.e, haploinsufficiency; Figure S3D) or more (i.e., triplosensitivity; 583 \nFigure S3E) than two functional copies.  No significant difference in pHaplo score s were 584 \nobserved but genes overlapping regions whose duplication (pWilcoxon = 9.0 x 10-19) and deletion 585 \n(pWilcoxon = 1.0 x 10-23) have been linked to diseases were more likely to be triplosensitive than 586 \nbackground genes . Similar tr ends were observed considering genes overlapping CNVRs 587 \ninvolved uniquely through the duplication -only and deletion -only models and not the other 588 \nCNV type -specific model  (Figure S 3F-G). Overall, our results indicate that a CNVR’s 589 \npathogenicity is determined both by the number and characteristics of affected genes. 590 \n 591 \nNew insights in known disease genes 592 \nTwo out of 12 female BRCA1 deletion carriers were diagnosed with ovarian cancer 593 \n(chr17:41,197,733-41,276,111; ORdel = 284.3; 95%-CI [24.6; 3290.8];  p = 6.1 x 10 -6; Figure 594 \n4A). BRCA1 [MIM: 113705] is a tumor suppressor gene whose LoF represents a major genetic 595 \nrisk factor for the development of hereditary breast and ovarian cancer (HBOC) [MIM: 604370] 596 \n[64]. Exploring the clinical records of the 12 deletion carriers, we found five diagnoses of breast 597 \ncancer (a trait assessed by CNV-GWAS but that did not yield a GW -significant association), 598 \none of endometrial cancer, and one of Fallopian tube cancer, so that eight carriers (67%) had 599 \nreceived a HBOC diagnosis (Figure 4B). Not only was prevalence of HBOC higher among 600 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 24 \nBRCA1 deletion carriers (ORFisher = 31.0; p = 1.1 x 10-6), but disease onset was earlier (HR = 601 \n17.0; p = 1.3 x 10 -15; Figure 4C). Among the four carriers with no HBOC, two had received 602 \ncancer prophylactic surgery, de facto reducing the penetrance of the deletion. Surgeries were 603 \nlikely carried out based on family history of HBOC, which was reported for 6  carriers (50%), 604 \nsuggesting that these deletions are inherited. We did not observe higher prevalence of other 605 \ncancer types (Figure 4B).  606 \n 607 \nHigh abundance of Alu repeats make the low-density lipoprotein (LDL) receptor (LDLR) [MIM: 608 \n606945] susceptible to CNVs  [65]. We found that deletion of  exon 2 -6 increased risk for 609 \nischemic heart disease (IHD; chr19:11,210,904-11,218,188; ORdel = 31.2; 95%-CI [7.1; 137.8]; 610 \np = 5.6 x 10-6), a condition present in 8 of 14 deletion carriers (Figure 4D). Heterozygous - and 611 \nless frequently homozygous – mutations in LDLR represent the main genetic etiology for 612 \nfamilial hypercholesterolemia [66], which is characterized by elevated LDL cholesterol and 613 \npredisposition for adverse cardiovascular outcomes [67]. Previously identified in clinical 614 \nstudies of familial hypercholesterolemia [68], the CNVR implicated by our analysis specifically 615 \nencompasses the ligand-binding domain of LDLR [66]. Confirming widespread prevalence and 616 \nfamily history (43%) of cardiovascular diseases  (Figure 4E), medical records of deletion 617 \ncarriers further revealed higher prevalence (ORFisher = 11.6; p = 7.9 x 10-5) and earlier onset 618 \n(HR = 5.8; p = 1.4 x 10-7; Figure 4F) of pure hypercholesterolemia (E78.0), a code included in 619 \nour lipidemia definition but that did not yield a signal pick -up by the CNV -GWAS. As we 620 \npreviously did not find the CNVR to associate with standardized blood biochemistry LDL levels 621 \n[23], we hypothesized that the latter  were lowered by hypolipidemic agents. Ten (71%)  622 \ndeletion carriers were on statins and six (43%) were additionally using cholesterol absorption 623 \ninhibitors, while the  remaining four did not receive a dyslipidemia  or IHD diagnosis and 624 \nharbored smaller deletions (i.e., P12-14; Figure 4E). We concluded that drugs likely masked 625 \ngenetically determined LDL levels, as shown by higher  LDL levels in the first primary care 626 \nmeasurement on record, measured prior to the standardized LDL measurement (pt-test = 0.03; 627 \nFigure 4G). Despite this, the recommended target of ≤ 1.8 mmol/L for high-risk individuals [69] 628 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 25 \nwas never met. By recovering known gene-disease pairs typically studied in clinical cohorts, 629 \nwe showcase how the rich phenotypic data from biobanks can generate insights into the 630 \nmechanisms, epidemiology, and comorbidities of these diseases, implicating CNVs as 631 \nimportant genetic risk factors. 632 \n \n \n \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 26 \nFigure 4. Refining contribution of CNVs to gene-disease pairs 633 \n(A) Genomic coordinates of the 12 females (P1 -12) carrying a BRCA1 deletion (CNVR delimited by 634 \nvertical dashed lines ), colored according to ovarian cancer diagnosis. ( B) Left: Cancer and related 635 \nfamily/personal diagnoses received by individuals in (A). Color indicates age at diagnosis. Right: Counts 636 \nper ICD-10 code. (C) Kaplan-Meier curve depicting the percentage, with 95% confidence interval, of 637 \nfemales free of female-specific cancers over time among copy -neutral and BRCA1 deletion carriers. 638 \nHazard ratio (HR ) and p -value for the BRCA1 deletion are given (CoxPH model). ( D) Genomic 639 \ncoordinates of the 14 individuals (P1-14) carrying an LDLR deletion (CNVR delimited by vertical dashed 640 \nlines), colored according to ischemic heart disease (IHD) diagnosis. ( E) Left: Medical conditions and 641 \nfamily/personal diagnoses and medication received by ≥ 3  LDLR deletion carriers in (D) , following 642 \nlegend in (B). (F) Kaplan-Meier curve for pure hypercholesterolemia (E78.0) among copy-neutral and 643 \nLDLR deletion carriers, following legend as in (C). (G) Low density lipoprotein (LDL)-cholesterol levels 644 \n(y-axis) from primary care data (first available measurement) and blood biochemistry (average over 645 \ninstances) for six deletion carriers in (D) with at leas t one antecedent primary care LDL -cholesterol 646 \nmeasurement, colored according to IHD diagnosis. P -value compares the two data sources (paired 647 \none-sided t -test). Grey horizontal line represents median LDL -cholesterol value  (from blood 648 \nbiochemistry) in non -carriers. Light and darker green background represent recommended target 649 \nvalues for low (≤ 3 mmol/L) and high (≤ 1.8 mmol/L) risk individuals, respectively. (H) 17q12 association 650 \nlandscape. Top: Negative logarithm of association p -values of CNVs (dark grey; CNVR delimited by 651 \nvertical dashed lines) and SNPs  (orange) [70] with chronic kidney disease (CKD) and SNPs  with 652 \nestimated glomerular filtration rate (eGFR;  red) [71]. Lead SNPs are labeled. Red horizontal dashed 653 \nlines represent the genome-wide threshold for significance for CNV-GWAS (p ≤ 7.5 x 10-6) and SNP-654 \nGWAS (p ≤ 5 x 10 -8). Middle: Genomic coordinates of genes and DECIPHER CNV, with HNF1B, the 655 \nputative causal gene in red. Segmental duplications are represented as a gray gradient proportional to 656 \nthe degree of similarity. Bottom: Genomic coordinates of duplications (blue) and deletions (red) of UK 657 \nBiobank participants overlapping the region. (I) CKD prevalence (± standard error) according to 17q12 658 \ncopy-number (CN). P -values compare deletion (CN = 1) and duplication (CN = 3) carriers to copy -659 \nneutral (CN = 2) individuals (two-sided Fisher test). Number of cases and samples sizes are indicated 660 \n(N = cases/sample size). (J) eGFR levels according to 17q12 CN, shown as boxplots; outliers are not 661 \nshown. P-values comparisons as in (I) (two-sided t-test). Grey horizontal line represents median eGFR 662 \nin non-carriers. Light and darker green background represent mildly decreased (60-90 ml/min/1.73m2) 663 \nand normal (≥ 90 ml/min/1.73m2) kidney function, respectively. (K) Kaplan-Meier curve for CKD among 664 \n17q12 deletion and duplication carriers, following legend as in (C).  665 \n \nBiomarker CNV associations tag pathophysiological processes 666 \nIntegration of biomarker and disease CNV -GWAS signals can identify high -confidence, 667 \nclinically relevant associations. Heterozygous LoF of HNF1B [MIM: 189907] and 17q12 668 \ndeletions cause renal cyst and diabetes (RCAD) [MIM: 137920 ], a severe disorder 669 \ncharacterized by renal abnormalities and maturity-onset diabetes by the young [72,73]. While 670 \nwe previously showed that renal biomarkers were increased in duplication carriers [23], here, 671 \nwe demonstrate that both 17q12 deletions and  duplications increase CKD risk 672 \n(chr17:34,755,219-36,249,489; ORU-shape = 6.5; 95%-CI [3.4; 12.1]; p = 5.9 x 10-9; Figure 4H), 673 \nwith a prevalence of 33.3% (pt-test = 0.026) and 16.9% (pt-test = 6.8 x 10-5) among deletion and 674 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 27 \nduplication carriers respectively, versus 4.4% in copy-neutral individuals (Figure 4I). Results 675 \nreplicated in the EstBB (p = 1.8. x 10-7; Figure 3B) and are supported by 20% of CNV carriers 676 \nshowing signs of kidney disease based on estimated glomerular filtration rate (eGFR < 6 0 677 \nml/min/1.73m2), compared to 2.2% in copy -neutral individuals (Figure 4J). Importantly, both 678 \n17q12 deletion and duplication lower age of CKD onset (HR ≥ 4.6; p ≥ 1.3 x 10 -7; Figure 4K), 679 \nproviding strong evidence of the deleterious consequences on kidney health of altered dosage 680 \nof 17q12. 681 \n 682 \nSimilarly, the blood pressure -increasing 16p12.2 deletion (chr16: 21,946,523-22,440,319) 683 \n[19,23] increased risk for hypertension (OR del = 2.7; 95% -CI [1.9; 3.8]; p = 1.3 x 10 -8) and 684 \ncardiac conduction disorders (ORdel = 3.3; 95%-CI [2.2; 4.9]; p = 1.1 x 10-8), suggesting a role 685 \nin cardiovascular health (Figure S4A-D). Primarily associated with developmental delay and 686 \nintellectual disability [74,75] – proxied by decreased fluid intelligence (pt-test = 8.7 x 10 -5) and 687 \nincome (pt-test = 1.4 x 10-12) in the UKBB (Figure S4E-F) – cardiac malformations are reported 688 \nin ~38% of clinically ascertained cases [76]. Among 193 UKBB deletion carriers, two (1%) had 689 \ncongenital insufficiency of the aortic valve (Q23.1), correspo nding to a higher but not 690 \nsignificantly different prevalence of cardiovascular malformations (Q20 -28) than in copy -691 \nneutral individuals (ORFisher = 2.1; p = 0.251). The deletion also associated with pneumonia 692 \n(ORdel = 3.0; 95%-CI [1.9; 4.6]; p = 5.4 x 10 -7), coinciding with associations with decreased 693 \nforced vital capacity [23] (Figure S4G-H) and peak expiratory flow [19], together demonstrating 694 \nthe clinical relevance of CNV-biomarker associations. 695 \n 696 \nDissecting complex pleiotropic CNV regions 697 \nWhile some CNV signals converge onto the same underlying physiological processes, others 698 \ntie apparently unrelated traits to the same genetic region, suggesting genuine pleiotropy. 699 \n16p13.11 harbors multiple, partially overlapping recurrent groups of CNVs  allowing fine-700 \nmapping of  signals to different subregions of the CNVR  (Figure 5). Through different 701 \nassociation models, the CNVR was linked to uncorrelated traits including epilepsy, kidney 702 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 28 \nstones, hypertension, alkaline phosphatase (ALP), forced vital cap acity, and age at 703 \nmenopause and menarche . We previously proposed MARF1 as a candidate gene for the 704 \nfemale reproductive phenotypes [23] and will focus here on the remaining traits.  705 \n \nFigure 5. Dissection of complex pleiotropic patterns of recurrent CNVs at 16p13.11  706 \n(A) 16p13.11 genetic landscape. Coordinates of  UK Biobank duplications (shades of blue; top) and 707 \ndeletions (shades of red; bottom) overlapping the maximal CNV region (CNVR delimited by vertical 708 \ndashed lines) associated with epilepsy, kidney stones, hypertension, and alkaline phosphatase (ALP). 709 \nCNVs are divided and colored according to five categories (cat1-5) to reflect recurrent breakpoints, with 710 \natypical CNVs in grey. Breakpoints reflect segmental duplications, represented with a grey gradient 711 \nproportional to the degree of similarity. Middle: genomic coordinate s of genes and DECIPHER CNV. 712 \nInset: Overlap between ABCC6’s exonic structure and cat5 deletions. (B, D, F, H) Negative logarithm 713 \nof association p-values of CNVs with ( B) epilepsy; (D) kidney stones; ( F) hypertension; and (H) ALP 714 \n(dark grey; model in parenthesis; CNVR delimited by vertical dashed lines) and SNPs with (B) epilepsy 715 \n[77]; ( D) kidney stones  [78], calcium levels, and phosphate levels (y -axis; break: //); ( F) essential 716 \nhypertension and systolic blood pressure  [79]; and (H) ALP. Lead SNPs are labeled. Red horizontal 717 \ndashed lines represent genome-wide thresholds for significance for CNV -GWAS (p ≤ 7.5 x 10 -6) and 718 \nSNP-GWAS (p ≤ 5 x 10-8). (C, E, G) Prevalence (± standard error) of (C) epilepsy, (E) kidney stones, 719 \nand (G) hypertension according to 16p13.11 copy-number (CN) and CNV categories from (A). P-values 720 \ncompare deletion (CN = 1) and duplication (CN = 3) carriers to copy-neutral (CN = 2) individuals (two-721 \nsided Fisher test). Number of cases and samples sizes are indicated (N = cases/sample size). (I) ALP 722 \nlevels according to 16p13.11 CN and CNV category, shown as boxplots; outliers are not shown. P -723 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 29 \nvalues compare deletion (CN = 1) and duplication (CN = 3) carriers to copy-neutral (CN = 2) individuals 724 \n(two-sided t-test). Grey horizontal line represents median ALP value in non-carriers. 725 \n \n \nThe 654 duplications and 355 deletions overlapping the maximal CNVR (chr16 :15,070,916-726 \n16,353,166) were grouped into 5 categories (cat1-5) based on their breakpoints  (Figure 5A; 727 \nSupplemental Note 3). Risk for epilepsy was increased in deletion carriers (chr16:15,122,801-728 \n16,353,166; ORdel = 6.2; 95%-CI [2.8; 13.4]; p = 4.4 x 10 -6; Figure 5B), with a prevalence of 729 \n8.3% among cat1-4 deletion carriers compared to less than 1.5% among copy -neutral and 730 \nduplication carriers ( Figure 5C). Previously associated with epilepsy in clinical cohorts  731 \n[6,80,81], t he region harbors NDE1 [MIM: 609449], a ge ne associated with autosomal 732 \nrecessive lissencephaly [MIM: 614019] and microhydranencephaly [MIM: 605013] and whose 733 \nmutation has been linked to epilepsy [82,83]. Deletions also increased risk for kidney stones 734 \n(chr16:15,120,501-16,353,166; ORdel = 5.9; 95%-CI [2.9; 11.9]; p = 7.3 x 10-7), with the CNV-735 \nGWAS signal peaking close to a missense variant (rs41278174 G>A; FrequencyA: 2.1%) in 736 \nexon 23 of ABCC6 [MIM: 603234] associating with calcium and phosphate levels through 737 \nSNP-GWAS (Figure 5D). These signals coincide with the recurrent cat5 deletion that covers 738 \n29 probes spanning exons 23-29 from ABCC6 (Figure 5A). Kidney stones prevalence reaches 739 \n4.3% among cat5 deletion carriers, in-between estimates for larger cat1 -4 deletion carriers 740 \n(9.3%) and copy -neutral individuals (2.3%) (Figure 5E). A wide range of variants affecting 741 \nABCC6 have been identified and linked to the calcification disorder pseudoxanthoma 742 \nelasticum through recessive [MIM: 264800] – and more rarely dominant [MIM: 177850] – 743 \ninheritance [84–87], with the Alu-mediated cat5 deletion representing one of the most frequent 744 \nvariants [88,89]. ABCC6 is expressed in the kidney and recent estimates from clinical cohorts 745 \nsuggested that kidney stone s are an unrecognized  (i.e., not used to establish clinical 746 \ndiagnosis) but prevalent (11-40%) feature of pseudoxanthoma elasticum [90–92]. Our data 747 \nsupport kidney stones as a clinical outcome of ABCC6 disruption with partial gene deletions 748 \nleading to lower penetrance than large 16p13.11 deletions. Unlike epilepsy and kidney stones, 749 \nboth deletion (39.2%) and duplication (42%) carriers are at increased risk for  hypertension 750 \n(chr16:15,127,986-16,308,285; ORU-shape = 1.5; 95%-CI [1.3; 1.8]; p = 5.5 x 10-6; Figure 5F), 751 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 30 \ncompared to copy-neutral individuals (35.3%) (Figure 5G). The CNVR overlaps a SNP-GWAS 752 \nsignal for systolic blood pressure mapping to MYH11 [MIM: 160745] (Figure 5F). Expressed 753 \nin arteries, MYH11 encodes for smooth muscle myosin heavy chains and has been linked to 754 \ndominant familial thoracic aortic aneurysm [MIM: 132900], for which hypertension represents 755 \na leading risk factor. Increased prevalence (37.4%) of hypertension among cat5 deletions 756 \nimplicates ABCC6, suggesting cis-epistasis for hypertension risk at 16p13.11. Consistent with 757 \nthis model, ABCC6 plays a role in vascular calcification as the causal gene for generalized 758 \narterial calcification of infancy [MIM: 614473]  [93,94], typically diagnosed by hypertension in 759 \nnewborns. Interestingly, the previously described mirror association with ALP  760 \n(chr16:15,070,916-16,276,964; bmirror = 6.6 U/L; p = 3.5 x 10 -7) peaks at the distal end of the 761 \nCNVR [23], nearby a suggestive SNP-GWAS signal for ALP levels (Figure 5H). Splitting ALP 762 \nlevels by CNV category revealed that th is mirroring effect is driven by individuals with cat2 763 \ndeletion (mean = 76.4 U/L; pt-test = 9.7 x 10-3) and duplication (mean = 92.9 U/L; pt-test = 8.2 x 764 \n10-5), as other CNV carriers had ALP levels indistinguishable from those of copy -neutral 765 \nindividuals (mean = 83.6 U/L) (Figure 5I). Hence, we propose the distal region of the CNVR 766 \nto harbor the critical region regulating ALP levels, even if no obvious candidate gene could be 767 \nidentified through literature review.  768 \n 769 \nThe proximal 22q11.2 region , previously linked  to DiGeorge [MIM: 188400] and 770 \nvelocardiofacial [MIM: 192430] syndromes, harbors four low-copy repeat (LCR; labeled A-to-771 \nD) [95]. Building on evidence of complex association patterns with this CNVR [38], we report 772 \nnovel associations between CNVs spanning LCR A -D and IHD (chr22: 19,024,651-773 \n21,463,545; OR U-shape = 2.1; 95% -CI [1.6; 2.8]; p = 1.5 x 10 -7), LCR B -D and aneurysm 774 \n(chr22:20,708,685-21,460,008; ORdel = 41.8; 95%-CI [10.0; 175.1]; p = 3.2 x 10 -7), and LCR 775 \nA-C and headaches (chr22:19,024,651-21,110,240; ORmirror. = 3.7; 95%-CI [2.1; 6.5]; p = 4.8 776 \nx 10 -6) (Figure S 5A; Supplemental Note 3). Based on  3 LCR B -D deletion carriers  with 777 \naneurysm, this corresponds to a 22-times higher prevalence than in copy-neutral individuals 778 \n(Figure S5B). Association with IHD is better powered, with a prevalence of 12%, 21%, 16%, 779 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 31 \nand 20% among copy -neutral individuals and carriers of LCR C -D, B -D, and A -D CNVs, 780 \nrespectively (Figure S5C). This suggests that IHD risk scales with the amount of affected 781 \ngenetic content, supporting the presence of genetic driver (s) and/or modifier(s) in the C -D 782 \ninterval, beyond the prime candidate TBX1 [MIM: 602054] [95]. Collectively, our data indicates 783 \nthat altered 22q11.2 dosage can result in a spectrum of  cardiovascular afflictions of various 784 \ndegrees of severity, ranging from well-described congenital malformation [95,96] to adult -785 \nonset cardiovascular disorders. 786 \n 787 \n15q13 deletions spanning BP4-5 [MIM: 612001] – and to a lesser extent duplications – have 788 \nbeen associated with neuropsychiatric and developmental conditions [97,98], with the nicotinic 789 \nacetylcholine receptor ion channel CHRNA7 being proposed as the driver gene based on the 790 \npresence of similar phenotypes in individuals with a smaller deletion  (D-CHRNA7-BP5) only 791 \naffecting CHRNA7 [99] (Figure S 6A). While BP4 -5 duplication  carriers showed higher 792 \nprevalence of AKI (EstBB-replicated: Figures 3B, S6B), hemorrhagic stroke 793 \n(chr15:30,912,719-31,982,408; ORU-shape = 7.5; 95%-CI [3.2; 17.9]; p = 4.3 x 10-6; Figure S6C), 794 \nand anemia (chr15:30,912,719-31,094,479; ORdup = 4.9; 95% -CI [2.5; 9.7]; p = 3.2 x 10 -6; 795 \nFigure S6D), reminiscent of associations with pulse rate, mean corpuscular hemoglobin, and 796 \nred blood cell count  [19,23], this was not the case for  the ~10 -times more numerous D-797 \nCHRNA7-BP5 duplication carriers. Replicating an association with a sthma [20] 798 \n(chr15:30,912,719-32,516,949; OR mirror = 0.17; 95% -CI [0.08; 0.35]; p = 1.2 x 10 -6) which 799 \nparallels decreased forced vital capacity [23] and peak expiratory flow [19], this was the only 800 \ndeletion-driven signal for which the CNVR encompassed the  entire BP4-5 region. However, 801 \nonly BP4-5 (46.2%; p t-test = 1.8 x 10 -5) deletion carriers had higher asthma prevalence than 802 \ncopy-neutral individuals (12.1%) ( Figure S 6E). Overall, the non -neurological disorders we 803 \nassociate with 15q13 CNVs appear to specifically involve dosage of the genes within BP4-D-804 \nCHRNA7 and not CHRNA7.  805 \n 806 \nPathological consequences of an increased CNV burden  807 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 32 \nBy assessing the global pathogenic impact of CNVs we can  capture the effect of ultra-rare 808 \nvariants (frequency ≤ 0.01%), as well as those whose effect is not strong enough to reach GW 809 \nsignificance under current settings. Individual-level autosomal CNV (duplication + deletion), 810 \nduplication, and deletion burdens were calculated as the number of Mb or genes affected by 811 \nthe considered type of variation and their predictive value on the same 60 diseases (and the 812 \ndisease burden) previously assessed through CNV -GWAS was estimated ( Figure 6A; left). 813 \nDisease burden strongly associated with a high CNV load (Mb: bCNV = +0.08 disease/Mb; p = 814 \n3.7 x 10-27 | Gene: bdel = +0.03 disease/gene; p = 3.7 x 10-27) and risk for 20 individual disorders 815 \nwas increased by at least one type of CNV burden (p ≤ 0.05/61 = 8.2 x 10-4; Figure 6B; left; 816 \nTable S5). Strongest effect sizes were observed for the Mb burden for psychiatric disorders, 817 \nsuch as bipolar disorder (ORdel = 1.4; p = 6.9 x 10-4), schizophrenia (ORdel = 1.4; p = 4.1 x 10-818 \n5), or epilepsy (ORCNV = 1.1; p = 8.3 x 10 -5), in agreement with CNVs representing important 819 \nrisk factors for these complex and polygenic disorders. To ensure that we do  not merely 820 \ncapture the effect of individual CNV-disease associations previously isolated by CNV-GWAS, 821 \nwe excluded CNVs overlapping disease-associated CNVRs from the burden calculation and 822 \nre-estimated their predictive value (Figure 6A; right). Overall association strength dropped but 823 \nsignal was lost only for type 1 diabetes and chronic obstructive pulmonary disease (Figure 6B; 824 \nright; Table S 5), s uggesting that CNVs not captured by the CNV-GWAS contribute to 825 \nmodulating disease risk. Establishing the polygenic CNV architecture of a substantial number 826 \nof common diseases , this implies that increased power will likely lead to the discovery of 827 \nfurther CNV-disease associations. 828 \n 829 \nAs the “gene burden” captures more associations than the “Mb burden” and as the deletion 830 \nburden yields stronger associations than the duplication one (Figure 6B), we hypothesized 831 \nthat the pathogenic effect of CNVs is mainly driven by the number of deleted genes. Based 832 \non an average GW gene density of 8.4 genes/Mb, we assume the deletion burden measured 833 \nin Mb (𝛽:;) to be 8.4 times larger than the one measured in number of affected genes (𝛽<\"!\"), 834 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 33 \nexpecting a ratio of 8.4 if the type of affected genetic content does not impact disease risk 835 \n(see Methods; Figure 6C). Ratios significantly (p ≤ 0.05/21 = 2.4 x 10-3) smaller or larger than 836 \n8.4 indicate disproportionately large or small contribution of protein coding regions, 837 \nrespectively. We detected a larger than expected contribution of deleted genes to the disease 838 \nburden (ratio = 4.4; p = 1.3 x 10-6), and AKI (ratio = 3.5; p = 7.0. x 10-4) and lipidemia (ratio = 839 \n2.7; p = 1.4 x 10 -3) risk, along with seven additional nominally decreased ratios ( Figure 6D; 840 \nleft; Table S6), confirming that haploinsufficient genes mediate the burden’s effect on disease 841 \nrisk. Repeating the analysis with the duplication burden did not reveal any significant 842 \ndeviations ( Table S 6). After correcting for CNV -GWAS signals, only AKI (p = 7.0 x 10 -4), 843 \narthrosis (p = 0.030), and disease burden (p = 0.016) showed increased contribution of deleted 844 \ngenes on disease risk, indicating that a substantial fraction of effects capture by the CNV -845 \nGWAS are mediated by deletion of coding regions.  846 \n \n \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 34 \nFigure 6. The coding deletion burden increases overall disease risk 847 \n(A) Burden calculation. Left: Raw CNV burden is calculated by summing up the length (in Mb or number 848 \nof affected genes) of all CNVs in an individual. Burden values are used as a predictor for disease risk. 849 \nRight: Corrected CNV burden is calculated by summing up the length of all CNVs that do not overlap 850 \nwith a CNV region (CNVR) significantly associated with the investigated disease through CNV-GWAS. 851 \nCorrected burden values are used to re -estimate contribution of the CNV burden to disease risk (red 852 \ncurve). (B) Contribution of raw and corrected CNV burdens in Mb or affected genes (x-axis) to disease 853 \nrisk (y-axis). Only the most significantly associated burden, providing p ≤ 0.05/61 = 8.2 x 10-4, is shown. 854 \nColor indicates whether the CNV (duplication + deletion), duplication, or deletion burden was most 855 \nsignificantly associated, with size and transparency being proportional to the effect size (beta) and p -856 \nvalue, respectively. Grey horizontal bands mark t raits with no CNV-GWAS signal. (C) Top: Based on 857 \nan average genome-wide gene density of 8.4 genes/Mb, we expect the effect of the CNV burden in Mb 858 \nto be 8.4x larger than the one measured in number of affected genes. Bottom: If the ratio is smaller 859 \nthan 8.4, it indicates a n increased contribution of coding regions. Alternatively, if larger than 8.4, it 860 \nindicates a decreased contribution of coding regions. ( D) Ratio, with 95% confidence interval (CI), of 861 \nthe raw and corrected Mb/genes deletion burdens (x -axis) for all diseases showing at least one 862 \nsignificant burden association. Significant deviation from the expected ratio of 8.4 (grey vertical line) is 863 \nindicated as *** for p ≤ 0.05/21 = 2.4e-3 and * for p ≤ 0.05 (two-sided one sample t-test). 864 \n 865 \nDISCUSSION 866 \nUsing an adapted GWAS framework, we provide a detailed investigation of the contribution of 867 \nCNVs to the genetic architecture of 60 common diseases  and showcase how the rich 868 \nphenotypic data of the UKBB can be leveraged to gain new biological insights, highlighting the 869 \nrole of CNVs as modulators of common disease susceptibility in the general population. 870 \n 871 \nVarious strategies have been used to study CNV-disease associations in the UKBB. Focusing 872 \non diseases related to the ones assessed in the current study, we replicate 10 out of the 24 873 \ndetected associations (at FDR ≤ 0.1)  with 54 likely pathogenic CNVs  [20] and all four 874 \nassociations (at p ≤ 1 x 10-9) in a recent CNV-GWAS investigating 757 diseases [25]. Despite 875 \nanalyzing the same dataset,  we often obtained p-values orders of magnitude smaller  (e.g., 876 \n16p11.2 BP4-5 deletion and AKI: p = 5.6 x 10 -20; pCrawford = 6.0 x 10 -5; pHujoel = 3.3 x 10 -15). 877 \nBesides accruing case count from updated hospital records , careful case-control definition 878 \nand statistical handling of the binary outcomes , probe-level association analysis, and usage 879 \nof different association models to mimic  various dosage mechanisms could explain the 880 \nincreased power of our study. We consequently identified previously unreported CNV-disease 881 \nassociations whose relevance was asserted by follow-up analyses. Only one signal – 17q12 882 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 35 \nCNVs increasing CKD risk – was backed by all approaches, while others were supported by 883 \nsome analysis but not by others – e.g., two associations in the lowest statistical confidence 884 \ntier were replicated in the EstBB and harbored plausible candidate genes – emphasizing the 885 \nimportance of consid ering diverse lines of corroborative evidence. Substantial overlap with 886 \nrelevant SNP-GWAS signals and OMIM genes indicates shared genetic mechanisms  with 887 \nconvergent effects on the phenotype. Another line of evidence is coincidence with a disease-888 \nrelevant biomarker CNV signals, e.g., four (1q21.1-2, 15q13, 16p12.2, 16p11.2 BP4-5) out of 889 \nsix CNVRs decreasing forced vital capacity  [23] were found to increase risk for pulmonary 890 \ndiseases. This demonstrates that biomarkers are efficient proxies underlying (CNV-driven) 891 \npathological processes, often increasing the statistical power to detect associations  due to 892 \ntheir continuous nature. To render binary outcomes continuous, we regressed covariates out 893 \nof the disease status. With the same goal, more sophisticated approaches have recently been 894 \ndeveloped that transform binary outcomes in continuous liability scores while borrowing 895 \ninformation from age-at-disease onset, sex, and familial history [100]. Successfully applied to 896 \nSNP-based GWASs, future exploration is warranted to assess their benefit in the context of 897 \nCNV-GWASs. By coupling a CNV -GWAS framework designed to account for challenges 898 \nlinked to disease CNV association studies in population cohorts  to extensive validation, we 899 \ngenerated a list of 73 CNV-disease pairs with various levels of supporting evidence that can 900 \ninform follow-up studies. 901 \n 902 \nDisease-associated CNVRs harbored genes und er stronger evolutionary constraint than 903 \nthose lacking associations and their length correlated with their propensity for pleiotropy , 904 \nindicating that as previously observed [8], both the number and the nature of genes affected 905 \nby CNVs influence their pathogenicity. Consequently, large multi -gene recurrent CNVs 906 \nexhibited the strongest pleiotropy . A  longstanding question relates to the identification of 907 \ncausal genes whose altered dosage drives the phenotypic alterations observed in carriers. 908 \nModels with various levels of complexity have been proposed, ranging from a single driver 909 \ngene to multiple driver genes modulated by epistatic interactions with other genes in the CNVR 910 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 36 \n[101]. By analyzing disease prevalence in subsets of CNV carriers, association signals could 911 \nbe fine-mapped to narrower regions, pinpointing candidate drivers - such as ABCC6 for kidney 912 \nstones. In other cases, our data suggests that multiple subregions of the CNVR contribute to 913 \nincreased risk for a given disease  by cis-epistasis, as observed for 22q11.2 and ischemic 914 \nheart disease or 16p13.11 and hypertension. Interestingly, the putative driver for phenotypes 915 \noriginally associated with a CNVR might not be driving our newly identified associations, as 916 \nshown for the 15q13 CNVR, whose non-neurological phenotypes do not appear to be linked 917 \nto altered dosage of CHRNA7. Beyond characterizing the pleiotropic pathological 918 \nconsequences of recurrent CNVRs, we demonstrate that dissection of CNV -GWAS signals 919 \ncan fine-map associations and provide mechanistic insights into their phenotypic expression. 920 \n 921 \nAll CNVs increased disease risk and led to an earlier age at onset. Incorporating age at onset 922 \ninformation has been shown to  improve power to detect associations  [100], and more 923 \nimportantly, represents the ultimate proof of clinical relevance. Indeed, many signals mapped 924 \nto regions whose genetic perturbation ha s been reported to be pathogenic. These include 925 \nassociations between well-described, clinically relevant gene-disease pairs – such as BRCA1 926 \nand LDLR deletions increasing the risk for ovarian cancer and IHD, respectively – but for which 927 \nthe role of CNVs in a large population cohort had not been previously investigated. CNVs in 928 \nthese genes have high penetrance but are extremely rare in the UKBB, so that association 929 \nbarely reached GW significanc e. Follow-up analyses based on the medical record s, family 930 \nhistory, medication use, and biomarkers could recapitulate additional clinical associations and 931 \nestablish that these deletions were most likely inherited, thereby generating insights into their 932 \nrole in the general population. We further show that many CNVRs previously linked to pediatric 933 \ngenomic disorders also increased risk for a broad spectrum of adult-onset common diseases. 934 \nThese associations were probably overlooked as the  medical consequences in adulthood of 935 \nthese etiologies is often poorly characterized owing to ascertainment bias and difficulty to 936 \ngather large cohorts . While awaiting validation in clinical cohorts of CNV  carriers, we hope 937 \nthat these findings will improve clinical characterization of genomic disorders, thereby 938 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 37 \nfacilitating diagnosis and allowing physicians to anticipate later-onset comorbidities . For 939 \ninstance, we found carriers of 16p13.11 deletions affect ing ABCC6, the causal gene for 940 \npseudoxanthoma elasticum, to be at increased risk for kidney stones, paralleling reports from 941 \nclinical cohorts showing that kidney stones represent an unrecognized feature of the disease 942 \n[90–92]. Awareness of this disease feature can mitigate kidney stone risk through adapted 943 \ndiet and sufficient water intake. Together, our results advocate for a complex model of variable 944 \nCNV expressivity and penetrance that can result in a broad range of phenotypes along the 945 \nrare-to-common disease spectrum and represent fertile ground for in -depth, phenome-wide 946 \nstudies aiming at better characterizing specific CNV regions [38].  947 \n 948 \nCorroborating the deleterious impact of CNVs on an individual’s health parameters, socio -949 \neconomic status, and lifespan [18,22,23,25,31,102–105], we here speculate that it acts by 950 \nincreasing risk for a broad range of common diseases beyond their known role in  951 \nneuropsychiatric disorders  [4–7]. While both duplication s and deletion s contributed to 952 \nincreased disease risk, the deletion burden’s impact  was much stronger – especially for 953 \nmetabolic, psychiatric, and musculoskeletal diseases  – and for about half the diseases, 954 \npredominantly driven by deletion of coding region, in line  with the commonly accepted view 955 \nthat deletions tend to be more deleterious than duplications and that genetic variation affecting 956 \ncoding regions have stronger phenotypic consequences. Only a marginal fraction of the CNV 957 \nburden’s contribution to disease risk was captured by disease-associated CNVRs and risk for 958 \nfour diseases, i.e., hypothyroidism, cataract, arthrosis, and osteoporosis, associated with the 959 \nburden despite lacking CNV-GWAS signals. With over 10,000 cases, these diseases were not 960 \nunderpowered compared to others, suggesting genuine differences in the genetic architecture, 961 \nand illustrating the added value of burden tests. Collectively, these results predict that better 962 \npowered CNV-GWAS are likely to isolate further associations currently captured by the CNV 963 \nburden.  964 \n 965 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 38 \nA major limitation of our study is the reliance on microarray CNV calls, which allows to assess 966 \nonly a fraction of the CNV landscape, i.e., mostly large CNVs or in regions with high probe 967 \ncoverage. We speculate that small and/or multiallelic CNVs that can only be uncovered by  968 \nsequencing, will have a genetic architecture closer to the one of SNPs and indels, with higher 969 \nfrequencies and more subtle effect sizes, resembling those of SNPs and indels. These effects, 970 \nhowever, are more likely tagged by common variants, limiting novel discoveries. Furthermore, 971 \nby detecting more events, sequencing-based studies require adapted and more stringent 972 \nsignificance thresholds. Still, having improved breakpoint resolution, such CNV calls are also 973 \nlikely to enhance fine-mapping strategies. Microarray CNV calls also exhibit high false positive 974 \nrates [47]. By using stringent CNV selection criteria, we decrease the latter at the cost of 975 \ndecreasing power to detect true associations. This aspect is particularly relevant given that 976 \nthe type of CNVs we assess are rare and that the UKBB is depleted for disease cases  [34], 977 \nresulting in low powered GWASs. While we adopt strategies to counter the lack of power, our 978 \nresults are likely subject to Winner’s curse, only capturing a fraction of the strongest, possibly 979 \noverestimated effects. This phenomenon might be compensated by UKBB CNV carriers being 980 \nat the milder end of the clinical spectrum, leading to effect underestimation. An interesting 981 \nquestion will be to compare effect sizes from population-based studies to those emerging from 982 \nclinical cohort . In the future, longitudinal follow -up of UKBB participants will increas e the  983 \nnumber of cases – especially for late -onset diseases such as Alzheimer’s or Parkinson’ 984 \ndiseases – allowing better powered CNV -GWASs. Alternatively, meta-analyses can boost 985 \ncase number through inclusion of clinical cohorts, at the cost of poorer disease definition due 986 \nto imperfect data harmonization [8]. Larger and more diverse  biobanks linking genotype to 987 \nphenotype data [106–108] should both validate reported associations and identify new ones.  988 \n 989 \nCONCLUSIONS 990 \nOur study provides in -depth analysis of the role of CNVs in modulating susceptibility to 60 991 \ncommon diseases in the general population, broadening our view on how this mutational class 992 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 39 \nimpacts human health. Besides describing clinically relevant and actionable associations, we 993 \nillustrate how complex pleiotropic patterns can be dissected to gain new insights into the 994 \npathological mechanisms of large recurrent CNVs, providing a framework that can be applied 995 \nto an even larger spectrum of diseases.  996 \n 997 \nSUPPLEMENTAL DATA 998 \nSupplemental data include 6 figures, 6 tables, and 3 notes. 999 \n 1000 \nDECLARATIONS 1001 \nAvailability of data and materials : All data used in this study are publicly available, as 1002 \ndescribed in the methods. CNV -GWAS summary statistics (UKBB) will be deposited on the 1003 \nGWAS Catalog upon publication and are available upon request until then.  1004 \n 1005 \nCompeting interests: SO is an employee of MSD at the time of the submission; contribution 1006 \nto the research occurred during the affiliation at the University of Lausanne. 1007 \n 1008 \nFunding: The study was funded by the Swiss National Science Foundation (31003A_182632, 1009 \nAR; 310030_189147, ZK ), Horizon2020 Twinning projects (ePerMed 692145, AR), the 1010 \nEstonian Research Council  (PRG687, MJ and RM), and the Department of Computational 1011 \nBiology (ZK) and the Center for Integrative Genomics (AR) from the University of Lausanne. 1012 \n 1013 \nAuthor’s contributions: CA, AR and ZK conceived the study; CA carried out the analyses 1014 \nwith contributions from MCS, NT, and CJC; The Estonian Biobank Research Team 1015 \ncoordinated genotyping and sequencing data acquisition in the EstBB; MJ performed the 1016 \nreplication study in the EstBB under the supervision of RM;  ZK supervised statistical analyses; 1017 \nSO provided guidance for time -to-event analysis; CA drafted the manuscript and generated 1018 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 40 \nthe figures; AR and ZK made critical revisions; All authors read, approved, and provided 1019 \nfeedback on the final manuscript.   1020 \n 1021 \nAcknowledgments: We thank all biobank participants for sharing their data. UKBB and 1022 \nEstBB computations were performed on the JURA server (University of Lausanne) and the 1023 \nHigh-Performance Computing Center (University of Tartu), respectively. 1024 \n \n \nREFERENCES \n \n1. Sudmant PH, Rausch T, Gardner EJ, Handsaker RE, Abyzov A, Huddleston J, et al. An integrated map of 1025 \nstructural variation in 2,504 human genomes. Nature. 2015;526:75–81.  1026 \n2. Conrad DF, Pinto D, Redon R, Feuk L, Gokcumen O, Zhang Y, et al. Origins and fun ctional impact of copy 1027 \nnumber variation in the human genome. Nature. 2010;464:704–12.  1028 \n3. Zhang F, Gu W, Hurles ME, Lupski JR. Copy Number Variationin Human Health, Disease, and Evolution. Annu 1029 \nRev Genomics Hum Genet. 2009;10:451–81.  1030 \n4. Sebat J, Lakshmi B, Malhotra D, Troge J, Lese-martin C, Walsh T, et al. Strong Association of De Novo Copy 1031 \nNumber Mutations with Autism. Science (1979). 2007;316:445–9.  1032 \n5. Walsh T, McClellan JM, McCarthy SE, Addington AM, Pierce SB, Cooper GM, et al. Rare structural variants 1033 \ndisrupt multiple genes in neurodevelopmental pathways in schizophrenia. Science (1979). 2008;320:539–43.  1034 \n6. Mefford HC, Muhle H, Ostertag P, von Spi czak S, Buysse K, Baker C, et al. Genome -Wide Copy Number 1035 \nVariation in Epilepsy: Novel Susceptibility Loci in Idiopathic Generalized and Focal Epilepsies. PLoS Genet. 1036 \n2010;6:e1000962.  1037 \n7. Cooper GM, Coe BP, Girirajan S, Rosenfeld JA, Vu TH, Baker C, et al. A copy number variation morbidity map 1038 \nof developmental delay. Nat Genet. 2011;43:838–46.  1039 \n8. Collins RL, Glessner JT, Porcu E, Lepamets M, Brandon R, Lauricella C, et al. A cross -disorder dosage 1040 \nsensitivity map of the human genome. Cell. 2022;185:3041-3055.e25.  1041 \n9. Firth H V., Richards SM, Bevan AP, Clayton S, Corpas M, Rajan D, et al. DECIPHER: Database of Chromosomal 1042 \nImbalance and Phenotype in Humans Using Ensembl Resources. Am J Hum Genet. 2009;84:524–33.  1043 \n10. Carvalho CMB, Lupski JR. Mechanisms underlying structural variant formation in genomic disorders. Nat Rev 1044 \nGenet. 2016;17:224–38.  1045 \n11. Collins RL, Brand H, Karczewski KJ, Zhao X, Alföldi J, Francioli LC, et al. A structural variation reference for 1046 \nmedical and population genetics. Nature. 2020;581:444–51.  1047 \n12. Abel HJ, Larson DE, Chiang C, Das I, Kanchi KL, Layer RM, et al. Mapping and characterization of structural 1048 \nvariation in 17,795 deeply sequenced human genomes. Nature. 2020;583:83–9.  1049 \n13. Halvorsen M, Huh R, Oskolkov N, Wen J, Netotea S, Giusti-Rodriguez P, et al. Increased burden of ultra-rare 1050 \nstructural variants localizing to boundaries of topologically associated domains in schizophrenia. Nat Commun. 1051 \n2020;11:1–13.  1052 \n14. Chen L, Abel HJ, Das I, Larson DE, Ganel L, Kanchi KL, et al. Association of structural variation with 1053 \ncardiometabolic traits in Finns. Am J Hum Genet. 2021;108:583–96.  1054 \n15. Beyter D, Ingimundardottir H, Oddsson A, Eggertsson HP, Bjornsson E, Jonsson H , et al. Long -read 1055 \nsequencing of 3,622 Icelanders provides insight into the role of structural variants in human diseases and other 1056 \ntraits. Nat Genet. 2021;53:779–86.  1057 \n16. Babadi M, Fu JM, Lee SK, Smirnov AN, Gauthier LD, Walker M, et al. GATK -gCNV: A Rare Copy Number 1058 \nVariant Discovery Algorithm and Its Application to Exome Sequencing in the UK Biobank. bioRxiv. 1059 \n2022;2022.08.25.504851.  1060 \n17. Fitzgerald T, Birney E. CNest: A novel copy number association discovery method uncovers 862 new 1061 \nassociations from 200 ,629 whole -exome sequence datasets in the UK Biobank. Cell Genomics. 1062 \n2022;2:100167.  1063 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 41 \n18. Kendall KM, Rees E, Escott -Price V, Einon M, Thomas R, Hewitt J, et al. Cognitive Performance Among 1064 \nCarriers of Pathogenic Copy Number Variants: Analysis of 152,000 UK  Biobank Subjects. Biol Psychiatry. 1065 \n2017;82:103–10.  1066 \n19. Owen D, Bracher-Smith M, Kendall KM, Rees E, Einon M, Escott-Price V, et al. Effects of pathogenic CNVs on 1067 \nphysical traits in participants of the UK Biobank. BMC Genomics. 2018;19:1–9.  1068 \n20. Crawford K, Bracher-Smith M, Owen D, Kendall KM, Rees E, Pardiñas AF, et al. Medical consequences of 1069 \npathogenic CNVs in adults: Analysis of the UK Biobank. J Med Genet. 2019;56:131–8.  1070 \n21. Kendall KM, Rees E, Bracher-Smith M, Legge S, Riglin L, Zammit S, et al. Association of Rare Copy Number 1071 \nVariants with Risk of Depression. JAMA Psychiatry. 2019;76:818–25.  1072 \n22. Macé A, Tuke MA, Deelen P, Kristiansson K, Mattsson H, Nõukas M, et al. CNV-association meta-analysis in 1073 \n191,161 European adults reveals new loci associated with anthropometric traits. Nat Commun. 2017;8:1–11.  1074 \n23. Auwerx C, Lepamets M, Sadler MC, Patxot M, Stojanov M, Baud D, et al. The individual and global impact of 1075 \ncopy-number variants on complex human traits. Am J Hum Genet. 2022;109:647–68.  1076 \n24. Aguirre M, Rivas MA, Priest J. Phenome-wide Burden of Copy-Number Variation in the UK Biobank. Am J Hum 1077 \nGenet. 2019;105:373–83.  1078 \n25. Hujoel MLA, Sherman MA, Barton AR, Mukamel RE, Sankaran VG, Terao C, et al. Influences of rare copy -1079 \nnumber variation on human complex traits. Cell. 2022;185:4233-4248.e27.  1080 \n26. Sinnott-Armstrong N, Tanigawa Y, Amar D, Mars N, Benner C, Aguirre M, et al. Genetics of 35 blood and urine 1081 \nbiomarkers in the UK Biobank. Nat Genet. 2021;53:185–94.  1082 \n27. Li YR, Glessner JT, Coe BP, Li J, Mohebnasab M, Chang X, et al. Rare copy number variants in over 100,000 1083 \nEuropean ancestry subjects reveal multiple disease associations. Nat Commun. 2020;11:1–9.  1084 \n28. Kopal J, Kumar K, Saltoun K, Modenato C, Moreau CA, Martin-Brevet S, et al. Rare CNVs and phenome-wide 1085 \nprofiling highlight brain structural divergence and phenotypical convergence. Nat Hum Behav. 2023;  1086 \n29. Bycroft C, Freeman C, Petkova D, Band G, Elliott LT, Sharp K, et al. The UK Biobank resource with deep 1087 \nphenotyping and genomic data. Nature. 2018;562:203–9.  1088 \n30. Wright CF, West B, Tuke M, Jones SE, Patel K, Laver TW, et al. Assessing the Pathogenicity, Penetrance, and 1089 \nExpressivity of Putative Disease-Causing Variants in a Population Setting. Am J Hum Genet. 2019;104:275–1090 \n86.  1091 \n31. Kingdom R, Tuke M, Wood A, Beaumont RN, Frayling TM, Weedon MN, et al. Rare genetic variants in genes 1092 \nand loci linked to dominant monogenic developmental disorders cause milder related phenotypes in the general 1093 \npopulation. Am J Hum Genet. 2022;109:1308–16.  1094 \n32. Chen R, Shi L, Hakenberg J, Naughton B, Sklar P, Zhang J, et al. Analysis of 589,306 genomes identifies 1095 \nindividuals resilient to severe Mendelian childhood diseases. Nat Biotechnol. 2016;34:531–8.  1096 \n33. Goodrich JK, Singer-Berk M, Son R, Sveden A, Wood J, England E, et al. Determinants of penetrance and 1097 \nvariable expressivity in monogenic metabolic conditions across 77,184 exomes. Nat Commun. 2021;12:1–15.  1098 \n34. Fry A, Littlejohns TJ, Sudlow C, Doherty N, Adamska L, Sprosen T, et al. Comparison of Sociodem ographic 1099 \nand Health-Related Characteristics of UK Biobank Participants with Those of the General Population. Am J 1100 \nEpidemiol. 2017;186:1026–34.  1101 \n35. Falconer D. The inheritance of liability to certain diseases, estimated from the incidence among relatives. Ann 1102 \nHum Genet. 1965;29:51–76.  1103 \n36. Senn StephenS. Statistical issues in drug development. Wiley; 2021.  1104 \n37. Mollon J, Schultz LM, Huguet G, Knowles EEM, Mathias SR, Rodrigue A, et al. Impact of Copy Number Variants 1105 \nand Polygenic Risk Scores on Psychopathology in the UK Biobank. Biol Psychiatry. 2023;  1106 \n38. Zamariolli M, Auwerx C, Sadler MC, Graaf A Van Der, Lepik K, Schoeler T, et al. The impact of 22q11 . 2 copy-1107 \nnumber variants on human traits in the general population. Am J Hum Genet. 2023;110:1–14.  1108 \n39. Giannuzzi G, Schmidt PJ, Porcu E, Willemin G, Munson KM, Nuttle X, et al. The Human -Specific BOLA2 1109 \nDuplication Modifies Iron Homeostasis and Anemia Predisposition in Chromosome 16p11.2 Autism Individuals. 1110 \nAm J Hum Genet. 2019;105:947–58.  1111 \n40. Leitsalu L, Haller T, Esko T, Tammesoo ML, Alavere H, Snieder H, et al. Cohort profile: Estonian biobank of 1112 \nthe Estonian genome center, university of Tartu. Int J Epidemiol. 2015;44:1137–47.  1113 \n41. Wu P, Gifford A, Meng X, Li X, Campbell H, Varley T, et al. Mapping ICD -10 and ICD -10-CM Codes to 1114 \nPhecodes: Workflow Development and Initial Evaluation. JMIR Med Inform. 2019;7:e14325.  1115 \n42. Sollis E, Mosaku A, Abid A, Buniello A, Cerezo M, Gil L, et al. The NHGRI-EBI GWAS Catalog: knowledgebase 1116 \nand deposition resource. Nucleic Acids Res. 2023;51:D977–85.  1117 \n43. Hamosh A, Scott AF, Amberger JS, Bocchini CA, McKusick VA. Online Mendelian Inheritance in Man (OMIM), 1118 \na knowledgebase of human genes and genetic disorders. Nucleic Acids Res. 2005;33.  1119 \n44. Karczewski KJ., Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q, et al. The mutational constraint 1120 \nspectrum quantified from variation in 141,456 humans. Nature. 2020;581:434–43.  1121 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 42 \n45. Aguet F, Barbeira A, Bona zzola R, Brown A, Castel S, Jo B, et al. The GTEx Consortium atlas of genetic 1122 \nregulatory effects across human tissues. Science (1979). 2020;369:1318–30.  1123 \n46. Wang K, Li M, Hadley D, Liu R, Glessner J, Grant SFA, et al. PennCNV: An integrated hidden Markov model 1124 \ndesigned for high-resolution copy number variation detection in whole-genome SNP genotyping data. Genome 1125 \nRes. 2007;17:1665–74.  1126 \n47. Macé A, Tuke MA, Beckmann JS, Lin L, Jacquemont S, Weedon MN, et al. New quality measure for SNP array 1127 \nbased CNV detection. Bioinformatics. 2016;32:3298–305.  1128 \n48. Chang CC, Chow CC, Tellier LCAM, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: Rising to the 1129 \nchallenge of larger and richer datasets. Gigascience. 2015;4:1–16.  1130 \n49. Wang K, Li M, Hakonarson H. ANNOVA R: functional annotation of genetic variants from high -throughput 1131 \nsequencing data. Nucleic Acids Res. 2010;38:e164–e164.  1132 \n50. Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, et al. The human genome browser at 1133 \nUCSC. Genome Res. 2002;12:996–1006.  1134 \n51. Gao X, Starmer J, Martin ER. A multiple testing correction method for genetic association studies using 1135 \ncorrelated single nucleotide polymorphisms. Genet Epidemiol. 2008;32:361–9.  1136 \n52. Therneau TM. Survival Analysis [R package survival version 3.5-3]. 2022 [cited 2023 Apr 17]; Available from: 1137 \nhttps://CRAN.R-project.org/package=survival 1138 \n53. Berman JJ. Rare Diseases and Orphan Drugs: Keys to Understanding and Treating the Common Diseases. 1139 \nElsevier Inc.; 2014.  1140 \n54. Li A, Jiao X, Munier FL, Schorderet DF, Yao W, Iwata F, et al. Bietti Crystalline Corneoretinal Dystrophy Is 1141 \nCaused by Mutations in the Novel Gene CYP4V2. Am J Hum Genet. 2004;74:817–26.  1142 \n55. Zhou W, Otto EA, Cluckey A, Airik R, Hurd TW, Chaki M, et al. FAN1 mutations cause karyomegalic interstitial 1143 \nnephritis, linking chronic kidney failure to defective DNA damage repair. Nat Genet. 2012;44:910–5.  1144 \n56. Kryger MH, Roth T, Goldstein CA. Principles and Practice of Sleep Medicine. Elsevier Health Sciences; 2021.  1145 \n57. Shinawi M, Liu P, Kang SHL, Shen J, Belmont JW, Scott DA, et al. Recurrent reciprocal 16p11.2 1146 \nrearrangements associated with global developmental delay, behavioural problems, dysmorphism, epilepsy, 1147 \nand abnormal head size. J Med Genet. 2010;47:332–41.  1148 \n58. Weiss LA, Shen Y, Korn JM, Arking DE, Miller DT, Fossdal R, et al. Association between Microdeletion and 1149 \nMicroduplication at 16p11.2 and Autism. N Engl J Med. 2008;358:667–75.  1150 \n59. D’Angelo D, Lebon S, Chen Q, Martin -Brevet S, Snyder LAG, Hippolyte L, et al. Defining the Effect of the 1151 \n16p11.2 Duplication on Cognition, Behavior, and Medical Comorbidities. JAMA Psychiatry. 2016;73:20–30.  1152 \n60. Reinthaler EM, Lal D, Lebon S, Hildebra nd MS, Dahl HHM, Regan BM, et al. 16p11.2 600 kb Duplications 1153 \nconfer risk for typical and atypical Rolandic epilepsy. Hum Mol Genet. 2014;23:6069–80.  1154 \n61. Jacquemont S, Reymond A, Zufferey F, Harewood L, Walters RG, Kutalik Z, et al. Mirror extreme BMI 1155 \nphenotypes associated with gene dosage at the chromosome 16p11.2 locus. Nature. 2011;478:97–102.  1156 \n62. McCarthy SE, Makarov V, Kirov G, Addington AM, McClellan J, Yoon S, et al. Microduplications of 16p11.2 are 1157 \nassociated with schizophrenia. Nat Genet. 2009;41:1223–7.  1158 \n63. Walters RG, Jacquemont S, Valsesia A, De Smith AJ, Martinet D, Andersson J, et al. A new highly penetrant 1159 \nform of obesity due to deletions on chromosome 16p11.2. Nature. 2010;463:671–5.  1160 \n64. Miki Y, Swensen J, Shattuck-Eidens D, Futreal PA, Harshman K, Tavtigian S, et al. A strong candidate for the 1161 \nbreast and ovarian cancer susceptibility gene BRCA1. Science (1979). 1994;266:66–71.  1162 \n65. Iacocca MA, Hegele RA. Role of DNA copy number variation in dyslipidemias. Curr Opin Lipidol. 2018;29:125–1163 \n32.  1164 \n66. Hobbs HH, Russell DW, Brown MS, Goldstein JL. The LDL receptor locus in familial hypercholesterolemia: 1165 \nmutational analysis of a membrane protein. Annu Rev Genet. 1990;24:133–70.  1166 \n67. Defesche JC, Gidding SS, Harada -Shiba M, Hegele RA, Santos RD, Wie rzbicki AS. Familial 1167 \nhypercholesterolaemia. Nat Rev Dis Primers. 2017;3:17093.  1168 \n68. Iacocca MA, Wang J, Dron JS, Robinson JF, McIntyre AD, Cao H, et al. Use of next-generation sequencing to 1169 \ndetect LDLR gene copy number variation in familial hypercholesterolemia. J Lipid Res. 2017;58:2202–9.  1170 \n69. Mach F, Baigent C, Catapano AL, Koskina s KC, Casula M, Badimon L, et al. 2019 ESC/EAS Guidelines for 1171 \nthe management of dyslipidaemias: lipid modification to reduce cardiovascular riskThe Task Force for the 1172 \nmanagement of dyslipidaemias of the European Society of Cardiology (ESC) and European Ath erosclerosis 1173 \nSociety (EAS). Eur Heart J. 2020;41:111–88.  1174 \n70. Wuttke M, Li Y, Li M, Sieber KB, Feitosa MF, Gorski M, et al. A catalog of genetic loci associated with kidney 1175 \nfunction from analyses of a million individuals. Nat Genet. 2019;51:957–72.  1176 \n71. Stanzick KJ, Li Y, Schlosser P, Gorski M, Wuttke M, Thomas LF, et al. Discovery and prioritization of variants 1177 \nand genes for kidney function in >1.2 million individuals. Nat Commun. 2021;12:1–17.  1178 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 43 \n72. Fajans SS, Bell GI, Polonsky KS. Molecular mechanisms and  clinical pathophysiology of maturity -onset 1179 \ndiabetes of the young. N Engl J Med. 2001;345:97180.  1180 \n73. Mefford HC, Clauin S, Sharp AJ, Moller RS, Ullmann R, Kapur R, et al. Recurrent Reciprocal Genomic 1181 \nRearrangements of 17q12 Are Associated with Renal Disea se, Diabetes, and Epilepsy. Am J Hum Genet. 1182 \n2007;81:1057–69.  1183 \n74. Girirajan S, Rosenfeld JA, Cooper GM, Antonacci F, Siswara P, Itsara A, et al. A recurrent 16p12.1 1184 \nmicrodeletion supports a two-hit model for severe developmental delay. Nat Genet. 2010;42:203–9.  1185 \n75. Stefansson H, Meyer -Lindenberg A, Steinberg S, Magnusdottir B, Morgen K, Arnarsdottir S, et al. CNVs 1186 \nconferring risk of autism or schizophrenia affect cognition in controls. Nature. 2013;505:361–6.  1187 \n76. Girirajan S, Pizzo L, Moeschler J, Rosenfeld J. 16p12.2 Recurrent Deletion. GeneReviews® [Internet]. 2018 1188 \n[cited 2023 Apr 4]; Available from: https://www.ncbi.nlm.nih.gov/books/NBK274565/ 1189 \n77. Stevelink R, Campbell C, Chen S, The International League Against Epilepsy Consortium on Complex 1190 \nEpilepsies. Genome-wide meta-analysis of over 29,000 people with epilepsy reveals 26 loci and subtype -1191 \nspecific genetic architecture. medRxiv. 2022;2022.06.08.22276120.  1192 \n78. Howles SA, Wiberg A, Goldsworthy M, Bayliss AL, Gluck AK, Ng M, et al. Genetic variants of calcium and 1193 \nvitamin D metabolism in kidney stone disease. Nat Commun. 2019;10:1–10.  1194 \n79. Evangelou E, Warren HR, Mosen-Ansorena D, Mifsud B, Pazoki R, Gao H, et al. Genetic analysis of over 1 1195 \nmillion people identifies 535 new loci associated with blood pressure traits. Nat Genet. 2018;50:1412–25.  1196 \n80. De Kovel CGF, Trucks H, Helbig I, Mefford HC, Baker C, Leu C, et al. Recurrent microdeletions at 15q11.2 and 1197 \n16p13.11 predispose to idiopathic generalized epilepsies. Brain. 2010;133:23–32.  1198 \n81. Heinzen EL, Radtke RA, Urban TJ, Cavalleri GL, Depondt C, Need AC, et al. Rare Deletions at 16p13.11 1199 \nPredispose to a Diverse Spectrum of Sporadic Epilepsy Syndromes. Am J Hum Genet. 2010;86:707–18.  1200 \n82. Alkuraya FS, Cai X, Emery C, Mochida GH, Al-Dosari MS, Felie JM, et al. Human Mutations in NDE1 Cause 1201 \nExtreme Microcephaly with Lissencephaly. Am J Hum Genet. 2011;88:536–47.  1202 \n83. Bakircioglu M, Carvalho OP, Khurshid M, Cox JJ, Tuysuz B, Barak T, et al. The Essential Role of Centrosomal 1203 \nNDE1 in Human Cerebral Cortex Neurogenesis. Am J Hum Genet. 2011;88:523–35.  1204 \n84. Ringpfeil F, Lebwohl MG, Christiano AM, Uitto J. Pseudoxanthoma elasticum: Mutations in the MRP6 gene  1205 \nencoding a transmembrane ATP-binding cassette (ABC) transporter. Proc Natl Acad Sci U S A. 2000;97:6001–1206 \n6.  1207 \n85. Struk B, Cai L, Zäch S, Ji W, Chung J, Lumsden A, et al. Mutations of the gene encoding the transmembrane 1208 \ntransporter protein ABC-C6 cause pseudoxanthoma elasticum. J Mol Med. 2000;78:282–6.  1209 \n86. Le Saux O, Urban Z, Tschuch C, Csiszar K, Bacchelli B, Quaglino D, et al. Mutations in a gene encoding an 1210 \nABC transporter cause pseudoxanthoma elasticum. Nat Genet. 2000;25:223–7.  1211 \n87. Bergen AAB, Plomp AS, Schuurman EJ, Terry S, Breuning M, Dauwerse H, et al. Mutations in ABCC6 cause 1212 \npseudoxanthoma elasticum. Nat Genet. 2000;25:228–31.  1213 \n88. Le Saux O, Beck K, Sachsinger C, Silvestri C, Treiber C, Goöring HHH, et al. A Spectrum of ABCC6 Mutations 1214 \nIs Responsible for Pseudoxanthoma Elasticum. Am J Hum Genet. 2001;69:749–64.  1215 \n89. Ringpfeil F, Nakano A, Uitto J, Pulkkinen L. Compound Heterozygosity for a Recurrent 16.5-kb Alu-Mediated 1216 \nDeletion Mutation and Single -Base-Pair Substitutions in the ABCC6 Gene Results in Pseudoxanthoma 1217 \nElasticum. Am J Hum Genet. 2001;68:642–52.  1218 \n90. Ralph D, Allawh R, Terry IF, Terry SF, Uitto J, Li QL. Kidney Stones Are Prevalent in Individuals with 1219 \nPseudoxanthoma Elasticum, a Genetic Ectopic Mineralization Disorder. Int J Dermatol Venereol. 2020;3:198–1220 \n204.  1221 \n91. Legrand A, Cornez L, Samkari W, Mazzella JM, Venisse A, Boccio V, et al. Mutation spectrum in the ABCC6 1222 \ngene and genotype–phenotype correlations in a French cohort with pseudoxanthoma elasticum. Genetics in 1223 \nMedicine. 2017;19:909–17.  1224 \n92. Letavernier E, Kauffenstein G, Hugue t L, Navasiolava N, Bouderlique E, Tang E, et al. ABCC6 deficiency 1225 \npromotes development of randall plaque. Journal of the American Society of Nephrology. 2018;29:2337–47.  1226 \n93. Nitschke Y, Baujat G, Botschen U, Wittkampf T, Du Moulin M, Stella J, et al. Generalized Arterial Calcification 1227 \nof Infancy and Pseudoxanthoma Elasticum Can Be Caused by Mutations in Either ENPP1 or ABCC6. Am J 1228 \nHum Genet. 2012;90:25–39.  1229 \n94. Le Boulanger G, Labrèze C, Croué A, Schurgers LJ, Chassaing N, Wittkampf T, et al. An unusual s evere 1230 \nvascular case of pseudoxanthoma elasticum presenting as generalized arterial calcification of infancy. Am J 1231 \nMed Genet A. 2010;152A:118–23.  1232 \n95. McDonald-McGinn DM, Sullivan KE, Marino B, Philip N, Swillen A, Vorstman JAS, et al. 22q11.2 deletion 1233 \nsyndrome. Nat Rev Dis Primers. 2015;1:1–19.  1234 \n96. Bartik LE, Hughes SS, Tracy M, Feldt MM, Zhang L, Arganbright J, et al. 22q11.2 duplications: Expanding the 1235 \nclinical presentation. Am J Med Genet A. 2022;188:779–87.  1236 \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint \n\nAuwerx et al., 2023 44 \n97. Sharp AJ, Mefford HC, Li K, Baker C, Skinner C, Stevenson RE, et al. A recurrent 15q13.3 microdeletion 1237 \nsyndrome associated with mental retardation and seizures. Nat Genet. 2008;40:322–8.  1238 \n98. Lowther C, Costain G, Stavropoulos DJ, Melvin R, Silversides CK, Andrade DM, et al. Delineating the 15q13.3 1239 \nmicrodeletion phenotype: a case series and comprehensive review of the literature. Genetics in Medicine. 1240 \n2015;17:149–57.  1241 \n99. Gillentine MA, S chaaf CP. The Human Clinical Phenotypes of Altered CHRNA7 Copy Number. Biochem 1242 \nPharmacol. 2015;97:352.  1243 \n100. Pedersen EM, Agerbo E, Plana-Ripoll O, Grove J, Dreier JW, Musliner KL, et al. Accounting for age of onset 1244 \nand family history improves power in genome-wide association studies. Am J Hum Genet. 2022;109:417–32.  1245 \n101. Golzio C, Katsanis N. Genetic architecture of reciprocal CNVs. Curr Opin Genet Dev. 2013;23:240–8.  1246 \n102. Männik K, Mägi R, Macé A, Cole B, Guyatt AL, Shihab HA, et al. Copy number variati ons and cognitive 1247 \nphenotypes in unselected populations. JAMA. 2015;313:2044–54.  1248 \n103. Wheeler E, Huang N, Bochukova EG, Keogh JM, Lindsay S, Garg S, et al. Genome -wide SNP and CNV 1249 \nanalysis identifies common and low-frequency variants associated with severe early-onset obesity. Nat Genet. 1250 \n2013;45:513–7.  1251 \n104. Dauber A, Yu Y, Turchin MC, Chiang CW, Meng YA, Demerath EW, et al. Genome-wide association of copy-1252 \nnumber variation reveals an association between short stature and the presence of low -frequency genomic 1253 \ndeletions. Am J Hum Genet. 2011;89:751–9.  1254 \n105. Saarentaus EC, Havulinna AS, Mars N, Ahola-Olli A, Kiiskinen TTJ, Partanen J, et al. Polygenic burden has 1255 \nbroader impact on health, cognition, and socioeconomic outcomes than most rare and high-risk copy number 1256 \nvariants. Mol Psychiatry. 2021;26:4884–95.  1257 \n106. Denny JC, Rutter JL, Goldstein DB, Philippakis A, Smoller JW, Jenkins G, et al. The ‘All of Us’ Research 1258 \nProgram. N Engl J Med. 2019;381:668–76.  1259 \n107. Hunter-Zinck H, Shi Y, Li M, Gorman BR, Ji SG, Sun N, et al. Genotyping Array Design and Data Qualit y 1260 \nControl in the Million Veteran Program. Am J Hum Genet. 2020;106:535–48.  1261 \n108. Kurki MI, Karjalainen J, Palta P, Sipilä TP, Kristiansson K, Donner KM, et al. FinnGen provides genetic insights 1262 \nfrom a well-phenotyped isolated population. Nature. 2023;613:508–18.  1263 \n  \n \n . CC-BY 4.0 International licenseIt is made available under a \nperpetuity. \n is the author/funder, who has granted medRxiv a license to display the preprint in(which was not certified by peer review)preprint \nThe copyright holder for thisthis version posted August 5, 2023. ; https://doi.org/10.1101/2023.07.31.23293408doi: medRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}