Genome-wide associations for birth weight and correlations with adult disease.

Horikoshi M, Robin N. Beaumont, Felix R. Day, Warrington NM, Kooijman MN, Fernandez-Tajes J, Feenstra B, Van Zuydam NR, Gaulton KJ, Grarup N, Bradfield JP, Strachan DP, Li-Gao R, Ahluwalia TS, Kreiner E, Rueedi R, Lyytikäinen LP, Cousminer DL, Wu Y, Thiering E, Wang CA, Have CT, Jouke-Jan Hottenga, Vilor-Tejedor N, Joshi PK, Boh ETH, Ntalla I, Pitkänen N, Anubha Mahajan, van Leeuwen EM, Joro R, Lagou V, Nodzenski M, Diver LA, Krina T. Zondervan, Bustamante M, Marques-Vidal P, Mercader JM, Bennett AJ, Nilüfer Rahmioğlu, Dale R. Nyholt, Ma RCW, Tam CHT, Tam WH, Ganesh SK, van Rooij FJ, Samuel E. Jones, Loh PR, Ruth KS, Tuke MA, Jessica Tyrrell, Andrew R. Wood, Hanieh Yaghootkar, Yaghootkar H, Scholtens DM, Paternoster L, Inga Prokopenko, Kovacs P, Atalay M, Willems SM, Panoutsopoulou K, Wang X, Carstensen L, Geller F, Schraut KE, Murcia M, van Beijsterveldt CE, Gonneke Willemsen, Appel EVR, Fonvig CE, Trier C, Tiesler CM, Standl M, Kutalik Z, Bonas-Guarch S, Hougaard DM, Sánchez F, Torrents D, Waage J, Hollegaard MV, de Haan HG, Rosendaal FR, Carolina Medina-Gomez, Ring SM, Gibran Hemani, McMahon G, Robertson NR, Groves CJ, Claudia Langenberg, Luan J, Scott RA, Zhao JH, Mentch FD, MacKenzie SM, Reynolds RM, Lowe WL Jr, Tönjes A, Stumvoll M, Lindi V, Lakka TA, van Duijn CM, Kiess W, Körner A, Sørensen TI, Niinikoski H, Pahkala K, Raitakari OT, Zeggini E, Dedoussis GV, Teo YY, Saw SM, Melbye M, Campbell H, Wilson JF, Vrijheid M, de Geus EJC, Boomsma DI, Kadarmideen HN, Holm JC, Sylvain Sebért, Sebert S, Hattersley AT, Beilin LJ, Newnham JP, Pennell CE, Heinrich J, Adair LS, Borja JB, Mohlke KL, Eriksson JG, Widen E, Kähönen M, Viikari JS, Lehtimäki T, Vollenweider P, Bønnelykke K, Bisgaard H, Mook-Kanamori DO, Hofman A, Fernando Rivadeneira, André G Uitterlinden, Pisinger C, Pedersen O, Power C, Elina Hyppӧnen, Wareham NJ, Håkon Håkonarson, Davies E, Walker BR, Jaddoe VW, Jarvelin MR, Grant SFA, Vaag A, Debbie A. Lawlor, George Davey Smith, Timothy M. Frayling, Davey Smith G, Andrew P. Morris, Ong KK, Felix JF, Nicholas J Timpson, Perry JR, Evans DM, Mark I. McCarthy, Rachel M. Freathy, Freathy RM
OA: closed

Abstract

Birth weight (BW) has been shown to be influenced by both fetal and maternal factors and in observational studies is reproducibly associated with future risk of adult metabolic diseases including type 2 diabetes (T2D) and cardiovascular disease. These life-course associations have often been attributed to the impact of an adverse early life environment. Here, we performed a multi-ancestry genome-wide association study (GWAS) meta-analysis of BW in 153,781 individuals, identifying 60 loci where fetal genotype was associated with BW (P < 5 × 10-8). Overall, approximately 15% of variance in BW was captured by assays of fetal genetic variation. Using genetic association alone, we found strong inverse genetic correlations between BW and systolic blood pressure (Rg = -0.22, P = 5.5 × 10-13), T2D (Rg = -0.27, P = 1.1 × 10-6) and coronary artery disease (Rg = -0.30, P = 6.5 × 10-9). In addition, using large -cohort datasets, we demonstrated that genetic factors were the major contributor to the negative covariance between BW and future cardiometabolic risk. Pathway analyses indicated that the protein products of genes within BW-associated regions were enriched for diverse processes including insulin signalling, glucose homeostasis, glycogen biosynthesis and chromatin remodelling. There was also enrichment of associations with BW in known imprinted regions (P = 1.9 × 10-4). We demonstrate that life-course associations between early growth phenotypes and adult cardiometabolic disease are in part the result of shared genetic effects and identify some of the pathways through which these causal genetic effects are mediated.
Full text 40,272 characters · extracted from pmc-nxml · 3 sections · click to expand

Methods

All human research was approved by the relevant institutional review boards and conducted according to the Declaration of Helsinki. All participants provided written informed consent. Ethical approval for the study was obtained from the ALSPAC Ethics and Law Committee and the Local Research Ethics Committees. Within each study, BW was collected from a variety of sources, including measurements at birth by medical practitioners, obstetric records, medical registers, interviews with the mother and self-report as adults ( Supplementary Table 1 ). BW was z -score transformed, separately in males and females. Individuals with extreme BW (>5 SD from the sex-specific study mean), monozygotic or polyzygotic siblings, or preterm births (gestational age <37 weeks, where this information was available) were excluded from downstream association analyses ( Supplementary Table 1 ). Within each study, stringent quality control of the GWAS genotype scaffold was undertaken, prior to imputation ( Supplementary Table 2 ). Each scaffold was then pre-phased and imputed 30 , 31 up to reference panels from the 1000G 2 or 1000G and UK10K Project 3 ( Supplementary Table 2 ). Association of BW with each variant passing established GWAS quality control filters 32 was tested in a linear regression framework, under an additive model for the allelic effect, after adjustment for study-specific covariates, including gestational age, where available ( Supplementary Table 2 ). Where necessary, population structure was accounted for by adjustment for axes of genetic variation from principal components analysis 33 and subsequent genomic control correction 34 , or inclusion of a genetic relationship matrix in a mixed model 35 ( Supplementary Table 2 ). We calculated the genomic control inflation factor (λ) in each study to confirm that study-level population structure was accounted for prior to meta-analysis. UK Biobank phenotype data were available for 502,655 participants 36 . All participants in the UK Biobank were asked to recall their BW, of which 279,971 did so at either the baseline or follow-up assessment visit. Of these, 7,686 participants reported being part of multiple births and were excluded from downstream analyses. Ancestry checks, based on self-reported ancestry, resulted in the exclusion of 8,998 additional participants reported not to be white European. Of those individuals reporting BW at baseline and follow-up assessments, 393 were excluded because the two reported values differed by more than 0.5 kg. For those reporting different values (≤0.5 kg) between baseline and follow-up, we took the baseline measure forward for downstream analyses. We then excluded 36,716 individuals reporting values 4.5 kg as implausible for live term births before 1970. In total 226,178 participants had data relating to BW that matched these inclusion criteria. Genotype data from the May 2015 release were available for a subset of 152,249 participants from UK Biobank. In addition to the quality control metrics performed centrally by UK Biobank, we defined a subset of “white European” ancestry samples using a K-means (K=4) clustering approach based on the first four genetically determined principal components. A maximum of 67,786 individuals (40,425 females and 27,361 males) with genotype and valid BW measures were available for downstream analyses. We tested for association with BW, assuming an additive allelic effect, in a linear mixed model implemented in BOLT-LMM 37 to account for cryptic population structure and relatedness. Genotyping array was included as a binary covariate in all models. Total chip heritability (i.e. the variance explained by all autosomal polymorphic genotyped SNPs passing quality control) was calculated using Restricted Maximum Likelihood (REML) implemented in BOLT-LMM 37 . We additionally analysed the association between BW and directly genotyped SNPs on the X chromosome: for this analysis, we used 57,715 unrelated individuals with BW available and identified by UK Biobank as white British. We excluded SNPs with evidence of deviation from Hardy-Weinberg Equilibrium ( P <1x10 -6 ), MAF0.015, resulting in 19,423 SNPs for analysis in Plink v1.07 ( http://pngu.mgh.harvard.edu/purcell/plink/ ) 38 , with the first five ancestry principal components as covariates. In both the full UK Biobank sample and our refined sample, we observed that BW was associated with sex, year of birth and maternal smoking ( P <0.0015, all in the expected directions), confirming more comprehensive previous validation of self-reported BW 4 . We additionally verified that BW associations with lead SNPs at seven established loci 5 based on self-report in UK Biobank were consistent with those previously published. The European ancestry meta-analysis consisted of two components: (i) 75,891 individuals from 30 GWAS from Europe, USA and Australia; and (ii) 67,786 individuals of white European origin from UK Biobank. In the first component, we combined sex-specific BW association summary statistics across studies in a fixed-effects meta-analysis, implemented in GWAMA 39 and applied a second round of genomic control 34 (λ GC = 1.001). Subsequently, we combined association summary statistics from this component with UK Biobank in a European ancestry fixed-effects meta-analysis, implemented in GWAMA 39 . Variants failing GWAS quality control filters in UK Biobank, reported in less than 50% of the total sample size in the first component, or with MAF<0.1%, were excluded from the European ancestry meta-analysis. We aggregated X-Chromosome association summary statistics from UK Biobank (19,423 SNPs) with corresponding statistics from the European GWAS component using fixed effects P -value based meta-analysis in METAL 40 (max N=99,152). We were concerned that self-reported BW as adults in UK Biobank would not be comparable with that obtained from more stringent collection methods used in other European ancestry GWAS. In addition, UK Biobank lacked information on gestational age for adjustment, which could have an impact on difference in strength of association compared to the results obtained from other European ancestry GWAS. However, we observed no evidence of heterogeneity in BW allelic effects at lead SNPs between the two components of European ancestry meta-analysis, using Cochran’s Q statistic 41 , implemented in GWAMA 39 , after Bonferroni correction ( P >0.00083) ( Supplementary Table 3 ). We tested for heterogeneity in allelic effects between studies within the European component using Cochran’s Q . At loci demonstrating evidence of heterogeneity, we confirmed that association signals were not being driven by outlying studies by visual inspection of forest plots. We performed sensitivity analyses to assess the impact of covariate adjustment (gestational age and population structure) on heterogeneity. We were also concerned that overlap of individuals (duplicated or related) between the two components of the European ancestry meta-analysis might lead to false positive association signals. We performed bivariate LD Score regression 8 using the two components of the European ancestry meta-analysis and observed a genetic covariance intercept of 0.0156 (SE 0.0058), indicating a maximum of 1,119 duplicate individuals. Univariate LD Score regression 8 of the European ancestry meta-analysis estimated the intercept as 1.0426, which may indicate population structure or relatedness that is not adequately accounted for in the analysis. To assess the impact of this inflation on the European ancestry meta-analysis, we expanded the standard errors of BW allelic effect size estimates and re-calculated association P -values. On the basis of this adjusted analysis, the lead SNP only at MTNR1B dropped below genome-wide significance (rs10830963, P =5.5x10 -8 ). The trans-ancestry meta-analysis combined the two European ancestry components with an additional 10,104 individuals from six GWAS from diverse ancestry groups: African American, Chinese, Filipino, Surinamese, Turkish and Moroccan. Within each GWAS, we first combined sex-specific BW association summary statistics in a fixed-effects meta-analysis, implemented in GWAMA 39 and applied a second round of genomic control 34 . Subsequently, we combined association summary statistics from the six non-European GWAS and the two European ancestry components in a trans-ancestry fixed-effects meta-analysis, implemented in GWAMA 39 . Variants failing GWAS quality control filters in UK Biobank, reported in less than 50% of the total sample size in the first component, or with MAF<0.1%, were excluded from the trans-ancestry meta-analysis. We tested for heterogeneity in allelic effects between ancestries using Cochran’s Q 41 . We searched for multiple distinct BW association signals in each of the established and novel loci, defined as 1Mb up- and down-stream of the lead SNP from the trans-ancestry meta-analysis, through approximate conditional analysis. We applied GCTA 42 to identify “index SNPs” for distinct association signals attaining genome-wide significance ( P <5x10 -8 ) in the European ancestry meta-analysis using a reference sample of 5,000 individuals of white British origin, randomly selected from UK Biobank, to approximate patterns of linkage disequilibrium (LD) between variants in these regions. Note that we performed approximate conditioning on the basis of only the European ancestry meta-analysis because GCTA cannot accommodate LD variation between diverse populations. We combined a number of approaches to prioritise the most likely candidate gene(s) in each BW locus. Expression quantitative trait loci (eQTLs) were obtained from the Genotype Tissue Expression (GTEx) Project 43 , the GEUVADIS Project 44 and eleven other studies 45 – 55 using HaploReg v4 56 . We interrogated coding variants for each BW lead SNP and its proxies (EUR r 2 >0.8) using Ensembl 57 and HaploReg. Their likely functional consequences were predicted by SIFT 58 and PolyPhen2 59 . Biological candidacy was assessed by presence in significantly enriched gene set pathways from MAGENTA analyses (see below for details). We extracted all genes within 300 kb of all lead BW SNPs and searched for connectivity between any genes using STRING 60 . If two or more genes between two separate BW loci were connected, they were given an increased prior for both being plausible candidates. We also applied protein-protein interaction (PPI) analysis (see below for details) to all genes within 300 kb of each lead BW SNPs and ranked the genes based on the score for connectivity with the surrounding genes. At the YKT6-GCK locus, the lead SNP (rs138715366) is of low-frequency in European ancestry populations (MAF=0.92%) and even rarer in other ancestry groups (MAF=0.23% in African Americans, otherwise monomorphic) and is not present in the HapMap reference panel 61 . To assess the accuracy of imputation for this low-frequent variant, we genotyped rs138715366 in the Northern Finland Birth Cohort (NFBC) 1966 ( Supplementary Table 1 ). Of the 5,009 samples in the study, 4,704 were successfully imputed and genotyped (or sequenced) for rs138715366. The overall concordance rate between imputed and directly assayed genotypes was 99.8% and for directly assayed heterozygote calls was 75.0%. We sought to leverage LD differences between populations contributing to the trans-ancestry meta-analysis and to take advantage of the improved coverage of common and low-frequency variation offered by 1000G or 1000G and UK10K combined imputation to localise variants driving each distinct association signal achieving locus-wide significance. For each distinct signal, we used MANTRA 62 to construct 99% credible sets of variants 63 that together account for 99% of the posterior probability of driving the association. MANTRA incorporates a prior model of relatedness between studies, based on mean pair-wise allele frequency differences across loci, to account for heterogeneity in allelic effects ( Supplementary Table 3 ). MANTRA has been demonstrated, by simulation, to improve localisation of causal variants compared with either a fixed- or random-effects trans-ancestry meta-analysis 62 , 64 . For loci with only one signal of association, we used MANTRA to combine summary statistics from the six non-European GWAS and the two European ancestry components. However, for loci with multiple distinct association signals, we used MANTRA to combine summary statistics from approximate conditioning for the two European components, separately for each signal. For each distinct signal, we calculated the posterior probability that the j th variant, π C j , is driving the association, given by π C j = Λ j ∑ k Λ k , where the summation is over all variants mapping within the (conditional) meta-analysis across the locus. In this expression, Λ j is the Bayes’ factor (BF) in favour of association from the MANTRA analysis. A 99% credible set 63 was then constructed by: (i) ranking all variants according to their BF, Λ j ; and (ii) including ranked variants until their cumulative posterior probability exceeds 0.99. We used genomic annotations of DNaseI hypersensitive sites (DHS) from the ENCODE 65 project and protein coding genes from GENCODE 66 . We filtered cell types that are cancer cell lines (karyotype ‘cancer’ from https://genome.ucsc.edu/ENCODE/cellTypes.html ), and merged data from multiple samples from the same cell type. This resulted in 128 DHS cell-type annotations, as well as 4 gene-based annotations (coding exon, 5UTR, 3UTR and 1kb upstream of TSS). First, we tested for the effect of each cell type DHS and gene annotation individually using the Bayes’ factors for all variants in the 62 credible sets using fgwas 67 . Second, we categorised the annotations into ‘genic’, ‘foetal DHS’, ‘embryonic DHS’, ‘stem cell DHS’, ‘neonatal DHS’ and ‘adult DHS’ based on the description fields from ENCODE, and tested for the effect of each category individually as described above using fgwas. Third, we then tested the effect of each category by including all categories in a joint model using fgwas. For each of the three analyses, we obtained the estimated effects and 95% confidence intervals (CI) for each annotation, and considered an annotation enriched if the 95% CI did not overlap zero. Variance explained was calculated using the REML method implemented in GCTA 68 . We considered the variance explained by two sets of SNPs: (i) lead SNPs of all 62 distinct association signals at the 59 established and novel autosomal BW loci identified in the European-specific or trans-ancestry meta-analyses; (ii) lead SNPs of 55 distinct association signals at the 52 novel autosomal BW loci ( Extended Data Table 1a and Supplementary Table 7 ). Variance explained was calculated in samples of European ancestry in the Hyperglycemia and Adverse Pregnancy Outcome (HAPO) study 69 (independent of the meta-analysis) and two studies that were part of the European ancestry meta-analysis: NFBC1966 and Generation R ( Supplementary Table 1 ). In each study, the genetic relationship matrix was estimated for each set of SNPs and was tested individually against BW (males and females combined) with study specific covariates. These analyses provided an estimate and standard error for the variance explained by each of the given sets of SNPs. We performed four sets of analyses. First, we used GWAS data from 4,382 mother-child pairs in the Avon Longitudinal Study of Parents and Children (ALSPAC) study to fit a “maternal-GCTA model” 6 to estimate the extent to which the maternal genome might influence offspring BW independent of the foetal genome. The m-GCTA model uses genome-wide genetic similarity between mothers and offspring to partition the phenotypic variance in BW into components due to the maternal genotype, the child’s genotype, the covariance between the two and environmental sources of variation. Second, we compared associations with BW of the foetal versus maternal genotype at each of the 60 BW loci. The maternal allelic effect on offspring BW was obtained from a maternal GWAS meta-analysis of 68,254 European mothers from the EGG Consortium (n=19,626) 7 and the UK Biobank (n=48,628). In the UK Biobank, mothers were asked to report the BW of their first child. Women of European ancestry with genotype data available in the May 2015 data release were included, and those with reported BW equivalent to 4.5 kg were excluded. No information on gestational age or gender of child was available. BW of first child was associated with maternal factors such as smoking status, BMI and height in the expected directions. Of the 68,254 women included in the maternal GWAS, 13% were mothers of individuals included in the current foetal European ancestry GWAS, and a further approximately 45% were themselves (with their own BW) included in the foetal GWAS. Third, we additionally conducted analyses in 12,909 mother-child pairs from nine contributing studies: at each of the 60 loci, we compared the effect of the foetal genotype on BW adjusted for sex and gestational age, with and without adjustment for maternal genotype. We reciprocally compared the association between the maternal genotype and BW with and without adjustment for foetal genotype. Fourth, we used the method of Zhang et al 15 to test associations between BW and the maternal untransmitted, maternal transmitted and inferred paternal transmitted haplotype score of 422 height SNPs 25 , 30 SBP SNPs 13 , 14 and 84 T2D SNPs 24 in 5,201 mother-child pairs from the ALSPAC study. The use of LD Score regression to estimate the genetic correlation between two traits/diseases has been described in detail elsewhere 70 . Briefly, “LD Score” is a measure of how much genetic variation each variant tags; if a variant has a high LD Score then it is in high LD with many nearby polymorphisms. Variants with high LD Scores are more likely to contain more true signals and hence provide more chance of overlap with genuine signals between GWAS. The LD score regression method uses summary statistics from the GWAS meta-analysis of BW and the other traits of interest, calculates the cross-product of test statistics at each SNP, and then regresses the cross-product on the LD Score. Bulik-Sullivan et al 70 show that the slope of the regression is a function of the genetic covariance between traits: E ( z 1 j z 2 j ) = N 1 N 2 ρ g M l j + ρ N s N 1 N 2 where N i is the sample size for study i , ρ g is the genetic covariance, M is the number of SNPs in the reference panel with MAF between 5% and 50%, l j is the LD score for SNP j , N s quantifies the number of individuals that overlap both studies, and ρ is the phenotypic correlation amongst the Ns overlapping samples. Thus, if there is sample overlap (or cryptic relatedness between samples), it will only affect the intercept from the regression (i.e. the term ρ N s N 1 N 2 ) and not the slope, and hence estimates of the genetic covariance will not be biased by sample overlap. Likewise, population stratification will affect the intercept but will have minimal impact on the slope (i.e. intuitively since population stratification does not correlate with linkage disequilibrium between nearby markers). Summary statistics from the GWAS meta-analysis for traits and diseases of interest were downloaded from the relevant consortium website. The summary statistics files were reformatted for LD Score regression analysis using the munge_sumstats.py python script provided on the developer’s website ( https://github.com/bulik/ldsc ). For each trait, we filtered the summary statistics to the subset of HapMap 3 SNPs 71 , as advised by the developers, to ensure that no bias was introduced due to poor imputation quality. Summary statistics from the European-specific BW meta-analysis were used because of the variable LD structure between ancestry groups. Where the sample size for each SNP was included in the results file this was flagged using --N-col; if no sample size was available then the maximum sample size reported in the reference for the GWAS meta-analysis was used. SNPs were excluded for the following reasons: MAF<0.01; ambiguous strand; duplicate rsID; non-autosomal SNPs; reported sample size less than 60% of the total available. Once all files were reformatted, we used the ldsc.py python script, also on the developers’ website, to calculate the genetic correlation between BW and each of the traits and diseases. The European LD Score files that were calculated from the 1000G reference panel and provided by the developers were used for the analysis. Where multiple GWAS meta-analyses had been conducted on the same phenotype (i.e. over a period of years), the genetic correlation with BW was estimated using each set of summary statistics and presented in Supplementary Table 12 . The phenotypes with multiple GWAS included height, BMI, waist-hip ratio (adjusted for BMI), total cholesterol, triglycerides, high density lipoprotein (HDL) and low density lipoprotein (LDL). The estimate of the genetic correlation between the multiple GWAS meta-analyses on the same phenotype were comparable and the later GWAS had a smaller standard error due to the increased sample size, so only the genetic correlation between BW and the most recent meta-analyses were presented in Fig. 2 . In the published GWAS for BP 14 the phenotype was adjusted for BMI. Caution is needed when interpreting the genetic correlation between BW and BMI-adjusted SBP due to the potential for collider bias 72 . Since BMI is associated with both BP and BW, it is possible that the use of a BP genetic score adjusted for BMI might bias the genetic correlation estimate towards a more negative value. To verify that the inverse genetic correlation with BW (r g =-0.26, SE=0.05, P =6.5x10 -9 ) was not due to collider bias caused by the BMI adjustment of the phenotype, we obtained an alternative estimate using UK Biobank GWAS data for SBP that was unadjusted for BMI and obtained a similar result (r g =-0.22, SE=0.03, P =5.5x10 -13 ). The SBP phenotype in UK Biobank was prepared as follows. Two BP readings were taken at assessment, approximately 5 minutes apart. We included all individuals with an automated BP reading (taken using an automated Omron BP monitor). Two valid measurements were available for most participants (averaged to create a BP variable, or alternatively a single reading was used if only one was available). Individuals were excluded if the two readings differed by more than 4.56 SD. BP measurements more than 4.56 SD away from the mean were excluded. We accounted for BP medication use by adding 15 mmHg to the SBP measure. BP was adjusted for age, sex and centre location and then inverse rank normalised. We performed the GWAS on 127,698 individuals of British descent using BOLT-LMM 37 , with genotyping array as covariate. We estimated the phenotypic, genetic and residual correlations as well as the genetic and residual covariance between BW and several quantitative traits/disease outcomes in UK Biobank using directly genotyped SNPs and the REML method implemented in BOLT-LMM 37 . The traits examined included T2D, SBP, diastolic BP, CAD, height, BMI, weight, waist-hip ratio, hip circumference, waist circumference, obesity, overweight, age at menarche, asthma, and smoking. Where phenotypes were not available (e.g. serum blood measures are not currently available in UK Biobank), we obtained estimates using the NFBC1966 study (for correlations/covariance between BW and triglycerides, total cholesterol, HDL, LDL, fasting glucose and fasting insulin). In the UK Biobank analysis, we used 57,715 unrelated individuals with BW available and identified by UK Biobank as white British. SNPs with evidence of deviation from Hardy-Weinberg Equilibrium ( P <1x10 -6 ), MAF0.015 were excluded, resulting in 328,928 SNPs for analysis. We included the first five ancestry principal components as covariates. In the NFBC1966 analysis, 5,009 individuals with BW were enrolled. Genotyped SNPs that passed quality control ( Supplementary Table 2 ) were included, resulting in 324,895 SNPs for analysis. The first three ancestry principal components and sex were included as covariates. Meta-Analysis Gene-set Enrichment of variaNT Associations (MAGENTA) was used to explore pathway-based associations using summary statistics from the trans-ancestry meta-analysis. MAGENTA implements a gene set enrichment analysis (GSEA) based approach, as previously described 9 . Briefly, each gene in the genome is mapped to a single index SNP with the lowest P -value within a 110 kb upstream and 40 kb downstream window. This P -value, representing a gene score, is then corrected for confounding factors such as gene size, SNP density and LD-related properties in a regression model. Genes within the HLA-region were excluded from analysis due to difficulties in accounting for gene density and LD patterns. Each mapped gene in the genome is then ranked by its adjusted gene score. At a given significance threshold (95th and 75th percentiles of all gene scores), the observed number of gene scores in a given pathway, with a ranked score above the specified threshold percentile, is calculated. This observed statistic is then compared to 1,000,000 randomly permuted pathways of identical size. This generates an empirical GSEA P -value for each pathway. Significance was attained when an individual pathway reached a false discovery rate (FDR) <0.05 in either analysis. In total, 3,216 pre-defined biological pathways from Gene Ontology, PANTHER, KEGG and Ingenuity were tested for enrichment of multiple modest associations with BW. The MAGENTA software was also used for enrichment testing of custom gene sets. We used the integrative Protein-Interaction-Network-Based Pathway Analysis (iPINBPA) method 73 . Briefly, we generated gene-wise P -values from the trans-ancestry meta-analysis using VEGAS2 74 , which map the SNPs to genes and account for possible cofounders, such as LD between markers. The empirical gene-wise P -values are calculated using simulations from the multivariate normal distribution. Those that were nominally significant ( P ≤0.01) were selected as “seed genes”, and were collated within high confidence version of inweb3 75 , to weight the nodes in the network following a guilt-by-association approach. In a second step, a network score was defined by the combination of the z -scores derived from the gene-wise P -values with node weights using the Liptak-Stouffer method 76 . A heuristic algorithm was then applied to extensively search for modules enriched in genes with low P -values. The modules were further normalised using a null distribution of 10,000 random networks. Only those modules with z -score >5 were selected. Finally, the union of all modules constructed a BW-overall PPI network. Both the proteins on the individual modules and on the overall BW-PPI were interrogated for enrichment in Gene Ontology Terms (Biological Processes) using a Hypergeometric test. Terms were considered as significant when adjusted P -value, following Benjamini-Hochberg procedure, was below 0.05. The same methodology described above was applied to 16 different adult traits resulting in a number of enriched modules per trait. Different modules for each trait were combined in a single component and the intersection between these trait-specific components and the BW component was calculated. This intersection is defined as the PoC network. We used the resulting PoC networks in downstream analyses to interrogate which set of proteins connects BW variation and adult trait variation via pathways enriched in the overall BW analysis. We first searched for evidence of parent of origin effects in the UK Biobank samples by comparing variance between heterozygotes and homozygotes using Quicktest 77 . In this analysis, we used only unrelated individuals identified genetically as of white British origin (n=57,715). Principal components were generated using these individuals and the first five were used to adjust for population structure as covariates in the analysis, in addition to a binary indicator for genotyping array. We also examined 4,908 mother-child pairs in ALSPAC and determined the parental origin of the alleles where possible 78 . Briefly, the method uses mother-child pairs to determine the parent of origin of each allele. For example, if the mother/child genotypes are AA/Aa, the child’s maternal/paternal allele combination is A/a. For the situation where both mother and child are heterozygous, the child’s maternal/paternal alleles cannot be directly specified. However, the parental origin of the alleles can be determined by phasing the genotype data and comparing maternal and child haplotypes. We then tested these alleles for association with BW adjusting for sex and gestational age. Statistical power in these currently available sample sizes is insufficient to rule out widespread parent-of-origin effects across the regions tested. Using the mean beta (0.034 SD) and MAF (0.28) of the identified loci, we estimate that we would need at least 200,000 unrelated individuals or 70,000 mother-child pairs for 80% power to detect parent-of-origin effects at P<0.00085. To explore the different patterns of association between BW and other anthropometric/metabolic/endocrine traits and diseases, we performed hierarchical clustering analysis. The lead SNP (or proxy, EUR r 2 >0.6) at the 60 BW loci was queried in publicly available GWAS meta-analysis datasets or in GWAS result obtained through collaboration 79 . Results were available for 53 of those loci and the extracted z -score (allelic effect/SE, Supplementary Table 17 ) was aligned to the BW-raising allele. We performed two dimensional clustering by trait and by locus. We computed the Euclidean distance amongst z -scores of the extracted traits/loci and performed complete hierarchical clustering implemented in the pvclust package ( http://www.sigmath.es.osaka-u.ac.jp/shimo-lab/prog/pvclust/ ) in R v3.2.0 ( http://www.R-project.org/ ). Clustering uncertainty was measured by multiscale bootstrap resampling estimated from 1,000 replicates. We used α=0.05 to define distinct clusters and, based on the bootstrap analysis, calculated the Calinski index to identify the number of well-supported clusters (cascadeKM function, Vegan package, http://CRAN.R-project.org/package=vegan ). Clustering was visualised by constructing dendrograms and a heatmap. Separately from the hierarchical clustering analysis, we queried the lead SNP at EPAS1 in a GWAS of haematological traits 80 because variation at that locus has previously been implicated in BW and adaptation to hypoxia at high altitudes in Tibetans 81 , 82 ( Supplementary Table 17 ).

Extended Data

a, Manhattan (main panel) and QQ (top right) plots of genome-wide association results for BW from trans-ancestry meta-analysis of up to 153,781 individuals. The association P- value (on -log 10 scale) for each of up to 22,434,434 SNPs ( y axis) is plotted against the genomic position (NCBI Build 37; x axis). Association signals that reached genome-wide significance ( P <5x10 -8 ) are shown in green if novel and pink if previously reported. In the QQ plot, the black dots represent observed P -values and the grey line represents expected P- values under the null distribution. The red dots represent observed P -values after excluding the previously identified signals 5 . b, Manhattan (main panel) and QQ (top right) plots of trans-ethnic GWAS meta-analysis for BW highlighting the reported imprinted regions described in Supplementary Table 14 . Novel association signals that reached genome-wide significance ( P <5x10 -8 ) and mapped to imprinted regions are shown in green. Genomic regions outside imprinted regions are shaded in grey. SNPs in the imprinted regions are shown in light blue or dark blue, depending on chromosome number (odd or even). In the QQ plot, the black dots represent observed P values and the grey lines represent expected P- values and their 95% confidence intervals under the null distribution for the SNPs within the imprinted regions. Regional plots for each locus are displayed from: the unconditional European-specific meta-analysis of up to 143,677 individuals (left); the approximate conditional meta-analysis for the primary signal after adjustment for the index variant for the secondary signal (middle); and the approximate conditional meta-analysis for the secondary signal after adjustment for the index variant for the primary signal (right). Directly genotyped or imputed SNPs are plotted with their association P- values (on a -log 10 scale) as a function of genomic position (NCBI Build 37). Estimated recombination rates (blue lines) are plotted to reflect the local LD structure around the index SNPs and their correlated proxies. SNPs are coloured in reference to LD with the particular index SNP according to a blue to red scale from r 2 = 0 to 1, based on pairwise r 2 values estimated from a reference of 5,000 individuals of white British origin, randomly selected from the UK Biobank. For each BW locus, the following six effect sizes (with 95% CI) are shown, all aligned to the same BW-raising allele: foetal_GWAS = foetal allelic effect on BW (from European ancestry meta-analysis of up to n=143,677 individuals); foetal_unadjusted = foetal allelic effect on BW (unconditioned in n=12,909 mother-child pairs); foetal_adjusted = foetal effect (conditioned on maternal genotype, n=12,909); maternal_GWAS = maternal allelic effect on offspring BW (from meta-analysis of up to n=68,254 European mothers) 7 ; maternal_unadjusted = maternal allelic effect on offspring BW (unconditioned, n=12,909); maternal_adjusted = maternal effect (conditioned on foetal genotype, n=12,909). The 60 BW loci are ordered by chromosome and position ( Supplementary Tables 10, 11 ). These plots illustrate that in large GWAS of BW, foetal effect size estimates are larger than those of maternal at 55/60 identified loci (binomial P =1x10 -11 ), suggesting that most of the associations are driven by the foetal genotype. In conditional analyses that modelled the effects of both maternal and foetal genotypes (n=12,909 mother-child pairs), confidence intervals around the estimates were wide, precluding inference about the likely contribution of maternal vs. foetal genotype at individual loci. a, Continued from Extended Data Figure 4 . b, The scatterplot illustrates the difference between the foetal ( x axis) and maternal ( y axis) effect sizes in the overall maternal vs. foetal GWAS results. a, Illustrates the largest global component of birth weight (BW) PPI network containing 13 modules. b, The histogram shows the null distribution of z-scores of BW PPI networks based on 10,000 random networks, and where the z-scores for the 13 BW modules (M1-13) lie. For each module, the two most significant GO terms are depicted. c, Illustrates a heatmap which takes the top 50 biological processes over-represented in the global BW PPI network (listed at the right of the plot), and displays extent of enrichment for the various trait-specific “point of contact“ (PoC) PPI networks. d-e, Trait-specific PoC PPI networks composed of proteins that are shared in both the global BW PPI network and networks generated using the same pipeline for each of the adult traits: d, canonical Wnt signalling pathway enriched for PoC PPI between BW and blood pressure (BP)-related phenotypes; and e, regulation of insulin secretion pathway enriched for PoC between BW and type 2 diabetes (T2D)/fasting glucose (FG). Red nodes are those that are present in PoC for BW and traits of interest; blue nodes correspond to the pathway nodes; purple nodes are those present in both the pathway and PoC; orange nodes are genes in BW loci that overlap with both the pathway and PoC. Large nodes correspond to genes in BW loci (within 300kb from the lead SNP), and have black border if they, amongst all BW loci, have a stronger (top 5) association with at least one of the pairing adult traits. a, QQ plot from the Quicktest 77 analysis comparing the BW variance of heterozygotes with homozygotes in 57,715 UK Biobank samples. b, QQ plot from the parent-of-origin specific analysis testing the association between BW and maternally transmitted vs. paternally transmitted alleles in 4,908 mother-child pairs from the ALSPAC study ( Methods , Supplementary Tables 15, 16 ). In both panels, the black dots represent lead SNPs at 59 identified autosomal BW loci and a further sub-genome-wide significant signal for BW near DLK1 (rs6575803; P =5.6x10 -8 ). The grey lines represent expected P values and their 95% confidence intervals under the null distribution for the 60 SNPs. Both results show trends in favour of imprinting effects at BW loci: however, despite the large sample size, these analyses were underpowered (see Methods ) and much larger sample sizes are required for definitive analysis. a-d, Effect sizes (left y axis) of previously reported 30 SBP loci 13 , 14 , 45 CAD loci 23 , 84 T2D loci 24 and 422 adult height loci 25 are plotted against effects on BW ( x axis). Effect sizes are aligned to the adult trait-raising allele. The colour of each dot indicates BW association P value: red, P <5×10 −8 ; orange, 5×10 −8 ≤ P <0.001; yellow, 0.001≤ P <0.01; white, P ≥0.01. The superimposed grey frequency histogram shows the number of SNPs (right y axis) in each category of BW effect size. e, Effect sizes (with 95% CI) on BW of 45 known CAD loci are plotted arranged in the order of CAD effect size from highest to lowest, separating out the known SBP loci. CAD loci with a larger effect on BW concentrated amongst loci with primary BP association. f, Effect sizes (with 95% CI) on BW of 32 known T2D loci are plotted, subdivided by previously reported categories derived from detailed adult physiological data 27 . Heterogeneity in BW effect sizes between five T2D loci groups with different mechanistic categories was substantial ( P het =1.2x10 -9 ). In pairwise comparisons, the “beta cell” group of variants differed from the other four groups: fasting hyperglycaemia ( P het =3x10 -11 ), insulin resistance ( P het =0.002), proinsulin ( P het =0.78) and unclassified ( P het =0.02) groups. All of the BW effect sizes plotted in the forest plots are aligned to the trait (or risk)-raising allele. a, Effects (beta values) are aligned to the BW-raising allele. EAF was obtained from the trans-ancestry meta-analysis, except for PLAC1 , for which the EAF was obtained from the European ancestry meta-analysis due to lack of X chromosome data from the non-European studies. Chr., chromosome; bp, base pair; EAF, effect allele frequency; SE, standard error. b, The effect of the lead SNP (absolute value of beta, y axis) is given as a function of minor allele frequency ( x axis) for 60 known (pink) and novel (green) BW loci from the trans-ancestry meta-analysis. Error bars are proportional to the standard error of the effect size. The dashed line indicates 80% power to detect association at genome-wide significance level for the sample size in trans-ancestry meta-analysis. Two complementary analyses of the overall GWAS summary data identified enrichment of BW associations in biological pathways related to metabolism, growth and development. a, The top results (FDR<0.05 at the 95 th percentile enrichment threshold) from a total of 3,216 biological pathways tested for enrichment of multiple modest associations with BW. Additionally, results are presented for custom sets of imprinted genes. b, The results of a complementary analysis of empirical PPI data, displaying the top 10 most significant pathways enriched for BW-association scores. P -value is adjusted for multiple correction using Benjamini and Hochberg method.

Supplementary Material

Supplementary Information is linked to the online version of the paper.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-08-11T06:11:44.160905+00:00