Methods
We used FinnGen 50 GWAS summary statistics from release 6, covering 2861 disease endpoints with a total sample size of 260,405 individuals (Supplementary Data 1 ). Medication-related phenotypes, dental endpoints, injury endpoints, pregnancy and delivery were not included in the analysis. FinnGen individuals were genotyped with Illumina and Affymetrix chip arrays. The SNP array data were imputed to a Finnish SISu reference panel, which consists of 3775 whole genomes. The genomic positions of variants are represented by the hg38 assembly.
The GWAS summary statistics of UK Biobank were acquired from Neale Lab ( http://www.nealelab.is/uk-biobank ), round 2 version. The sample QC and variant QC criteria are described in their documentation ( https://github.com/Nealelab/UK_Biobank_GWAS ).
We additionally downloaded summary statistics of binary traits carried out by using SAIGE 51 . Some of these GWASs traits overlap with those from FinnGen. We used these as supplement traits to confirm the discovery of antagonistic variants. Briefly, GWASs were conducted on UKB White British participants. All individuals were imputed to the Haplotype Reference Consortium (HRC) panel. The imputed genotype data were downloaded from the UK Biobank release 2. The first four principal components, sex and age were included as covariates in the association analysis.
Summary statistics of psychiatric disorders were downloaded from PGC , including attention-deficit/hyperactivity disorder (ADHD), bipolar disorder (BIP), anorexia nervosa (AN), cannabis use disorder (CUD), autism spectrum disorder (ASD), major depression (MDD) and schizophrenia (SCZ).
We also considered other large-scale GWA Studies, with summary statistics available in the GWAS catalogue. Studies conducted on European individuals were included.
The full list of GWAS summary statistics used in our study is summarised in Supplementary Data 1 .
We first defined sentinel SNPs for each trait. Different LD reference panels were used for these summary statistics according to the samples used for conducting GWAS. For FinnGen, the LD estimation was acquired from SISu v3 reference panel as described in FinnGen documentation ( https://finngen.gitbook.io/documentation/methods/genotype-imputation/sisu-reference-panel ). For UKB studies, we randomly selected 20,000 unrelated individuals from UKB as the LD reference panel. The variants were matched to each study. Specifically, for GWASs carried out by SAIGE, the imputed UK Biobank release 2 data were used as the LD reference panel. For GWASs acquired from the Neale lab, the UKB release 3 was used as the LD reference panel. For other GWASs, 1000 Genomes European data were used as the LD reference panel, which consists of 503 individuals. For each trait, we defined the sentinel SNPs with plink1.9 52 (--clump-kb 500 –clump-p1 1e−7 –clump-r 2 0.1).
We simply defined an antagonistic variant as having a discordant effect on disease traits if the variant increases the risk of one disease but decreases the risk of another disease. For some quantitative traits, we allocated direction according to their relationship with fitness 22 . A reduction in fitness was considered a disease state when defining antagonistic variants with these traits.
As the terminology “antagonistic pleiotropy” specifically refers to the theory of genetic conflict at different points in the life cycle, we use “pleiotropic conflict” or “antagonistic variants” interchangeably in this study to refer to the wider phenomenon. The genomic positions represented in hg19 were converted to the hg38 assembly. We took a 2 Mb window around each sentinel SNP, then physically overlapped regions across all traits using BEDTools 53 (Fig. 1 ). If the genomic region overlaps across traits, then colocalization analysis was performed using HyPrColoc 54 , which allows for a large number of traits to be analysed simultaneously. We extracted a 2 Mb genomic region around each sentinel SNP in each trait (1 Mb on each side of a sentinel SNP). The overlapped SNPs in a trait set were used in colocalization analysis. Several scenarios were observed: 1) low LD region, high posterior probability and regional probability from colocalization analysis. 2) high LD region, high regional probability, but the shared causal SNP is not distinguishable due to strong LD. 3) different association peaks within the overlapping region, and low regional probability (not colocalized) (Fig. 1 ). Here, we considered the first two scenarios as sharing the associated region. In summary, if the genomic region meets these criteria: 1) it is discordantly associated with different traits, 2) regional probability > 0.8 and posterior probability > 0.5 in colocalization analysis, we consider it as a pleiotropic conflict locus.
To aid interpretation, quantitative traits were excluded from the genetic correlation analysis. We first identified antagonistic loci involving disease traits, and for each antagonistic locus, we randomly selected a pair of diseases. Genetic correlations between these disease pairs were then estimated using GWAS summary statistics of diseases and LD score regression 10 .
We took the antagonistic variants identified in the European ancestry study and sought replication in Biobank Japan. We first identified trait pairs that can be potentially replicated (at least one trait in each direction available in Biobank Japan phenotypes). This results in 62 trait pairs with at least one trait in each direction matched to phenotypes available in Biobank Japan. A significance of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$8.06\times {10}^{-4}$$\end{document} 8.06 × 10 − 4 (0.05/62) was used for a successful replication.
To compare the allele frequency distributions between pleiotropic variants with concordant effects and antagonistic pleiotropic variants, we included all the putative causal SNPs in pleiotropic regions identified from colocalization analysis. The allele frequencies were calculated using 1KG phase 3 EUR data. We only kept SNPs that are in the 1000 G EUR panel. To ensure the variants included in the comparison are independent of each other, we removed variants that are in LD with each other ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}^{2} > 0.1$$\end{document} r 2 > 0.1 ) and only kept one variant in each LD region. We used 1000 G EUR as a reference panel for LD calculation. Finally, the remaining variants were classified into pleiotropic concordant SNPs and discordant SNPs according to the direction of effect on traits.
We selected traits/diseases that have at least 3 SNPs in each category: (1) antagonistic SNPs, (2) SNPs with concordant effect on traits from different domains and (3) SNPs that affect traits in the same domain. SNPs in the MHC region were excluded from the analysis. We calculated \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$f \times (1-f) \times {\beta }^{2}$$\end{document} f × ( 1 − f ) × β 2 to represent genetic variance.
We tested the \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T at pleiotropic conflict SNPs with a matched set sampled from the genome following the procedure described in Guo et al 55 . We first calculated the MAF and LD score of all the SNPs on 1KG EUR data. \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T was calculated using 1KG EUR, EAS and AFR populations. The \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T method described in Weir, B.S. 1996 56 was implemented in GCTA 57 . MAF was set from 0 to 0.5, with an increment of 0.025. In each MAF bin, the LD scores were divided into 20 groups by an increment of 5% percentile. Then each putatively causal antagonistic SNP was allocated into a group of SNPs with MAF and LD scores matched. These groups with antagonistic SNPs allocated were combined as the full control SNP list for the permutation test. Next, for each antagonistic SNP, we sampled one SNP from its matched control group (MAF and LD score matched). For each permutation, we calculated the mean \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T of the sampled control SNPs. We repeated the sampling procedure 10,000 times. The \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T distribution of the control SNPs was generated from the permutations.
The causal genes were looked up from the Open Targets Platform 28 , a database covering all the putatively causal genes for published GWAS studies. In general, the causal genes were prioritised by leveraging proximity and functional information. For each associated locus, the causal genes were acquired from the results prioritised by the locus-to-gene (L2G) pipeline. We followed these customised rules to link the shared associated locus to genes: 1) If the locus is not colocalized with a QTL locus, the locus was linked to the gene that shows the highest overall L2G score. 2) If the locus is colocalized with QTL, the locus was linked to the colocalized gene(s) and the gene that shows the highest overall L2G score if these were different. 3) For a locus that is not on Open Target, the prioritised variant was linked to eQTL gene(s) via the variant-to-gene (V2G) pipeline, which was designed by Open Targets Genetics 58 . 4) If it is not an eQTL SNP, it was linked to the nearest gene.
We take all the linked putative causal antagonistic genes for gene-set enrichment analysis. Gene-set enrichment analysis was conducted using FUMA 29 , with the hallmark gene sets obtained from MsigDB 59 . Gene-set P -value was corrected for multiple testing using the Benjamini-Hochberg (FDR) method. Molecular function enrichment was analysed using the GO enrichment analysis tool ( http://geneontology.org/ ). The enriched pathways were reported at \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${P}_{{FDR}} < 0.05$$\end{document} P F D R < 0.05 .
To detect recent positive selection, we performed both iHS and nS L analysis using selscan 60 on the UK10K data. Briefly, we selected 3647 unrelated individuals using GCTA 57 (--rel-cut-off 0.05). SNPs with MAF < 0.001, HWE \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$p < 1\times {10}^{-6}$$\end{document} p 0.05 were removed. As iHS is a haplotype-based method, which uses phased genetic data as input, statistical phasing was conducted using Shapeit4 61 with the genetic recombination map downloaded from the 1000 genomes resource ( ftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/working/20130507_omni_recombination_rates/ ). Genetic variants with MAF < 0.05 were excluded from the iHS and nS L analyses. The ancestral state of each SNP was acquired from the dbSNP FTP site ( https://ftp.ncbi.nlm.nih.gov/snp/organisms/human_9606/database/organism_data/SNPAncestralAllele.bcp.gz ). In selscan, the unstandardised iHS is calculated as \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\rm{ln}}}(\frac{{{iHH}}_{1}}{{{iHH}}_{0}})$$\end{document} ln ( i H H 1 i H H 0 ) , where \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{iHH}}_{1}$$\end{document} i H H 1 is the integrated haplotype homozygosity for the derived haplotype, while \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{iHH}}_{0}$$\end{document} i H H 0 is the integrated haplotype homozygosity of the ancestral haplotype. Therefore, a large positive value indicates a long-range haplotype carrying the derived allele; a large negative value indicates a long-range haplotype carrying the ancestral allele. Both positive and negative values were considered in this study since we assumed that selection could act on both ancestral and derived alleles. We performed nS L analysis as a complementary validation, as it leverages haplotype lengths based on segregating sites, which makes it less sensitive to the variation of recombination and/or mutation rate 35 .
Tajima’s D in 1 kb and 10 kb window sizes were also calculated using the UK10K data. VCFtools 62 was used for the calculation.
To detect signals of balancing selection, we used BetaScan2, which utilises both polymorphism and substitution data 37 . This can slightly improve detection power compared to the Beta1 statistic, which only uses polymorphism data 37 , 41 . To determine ancestral alleles, we downloaded the epo file, which gives variant states information in the outgroup as the documentation suggested ( https://github.com/ksiewert/BetaScan/wiki/Tutorial ). This file was also used to call substitutions where the allele of humans differs from the outgroup allele. Then glactools was used to convert UK10K VCF files to ACF files needed by BetaScan. We calculated Beta2 with the divergence time set to 12.5 (in coalescent units).
For disease traits collected from FinnGen, the information on the age of onset was acquired from https://risteys.finregistry.fi/ , which is used for exploring FinnGen data at the phenotype level. The median age of the first event was used in the analysis. For disease traits from other resources, the age of onset was also looked up in FinnGen if the same disease exists in FinnGen. If a disease trait is not available in FinnGen, we looked it up from UpToDate ( https://www.uptodate.com ), which provides information on disease epidemiology. In general, we considered schizophrenia, eating disorders, type 1 diabetes, IBS, hayfever and some other respiratory diseases as early onset diseases, while age at first birth, PCOS and endometriosis are related to reproduction success. The details of the age of onset of diseases are shown in Supplementary Data 11 .
Further information on research design is available in the Nature Portfolio Reporting Summary linked to this article.
Results
The overview of the study is shown in Fig. 1 . Disease/trait-associated genomic regions were identified using GWAS summary statistics from FinnGen R6, UK Biobank and other large GWAS studies conducted using individuals of European genetic ancestry (inferred from genetic principal components), including binary and quantitative traits (Supplementary Data 1 ). Each quantitative trait was linked to human fitness using the evidence of natural selection observed with modern traits in a previous study 22 to provide a direction of effect to enable comparison with disease traits. For example, higher heel bone density, early age at first birth and later menopause were assumed to be linked to higher fitness. To aid the discussion of results, we have broadly classified diseases as autoimmune diseases and non-autoimmune diseases. For non-autoimmune diseases, ICD-10 was used to assign each disease to a biological domain (for example, respiratory system, Supplementary Data 2 ). We took a systematic approach to identify pleiotropic regions and pleiotropic SNPs. We define a sentinel association as the most associated SNP in a genomic region. It is well recognised that GWAS sentinel SNPs do not necessarily represent causal variants, but they are expected to be correlated with causal variants. Hence, a pleiotropic causal variant could be tagged by different sentinel SNPs in different GWASs. Thus, for each GWAS study, we first performed clumping to identify a set of roughly independent trait-associated sentinel SNPs ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P < 1\times {10}^{-7},\,{r}^{2} < 0.1$$\end{document} P < 1 × 10 − 7 , r 2 < 0.1 ). Our analysis found the MHC region was the most pleiotropic region in the genome. Multiple studies have reported the pleiotropic effect of classical HLA alleles across phenotypes 23 – 25 . Given that this region has been well studied and the existence of long LD making the identification of shared and independent associations difficult, the MHC region (chr6:24 Mb~34 Mb) was excluded from downstream analysis. Rare variants (MAF < 0.01) were also removed from our study. Potential pleiotropic regions were identified by considering a 2 Mb region centred around each sentinel SNP across all traits, followed by colocalization analysis on regions associated with two or more traits. SNPs prioritised by colocalization analysis having opposite effects on different traits were considered candidate antagonistic SNPs. Here, we report those loci showing high regional posterior probability, suggesting it is highly likely that the same associated region is shared by different traits. In other words, we included both high linkage disequilibrium (LD) regions, in which a putative causal variant cannot be pinpointed by colocalization analysis precisely, and regions where traits share a single causal variant with high posterior probability. Overall, we identified 71,257 significant trait-associated loci. These associations were mapped to 10,607 candidate pleiotropic regions, defined as genomic regions (constructed by creating a 2 Mb window around each sentinel SNP and merging overlapping windows across all traits that are associated with more than one trait). Among these regions, 219 regions (spanning 194 LD blocks) passed the colocalization test and discordantly affect different traits and diseases, covering approximately 11.4% of the independent LD regions in the human genome (assuming that there are 1703 independent LD blocks in the European genetic ancestry 26 ) (Supplementary Data 3 ) . Four hundred and four regions passed the colocalization test and showed concordant associations across traits from different domains (Supplementary Data 4 ). Fig. 1 The overview of the study. GWAS summary statistics were collected from UK Biobank (Neale lab round 2 summary statistics), FinnGen R6 and the GWAS catalogue. Sentinel SNPs for each GWAS study were defined with an LD reference matched to the population. For each GWAS, an associated region was defined as 1 Mb nearby the sentinel SNP. Overlapped associated regions and associated regions close to each other (<250 kb) were merged into one region. Pleiotropic regions were identified by intersecting associated regions of all diseases/traits. This was followed by colocalization analysis. The first two scenarios were considered as shared associated regions. Finally, we investigated the characteristics of antagonistic variants (eg., minor allele frequency and effect size) as well as the signatures of selection on each region. COPD, Chronic Obstructive Pulmonary Disease; PIP, Posterior Inclusion Probability; EAS, East Asian; EUR, European; AFR, African; Fst, fixation index; iHS, integrated haplotype homozygosity score; nS L , number of segregating sites by length.
GWAS summary statistics were collected from UK Biobank (Neale lab round 2 summary statistics), FinnGen R6 and the GWAS catalogue. Sentinel SNPs for each GWAS study were defined with an LD reference matched to the population. For each GWAS, an associated region was defined as 1 Mb nearby the sentinel SNP. Overlapped associated regions and associated regions close to each other (<250 kb) were merged into one region. Pleiotropic regions were identified by intersecting associated regions of all diseases/traits. This was followed by colocalization analysis. The first two scenarios were considered as shared associated regions. Finally, we investigated the characteristics of antagonistic variants (eg., minor allele frequency and effect size) as well as the signatures of selection on each region. COPD, Chronic Obstructive Pulmonary Disease; PIP, Posterior Inclusion Probability; EAS, East Asian; EUR, European; AFR, African; Fst, fixation index; iHS, integrated haplotype homozygosity score; nS L , number of segregating sites by length.
We observed that most of the identified loci (211 of 219) affect traits from different biological domains, highlighting that these genomic regions have effects on a wide range of phenotypes. To further characterise these antagonistic regions, we assessed genetic correlations for 139 disease pairs corresponding to the 139 antagonistic loci. Of these pairs, 23 showed significant genetic correlations ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P < 3.6\times {10}^{-4}$$\end{document} P < 3.6 × 10 − 4 ), including two with significantly negative correlations at this threshold. At a nominal significance level ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P < 0.05$$\end{document} P < 0.05 ), 54 pairs were significantly correlated, nine of which were negative. Overall, most disease pairs were either positively correlated or showed no correlation, suggesting predominantly horizontal antagonistic effects (Supplementary Data 5 ).
Pleiotropic conflicts were identified both within and across biological domains. The domains with the highest numbers of conflicts were malignant neoplasm (77), autoimmune (68), circulatory (66), endocrine diseases (62) and digestive diseases (60) having the largest number of conflicts with other diseases within and across domains (Fig. 2a ). To assess whether pleiotropic conflicts were enriched in specific domains, we performed a one-sided Fisher’s exact test comparing the number of conflicts in each domain to all other domains by accounting for the number of diseases in each domain. A one-sided test was used because our hypothesis was directional: we specifically aimed to identify domains enriched for antagonistic pleiotropy rather than testing for enrichment or depletion. We observed that autoimmune, circulatory, respiratory, endocrine, metabolic, nutritional diseases and cancers more frequently exhibited antagonistic associations with diseases in other domains ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=2.96\times {10}^{-18},{P}=3.71\times {10}^{-9},{P}=8.04\times {10}^{-7},{P}=8.37\times {10}^{-6},{P}=1.06\times {10}^{-2},{P}=3.64\times {10}^{-2}\, {{\rm{and}}}{P}=3.91\times {10}^{-2}$$\end{document} P = 2.96 × 10 − 18 , P = 3.71 × 10 − 9 , P = 8.04 × 10 − 7 , P = 8.37 × 10 − 6 , P = 1.06 × 10 − 2 , P = 3.64 × 10 − 2 and P = 3.91 × 10 − 2 , respectively, one-sided Fisher’s exact test). Conflicts occurred more frequently in diseases from different domains (196 out of 219 (~90%)) than within domains (23 out of 219 (~10%)). Genetic variants discordantly affecting traits in the same domain were observed across autoimmune diseases (9 loci), female reproductive traits (5 loci), diseases of the digestive system (3 loci), diseases of the nervous system (3 loci) and some others. Many conflicts were identified across domains. For example, we observed that 16 loci are shared between autoimmune diseases and endocrine diseases (11 occur between thyroid diseases with and without an immune component), and 11 loci are shared between autoimmune diseases and respiratory diseases. Diseases of the circulatory and digestive systems shared 10 antagonistic loci. Migraine and cardiovascular diseases, which partly share genetic aetiology 27 , were observed to be discordantly affected by 4 genetic loci. Additionally, some genetic loci were identified to be discordantly associated with the risk of cardiovascular diseases, lipid metabolic disorders and disorders of the gallbladder. Inverse associations were also observed between type 2 diabetes and malignant neoplasms such as breast and prostate cancer. Compared with SNPs showing non-antagonistic associations with multiple traits, the antagonistic SNPs were found to be associated with traits across more biological domains ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P < 2.2\times {10}^{-16}$$\end{document} P < 2.2 × 10 − 16 , one-sided Mann-Whitney U -test) (Fig. 2b ). Fig. 2 Antagonistic variants affect traits within and across biological domains. a The number of pleiotropic conflict events across each domain pair. Different domains are given along the x-axis and y-axis. The number in each bubble indicates the number of conflicts identified in each domain pair. The total number of identified conflicts for each biological domain is given along the y-axis (right). Note that some single diseases/traits are presented as a domain. b Distribution of antagonistic versus other genetic variants across the number of affected biological domains. The left panel shows the number of antagonistic variants stratified by how many biological domains each variant is associated with. The right panel shows the number of all other genetic variants stratified by how many biological domains each variant is associated with. Pregnancy: Pregnancy, childbirth and the puerperium; Skin: Diseases of the skin; Perinatal period: Certain conditions originating in the perinatal period; Blood: Diseases of the blood and blood-forming organs and certain disorders involving the immune mechanism; ANM: Age at natural menopause.
a The number of pleiotropic conflict events across each domain pair. Different domains are given along the x-axis and y-axis. The number in each bubble indicates the number of conflicts identified in each domain pair. The total number of identified conflicts for each biological domain is given along the y-axis (right). Note that some single diseases/traits are presented as a domain. b Distribution of antagonistic versus other genetic variants across the number of affected biological domains. The left panel shows the number of antagonistic variants stratified by how many biological domains each variant is associated with. The right panel shows the number of all other genetic variants stratified by how many biological domains each variant is associated with. Pregnancy: Pregnancy, childbirth and the puerperium; Skin: Diseases of the skin; Perinatal period: Certain conditions originating in the perinatal period; Blood: Diseases of the blood and blood-forming organs and certain disorders involving the immune mechanism; ANM: Age at natural menopause.
To investigate whether the identified pleiotropic conflicts are shared among ancestries, we looked up the 219 identified antagonistic variants in Biobank Japan. Three pairs of associations were replicated at the level of genome-wide significance. These include GCKR , which discordantly affects type 2 diabetes and cholesterol metabolic disorder; PSCA , a locus that has an opposite effect on gastric ulcer and gastric cancer; HNF1B , the associated variants in this locus have an opposing effect on the risk of prostate cancer and type 2 diabetes. We identified 62 trait pairs with at least one trait that can be found in Biobank Japan for each direction. Of these, 11 (17.7%) antagonistic variants were replicated at the intermediate significance of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$8.06\times {10}^{-4}$$\end{document} 8.06 × 10 − 4 (0.05/62) and 19 (30.6%) of them were replicated at 0.05 significance (Supplementary Data 6 ).
Putative causal genes for the identified pleiotropic conflict loci were linked using the locus-to-gene and variant-to-gene models provided by the Open Targets Genetics 28 (Supplementary Data 7 ). Open Targets Genetics leveraged functional data and QTLs data to prioritise causal genes and identified a total of 354 putative causal genes for the 219 investigated regions (160 regions were linked to a single gene). Gene-set enrichment analysis was performed on the putative causal genes by FUMA 29 , which highlighted 24 gene sets (Supplementary Fig. 1a ). These are mainly related to immune, cancer, cell proliferation and ageing. Specifically, ten genes highlighted by FUMA are involved in TNF alpha signalling via NF-KB, an inflammatory response pathway regulated by NF-KB. Other immune-related pathways are also implicated, such as interferon-gamma response, IL-2-STAT5, and complement. The enrichment in the p53 pathway is driven by six genes. These genes are particularly involved in DNA maintenance, controlling cell proliferation and the cell cycle. The IL6/JAK/STAT3 signalling pathway, which is involved in both immune responses and cancer, was also highlighted. In addition, we used the specific gene lists (cancer-related genes, oncogenes, tumour suppressor genes and housekeeping genes) from Voskarides 30 to test if antagonistic genes are enriched for any of these gene sets. We found that AP genes are enriched for cancer-related genes (Fisher’s exact test \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=1.12\times {10}^{-12}$$\end{document} P = 1.12 × 10 − 12 ), tumour suppressor genes ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=8.99\times {10}^{-4}$$\end{document} P = 8.99 × 10 − 4 ), but not oncogenes ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=0.07$$\end{document} P = 0.07 ) and housekeeping genes ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=0.58$$\end{document} P = 0.58 ), broadly consistent with the FUMA gene set analysis. In terms of molecular function, we observed that these genes are strongly enriched in the DNA binding pathway, suggesting their roles in gene regulation (Supplementary Fig. 1b ). Notably, of the 354 linked genes, 58 were linked to transcription factor activity ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${P}_{{adjusted}}=1.76\times {10}^{-14}$$\end{document} P a d j u s t e d = 1.76 × 10 − 14 ).
We compared the minor allele frequency (MAF) distribution of the putatively causal antagonistic SNPs to SNPs that have concordant effects on different traits from more than one biological domain. Different MAF distributions were observed between these two groups (Kolmogorov–Smirnov test \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=2.78\times {10}^{-4}$$\end{document} P = 2.78 × 10 − 4 , Fig. 3a ). For antagonistic variants, the proportion of variants in each MAF bin increased with MAF. In contrast, variants that have concordant effects on traits tend to be uniformly distributed in different MAF bins. However, the detection power is determined by allele frequency of the risk allele, genotype relative risk, number of cases and controls, significance threshold and lifetime risk of disease. For antagonistic variants detection, the risk minor alleles will have more power to be detected than the risk major alleles, assuming all the other parameters are the same constant for different diseases. Given that the GWAS summary statistics were collected from large biobanks and consortiums, we expected ascertainment bias in our comparison to be limited. Fig. 3 Characteristics of antagonistic variants and the comparison between antagonistic variants and non-antagonistic variants. a A comparison of minor allele frequency probability density distribution between SNPs that have a concordant effect on traits in different domains and antagonistic SNPs. The left panel is the allele frequency distribution of putatively causal SNPs that have concordant effects on traits in different domains. The right panel is the allele frequency distribution of putatively antagonistic causal SNPs. b Comparison of genetic variance explained by antagonistic, concordant (SNPs affect diseases from >=2 domains) and single domain SNPs. Violin plots show kernel density estimates. Embedded boxplots display the median (centre line), the 25th and 75th percentiles (box), and whiskers extending to 1.5× the interquartile range. Sample size (n) corresponds to the number of SNP-disease associations: concordant = 499; discordant = 325; single domain = 6467. p -values are calculated using two-sided Welch’s t -tests comparing mean genetic variance explained between groups and are not corrected for multiple testing. Significance level: p < 0.0001 (****). c The differences between the mean value of observed \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t and permutations. The mean value of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t is given along the x-axis. Y-axis represents the frequency of occurrences over a range of mean \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t in 10,000 permutations. The green bars show the permutation distribution. The vertical pink line is the mean value of the observed \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t .
a A comparison of minor allele frequency probability density distribution between SNPs that have a concordant effect on traits in different domains and antagonistic SNPs. The left panel is the allele frequency distribution of putatively causal SNPs that have concordant effects on traits in different domains. The right panel is the allele frequency distribution of putatively antagonistic causal SNPs. b Comparison of genetic variance explained by antagonistic, concordant (SNPs affect diseases from >=2 domains) and single domain SNPs. Violin plots show kernel density estimates. Embedded boxplots display the median (centre line), the 25th and 75th percentiles (box), and whiskers extending to 1.5× the interquartile range. Sample size (n) corresponds to the number of SNP-disease associations: concordant = 499; discordant = 325; single domain = 6467. p -values are calculated using two-sided Welch’s t -tests comparing mean genetic variance explained between groups and are not corrected for multiple testing. Significance level: p < 0.0001 (****). c The differences between the mean value of observed \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t and permutations. The mean value of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t is given along the x-axis. Y-axis represents the frequency of occurrences over a range of mean \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t in 10,000 permutations. The green bars show the permutation distribution. The vertical pink line is the mean value of the observed \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{st}}$$\end{document} F s t .
We also compared the genetic variance explained by (1) antagonistic SNPs, (2) SNPs with concordant effect on traits from different domains and (3) SNPs that affect traits in the same domain. Across 33 selected diseases that have at least 3 SNPs in each of the 3 classes, we observed that the genetic variance explained by antagonistic SNPs was significantly larger than that of SNPs showing concordant effects across domains and SNPs affecting traits in a single domain (Fig. 3b ).
Considering the identified antagonistic SNPs were enriched at intermediate allele frequency, we then asked whether these regions were under selective pressure. We first calculated Wright’s fixation index \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T for each SNP using three populations (EUR, AFR and EAS) from 1000 Genomes data. Briefly, \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T is a means of assessing population structure and an estimate of the variance of allele frequency among populations. This indicator reflects multiple evolutionary processes, including selection, migration, mutation and genetic drift 31 . It has been commonly used for identifying signatures of divergent selection, with loci exhibiting larger \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T being candidates that are subject to natural selection 32 , 33 . Here, we generated 10,000 sets of matched control SNPs by sampling SNPs from the 1000 Genomes EUR data, with MAF and LD scores matched to the observed antagonistic SNPs. We then sampled each control SNP from its matched control list and calculated the average \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T among the control SNPs. We observed strong evidence that the observed antagonistic SNPs have larger \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T compared to the control SNPs (Fig. 3c ).
The larger effect size, MAF and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${F}_{{ST}}$$\end{document} F S T of antagonistic variants indicate that some antagonistic regions are putatively under natural selection. We further used other population genetics metrics, including the integrated haplotype score ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${iHS}$$\end{document} i H S ) 34 , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$n{S}_{L}$$\end{document} n S L 35 , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\rm{Tajima}}}{\prime} {{\rm{s\; D}}}$$\end{document} Tajima ′ s D 36 and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}$$\end{document} β ( 2 ) statistics 37 to detect signatures of natural selection in antagonistic regions. In brief, \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${iHS}$$\end{document} i H S is a metric that quantifies the signature of positive selection using haplotype data. Compared to \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${iHS}$$\end{document} i H S , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$n{S}_{L}$$\end{document} n S L is less sensitive to the variation of recombination rates and different demographic models. A larger \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|$$\end{document} i H S and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|$$\end{document} n S L indicate a signal of positive selection. \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\rm{Tajima}}}{\prime} {{\rm{s\; D}}}$$\end{document} Tajima ′ s D is a commonly used test that indicates a signal of positive selection. \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\rm{Tajima}}}{\prime} {{\rm{s\; D}}}$$\end{document} Tajima ′ s D is a commonly used test of neutrality. \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}$$\end{document} β ( 2 ) score is used to test long-term balancing selection using allele frequency correlation. These analyses were performed on UK10K unrelated individuals. For each antagonistic SNP, we extracted \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${|iHS|}$$\end{document} ∣ i H S ∣ , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|$$\end{document} n S L and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}$$\end{document} β ( 2 ) values of SNPs in LD ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}^{2} > 0.3$$\end{document} r 2 > 0.3 ) with it (Supplementary Data 8 , 9 ) and calculated \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\rm{Tajima}}}{\prime} {{\rm{s\; D}}}$$\end{document} Tajima ′ s D around the identified antagonistic SNPs (Supplementary Data 10 ). The empirical distribution of each metric was generated from the whole genome results. We observed that 35 (16.5%) of the tested 212 regions (MAF of antagonistic variants greater than 0.05) have \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|$$\end{document} i H S greater than the 1% threshold of the empirical distribution ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|$$\end{document} i H S = 2.6), 19 (8.9%) of the tested 212 regions have \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|$$\end{document} n S L greater than the 1% threshold of the empirical distribution ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|$$\end{document} n S L = 2.54) and 17 (7.8%) regions have a \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}$$\end{document} β ( 2 ) greater than the 1% threshold of the empirical distribution ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}=9.43$$\end{document} β ( 2 ) = 9.43 ). We compared the distribution of \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|$$\end{document} i H S , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|$$\end{document} n S L and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}$$\end{document} β ( 2 ) of SNPs in antagonistic regions to their respective genome-wide distributions and found that SNPs in antagonistic regions tend to exhibit larger values (Supplementary Figs. 2 – 4 ) .
We identified several antagonistic variants associated with diseases regulated by the immune system that are under selection, suggesting that the genetic trade-off might be due to adaptation to pathogens in the past. For example, the TLR gene cluster was identified as a candidate region under positive selection, with rs10008032 showing strong selection signal ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|=2.59$$\end{document} n S L = 2.59 , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|=3.6$$\end{document} i H S = 3.6 ), which is in LD with rs5743618 ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${r}^{2}=0.34$$\end{document} r 2 = 0.34 ), a missense variant in the innate immune gene TLR1 (Fig. 4a ). This SNP was identified to be associated with hay fever and breast cancer in opposite direction. Additionally, previously reported loci such as SH2B3/ATXN2 (rs10850001: \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|=2.03$$\end{document} n S L = 2.03 , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|=4.41$$\end{document} i H S = 4.41 ) and FUT2 (rs2617801: \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|n{S}_{L}\right|=2.89$$\end{document} n S L = 2.89 , \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\left|{iHS}\right|=4.00$$\end{document} i H S = 4.00 ) also showed a signature of positive selection with the UK10K data (Supplementary Fig. 5a, b ). We also identified some immune-related loci under balancing selection. A striking example is the NFKB1 locus, a transcription factor that can be activated by bacteria and viruses and regulates inflammatory responses, showing larger Tajima’s D value (regional largest \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${{\rm{Tajima}}}{\prime} {{\rm{s\; D}}}=\,3.21$$\end{document} Tajima ′ s D = 3.21 ) and \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}$$\end{document} β ( 2 ) statistics (largest \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$${\beta }^{(2)}=12.95$$\end{document} β ( 2 ) = 12.95 ) (Fig. 4c ). The signature of balancing selection at this locus is colocalized with GWAS signals of hay fever and primary biliary cirrhosis. We also detected that the CTLA4 locus, which regulates T-cell responses, is under balancing selection (Fig. 4d ). This suggests that the conflict effects on different diseases might be due to different haplotypes that have been maintained for a long time. Fig. 4 Selected examples of signatures of selection observed in antagonistic regions. The top two figures are GWAS signals of two traits, which are affected by the locus in opposite directions. Each dot represents an SNP. The x-axis is the genomic position, while the y-axis is \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$-{\log }_{10}(P)$$\end{document} − log 10 ( P ) . a , b The bottom figure shows |nS L | in the genomic region. Each dot indicates |nS L | of a single SNP, while the line connects the average |nS L | within 500 bp. c , d The third row is Tajima’D in the genomic region. Each dot indicates Tajima’s D value calculated with a window size of 1 kb, while the line connects Tajima’s D value calculated with a window size of 10 kb. The bottom figure shows the beta score in the genomic region. Each dot indicates the beta score of a single SNP. The colocalized GWAS and selection signals are highlighted in a rectangle.
The top two figures are GWAS signals of two traits, which are affected by the locus in opposite directions. Each dot represents an SNP. The x-axis is the genomic position, while the y-axis is \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$-{\log }_{10}(P)$$\end{document} − log 10 ( P ) . a , b The bottom figure shows |nS L | in the genomic region. Each dot indicates |nS L | of a single SNP, while the line connects the average |nS L | within 500 bp. c , d The third row is Tajima’D in the genomic region. Each dot indicates Tajima’s D value calculated with a window size of 1 kb, while the line connects Tajima’s D value calculated with a window size of 10 kb. The bottom figure shows the beta score in the genomic region. Each dot indicates the beta score of a single SNP. The colocalized GWAS and selection signals are highlighted in a rectangle.
In addition to immune-related loci, we detected several essential genes, including transcription factors and other genes that participate in basic activities throughout life. For example, we identified that the CD44 locus is under positive selection and discordantly affects heel bone density and uteri leiomyoma (Fig. 4b ). The CD44 gene is involved in cell-cell interaction, cell adhesion and migration, and is highly expressed in many cancers 38 . It was also identified for bone loss 39 . Another example is the BNC2 locus, which encodes a conserved zinc finger protein and has opposite effects on age at first birth and lifespan. Some transcription factors and genes regulating gene expression showed signatures of balancing selection. These include the NFKB1 gene, BICC1 gene and SOX9 gene (Fig. 4c , Supplementary Fig. 6a, b ). The BICC1 gene encodes an RNA-binding protein and plays a fundamental role in regulating gene expression, which is associated with heel bone density and glaucoma in the opposite direction.
Our results also showed some other interesting examples, including the MAPT/CRHR1 cluster, PSCA , INHBB , CFDP1 , TSPAN10 and ADCY3 genes. We observed a signature of positive selection at the INHBB locus (Supplementary Fig. 5c ), which encodes a subunit of both activin and inhibin participating in stimulating and inhibiting FSH secretion. The SNP rs4849879(G) was associated with a lower risk of breast hypertrophy, which is related to hormonal changes and may occur during puberty, but a higher risk of breast cancer. The region containing the TSPAN10 locus has an opposing effect on the risk of basal cell carcinoma and eye disorder (Supplementary Fig. 5d ). Interestingly, this locus is strongly associated with hair colour and skin colour. The CFDP1 locus, which has an opposing effect on cardiovascular diseases and migraine, was identified under positive selection (Supplementary Fig. 5e ). The MAPT/CRHR1 cluster was found to be under natural selection 40 and has been discordantly associated with Parkinson’s disease and neuroticism (Supplementary Fig. 6c ). Similarly, we observed signature of balancing selection at the ADCY3 locus, with a putative causal SNP rs4343432(A) associated with a higher risk of breast cancer but lower risk of Crohn’s disease (Supplementary Fig. 6d ). Another example is that the PSCA locus (rs1045547) has an opposite effect on the risk of stomach cancer and duodenal ulcer, which was detected to be under balancing selection 41 (Supplementary Fig. 6e ).
We examined the age of onset for each disease within antagonistic trait pairs and classified antagonistic regions into three categories: early-early, early-late and late-late. Among the 168 disease/trait pairs, 115 (68%) were affected by pleiotropic conflicts occurring late in life, while 51 (30%) involved conflicts between early-onset and late-onset diseases. After accounting for the number of diseases in early- and late-onset categories, we observed that antagonistic SNPs were significantly enriched in early-late trait pairs ( \documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$P=6.55\times {10}^{-18}$$\end{document} P = 6.55 × 10 − 18 , Fisher’s exact test), consistent with the classical model of antagonistic pleiotropy (Fig. 5a ). When focusing on regions affecting pairs of diseases with onset early and late in life, approximately one-third of the loci were linked to genes related to the immune system. Representative examples are shown in Fig. 5b . Notably, several antagonistic loci discordantly affect immune-related diseases and malignant neoplasms. The SNP rs3087243 (A) in the CTLA4 gene, which encodes a protein on T cells that regulates immune responses, is significantly associated with a lower risk of autoimmune diseases, including Type 1 diabetes and myxoedema, but a higher risk of basal cell carcinoma. Similarly, this was also observed in the PTPN22 gene (SNP rs2476601), which encodes a key regulator of immune homoeostasis. Fig. 5 Antagonistic regions affect disease pairs in early-late and late-late age of onset patterns. a The heatmap shows the proportion of disease pairs falling into each age-of-onset combination. Values were normalised by dividing the number of conflict pairs in each age bin by the total number of disease pairs involving that age-of-onset group. b Selected examples showing antagonistic regions affect life-history traits. Diseases/traits affected by the genomic region in opposite directions are distinguished in black and blue.
a The heatmap shows the proportion of disease pairs falling into each age-of-onset combination. Values were normalised by dividing the number of conflict pairs in each age bin by the total number of disease pairs involving that age-of-onset group. b Selected examples showing antagonistic regions affect life-history traits. Diseases/traits affected by the genomic region in opposite directions are distinguished in black and blue.
Hormones have been implicated to have conflicting effects and exhibit trade-offs on traits that occur at different life stages 42 . Among the identified antagonistic genes, 10 genes ( WNT4, GCKR, EEFSEC, BMPR1B, TERT, FSHB, ESR1, MYRF, WT1 and SALL1 ) were also associated with reproductive traits, including the number of children ever born and diagnoses related to reproductive health 43 , 44 . Several variants involved in the reproductive system have opposite effects on traits occurring at different life stages. For example, rs11031005 (T), located near FSHB , was associated with an earlier age at first birth but also with increased risk of genitourinary system diseases and earlier age at natural menopause. This suggests that a genetic variant associated with earlier reproductive behaviour may incur costs later in life, potentially mediated by hormones. Genetic variants at or near the WNT4 locus were associated with discordant risks of endometriosis and female genital prolapse, which occur at early and late stages of life. The WNT4 protein is responsible for the production of androgens and the development of the female reproductive system. Taken together, these findings highlight that hormone-related genes contribute to antagonistic effects across different stages of the human life course.
Discussion
Understanding pleiotropic regions with conflicting effects across diseases is important for human health, particularly for drug development 45 . In this study, we demonstrated that antagonistic variants influencing contemporary human diseases are widespread across the genome. By leveraging GWAS summary statistics with multiple population genetic metrics, we characterised antagonistic regions and found that putatively causal antagonistic SNPs tend to show greater population differentiation. Many of these regions are also associated with life-history traits.
We observed strong evidence for the involvement of immune-related loci, several of which affect life-history traits (Fig. 5b ). Previous studies have suggested that early-life infections and inflammation play an important role in shaping disease risk later in life 46 . In particular, alleles conferring resistance to infections in early life have been shown to have opposite effects on longevity and adult-onset diseases 47 . Such conflicts may reflect evolutionary forces, including positive selection and balancing selection, which could explain the high proportion of common variants observed in this study. Analyses of ancient DNA have shown that alleles adapted to past environments can contribute to the risk of autoimmune diseases in present 48 .
Our analysis of disease age of onset further revealed that a large proportion of pleiotropic conflicts occur between early- and late-onset diseases as well as between late-onset diseases. Notably, we identified hormone-related loci with antagonistic effects on life-history traits, consistent with the role of hormones in regulating physiological processes across different stages of life. In addition, putatively causal genes were highly enriched for transcription-related functions. Previous work examining pleiotropy at the cellular network level has shown that the pleiotropic gene modules are enriched for core physiological processes 49 , however, the directionality of effects of genes remains to be investigated.
Despite providing a comprehensive analysis of antagonistic variants in the human genome, our study has several limitations. First, identifying causal variants from GWAS remains challenging, and we were only able to identify the shared antagonistic regions rather than causal variants. As a result, whether opposite effects on target traits are driven by the same variant cannot be determined. Secondly, limited by the range of diseases examined in GWAS, we could not infer the selective pressures underlying these antagonistic effects. For example, genetic adaptations to ancestral pathogen exposure cannot be quantified, and we rely on modern-day diseases as proxies. Thus, there are likely many more regions of antagonistic effects in the genome. Thirdly, using Biobank Japan to replicate antagonistic variants may reduce the consistency due to ancestry-specific LD and allele frequency differences, and gene-environment interactions. Such factors may influence the direction and magnitude of genetic effects.
Overall, our study provides a picture of widespread antagonistic variants across human diseases, highlighting the complexity of the human genome. We expect the number of identified antagonistic regions to increase as the range of phenotypes examined via GWAS increases, and with biobank-scale studies across ancestral diversity. Integrating large -scale GWAS resources and population genetic approaches will further advance our understanding of how genetic and phenotypic diversity is maintained in human complex traits.
Introduction
Pleiotropy is defined as a single genetic unit (either a variant or a gene) that influences multiple traits. It is a pervasive phenomenon observed in many species 1 – 3 . In humans, pleiotropy is observed as a characteristic of many Mendelian diseases 4 , 5 , and more recently, of human complex traits 6 , 7 , demonstrating that the human genome is highly pleiotropic 7 , 8 . For example, sickle cell anaemia is caused by a mutation in the HBB gene in chromosome 11, and the abnormal red blood cells further influence other tissue functions, including the eyes, liver and heart 9 . A survey of GWAS-associated loci showed 90% of GWAS-associated loci overlap with more than one trait 7 . The extent of pleiotropy in the human genome is further highlighted by extensive genetic correlations observed between pairs of traits 7 , 10 .
While most pleiotropic loci have concordant effects for disease risk, a subset of pleiotropic loci have opposing effects on human diseases, resulting in pleiotropic conflict. Such conflict has been reported in several empirical studies of human diseases. For example, discordant genetic effects have been reported across brain disorders 11 , with a meta-analysis across multiple disorders identifying three loci that have opposite effects on schizophrenia and depression, and two between schizophrenia and autism 11 . The shared genetic variants associated with immune-mediated diseases often have opposite effects, with a risk variant for one disease being protective for another 12 , 13 . Interestingly, many opposite directions of genetic effects on gene expression (eQTLs) have been identified across tissues in humans 14 , suggesting a potential mechanistic explanation for the discordant pleiotropy observed between diseases.
Antagonistic pleiotropy (AP), where a genetic variant controls a trait that is beneficial to an organism early in life but also controls another trait that is detrimental later in life, is the most well-studied example of pleiotropic conflict, with examples seen in humans. AP was first proposed by George Williams in 1957 as an explanation for senescence 15 , and recent evidence has continued to support the theory. For instance, as a tumour suppressor, p53 can prevent cancer by interrupting abnormal cell proliferation, but the increased activity of p53 can simultaneously contribute to ageing 16 . Another study showed that SNPs in the p53 protein are associated with human fertility in in vitro fertilisation (IVF) patients, but the effect was reduced or absent for older IVF patients, suggesting an antagonistic pleiotropic effect of p53 17 . A comparison of genetic variants associated with diseases that occur at different life stages supported the role of AP in ageing 18 . For example, coronary artery disease (CAD) risk loci have been identified to be antagonistically related to reproductive success 19 .
While the study of pleiotropy in humans is limited experimentally, the increasing availability of biobank cohorts and population-based disease collection for GWAS 20 , 21 provides an avenue to assess the extent of such conflict empirically. Here, we use GWAS summary statistics to identify antagonistic regions in the human genome through multi-trait colocalization analysis. We demonstrate extensive pleiotropic conflict both within and across trait biological domains (International Statistical Classification of Diseases 10th Revision codes). We observe signatures of positive and balancing selection at or near the identified antagonistic regions.