Large-scale genome-wide association study in a Japanese population identifies novel susceptibility loci across different diseases.

OA: closed
⚙ AI-generated summary by qwen3.7-flash, 2026-08-20 ⓘ

This large-scale genome-wide association study in 212,453 Japanese individuals identified 320 independent signals across 27 diseases, highlighting East Asian-specific variants and emphasizing the importance of non-European genetic research.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-08-20 · read from full text ⓘ

This large-scale genome-wide association study analyzed genetic data from approximately 200,000 individuals in the BioBank Japan Project to identify susceptibility loci for 42 common diseases. The researchers discovered 25 novel autosomal loci and highlighted East Asian-specific variants, such as missense mutations in ATG16L2, POT1, and PHLDA3, which were monomorphic or rare in European populations. By demonstrating that non-European cohorts can uncover unique genetic signals missed in Western studies, the paper emphasizes the critical need for ethnic diversity in precision medicine and polygenic risk score development. The paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

The overwhelming majority of participants in current genetic studies are of European ancestry. To elucidate disease biology in the East Asian population, we conducted a genome-wide association study (GWAS) with 212,453 Japanese individuals across 42 diseases. We detected 320 independent signals in 276 loci for 27 diseases, with 25 novel loci (P < 9.58 × 10-9). East Asian-specific missense variants were identified as candidate causal variants for three novel loci, and we successfully replicated two of them by analyzing independent Japanese cohorts; p.R220W of ATG16L2 (associated with coronary artery disease) and p.V326A of POT1 (associated with lung cancer). We further investigated enrichment of heritability within 2,868 annotations of genome-wide transcription factor occupancy, and identified 378 significant enrichments across nine diseases (false discovery rate < 0.05) (for example, NKX3-1 for prostate cancer). This large-scale GWAS in a Japanese population provides insights into the etiology of complex diseases and highlights the importance of performing GWAS in non-European populations.
Full text 48,672 characters · extracted from pmc-nxml · 5 sections · click to expand

Online

All case samples in this GWAS were collected in the BioBank Japan Project (BBJ; https://biobankjp.org/english/index.html ) 15 , 16 , which is a biobank that collaboratively collects DNA and serum samples from 12 medical institutions in Japan and recruited approximately 200,000 patients with the diagnosis of at least one of 47 diseases. Among them, cases with dyslipidemia were not analyzed in this study because it was already reported as a quantitative trait in our previous study 23 . Amyotrophic lateral sclerosis and febrile seizure were also not analyzed due to limited sample size. Cases with myocardial infarction, stable angina, and unstable angina were re-classified into a single disease category (coronary artery disease). Thus, we analyzed 42 disease in this study. For control samples, we used samples from the population-based prospective cohorts; the Tohoku University Tohoku Medical Megabank Organization (ToMMo), Iwate Medical University Iwate Tohoku Medical Megabank Organization (IMM) 53 , the Japan Public Health Center–based Prospective Study and the Japan Multi-institutional Collaborative Cohort Study. In addition, we also included samples in BBJ without related diagnoses into control group ( Extended Data Figure 1 and Supplementary Table 1 ). The sample sizes and the demographic data are provided in Supplementary Table 1 . All participating studies obtained informed consent from all participants by following the protocols approved by their institutional ethical committees. We obtained approval from ethics committees of RIKEN Center for Integrative Medical Sciences, and the Institute of Medical Sciences, The University of Tokyo. We have complied with all relevant ethical regulations. We genotyped samples with the Illumina HumanOmniExpressExome BeadChip or a combination of the Illumina HumanOmniExpress and HumanExome BeadChips. For quality control (QC) of samples, we excluded those with (i) sample call rate < 0.98 and (ii) outliers from East Asian clusters identified by principal component analysis using the genotyped samples and the three major reference populations (Africans, Europeans, and East Asians) in the International HapMap Project 54 . For QC of genotypes, we excluded variants meeting any of the following criteria: (i) call rate < 99%, (ii) P value for Hardy Weinberg equilibrium (HWE) < 1.0 × 10 −6 , and (iii) number of heterozygotes less than five. Using 939 samples whose genotypes were also analyzed by whole genome sequencing (WGS), we added additional QC based on the concordance rate between genotyping array and WGS. Variants with a concordance rate < 99.5% or a non-reference discordance rate ≥ 0.5% were excluded. We note that the allele frequency of rs671 (the East Asian-specific functional missense variant at ALDH2) substantially varies among the domestic regions within Japan due to strong selection pressure 55 and that genotypes of rs671 did not follow HWE. We thus did not apply the HWE QC for rs671. We had confirmed the 100% concordance of rs671 genotypes between the SNP microarray data used in this study and our internal WGS data (n = 2,798; see details in the discussion in ref 56 ). We utilized all samples in the 1000 Genomes Project Phase 3 (version 5; www.1000genomes.org/ ) 18 as a reference for imputation. We first pre-phased the genotypes with SHAPEIT2 (v2.778; https://mathgen.stats.ox.ac.uk/genetics_software/shapeit/shapeit.html ) and then imputed dosages with minimac3 (v2.0.1; https://genome.sph.umich.edu/wiki/Minimac ). After imputation, we excluded variants with imputation quality of Rsq < 0.7. For the X chromosome, we performed prephasing and imputation separately for males and females, and we excluded variants with imputation quality of Rsq < 0.7 in either of them. We conducted GWAS by employing a generalized linear mixed model (GLMM) using SAIGE (v0.29.4.2; https://github.com/weizhouUMICH/SAIGE ) 17 . This strategy enabled us to maintain related samples in our GWAS, and the sample sizes were increased by 6% on average compared to removing related samples. Briefly, there are two steps in SAIGE. In step 1, we fit a null logistic mixed model using genotype data, and we added covariates in this step (see below). In step 2, we performed the single-variant association tests using imputed variant dosages. We applied the leave-one-chromosome-out (LOCO) approach. For the X chromosome, we conducted GWAS separately for males and females, and merged their results by inverse-variance fixed-effect meta-analysis. We used only female control samples for GWAS of female-specific diseases; breast cancer, cervical cancer, endometrial cancer, ovarian cancer, endometriosis, and uterine fibroids. Similarly, we used only male control samples for GWAS of prostate cancer. We incorporated age and top 5 principal component (PC) as covariates. We also used sex as covariate for GWAS of diseases which include both of male and female samples. We also conducted male-specific and female-specific GWAS using the same pipeline as described above, and estimated heterogeneity in the effect size estimates using Cochran’s Q test. In each GWAS, we excluded variants with minor allele count (MAC) < 10 based on the recommendation from thee developers of SAIGE. We created regional association plots by LocusZoom (v1.2; http://locuszoom.sph.umich.edu/locuszoom/ ) 57 . We performed stepwise conditional analysis within ± 1 Mb from the lead variant; we repeated the association test by additionally incorporating the dosages of the identified variants as covariates in SAIGE step 1 until we do not detect any significant associations. For each disease, we defined a significantly associated locus as a genomic region within ± 1 Mb from the lead variant. When a locus did not include any variants which were previously reported to be significantly associated with the same disease ( P < 5.0 × 10 −8 ), we defined it as a novel locus. Since we tested each variant for disease association three times (sex-combined, female-specific, and male-specific analysis), we considered multiple-testing burden on the empirical significance threshold ( P = 2.87 x 10 −8 , see next paragraph), and we set the genome-wide significance threshold for our study at P = 2.87 x 10 −8 / 3 (= 9.58 x 10 −9 ). Using the identical statistical method and imputed genotype data as used in the main analysis, we conducted GWAS using 1,000 simulated phenotypes. We utilized down-sampled individuals (n=10,000) because permutation test using all samples (~200,000) was not computationally tractable. We simulated binary phenotypes with 1,920 cases and 8,080 controls; the same case-control ratio as in T2D GWAS in our study. For each of the 1,000 simulated phenotypes, the minimum P values ( P min ) were recorded, and the distributions of 1,000 P min were analyzed. This analysis showed that the 95-th percentile of P min is 2.87 x 10 −8 ( Extended Data Figure 4 ). We defined this value as an empirical genome-wide significance threshold at a significance level of α=0.05. 95% confidence interval was estimated by 1,000 bootstraps using the R package boot (v1.3-20). To test the potential effect of down-sampling on the P min distributions, we compared the P min distributions using all samples (n=198,137) with those using 10,000 samples. To increase computational efficiency, we restricted this analysis to imputed genotype data in chromosome 22. For this analysis, we utilized Plink2 ( https://www.cog-genomics.org/plink/2.0/ ) 58 because SAIGE requires whole genotype data to estimate relatedness even when we restrict the analysis to chromosome 22. This analysis confirmed that down-sampling does not have substantial impact on the P min distributions ( Extended Data Figure 4 ). We estimated heritability and confounding bias in our GWAS results with LDSC (v1.0.0; https://github.com/bulik/ldsc/ ) 19 using the baselineLD model (v2.1; https://data.broadinstitute.org/alkesgroup/LDSCORE/ ) 21 which includes 86 annotations, including 10 MAF- and 6 LD-related annotations that correct for bias in heritability estimates 20 , and were calculated using 481 East Asian samples in 1KG Phase3. For the analysis using LDSC, we excluded variants in the HLA region (chr6:26 Mb-34 Mb). We also calculated heritability Z-score to assess the reliability of heritability estimation. Absolute quantification of heritability estimation using GWAS results using GLMM can be biased because effective sample size could be different from the true sample size (relative quantification is not biased, and hence GWAS results using GLMM can be applied for genetic correlation analysis and S-LDSC safely). Therefore, to confirm the robustness of heritability estimation in our analysis, we also performed GWAS using generalized linear regression model (GLM). As simple GLM does not account for the bias caused by genetic relationships, we further excluded related samples (Pi-hat by > 0.187), and we analyzed genotype data with Plink2 using the same covariates as described above. Heritability estimates based on GWAS using two different methods (SAIGE vs PLINK) were comparable ( Supplementary Table 2 ). We included data in the GWAS Catalog ( https://www.ebi.ac.uk/gwas/ ) that satisfy the following criteria; (i) P in previous GWAS < 5 x 10 −8 , (ii) risk allele information is reported, (iii) outside of MHC region (Chr6: 23Mb-37Mb), and (iv) variants were analyzed in this study. When multiple variants were reported within 1Mb window, we included one variant for each disease. We considered a previous GWAS signal as replicated when the signal in the previous GWAS has the same effect direction in our GWAS. We included an independent Japanese cohort of CAD and controls who enrolled in the Osaka Acute Coronary Insufficiency Study (OACIS) 59 and the National Center for Geriatrics and Gerontology (NCGG) Biobank 60 . OACIS is a study that examined patients with myocardial infarction at 25 collaborating hospitals in Osaka, Japan, from April 1998 to April 2006. The NCGG Biobank is one of the facilities belonging to the National Center Biobank Network (NCBN; https://ncbiobank.org/en/home.php ). It has been running since 2012. The participants were recruited from NCGG hospital, which is located in Obu city, and the other nearby medical institutes. We also included 1,392 control DNAs from the Health Science Research Resources Bank (HSRRB), Osaka, Japan. Samples in NCGG were genotyped by Infinium Asian Screening Array-24 v1.0 (Illumina), and samples in OACIS were genotyped using the same platform as in BBJ samples. We extracted bi-allelic, shared variants genotyped in these studies. We excluded variants with 1) hardy Weinberg disequilibrium ( P < 1 x 10 −6 ), 2) low call rate (< 99%). We excluded samples using the following criteria: samples with low call rate (< 99%), PCA outliers, heterozygosity outliers, and sex discordant samples. After QC, 2,855 CAD cases, 15,211 controls, and 111,041 SNPs remained. After pre-phasing with Eagle (v2.3), we performed imputation by minimac4 (v1.0.0) using 1KG phase3 reference panel. Association test was conducted using SAIGE (v0.36.3) including age, sex, top 5 PCs as covariates. We tested the influence of bias using LDSC; intercept was 1.008 (S.E. = 0.014), and lambda GC was 1.053, suggesting there is no substantial bias in the association results. We also included a Japanese cohort with 2,440 female lung cancer cases and 467 female controls enrolled in the study of the National Cancer Center Hospital (NCCH). All cases are adenocarcinoma. Genotyping of rs75932146 was conducted by invader assay. Association test was conducted by logistic regression. Meta-analysis was conducted using fixed effect model via inverse-variance weighting; heterogeneity of effect size estimates was tested by using Cochran’s Q test. We searched for European GWAS whose summary statistics are publicly available and whose disease affection status were based on physician diagnosis (excluding GWAS based on self-reported phenotypes). The latter criterion was added because all cases in BBJ were diagnosed by a physician, and we wanted to prepare European GWAS of comparable phenotypes. We were able to prepare European GWAS summary statistics for 10 diseases. Summary statistics for eight diseases were downloaded from GWAS Catalog ( https://www.ebi.ac.uk/gwas/ ) and their names and their PMIDs were as follows; atrial fibrillation (30061737), breast cancer (29059683), coronary artery disease (29212778), glaucoma (29891935), ischemic stroke (29531354), prostate cancer (29892016), rheumatoid arthritis (24390342), and type 2 diabetes (30054458). Summary statistics of two diseases were downloaded from UK Biobank GWAS summary statistics at Neale Lab ( http://www.nealelab.is/uk-biobank ) and their names and their phenotype code were as follows; asthma (22127), and congestive heart failure (I50). Meta-analysis was conducted using fixed effect model via inverse-variance weighting, and tested heterogeneity in effect size estimates using Cochran’s Q test. We utilized the following variants detected in GWAS for each disease; (i) lead variants in the significantly associated loci, (ii) independent signals detected by conditional analysis, and (iii) lead variants detected in sex-specific GWAS. We defined pleiotropic association when these variants were in LD (r 2 > 0.6). We calculated r 2 using East Asian samples in the 1KG Phase3 18 by PLINK 58 . We calculated r 2 using East Asian samples (r 2 EAS ) and European samples (r 2 EUR ) in the 1KG Phase3 18 by PLINK 58 . We also identified 95% credible sets using R package corrcoverage (v1.2.1). We linked the GWAS association and the missense variant when the lead variant and the missense variant are in LD (r 2 EAS > 0.6) and the missense variant is included in 95% credible set. For the annotation of nonsynonymous variants, we used ANNOVAR ( http://annovar.openbioinformatics.org/en/latest/ ) 61 . GRCh37 (hg19) coordinates were used in this study. We also annotated GWAS variants with eQTL detected in the European population (release v7 of the GTEx project) 33 in the following conditions; (i) the lead variants of the eQTL study are in LD (r 2 EAS > 0.6 and r 2 EUR > 0.6) with GWAS variants, (ii) the missense variant is included in 95% credible set, and (iii) Q values of the lead variants in the eQTL study are less than 0.05. We estimated genetic correlations between our GWAS results by LDSC (v1.0.0) 19 using East Asian LD scores which we presented in our previous study 23 . We excluded variants in the HLA region (chr6:26 Mb-34 Mb). We analyzed 20 diseases based on two criteria; (i) heritability was reliably estimated (heritability Z-score > 2; Supplementary Table 2 ); and (ii) both of male and female patients were included. We obtained 3,158 raw human ChIP-seq data files in SRA format from the GEO database. We converted them to FASTQ format using the fastq-dump function of SRA Toolkit ( https://www.ncbi.nlm.nih.gov/sra/ ). We performed QC of sequence reads using FastQC ( https://www.bioinformatics.babraham.ac.uk/projects/fastqc/ ). We mapped these reads to the genome assembly GRCh37 using Bowtie2 (v2.2.5; http://bowtie-bio.sourceforge.net/bowtie2/manual.shtml ) with default parameters. We called peaks using MACS (v2.1; https://github.com/taoliu/MACS ) with default parameters (q < 0.01) and defined them as TF binding sites. We excluded TF binding site tracks which do not have at least one binding region on every chromosome, and 2,868 genome-wide TF binding site tracks remained ( Supplementary Table 14 ). We conducted stratified LD score regression (S-LDSC) 38 to partition heritability. For S-LDSC analysis of sex-specific GWAS of asthma, we used 220 cell-type specific annotations used in previous articles 23 , 38 . For other S-LDSC analysis, we used TF binding site tracks which were described in the previous paragraph. For all sites of TF binding, we empirically extended sites by 500 bp at the both ends for this analysis. We computed annotation-specific LD scores using the 1000 Genomes Project Phase 3 (version 5) East Asian reference haplotypes 18 . We estimated heritability enrichment of binding sites of each TF, while controlling for the merged binding sites of all TFs and the 53 categories of the full baseline model available at the authors’ website ( https://data.broadinstitute.org/alkesgroup/LDSCORE/ ). We did not use the baselineLD model (v2.1) 21 in this analysis to increase the power of detecting significant enrichment. We excluded variants in the HLA region (chr6:26 Mb-34 Mb). We analyzed 24 diseases whose heritability was reliably estimated (heritability Z-score > 2; Supplementary Table 2 ). We calculated the P value of the regression coefficient. For each trait, we calculate FDR using the Benjamini-Hochberg method. We set a significance threshold at FDR < 0.05 for this analysis. There is a complex correlation structure among 2,868 TF binding site tracks used for S-LDSC analysis. In S-LDSC, we regress GWAS chi-squared statistics on LD-scores of each TF binding site (TF LD-score), and hence we focused on correlations between TF LD-scores, not correlations between TF binding sites. We first performed PCA using all TF LD-scores. To classify them into mutually correlated TF groups, we performed k-means clustering (k=15) using the top 15 PCs. We named each cluster by the most dominant TF in each cluster ( Figure 4 ). The list of each TF binding site and its assigned cluster name was provided in Supplementary Table 14 . We then performed uniform manifold approximation and projection (UMAP) 41 using the top 15 PCs to project all TF binding sites into a two-dimensional space. UMAP was conducted using the R package umap (v.0.2.0.0). Our workflow was illustrated in Supplementary Figure 7 .

Results

We conducted a genome-wide association study (GWAS) of 42 diseases in a Japanese population, comprising 179,660 patients who participated in BBJ and 32,793 population-based controls ( Table 1 and Supplementary Table 1 ). The 42 diseases encompassed a wide-range of disease categories; 13 neoplastic diseases, five cardiovascular diseases, four allergic diseases, three infectious diseases, two autoimmune diseases, one metabolic disease, and 14 uncategorized diseases. By including patients with unrelated diagnoses into control samples, we maximized the power of our GWAS ( Methods , Extended Data Figure 1 , and Supplementary Table 1 ). We employed a generalized linear mixed model in our association analysis using SAIGE 17 . After imputing our genotypes with 1000 Genomes Project Phase 3 reference data (1KG Phase3) 18 , we tested 8,712,794 autosomal variants and 207,198 X chromosome variants for association with 42 diseases. For 35 diseases for which we have both male and female patients, we also conducted male- and female-specific GWAS. To quantify the heritability and the bias in our GWAS results, we analyzed them using linkage disequilibrium score regression (LDSC) analysis 19 ( Supplementary Table 2 ). Consistent with a recent finding in the European population 20 , heritability estimation was improved by incorporating the baselineLD model 21 which includes functional annotations, LD-dependent architectures, and minor allele frequency (MAF)-dependent architectures ( Supplementary Figure 1 and Supplementary Table 2 ). Although we observed high genomic inflation factors ( λ GC ) for some diseases (e.g. λ GC = 1.3 for T2D; Supplementary Table 2 ), LDSC analysis indicated that the majority of the inflated chi-squared statistics originated from polygenic effects rather than confounding biases (e.g. intercept = 1.01 for T2D; Supplementary Table 2 ). To confirm that our GWAS produced reasonable signals, we examined how much of the previously identified risk alleles were replicable in our GWAS results ( Extended Data Figure 2 , Table 1 , and Supplementary Table 3 ). By analyzing all diseases together, 1,219 out of 1,396 previously reported risk alleles were replicated with the same effect direction (sign test P = 1.47x10 −191 ). In East Asian populations of 1KG Phase3, MAF of non-replicated alleles are significantly lower than those of replicated alleles ( Extended Data Figure 3 ). Therefore, the replication failures might be due to insufficient statistical power. The high replicability of previous GWAS signals suggested that genetic etiologies are generally shared across populations. Considering that more than 1.5 million variants in our study are rare variants (MAF < 1%) ( Supplementary Figure 2 ), applying the conventional genome-wide significance threshold ( P < 5 x 10 −8 ), which assumes 1 million independent tests, might increase type-I errors. Therefore, to empirically estimate the appropriate P value threshold, we conducted GWAS using 1,000 random binary phenotypes and analyzed distributions of minimum P values ( P min ) for each phenotype. The 95-th percentile of P min was 2.87 x 10 −8 , and we defined this P value as an empirical genome-wide significance threshold at a significance level of α = 0.05 ( Extended Data Figure 4 ). In addition, we considered the multiple testing burden of analyzing sex-specific GWAS; each variant was tested for sex-combined, male-specific, and female-specific analyses. Therefore, we set the significance threshold for our GWAS at P = 2.87 x 10 −8 / 3 (= 9.58 x 10 −9 ), and considered P = 5 x 10 −8 as a threshold of suggestive associations. We defined a locus as a genomic region within ± 1 Mb from the lead variant, and we considered a locus as novel when it does not include any previously reported variants ( P in previous GWAS < 5.0 x 10 −8 ). In sex-combined analysis, we detected significant associations for 27 diseases at 260 autosomal loci (outside of the HLA region) and nine loci on the X chromosome ( P < 9.58 x 10 −9 ; Supplementary Table 4 and 5 ). Associations at the HLA region have been investigated in detail in a separate article 22 . We further performed conditional analyses in these 269 loci to explore associations independent of the lead variants. We detected 44 additional independent signals for 9 diseases ( P < 9.58 x 10 −9 ; Supplementary Table 6 ). The largest number of independent signals in a single locus was seven, found in the FAM84B/POU5F1B locus associated with prostate cancer. In the sex-specific analysis (male- and female-cases were analyzed separately), we detected 4 additional loci for 3 diseases which were not identified in a sex-combined analysis ( P < 9.58 x 10 −9 ; Supplementary Table 7 ). We tested heterogeneity between effect size estimates for males and females using Cochran’s Q test. This analysis found all of the four loci showed nominally significant differences in effect size estimates between sexes ( P values of heterogeneity ( P het ) < 0.003). As we will introduce below, three variants with novel suggestive associations ( P < 5.0 x 10 −8 ) passed the significance threshold after meta-analyzing with independent replication studies ( P < 9.58 x 10 −9 ). In total, we detected 320 independent significant signals in 276 loci for 27 diseases, of which 25 loci were novel ( P < 9.58 x 10 −9 ; Figure 1a , Table 1 , and Table 2 ). At three novel significant loci, the lead variants are rare variants with large effect size (MAF 2; Figure 1b ), and two of them are missense variants. To understand the characteristics of novel and known disease-associated variants in our study, we examined their allele frequencies in East Asian and European populations of 1KG Phase3. Intra-population MAF comparison showed that novel variants have significantly lower allele frequencies than known variants in European populations but not in East Asian populations ( Extended Data Figure 5 ). Trans-ethnic MAF comparison showed that both novel and known variants have higher MAF in East Asian populations than in European populations ( Figure 1c and d ). However, trans-ethnic MAF differences are more pronounced in novel variants ( Figure 1e ). These observations suggested that the high allele frequencies of disease-associated variants in our cohorts increased the statistical power to detect their significance, especially for novel variants. This highlights the importance of performing GWAS in non-European populations. We sought to refine the previously identified association signals in European GWAS. We counted the number of variants in LD with the lead variants in our GWAS and those of previous European GWAS (r 2 > 0.8 in respective populations in 1KG Phase3) ( Supplementary Table 4 ). The average number of variants in LD with the lead variants is 25.9 in European GWAS and 29.3 in our GWAS. On the other hand, the average number of variants in LD with both lead variants is 12.9. Therefore, our study successfully limited the number of potential causative variants. Since a disproportionate number of patients with T2D and coronary artery disease (CAD) were included in the controls of GWAS for other diseases, our study design might create spurious associations mirroring the effects of risk alleles of T2D and CAD. However, this possibility was ruled out by the following observations; (i) excluding all patients from control samples did not affect effect size estimates ( Extended Data Figure 1 ); (ii) risk loci detected in our GWAS for other diseases were not enriched within T2D or CAD known loci ( Supplementary Figure 3 ); and (iii) effect directions of the known protective alleles of T2D or CAD were not significantly biased to positive values in our GWAS for other diseases ( Supplementary Figure 4 ). Thus, we confirmed that our study results were not biased by having many patient samples in control groups. We next investigated the potential impact of the disease-associated variants on protein functions ( Supplementary Table 8 ). We linked the GWAS association and the missense variant when the lead variant and the missense variant are in LD (r 2 > 0.6 in East Asians of 1KG Phase3) and the missense variant is included in 95% credible set (Methods). Using these criteria, seven novel significant signals ( P < 9.58 x 10 −9 ) are linked to missense variants. Although four missense variants are not the lead variant, conditioning on these missense variants cancelled the signal of the lead variant ( Figure 2a and Supplementary Figure 5 ). Importantly, three missense variants are monomorphic in Europeans and Africans (1KG Phase3); p.R220W of ATG16L2 (rs11235604) associated with CAD; p.V326A of POT1 (rs75932146) associated with lung cancer; and p.E62G of PHLDA3 (rs192314256) associated with keloid ( Figure 2 , Extended Data Figure 6 , and Table 3 ). Considering the relevance of these findings, we additionally included two independent cohorts in a Japanese population (2,855 CAD cases and 15,211 controls; and 2,440 lung cancer cases and 467 controls). This replication study successfully confirmed the associations at ATG16L2 and POT1 loci, and fixed-effect meta-analysis improved statistical significance; the suggestive association at POT1 locus passed significance threshold ( P < 9.58 x 10 −9 ) ( Supplementary Table 9 and 10 ). Here, we discuss each of the three East Asian-specific missense variants in detail. First, ATG16L2 is an autophagy-related gene highly expressed in immune cells, and previous studies reported that p.R220W of ATG16L2 is also associated with immune related traits; serum level of non-albumin protein in a Japanese population 23 and Crohn's disease in a Chinese population 24 . Previous GWAS for CAD in European populations did not detect significant associations at ATG16L2 locus 25 ( Figure 2a ), suggesting that p.R220W of ATG16L2, absent in Europeans, may be the causal variant. Therefore, dysregulated autophagy in immune cells might have an important role in CAD. Second, POT1 is a member of the telombin family and this protein binds to telomeres, regulating telomere length. Missense variants of POT1 have been described as being responsible for several familial cancers 26 - 28 . In addition, our study showed that p.V326A of POT1 is also positively associated with the risk of five other neoplastic diseases ( P < 0.05; Extended Data Figure 7 ). These findings suggest this variant might increase the risk of neoplastic diseases in general. p.V326A of POT1 is more strongly associated with lung cancer in females than males; OR for female is 2.29 and OR for male is 1.26 ( P het = 7.7x10 −4 ) ( Figure 2b and Supplementary Table 7 ). We sought to figure out whether the sex-dependent effect can be explained by other factors, and conducted an association test stratified by histological and smoking status ( Supplementary Table 10 ). However, we could not reach a definitive conclusion due to limited statistical power, and hence further large-scale studies will be required to answer this question. Together with a known association at the TERT locus ( Supplementary Table 4 ), we provide additional evidence that telomere dysregulation is pathogenic for lung cancer. Third, p.E62G of PHLDA3 is predicted to have a deleterious effect to its protein function (SIFT score 29 =0; CADD score 30 =33), and we detected a large effect size for keloid (odds ratio = 9.56; 95% CI 5.91-15.45). We confirmed that genotyping of rs192314256 (p.E62G of PHLDA3 ) was not biased by batches of genotyping experiments or geographic areas ( Supplementary Figure 6 ). PHLDA3 is known to be a suppressor of AKT 31 , and upregulated AKT signaling pathway is related to increased collagen production from dermal fibroblasts 32 . Therefore, damaged PHLDA3 may activate the AKT pathway, promoting the development of keloid. Together, our study successfully identified novel potential causal genes which would be hard to be discovered by GWAS in European populations due to restrictive European allele frequencies. We also investigated the potential impacts of the disease-associated variants on the mRNA levels using the GTEx database of expression quantitative trait loci (eQTL) 33 . Since the eQTL data are generated in European populations, we could not apply formal colocalization tests 34 , 35 which assume the same LD-structures between GWAS and eQTL studies. Therefore, we linked the GWAS association and the eQTL variant when the GWAS lead variant and the eQTL variant are in LD (r 2 > 0.6 both in East Asian and European populations of 1KG Phase3) and the eQTL variant is included in 95% credible set. We found that seven novel significant signals ( P < 9.58 x 10 −9 ) and five novel suggestive signals ( P < 5 x 10 −8 ) can be explained by at least one eQTL variant ( Supplementary Table 11 ). Among them, the eQTL signals for ATP2B1 which were linked to a novel, suggestive variant of cerebral aneurysm (rs11105352; P = 1.22 x 10 −8 ) is highly specific to arterial tissues ( Figure 3 ). Since the loss of ATP2B1 in vascular smooth muscle cells induced blood pressure elevation in mice 36 , decreased expression of ATP2B1 in arteries might induce hypertension, which leads to increased risk of cerebral aneurysm. Replication analysis in the same population is a critical part of genetic studies. Although we included two independent replication studies for CAD and lung cancer in a Japanese population, we were not able to prepare replication cohorts in a Japanese population for other diseases. Therefore, we conducted replication studies using previous European GWAS results. We utilized publicly available GWAS summary statistics of European populations for 10 diseases (asthma, atrial fibrillation, breast cancer, CAD, congestive heart failure, glaucoma, ischemic stroke, prostate cancer, rheumatoid arthritis, and T2D; see Methods for selection of diseases), and tested for consistency in direction of effect. For these 10 diseases, our GWAS detected suggestive associations at 218 known and 19 novel loci ( P < 5 x 10 −8 ); among them, statistics of European GWAS were available at 149 known and 15 novel loci. We first conducted replication analysis at the known loci. We restricted this analysis to 112 known loci with significant associations also in European GWAS ( P < 5 x 10 −8 ) to exclude loci where the European GWAS had insufficient power. Effect directions are consistent between BBJ- and European-GWAS at 109 out of 112 loci; but opposite at 3 loci ( Extended Data Figure 8 and Supplementary Table 12 ). These three replication failures are probably due to differences in LD structure between populations ( Extended Data Figure 8 ). We then conducted replication analysis at the novel loci. Among 15 novel variants, 12 were replicated with the same effect direction ( Supplementary Table 13 ). Meta-analysis using fixed-effect model increased the level of significance in six of them; and two suggestive novel variants passed significance threshold ( P < 9.58 x 10 −9 ) (rs2277339 and rs17105012 associated with T2D; Table 2 and Supplementary Table 13 ). Among the three variants that failed replication, rs13227841 is a missense variant originally identified as a potential causal variant at this locus (p.W78R of WBSCR28 ; Supplementary Table 8 ), which suggests that variants in LD with rs13227841, not rs13227841 itself, may be responsible for the observed associations. The other replication failures might be due to different LD-structures or the absence of the causal variants in European populations. Further efforts to conduct a replication analysis in a Japanese population will be required to confirm the associations which we failed to replicate in these European studies. To understand differences in the genetic risks between males and females, we assessed genetic correlations using LDSC 37 between the results of sex-specific GWAS for the 20 diseases (see Methods for selection of diseases). Although most correlations are close to one, the correlation of asthma was significantly smaller than one (genetic correlation = 0.63 (S.E. = 0.12) and P = 2.2 x 10 −3 < 0.05/20; Extended Data Figure 9 ). This finding suggested that genetic risks of asthma might be different between males and females. To explore the biological mechanism underlying this finding, we estimated the enrichment of the heritability of male or female asthma in the 220 cell-type specific regulatory regions using stratified LD-score regression (S-LDSC) 38 . We found significant enrichments for either male or female asthma in three annotations; Th0, Th1, and colonic mucosa ( P < 0.05/220; Extended Data Figure 9 ). Among them, the colonic mucosa annotation showed significant heterogeneity in the enrichment of heritability ( P het = 0.006 < 0.05/3). Recent studies suggested that host-microbiome interactions at intestinal mucosa (gut-lung axis) have important roles in the development of asthma 39 , 40 , and our study suggested that the gut-lung axis might have sex-dependent roles in asthma. Considering their marginal significance, a replication study will be required to confirm these findings. To acquire more insights to disease biology, we estimated the heritability enrichments in the binding sites of a variety of transcription factors (TFs) using S-LDSC. We included TF binding sites defined by 2,868 publicly available chromatin immunoprecipitation sequencing (ChIP-seq) datasets for 410 unique TFs ( Supplementary Table 14 ). To make mutually comparable data, we began our analysis from the raw sequencing data, and defined TF binding sites using a uniform protocol (Methods). Using LD-scores of all TF binding sites, we grouped them into 15 clusters (cluster name was defined by the most dominant TF), and performed uniform manifold approximation and projection (UMAP) 41 to project all TF binding sites into a two-dimensional space ( Methods ; Figure 4a and Supplementary Figure 7 ). To scale the performance of this analysis, we first analyzed previously reported GWAS for red blood cell-related traits 23 where the critical role of GATA1 was supported by multiple pieces of evidence 42 - 46 , and we successfully recapitulated this biology ( Figure 4b ). We then applied this analysis to our 24 GWAS results (see Methods for selection of diseases), and detected 378 significant enrichments for nine diseases (FDR < 0.05) ( Figure 4c , Extended Data Figure 10 , and Supplementary Table 15 ). Biologically plausible TFs were highlighted by this analysis; RELA , a subunit of NF-κB, for atopic dermatitis, rheumatoid arthritis (RA), and Graves’ disease; sex hormone receptors ( AR and ESR1 ) for prostate cancer; and FOXA2 , which regulates insulin secretion in pancreatic beta-cells 47 , for T2D ( Figure 4c ). This analysis also suggested that NKX3-1 , a prostate-specific homeobox gene, has an important role in the biology of prostate cancer ( Figure 4c ). In addition to this polygenic analysis, the importance of NKX3-1 was also suggested by the regional analysis integrating eQTL databases; the risk allele of prostate cancer at the NKX3-1 locus (rs4872174-C) was suggested to decrease the expression of NKX3-1 ( Supplementary Table 11 ). Consistently, loss of NKX3-1 expression in human prostate cancers was reported to be correlated with tumor progression 48 . Together, our results confirmed and expanded our current understanding of complex traits in the context of TF activity.

Discussion

Our study demonstrated the advantages of conducting genetic studies in non-European populations. Typically, LD acts as a major hurdle limiting the identification of causal variants in GWAS. However, jointly analyzing GWAS results from populations with different LD structures can narrow down causal variants 12 . Indeed, when we consider variants in LD with a lead variant as candidate causal variants (r 2 > 0.8), our study successfully reduced the number of candidate causal variants at 68 loci which were originally discovered in previous European GWAS ( Supplementary Table 4 ). In addition, some novel variants in our study have been missed in larger GWAS in European populations due to restrictive European allele frequencies. Therefore, diversifying the ethnicity of participants is important not only for the equality of genetic findings but also for the discovery of novel disease etiology. Although previous studies already reported important roles of TFs in the etiology of complex traits 49 - 51 , our TF enrichment analysis has two distinguishing features from previous studies. One feature is the comprehensiveness; we included 2,868 TF annotations, more than those used in most previous studies. The second feature is the method of the enrichment test; we utilized S-LDSC, whereas most previous studies utilized naïve enrichment tests using genome-wide significant variants. S-LDSC evaluates enrichment of GWAS signals irrespective of significance, and it is robust to the biases coming from the overlapping annotations. Therefore, by incorporating a comprehensive catalog of TF annotations with a sophisticated method to test heritability enrichment, we provided evidence of TF importance in complex diseases from a polygenic angle. The critical limitation of this study is insufficient replication analyses to validate novel signals. Among 25 novel loci ( P < 9.58 x 10 −9 ), we were able to prepare East-Asian replication datasets for only two of them; p.R220W of ATG16L2 associated with CAD and p.V326A of POT1 associated with lung cancer. To supplement this insufficiency, we utilized European GWAS results when data was available; we tested replicability of eight novel signals ( P < 9.58 x 10 −9 ) and observed evidence of heterogeneity in effect size estimates for three of them ( P het < 0.05; Supplementary Table 13 ). This may be the case for several reasons; the locus might possess different LD structures between populations and the variant might tag the causal variant only in East Asian populations (as illustrated in Extended Data Figure 8 ); effect sizes might be truly different between populations; or they might be false positives. Therefore, until further replication studies in East-Asian populations are conducted, we need to be cautious about the validity of these putatively novel variants since we were not able to provide evidence of replicability. In summary, we conducted a large-scale GWAS of 42 diseases in a non-European population and provided rich public resources for genetic studies. Our study provided multiple insights into the etiology of complex traits by integrating annotations of missense variants, eQTL variants, and transcription factor binding site tracks. Currently, genetic studies are overwhelmed by European-descent samples, making the clinical translation of genetic findings far more beneficial to European individuals than other populations 1 . Our study contributed to broaden the population diversity in genetic studies and should potentially mitigate the problems originating from this imbalance.

Introduction

Currently, large-scale genetic studies are dominated by European-descent samples, and fail to capture the level of diversity that exists globally 1 - 5 . Due to differential genetic architectures, transferability of genetic findings between populations is generally limited. Therefore, this imbalance poses a limitation in our understanding of the genetic architecture of complex diseases in non-European populations. Moreover, this imbalance could result in unequal benefits of precision medicine, as polygenic risk sores (PRS) based on large-scale genetic studies in European populations have high predictive power of clinical outcomes in European samples 6 - 10 but poor predictive power in non-European samples 1 , 11 . Therefore, increasing the ethnic diversity of participants is an essential direction of genetic studies for the equality of genetic findings. In addition, diversifying the ethnicity of participants is important for the discovery of novel disease etiology 12 . Even in large-scale European studies, causal variants might be missed if they have low frequencies or are monomorphic in European populations; such examples include p.E508K of HNF1A identified in Latino populations 13 and p.R684* of TBC1D4 identified in a Greenlandic population 14 , both associated with type 2 diabetes (T2D). Therefore, differences in allele frequencies across populations can be an advantage for discovering genetic signals which were failed to be identified in European populations. Here we report a GWAS of 42 common diseases in the BioBank Japan Project (BBJ) 15 , 16 , one of the largest non-European biobanks consisting of around 200,000 individuals. We provide detailed discussion of the biology of these diseases using multiple genomic annotations. We also examined inter-sex differences in genetic signals. Moreover, by incorporating previous genetic findings, we discussed the extent to which genetic signals are shared across populations while also investigating East Asian-specific genetic signals. Our study provided multiple insights into the etiology of complex traits, and highlighted the importance of conducing genetic studies in non-European populations.

Extended Data

a, Study designs in this GWAS. Study design 1 (top) was used in the main analysis. An example of study design 1 is provided; in GWAS of disease 3, we included all other patients (except those have related diseases) into control group. The definition of related diseases is provided in Supplementary Table 1 . Study design 2 (bottom) was used to discuss the appropriateness of study design selection. b, Effect size estimates and S.E. at the 309 autosomal disease-associated variants detected in sex-combined analysis ( P < 5 x 10 −8 ). We compared the effect size estimates in study design 1 with those in study design 2. Heterogeneity between two studies was tested using Cochran’s Q test. The identity line is shown in blue. The red dot (rs373205748 associated with arrhythmia) indicates a variant with significant heterogeneity in effect size estimates between two study designs ( P = 0.00012 < 0.05/309). We compared effect sizes reported in the previous GWAS with those in this GWAS. Effect size and S.E. are shown. The identity line is shown in blue. The sample size of GWAS is provided in Table 1 . We utilized a generalized linear mixed model in our GWAS. We first compared effect sizes reported in the previous GWAS with those in our GWAS ( Supplementary Table 3 and Extended Data Figure 2 ); 1,219 out of 1,396 previously reported risk alleles were replicated with the same effect direction (177 alleles were not replicated). We compared MAF of replicated variants (n=1,219) and MAF of not replicated variants (n=177). Mann-Whitney U test P value is provided (two-sided test). Using 1,000 simulated binary phenotypes with down-sampled samples (n=10,000), we conducted GWAS utilizing the same strategy as used in the main analysis. a, The distribution of minimum P values in each phenotype ( P min ). The 95-th percentile of P min was 2.87 x 10 -8 . The 95% confidence interval was estimated by 1,000 bootstraps. b, The distributions of P min using all samples (n=198,137) and those using 10,000 samples. To increase computational efficiency, we restricted this analysis to imputed genotype data in chromosome 22. For this analysis in b , we utilized Plink2. MAF comparison at disease-associated variants at novel (n=41) and known loci (n=153) with suggestive significance ( P < 5 x 10 −8 ) ( a , East Asian populations; b , European populations in 1KG phase3). For known loci, we restricted this analysis to loci where the closest reported variants were discovered by GWAS in European populations. Mann-Whitney U test P value is provided (two-sided test). A regional association plot for keloid (812 cases vs 211,641 controls) at the PHLDA3 region is provided. We utilized a generalized linear mixed model in our GWAS. Effect size and S.E. are provided for neoplastic diseases ( a ) and non-neoplastic diseases ( b ). The sample size of GWAS is provided in Table 1 . We utilized a generalized linear mixed model in our GWAS. a, Schematic explanations how we compared statistics between BBJ-GWAS and GWAS conducted in European populations (EUR-GWAS). We utilized two inclusion criteria of known loci: (i) EUR-GWAS has significant associations ( P 0.4 in European samples in 1KG phase3). The first criterion was added to exclude loci where EUR-GWAS has insufficient power (112 known loci remained after applying the first criterion). The second criterion was added because EUR-GWAS statistics at the BBJ-lead variant is not representing those at the EUR-lead variant when they are not in LD. b, effect sizes of BBJ- and EUR-GWAS at the BBJ-lead variants. All variants which passed the first criterion were used (n=112). Variants which passed the second criterion are shown in red (n=65). Since two variants have extremely large effect size, we provided two plots in different scales. The three variants with the opposite effect directions are marked by large dots, and their details are also provided. c, Regional association of T2D around rs12031188. Variants in LD (r 2 > 0.4) with BBJ-lead variant (rs12031188) but not with EUR-lead variant are shown in red; Variants in LD (r 2 > 0.4) with both lead variants are shown in blue. East Asians and Europeans in 1KG phase3 were used for LD calculation of the BBJ- and the EUR-lead variant, respectively. a. Genetic correlations between male- and female-specific GWAS. Estimates of genetic correlation and standard errors are provided. *: genetic correlation was significantly different from one (two-sided t test P = 2.2 x 10 −3 < 0.05/20). b. The results of S-LDSC analysis based on sex-specific GWAS of asthma using 220 cell-type specific annotations. Significant annotations in either male or female asthma were shown ( P < 0.05/220). Heterogeneity was tested by Cochran’s Q test, and its P values ( P het ) were also provided. Black dashed line indicates P value = 0.05/220; grey dashed line indicates P value = 0.05. The results of S-LDSC were plotted on the UMAP space. The significant results (FDR<0.05) were highlighted by cluster-specific colors (the same colors as used in Figure 4 ). The names of the top five most significant TFs were also shown on the plot. The results of diseases with less than five significant TF binding site tracks were shown.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-09-27T09:11:36.575535+00:00
unpaywall
last seen: 2026-09-28T06:20:55.906821+00:00