SEAD: an augmented reference panel with 22,134 haplotypes boosts the rare variants imputation and GWAS analysis in Asian population

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

The SEAD reference panel, with 22,134 haplotypes, improves rare variant imputation accuracy in Asian populations and identified a novel association between rare variants near the SNTG1 gene and hip bone mineral density.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-17 · read from full text

The study developed the South and East Asian Reference Database (SEAD) by integrating whole-genome sequencing data from 11,067 individuals across 17 Asian countries, creating a reference panel of 22,134 haplotypes and 80,367,720 variants (excluding singleton variants). The authors evaluated genotype imputation performance by comparing concordance rates across global populations using Human Genome Diversity Project (HGDP) data and then assessed rare/low-frequency variant imputation in Asian populations, finding SEAD to improve imputation accuracy and the number of well-imputed sites, particularly for low-frequency and rare variants. They additionally applied the augmented SEAD panel to impute and run discovery/replication GWAS for hip and femoral neck bone mineral density in 5,369 Chinese samples from the Westlake BioBank, reporting an association signal near SNTG1 for hip BMD and suggestive results for femoral neck BMD, with a preliminary functional experiment implicating SNTG1 regulation of preosteoblast proliferation and differentiation. A major caveat is that the preprint notes the work is not peer reviewed. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Here, we present the South and East Asian Reference Database (SEAD) reference panel (https://imputationserver.westlake.edu.cn/), which comprises whole genome sequencing data from 11,067 individuals across 17 countries in Asia. The SEAD panel, which excludes singleton variants, consists of 22,134 haplotypes and 80,367,720 variants. Firstly, we assessed the concordance rate in global populations using HGDP datasets, notably, the SEAD panel showed advantage in East Asia, Central and South Asia, and Oceania populations. When imputing the disease-associated variants of Asian population, the SEAD panel displayed a distinct preponderance in imputing low-frequency and rare variants. In imputation of Chinese population, the SEAD panel imputed a larger number of well-imputed sites across all minor allele frequency (MAF) bins. Additionally, the SEAD panel exhibited higher imputation accuracy for shared sites in all MAF bins. Finally, we applied the augmented SEAD panel to conduct a discovery and replication genome-wide association study (GWAS) for hip and femoral neck (FN) bone mineral density (BMD) traits within the 5,369 Westlake BioBank for Chinese (WBBC) samples. The single-variant test suggests that rare variants near SNTG1 gene are associated with hip BMD (rs60103302, MAF = 0.0091, P = 4.79×10− 8). The spatial clustering analysis also suggests the association of this gene (Pslide_window=1.08×10− 8, Pgene_centric=4.72×10− 8). The gene and variants achieved a suggestive level for FN BMD. This gene was not reported previously, and the preliminary experiment demonstrated that the identified rare variant can upregulate the SNTG1 expression, which in turn inhibits the proliferation and differentiation of preosteoblast.
Full text 155,018 characters · extracted from preprint-html · click to expand
SEAD: an augmented reference panel with 22,134 haplotypes boosts the rare variants imputation and GWAS analysis in Asian population | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article SEAD: an augmented reference panel with 22,134 haplotypes boosts the rare variants imputation and GWAS analysis in Asian population Hou-Feng Zheng, Meng-yuan Yang, Jia-Dong Zhong, Xin Li, Wei-Yang Bai, and 32 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3680930/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 30 Dec, 2024 Read the published version in Nature Communications → Version 1 posted You are reading this latest preprint version Abstract Here, we present the South and East Asian Reference Database (SEAD) reference panel ( https://imputationserver.westlake.edu.cn/ ), which comprises whole genome sequencing data from 11,067 individuals across 17 countries in Asia. The SEAD panel, which excludes singleton variants, consists of 22,134 haplotypes and 80,367,720 variants. Firstly, we assessed the concordance rate in global populations using HGDP datasets, notably, the SEAD panel showed advantage in East Asia, Central and South Asia, and Oceania populations. When imputing the disease-associated variants of Asian population, the SEAD panel displayed a distinct preponderance in imputing low-frequency and rare variants. In imputation of Chinese population, the SEAD panel imputed a larger number of well-imputed sites across all minor allele frequency (MAF) bins. Additionally, the SEAD panel exhibited higher imputation accuracy for shared sites in all MAF bins. Finally, we applied the augmented SEAD panel to conduct a discovery and replication genome-wide association study (GWAS) for hip and femoral neck (FN) bone mineral density (BMD) traits within the 5,369 Westlake BioBank for Chinese (WBBC) samples. The single-variant test suggests that rare variants near SNTG1 gene are associated with hip BMD (rs60103302, MAF = 0.0091, P = 4.79×10 − 8 ). The spatial clustering analysis also suggests the association of this gene ( P slide_window =1.08×10 − 8 , P gene_centric =4.72×10 − 8 ). The gene and variants achieved a suggestive level for FN BMD. This gene was not reported previously, and the preliminary experiment demonstrated that the identified rare variant can upregulate the SNTG1 expression, which in turn inhibits the proliferation and differentiation of preosteoblast. Biological sciences/Computational biology and bioinformatics/Genome informatics Biological sciences/Genetics/Genetic association study/Genome-wide association studies reference panel GWAS rare variants hip BMD FN BMD Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Genotype imputation is an effective way to save the cost of sequencing and has become a key step in genome-wide association studies (GWAS). It can increase the chance of identifying likely causal variants ( 1 ), facilitate meta-analysis ( 2 ), and aid in finding pleiotropic effects of risk variants ( 3 , 4 ). Factors such as the haplotype size, ancestry diversity and sequencing depth can affect the imputation accuracy ( 5 ). Generally, a diverse reference panel can improve accuracy in genetically diverse populations ( 6 ), while an ancestry-specific reference panel can benefit the corresponding population ( 7 , 8 ). In our pervious study, we found that the imputation accuracy for European population benefit from an increasing in haplotype size and population diversity of the reference panel, while the accuracy for the Han Chinese population reached its peak with a fraction of an extra diverse sample (8–21%) in the reference panel ( 9 ). Given that imputation accuracy directly impacts the credibility of subsequent analysis, it is crucial to select an appropriate reference panel before imputation. The 1000 Genomes Project (1kGP) is one of the most renowned and widely used reference panel for genotype imputation, comprising 2,504 individuals from 26 populations ( 10 ). The latest update to 1kGP phase 3 sequenced 3,202 samples (including an additional 698 samples related to the previous participants) and identified 70,594,286 single nucleotide polymorphisms (SNPs) with 30× depth. As many rare and low-frequency variants associated with diseases tend to be population-specific ( 10 ), an increasing number of human genome projects are focusing on specific populations to provide population-specific reference panels. However, most of the whole-genome sequencing (WGS) efforts were carried out in European-descent individuals, such as the GoNL project (the Genome of the Netherlands project, n = 769) ( 11 ), UK10K project (~ 4,000 WGS and ~ 6,000 whole exome sequencing samples in UK)( 12 , 13 ) and the TOPMed program (Trans-Omics for Precision Medicine, n = ~ 138,000) ( 14 ). These efforts have enabled the provision of a large-scale combined reference panel for the European population, such as the HRC panel (Haplotype Reference Consortium, n = 32,470) ( 15 ). The update of analytical methods and the emergence of various imputation reference panels have improved the imputation quality for European population. However, one of the complicated legacies of human genome project after 20 years’ development is that lacking of diversity might hinder the promise of genome science ( 16 ). As the largest continent, Asia accounts for 59.5% of the worldwide population ( https://www.worldometers.info/world-population/ ). Fortunately, the Asian ancestry-dominant populations have been recently sequenced and analyzed to understand the genetic basis of the Asian, such as the Japanese ( 17 ), Korean ( 18 ), Singaporean ( 19 ), Indian ( 20 ) and Chinese ( 21 – 24 ). Among these efforts, the Westlake BioBank for Chinese (WBBC) project ( 25 ) was initiated in 2017 by our team. Up to now, we have whole-genome sequenced 4480 Chinese samples, covering 29 out of 34 administrative divisions of China ( 23 , 26 ). Rare variants, which may have low level of pairwise linkage disequilibrium with common variants, can potentially result in significant functional consequences ( 27 ). Next-generation sequencing, with adequate coverage, enables the accurate detection of rare variants ( 28 ). Rare variants are harder to impute than common variants because rare variants only appear a few times in the reference panel. Further, the evaluation of rare variant imputation in Asian populations is hindered by the limited sample size of currently published studies, which typically comprise only a few thousand individuals. In this study, we integrated the WGS data from multiple sources, including Singapore SG10K pilot project (13.7×, 4563 samples), GenomeAsia pilot project (36×, 1031 samples), WBBC pilot project (13.9×, 4480 samples), and the high-coverage 1kGP-Asian phase 3 (30×, 993 East and South Asian from 3202 samples). These data were combined to create the South and East Asian Reference Database (SEAD), encompassing a total of 11,067 individuals, making it a large-scale reference panel for Asian populations. We compared the concordance rate of the genotypes imputed from the SEAD panel across populations around the world. We then focused on evaluating the imputation quality of this panel to impute rare variants in Asian populations. Finally, we applied this augmented panel to the WBBC genome-wide genotyping data to explore its role in imputing possible likely causal rare variants for bone-related traits. Methods Quality control of the four Asian reference panels We collected Asian haplotypes from the SG10K pilot project ( 19 ), the GenomeAsia pilot project (GAsP) ( 20 ), the newly released high-coverage 1000 Genomes project (1kGP) ( 29 ) and Westlake Biobank for Chinese pilot project (WBBC) ( 23 ). The SG10K dataset comprises 4,810 whole-genome sequencings (WGS) of Singaporean Chinese, Malays, and Indians, with an average depth of 13.7×. The GAsP dataset includes 1,163 WGS samples with an average depth of 36×, mainly covering Indians, Korean, Pakistanis, etc. These two datasets were obtained by applying to the consortium. Since the SG10K and GAsP datasets were based on Genome Reference Consortium Human build 37 (GRCh37), we employed liftover ( 30 ) to update the genome assembly version from GRCh37 to GRCh38. For the GAsP dataset, we transformed multi-allelic variants into multiple bi-allelic variants using Bcftools ( 31 ) and performed haplotype re-phasing using SHAPEIT v2 with parameters including windows of size 0.5, 200 states, and an effective size of 14,269. The high-coverage 1kGP dataset with an average depth of 30×, based on the GRCh38 assembly, were downloaded from the https://www.internationalgenome.org/data-portal/data-collection/30x-grch38 . Then we extracted East Asian (EAS) and South Asian (SAS) dataset from 1kGP and removed sites with minor allele count (MAC) equals zero. The WBBC dataset, which were generated by our own, included 4,480 Chinese samples with an average sequencing depth of 13.9×, supporting both GRCh37 and GRCh38. We identified relative pairs using KING ( 32 ) and excluded sample pairs with estimated kinship coefficients restricted to the range of 0.177–0.354. Ultimately, we retained 4,480 samples from WBBC 4,563 samples from SG10K, 1,031 samples from GAsP, and 993 samples from the EAS/SAS subset of the 1kGP dataset (1kGP-Asian) (Supplementary Table 1). Evaluation of genotype concordance in HGDP dataset. The 929 genomes from 55 diverse human populations from Human Genome Diversity Project (HGDP) ( 33 ) dataset was utilized to estimate the genotype concordance of imputation across global populations. The HGDP dataset includes 104 samples from African populations, 61 samples from American populations, 197 samples from Central and South Asia populations, 223 samples from East Asia populations, 155 samples from European populations, 161 samples from Middle East populations, and 28 samples from Oceania populations (Supplementary Table 2). We obtained the HGDP data from the following source: ftp://ngs.sanger.ac.uk/production/hgdp . The phased autosome VCF file was split into 22 chromosomes, which were considered as the true set. To compare the imputation concordance of the reference panels (WBBC, SG10K, 1kGP, GAsP and SEAD), we extracted the genetic variants from the 929 HGDP genome data with Axiom array loci (820,967 variants) to create pseudo array, which was then used as the test set for imputation. We downloaded Axiom array SNP list from https://www.ukbiobank.ac.uk/enable-your-research/about-our-data/genetic-data . We imputed the genotypes from the 5 panels, then compared the concordance between the imputed genotypes and true set. We also extracted Asian Screening Array sites (659,184 variants) in Central and South Asia populations and East Asia populations (420 samples in total) as Asian array for comparison purpose. Before imputation, variants in the pseudo arrays with MAF less than 1% and samples with calling rate below 95% were excluded. Here, the singletons were excluded in the 5 reference panels, consists of 35,616,674 variants for WBBC, 50,246,865 variants for SG10K, 70,594,286 variants for 1kGP (not just Asian), 20,203,158 variants for GAsP, 80,367,720 variants for SEAD. We calculated the non-reference genotype concordance rate, the non-reference heterozygote concordance rate and non-reference homozygote concordance rate between true genotypes and imputed genotypes pseudo arrays for each individual (Supplementary Fig. 1), as what we did before ( 23 ). The average concordance rate in single population was calculated as the mean of all the individuals’ concordance rate in this population. We quantified the ratio of NR concordance rate of the SEAD panel to other four panels by using the following formula, respectively. $$\text{l}\text{o}\text{g}\text{F}\text{C}=\text{l}\text{o}\text{g}\left(\frac{\text{N}\text{R} \text{c}\text{o}\text{n}\text{c}\text{o}\text{r}\text{d}\text{a}\text{n}\text{c}\text{e} \text{r}\text{a}\text{t}\text{e}1}{\text{N}\text{R} \text{c}\text{o}\text{n}\text{c}\text{o}\text{r}\text{d}\text{a}\text{n}\text{c}\text{e} \text{r}\text{a}\text{t}\text{e}2}\right)$$ All these analyses were performed on chromosome 2. Quality control of the Chinese genotyping array data Besides the whole genome sequencing data, the WBBC study also genotyped 6,080 individuals with the high-density Illumina Asian Screening Array (ASA, based on GRCh37), resulting in the identification of a total of 659,184 SNPs ( 23 , 25 ). Of which, 184 samples underwent both whole-genome sequencing and array genotyping. As part of quality control, we employed GCTA ( 34 ) to calculate the pairwise genetic relationship matrix using common variants and remove samples with a coefficient > 0.025 (Supplementary Table 3). We then excluded samples with missing call rates ≥ 2% and excluded SNPs with missing call rates ≥ 5%, MAF < 1% and Hardy-Weinberg equilibrium at P < 1 × 10 − 6 by using PLINK1.9 ( 35 ). The genotype assembly version of the WBBC array data was updated from GRCh37 to GRCh38. Finally, we retained a total of 470,242 variants and 5,679 samples in the WBBC array dataset. The data were phased using SHAPEIT with a window size of 0.5, 200 states, and an effective size of 14,269. Evaluation of reference panel for imputation in the Chinese population The procedure to assess the imputation performance in the Chinese population using the five reference panels (WBBC, SG10K, 1kGP, GAsP, and SEAD) was similar to the evaluation in HGDP dataset. As we have thousands of ASA array samples, we could assess the imputation performance for low-frequency and rare variants. We grouped the variants into seven Minor Allele Frequency (MAF) bins: < 0.1%, 0.1%-0.3%, 0.3%-0.5%, 0.5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%. Genotype imputation was performed using Minimac4 ( 36 , 37 ), with a chunk length of 20 Mb and a chunk overlap of 4 Mb. We used R-square as the estimated value, defined as the squared correlation between imputed genotypes and observed genotypes, produced by Minimac4. Sites with an R-square value greater than 0.8 were considered well-imputed variants. We counted the number of well-imputed variants and calculated the average R 2 within each MAF bin. To compare the imputation accuracy among different reference panels, we extracted 8,525,136 variants that were shared by all five panels and evaluated the imputation performance again. In the WBBC ASA array data, 184 samples underwent both whole-genome sequencing and array genotyping. We then utilized these samples to evaluate the imputation concordance rate. We randomly masked 2,600 variants on chromosome 2 from the array data and subsequently imputed these SNPs using the five Asian panels. We extracted the imputed 2,600 loci as the test set, while the original genotype of the 2,600 loci from the array data served as the true set. Employing the methods from the previous section, we calculated both the non-reference homozygous and heterozygous concordance rates. Genome wide association study of BMD traits As we mentioned above, the WBBC pilot study genotyped 6,080 individuals with ASA array, and a bunch of bone-related phenotypes were collected within WBBC ( 25 ). After quality control, a total of 470,242 variants and 5,679 samples were retained. Here, we took two correlated bone mineral density (BMD) traits (total hip and femoral neck BMD) as example, we removed individuals with missing phenotype data, and excluded outliers using the mean ± 4 standard deviations (n = 5,369 left). We then grouped the ASA samples according to the assessment place of collection: 2,332 samples recruited from Jiangxi province (samples were mainly from Southern China) were taken as discovery cohort (cohort 1), and 3,037 samples recruited from Shandong province (samples were mainly from Northern China) were taken as replication cohort (cohort 2), and vice versa . The baseline statistics of the study samples were shown in Supplementary Table 4. Comprehensive GWAS analyses were conducted with the imputed genotypes from the augmented SEAD reference panel. We kept the imputed variants with Rsq > 0.5 in the GWAS analysis. The imputation boosted the analyzed genetic variants from 470,242 to 19,235,129, with 3,195,758 variants with MAF between 0.1% and 1%. The BMD phenotypes (total hip and femoral neck BMD) in this study was analyzed as continuous outcomes, adjusting for ‘sex’, ‘age’, ‘BMI’, ‘geographical region’, and ‘10 principal components (PCs)’. The parameters applied in PLINK were ‘--geno’ of 0.05, ‘--mind’ of 0.05, ‘--hwe’ of 1 × 10 − 6 and ‘--maf’ of 0.001. Furthermore, we performed the meta-analysis of GWAS summary statistics in cohort 1 and cohort 2 using METAL software ( 38 ). Rare variants (0.001 < MAF < 0.01) were considered as candidates if they met the following criteria: a P -value less than 1e-04 in either cohort, a P -value less than 0.05 in the other cohort, and a P -value less than 1e-04 in the meta-analysis. Spatial clustering analysis We also performed the STAAR (variant-set test for association using annotation information) ( 39 ) framework to identify rare variants that would be associated with BMD traits (Hip and FN BMD). The STAAR pipeline facilitates rare variants spatial clustering analyses, including sliding window-based analysis and gene-centric analysis ( 39 ). In each test, we included variants with a minor allele frequency (MAF) of 0.1–1%. We employed fixed-size sliding window analysis, allowing for the systematic examination of distinct genomic regions by moving a 2kb window every 1kb across the genome, a total of 1,392,296 windows was divided. The gene-centric analysis method is the variant-set test for association using annotation information for a gene (STAAR-O), enhancing the power of rare variant association test by incorporating multiple variant functional annotations. For the gene-centric non-coding variants, aggregation was performed based on 7 categories: downstream, enhancer variants overlaid with Cap Analysis of Gene Expression (CAGE) sites, promoter CAGE, enhancer variants overlaid with DNase hypersensitivity (DHS), promoter DHS, upstream, and UTR2, and clustering regions with rare variations less than 2 were excluded. Both the sliding window and gene-centric methods were performed separately on the cohort 1 and cohort 2. Cell culture The HEK293T cells were cultured at 37°C in a humidified atmosphere with 5% CO2 using DMEM basal media supplemented with 10% fetal bovine serum. The MC3T3-E1 subclone 14 cells were cultured with similar condition with ascorbate-free αMEM basal media supplemented with 10% fetal bovine serum. The media were refreshed every 3–4 days, and cells were passaged when they reached 90% confluence. To initiate differentiation, the media were replaced with a calcification-inducing medium containing 50 ng/L vitamin C (VC), 10 nM dexamethasone, and 10 mM beta-glycerophosphate. Luciferase Reporter assays Dual-Luciferase reporter assays were conducted following previously described methods ( 40 ). In summary, the wild type luciferase plasmid was constructed by inserting the sequence of hg38 chr8: 49957722–49958204 into pGL6-TA firefly luciferase reporter vector (Beyotime). The risk luciferase plasmid was designed according to the SNTG1 intronic SNP rs111829635 (hg38, chr8:49958116 C-T). The 293T or MC3T3-E1 cells were seeded in a 24-well plate for 16 hours and subsequently transfected with luciferase reporter plasmids along with the control Renilla plasmid (Beyotime), using lipofectamine 3000 (Thermo Fisher Scientific). After 48 hours of transfection, cells were collected and lysed using the lysis buffer provided in the Dual-Luciferase Reporter Assay System kit (Beyotime). Luciferase activity was analyzed in accordance with the guidelines outlined in the technique manual (Beyotime). SNTG1 overexpression and qRT-PCR analysis The cDNA of mouse SNTG1 was amplified via polymerase chain reaction (PCR) and subsequently cloned into the pEF-GFP vector (a gift from Connie Cepko, Addgene plasmid # 11154 ( 41 )), by replacing GFP to create vector pEF-SNTG1. To initiate the overexpression of SNTG1 , MC3T3-E1 cells were plated in 6-well and 24-well plates at a density of 6,000 cells per square centimeter. Following an incubation period of 16–20 hours, transfection of pEF-SNTG1 was performed using Lipofectamine 3000 transfection reagent (Thermo Fisher Scientific). As a negative control, pEF-GFP was utilized. Additionally, a set of duplicate wells without any transfection was included and labeled as "WT". After a further 4–6 hours, the cell culture medium was replaced with differentiation-induction medium. For quantitative real-time polymerase chain reaction (qRT-PCR) analysis, cells were harvested 72 hours after transfection. Total RNA was extracted from the target cells using TRIZOL (Invitrogen) according to the manufacturer's protocol. The isolated RNA was then reverse transcribed into complementary DNA (cDNA) using a reverse transcription kit (TransGen Biotech). For qRT-PCR analysis, a 2xSYBR Green Mix (TransGen Biotech) was utilized. The expression level of the target gene was normalized to that of GAPDH, which served as an endogenous control, enabling the comparison of samples. Cells intended for alkaline phosphatase (ALP) activity measurements were harvested after 6 days following transfection. Protein-protein interaction network analysis To uncover potential pathways and functions of candidate genes, we explored their direct relationships with previously reported genes for bone mineral density (BMD) using a protein network. The known genes were derived from 1,103 independent genome-wide significant SNPs annotated to 863 genes, as reported in the study by Morris et al. ( 42 ). We utilized the STRING database ( 43 ) online resource to predict the relationships between these 863 known genes and the 10 suggestive candidate genes (including three lincRNAs), ultimately identifying 19 known genes directly associated with the candidate genes. The interconnections between these genes were visualized using Cytoscape ( 44 ). We mainly considered two databases, the Gene Ontology (GO) database (for biological processes and cellular components) and the KEGG database, to perform pathway enrichment and functional annotation analyses on the selected 29 proteins. We manually selected pathways related to bone formation and development with a false discovery rate (FDR) of less than 0.05 for annotation. Results Integration of the SEAD reference panel After quality control of each reference panel (see Methods), we retained 4,480 samples and 78,429,408 variants from WBBC, 4,563 samples and 95,597,234 variants from SG10K, 1,031 samples and 63,925,145 variants from GAsP, and 993 samples and 35,157,155 variants from 1kGP-Asian. To maximize the number of variants as we did before ( 45 ), we applied reciprocal imputation approach step by step with minimac4 ( 37 ): first, we merged the reference panel of SG10K and WBBC by imputing with each other, we then combined haplotypes from 1kGP-Asian to the WBBC-SG10K panel, next, we obtained the SEAD (South and East Asian Reference Database) panel by combining all the 4 panels together (Fig. 1 A). After removing the singletons, we finally got 80,367,720 variants and 22,134 haplotypes for SEAD panel. The SEAD reference panel is now available online ( https://imputationserver.westlake.edu.cn/ ). Imputation performance in global populations The imputation concordance rate of NR-alleles were assessed across five reference panels using data from 929 global samples representing 55 populations. The samples were then categorized into regions, including Africa, America, Europe, Middle East, Oceania, Central and South Asia, and East Asia (Supplementary Table 2). Two types of pseudo arrays were generated: the Axiom array loci (820,967 variants) for all populations, and the Asian Screening array loci (659,184 variants) for East Asian and Central and South Asian population (Fig. 2 A). Among the non-Asian populations, the NR allele concordance rates from the 1kGP panel (this panel contains all 1kGP populations, not just Asian) were higher than that of the other reference panels in most of the populations, particularly in African populations (Supplementary Table 5 and Supplementary Fig. 2). To visualize these differences among the panels, we calculated the fold change of the NR allele concordance rate between the panels. When compared to the 1kGP panel, the SEAD reference panel demonstrated a lower level of NR-allele concordance for most of the non-Asian populations, however, it is worth noting that the SEAD reference panel demonstrated a slight advantage in imputing Oceanian populations (Fig. 2 A). Although SEAD might outperform the other three Asian reference panels (Supplementary Fig. 3), we did not suggest to use SEAD to impute the populations outside Asian. As for East Asian populations, the distributions of the non-reference homozygous allele concordance rate and the non-reference heterozygous allele concordance rate imputed by the SEAD reference panel were concentrated on higher values. The higher peak of the SEAD reference panel in the density indicated that the NR-allele concordance rate of this panel was more concentrated (Supplementary Fig. 4A). The SEAD reference panel consistently demonstrated higher NR-allele concordance when compared to the 1kGP panel (Fig. 2 A and Supplementary Fig. 2). In comparison to three other Asian reference panels, WBBC and SG10K might be slightly better in some of the populations (Supplementary Fig. 3). For the Central and South Asian populations, the NR-allele concordance rate corresponding to the peak of the SEAD reference panel also showed slightly higher than that of other reference panels (Supplementary Fig. 4B). The SEAD panel outperformed GAsP, WBBC and SG10K in all populations (Supplementary Fig. 3). Compared with the 1kGP panel, the SEAD panel also demonstrated its advantage in imputing most Central and South Asian populations (Fig. 2 A). These results indicated that SEAD panel was a suitable imputation reference panel for Asian populations. Imputation of disease-associated variants for Asian population We downloaded the complete dataset of genome-wide association data from the GWAS catalog, ( 46 ) which comprised 331,513 association sites. As these association loci were predominantly identified in European population data, we manually examined the population ancestry for each catalog ID, resulting in the identification of 9,810 association loci specific to Asian population (Supplementary Table 6). We assessed the imputation performance for the reference panels of the 9,810 loci extracted from WBBC array imputation (Fig. 1 B). The SEAD panel exhibited the highest number of well imputed variants, amounting to 4,424 (45.10%), followed by SG10K (Fig. 2 B and Supplementary Table 7). Only 25.97% of the disease-associated variants were imputed in GAsP panel (Fig. 2 B and Supplementary Table 7). The main advantage of SEAD panel lied in the imputation of low-frequency and rare variants, encompassing a total of 356 low-frequency variants and 133 rare variants (Supplementary Table 7). When categorizing the variants into seven Minor Allele Frequency (MAF) bins: < 0.1%, 0.1%-0.3%, 0.3%-0.5%, 0. 5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%, the SEAD panel demonstrated high number of well imputed sites compared to other panels (Fig. 2 C and Supplementary table 7). The WBBC and SG10K also performed well in some of the MAF bins (Fig. 2 C and Supplementary table 7). These results suggested SEAD panel performed well to impute the disease-associated variants in Asian population. Improvement in detecting rare variants in Chinese population To further assess the imputation performance of the reference panels in Chinese population (Fig. 1 B), we utilized GAsP, 1kGP, WBBC, SG10K, and the SEAD reference panels to impute WBBC array data. The SEAD reference panel imputed a total of 7,976,019 well-imputed sites, surpassing all the other panels, with SG10K ranking second with 6,725,695 sites (Fig. 3 A and Supplementary Table 8). GAsP panel imputed fewest well-imputed variants, because only a thousand samples were included in this panel (Fig. 3 A and Supplementary Table 8). Given that common variants have been extensively studied, our primary focus was on the imputation effect of reference panels for low-frequency (1% < MAF < 5%) and rare variants (MAF < 1%). The SEAD reference panel had obtained 3,209,490 well-imputed low/rare sites, and as many as twice rare mutations compared to the 1kGP panel (Fig. 3 A and Supplementary Table 8). Across the 7 MAF bins for rare and low-frequency variants, as MAF decreased, the SEAD reference panel exhibited a higher number of well-imputed sites compared to other panels (Fig. 3 B and Supplementary Table 8). Especially in the intervals of MAF < 0.1% and 0.1% < MAF < 0.3%, the imputation advantage of SEAD was most pronounced. (Fig. 3 B and Supplementary Table 8). To make the results comparable across the five reference panels, we also extracted a total of 8,525,136 shared SNPs for analysis. Within the 7 MAF bins, the SEAD reference panel always exhibited the highest Rsq value and largest number of well-quality sites (Fig. 3 C and Supplementary Table 9–10). As 184 samples underwent both whole-genome sequencing and array genotyping, we further assessed the imputation concordance of the five panels. The imputation accuracy of the SEAD, WBBC, and SG10K reference panels was considerably superior to that of the GASP and 1KGP panels. The concordance rate values corresponding to the peak of the SEAD imputation panel were slightly higher than those of WBBC and SG10K, but the differences were minimal (Fig. 3 D). Employment of the SEAD reference panel in bone mineral density GWAS analysis After imputation with SEAD panel for WBBC array data, we conducted GWAS analysis on two BMD traits (total hip and femoral neck BMD) (Fig. 1 C). These two BMD traits were highly correlated in either Pearson correlation ( r = 0.918) or genetic correlation ( r = 0.932, SD = 0.025). With genome-wide significance threshold ( P 0.01) that have been previously reported in GWAS were identified to be associated with both BMD traits, with the top signals at chr1: rs9659023 (an intronic variant of FMN2 , this locus was reported to be associated with heel ( 42 ) and total body BMD ( 47 )) and at chr6: rs138367190 (near SOX4 , this locus was reported to be associated with heel ( 42 ) and FN BMD ( 12 )) (Supplementary Fig. 5). For the relative rare variants (0.001 < MAF < 0.01), we utilized two analytical strategies: single-variant test with PLINK and spatial clustering analysis via the STAARpipeline (incorporating both sliding-window and gene-centric methods). In the single-variant test analysis, we prioritized the variants with small P -value and with replication (see Methods), and identified 95 suggestive variants for Hip BMD (Supplementary Table 11). Among these variants, 70 (73.7%) were annotated to SNTG1 gene by ANNOVAR (Supplementary Table 11). Of which, four rare variants (rs60103302, rs61260287, rs60600379, and rs57319781), in complete linkage disequilibrium (LD r 2 = 1.00), were identified for Hip BMD at genome-wide significance level ( P = 4.79×10 − 8 , MAF = 0.0091), these variants located in the intergenic region near SNTG1 on chromosome 8 (Fig. 4 A and 4 F, and Supplementary Table 11). As for FN BMD, 18 (46.2%) out of the 39 suggestive variants were annotated to SNTG1 gene (Fig. 4 C and 4 F, and Supplementary Table 11). In the spatial clustering analysis, we first implemented a sliding window of 2kb in STAAR-O test and adhered to the previous “discovery and replication" strategy. The results showed that 5 clustering regions were identified as the suggestive signals for Hip BMD (Fig. 4 B and Supplementary Table 12). In the cluster on chromosome 8 (chr8:49873459–49958458), a total of 22 windows were suggested, of which, 3 windows annotated near SNTG1 were discovered at bonferroni-corrected level P = 3.59×10 − 8 (0.05 divided by 1,392,296 windows), with top window at chr8:49898459–49900458 ( P = 1.08×10 − 8 ) (Fig. 4 B and Supplementary Table 12). Moreover, in the annotation-based gene-centric analysis, of the 7 non-coding categories of SNTG1 gene, 5 categories showed suggestive signal for Hip BMD, with the top signal at promoter region ( P promoter_CAGE = 4.72×10 − 8 ) (Fig. 4 E and Supplementary Table 13), showcasing at least one annotation reaching the genome-wide bonferroni-corrected level P = 3.57×10 − 7 (0.05 divided by 20,000 genes then divided by 7 categories). In summary, SNTG1 locus was highlighted in both analytical strategies with the strictest threshold applied to Hip BMD, and also achieved a suggestive level in FN BMD ( P single_variant = 4.85×10 − 6 , P slide_window = 2.36×10 − 6 , P gene_centric = 5.43×10 − 6 ) (Fig. 4 C, 4 D, 4 E, and Supplementary Tables 12 and 13), We further exploited the preliminary function of SNTG1 gene. Dual-luciferase assays were then performed in two cell lines, 293T and MC3T3-E1. It was observed that the intronic SNP rs111829635 C-T mutation significantly increased luciferase activity in both 293T cells and MC3T3-E1 cells (Fig. 4 G), suggesting that the identified locus would alter the SNTG1 gene expression. Moreover, the results from the CCK-8 proliferation assay showed that the absorbance at OD 450 nm was decreased after SNTG1 overexpression, and cell density was also decreased ( P < 0.05, Fig. 4 H and Fig. 4 I), revealing that overexpression of SNTG1 led to a significant reduction in cell proliferation. At the 72 hours, a noticeable decrease in the expression of osteogenic marker genes ( RUNX2 , COL1A1 and OCN) confirmed that SNTG1 inhibited the differentiation of preosteoblast MC3T3-E1 cells ( P < 0.05, Fig. 4 J). Furthermore, there was also a significant reduction in alkaline phosphatase (ALP) activity observed after six days. Overall, these results supported the notion that SNTG1 overexpression might inhibit human osteogenic proliferation and differentiation. Other suggestive candidate genes and loci Besides the SNTG1 locus, other suggestive loci were ranked in a credible priority order: 3 loci ( GRM7 , CPT2 , LINC00407 ) were identified by both single-variant test and spatial clustering analyses (level A), and 6 loci ( SH3TC2 , C4BPA , LINC02572 , LINC00461 , MORC2 , ADIG ) were identified in both BMD traits by either method (level B) (Fig. 4 A, 4 B, 4 C, 4 D, and Supplementary Table 14). Of these, 7 were annotated near genes, three candidates were lincRNAs. We then employed the protein-protein interaction analysis to include the known BMD-associated genes reported in previously GWAS ( 42 ) and suggested that 19 known genes directly associated with the 7 candidate genes (Supplementary Table 15). Functional enrichment of directly associated proteins revealed 24 potential categories with a false discovery rate (FDR) < = 0.05 (Supplementary Table 16). Among these enrichment analysis, we prioritized four functional annotations closely related to bone biology (Fig. 5 A), including “abnormal skeletal muscle morphology” (‘Monarch Human Phenotype’, FDR = 0.0246), “Dystrophin-associated glycoprotein complex” (‘GO Cellular Component’, FDR = 0.0047), “Fatty acid metabolic process” (‘GO Biological Process’, FDR = 0.0493) and “Complement and coagulation cascades” (‘KEGG pathway’, FDR = 0.0011). We compared the frequency distribution of all the rare variants identified in level-A and -B (84 SNPs) with the gnomAD and TOPMed databases across various populations (Fig. 5 B, Supplementary Table 17). The MAF of these rare variants most closely aligned with the MAF value in the East Asian (EAS) population (2,604 genomes) within gnomAD, validating the appropriateness of our panel for imputing genotype data in East Asian populations. Furthermore, the SNTG1 variants were predominantly rare in the Non-Finnish European (NFE: 34,029 genomes) and South Asian (SAS: 2,419 genomes) populations, while they were more common in the African/African American (AFR: 20,744 genomes) and Latino/Admixed American (AMR: 7,647 genomes) populations, highlighting population-specific patterns. Additionally, the other variants were present in either the SAS or EAS population but were largely absent in other populations, further illustrating the panel's ability to detect variations specific to East Asian and South Asian populations. Discussion In this study, we integrated whole-genome sequencing data from SG10K, GenomeAsia, WBBC, and 1kGP-Asian to create a combined reference panel for genotype imputation, the South and East Asian Reference Database (SEAD). It comprised a diverse range of populations across Asia, including 11 populations from GenomeAsia, one population from WBBC, three populations from SG10K, and 8 populations from 1kGP-Asian. With a sample size of 22,134 haplotypes, the SEAD panel stands as one of the most comprehensive panels in terms of coverage across Asia. We compared the concordance rate of the genotypes imputed from the SEAD panel across populations around the world, and suggested that the SEAD panel performed the best for the rare variants imputation in Asian populations. The SEAD panel also had advantage in detecting disease-associated low/rare variants. Subsequently, we applied the SEAD panel to the bone mineral density GWAS analyses in WBBC array cohorts, and identified a rare locus, SNTG1 , that was not reported even in the large biobank-scale GWAS. Asia, being the largest and most populous continent worldwide, boasts a wealth of human genetic resources. However, most of the whole genome-sequencing (WGS) efforts were carried out in Caucasian populations in the last decade ( 16 ). In our previous study, we conducted a thorough evaluation and discussion on the imputation of rare variants, highlighting the necessity of constructing a haplotype reference panel for Asian populations ( 9 , 48 ). In recent years, there has been a notable surge WGS data across Asia, particularly in regions such as Japan, Singapore, Korea and China. Despite the abundance of WGS projects, large-scale reference panel with broad geographical coverage across Asia is still needed. The accumulation of such vast datasets from diverse populations presents a unique opportunity: the combination of multiple WGS datasets to form a more comprehensive and expansive reference panel. The recently published Northeast Asian Reference Database, with over half of the samples deriving from Japan and Korea, represented populations from Northeast Asia ( 49 ). In our study, we integrated the data from WBBC with other Asian populations to construct the South and East Asian Reference Database (SEAD) reference panel, covering the widest geographical area in Asia, with the majority of its population originating from South and East Asia. With the SEAD panel, we first assessed the imputation performance across populations at the global scale by using HGDP database. We extracted the UK Biobank Axiom array sites as pseudo array from all HGDP populations, and extracted the Illumina Asian Screening Array sites from HGDP East Asian and Central and South Asian populations as pseudo array. Then the non-reference (NR) allele concordance rates were compared between true and imputed genotypes. It was not surprising that the SEAD reference panel outperformed other reference panels in imputing all East Asian populations and some of the Central and South Asian populations. Interestingly, the relative high concordance rate was achieved by the SEAD reference panel in the Oceania population, this might be attributable to shared haplotypes between the Oceania population and Singaporean Malays ( 19 ). Among all the Asian populations imputed by the SEAD panel, the NR concordance rate obtained from the pseudo array based on the Axiom array was consistently higher compared to that of the pseudo array extracted from Asian Screening array sites. These higher concordance rates could be attributed to the fact that the Axiom array contains approximately 150,000 more variants compared to the Asian Screening array. As WBBC cohort had samples that were both sequenced and genotyped, we then assessed the performance of the panels on Chinese population imputation. We found that the concordance rate values of the SEAD panel were slightly higher than those of WBBC and SG10K, but the differences were minimal. This suggested that the accuracy did not significantly improve after the combination, this is consistent with our pervious study that the imputation accuracy would decrease if the diverse sample exceed a set amount ( 9 ). However, the SEAD panel outperformed the other four panels in terms of well-imputed variants, with almost twice the number of variants at MAF bin less than 0.1% compared to SG10K and WBBC imputation. Finally, if we focused on the variants that had been identified to be associated with diseases and traits in Asian population in GWAS catalog, still, the SEAD panel performed the best. Besides the systematical and comprehensive imputation evaluation for the panels, we successfully applied the SEAD panel to a GWAS analysis for BMD traits at hip and femoral neck (FN). These two traits were highly correlated in either phenotypic or genetic correlation. Therefore, the association signals reported in both traits would reduce the random error of the results. Through both single-variant association test and spatial clustering analysis, we detected that rare variants near SNTG1 gene were associated with hip BMD. The SNTG1 gene also reached the suggestive threshold in FN BMD analysis using both methods. The gene and variants were not reported even in large-scale biobank GWAS for BMD ( 42 ), and the SNTG1 variants were predominantly rare in the Non-Finnish European and more common in the African/African American and Latino/Admixed American. The SNTG1 (Syntrophin Gamma 1) gene is known to encode a cytoplasmic peripheral membrane protein, which a candidate gene for idiopathic scoliosis ( 50 ). A southern Chinese cohort of patients with congenital scoliosis also identified copy number variants of SNTG1 ( 51 ). It was also noteworthy that SNTG1 genetic variations had been associated with neurodevelopmental disorders, suggesting a potential role in developmental processes ( 52 ). The SNTG1 protein was involved in the Dystrophin-associated glycoprotein complex (DGC) as shown in our pathway enrichment analysis. The DGC was known to play a significant role in maintaining the structural integrity of muscle fibers ( 53 ). In our study, we investigated the preliminary function of this gene, and suggested that overexpression of SNTG1 inhibited the proliferation and differentiation of preosteoblasts. This observation aligned with the direction of effect on BMD in our GWAS results. In summary, we constructed the haplotype reference panel for Asian population (SEAD panel) with the largest number of samples and the most abundant population diversity in Asia. The reference panel demonstrated excellent performance on imputing East Asian and Central and South Asian populations, especially in detecting rare variants. We provided this optimal imputation service online for free ( https://imputationserver.westlake.edu.cn/ ) for genetic studies in Asian populations. By applying the SEAD panel to impute the genotyping array data in Chinese population, we, for the first time, successfully identified rare variants near SNTG1 gene showing association with bone mineral density. Declarations Data availability GWAS summary statistics for the Hip and FN BMD from our analysis of the WBBC data are fully available at https://wbbc.westlake.edu.cn/downloads.html Acknowledgements We thank the “SG10K_Pilot Investigators” for providing the SG10K_Pilot data (EGAD00001005337). The data from the “SG10K_Pilot Study” reported here were obtained from EGA. This manuscript was not prepared in collaboration with the “SG10K_Pilot Study” and does not necessarily reflect the opinions or views of the “SG10K_Pilot Study”. We also thank the “Genome Asia 100K consortium” for providing the “GenomeAsia pilot project” (EGAS00001002921). We thankfully acknowledge the High-performance Computing Center at Westlake University. Funds This work was supported by the "Pioneer" and "Leading Goose" R&D Program of Zhejiang (#2023C03164), the National Natural Science Foundation of China (#82370887), the Chinese National Key Technology R&D Program, Ministry of Science and Technology (#2021YFC2501702), and the funds from the Westlake Laboratory of Life Sciences and Biomedicine (#202208014). Author contributions H.-F.Z. conceptualized and designed the study. M.-Y.Y., J.-D.Z., and W.-Y.B. conducted the data analysis. X.L. conducted experiments. S.-H.Y., W.-W.Z., J.-Q.L. and Y. S. conducted the whole sequencing experiments. C.-D.Y., M.-C.Q., K.-Q.L., C.-F.W., P.-K.C., K.S., S.-R.G., P.-P.Z., P.-L.G., Y.Q., J.-G.T., X.-J.Y., J.-X.G., X.C., M.-M.L., L.-X.L., G.T., S.-Y.X., L.X., F.H., J.-C.L., J.-F.G., B.-S.T., L.Y. and D.K. contributed to the sample collection, processing and preliminary data analysis. J.-J.Y., Y.-H.L. and N.L. designed the online website resource. M.-Y.Y. and J.-D.Z. drafted the manuscript, H.-F.Z. reviewed and edited manuscript. All authors contributed, discussed and approved manuscript. Competing interests S.-H.Y., W.-W.Z. and J.-Q.L. Y.S are employee of KingMed Diagnostics Co., Ltd. The other authors have no conflict of interest to declare. References Mahajan A, Taliun D, Thurner M, Robertson NR, Torres JM, Rayner NW, et al. Fine-mapping type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps. Nat Genet. 2018;50(11):1505-13. Zheng HF, Duncan EL, Yerges-Armstrong LM, Eriksson J, Bergstrom U, Leo PJ, et al. Meta-analysis of genome-wide studies identifies MEF2C SNPs associated with bone mineral density at forearm. J Med Genet. 2013;50(7):473-8. Hoffmann TJ, Sakoda LC, Shen L, Jorgenson E, Habel LA, Liu J, et al. Imputation of the rare HOXB13 G84E mutation and cancer risk in a large population-based cohort. PLoS Genet. 2015;11(1):e1004930. Handsaker RE, Van Doren V, Berman JR, Genovese G, Kashin S, Boettger LM, et al. Large multiallelic copy number variations in humans. Nat Genet. 2015;47(3):296-303. Das S, Abecasis GR, Browning BL. Genotype Imputation from Large Reference Panels. Annu Rev Genomics Hum Genet. 2018;19:73-96. Nelson SC, Stilp AM, Papanicolaou GJ, Taylor KD, Rotter JI, Thornton TA, et al. Improved imputation accuracy in Hispanic/Latino populations with larger and more diverse reference panels: applications in the Hispanic Community Health Study/Study of Latinos (HCHS/SOL). Hum Mol Genet. 2016;25(15):3245-54. Lert-Itthiporn W, Suktitipat B, Grove H, Sakuntabhai A, Malasit P, Tangthawornchaikul N, et al. Validation of genotype imputation in Southeast Asian populations and the effect of single nucleotide polymorphism annotation on imputation outcome. BMC Med Genet. 2018;19(1):23. Vergara C, Parker MM, Franco L, Cho MH, Valencia-Duarte AV, Beaty TH, et al. Genotype imputation performance of three reference panels using African ancestry individuals. Hum Genet. 2018;137(4):281-92. Bai WY, Zhu XW, Cong PK, Zhang XJ, Richards JB, Zheng HF. Genotype imputation and reference panel: a systematic evaluation on haplotype size and diversity. Brief Bioinform. 2019. Genomes Project C, Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68-74. Genome of the Netherlands C. Whole-genome sequence variation, population structure and demographic history of the Dutch population. Nat Genet. 2014;46(8):818-25. Zheng HF, Forgetta V, Hsu YH, Estrada K, Rosello-Diez A, Leo PJ, et al. Whole-genome sequencing identifies EN1 as a determinant of bone density and fracture. Nature. 2015;526(7571):112-7. Consortium UK, Walter K, Min JL, Huang J, Crooks L, Memari Y, et al. The UK10K project identifies rare variants in health and disease. Nature. 2015;526(7571):82-90. Jun G, English AC, Metcalf GA, Yang J, Chaisson MJ, Pankratz N, et al. Structural variation across 138,134 samples in the TOPMed consortium. bioRxiv. 2023. McCarthy S, Das S, Kretzschmar W, Delaneau O, Wood AR, Teumer A, et al. A reference panel of 64,976 haplotypes for genotype imputation. Nat Genet. 2016;48(10):1279-83. Jones KM, Cook-Deegan R. Complicated legacies: The human genome at 20. Science. 2021;371(6529):564-9. Nagasaki M, Yasuda J, Katsuoka F, Nariai N, Kojima K, Kawai Y, et al. Rare variant discovery by deep whole-genome sequencing of 1,070 Japanese individuals. Nat Commun. 2015;6:8018. Jeon S, Bhak Y, Choi Y, Jeon Y, Kim S, Jang J, et al. Korean Genome Project: 1094 Korean personal genomes with clinical information. Science Advances. 2020;6(eaaz7835). Wu D, Dou J, Chai X, Bellis C, Wilm A, Shih CC, et al. Large-Scale Whole-Genome Sequencing of Three Diverse Asian Populations in Singapore. Cell. 2019;179(3):736-49 e15. GenomeAsia KC. The GenomeAsia 100K Project enables genetic discoveries across Asia. Nature. 2019;576(7785):106-11. Zhang P, Luo H, Li Y, Wang Y, Wang J, Zheng Y, et al. NyuWa Genome resource: A deep whole-genome sequencing-based variation profile and reference panel for the Chinese population. Cell Rep. 2021;37(7):110017. Li L, Huang P, Sun X, Wang S, Xu M, Liu S, et al. The ChinaMAP reference panel for the accurate genotype imputation in Chinese populations. Cell Res. 2021;31(12):1308-10. Cong PK, Bai WY, Li JC, Yang MY, Khederzadeh S, Gai SR, et al. Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project. Nat Commun. 2022;13(1):2939. Wang C, Dai J, Qin N, Fan J, Ma H, Chen C, et al. Analyses of rare predisposing variants of lung cancer in 6,004 whole genomes in Chinese. Cancer Cell. 2022;40(10):1223-39 e6. Zhu XW, Liu KQ, Wang PY, Liu JQ, Chen JY, Xu XJ, et al. Cohort profile: the Westlake BioBank for Chinese (WBBC) pilot project. BMJ Open. 2021;11(6):e045564. Cong PK, Khederzadeh S, Yuan CD, Ma RJ, Zhang YY, Liu JQ, et al. Identification of clinically actionable secondary genetic variants from whole-genome sequencing in a large-scale Chinese population. Clin Transl Med. 2022;12(5):e866. Gibson G. Rare and common variants: twenty arguments. Nat Rev Genet. 2012;13(2):135-45. Alioto TS, Buchhalter I, Derdak S, Hutter B, Eldridge MD, Hovig E, et al. A comprehensive assessment of somatic mutation detection in cancer using whole-genome sequencing. Nat Commun. 2015;6:10001. Byrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022;185(18):3426-40 e19. Hinrichs AS, Karolchik D, Baertsch R, Barber GP, Bejerano G, Clawson H, et al. The UCSC Genome Browser Database: update 2006. Nucleic Acids Res. 2006;34(Database issue):D590-8. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25(16):2078-9. Manichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26(22):2867-73. Bergstrom A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, et al. Insights into human genetic variation and population history from 929 diverse genomes. Science. 2020;367(6484). Yang J, Lee SH, Goddard ME, Visscher PM. GCTA: a tool for genome-wide complex trait analysis. Am J Hum Genet. 2011;88(1):76-82. Chang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015;4:7. Das S, Forer L, Schonherr S, Sidore C, Locke AE, Kwong A, et al. Next-generation genotype imputation service and methods. Nat Genet. 2016;48(10):1284-7. Fuchsberger C, Abecasis GR, Hinds DA. minimac2: faster genotype imputation. Bioinformatics. 2015;31(5):782-4. Willer CJ, Li Y, Abecasis GR. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics. 2010;26(17):2190-1. Li Z, Li X, Zhou H, Gaynor SM, Selvaraj MS, Arapoglou T, et al. A framework for detecting noncoding rare-variant associations of large-scale whole-genome sequencing studies. Nat Methods. 2022;19(12):1599-611. Bai WY, Wang L, Ying ZM, Hu B, Xu L, Zhang GQ, et al. Identification of PIEZO1 polymorphisms for human bone mineral density. Bone. 2020;133:115247. Matsuda T, Cepko CL. Electroporation and RNA interference in the rodent retina in vivo and in vitro. Proceedings of the National Academy of Sciences of the United States of America. 2004;101(1):16-22. Morris JA, Kemp JP, Youlten SE, Laurent L, Logan JG, Chai RC, et al. An atlas of genetic influences on osteoporosis in humans and mice. Nat Genet. 2019;51(2):258-66. Szklarczyk D, Kirsch R, Koutrouli M, Nastou K, Mehryary F, Hachilif R, et al. The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51(D1):D638-D46. Shannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13(11):2498-504. Huang J, Howie B, McCarthy S, Memari Y, Walter K, Min JL, et al. Improved imputation of low-frequency and rare variants using the UK10K haplotype reference panel. Nat Commun. 2015;6:8111. Sollis E, Mosaku A, Abid A, Buniello A, Cerezo M, Gil L, et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51(D1):D977-D85. Medina-Gomez C, Kemp JP, Trajanoska K, Luan J, Chesi A, Ahluwalia TS, et al. Life-Course Genome-wide Association Study Meta-analysis of Total Body BMD and Assessment of Age-Specific Effects. Am J Hum Genet. 2018;102(1):88-102. Zheng HF, Ladouceur M, Greenwood CM, Richards JB. Effect of genome-wide genotyping and reference panels on rare variants imputation. J Genet Genomics. 2012;39(10):545-50. Choi J, Kim S, Kim J, Son H-Y, Yoo S-K, Kim C-U, et al. A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants. SCIENCE ADVANCES. 2023;9(eadg6319 ). Bashiardes S, Veile R, Allen M, Wise CA, Dobbs M, Morcuende JA, et al. SNTG1, the gene encoding gamma1-syntrophin: a candidate gene for idiopathic scoliosis. Hum Genet. 2004;115(1):81-9. Lai W, Feng X, Yue M, Cheung PWH, Choi VNT, Song YQ, et al. Identification of Copy Number Variants in a Southern Chinese Cohort of Patients with Congenital Scoliosis. Genes (Basel). 2021;12(8). Lemos RR, Oliveira DF, Zatz M, Oliveira JR. Population and computational analysis of the MGEA6 P521A variation as a risk factor for familial idiopathic basal ganglia calcification (Fahr's disease). J Mol Neurosci. 2011;43(3):333-6. Omairi S, Hau KL, Collins-Hooper H, Scott C, Vaiyapuri S, Torelli S, et al. Regulation of the dystrophin-associated glycoprotein complex composition by the metabolic properties of muscle fibres. Sci Rep. 2019;9(1):2770. Additional Declarations Yes there is potential Competing Interest. S.-H.Y., W.-W.Z. and J.-Q.L. Y.S are employee of KingMed Diagnostics Co., Ltd. The other authors have no conflict of interest to declare. Supplementary Files Supplementarytable12XXX45XXX710XXX1217.pdf Supplementary table 1-2,4-5,7-10,12-17 Supplementarytable3.pdf Supplementary table 3 Supplementarytable6.pdf Supplementary table 6 Supplymentarytable11.pdf Supplementary table 11 supplementaryfigure.docx Supplementary_figures Cite Share Download PDF Status: Published Journal Publication published 30 Dec, 2024 Read the published version in Nature Communications → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3680930","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":256566223,"identity":"3bafaa7d-0ca6-4ca0-8355-c6e2bd269be3","order_by":0,"name":"Hou-Feng Zheng","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABLUlEQVRIie2QsUrEQBCGZ1nYNPMAG46QV4gsRMR4eRVDIDY5LGwshYNcs2cd0YcICMHOwMLZBJ9BObhGi4icRBDPzXFik4h2gvlg2Nmf+WB3AHp6/iJcFzkBMHRB1SQGUNiE3ytY6DPVF6S/USj+RLHPx7P75ZUXSGOqHvYSz/cpnTsVeFZW0MVdi0IuZgfCKqNA4m20O0p0Q5kIUohEVrBtp0WhPHYHZqKCa92IUa729cOEQlBBViDjLQrjhy9rRdqPrtjJVz5S41m9wapTQR4z86lROIo5yQsiKYoQoOhUOI/cAUkiITF2yfQ9DKTCoy3phOJMMbdNsdNwYb4mniWNUlR1OfSNyeSS18dD6/RmvGhT1ivAz399Zc2qaMe8htQbteqe6enp6fnPfAC/GWAgYoavfAAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0001-5681-8598","institution":"Westlake University","correspondingAuthor":true,"prefix":"","firstName":"Hou-Feng","middleName":"","lastName":"Zheng","suffix":""},{"id":256566224,"identity":"a6b3b63b-3acf-4a8a-a137-932d0f8f911a","order_by":1,"name":"Meng-yuan Yang","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Meng-yuan","middleName":"","lastName":"Yang","suffix":""},{"id":256566225,"identity":"bcc1f600-98bb-4944-a762-0ead1974e335","order_by":2,"name":"Jia-Dong Zhong","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Jia-Dong","middleName":"","lastName":"Zhong","suffix":""},{"id":256566226,"identity":"6f94189d-2d78-4f84-96cc-bd5e519a2780","order_by":3,"name":"Xin Li","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Xin","middleName":"","lastName":"Li","suffix":""},{"id":256566227,"identity":"3b662943-a8f1-412d-8849-e13e26f28f36","order_by":4,"name":"Wei-Yang Bai","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Wei-Yang","middleName":"","lastName":"Bai","suffix":""},{"id":256566228,"identity":"6732663a-e053-466f-ac1a-b481b061cb82","order_by":5,"name":"Cheng-Da Yuan","email":"","orcid":"","institution":"Hangzhou Hospital of Traditional Chinese Medicine","correspondingAuthor":false,"prefix":"","firstName":"Cheng-Da","middleName":"","lastName":"Yuan","suffix":""},{"id":256566229,"identity":"619a4da0-ded8-45bb-ba35-7dd04bd601dd","order_by":6,"name":"Mo-Chang Qiu","email":"","orcid":"","institution":"Jiangxi Medical College","correspondingAuthor":false,"prefix":"","firstName":"Mo-Chang","middleName":"","lastName":"Qiu","suffix":""},{"id":256566230,"identity":"8c111c5c-ab56-4667-afaa-90eb7386f386","order_by":7,"name":"Ke-Qi Liu","email":"","orcid":"","institution":"Jiangxi Medical College","correspondingAuthor":false,"prefix":"","firstName":"Ke-Qi","middleName":"","lastName":"Liu","suffix":""},{"id":256566231,"identity":"06a4cea9-9369-42b6-af35-fbb781056f42","order_by":8,"name":"Chun-Fu Yu","email":"","orcid":"","institution":"Shangrao Municipal Hospital","correspondingAuthor":false,"prefix":"","firstName":"Chun-Fu","middleName":"","lastName":"Yu","suffix":""},{"id":256566232,"identity":"4d04ae3c-ba76-4fb8-adc8-4f7ccd89360b","order_by":9,"name":"Nan Li","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Nan","middleName":"","lastName":"Li","suffix":""},{"id":256566233,"identity":"aebee617-428d-4886-b436-4013db87f7ea","order_by":10,"name":"Ji-Jian Yang","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Ji-Jian","middleName":"","lastName":"Yang","suffix":""},{"id":256566234,"identity":"eae4a770-8a76-4386-8eed-68efee56680f","order_by":11,"name":"Yu-Heng Liu","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Yu-Heng","middleName":"","lastName":"Liu","suffix":""},{"id":256566235,"identity":"e2a2c272-3fb4-4067-ba00-947f88452919","order_by":12,"name":"Shi-Hui Yu","email":"","orcid":"","institution":"KingMed Diagnostics","correspondingAuthor":false,"prefix":"","firstName":"Shi-Hui","middleName":"","lastName":"Yu","suffix":""},{"id":256566236,"identity":"21e34622-1c3e-41c6-9026-9adcf9afcc40","order_by":13,"name":"Wei-Wei Zhao","email":"","orcid":"","institution":"KingMed Diagnostics","correspondingAuthor":false,"prefix":"","firstName":"Wei-Wei","middleName":"","lastName":"Zhao","suffix":""},{"id":256566237,"identity":"17c5084e-d1db-458e-80e1-a0d9b01a1a22","order_by":14,"name":"Jun-Quan Liu","email":"","orcid":"https://orcid.org/0000-0002-3697-8866","institution":"KingMed Diagnostics","correspondingAuthor":false,"prefix":"","firstName":"Jun-Quan","middleName":"","lastName":"Liu","suffix":""},{"id":256566238,"identity":"52325e10-f7e4-47d8-8fa0-f0df05b86509","order_by":15,"name":"Yi Sun","email":"","orcid":"","institution":"KingMed Diagnostics, Co., Ltd.","correspondingAuthor":false,"prefix":"","firstName":"Yi","middleName":"","lastName":"Sun","suffix":""},{"id":256566239,"identity":"9a283f13-ec28-4413-9428-96cc4bbe833d","order_by":16,"name":"Peikuan Cong","email":"","orcid":"https://orcid.org/0000-0002-4921-5657","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Peikuan","middleName":"","lastName":"Cong","suffix":""},{"id":256566240,"identity":"1426bcca-11db-414b-81c6-c4efa829ef36","order_by":17,"name":"Saber Khederzadeh","email":"","orcid":"https://orcid.org/0000-0002-0115-8710","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Saber","middleName":"","lastName":"Khederzadeh","suffix":""},{"id":256566241,"identity":"6a867e54-cad1-4633-a5c5-78fa6b3bbd60","order_by":18,"name":"Pianpian Zhao","email":"","orcid":"https://orcid.org/0009-0009-1676-1748","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Pianpian","middleName":"","lastName":"Zhao","suffix":""},{"id":256566242,"identity":"e48098b4-bffc-4230-a0b3-9b7be15d9681","order_by":19,"name":"Yu Qian","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Yu","middleName":"","lastName":"Qian","suffix":""},{"id":256566243,"identity":"f5890753-2bfd-4070-9f0f-49ae8c71def0","order_by":20,"name":"Peng-Lin Guan","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Peng-Lin","middleName":"","lastName":"Guan","suffix":""},{"id":256566244,"identity":"aaf0a865-0aaa-4af1-b570-1ea72f5658fc","order_by":21,"name":"Jia-Xuan Gu","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Jia-Xuan","middleName":"","lastName":"Gu","suffix":""},{"id":256566245,"identity":"3196087a-bba6-451e-a6b3-f97c43fd778d","order_by":22,"name":"Si-Rui Gai","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Si-Rui","middleName":"","lastName":"Gai","suffix":""},{"id":256566246,"identity":"a7328bee-74d5-4cfc-82d0-e71de83e494d","order_by":23,"name":"Xiang-Jiao Yi","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Xiang-Jiao","middleName":"","lastName":"Yi","suffix":""},{"id":256566247,"identity":"06303572-8ab6-4480-8dd4-e6b6ef6b0cc9","order_by":24,"name":"Jianguo Tao","email":"","orcid":"https://orcid.org/0000-0002-3473-9212","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Jianguo","middleName":"","lastName":"Tao","suffix":""},{"id":256566248,"identity":"4d512441-3cb9-43fc-a444-c3415346c1ac","order_by":25,"name":"Xiang Chen","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Xiang","middleName":"","lastName":"Chen","suffix":""},{"id":256566249,"identity":"4e4b837c-515e-4094-966a-2de49e5d6c30","order_by":26,"name":"Mao-Mao Miao","email":"","orcid":"","institution":"Westlake University","correspondingAuthor":false,"prefix":"","firstName":"Mao-Mao","middleName":"","lastName":"Miao","suffix":""},{"id":256566250,"identity":"93691679-520f-42ee-a078-1ca251acc901","order_by":27,"name":"Lan-Xin Lei","email":"","orcid":"","institution":"Imperial College London","correspondingAuthor":false,"prefix":"","firstName":"Lan-Xin","middleName":"","lastName":"Lei","suffix":""},{"id":256566251,"identity":"968c62b3-e63d-43e0-96d6-00b04fa0127e","order_by":28,"name":"Lin Xu","email":"","orcid":"","institution":"Binzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Lin","middleName":"","lastName":"Xu","suffix":""},{"id":256566252,"identity":"8a088b98-c149-4220-88b5-8ed91863f9af","order_by":29,"name":"Shu-Yang Xie","email":"","orcid":"","institution":"Binzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Shu-Yang","middleName":"","lastName":"Xie","suffix":""},{"id":256566253,"identity":"516d9fa9-4105-4eff-baf4-b7c29c5a2739","order_by":30,"name":"Geng Tian","email":"","orcid":"","institution":"Binzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Geng","middleName":"","lastName":"Tian","suffix":""},{"id":256566254,"identity":"183bc6ab-599e-422f-81ed-bd40f27c7ee9","order_by":31,"name":"Jinchen Li","email":"","orcid":"https://orcid.org/0000-0003-3335-9303","institution":"Xiangya hospital Central South University","correspondingAuthor":false,"prefix":"","firstName":"Jinchen","middleName":"","lastName":"Li","suffix":""},{"id":256566255,"identity":"279d2916-40fd-4389-bd35-e516fef03817","order_by":32,"name":"Jifeng Guo","email":"","orcid":"","institution":"Central South University","correspondingAuthor":false,"prefix":"","firstName":"Jifeng","middleName":"","lastName":"Guo","suffix":""},{"id":256566256,"identity":"a140b627-bbcd-448b-b0ca-5c354646f97b","order_by":33,"name":"David Karasik","email":"","orcid":"https://orcid.org/0000-0002-8826-0530","institution":"Hebrew SeniorLife","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"","lastName":"Karasik","suffix":""},{"id":256566257,"identity":"b043480d-fb8d-4217-ba15-83d31549fa48","order_by":34,"name":"Liu Yang","email":"","orcid":"","institution":"Fourth Military Medical University","correspondingAuthor":false,"prefix":"","firstName":"Liu","middleName":"","lastName":"Yang","suffix":""},{"id":256566258,"identity":"b8c2b117-c889-471b-925e-6f0edbdf85ef","order_by":35,"name":"Beisha Tang","email":"","orcid":"https://orcid.org/0000-0003-2120-1576","institution":"Xiangya Hospital, Central South University","correspondingAuthor":false,"prefix":"","firstName":"Beisha","middleName":"","lastName":"Tang","suffix":""},{"id":256566259,"identity":"905edcb4-4cd3-49b9-8286-f2647ce00496","order_by":36,"name":"Fei Huang","email":"","orcid":"","institution":"Binzhou Medical University","correspondingAuthor":false,"prefix":"","firstName":"Fei","middleName":"","lastName":"Huang","suffix":""}],"badges":[],"createdAt":"2023-11-29 09:56:48","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3680930/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3680930/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41467-024-55147-4","type":"published","date":"2024-12-30T05:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":50332225,"identity":"6a787330-438d-46fb-99eb-efb42293e8d0","added_by":"auto","created_at":"2024-01-29 21:59:35","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1682740,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eResearch design.\u003c/strong\u003e A. The construction process of the SEAD reference panel. Each panel first removes related samples. Impute and merge WBBC and SG10K each other first, then impute and merge the obtained panel with 1000G EAS and SAS samples, and finally impute and merge with GAsP. The resulting SEAD reference panel had 11,067 samples containing 80,367,720 loci. B. The evaluation of imputation performance in global populations from HGDP, in disease-associated variants for Asian population, and in Chinese population. C. The application of SEAD reference panel in and hip and femoral neck (FN) bone mineral density (BMD) GWAS analysis.\u003c/p\u003e","description":"","filename":"FIG1.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/1639f4f452a1f4c521f2c599.jpg"},{"id":50331971,"identity":"d0ffc428-1ce5-47b7-9d60-d043dd3afbee","added_by":"auto","created_at":"2024-01-29 21:51:35","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1041121,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe imputation performance in imputing HGDP samples and Asian-descent GWAS catalog sites.\u003c/strong\u003e A. The fold change of mean NR-allele concordance rate of the SEAD panel to 1kGP panel in Africa, America, Europe, Middle East and Oceania populations of HGDP. For all populations of HGDP, we used Axiom array sites to make preudo array. We also used the Asia screening array loci to make pseudo array only in East Asian and Central and South Asian populations. B. The total number of common, low frequency and rare well imputed variants in imputing Asian-descent GWAS association sites. C. The number of well-imputed rare/low variants imputed by five reference panels grouped into 7 MAF bins including \u0026lt; 0.1%, 0.1%-0.3%, 0.3%-0.5%, 0.5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%.\u003c/p\u003e","description":"","filename":"FIG2.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/446fc003617d2cc5de12be13.jpg"},{"id":50331974,"identity":"0130c2ee-49d8-48a0-ae84-4275d65a1ca4","added_by":"auto","created_at":"2024-01-29 21:51:35","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":1324287,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe imputation performance in imputing Chinese population.\u003c/strong\u003eA. The total number of common, low frequency and rare well-imputed variants in imputing WBBC array data. B. The well-imputed number of variants imputed by five reference panels in 7 MAF bins including \u0026lt; 0.1%, 0.1%-0.3%, 0.3%-0.5%, 0.5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%. C. The average R-square (Rsq) and number of well-imputed (Rsq ≥ 0.8) variants in shared sites of five reference panels (8,525,136 SNPs) among 7 MAF bins including \u0026lt; 0.1%, 0.1%-0.3%, 0.3%-0. 5%, 0. 5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%. D. Non-reference allele (NR-allele) concordance rate distribution (imputed variants vs. WGS variants) in 184 replication samples. Each dot represents an individual. The plots on the top and right are the corresponding density distributions.\u003c/p\u003e","description":"","filename":"FIG3.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/c7f75a81ee0fb20aadb019bc.jpg"},{"id":50332228,"identity":"ebfe4364-5695-426f-8705-57cea8aed039","added_by":"auto","created_at":"2024-01-29 21:59:35","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":3633071,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGWAS analysis of Hip and FN BMD in WBBC cohort.\u003c/strong\u003e A. and C. Meta-analysis results of the single-variant test of Hip and FN BMD. The P-value of significance threshold of 5 × 10-8 delineated by red lines, while the P-value of significance threshold of 1 × 10-4 delineated by blue lines. B and D. The sliding window analysis of Hip and FN BMD. The P-value of significance threshold of 5 × 10-8 delineated by red lines, while the P-value of significance threshold of 3.59 × 10-8 (0.05 divided by 1,392,296 windows) delineated by blue lines. Ten candidate genes which could be detected by more than two methods are marked by blue color. The SNTG1 gene which could be detected by four methods is labeled with a rectangle. E. Noncoding Gene-centric analysis of Hip and FN BMD. F. Locuszoom of SNTG1 gene in Hip and FN BMD individual analysis. G. The impact of rs111829635 alleles C and T on the expression of SNTG1 in 293T and MC3T3-E1 cells. H and I. The overexpression of SNTG1 alone inhibits cell proliferation. J. The overexpression of SNTG1 inhibits cell differentiation. With COL1A1, RUNX2, and Osteocalcin serving as indicators of cell differentiation. ALP refers to alkaline phosphatase.\u003c/p\u003e","description":"","filename":"FIG4.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/8de8655b05c2b5cbd62d7834.jpg"},{"id":50332227,"identity":"af2b6dc1-878e-4549-bab8-0a61a5bd2f5e","added_by":"auto","created_at":"2024-01-29 21:59:35","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":722542,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eNetwork of direct interactions between candidate rare variants and known BMD associated genes, and the MAF of candidate rare variants in TOPMed and GnomAD.\u003c/strong\u003e A. The graph marked orange are candidate genes, the graph marked blue are known BMD-related genes. The octagon are prior candidate genes (protein), the triangle are prior candidate genes (lincRNA). The pathway marked by the blue square are those of Cellular component (GO): Dystrophin-associated glycoprotein complex. The pathway marked by the yellow square are those of KEGG Pathways: Complement and coagulation cascades. The pathway marked by the pink square are those of Human Phenotype (Monarch): Abnormal skeletal muscle morphology. The pathway marked by the green square are those of Biological Process (GO): Fatty acid metabolic process. B. Candidate genes-related SNV frequency comparison. The gene marked by the black frame is the SNTG1 gene. The databases contained TOPMed and GnomAD.\u003c/p\u003e","description":"","filename":"FIG5.jpg","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/a8e64cd8d8e5e5734a01b06a.jpg"},{"id":72685560,"identity":"a4b80b36-413b-4492-8491-b1d030a27cc8","added_by":"auto","created_at":"2024-12-31 08:11:04","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":9018554,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/326374e5-0f69-4d94-80f3-6706518f0808.pdf"},{"id":50332371,"identity":"ba6bb70d-78da-4c33-b180-ba89b2c55a79","added_by":"auto","created_at":"2024-01-29 22:07:35","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1825535,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary table 1-2,4-5,7-10,12-17\u003c/p\u003e","description":"","filename":"Supplementarytable12XXX45XXX710XXX1217.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/c6218f8d92995b5e87beb5c5.pdf"},{"id":50331975,"identity":"de0f77b8-5745-4ece-b669-faf6e8cbc245","added_by":"auto","created_at":"2024-01-29 21:51:35","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":3995604,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary table 3\u003c/p\u003e","description":"","filename":"Supplementarytable3.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/51ebd025e72bb09acc12e93d.pdf"},{"id":50331980,"identity":"29723393-3bb5-4162-bfb4-f9ef83f88cb4","added_by":"auto","created_at":"2024-01-29 21:51:35","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":8048420,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary table 6\u003c/p\u003e","description":"","filename":"Supplementarytable6.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/b2fd2caf48981cfbf7dba8b5.pdf"},{"id":50331977,"identity":"bdc0a381-a5cb-48dd-8ddc-f7eda03f3931","added_by":"auto","created_at":"2024-01-29 21:51:35","extension":"pdf","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":612739,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary table 11\u003c/p\u003e","description":"","filename":"Supplymentarytable11.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/f6673676302c19d3e491f05b.pdf"},{"id":50331979,"identity":"26b6ffaf-fc19-4f01-bd47-8f695e771788","added_by":"auto","created_at":"2024-01-29 21:51:35","extension":"docx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":1394056,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary_figures\u003c/p\u003e","description":"","filename":"supplementaryfigure.docx","url":"https://assets-eu.researchsquare.com/files/rs-3680930/v1/1b7a916ef55ef5cea7e3a0c1.docx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential Competing Interest.\nS.-H.Y., W.-W.Z. and J.-Q.L. Y.S are employee of KingMed Diagnostics Co., Ltd. The other authors have no conflict of interest to declare.","formattedTitle":"SEAD: an augmented reference panel with 22,134 haplotypes boosts the rare variants imputation and GWAS analysis in Asian population","fulltext":[{"header":"Introduction","content":"\u003cp\u003eGenotype imputation is an effective way to save the cost of sequencing and has become a key step in genome-wide association studies (GWAS). It can increase the chance of identifying likely causal variants (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e), facilitate meta-analysis (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e), and aid in finding pleiotropic effects of risk variants (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e). Factors such as the haplotype size, ancestry diversity and sequencing depth can affect the imputation accuracy (\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e). Generally, a diverse reference panel can improve accuracy in genetically diverse populations (\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e), while an ancestry-specific reference panel can benefit the corresponding population (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). In our pervious study, we found that the imputation accuracy for European population benefit from an increasing in haplotype size and population diversity of the reference panel, while the accuracy for the Han Chinese population reached its peak with a fraction of an extra diverse sample (8\u0026ndash;21%) in the reference panel (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). Given that imputation accuracy directly impacts the credibility of subsequent analysis, it is crucial to select an appropriate reference panel before imputation.\u003c/p\u003e \u003cp\u003eThe 1000 Genomes Project (1kGP) is one of the most renowned and widely used reference panel for genotype imputation, comprising 2,504 individuals from 26 populations (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). The latest update to 1kGP phase 3 sequenced 3,202 samples (including an additional 698 samples related to the previous participants) and identified 70,594,286 single nucleotide polymorphisms (SNPs) with 30\u0026times; depth. As many rare and low-frequency variants associated with diseases tend to be population-specific (\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e), an increasing number of human genome projects are focusing on specific populations to provide population-specific reference panels. However, most of the whole-genome sequencing (WGS) efforts were carried out in European-descent individuals, such as the GoNL project (the Genome of the Netherlands project, n\u0026thinsp;=\u0026thinsp;769) (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e), UK10K project (~\u0026thinsp;4,000 WGS and ~\u0026thinsp;6,000 whole exome sequencing samples in UK)(\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e) and the TOPMed program (Trans-Omics for Precision Medicine, n\u0026thinsp;=\u0026thinsp;~\u0026thinsp;138,000) (\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e). These efforts have enabled the provision of a large-scale combined reference panel for the European population, such as the HRC panel (Haplotype Reference Consortium, n\u0026thinsp;=\u0026thinsp;32,470) (\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e). The update of analytical methods and the emergence of various imputation reference panels have improved the imputation quality for European population. However, one of the complicated legacies of human genome project after 20 years\u0026rsquo; development is that lacking of diversity might hinder the promise of genome science (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). As the largest continent, Asia accounts for 59.5% of the worldwide population (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.worldometers.info/world-population/\u003c/span\u003e\u003cspan address=\"https://www.worldometers.info/world-population/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). Fortunately, the Asian ancestry-dominant populations have been recently sequenced and analyzed to understand the genetic basis of the Asian, such as the Japanese (\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e), Korean (\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e), Singaporean (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e), Indian (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e) and Chinese (\u003cspan additionalcitationids=\"CR22 CR23\" citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e). Among these efforts, the Westlake BioBank for Chinese (WBBC) project (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e) was initiated in 2017 by our team. Up to now, we have whole-genome sequenced 4480 Chinese samples, covering 29 out of 34 administrative divisions of China (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eRare variants, which may have low level of pairwise linkage disequilibrium with common variants, can potentially result in significant functional consequences (\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e). Next-generation sequencing, with adequate coverage, enables the accurate detection of rare variants (\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e). Rare variants are harder to impute than common variants because rare variants only appear a few times in the reference panel. Further, the evaluation of rare variant imputation in Asian populations is hindered by the limited sample size of currently published studies, which typically comprise only a few thousand individuals.\u003c/p\u003e \u003cp\u003eIn this study, we integrated the WGS data from multiple sources, including Singapore SG10K pilot project (13.7\u0026times;, 4563 samples), GenomeAsia pilot project (36\u0026times;, 1031 samples), WBBC pilot project (13.9\u0026times;, 4480 samples), and the high-coverage 1kGP-Asian phase 3 (30\u0026times;, 993 East and South Asian from 3202 samples). These data were combined to create the South and East Asian Reference Database (SEAD), encompassing a total of 11,067 individuals, making it a large-scale reference panel for Asian populations. We compared the concordance rate of the genotypes imputed from the SEAD panel across populations around the world. We then focused on evaluating the imputation quality of this panel to impute rare variants in Asian populations. Finally, we applied this augmented panel to the WBBC genome-wide genotyping data to explore its role in imputing possible likely causal rare variants for bone-related traits.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eQuality control of the four Asian reference panels\u003c/p\u003e \u003cp\u003eWe collected Asian haplotypes from the SG10K pilot project (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e), the GenomeAsia pilot project (GAsP) (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e), the newly released high-coverage 1000 Genomes project (1kGP) (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e) and Westlake Biobank for Chinese pilot project (WBBC) (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). The SG10K dataset comprises 4,810 whole-genome sequencings (WGS) of Singaporean Chinese, Malays, and Indians, with an average depth of 13.7\u0026times;. The GAsP dataset includes 1,163 WGS samples with an average depth of 36\u0026times;, mainly covering Indians, Korean, Pakistanis, etc. These two datasets were obtained by applying to the consortium. Since the SG10K and GAsP datasets were based on Genome Reference Consortium Human build 37 (GRCh37), we employed liftover (\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e) to update the genome assembly version from GRCh37 to GRCh38. For the GAsP dataset, we transformed multi-allelic variants into multiple bi-allelic variants using Bcftools (\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e) and performed haplotype re-phasing using SHAPEIT v2 with parameters including windows of size 0.5, 200 states, and an effective size of 14,269. The high-coverage 1kGP dataset with an average depth of 30\u0026times;, based on the GRCh38 assembly, were downloaded from the \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.internationalgenome.org/data-portal/data-collection/30x-grch38\u003c/span\u003e\u003cspan address=\"https://www.internationalgenome.org/data-portal/data-collection/30x-grch38\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Then we extracted East Asian (EAS) and South Asian (SAS) dataset from 1kGP and removed sites with minor allele count (MAC) equals zero. The WBBC dataset, which were generated by our own, included 4,480 Chinese samples with an average sequencing depth of 13.9\u0026times;, supporting both GRCh37 and GRCh38. We identified relative pairs using KING (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e) and excluded sample pairs with estimated kinship coefficients restricted to the range of 0.177\u0026ndash;0.354. Ultimately, we retained 4,480 samples from WBBC 4,563 samples from SG10K, 1,031 samples from GAsP, and 993 samples from the EAS/SAS subset of the 1kGP dataset (1kGP-Asian) (Supplementary Table\u0026nbsp;1).\u003c/p\u003e \u003cp\u003eEvaluation of genotype concordance in HGDP dataset.\u003c/p\u003e \u003cp\u003eThe 929 genomes from 55 diverse human populations from Human Genome Diversity Project (HGDP) (\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e) dataset was utilized to estimate the genotype concordance of imputation across global populations. The HGDP dataset includes 104 samples from African populations, 61 samples from American populations, 197 samples from Central and South Asia populations, 223 samples from East Asia populations, 155 samples from European populations, 161 samples from Middle East populations, and 28 samples from Oceania populations (Supplementary Table\u0026nbsp;2). We obtained the HGDP data from the following source: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003eftp://ngs.sanger.ac.uk/production/hgdp\u003c/span\u003e\u003cspan address=\"http://ftp://ngs.sanger.ac.uk/production/hgdp\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. The phased autosome VCF file was split into 22 chromosomes, which were considered as the true set. To compare the imputation concordance of the reference panels (WBBC, SG10K, 1kGP, GAsP and SEAD), we extracted the genetic variants from the 929 HGDP genome data with Axiom array loci (820,967 variants) to create pseudo array, which was then used as the test set for imputation. We downloaded Axiom array SNP list from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ukbiobank.ac.uk/enable-your-research/about-our-data/genetic-data\u003c/span\u003e\u003cspan address=\"https://www.ukbiobank.ac.uk/enable-your-research/about-our-data/genetic-data\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. We imputed the genotypes from the 5 panels, then compared the concordance between the imputed genotypes and true set. We also extracted Asian Screening Array sites (659,184 variants) in Central and South Asia populations and East Asia populations (420 samples in total) as Asian array for comparison purpose. Before imputation, variants in the pseudo arrays with MAF less than 1% and samples with calling rate below 95% were excluded. Here, the singletons were excluded in the 5 reference panels, consists of 35,616,674 variants for WBBC, 50,246,865 variants for SG10K, 70,594,286 variants for 1kGP (not just Asian), 20,203,158 variants for GAsP, 80,367,720 variants for SEAD. We calculated the non-reference genotype concordance rate, the non-reference heterozygote concordance rate and non-reference homozygote concordance rate between true genotypes and imputed genotypes pseudo arrays for each individual (Supplementary Fig.\u0026nbsp;1), as what we did before (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e). The average concordance rate in single population was calculated as the mean of all the individuals\u0026rsquo; concordance rate in this population. We quantified the ratio of NR concordance rate of the SEAD panel to other four panels by using the following formula, respectively.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\text{l}\\text{o}\\text{g}\\text{F}\\text{C}=\\text{l}\\text{o}\\text{g}\\left(\\frac{\\text{N}\\text{R} \\text{c}\\text{o}\\text{n}\\text{c}\\text{o}\\text{r}\\text{d}\\text{a}\\text{n}\\text{c}\\text{e} \\text{r}\\text{a}\\text{t}\\text{e}1}{\\text{N}\\text{R} \\text{c}\\text{o}\\text{n}\\text{c}\\text{o}\\text{r}\\text{d}\\text{a}\\text{n}\\text{c}\\text{e} \\text{r}\\text{a}\\text{t}\\text{e}2}\\right)$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eAll these analyses were performed on chromosome 2.\u003c/p\u003e \u003cp\u003eQuality control of the Chinese genotyping array data\u003c/p\u003e \u003cp\u003eBesides the whole genome sequencing data, the WBBC study also genotyped 6,080 individuals with the high-density Illumina Asian Screening Array (ASA, based on GRCh37), resulting in the identification of a total of 659,184 SNPs (\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). Of which, 184 samples underwent both whole-genome sequencing and array genotyping.\u003c/p\u003e \u003cp\u003eAs part of quality control, we employed GCTA (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e) to calculate the pairwise genetic relationship matrix using common variants and remove samples with a coefficient\u0026thinsp;\u0026gt;\u0026thinsp;0.025 (Supplementary Table\u0026nbsp;3). We then excluded samples with missing call rates \u0026ge; 2% and excluded SNPs with missing call rates \u0026ge; 5%, MAF\u0026thinsp;\u0026lt;\u0026thinsp;1% and Hardy-Weinberg equilibrium at \u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;1 \u0026times; 10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e by using PLINK1.9 (\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e). The genotype assembly version of the WBBC array data was updated from GRCh37 to GRCh38. Finally, we retained a total of 470,242 variants and 5,679 samples in the WBBC array dataset. The data were phased using SHAPEIT with a window size of 0.5, 200 states, and an effective size of 14,269.\u003c/p\u003e \u003cp\u003eEvaluation of reference panel for imputation in the Chinese population\u003c/p\u003e \u003cp\u003eThe procedure to assess the imputation performance in the Chinese population using the five reference panels (WBBC, SG10K, 1kGP, GAsP, and SEAD) was similar to the evaluation in HGDP dataset. As we have thousands of ASA array samples, we could assess the imputation performance for low-frequency and rare variants. We grouped the variants into seven Minor Allele Frequency (MAF) bins: \u0026lt; 0.1%, 0.1%-0.3%, 0.3%-0.5%, 0.5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%. Genotype imputation was performed using Minimac4 (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e, \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e), with a chunk length of 20 Mb and a chunk overlap of 4 Mb. We used R-square as the estimated value, defined as the squared correlation between imputed genotypes and observed genotypes, produced by Minimac4. Sites with an R-square value greater than 0.8 were considered well-imputed variants. We counted the number of well-imputed variants and calculated the average R\u003csup\u003e2\u003c/sup\u003e within each MAF bin. To compare the imputation accuracy among different reference panels, we extracted 8,525,136 variants that were shared by all five panels and evaluated the imputation performance again.\u003c/p\u003e \u003cp\u003eIn the WBBC ASA array data, 184 samples underwent both whole-genome sequencing and array genotyping. We then utilized these samples to evaluate the imputation concordance rate. We randomly masked 2,600 variants on chromosome 2 from the array data and subsequently imputed these SNPs using the five Asian panels. We extracted the imputed 2,600 loci as the test set, while the original genotype of the 2,600 loci from the array data served as the true set. Employing the methods from the previous section, we calculated both the non-reference homozygous and heterozygous concordance rates.\u003c/p\u003e \u003cp\u003eGenome wide association study of BMD traits\u003c/p\u003e \u003cp\u003eAs we mentioned above, the WBBC pilot study genotyped 6,080 individuals with ASA array, and a bunch of bone-related phenotypes were collected within WBBC (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e). After quality control, a total of 470,242 variants and 5,679 samples were retained. Here, we took two correlated bone mineral density (BMD) traits (total hip and femoral neck BMD) as example, we removed individuals with missing phenotype data, and excluded outliers using the mean\u0026thinsp;\u0026plusmn;\u0026thinsp;4 standard deviations (n\u0026thinsp;=\u0026thinsp;5,369 left). We then grouped the ASA samples according to the assessment place of collection: 2,332 samples recruited from Jiangxi province (samples were mainly from Southern China) were taken as discovery cohort (cohort 1), and 3,037 samples recruited from Shandong province (samples were mainly from Northern China) were taken as replication cohort (cohort 2), and \u003cb\u003evice versa\u003c/b\u003e. The baseline statistics of the study samples were shown in Supplementary Table\u0026nbsp;4. Comprehensive GWAS analyses were conducted with the imputed genotypes from the augmented SEAD reference panel. We kept the imputed variants with Rsq\u0026thinsp;\u0026gt;\u0026thinsp;0.5 in the GWAS analysis. The imputation boosted the analyzed genetic variants from 470,242 to 19,235,129, with 3,195,758 variants with MAF between 0.1% and 1%. The BMD phenotypes (total hip and femoral neck BMD) in this study was analyzed as continuous outcomes, adjusting for \u0026lsquo;sex\u0026rsquo;, \u0026lsquo;age\u0026rsquo;, \u0026lsquo;BMI\u0026rsquo;, \u0026lsquo;geographical region\u0026rsquo;, and \u0026lsquo;10 principal components (PCs)\u0026rsquo;. The parameters applied in PLINK were \u0026lsquo;--geno\u0026rsquo; of 0.05, \u0026lsquo;--mind\u0026rsquo; of 0.05, \u0026lsquo;--hwe\u0026rsquo; of 1 \u0026times; 10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e and \u0026lsquo;--maf\u0026rsquo; of 0.001. Furthermore, we performed the meta-analysis of GWAS summary statistics in cohort 1 and cohort 2 using METAL software (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e). Rare variants (0.001\u0026thinsp;\u0026lt;\u0026thinsp;MAF\u0026thinsp;\u0026lt;\u0026thinsp;0.01) were considered as candidates if they met the following criteria: a \u003cem\u003eP\u003c/em\u003e-value less than 1e-04 in either cohort, a \u003cem\u003eP\u003c/em\u003e-value less than 0.05 in the other cohort, and a \u003cem\u003eP\u003c/em\u003e-value less than 1e-04 in the meta-analysis.\u003c/p\u003e \u003cp\u003eSpatial clustering analysis\u003c/p\u003e \u003cp\u003eWe also performed the STAAR (variant-set test for association using annotation information) (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e) framework to identify rare variants that would be associated with BMD traits (Hip and FN BMD). The STAAR pipeline facilitates rare variants spatial clustering analyses, including sliding window-based analysis and gene-centric analysis (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e). In each test, we included variants with a minor allele frequency (MAF) of 0.1\u0026ndash;1%. We employed fixed-size sliding window analysis, allowing for the systematic examination of distinct genomic regions by moving a 2kb window every 1kb across the genome, a total of 1,392,296 windows was divided. The gene-centric analysis method is the variant-set test for association using annotation information for a gene (STAAR-O), enhancing the power of rare variant association test by incorporating multiple variant functional annotations. For the gene-centric non-coding variants, aggregation was performed based on 7 categories: downstream, enhancer variants overlaid with Cap Analysis of Gene Expression (CAGE) sites, promoter CAGE, enhancer variants overlaid with DNase hypersensitivity (DHS), promoter DHS, upstream, and UTR2, and clustering regions with rare variations less than 2 were excluded. Both the sliding window and gene-centric methods were performed separately on the cohort 1 and cohort 2.\u003c/p\u003e \u003cp\u003eCell culture\u003c/p\u003e \u003cp\u003eThe HEK293T cells were cultured at 37\u0026deg;C in a humidified atmosphere with 5% CO2 using DMEM basal media supplemented with 10% fetal bovine serum. The MC3T3-E1 subclone 14 cells were cultured with similar condition with ascorbate-free αMEM basal media supplemented with 10% fetal bovine serum. The media were refreshed every 3\u0026ndash;4 days, and cells were passaged when they reached 90% confluence. To initiate differentiation, the media were replaced with a calcification-inducing medium containing 50 ng/L vitamin C (VC), 10 nM dexamethasone, and 10 mM beta-glycerophosphate.\u003c/p\u003e \u003cp\u003eLuciferase Reporter assays\u003c/p\u003e \u003cp\u003eDual-Luciferase reporter assays were conducted following previously described methods (\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e). In summary, the wild type luciferase plasmid was constructed by inserting the sequence of hg38 chr8: 49957722\u0026ndash;49958204 into pGL6-TA firefly luciferase reporter vector (Beyotime). The risk luciferase plasmid was designed according to the \u003cem\u003eSNTG1\u003c/em\u003e intronic SNP rs111829635 (hg38, chr8:49958116 C-T). The 293T or MC3T3-E1 cells were seeded in a 24-well plate for 16 hours and subsequently transfected with luciferase reporter plasmids along with the control Renilla plasmid (Beyotime), using lipofectamine 3000 (Thermo Fisher Scientific). After 48 hours of transfection, cells were collected and lysed using the lysis buffer provided in the Dual-Luciferase Reporter Assay System kit (Beyotime). Luciferase activity was analyzed in accordance with the guidelines outlined in the technique manual (Beyotime).\u003c/p\u003e \u003cp\u003e \u003cem\u003eSNTG1\u003c/em\u003e overexpression and qRT-PCR analysis\u003c/p\u003e \u003cp\u003eThe cDNA of mouse \u003cem\u003eSNTG1\u003c/em\u003e was amplified via polymerase chain reaction (PCR) and subsequently cloned into the pEF-GFP vector (a gift from Connie Cepko, Addgene plasmid # 11154 (\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e)), by replacing GFP to create vector pEF-SNTG1. To initiate the overexpression of \u003cem\u003eSNTG1\u003c/em\u003e, MC3T3-E1 cells were plated in 6-well and 24-well plates at a density of 6,000 cells per square centimeter. Following an incubation period of 16\u0026ndash;20 hours, transfection of pEF-SNTG1 was performed using Lipofectamine 3000 transfection reagent (Thermo Fisher Scientific). As a negative control, pEF-GFP was utilized. Additionally, a set of duplicate wells without any transfection was included and labeled as \"WT\". After a further 4\u0026ndash;6 hours, the cell culture medium was replaced with differentiation-induction medium.\u003c/p\u003e \u003cp\u003eFor quantitative real-time polymerase chain reaction (qRT-PCR) analysis, cells were harvested 72 hours after transfection. Total RNA was extracted from the target cells using TRIZOL (Invitrogen) according to the manufacturer's protocol. The isolated RNA was then reverse transcribed into complementary DNA (cDNA) using a reverse transcription kit (TransGen Biotech). For qRT-PCR analysis, a 2xSYBR Green Mix (TransGen Biotech) was utilized. The expression level of the target gene was normalized to that of GAPDH, which served as an endogenous control, enabling the comparison of samples. Cells intended for alkaline phosphatase (ALP) activity measurements were harvested after 6 days following transfection.\u003c/p\u003e \u003cp\u003eProtein-protein interaction network analysis\u003c/p\u003e \u003cp\u003eTo uncover potential pathways and functions of candidate genes, we explored their direct relationships with previously reported genes for bone mineral density (BMD) using a protein network. The known genes were derived from 1,103 independent genome-wide significant SNPs annotated to 863 genes, as reported in the study by Morris et al. (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e). We utilized the STRING database (\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e) online resource to predict the relationships between these 863 known genes and the 10 suggestive candidate genes (including three lincRNAs), ultimately identifying 19 known genes directly associated with the candidate genes. The interconnections between these genes were visualized using Cytoscape (\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e). We mainly considered two databases, the Gene Ontology (GO) database (for biological processes and cellular components) and the KEGG database, to perform pathway enrichment and functional annotation analyses on the selected 29 proteins. We manually selected pathways related to bone formation and development with a false discovery rate (FDR) of less than 0.05 for annotation.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eIntegration of the SEAD reference panel\u003c/p\u003e\n\u003cp\u003eAfter quality control of each reference panel (see Methods), we retained 4,480 samples and 78,429,408 variants from WBBC, 4,563 samples and 95,597,234 variants from SG10K, 1,031 samples and 63,925,145 variants from GAsP, and 993 samples and 35,157,155 variants from 1kGP-Asian. To maximize the number of variants as we did before (\u003cspan class=\"CitationRef\"\u003e45\u003c/span\u003e), we applied reciprocal imputation approach step by step with minimac4 (\u003cspan class=\"CitationRef\"\u003e37\u003c/span\u003e): first, we merged the reference panel of SG10K and WBBC by imputing with each other, we then combined haplotypes from 1kGP-Asian to the WBBC-SG10K panel, next, we obtained the SEAD (South and East Asian Reference Database) panel by combining all the 4 panels together (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eA). After removing the singletons, we finally got 80,367,720 variants and 22,134 haplotypes for SEAD panel. The SEAD reference panel is now available online (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://imputationserver.westlake.edu.cn/\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003eImputation performance in global populations\u003c/p\u003e\n\u003cp\u003eThe imputation concordance rate of NR-alleles were assessed across five reference panels using data from 929 global samples representing 55 populations. The samples were then categorized into regions, including Africa, America, Europe, Middle East, Oceania, Central and South Asia, and East Asia (Supplementary Table\u0026nbsp;2). Two types of pseudo arrays were generated: the Axiom array loci (820,967 variants) for all populations, and the Asian Screening array loci (659,184 variants) for East Asian and Central and South Asian population (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eA).\u003c/p\u003e\n\u003cp\u003eAmong the non-Asian populations, the NR allele concordance rates from the 1kGP panel (this panel contains all 1kGP populations, not just Asian) were higher than that of the other reference panels in most of the populations, particularly in African populations (Supplementary Table\u0026nbsp;5 and Supplementary Fig.\u0026nbsp;2). To visualize these differences among the panels, we calculated the fold change of the NR allele concordance rate between the panels. When compared to the 1kGP panel, the SEAD reference panel demonstrated a lower level of NR-allele concordance for most of the non-Asian populations, however, it is worth noting that the SEAD reference panel demonstrated a slight advantage in imputing Oceanian populations (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eA). Although SEAD might outperform the other three Asian reference panels (Supplementary Fig.\u0026nbsp;3), we did not suggest to use SEAD to impute the populations outside Asian.\u003c/p\u003e\n\u003cp\u003eAs for East Asian populations, the distributions of the non-reference homozygous allele concordance rate and the non-reference heterozygous allele concordance rate imputed by the SEAD reference panel were concentrated on higher values. The higher peak of the SEAD reference panel in the density indicated that the NR-allele concordance rate of this panel was more concentrated (Supplementary Fig.\u0026nbsp;4A). The SEAD reference panel consistently demonstrated higher NR-allele concordance when compared to the 1kGP panel (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eA and Supplementary Fig.\u0026nbsp;2). In comparison to three other Asian reference panels, WBBC and SG10K might be slightly better in some of the populations (Supplementary Fig.\u0026nbsp;3). For the Central and South Asian populations, the NR-allele concordance rate corresponding to the peak of the SEAD reference panel also showed slightly higher than that of other reference panels (Supplementary Fig.\u0026nbsp;4B). The SEAD panel outperformed GAsP, WBBC and SG10K in all populations (Supplementary Fig.\u0026nbsp;3). Compared with the 1kGP panel, the SEAD panel also demonstrated its advantage in imputing most Central and South Asian populations (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eA).\u003c/p\u003e\n\u003cp\u003eThese results indicated that SEAD panel was a suitable imputation reference panel for Asian populations.\u003c/p\u003e\n\u003cp\u003eImputation of disease-associated variants for Asian population\u003c/p\u003e\n\u003cp\u003eWe downloaded the complete dataset of genome-wide association data from the GWAS catalog, (\u003cspan class=\"CitationRef\"\u003e46\u003c/span\u003e) which comprised 331,513 association sites. As these association loci were predominantly identified in European population data, we manually examined the population ancestry for each catalog ID, resulting in the identification of 9,810 association loci specific to Asian population (Supplementary Table\u0026nbsp;6). We assessed the imputation performance for the reference panels of the 9,810 loci extracted from WBBC array imputation (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eB). The SEAD panel exhibited the highest number of well imputed variants, amounting to 4,424 (45.10%), followed by SG10K (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eB and Supplementary Table\u0026nbsp;7). Only 25.97% of the disease-associated variants were imputed in GAsP panel (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eB and Supplementary Table\u0026nbsp;7). The main advantage of SEAD panel lied in the imputation of low-frequency and rare variants, encompassing a total of 356 low-frequency variants and 133 rare variants (Supplementary Table\u0026nbsp;7). When categorizing the variants into seven Minor Allele Frequency (MAF) bins: \u0026lt; 0.1%, 0.1%-0.3%, 0.3%-0.5%, 0. 5%-0.7%, 0.7%-1%, 1%-2%, and 2%-5%, the SEAD panel demonstrated high number of well imputed sites compared to other panels (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eC and Supplementary table 7). The WBBC and SG10K also performed well in some of the MAF bins (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003eC and Supplementary table 7). These results suggested SEAD panel performed well to impute the disease-associated variants in Asian population.\u003c/p\u003e\n\u003cp\u003eImprovement in detecting rare variants in Chinese population\u003c/p\u003e\n\u003cp\u003eTo further assess the imputation performance of the reference panels in Chinese population (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eB), we utilized GAsP, 1kGP, WBBC, SG10K, and the SEAD reference panels to impute WBBC array data. The SEAD reference panel imputed a total of 7,976,019 well-imputed sites, surpassing all the other panels, with SG10K ranking second with 6,725,695 sites (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eA and Supplementary Table\u0026nbsp;8). GAsP panel imputed fewest well-imputed variants, because only a thousand samples were included in this panel (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eA and Supplementary Table\u0026nbsp;8).\u003c/p\u003e\n\u003cp\u003eGiven that common variants have been extensively studied, our primary focus was on the imputation effect of reference panels for low-frequency (1% \u0026lt; MAF\u0026thinsp;\u0026lt;\u0026thinsp;5%) and rare variants (MAF\u0026thinsp;\u0026lt;\u0026thinsp;1%). The SEAD reference panel had obtained 3,209,490 well-imputed low/rare sites, and as many as twice rare mutations compared to the 1kGP panel (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eA and Supplementary Table\u0026nbsp;8). Across the 7 MAF bins for rare and low-frequency variants, as MAF decreased, the SEAD reference panel exhibited a higher number of well-imputed sites compared to other panels (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eB and Supplementary Table\u0026nbsp;8). Especially in the intervals of MAF\u0026thinsp;\u0026lt;\u0026thinsp;0.1% and 0.1% \u0026lt; MAF\u0026thinsp;\u0026lt;\u0026thinsp;0.3%, the imputation advantage of SEAD was most pronounced. (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eB and Supplementary Table\u0026nbsp;8). To make the results comparable across the five reference panels, we also extracted a total of 8,525,136 shared SNPs for analysis. Within the 7 MAF bins, the SEAD reference panel always exhibited the highest Rsq value and largest number of well-quality sites (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eC and Supplementary Table\u0026nbsp;9\u0026ndash;10).\u003c/p\u003e\n\u003cp\u003eAs 184 samples underwent both whole-genome sequencing and array genotyping, we further assessed the imputation concordance of the five panels. The imputation accuracy of the SEAD, WBBC, and SG10K reference panels was considerably superior to that of the GASP and 1KGP panels. The concordance rate values corresponding to the peak of the SEAD imputation panel were slightly higher than those of WBBC and SG10K, but the differences were minimal (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003eD).\u003c/p\u003e\n\u003cp\u003eEmployment of the SEAD reference panel in bone mineral density GWAS analysis\u003c/p\u003e\n\u003cp\u003eAfter imputation with SEAD panel for WBBC array data, we conducted GWAS analysis on two BMD traits (total hip and femoral neck BMD) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003eC). These two BMD traits were highly correlated in either Pearson correlation (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.918) or genetic correlation (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.932, SD\u0026thinsp;=\u0026thinsp;0.025). With genome-wide significance threshold (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;5\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e), two common loci (MAF\u0026thinsp;\u0026gt;\u0026thinsp;0.01) that have been previously reported in GWAS were identified to be associated with both BMD traits, with the top signals at chr1: rs9659023 (an intronic variant of \u003cem\u003eFMN2\u003c/em\u003e, this locus was reported to be associated with heel (\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e) and total body BMD (\u003cspan class=\"CitationRef\"\u003e47\u003c/span\u003e)) and at chr6: rs138367190 (near \u003cem\u003eSOX4\u003c/em\u003e, this locus was reported to be associated with heel (\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e) and FN BMD (\u003cspan class=\"CitationRef\"\u003e12\u003c/span\u003e)) (Supplementary Fig.\u0026nbsp;5).\u003c/p\u003e\n\u003cp\u003eFor the relative rare variants (0.001\u0026thinsp;\u0026lt;\u0026thinsp;MAF\u0026thinsp;\u0026lt;\u0026thinsp;0.01), we utilized two analytical strategies: single-variant test with PLINK and spatial clustering analysis via the STAARpipeline (incorporating both sliding-window and gene-centric methods). In the single-variant test analysis, we prioritized the variants with small \u003cem\u003eP\u003c/em\u003e-value and with replication (see Methods), and identified 95 suggestive variants for Hip BMD (Supplementary Table\u0026nbsp;11). Among these variants, 70 (73.7%) were annotated to \u003cem\u003eSNTG1\u003c/em\u003e gene by ANNOVAR (Supplementary Table\u0026nbsp;11). Of which, four rare variants (rs60103302, rs61260287, rs60600379, and rs57319781), in complete linkage disequilibrium (LD r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;1.00), were identified for Hip BMD at genome-wide significance level (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4.79\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e, MAF\u0026thinsp;=\u0026thinsp;0.0091), these variants located in the intergenic region near \u003cem\u003eSNTG1\u003c/em\u003e on chromosome 8 (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eA and \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eF, and Supplementary Table\u0026nbsp;11). As for FN BMD, 18 (46.2%) out of the 39 suggestive variants were annotated to \u003cem\u003eSNTG1\u003c/em\u003e gene (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eC and \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eF, and Supplementary Table\u0026nbsp;11). In the spatial clustering analysis, we first implemented a sliding window of 2kb in STAAR-O test and adhered to the previous \u0026ldquo;discovery and replication\" strategy. The results showed that 5 clustering regions were identified as the suggestive signals for Hip BMD (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eB and Supplementary Table\u0026nbsp;12). In the cluster on chromosome 8 (chr8:49873459\u0026ndash;49958458), a total of 22 windows were suggested, of which, 3 windows annotated near \u003cem\u003eSNTG1\u003c/em\u003e were discovered at bonferroni-corrected level \u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;3.59\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e (0.05 divided by 1,392,296 windows), with top window at chr8:49898459\u0026ndash;49900458 (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;1.08\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eB and Supplementary Table\u0026nbsp;12). Moreover, in the annotation-based gene-centric analysis, of the 7 non-coding categories of \u003cem\u003eSNTG1\u003c/em\u003e gene, 5 categories showed suggestive signal for Hip BMD, with the top signal at promoter region (\u003cem\u003eP\u003c/em\u003e\u003csub\u003e\u003cem\u003epromoter_CAGE\u003c/em\u003e\u003c/sub\u003e = 4.72\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eE and Supplementary Table\u0026nbsp;13), showcasing at least one annotation reaching the genome-wide bonferroni-corrected level \u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;3.57\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e (0.05 divided by 20,000 genes then divided by 7 categories). In summary, \u003cem\u003eSNTG1\u003c/em\u003e locus was highlighted in both analytical strategies with the strictest threshold applied to Hip BMD, and also achieved a suggestive level in FN BMD (\u003cem\u003eP\u003c/em\u003e\u003csub\u003e\u003cem\u003esingle_variant\u003c/em\u003e\u003c/sub\u003e = 4.85\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e, \u003cem\u003eP\u003c/em\u003e\u003csub\u003e\u003cem\u003eslide_window\u003c/em\u003e\u003c/sub\u003e = 2.36\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e, \u003cem\u003eP\u003c/em\u003e\u003csub\u003e\u003cem\u003egene_centric\u003c/em\u003e\u003c/sub\u003e = 5.43\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eC, \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eD, \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eE, and Supplementary Tables\u0026nbsp;12 and 13),\u003c/p\u003e\n\u003cp\u003eWe further exploited the preliminary function of \u003cem\u003eSNTG1\u003c/em\u003e gene. Dual-luciferase assays were then performed in two cell lines, 293T and MC3T3-E1. It was observed that the intronic SNP rs111829635 C-T mutation significantly increased luciferase activity in both 293T cells and MC3T3-E1 cells (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eG), suggesting that the identified locus would alter the \u003cem\u003eSNTG1\u003c/em\u003e gene expression. Moreover, the results from the CCK-8 proliferation assay showed that the absorbance at OD 450 nm was decreased after \u003cem\u003eSNTG1\u003c/em\u003e overexpression, and cell density was also decreased (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eH and Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eI), revealing that overexpression of \u003cem\u003eSNTG1\u003c/em\u003e led to a significant reduction in cell proliferation. At the 72 hours, a noticeable decrease in the expression of osteogenic marker genes (\u003cem\u003eRUNX2\u003c/em\u003e, \u003cem\u003eCOL1A1\u003c/em\u003e and OCN) confirmed that \u003cem\u003eSNTG1\u003c/em\u003e inhibited the differentiation of preosteoblast MC3T3-E1 cells (\u003cem\u003eP\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05, Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eJ). Furthermore, there was also a significant reduction in alkaline phosphatase (ALP) activity observed after six days. Overall, these results supported the notion that \u003cem\u003eSNTG1\u003c/em\u003e overexpression might inhibit human osteogenic proliferation and differentiation.\u003c/p\u003e\n\u003cp\u003eOther suggestive candidate genes and loci\u003c/p\u003e\n\u003cp\u003eBesides the \u003cem\u003eSNTG1\u003c/em\u003e locus, other suggestive loci were ranked in a credible priority order: 3 loci (\u003cem\u003eGRM7\u003c/em\u003e, \u003cem\u003eCPT2\u003c/em\u003e, \u003cem\u003eLINC00407\u003c/em\u003e) were identified by both single-variant test and spatial clustering analyses (level A), and 6 loci (\u003cem\u003eSH3TC2\u003c/em\u003e, \u003cem\u003eC4BPA\u003c/em\u003e, \u003cem\u003eLINC02572\u003c/em\u003e, \u003cem\u003eLINC00461\u003c/em\u003e, \u003cem\u003eMORC2\u003c/em\u003e, \u003cem\u003eADIG\u003c/em\u003e) were identified in both BMD traits by either method (level B) (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eA, \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eB, \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eC, \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003eD, and Supplementary Table\u0026nbsp;14). Of these, 7 were annotated near genes, three candidates were lincRNAs. We then employed the protein-protein interaction analysis to include the known BMD-associated genes reported in previously GWAS (\u003cspan class=\"CitationRef\"\u003e42\u003c/span\u003e) and suggested that 19 known genes directly associated with the 7 candidate genes (Supplementary Table\u0026nbsp;15). Functional enrichment of directly associated proteins revealed 24 potential categories with a false discovery rate (FDR)\u0026thinsp;\u0026lt;\u0026thinsp;=\u0026thinsp;0.05 (Supplementary Table\u0026nbsp;16). Among these enrichment analysis, we prioritized four functional annotations closely related to bone biology (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003eA), including \u0026ldquo;abnormal skeletal muscle morphology\u0026rdquo; (\u0026lsquo;Monarch Human Phenotype\u0026rsquo;, FDR\u0026thinsp;=\u0026thinsp;0.0246), \u0026ldquo;Dystrophin-associated glycoprotein complex\u0026rdquo; (\u0026lsquo;GO Cellular Component\u0026rsquo;, FDR\u0026thinsp;=\u0026thinsp;0.0047), \u0026ldquo;Fatty acid metabolic process\u0026rdquo; (\u0026lsquo;GO Biological Process\u0026rsquo;, FDR\u0026thinsp;=\u0026thinsp;0.0493) and \u0026ldquo;Complement and coagulation cascades\u0026rdquo; (\u0026lsquo;KEGG pathway\u0026rsquo;, FDR\u0026thinsp;=\u0026thinsp;0.0011).\u003c/p\u003e\n\u003cp\u003eWe compared the frequency distribution of all the rare variants identified in level-A and -B (84 SNPs) with the gnomAD and TOPMed databases across various populations (Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003eB, Supplementary Table\u0026nbsp;17). The MAF of these rare variants most closely aligned with the MAF value in the East Asian (EAS) population (2,604 genomes) within gnomAD, validating the appropriateness of our panel for imputing genotype data in East Asian populations. Furthermore, the \u003cem\u003eSNTG1\u003c/em\u003e variants were predominantly rare in the Non-Finnish European (NFE: 34,029 genomes) and South Asian (SAS: 2,419 genomes) populations, while they were more common in the African/African American (AFR: 20,744 genomes) and Latino/Admixed American (AMR: 7,647 genomes) populations, highlighting population-specific patterns. Additionally, the other variants were present in either the SAS or EAS population but were largely absent in other populations, further illustrating the panel's ability to detect variations specific to East Asian and South Asian populations.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eIn this study, we integrated whole-genome sequencing data from SG10K, GenomeAsia, WBBC, and 1kGP-Asian to create a combined reference panel for genotype imputation, the South and East Asian Reference Database (SEAD). It comprised a diverse range of populations across Asia, including 11 populations from GenomeAsia, one population from WBBC, three populations from SG10K, and 8 populations from 1kGP-Asian. With a sample size of 22,134 haplotypes, the SEAD panel stands as one of the most comprehensive panels in terms of coverage across Asia. We compared the concordance rate of the genotypes imputed from the SEAD panel across populations around the world, and suggested that the SEAD panel performed the best for the rare variants imputation in Asian populations. The SEAD panel also had advantage in detecting disease-associated low/rare variants. Subsequently, we applied the SEAD panel to the bone mineral density GWAS analyses in WBBC array cohorts, and identified a rare locus, \u003cem\u003eSNTG1\u003c/em\u003e, that was not reported even in the large biobank-scale GWAS.\u003c/p\u003e \u003cp\u003eAsia, being the largest and most populous continent worldwide, boasts a wealth of human genetic resources. However, most of the whole genome-sequencing (WGS) efforts were carried out in Caucasian populations in the last decade (\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e). In our previous study, we conducted a thorough evaluation and discussion on the imputation of rare variants, highlighting the necessity of constructing a haplotype reference panel for Asian populations (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e). In recent years, there has been a notable surge WGS data across Asia, particularly in regions such as Japan, Singapore, Korea and China. Despite the abundance of WGS projects, large-scale reference panel with broad geographical coverage across Asia is still needed. The accumulation of such vast datasets from diverse populations presents a unique opportunity: the combination of multiple WGS datasets to form a more comprehensive and expansive reference panel. The recently published Northeast Asian Reference Database, with over half of the samples deriving from Japan and Korea, represented populations from Northeast Asia (\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e). In our study, we integrated the data from WBBC with other Asian populations to construct the South and East Asian Reference Database (SEAD) reference panel, covering the widest geographical area in Asia, with the majority of its population originating from South and East Asia.\u003c/p\u003e \u003cp\u003eWith the SEAD panel, we first assessed the imputation performance across populations at the global scale by using HGDP database. We extracted the UK Biobank Axiom array sites as pseudo array from all HGDP populations, and extracted the Illumina Asian Screening Array sites from HGDP East Asian and Central and South Asian populations as pseudo array. Then the non-reference (NR) allele concordance rates were compared between true and imputed genotypes. It was not surprising that the SEAD reference panel outperformed other reference panels in imputing all East Asian populations and some of the Central and South Asian populations. Interestingly, the relative high concordance rate was achieved by the SEAD reference panel in the Oceania population, this might be attributable to shared haplotypes between the Oceania population and Singaporean Malays (\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e). Among all the Asian populations imputed by the SEAD panel, the NR concordance rate obtained from the pseudo array based on the Axiom array was consistently higher compared to that of the pseudo array extracted from Asian Screening array sites. These higher concordance rates could be attributed to the fact that the Axiom array contains approximately 150,000 more variants compared to the Asian Screening array.\u003c/p\u003e \u003cp\u003eAs WBBC cohort had samples that were both sequenced and genotyped, we then assessed the performance of the panels on Chinese population imputation. We found that the concordance rate values of the SEAD panel were slightly higher than those of WBBC and SG10K, but the differences were minimal. This suggested that the accuracy did not significantly improve after the combination, this is consistent with our pervious study that the imputation accuracy would decrease if the diverse sample exceed a set amount (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e). However, the SEAD panel outperformed the other four panels in terms of well-imputed variants, with almost twice the number of variants at MAF bin less than 0.1% compared to SG10K and WBBC imputation. Finally, if we focused on the variants that had been identified to be associated with diseases and traits in Asian population in GWAS catalog, still, the SEAD panel performed the best.\u003c/p\u003e \u003cp\u003eBesides the systematical and comprehensive imputation evaluation for the panels, we successfully applied the SEAD panel to a GWAS analysis for BMD traits at hip and femoral neck (FN). These two traits were highly correlated in either phenotypic or genetic correlation. Therefore, the association signals reported in both traits would reduce the random error of the results. Through both single-variant association test and spatial clustering analysis, we detected that rare variants near \u003cem\u003eSNTG1\u003c/em\u003e gene were associated with hip BMD. The \u003cem\u003eSNTG1\u003c/em\u003e gene also reached the suggestive threshold in FN BMD analysis using both methods. The gene and variants were not reported even in large-scale biobank GWAS for BMD (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e), and the \u003cem\u003eSNTG1\u003c/em\u003e variants were predominantly rare in the Non-Finnish European and more common in the African/African American and Latino/Admixed American. The \u003cem\u003eSNTG1\u003c/em\u003e (Syntrophin Gamma 1) gene is known to encode a cytoplasmic peripheral membrane protein, which a candidate gene for idiopathic scoliosis (\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e). A southern Chinese cohort of patients with congenital scoliosis also identified copy number variants of \u003cem\u003eSNTG1\u003c/em\u003e (\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e). It was also noteworthy that \u003cem\u003eSNTG1\u003c/em\u003e genetic variations had been associated with neurodevelopmental disorders, suggesting a potential role in developmental processes (\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e). The SNTG1 protein was involved in the Dystrophin-associated glycoprotein complex (DGC) as shown in our pathway enrichment analysis. The DGC was known to play a significant role in maintaining the structural integrity of muscle fibers (\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e). In our study, we investigated the preliminary function of this gene, and suggested that overexpression of \u003cem\u003eSNTG1\u003c/em\u003e inhibited the proliferation and differentiation of preosteoblasts. This observation aligned with the direction of effect on BMD in our GWAS results.\u003c/p\u003e \u003cp\u003eIn summary, we constructed the haplotype reference panel for Asian population (SEAD panel) with the largest number of samples and the most abundant population diversity in Asia. The reference panel demonstrated excellent performance on imputing East Asian and Central and South Asian populations, especially in detecting rare variants. We provided this optimal imputation service online for free (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://imputationserver.westlake.edu.cn/\u003c/span\u003e\u003cspan address=\"https://imputationserver.westlake.edu.cn/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) for genetic studies in Asian populations. By applying the SEAD panel to impute the genotyping array data in Chinese population, we, for the first time, successfully identified rare variants near \u003cem\u003eSNTG1\u003c/em\u003e gene showing association with bone mineral density.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eData availability\u003c/p\u003e\n\u003cp\u003eGWAS summary statistics for the Hip and FN BMD from our analysis of the WBBC data are fully available at https://wbbc.westlake.edu.cn/downloads.html\u003c/p\u003e\n\u003cp\u003eAcknowledgements\u003c/p\u003e\n\u003cp\u003eWe thank the \u0026ldquo;SG10K_Pilot Investigators\u0026rdquo; for providing the SG10K_Pilot data (EGAD00001005337). The data from the \u0026ldquo;SG10K_Pilot Study\u0026rdquo; reported here were obtained from EGA. This manuscript was not prepared in collaboration with the \u0026ldquo;SG10K_Pilot Study\u0026rdquo; and does not necessarily reflect the opinions or views of the \u0026ldquo;SG10K_Pilot Study\u0026rdquo;. We also thank the \u0026ldquo;Genome Asia 100K consortium\u0026rdquo; for providing the \u0026ldquo;GenomeAsia pilot project\u0026rdquo; (EGAS00001002921). We thankfully acknowledge the High-performance Computing Center at Westlake University.\u003c/p\u003e\n\u003cp\u003eFunds\u003c/p\u003e\n\u003cp\u003eThis work was supported by the \u0026quot;Pioneer\u0026quot; and \u0026quot;Leading Goose\u0026quot; R\u0026amp;D Program of Zhejiang (#2023C03164), the National Natural Science Foundation of China (#82370887), the Chinese National Key Technology R\u0026amp;D Program, Ministry of Science and Technology (#2021YFC2501702), and the funds from the Westlake Laboratory of Life Sciences and Biomedicine (#202208014).\u003c/p\u003e\n\u003cp\u003eAuthor contributions\u003c/p\u003e\n\u003cp\u003eH.-F.Z. conceptualized and designed the study. M.-Y.Y., J.-D.Z., and W.-Y.B. conducted the data analysis. X.L. conducted experiments. S.-H.Y., W.-W.Z., J.-Q.L. and Y. S. conducted the whole sequencing experiments. C.-D.Y., M.-C.Q., K.-Q.L., C.-F.W., P.-K.C., K.S., S.-R.G., P.-P.Z., P.-L.G., Y.Q., J.-G.T., X.-J.Y., J.-X.G., X.C., M.-M.L., L.-X.L., G.T., S.-Y.X., L.X., F.H., J.-C.L., J.-F.G., B.-S.T., L.Y. and D.K. contributed to the sample collection, processing and preliminary data analysis. J.-J.Y., Y.-H.L. and N.L. designed the online website resource. M.-Y.Y. and J.-D.Z. drafted the manuscript, H.-F.Z. reviewed and edited manuscript. All authors contributed, discussed and approved manuscript.\u003c/p\u003e\n\u003cp\u003eCompeting interests\u003c/p\u003e\n\u003cp\u003eS.-H.Y., W.-W.Z. and J.-Q.L. Y.S are employee of KingMed Diagnostics Co., Ltd. The other authors have no conflict of interest to declare.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMahajan A, Taliun D, Thurner M, Robertson NR, Torres JM, Rayner NW, et al. Fine-mapping type 2 diabetes loci to single-variant resolution using high-density imputation and islet-specific epigenome maps. Nat Genet. 2018;50(11):1505-13.\u003c/li\u003e\n\u003cli\u003eZheng HF, Duncan EL, Yerges-Armstrong LM, Eriksson J, Bergstrom U, Leo PJ, et al. Meta-analysis of genome-wide studies identifies MEF2C SNPs associated with bone mineral density at forearm. J Med Genet. 2013;50(7):473-8.\u003c/li\u003e\n\u003cli\u003eHoffmann TJ, Sakoda LC, Shen L, Jorgenson E, Habel LA, Liu J, et al. Imputation of the rare HOXB13 G84E mutation and cancer risk in a large population-based cohort. PLoS Genet. 2015;11(1):e1004930.\u003c/li\u003e\n\u003cli\u003eHandsaker RE, Van Doren V, Berman JR, Genovese G, Kashin S, Boettger LM, et al. Large multiallelic copy number variations in humans. Nat Genet. 2015;47(3):296-303.\u003c/li\u003e\n\u003cli\u003eDas S, Abecasis GR, Browning BL. Genotype Imputation from Large Reference Panels. Annu Rev Genomics Hum Genet. 2018;19:73-96.\u003c/li\u003e\n\u003cli\u003eNelson SC, Stilp AM, Papanicolaou GJ, Taylor KD, Rotter JI, Thornton TA, et al. Improved imputation accuracy in Hispanic/Latino populations with larger and more diverse reference panels: applications in the Hispanic Community Health Study/Study of Latinos (HCHS/SOL). Hum Mol Genet. 2016;25(15):3245-54.\u003c/li\u003e\n\u003cli\u003eLert-Itthiporn W, Suktitipat B, Grove H, Sakuntabhai A, Malasit P, Tangthawornchaikul N, et al. Validation of genotype imputation in Southeast Asian populations and the effect of single nucleotide polymorphism annotation on imputation outcome. BMC Med Genet. 2018;19(1):23.\u003c/li\u003e\n\u003cli\u003eVergara C, Parker MM, Franco L, Cho MH, Valencia-Duarte AV, Beaty TH, et al. Genotype imputation performance of three reference panels using African ancestry individuals. Hum Genet. 2018;137(4):281-92.\u003c/li\u003e\n\u003cli\u003eBai WY, Zhu XW, Cong PK, Zhang XJ, Richards JB, Zheng HF. Genotype imputation and reference panel: a systematic evaluation on haplotype size and diversity. Brief Bioinform. 2019.\u003c/li\u003e\n\u003cli\u003eGenomes Project C, Auton A, Brooks LD, Durbin RM, Garrison EP, Kang HM, et al. A global reference for human genetic variation. Nature. 2015;526(7571):68-74.\u003c/li\u003e\n\u003cli\u003eGenome of the Netherlands C. Whole-genome sequence variation, population structure and demographic history of the Dutch population. Nat Genet. 2014;46(8):818-25.\u003c/li\u003e\n\u003cli\u003eZheng HF, Forgetta V, Hsu YH, Estrada K, Rosello-Diez A, Leo PJ, et al. Whole-genome sequencing identifies EN1 as a determinant of bone density and fracture. Nature. 2015;526(7571):112-7.\u003c/li\u003e\n\u003cli\u003eConsortium UK, Walter K, Min JL, Huang J, Crooks L, Memari Y, et al. The UK10K project identifies rare variants in health and disease. Nature. 2015;526(7571):82-90.\u003c/li\u003e\n\u003cli\u003eJun G, English AC, Metcalf GA, Yang J, Chaisson MJ, Pankratz N, et al. Structural variation across 138,134 samples in the TOPMed consortium. bioRxiv. 2023.\u003c/li\u003e\n\u003cli\u003eMcCarthy S, Das S, Kretzschmar W, Delaneau O, Wood AR, Teumer A, et al. A reference panel of 64,976 haplotypes for genotype imputation. Nat Genet. 2016;48(10):1279-83.\u003c/li\u003e\n\u003cli\u003eJones KM, Cook-Deegan R. Complicated legacies: The human genome at 20. Science. 2021;371(6529):564-9.\u003c/li\u003e\n\u003cli\u003eNagasaki M, Yasuda J, Katsuoka F, Nariai N, Kojima K, Kawai Y, et al. Rare variant discovery by deep whole-genome sequencing of 1,070 Japanese individuals. Nat Commun. 2015;6:8018.\u003c/li\u003e\n\u003cli\u003eJeon S, Bhak Y, Choi Y, Jeon Y, Kim S, Jang J, et al. Korean Genome Project: 1094 Korean personal genomes with clinical information. Science Advances. 2020;6(eaaz7835).\u003c/li\u003e\n\u003cli\u003eWu D, Dou J, Chai X, Bellis C, Wilm A, Shih CC, et al. Large-Scale Whole-Genome Sequencing of Three Diverse Asian Populations in Singapore. Cell. 2019;179(3):736-49 e15.\u003c/li\u003e\n\u003cli\u003eGenomeAsia KC. The GenomeAsia 100K Project enables genetic discoveries across Asia. Nature. 2019;576(7785):106-11.\u003c/li\u003e\n\u003cli\u003eZhang P, Luo H, Li Y, Wang Y, Wang J, Zheng Y, et al. NyuWa Genome resource: A deep whole-genome sequencing-based variation profile and reference panel for the Chinese population. Cell Rep. 2021;37(7):110017.\u003c/li\u003e\n\u003cli\u003eLi L, Huang P, Sun X, Wang S, Xu M, Liu S, et al. The ChinaMAP reference panel for the accurate genotype imputation in Chinese populations. Cell Res. 2021;31(12):1308-10.\u003c/li\u003e\n\u003cli\u003eCong PK, Bai WY, Li JC, Yang MY, Khederzadeh S, Gai SR, et al. Genomic analyses of 10,376 individuals in the Westlake BioBank for Chinese (WBBC) pilot project. Nat Commun. 2022;13(1):2939.\u003c/li\u003e\n\u003cli\u003eWang C, Dai J, Qin N, Fan J, Ma H, Chen C, et al. Analyses of rare predisposing variants of lung cancer in 6,004 whole genomes in Chinese. Cancer Cell. 2022;40(10):1223-39 e6.\u003c/li\u003e\n\u003cli\u003eZhu XW, Liu KQ, Wang PY, Liu JQ, Chen JY, Xu XJ, et al. Cohort profile: the Westlake BioBank for Chinese (WBBC) pilot project. BMJ Open. 2021;11(6):e045564.\u003c/li\u003e\n\u003cli\u003eCong PK, Khederzadeh S, Yuan CD, Ma RJ, Zhang YY, Liu JQ, et al. Identification of clinically actionable secondary genetic variants from whole-genome sequencing in a large-scale Chinese population. Clin Transl Med. 2022;12(5):e866.\u003c/li\u003e\n\u003cli\u003eGibson G. Rare and common variants: twenty arguments. Nat Rev Genet. 2012;13(2):135-45.\u003c/li\u003e\n\u003cli\u003eAlioto TS, Buchhalter I, Derdak S, Hutter B, Eldridge MD, Hovig E, et al. A comprehensive assessment of somatic mutation detection in cancer using whole-genome sequencing. Nat Commun. 2015;6:10001.\u003c/li\u003e\n\u003cli\u003eByrska-Bishop M, Evani US, Zhao X, Basile AO, Abel HJ, Regier AA, et al. High-coverage whole-genome sequencing of the expanded 1000 Genomes Project cohort including 602 trios. Cell. 2022;185(18):3426-40 e19.\u003c/li\u003e\n\u003cli\u003eHinrichs AS, Karolchik D, Baertsch R, Barber GP, Bejerano G, Clawson H, et al. The UCSC Genome Browser Database: update 2006. Nucleic Acids Res. 2006;34(Database issue):D590-8.\u003c/li\u003e\n\u003cli\u003eLi H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N, et al. The Sequence Alignment/Map format and SAMtools. Bioinformatics. 2009;25(16):2078-9.\u003c/li\u003e\n\u003cli\u003eManichaikul A, Mychaleckyj JC, Rich SS, Daly K, Sale M, Chen WM. Robust relationship inference in genome-wide association studies. Bioinformatics. 2010;26(22):2867-73.\u003c/li\u003e\n\u003cli\u003eBergstrom A, McCarthy SA, Hui R, Almarri MA, Ayub Q, Danecek P, et al. Insights into human genetic variation and population history from 929 diverse genomes. Science. 2020;367(6484).\u003c/li\u003e\n\u003cli\u003eYang J, Lee SH, Goddard ME, Visscher PM. GCTA: a tool for genome-wide complex trait analysis. Am J Hum Genet. 2011;88(1):76-82.\u003c/li\u003e\n\u003cli\u003eChang CC, Chow CC, Tellier LC, Vattikuti S, Purcell SM, Lee JJ. Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience. 2015;4:7.\u003c/li\u003e\n\u003cli\u003eDas S, Forer L, Schonherr S, Sidore C, Locke AE, Kwong A, et al. Next-generation genotype imputation service and methods. Nat Genet. 2016;48(10):1284-7.\u003c/li\u003e\n\u003cli\u003eFuchsberger C, Abecasis GR, Hinds DA. minimac2: faster genotype imputation. Bioinformatics. 2015;31(5):782-4.\u003c/li\u003e\n\u003cli\u003eWiller CJ, Li Y, Abecasis GR. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics. 2010;26(17):2190-1.\u003c/li\u003e\n\u003cli\u003eLi Z, Li X, Zhou H, Gaynor SM, Selvaraj MS, Arapoglou T, et al. A framework for detecting noncoding rare-variant associations of large-scale whole-genome sequencing studies. Nat Methods. 2022;19(12):1599-611.\u003c/li\u003e\n\u003cli\u003eBai WY, Wang L, Ying ZM, Hu B, Xu L, Zhang GQ, et al. Identification of PIEZO1 polymorphisms for human bone mineral density. Bone. 2020;133:115247.\u003c/li\u003e\n\u003cli\u003eMatsuda T, Cepko CL. Electroporation and RNA interference in the rodent retina in vivo and in vitro. Proceedings of the National Academy of Sciences of the United States of America. 2004;101(1):16-22.\u003c/li\u003e\n\u003cli\u003eMorris JA, Kemp JP, Youlten SE, Laurent L, Logan JG, Chai RC, et al. An atlas of genetic influences on osteoporosis in humans and mice. Nat Genet. 2019;51(2):258-66.\u003c/li\u003e\n\u003cli\u003eSzklarczyk D, Kirsch R, Koutrouli M, Nastou K, Mehryary F, Hachilif R, et al. The STRING database in 2023: protein-protein association networks and functional enrichment analyses for any sequenced genome of interest. Nucleic Acids Res. 2023;51(D1):D638-D46.\u003c/li\u003e\n\u003cli\u003eShannon P, Markiel A, Ozier O, Baliga NS, Wang JT, Ramage D, et al. Cytoscape: a software environment for integrated models of biomolecular interaction networks. Genome Res. 2003;13(11):2498-504.\u003c/li\u003e\n\u003cli\u003eHuang J, Howie B, McCarthy S, Memari Y, Walter K, Min JL, et al. Improved imputation of low-frequency and rare variants using the UK10K haplotype reference panel. Nat Commun. 2015;6:8111.\u003c/li\u003e\n\u003cli\u003eSollis E, Mosaku A, Abid A, Buniello A, Cerezo M, Gil L, et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51(D1):D977-D85.\u003c/li\u003e\n\u003cli\u003eMedina-Gomez C, Kemp JP, Trajanoska K, Luan J, Chesi A, Ahluwalia TS, et al. Life-Course Genome-wide Association Study Meta-analysis of Total Body BMD and Assessment of Age-Specific Effects. Am J Hum Genet. 2018;102(1):88-102.\u003c/li\u003e\n\u003cli\u003eZheng HF, Ladouceur M, Greenwood CM, Richards JB. Effect of genome-wide genotyping and reference panels on rare variants imputation. J Genet Genomics. 2012;39(10):545-50.\u003c/li\u003e\n\u003cli\u003eChoi J, Kim S, Kim J, Son H-Y, Yoo S-K, Kim C-U, et al. A whole-genome reference panel of 14,393 individuals for East Asian populations accelerates discovery of rare functional variants. SCIENCE ADVANCES. 2023;9(eadg6319 ).\u003c/li\u003e\n\u003cli\u003eBashiardes S, Veile R, Allen M, Wise CA, Dobbs M, Morcuende JA, et al. SNTG1, the gene encoding gamma1-syntrophin: a candidate gene for idiopathic scoliosis. Hum Genet. 2004;115(1):81-9.\u003c/li\u003e\n\u003cli\u003eLai W, Feng X, Yue M, Cheung PWH, Choi VNT, Song YQ, et al. Identification of Copy Number Variants in a Southern Chinese Cohort of Patients with Congenital Scoliosis. Genes (Basel). 2021;12(8).\u003c/li\u003e\n\u003cli\u003eLemos RR, Oliveira DF, Zatz M, Oliveira JR. Population and computational analysis of the MGEA6 P521A variation as a risk factor for familial idiopathic basal ganglia calcification (Fahr\u0026apos;s disease). J Mol Neurosci. 2011;43(3):333-6.\u003c/li\u003e\n\u003cli\u003eOmairi S, Hau KL, Collins-Hooper H, Scott C, Vaiyapuri S, Torelli S, et al. Regulation of the dystrophin-associated glycoprotein complex composition by the metabolic properties of muscle fibres. Sci Rep. 2019;9(1):2770.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"reference panel, GWAS, rare variants, hip BMD, FN BMD","lastPublishedDoi":"10.21203/rs.3.rs-3680930/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3680930/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHere, we present the South and East Asian Reference Database (SEAD) reference panel (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://imputationserver.westlake.edu.cn/\u003c/span\u003e\u003cspan address=\"https://imputationserver.westlake.edu.cn/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), which comprises whole genome sequencing data from 11,067 individuals across 17 countries in Asia. The SEAD panel, which excludes singleton variants, consists of 22,134 haplotypes and 80,367,720 variants. Firstly, we assessed the concordance rate in global populations using HGDP datasets, notably, the SEAD panel showed advantage in East Asia, Central and South Asia, and Oceania populations. When imputing the disease-associated variants of Asian population, the SEAD panel displayed a distinct preponderance in imputing low-frequency and rare variants. In imputation of Chinese population, the SEAD panel imputed a larger number of well-imputed sites across all minor allele frequency (MAF) bins. Additionally, the SEAD panel exhibited higher imputation accuracy for shared sites in all MAF bins. Finally, we applied the augmented SEAD panel to conduct a discovery and replication genome-wide association study (GWAS) for hip and femoral neck (FN) bone mineral density (BMD) traits within the 5,369 Westlake BioBank for Chinese (WBBC) samples. The single-variant test suggests that rare variants near \u003cem\u003eSNTG1\u003c/em\u003e gene are associated with hip BMD (rs60103302, MAF\u0026thinsp;=\u0026thinsp;0.0091, \u003cem\u003eP\u003c/em\u003e\u0026thinsp;=\u0026thinsp;4.79\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e). The spatial clustering analysis also suggests the association of this gene (\u003cem\u003eP\u003c/em\u003e\u003csub\u003eslide_window\u003c/sub\u003e=1.08\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e, \u003cem\u003eP\u003c/em\u003e\u003csub\u003egene_centric\u003c/sub\u003e=4.72\u0026times;10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e). The gene and variants achieved a suggestive level for FN BMD. This gene was not reported previously, and the preliminary experiment demonstrated that the identified rare variant can upregulate the \u003cem\u003eSNTG1\u003c/em\u003e expression, which in turn inhibits the proliferation and differentiation of preosteoblast.\u003c/p\u003e","manuscriptTitle":"SEAD: an augmented reference panel with 22,134 haplotypes boosts the rare variants imputation and GWAS analysis in Asian population","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-01-29 21:51:30","doi":"10.21203/rs.3.rs-3680930/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"40abd18e-8277-45cc-bbb0-af364fe2d6b3","owner":[],"postedDate":"January 29th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":27101436,"name":"Biological sciences/Computational biology and bioinformatics/Genome informatics"},{"id":27101437,"name":"Biological sciences/Genetics/Genetic association study/Genome-wide association studies"}],"tags":[],"updatedAt":"2024-12-31T08:10:57+00:00","versionOfRecord":{"articleIdentity":"rs-3680930","link":"https://doi.org/10.1038/s41467-024-55147-4","journal":{"identity":"nature-communications","isVorOnly":false,"title":"Nature Communications"},"publishedOn":"2024-12-30 05:00:00","publishedOnDateReadable":"December 30th, 2024"},"versionCreatedAt":"2024-01-29 21:51:30","video":"","vorDoi":"10.1038/s41467-024-55147-4","vorDoiUrl":"https://doi.org/10.1038/s41467-024-55147-4","workflowStages":[]},"version":"v1","identity":"rs-3680930","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3680930","identity":"rs-3680930","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00