Keywords
TWAS, admixture, expression quantitative trait loci, eQTL, transcriptome, complex traits, polygenic score, local ancestry, GReX, cross-population
CADET is a transcriptome-wide association test for admixed cohorts that leverages local ancestry information to improve detection of trait-influencing genes. CADET outperformed existing methods using simulated data and, in UK Biobank data, the method uniquely identified 18 gene-trait hits supported by GWAS. Our findings highlight the importance of ancestry-aware analysis.
Introduction
While genome-wide association studies (GWASs) have been largely successful in identifying genetic variants associated with a wide range of complex traits and diseases, many of the top variants lie in non-protein coding regions of the genome. It has been estimated that up to 90% of GWAS-identified single-nucleotide polymorphisms (SNPs) are non-coding variants, and thus the biological mechanisms by which these variants exert their effects on a phenotype remain unclear.1 Transcriptome-wide association studies (TWASs) seek to elucidate whether the underlying biological mechanisms of such GWAS hits are due to complex gene-regulatory effects. TWASs have achieved notable successes in the studies of various complex traits and diseases, including cancers, nervous and immune system disorders, anthropomorphic measurements, and many others.2,3,4,5,6,7,8,9,10,11
A typical TWAS assumes two datasets: a training dataset (usually of modest sample size) that possesses genotype and gene expression data from a tissue related to the outcome of interest, and a target (GWAS) dataset that possesses genotype and outcome data but lacks expression data. TWASs involve a two-stage process. Stage I constructs a model on the training dataset to estimate the effects of genetic variants on expression levels of a target gene (expression quantitative trait locus weight, or eQTL weight). While traditional TWAS methods assume individual-level training data,7,10,12,13,14,15 we note that recent methods such as OTTERS allow more flexible use of summary-level eQTL training data for model fitting based on polygenic risk score (PRS) methodology.16 After one fits training models for stage I (using individual- or summary-level eQTL data), stage II uses the resulting eQTL weights to impute genetically regulated gene expression (GReX) within subjects from the target (GWAS) dataset and then tests for association between the imputed expression and the outcome of interest.
Existing TWASs generally assume that the training and target datasets are compiled from the same ancestrally homogeneous population. Given the increased collection and analysis of multi-ancestry subjects for the target GWAS data, there is need to expand the TWAS framework to handle genetic and genomic data from diverse groups. However, to develop methodology for multi-ancestry groups, we must consider that the underlying genetic architecture of gene expression may not be the same across populations and may differ by ancestral group.17 Applied work has shown that gene expression prediction models trained in one ancestral population do not generalize well to other populations.18 The gene expression prediction models used in such work are similar to PRS methods, which are known to have poor transferability across ancestral groups.19,20,21,22,23 It is hypothesized that this poor portability may be a function of population-specific effect sizes, which can arise from a multitude of factors including differences in allele frequencies or linkage disequilibrium (LD) patterns across populations, as well as gene-gene or gene-environment interaction effects.24,25,26,27
The development of sufficiently powered TWAS methods for diverse ancestries is particularly important for admixed individuals, whose genomes are a unique mosaic of multiple continental ancestral groups. Admixed groups account for a large and growing proportion of the United States population, with more than 33.8 million people identifying as multi-racial in the 2020 Census.28 It has been common practice over the past few decades in GWASs to exclude admixed individuals from consideration due to their complex ancestral makeup, and, as such, they are a historically under-represented group in genetic studies. In response to this realization, researchers have recently considered the utility of including local ancestry information in variant-level genetic association analyses of admixed populations.29,30,31,32,33,34,35,36,37,38,39 Building on these advancements, there further has been increased development of novel PRS approaches specifically designed for admixed populations that we can leverage for related TWASs.40,41,42 For example, Marnetto et al. developed an ancestry-aware approach for PRS that first deconvoluted admixed haplotypes for a target subject and then computed the subject’s ancestry-specific components of the PRS using GWAS summary data from the appropriate reference ancestral populations. The authors demonstrated that this approach can yield not only improved phenotype predictability over standard methods but also an unbiased distribution of PRS in recently admixed populations.40
Given the existing evidence in the literature for differential eQTL architecture between admixed groups and more ancestrally homogeneous subjects,17,43 we propose CADET (combining ancestry-deconvoluted expression in TWAS) for enhanced TWAS in admixed samples, which leverages local ancestry information as well as summary-level eQTL data from multiple reference datasets of differing ancestry. With this method, we first apply the gene expression training models used by OTTERS separately to each reference dataset. We then apply a local ancestry deconvolution method to admixed target GWAS samples and, for a given gene, impute ancestry-specific partial GReX following the method of Marnetto et al.40 We combine the vectors of ancestry-specific GReX into an aggregate ancestry-aware measure of GReX. We can then test whether the aggregate ancestry GReX (as well as ancestry-specific GReX) are associated with outcome. We further combine the aggregate and ancestry-specific results together into an omnibus test using the aggregated Cauchy association test (ACAT).44 We evaluate expression imputation accuracy and power of CADET in simulations. The procedure is well calibrated under the null hypothesis of GReX-phenotype independence and achieves optimal power in the majority of genetic architecture settings. We conclude with an application of our method to 29 blood biochemistry phenotypes in admixed individuals of African and European ancestry in the UK Biobank (UKB) and identify 18 significant gene-trait associations unique to our ancestry-aware strategy.
Material and methods
Overview
The overarching goal of CADET is to impute gene expression in GWAS data of admixed individuals using summary-level eQTL data from individuals of the corresponding founder (continental) ancestral groups. Like most established TWASs, this involves two stages. Stage I focuses on using OTTERS to train models of predicted gene expression in each of the reference summary eQTL datasets under consideration. Stage II then uses estimates from these trained models to construct both local ancestry (LA)-aware and LA-unaware estimates of GReX (see overview of imputation workflow in Figure 1). We then test each LA-aware and LA-unaware GReX with an outcome of interest and aggregate results across tests into an optimal omnibus statistic based on the Cauchy distribution (see overview of downstream workflow in Figure 2). We now describe how we construct our LA-aware and -unaware measures of GReX for admixed target subjects.
Modeling expression in admixed individuals
We first describe how we model gene expression in admixed subjects, which forms the backbone of our technique for estimating LA-aware GReX in the target GWAS. Assume admixed individuals are two-way admixed, e.g., of African (AFR) and European (EUR) descent. For a given gene g, we assume there are S cis-eQTLs that are shared between the two ancestral groups and U cis-eQTLs that are unique to only one ancestral group. We define the total number of cis-eQTLs for a given gene as .
For , let be the number of minor alleles for the vth eQTL of subject i on the subject’s maternal (paternal) haplotype. Further, let be an indicator variable that takes the value of 1 when the maternal (paternal) haplotype of subject i is of AFR origin and likewise define as a similar indicator variable for denoting when the maternal (paternal) haplotype of subject i is of EUR origin. Based on this notation, we can define to represent the number of AFR-ancestry minor alleles of the vth eQTL and likewise define as the number of EUR-ancestry minor alleles at the eQTL for subject i. We can arrange these quantities across all subjects and all V eQTLs by creating matrices and , each of dimension . In these matrices, we further assume each column has been centered to mean zero. We can therefore model the gene expression vector of our target population as follows:
| (Equation 1) |
Here, is the vector of gene expression in our admixed target population. and are the diagonal scaling matrices with diagonal elements , where is the MAF of the vth eQTL in reference population . represents the gene expression heritability in our admixed subjects. are the vectors of causal eQTL effect sizes in each population. is the error vector, while is the identity matrix. Using these quantities, we can derive both the joint distribution of all eQTL effect sizes and an estimate for the heritability of gene expression in our admixed subjects (see supplemental information).
Stage I: Reference expression model training via OTTERS
Let us now assume, for our stage I training datasets, we have summary-level eQTL data derived from individuals in each of the two reference populations that represent the source ancestries among our two-way admixed target population. In other words, we require eQTL summary data from an AFR cohort and an EUR cohort for imputing gene expression in African American individuals. We note that need not be the same for both AFR and EUR groups, but we assume so for ease of presentation.
In each homogeneous ancestral population (AFR and EUR), we can model GReX of each gene g as we did for our admixed subjects in the preceding section. As above, expression is a function of the V cis-eQTLs with non-zero effect sizes in each population.
| (Equation 2) |
| (Equation 3) |
In the above equations, and are the vectors of gene expression in our AFR and EUR reference populations. and are the matrices of genotypes (0/1/2) with columns centered and standardized by minor allele frequency (MAF) in each population. are the vectors of ancestry-specific eQTL effect sizes. are the error vectors, while is the identity matrix. represent the gene expression heritability in each population.
A priori, we do not know the true V cis-eQTLs described in Equations 2 and 3. Instead, we attempt to model gene expression using eQTL summary information available from cis-SNPs found in and around gene g. That is, we have summary-level results (estimated effect sizes and p values) from the following single-variant linear regression models for each cis-SNP within 1 Mb of the transcription start site and end site of each gene g:
| (Equation 4) |
| (Equation 5) |
These linear regression models are fitted separately in each population l. In the above equations, and are the vectors of genotypes (0/1/2) for the jth variant, , in the AFR and EUR reference populations, respectively. Similarly, we have error terms and . and are the standardized marginal eQTL effect sizes for this variant. Derived from these fitted models, for each gene, we have summary statistics in the form of (the marginal least-squares effect estimates) and corresponding p values. These summarize the marginal association of variant j with the expression of the gene of interest in AFR and EUR populations.
Using marginal estimated eQTL effect-size vectors and corresponding marginal p value vectors from the single-variant eQTL data, we use OTTERS to train PRS models to impute gene expression separately in each population. While the OTTERS pipeline includes multiple frequentist and Bayesian PRS methods as options, here we focus on two as illustrative examples. We briefly summarize the methodology of these two PRS models below.
Pruning and thresholding
This approach includes two steps: pruning, or clumping, of variants to exclude correlated SNPs, and thresholding to keep only those SNPs significantly associated with gene expression.45 First, variants are filtered to include only those with p values below a pre-specified threshold . Next, among these remaining variants, a set of pairwise-independent variants are selected such that LD among such variants is less than some pre-specified threshold , preferentially keeping those with the smallest p value. These pruning and thresholding steps are performed in the OTTERS pipeline using PLINK 1.9.46 For our analysis, we used thresholds and . We chose a liberal LD threshold, as LD patterns differ considerably between populations, and OTTERS previously indicated that LD pruning did not significantly impact the performance of their method. We used the marginal standardized eQTL effect sizes from SNPs meeting these criteria to predict expression.
Lassosum
This method represents a summary-statistics-based version of the least absolute shrinkage and selection operation (lasso) pipeline; a penalized variable selection approach for dimension reduction with a large number of predictors (variants).47 Full details on lassosum are provided elsewhere.48 In contrast to pruning and thresholding (P + T) methods, lassosum requires LD blocks from an external reference panel. LD blocks were pre-calculated using lddetect49 and data from 1000 Genomes (AFR and EUR populations, respectively).50 Similar to the above, we performed LD clumping with prior to training PRS via lassosum.
Using the two pruning and threshold approaches (P + T0.05, P + T0.001) and the lassosum approach, we have three sets of estimated eQTL effect-size vectors for each training population for a given gene g: , where {P + T0.05, P + T0.001, lassosum}.
Stage II: Imputing expression in admixed individuals
LA-aware approaches: aspPS and casPS
As the true causal eQTLs in each ancestry are unknown, we propose to impute ancestry-specific components of GReX in the admixed target set using the trained PRS models of gene expression in the two reference eQTL datasets. Let us assume that we have phased genotype information and have performed ancestry deconvolution of our admixed target dataset haplotypes using existing methods like FLARE, RFMix, ELAI, or MOSAIC, among others.51,52,53,54 For each PRS model considered {P + T0.05, P + T0.001, lassosum}, we use the estimated eQTL effect-size vectors from stage I, and , to impute ancestry-specific partial gene expression PRS (aspPS) as outlined in Figure 1 in the manner of Marnetto et al.40 We can then add together the AFR- and EUR-specific partial components of GReX to create a combined ancestry-specific PRS (casPS).
Assume there are J estimated cis variants in the region of gene g and admixed target subjects. Motivated by our underlying expression model for admixed subjects described in “modeling expression in admixed individuals,” let be a matrix where the (i,j) element equals 1 if individual i has an AFR allele at jth SNP on his/her maternal (paternal) haplotype and 0 otherwise. We further define as the ancestry-standardized maternal (paternal) haplotype (0/1) matrix with . Here, is the minor allele count of the ith individual at the jth variant on the maternal haplotype, and f is the MAF of that variant in EUR or AFR, depending on the local ancestry of that SNP. Let be the estimated eQTL effect-size vector from a given PRS method using reference AFR (EUR) eQTL data. In CADET, we impute the AFR component of GReX and the EUR component of GReX as follows:
| (Equation 6) |
| (Equation 7) |
In the above equations, the symbol indicates element-wise multiplication of matrices. For each PRS model , we then impute the total GReX in our admixed target samples as the sum of these two ancestry-specific components:
| (Equation 8) |
LA-unaware approaches: PS
We can also compare the performance of the ancestry-aware approaches to imputing GReX () to LA-unaware approaches that simply apply the standard PRS methodologies to admixed individuals. For this, let be the standardized (columns centered and scaled to unit variance, not ancestry-specific standardization) matrix of minor allele counts (0/1/2). We can impute GReX in our admixed target data as
| (Equation 9) |
Here, is the estimated eQTL effect-size vector from a given PRS method for a given set of eQTL summary statistics z. This standard approach could use the same reference population eQTL summary data as those used in the LA-aware approaches above (z = AFR, EUR), or it could be trained using summary eQTL data from a sample of admixed individuals that is independent of our testing dataset (z = ADMIX).
Stage II: Gene-trait association test
We assume that our phenotype vector , for whom we want to estimate the association with GReX, has already been adjusted for the effects of important non-genetic confounders. Once we have our all of our vectors of imputed gene expression in our admixed testing dataset for {P + T0.05, P + T0.001, lassosum} and z {AFR, EUR, ADMIX}, we perform simple linear or logistic regression of y on each of the aspPS, casPS, and PS imputed GReX vectors individually. For each and z, we thus obtain an ordinary least-squares p value. Using the approach employed by OTTERS,16 CADET then combines these p values in multiple ways (as outlined in Figure 2) using the ACAT.44 First, we can aggregate across (our PRS models) to generate one p value for each of our LA-aware approaches , as well as our LA-unaware approaches (PS). We refer to this as level 1 aggregation, and this is the typical p-value aggregation performed in OTTERS. For example, to obtain the level 1 p value for casPS, let be the p value from the simple regression of y on for {P + T0.05, P + T0.001, lassosum}. We then calculate the level 1 p value for casPS ( as
| (Equation 10) |
| (Equation 11) |
Here, T is the ACAT test statistic, which is assumed to follow a Cauchy distribution. In Equations 10 and 11, we assume the weights (1/3) to be equal across all PRS methods. We can then repeat this level 1 aggregation process to further calculate .
As shown in Figure 2, CADET then performs a second round of p-value aggregation (which we refer to as level 2 aggregation) that combines all LA-aware p values from level 1 analyses and can further incorporate LA-unaware p values , if desired. We combine such p values using the same ACAT framework outlined for level 1 analysis as shown in Equations 10 and 11.
Simulations
To evaluate the accuracy of our proposed ancestry-aware gene expression imputation approach, we performed extensive simulations. First, we chose our simulated admixed testing dataset to be two-way African and European recently admixed individuals. We used the 1000 Genomes (1KG) phase 3 biallelic data in GRCh38 as source genomes to simulate multiple sets of admixed genomes.50 We specifically chose CEU (Utah Residents with Northern and Western European ancestry) as the European source population and YRI (Yoruba in Ibadan, Nigeria) as the African source population. To ensure our source populations were large enough for sampling haplotypes to generate our admixed testing set, we first expanded the YRI and CEU populations to size 10,000 each using admix-kit.55 This tool generates additional sets of 1KG populations using the HAPGEN2 framework.56
We selected one gene, ABCA7 (MIM: 605414), on chromosome 19 (GRCh38 1040101:1065571) as our target gene for simulations, as this gene has been used previously for the simulation stage of TWASs.13 We assumed a window of 1 Mb upstream and downstream as the simulation region (66,358 variants). We generated different training datasets in which we first simulate gene expression and then subsequently calculate eQTL summary data. We generated AFR and EUR training datasets by randomly selecting 500 of our expanded sample of YRI and CEU subjects, respectively. We further generated several admixed training datasets of size 500 using the expanded sets of 10,000 YRI and CEU source haplotypes and the tool haptools.57 We simulated three admixed training datasets according to different admixture generation parameters. First, we assumed an initial realistic African American demographic model of one pulse of admixture ten generations ago with 80% contribution from YRI and 20% from CEU (ADMIX 80 10g). We also considered two additional admixed training sets; the first assumed 5 admixture generations plus an initial AFR population frequency of 80% (ADMIX 80 5g), while the other assumed 5 admixture generations plus an initial AFR population frequency of 50% (ADMIX 50 5g). We considered these additional training sets, as previous work in this field has shown that power of TWAS association tests can be negatively impacted not only by differing training and target populations (e.g., EUR vs. admixed), but it can also vary according to admixture generation parameters in the admixed target genomes, such as the initial AFR population frequency.18
To generate our admixed testing datasets, we again assumed the realistic African American demographic model of ten admixture generations with 80% contribution from YRI (AFR) and 20% from CEU (EUR). These admixture generation settings for our testing dataset match those used to generate the ADMIX 80 10g training dataset. We simulated a total of 10,000 individuals to serve as our pool of admixed testing samples. These simulation settings are similar to those used in previous methodological work in admixed individuals.29
Using the simulated haplotypes and genotypes from the non-admixed reference AFR/EUR populations, we next simulated gene expression. In our training datasets (admixed and reference AFR/EUR of size 500), we first simulated gene expression according to the models described in “modeling expression in admixed individuals” and “stage I: reference expression model training via OTTERS.” We varied gene-expression heritability in the AFR population and in the EUR population () among {0.1,0.2}, excluding the limited utility scenario where both are 0.1. We also varied the number of eQTLs (V) in both populations among {2,10,100}, the proportion of eQTLs that overlap (OP) between AFR and EUR populations (OP = S/V) among {0.5, 1}, and the correlation of effect sizes among shared AFR and EUR eQTLs () among {0.5, 1}. We also ensured that all SNPs selected as eQTLs in each population have MAF > 0.05 in the corresponding training datasets. In total, we have 36 combinations of expression simulation parameters.
Since all settings consider overlap in eQTLs across ancestries, we took care when simulating correlated eQTL effect-size vectors in each ancestry. Using the notation of Equations 2 and 3, we first randomly drew a temporary ) and calculated the scale factor . To simulate eQTL effect sizes in EUR, note that the first S are correlated with the first S elements of the effect-size vector in AFR. Thus, we sampled a temporary as
| (Equation 12) |
Let the scale factor for EUR be . Thus, for each simulation, the eQTL effect-size vectors achieving the desired and OP are simply and . We then used these ancestry-specific eQTL effect-size vectors to simulate gene expression in our reference AFR/EUR training sets and, also, our admixed training/testing sets according to Equation 1.
Next, for each simulated training dataset (reference AFR/EUR and our independent admixed training samples), we calculated single-variant eQTL summary statistics and used the OTTERS pipeline to train PRS models. For each dataset, we retained summary data for only those variants with MAF > 0.05 in the corresponding sample. Using these eQTL summary statistics, we then imputed gene expression in our admixed testing set in the manner described in “Stage II: imputing expression in admixed individuals.” We compared the imputation R2 (squared correlation between imputed and true expression in our admixed testing set) between our proposed LA-aware approaches {P + T0.05, P + T0.001, lassosum}, and LA-unaware approaches for the AFR, EUR, ADMIX 80 10g, ADMIX 80 5g, and ADMIX 50 5g training datasets. For each round of GReX imputation using LA-aware approaches, we also performed another round of imputation where we assumed that 10% of the cis variants in the region on both haplotypes of each individual in the testing set had incorrect local ancestry information.
Next, we performed power and type I error simulations for the stage II test of GReX-trait association. We simulated the trait according to following equation, where we assume the trait is a function of the total gene expression and is not a function of the ancestry-specific components of gene expression in Equation 1.
| (Equation 13) |
For both our power and type I error simulations, we calculated the level 1 ACAT p values and . We then also calculated CADET’s level 2 ACAT p values that aggregated over various combinations of these level 1 ACAT p values. For our power simulations, we considered such that = 0.025. For our null simulations to evaluate type I error rate, we assume this value was 0. For each of the 1,000 simulations performed for each of the 36 combinations of expression simulation parameters, we performed ten trait simulations per . Thus, for each expression simulation parameter combination, we performed 10,000 power and 10,000 type I error simulations.
Additional sensitivity simulations
In addition to our main simulation design listed above, we also evaluated CADET under additional sensitivity simulations that varied our original design. We first modified our admixture generation model. Our selection of admixture generation parameters for our test genomes (one pulse of admixture ten generations ago with 80% contribution from YRI and 20% from CEU) was guided by recent methodological research on GWASs in African American populations.29 However, this approach may not reflect realistic admixture patterns observed in real African American genomes, which result from both historical admixture events and demographic trends. To address this, we conducted an extensive round of additional simulations, expanding the AFR donor genomes to include not only YRI but also Gambian in Western Division - Mandinka (GWD) and Mende in Sierra Leone (MSL) populations. Similarly, we broadened the EUR donor genomes to include both CEU and British from England and Scotland (GBR) populations. In these simulations, we further refined the admixture generation parameters used in Haptools. The updated model is provided in Table S1. Instead of a single pulse event of admixture between EUR and AFR, this model simulates a gradual shift in ancestry contributions while maintaining a predominant AFR genomic contribution. Using this model, we generated genetic, expression, and phenotype data for our test samples and assessed CADET performance assuming summary-level EUR and AFR training datasets as before. Additionally, we also considered the situation where the AFR training data contained admixed individuals themselves (rather than distinct non-admixed populations), with such admixture mirroring that in the test sample. In this situation, we generated the AFR admixed training data in a manner similar to that of the testing data above.
We also evaluated the sensitivity of our findings to the selected gene (ABCA7) of interest. To do so, we randomly selected nine genes according to gene length (three from first quartile of gene length [30,538 bp]). For each gene, we generated genetic, expression, and phenotype data for training and testing samples as described above.
CADET analysis of admixed subjects from UK Biobank
UK Biobank data
To assess the utility of CADET for TWASs in practice, we obtained individual-level genotype and phenotype data from admixed individuals in the UKB, which is a large-scale biomedical database housing data collected from approximately 500,000 individuals across the United Kingdom. To best mimic the settings of our simulation study, we elected to focus our analysis on two-way admixed individuals of African and European ancestry in the UKB. To identify these individuals comprising our target dataset, we first performed preliminary subject filtering. Specifically, we excluded subjects who had subsequently withdrawn from the study, those who were marked as “redacted,” and those who were marked as outliers based on pre-calculated metrics of heterozygosity and missing rates. We also excluded subjects with putative sex aneuploidy, those with high pre-calculated estimates of relatedness, and those whose genetically inferred gender differed from their submitted gender. We then removed individuals falling in the “White British subset.” These individuals were previously identified by a combination of both self-reported ancestry and genetic principal components (PCs). We also excluded individuals who have missing self-reported ethnicity or whose self-reported ethnicity fell among the following: White, Irish, British, and any other white background. Following this, 27,491 subjects remained.
We then performed principal component analysis (PCA), projecting these filtered UKB individuals onto genetic PCs anchored in 1000 Genomes (1KG) data.50 For this, we first subset 1KG genotype data to include unrelated individuals from the following populations: African (ACB [African Caribbeans in Barbados] and ASW [Americans of African Ancestry in SW USA] excluded) (503), Admixed American (347), East Asian (504), European (503), and South Asian (489). We restricted variants to non-ambiguous SNPs, those found in the UKB GWAS data, those with MAF > 0.05 in each population, and those with Hardy-Weinberg equilibrium (HWE) p > 1E−6. We then pruned remaining variants (window size 1,000 bp, step size 50 variants, and R2 threshold 0.1). Using the loadings for the top ten PCs trained in 1KG samples, we projected the UKB self-reported non-White individuals (27,491) into this space (Figure S1). Following the approach of Atkinson et al.,29 using the 1KG data, we then trained a random forest classifier to predict continental ancestry (1KG population) from the top ten PCs. We then applied this random forest model to our UKB sample. We excluded any individuals with <50% estimated probability of African ancestry. Using the top three PCs in the 9,187 meeting these criteria, we constructed a 95% ellipsoid along the African-European cline (Figure S2). We kept the 8,752 UKB individuals lying within the ellipsoid. Finally, we excluded two additional subjects with self-reported Asian/Asian British and White and Asian ethnicity. The remaining 8,750 individuals made up our final two-way AFR and EUR admixed target dataset.
Genotype data on these subjects were generated using either the UK BiLEVE or UKB Axiom arrays. Prior to imputation, we followed the pre-imputation quality control pipeline provided at https://www.well.ox.ac.uk/∼wrayner/tools/\#Checking, using variant data from TOPMed Freeze3a on GRCh37/hg19. We performed imputation, liftover to GRCh38, and phasing using the TOPMed Imputation Server.58,59,60 Next, we prepared our AFR and EUR reference data for local ancestry deconvolution of our UKB genotypes. First, we imputed missing genotypes in the 1KG phase 3 biallelic phased GRCh38 data of AFR and EUR subjects using BEAGLE, again excluding ASW and ACB.61 Using this as our reference population genotype data, we performed local ancestry inference in our admixed UKB target set using FLARE.51 We assumed ten generations since admixture (10 admg).
eQTL reference summary data
For our EUR reference eQTL dataset, we downloaded cis-eQTL summary statistics in whole blood from the Genotype-Tissue Expression (GTEx) V8 dataset (dbGaP phs000424.v8.p2), where cis-eQTL analysis was performed in 570 European-American subjects. To briefly summarize the analysis performed, the authors adjusted RNA sequencing for the effects of top five PCs, top 60 PEER factors, sequencing platform (Illumina HiSeq 2000 or HiSeq X), sequencing protocol (PCR-based or PCR-free), and sex. For the eQTL analysis, they then restricted genes to those with >0.1 transcripts per million and ≥6 reads in at least 20% of the data samples, and they normalized expression vectors using TMM62 and inverse normal transformation. SNPs from whole-genome sequencing (WGS) data with MAF ≥ 1% were retained. Authors performed single-variant cis-eQTL analysis using FastQTL63 and a 1-Mb window from the transcription start site of each gene.
For our African (AFR) reference eQTL dataset, we used publicly available whole-blood cis-eQTL summary statistics from a subset of high-African ancestry admixed individuals.17 In this study, the authors mapped ancestry-specific gene expression signatures in 2,733 individuals (African American, Puerto Rican, and Mexican American) from the Genes-Environments and Admixture in Latino Asthmatics (GALA II) study and the Study of African Americans, Asthma, Genes, and Environments (SAGE). They obtained paired RNA-sequencing data and WGS data. RNA and WGS data processing are described elsewhere.17 The authors used CEU and YRI HapMap reference genotypes, as well as Indigenous American ancestry reference genotypes, to estimate global measures of ancestry with ADMIXTURE.64 They defined a high global African ancestry subset as those with >50% estimated global African ancestry (721). They then performed ancestry-specific eQTL analysis in these subjects using a 1-Mb cis window from the transcription start site of each gene and FastQTL.63 Analyses adjusted for expression for age, sex, asthma status, top five PCs, and 60 PEER factors.
TWASs of blood biomarkers in UKB
We used CADET to perform TWASs of 29 widely collected blood biomarkers in our UKB testing dataset. For each blood biomarker, we first log-normalized raw trait measurements and then obtained covariate-adjusted phenotypes by taking the residuals of linear regression models of each log trait on the top 20 PCs, sex, age at recruitment, and smoking status (prefer not to answer, never, previous, current). Using our AFR and EUR reference eQTL summary data described above, we first removed ambiguous SNPs (A/T, T/A, G/C, and C/G) and only kept eQTL data from SNPs with MAF > 0.05 in each respective sample. Next, we trained AFR and EUR ancestry-specific PRS models of gene expression using OTTERS and P + T0.001, P + T0.05, and lassosum models. We then imputed gene expression in UKB samples using LA-aware approaches (casPS and aspPS) and standard PS approaches (PS AFR and PS EUR). We concluded our analyses by performing simple linear regression analysis for the association of each imputed gene expression vector with each of the 29 adjusted blood biomarker traits. We considered multiple level 2 p-value aggregation approaches via ACAT.
Results
Expression imputation accuracy
We first compared GReX imputation accuracy of proposed LA-aware approaches to LA-unaware approaches assuming = 10,000 simulated admixed individuals (ten admixture generations, 80% initial contribution of AFR [1KG YRI] haplotypes, 20% contribution of EUR [1KG CEU] haplotypes). We computed CADET’s LA-aware measures of GReX using simulated eQTL summary data from a reference AFR and a reference EUR sample (each = 500) and computed the LA-unaware GReX using each of these reference samples (AFR PS and EUR PS). We also computed additional LA-unaware GReX assuming eQTL summary datasets from independent simulated admixed samples of varying admixture parameters. We considered admixed reference samples with the same configuration as our test GWAS (10 admg, 80% initial AFR contribution) as well as other admixed reference samples that differed from the test GWAS either in admg (5 vs. 10), AFR contribution (50% vs. 80%), or both to examine the impact of mismatch between admixed training and test datasets on imputation accuracy. Across our training datasets, we provide histograms showing a representative distribution of the number of SNPs selected by each PRS method in a subsample of 1,000 simulations in Figure S3. As the p-value threshold becomes more stringent, P + T select fewer SNPs, whereas lassosum selects more SNPs, since its penalized regression approach shrinks effect sizes rather than eliminating SNPs outright.
Assuming an admixed testing sample size of 10,000, we present the squared correlation (R2) between imputed GReX and true gene expression assuming gene expression heritability of 0.2 in both Africans and Europeans () under the existence of two true eQTLs (Figure 3) or ten true eQTLs (Figure S4). We provide imputation R2 results for other expression heritability and eQTL architecture settings in Figures S4–S6. From these figures, we see a few general trends. Across all LA-aware and LA-unaware imputation approaches, as well as across all PRS methods (P + T0.001, P + T0.05, lassosum), we see higher R2 values for the sparse eQTL scenario (2) compared to the scenarios where the number of causal eQTLs is 10 or 100. There may exist other PRS approaches (PRScs, for example65) that may perform better for the scenario in which we expect larger number of true eQTLs.
Figures 3 and S4 also demonstrate that the accuracy of our proposed casPS approach tends to be slightly higher than the optimal standard PS approach using admixed eQTL summary data from a perfect ancestry-matched admixture cohort (ADMIX 80 10g PS) when the eQTLs are not identical between AFR and EUR or when the correlation of shared eQTL effect sizes is less than 1. In other words, these simulations suggest that even if we had access to eQTL data from a cohort exactly matched for ancestry to our admixed testing cohort, we would still achieve as high or higher imputation accuracy using casPS with reference population eQTL datasets. In the scenario where eQTL architecture for the given gene is identical between AFR and EUR populations (OP = 1, = 1), the imputation R2 appears quite similar between the casPS LA-aware approach and ADMIX 80 10g PS. Furthermore, across all settings, we observe that R2 decreases with increasing differences in admixture generation parameters between training and testing datasets. As the number of admixture generations decreases (from ten to five) and the proportion of initial AFR donors decreases (from 80% to 50%) within the training datasets, we generally see less successful GReX imputation in the test dataset generated assuming ten admixture generations and initial AFR population frequency of 80%.
These initial results assume correct inference of local ancestry in the admixed test samples. To assess the impact of LA misclassification on imputation accuracy, we conducted additional simulations where 10% of the eligible cis-SNPs of the test gene with MAF > 0.05 in the corresponding training datasets (Ref AFR or Ref EUR) had incorrect LA tags (AFR or EUR). We further assumed misclassification of these SNPs on both maternal and paternal haplotypes. We note that this assumption is likely on the high end for misclassification rates, as popular local ancestry imputation algorithms (e.g., FLARE,51 MOSAIC,54 and RFMix52) have high imputation accuracy for reasonably sized training panels used to infer LA.51,66 Nevertheless, as shown in Figures S7–S9, we do not see a marked drop in the quantiles of imputation R2 of our proposed LA-aware GReX imputation approaches (aspPS AFR, aspPS EUR, and casPS) across PRS models.
Type I error rate
Next, we assessed the type I error rate of the level 1 and level 2 TWAS association tests. For these simulations, we assumed a null association between true gene expression and each simulated phenotype, i.e., a phenotypic heritability . As the type I error rates did not differ dramatically by the gene-expression simulation parameters (, OP, number of eQTLs), we present quantile-quantile (QQ) plots for p values across each of the 36 simulation settings, corresponding to a total of 36,000 null simulations per plot.
As we see, the p values resulting from level 1 aggregation across PRS models (P + T0.001, P + T0.05, lassosum) for the LA-aware approaches (aspPS AFR, aspPS EUR, casPS) show the expected distribution under the null when we assume no LA misclassification (Figure S10) and 10% LA misclassification (Figure S11). Similarly, the level 1 p-value aggregation for the standard imputation approaches using reference eQTL data (PS AFR, PS EUR) also appear to maintain appropriate rates of type I error. To assess whether we see any inflation of our gene-trait association p values at the second level of p-value aggregation, i.e., aggregating level 1 p values for our LA-aware approaches and/or standard PS approaches, we constructed a second set of QQ plots in Figure S12. Again, we see the expected distribution under the null even when we combine aspPS and casPS p values with those from the reference-derived PSs (AFR PS, EUR PS) and, further, with the p values from the independent admixed-derived PSs (Figure S12, middle and right panels, respectively). This applies to both assumptions of 0% (top row) or 10% LA misspecification (bottom row).
Power
Simulation results comparing the performance of our proposed LA-aware approaches to the competing standard GReX imputation approaches under the assumption of a true gene-trait association are summarized in Figure 4 (two eQTLs) and Figure 5 (ten eQTLs). These figures reflect normally distributed traits, a testing set sample size of 10,000, significance level 2.5E−6, and no LA misspecification. When comparing the performance of the level 1 approaches (aggregation across PRS), our proposed LA-aware casPS has higher or equal power compared to any of the LA-unaware standard PS approaches, including that trained using a perfectly matched admixed sample, when gene expression genetic architecture is different between populations. Of note, the power of level 1 casPS is also higher than that of an LA-unaware method with a perfectly matched admixed sample (ADMIX 80 10g) when the gene expression genetic architecture is exactly the same between AFR and EUR groups (OP = 1, = 1). Furthermore, across all simulation settings, results show that level 2 aggregation of LA-aware and LA-unaware methods using CADET achieved greater power than any individual level 1 LA-unaware method.
In Figures S13 (two eQTLs) and S14 (ten eQTLs), we illustrate the power of the LA-aware approaches assuming a high rate of LA misclassification (10%) for cis-SNPs in the gene region. With the exception of the scenario in which we have a higher level of gene expression heritability in Europeans ( = 0.2) and lower level of expression heritability in Africans ( = 0.1) and assume ten eQTLs, we still observe the level 2 ACAT LA-aware approach of CADET achieving highest power. For the anomalous setting mentioned, we note that gene expression imputation accuracy here is quite low across the board (Figure S6), and thus power is subsequently low for all LA-aware and LA-unaware approaches. This is similar to the scenario where we assume 100 eQTLs for the gene under study. We have comparatively less accurate GReX imputation, and therefore it is unsurprising that we see low power in the downstream association tests across all methods (Figures S15 and S16).
Sensitivity simulations
In our sensitivity simulation analyses, we first introduced a more complex admixture generation framework for our testing genomes (Table S1). We then evaluated how the performance of our proposed LA-aware approaches to gene expression imputation varied according to the length of our selected gene of interest. Across all settings, CADET maintains desired type I error (Figures S17–S22). The imputation accuracy of our proposed LA-informed casPS approach (Figures S23–S25) and power of CADET (Figures S26–S28) are consistently as high or higher than the optimal standard PS approach using admixed eQTL summary data from a perfect ancestry-matched admixture cohort using these new admixture generation parameters across each gene-length category. While figures are only presented here for the scenario in which two eQTLs exist for each ancestry, the trends shown in the main analyses comparing two and ten eQTLs hold throughout these additional simulations.
We also explored performance sensitivity when our AFR eQTL summary statistics were derived from an admixed training sample rather than an ancestrally homogeneous AFR group, as assumed previously. In Figures S17–S28, we refer to the casPS quantity in this scenario as “casPS w/AFR ADMIX” and the corresponding AFR-specific partial PS as “aspPS AFR w/AFR ADMIX.” While this version of casPS exhibits lower imputation accuracy compared to the original casPS, the second stage of CADET, in which we aggregate p values from both standard PS and casPS, helps mitigate this limitation. This omnibus approach enables CADET to retain power levels comparable to the optimal version based on non-admixed eQTL summary statistics (Figures S26–S28).
CADET analysis of UK Biobank
We applied CADET to detect genes associated with blood biomarkers by way of genetically regulated transcriptional activity in a subset of 8,750 two-way African/European admixed subjects in the UKB. We first performed LA deconvolution using FLARE and reference genotypes from AFR and EUR cohorts from 1000 Genomes. Through this, we estimated an overall global (genome-wide) proportion of AFR genotypes of 86.6% and an estimated proportion of EUR genotypes of 13.4%. Using our two sets of European-derived and high-African ancestry-derived eQTL summary data in whole blood, we trained our AFR and EUR GReX imputation models for 15,093 genes. Specifically, we used lassosum and pruning and thresholding models, keeping = 0.99 as in the simulations and filtering by p-value thresholds of = 0.05 and 0.001. For each PRS model, we then imputed GReX in our admixed UKB target set in the standard OTTERS approach separately using our AFR-trained models (PS AFR) and EUR-trained models (PS EUR). We also imputed GReX in our proposed LA-aware approach using both sets of PRS models and our inferred LA information. We calculated two ancestry-specific components of gene expression (aspPS AFR and aspPS EUR) and one vector representing the sum of these two components (casPS). In Table S2, we summarize the pairwise correlations between the vectors of imputed GReX by each approach and PRS model. PRS models derived from the same population (e.g., PS AFR vs. aspPS AFR) exhibit the highest correlations in imputed GReX, while cross-population comparisons (e.g., aspPS AFR vs. aspPS EUR or PS AFR vs. PS EUR) show lower correlations, particularly for lassosum, which has consistently lower mean correlations across approaches.
We tested for the association of each of the vectors of imputed GReX with 29 blood biomarkers: four bone and joint traits (alkaline phosphatase, calcium, rheumatoid factor, and vitamin D), eight cardiovascular traits (apolipoprotein A and B, C-reactive protein, cholesterol, high-density lipoprotein cholesterol, low-density lipoprotein [LDL] cholesterol, lipoprotein A, and triglycerides), two diabetes-related traits (glucose and HbA1c), three hormone traits (insulin-like growth factor 1 [IGF-1], sex hormone binding globulin [SHBG], and testosterone), six liver-related traits (alanine aminotransferase, albumin, aspartate aminotransferase, direct bilirubin, gamma-glutamyltransferase, and total bilirubin), and six traits related to renal function (creatinine, cystatin C, phosphate, total protein, urate, and urea). For each approach (PS AFR, PS EUR, aspPS AFR, aspPS EUR, and casPS), we calculated a level 1 p value aggregated across PRS models. We also aggregated across a second dimension beyond PRS model to calculate two level 2 p values (casPS + aspPS and casPS + aspPS + PS [CADET]).
We assessed significance at a Bonferroni-adjusted level of 0.05/(29 × 15,093) = 1.14E−07, adjusting for the total number of phenotypes and number of genes considered. Across all tests and imputation approaches, we identified 384 significant gene-trait associations. While 384 significant associations were found, most gene-trait associations were unsurprisingly picked up by more than one imputation approach, and thus we ultimately identified 92 unique gene-trait pairings. We identified significant associations for 15/29 traits (52%), with the most associations observed for SHBG (19), total bilirubin (11), apolipoprotein B (9), and lipoprotein A (9). To evaluate the utility of our LA-aware approaches, we compare and contrast the total number of unique gene-trait associations identified in Figure S29.
As an additional comparison, this figure includes results from standard LA-unaware PS imputation using eQTL summary statistics from all individuals of African ancestry, regardless of their estimated global African ancestry proportion (PS-AA). This broader set (N = 757) includes individuals with higher levels of admixture compared to the high-proportion African ancestry cohort used in our main analysis. Despite this difference, we observed strong consistency in the identified genes between PS-AA and CADET. Specifically, 35 out of 51 (69%) gene associations detected by PS-AA were also identified by PS in the high-African ancestry cohort (PS-AFR), and 37 out of 51 (73%) were captured by the LA-aware approach (casPS + aspPS). Notably, 21 significant associations were uniquely identified by casPS + aspPS but not by the LA-unaware PS-AA approach, highlighting the added value of incorporating LA information.
Leveraging our LA-derived ancestry-specific p values (aspPS AFR/EUR) and the p values from their combined component (casPS), this level 2 aggregation successfully identified 83/92 associations (90.2%). Eighteen gene-trait associations were identified by this casPS + aspPS approach that were not identified by standard GReX imputation using reference eQTL summary data (PS AFR or PS EUR). Some of the genes among these 18 gene-trait pairs are in close proximity to each other and likely do not reflect independent regions of association. Therefore, in Table 1, we present the most significant genes observed within a 1-Mb window for each phenotype. We provide the full lists of all gene-trait associations identified by each individual approach in Tables S3, S4, S5, S6, and S7).
Table 1.
| Phenotype | Group | Gene | Gene type | Chromosome | Position (hg38) | p value |
|---|---|---|---|---|---|---|
| Direct bilirubin | liver | ATG16L1 | protein coding | 2 | 233210051 | 3.11E−09 |
| Lipoprotein A | cardiovascular | ZDHHC14 | protein coding | 6 | 157381133 | 6.09E−08 |
| Lipoprotein A | cardiovascular | HNRNPH1P1 | processed pseudogene | 6 | 159712801 | 1.11E−10 |
| Apolipoprotein A | cardiovascular | CD36 | protein coding | 7 | 80369575 | 7.14E−08 |
| Apolipoprotein B | cardiovascular | IST1 | protein coding | 16 | 71845996 | 1.08E−07 |
| SHBG | hormone | WRAP53 | protein coding | 17 | 7686071 | 3.35E−17 |
| Cholesterol | cardiovascular | ZNF491 | protein coding | 19 | 11797667 | 7.96E−08 |
| LDL direct | cardiovascular | ZNF491 | protein coding | 19 | 11797667 | 6.26E−09 |
| Gamma-glutamyltransferase | liver | LRRC75B | protein coding | 22 | 24585620 | 5.39E−10 |
| Creatinine | renal | XRCC6 | protein coding | 22 | 41621119 | 1.40E−09 |
For genomic loci (1-Mb regions) containing multiple identified genes for the same phenotype, only the most significant gene is presented. The p values shown represent ACAT(casPS p, aspPS AFR p, aspPS EUR p).
Of the nine genes identified by the casPS + aspPS approach highlighted in Table 1, the vast majority have consistent evidence in the GWAS literature, harboring one or more significant GWAS variants (p < 5E−08) for the relevant traits. However, one association (HNRNPH [MIM: 601035] for lipoprotein A) is not located near known GWAS loci.67 While HNRNPH is a pseudogene, and the biological mechanisms behind this association are thus unclear, previous work has helped elucidate the role that pseudogenes can play in the development of cardiovascular disease.68
Next, we used the online database TWAS Atlas to examine prior documented TWAS associations for each of these nine genes.3 Some of the genes identified exclusively by the LA-aware approach have been implicated in previous TWASs of relevant traits, while others represent potentially new TWAS findings. For example, we did not find prior TWAS associations of ATG16L1 (MIM: 610767) with traits relevant to bilirubin levels. While some of the genes identified by our approach for lipoproteins have been implicated by previous TWASs of lipid biomarkers of cardiovascular disease (e.g., ZDHHC14 [MIM: 619295], CD36 [MIM: 173510], and IST1 [MIM: 616434]), HNRNPH1P1 does not have any reported association in the database. The gene associated with SHBG in Table 1 (WRAP53 [OMIM: 612661]) was identified in association with endometriosis in a recent TWAS.69 SHBG plays a role in the availability of sex hormones in the body, and increased levels of SHBG were observed among women with endometriosis compared to controls.70 While ZNF491 (MIM: not available) has been identified by previous TWASs of cholesterol, genes LRRC75B (MIM: not available) and XRCC6 (MIM: 152690) did not have any significant prior TWAS associations with biomarker-relevant phenotypes in the TWAS Atlas.
We note that the genes highlighted above were identified by level 2 p-value aggregation of casPS and the ancestry-specific partial GReX results (casPS + aspPS). We focus on these findings to underscore the additional insight that can be found by LA-aware approaches only. However, CADET optimally allows for further aggregation of results from the LA-unaware standard GReX imputation approaches (i.e., casPS + aspPS + PS). Using this method, we identify four additional significant gene-trait associations (Figure S29 and Table S4).
As a concluding sensitivity analysis, we repeated the applied TWAS analysis in UKB subjects for the 29 blood biomarker phenotypes using GReX imputation models trained with alternative EUR eQTL summary statistics from living participants of the Estonian Biobank (EB). eQTL summary statistics from blood in living donors may be preferred over those from cadavers (GTEx) due to postmortem RNA degradation and changes in gene expression that can affect eQTL mapping.71 We obtained single-variant eQTL summary statistics from Lepik et al.72 and downloaded whole-blood summary statistics derived from RNA-sequencing data on N = 471 individuals participating in the EB. Details on sample collection, RNA sequencing, and quality control procedures for these data are described elsewhere.72,73 For this analysis, we focused on 12,388 genes for which we successfully trained GReX imputation models using both AFR and EUR (Estonian) eQTL data. Restricting our comparison to genes overlapping between the GTEx and Lepik et al. datasets, CADET (casPS + aspPS + PS approach) originally identified 54 unique gene-trait pairs as significant in the GTEx-based analysis after Bonferroni correction. Of these, 35 (65%) were also identified as significant when using Estonian eQTL summary statistics, demonstrating substantial overlap. We include EB replication outcomes for our GTEx-identified associations in Table S8. Of note, of the genes highlighted in Table 1, we were able to evaluate six using EB data, and of these, four associations remained significant (ATG16L1, ZDHHC14, CD36, and WRAP53).
Discussion
In this project, we introduce CADET, a method for performing transcriptome-wide association analysis in admixed subjects. Genomes of admixed individuals are a mosaic of two or more distinct ancestral groups, and thus local ancestry tract information within each haplotype can be leveraged to improve power in genetic association analyses when causal variant effect sizes differ between populations. Our method contributes to the growing catalog of statistical methods dedicated to admixed individuals with the appealing feature that it does not require individual-level genotype and gene-expression training data, as is typically needed in stage 1 GReX model training in TWASs. Here, we build upon a recently published TWAS approach, OTTERS, that uses well-established PRS models designed for GWAS summary data and applies them to eQTL summary data to impute gene expression in a tissue of interest.16 In CADET, we use multiple sets of eQTL summary statistics, namely those derived from the distinct parent ancestral groups of our admixed testing sample. The framework is flexible enough to also incorporate eQTL summary data from admixed cohorts independent of the target set. The CADET software is available on GitHub (see web resources).
Through our simulations, we first demonstrate that p values of all variants of our proposed approach yield the expected distribution under the null assumption of no gene-trait association. These variants include two approaches of p-value aggregation: level 1, wherein we combine p values from the three PRS models, and level 2, in which we aggregate level 1 p values. Next, through our power simulations, we make several important observations. First, under all scenarios of ten or fewer eQTLs, both our level 1 casPS and level 2 (casPS + aspPS + PS) aggregation methods achieve superior power to any standard LA-unaware approach when the genetic architecture of gene expression differs between AFR and EUR populations. In fact, there is growing evidence for such ancestry-specific genetic architecture patterns, and estimated gene expression heritability has also been shown to differ significantly by local ancestry at the transcription start site of genes among admixed subjects.17 Second, even when genetic architecture patterns are identical between ancestries, our LA-aware level 1 casPS and level 2 (casPS + aspPS + PS) generally achieve greater power or comparable power if we employ standard GReX imputation using eQTL summary data from a perfectly matched (number of admixture generations, initial AFR/EUR contribution proportions) independent admixed sample. Third, we also see that our level 2 approach still performs competitively in these settings when we assume that a large number (10%) of local ancestries are misclassified in the gene region and is the preferred method for use in practice.
Finally, in our applied analysis, we use real-world eQTL summary data from a European sample and a high-African ancestry sample of African Americans to perform a TWAS of 29 blood biomarker traits. Here, our target set is two-way African and European admixed subjects from the UKB. We successfully identified 18 significant gene-trait associations using our level 2 p-value aggregation approach (casPS + aspPSs) that were not picked up using standard GReX PS imputation methods that ignore local ancestry (PS AFR and PS EUR). The majority of these genes are consistent with GWAS loci previously identified for the corresponding traits. However, one gene, HNRNPH on chromosome 6, represents a potentially novel locus for lipoprotein A and was not identified by standard GReX imputation.
We observe several limitations of our present work. First, the partial PSs generated through CADET do not have a direct biological interpretation. However, the utility of combining predicted gene expression stemming from district LA segments and eQTL data from different populations lies in its ability to capture ancestry-specific regulatory effects that may otherwise be obscured in traditional TWASs. Similar to how methods such as Tractor leverage LA to improve power in GWASs by accounting for ancestry-specific allele frequencies and effect sizes,29 our approach incorporates LA-informed predicted expression to deconvolute ancestry-specific regulatory contributions. By doing so, we circumvent any a priori assumptions regarding similarities or differences in the genetic regulation of gene expression across populations and local ancestries. We argue that this allows us to better predict gene expression in a variety of contexts. By incorporating admixture-dependent mechanisms, our approach provides a more nuanced understanding of the genetic architecture of complex traits.
We also note that while CADET demonstrates desirable power levels for modest trait heritability, the gene expression imputation accuracies are lower across all simulation settings than the simulated gene-expression heritability levels. We argue that our imputation models are trained using PRS approaches to leverage the widespread availability of eQTL summary data, and R2 has been shown to fall below true trait heritability for common PRS approaches,16,65 and we expect imputation accuracy to be even lower when ancestry-specific effects are involved. We argue that we may further improve imputation accuracy in CADET by including more PRS approaches to our method, and, importantly, those designed precisely for admixed samples (e.g., GAUDI42). We also note that our framework is flexible enough to incorporate individual-level training data (paired genotype and gene expression information), if available. However, previous work has shown that GReX models trained with individual-level reference data showed modest to no improvement in testing R2 over various PRS-based models trained with summary data.16 Additionally, in this project, we only consider two-way admixed individuals for both our simulated analyses and applied work in the UKB. We believe that we can easily extend this approach to allow for three-way or higher levels of admixture, and that the approach is computationally efficient enough to implement this in practice.
Another limitation of CADET includes the requirement of individual-level genotype data. While other TWAS methods may alternatively use summary statistics, this requirement may restrict the applicability of our method in scenarios where only summary-level GWAS data are available, particularly for large-scale meta-analyses or studies with strict data-access limitations. Future work could explore the feasibility of adapting our framework to settings where only summary statistics and summary-level LA information are available. An additional comment is that the method assumes substantial overlap in variants with non-missing eQTL summary statistics across parent ancestry groups to ensure accurate GReX imputation. When such overlap is limited, imputation accuracy can be impacted. However, this limitation can be addressed by applying existing GWAS summary statistic imputation tools74,75 to harmonize variant coverage prior to running the CADET pipeline.
Finally, we designed our method to utilize eQTL data from ancestrally homogeneous parent ancestral groups. However, when seeking to apply our method to admixed individuals of African ancestry, we note there is a marked paucity of eQTL studies performed in non-admixed African cohorts. The eQTL summary data from African American individuals with >50% African ancestry alleles currently represents the best surrogate dataset of reasonable sample size for our analysis, and this substitution has similarly been employed in another recent methodological paper.42 While using casPS with admixed reference eQTL data resulted in lower imputation accuracy than using non-admixed reference data in our simulations, the second stage of CADET where p values from both standard PS and casPS are aggregated helps counteract this limitation. We show that CADET thus maintains robust performance despite the constraints of available eQTL data.
Data and code availability
The GTEx V8 cis-eQTL summary statistics in Europeans used for the applied analysis of UKB subjects are publicly available and can be accessed at https://www.gtexportal.org/home/downloads/adult-gtex/qtl. The cis-eQTL summary statistics from high-African ancestry African Americans by Kachuri et al.17 also used in our applied UKB analysis can be accessed at https://zenodo.org/records/7735723. The eQTL summary statistics from the Estonian Biobank generated by Lepik et al.72 are available on the eQTL Catalog73 at https://www.ebi.ac.uk/eqtl/. All scripts and code used to generate figures shown in this paper are available from https://github.com/staylorhead/CADET.
Acknowledgments
This work was supported by the National Institutes of Health (AG071170, CA211574, and MH126449).
Declaration of interests
The authors declare no competing interests.
Published: June 13, 2025
Footnotes
Supplemental information can be found online at https://doi.org/10.1016/j.ajhg.2025.05.010.
Web resources
CADET source code, https://github.com/staylorhead/CADET
eQTL Estonian Biobank summary statistics, https://www.ebi.ac.uk/eqtl/
eQTL GTEx summary statistics for Europeans, https://www.gtexportal.org/home/downloads/adult-gtex/qtl
eQTL summary statistics for high-African ancestry African Americans, https://zenodo.org/records/7735723
OMIM, https://www.omim.org/
Imputation quality control pipeline, https://www.well.ox.ac.uk/∼wrayner/tools/\#Checking
Supplemental information
References
- 1.Giral H., Landmesser U., Kratzer A. Into the Wild: GWAS Exploration of Non-coding RNAs. Front. Cardiovasc. Med. 2018;5:181. doi: 10.3389/fcvm.2018.00181. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 2.Mancuso N., Shi H., Goddard P., Kichaev G., Gusev A., Pasaniuc B. Integrating Gene Expression with Summary Association Statistics to Identify Genes Associated with 30 Complex Traits. Am. J. Hum. Genet. 2017;100:473–487. doi: 10.1016/j.ajhg.2017.01.031. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 3.Lu M., Zhang Y., Yang F., Mai J., Gao Q., Xu X., Kang H., Hou L., Shang Y., Qain Q., et al. TWAS Atlas: a curated knowledgebase of transcriptome-wide association studies. Nucleic Acids Res. 2023;51:D1179–D1187. doi: 10.1093/nar/gkac821. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 4.Bhattacharya A., García-Closas M., Olshan A.F., Perou C.M., Troester M.A., Love M.I. A framework for transcriptome-wide association studies in breast cancer in diverse study populations. Genome Biol. 2020;21:42. doi: 10.1186/s13059-020-1942-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 5.Zhong J., Jermusyk A., Wu L., Hoskins J.W., Collins I., Mocci E., Zhang M., Song L., Chung C.C., Zhang T., et al. A Transcriptome-Wide Association Study Identifies Novel Candidate Susceptibility Genes for Pancreatic Cancer. J. Natl. Cancer Inst. 2020;112:1003–1012. doi: 10.1093/jnci/djz246. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 6.Parrish R.L., Gibson G.C., Epstein M.P., Yang J. TIGAR-V2: Efficient TWAS tool with nonparametric Bayesian eQTL weights of 49 tissue types from GTEx V8. HGG Adv. 2022;3 doi: 10.1016/j.xhgg.2021.100068. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 7.Tang S., Buchman A.S., De Jager P.L., Bennett D.A., Epstein M.P., Yang J. Novel Variance-Component TWAS method for studying complex human diseases with applications to Alzheimer’s dementia. PLoS Genet. 2021;17 doi: 10.1371/journal.pgen.1009482. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 8.Yao S., Zhang X., Zou S.-C., Zhu Y., Li B., Kuang W.-P., Guo Y., Li X.-S., Li L., Wang X.-Y. A transcriptome-wide association study identifies susceptibility genes for Parkinson’s disease. Npj Park. Dis. 2021;7:79. doi: 10.1038/s41531-021-00221-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 9.Dai Y., Pei G., Zhao Z., Jia P. A Convergent Study of Genetic Variants Associated With Crohn’s Disease: Evidence From GWAS, Gene Expression, Methylation, eQTL and TWAS. Front. Genet. 2019;10 doi: 10.3389/fgene.2019.00318. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 10.Gusev A., Ko A., Shi H., Bhatia G., Chung W., Penninx B.W.J.H., Jansen R., de Geus E.J.C., Boomsma D.I., Wright F.A., et al. Integrative approaches for large-scale transcriptome-wide association studies. Nat. Genet. 2016;48:245–252. doi: 10.1038/ng.3506. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 11.Rodriguez-Fontenla C., Carracedo A. UTMOST, a single and cross-tissue TWAS (Transcriptome Wide Association Study), reveals new ASD (Autism Spectrum Disorder) associated genes. Transl. Psychiatry. 2021;11:256. doi: 10.1038/s41398-021-01378-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 12.Gamazon E.R., Wheeler H.E., Shah K.P., Mozaffari S.V., Aquino-Michaels K., Carroll R.J., Eyler A.E., Denny J.C., Nicolae D.L., et al. GTEx Consortium A gene-based association method for mapping traits using reference transcriptome data. Nat. Genet. 2015;47:1091–1098. doi: 10.1038/ng.3367. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 13.Nagpal S., Meng X., Epstein M.P., Tsoi L.C., Patrick M., Gibson G., De Jager P.L., Bennett D.A., Wingo A.P., Wingo T.S., Yang J. TIGAR: An Improved Bayesian Tool for Transcriptomic Data Imputation Enhances Gene Mapping of Complex Traits. Am. J. Hum. Genet. 2019;105:258–266. doi: 10.1016/j.ajhg.2019.05.018. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 14.Mancuso N., Freund M.K., Johnson R., Shi H., Kichaev G., Gusev A., Pasaniuc B. Probabilistic fine-mapping of transcriptome-wide association studies. Nat. Genet. 2019;51:675–682. doi: 10.1038/s41588-019-0367-1. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 15.Yuan Z., Zhu H., Zeng P., Yang S., Sun S., Yang C., Liu J., Zhou X. Testing and controlling for horizontal pleiotropy with probabilistic Mendelian randomization in transcriptome-wide association studies. Nat. Commun. 2020;11:3861. doi: 10.1038/s41467-020-17668-6. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 16.Dai Q., Zhou G., Zhao H., Võsa U., Franke L., Battle A., Teumer A., Lehtimäki T., Raitakari O.T., Esko T., et al. OTTERS: a powerful TWAS framework leveraging summary-level reference data. Nat. Commun. 2023;14:1271. doi: 10.1038/s41467-023-36862-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 17.Kachuri L., Mak A.C.Y., Hu D., Eng C., Huntsman S., Elhawary J.R., Gupta N., Gabriel S., Xiao S., Keys K.L., et al. Gene expression in African Americans, Puerto Ricans and Mexican Americans reveals ancestry-specific patterns of genetic architecture. Nat. Genet. 2023;55:952–963. doi: 10.1038/s41588-023-01377-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 18.Keys K.L., Mak A.C.Y., White M.J., Eckalbar W.L., Dahl A.W., Mefford J., Mikhaylova A.V., Contreras M.G., Elhawary J.R., Eng C., et al. On the cross-population generalizability of gene expression prediction models. PLoS Genet. 2020;16 doi: 10.1371/journal.pgen.1008927. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 19.Martin A.R., Kanai M., Kamatani Y., Okada Y., Neale B.M., Daly M.J. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat. Genet. 2019;51:584–591. doi: 10.1038/s41588-019-0379-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 20.Manrai A.K., Funke B.H., Rehm H.L., Olesen M.S., Maron B.A., Szolovits P., Margulies D.M., Loscalzo J., Kohane I.S. Genetic Misdiagnoses and the Potential for Health Disparities. N. Engl. J. Med. 2016;375:655–665. doi: 10.1056/NEJMsa1507092. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 21.Reisberg S., Iljasenko T., Läll K., Fischer K., Vilo J. Comparing distributions of polygenic risk scores of type 2 diabetes and coronary heart disease within different populations. PLoS One. 2017;12 doi: 10.1371/journal.pone.0179238. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 22.De La Vega F.M., Bustamante C.D. Polygenic risk scores: a biased prediction? Genome Med. 2018;10:100. doi: 10.1186/s13073-018-0610-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 23.Kim M.S., Patel K.P., Teng A.K., Berens A.J., Lachance J. Genetic disease risks can be misestimated across global populations. Genome Biol. 2018;19:179. doi: 10.1186/s13059-018-1561-7. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 24.Visscher P.M., Wray N.R., Zhang Q., Sklar P., McCarthy M.I., Brown M.A., Yang J. 10 Years of GWAS Discovery: Biology, Function, and Translation. Am. J. Hum. Genet. 2017;101:5–22. doi: 10.1016/j.ajhg.2017.06.005. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 25.Scutari M., Mackay I., Balding D. Using Genetic Distance to Infer the Accuracy of Genomic Prediction. PLoS Genet. 2016;12 doi: 10.1371/journal.pgen.1006288. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 26.Skotte L., Jørsboe E., Korneliussen T.S., Moltke I., Albrechtsen A. Ancestry-specific association mapping in admixed populations. Genet. Epidemiol. 2019;43:506–521. doi: 10.1002/gepi.22200. [DOI] [PubMed] [Google Scholar]
- 27.Cutler D.J., Jodeiry K., Bass A.J., Epstein M.P. The Quantitative Genetics of Human Disease: 2 Polygenic Risk Scores. Hum. Popul. Genet. Genom. 2024;4:1–65. doi: 10.47248/hpgg2404030008. [DOI] [Google Scholar]
- 28.US Census Bureau . 2021. 2020 Census illuminates racial and ethnic composition of the country.www.census.gov/library/stories/2021/08/improved-race-ethnicity-measures-reveal-united-states-population-much-more-multiracial [Google Scholar]
- 29.Atkinson E.G., Maihofer A.X., Kanai M., Martin A.R., Karczewski K.J., Santoro M.L., Ulirsch J.C., Kamatani Y., Okada Y., Finucane H.K., et al. Tractor uses local ancestry to enable the inclusion of admixed individuals in GWAS and to boost power. Nat. Genet. 2021;53:195–204. doi: 10.1038/s41588-020-00766-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 30.Zhang J., Stram D.O. The Role of Local Ancestry Adjustment in Association Studies Using Admixed Populations. Genet. Epidemiol. 2014;38:502–515. doi: 10.1002/gepi.21835. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 31.Lachance J., Vernot B., Elbers C.C., Ferwerda B., Froment A., Bodo J.-M., Lema G., Fu W., Nyambo T.B., Rebbeck T.R., et al. Evolutionary History and Adaptation from High-Coverage Whole-Genome Sequences of Diverse African Hunter-Gatherers. Cell. 2012;150:457–469. doi: 10.1016/j.cell.2012.07.009. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 32.Tang H., Siegmund D.O., Johnson N.A., Romieu I., London S.J. Joint testing of genotype and ancestry association in admixed families. Genet. Epidemiol. 2010;34:783–791. doi: 10.1002/gepi.20520. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 33.Coram M.A., Duan Q., Hoffmann T.J., Thornton T., Knowles J.W., Johnson N.A., Ochs-Balcom H.M., Donlon T.A., Martin L.W., Eaton C.B., et al. Genome-wide Characterization of Shared and Distinct Genetic Components that Influence Blood Lipid Levels in Ethnically Diverse Human Populations. Am. J. Hum. Genet. 2013;92:904–916. doi: 10.1016/j.ajhg.2013.04.025. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 34.Aschard H., Gusev A., Brown R., Pasaniuc B. Leveraging local ancestry to detect gene-gene interactions in genome-wide data. BMC Genet. 2015;16:124. doi: 10.1186/s12863-015-0283-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 35.Zaitlen N., Paşaniuc B., Gur T., Ziv E., Halperin E. Leveraging genetic variability across populations for the identification of causal variants. Am. J. Hum. Genet. 2010;86:23–33. doi: 10.1016/j.ajhg.2009.11.016. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 36.Pasaniuc B., Zaitlen N., Lettre G., Chen G.K., Tandon A., Kao W.H.L., Ruczinski I., Fornage M., Siscovick D.S., Zhu X., et al. Enhanced Statistical Tests for GWAS in Admixed Populations: Assessment using African Americans from CARe and a Breast Cancer Consortium. PLoS Genet. 2011;7 doi: 10.1371/journal.pgen.1001371. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 37.Pasaniuc B., Sankararaman S., Torgerson D.G., Gignoux C., Zaitlen N., Eng C., Rodriguez-Cintron W., Chapela R., Ford J.G., Avila P.C., et al. Analysis of Latino populations from GALA and MEC studies reveals genomic loci with biased local ancestry estimation. Bioinformatics. 2013;29:1407–1415. doi: 10.1093/bioinformatics/btt166. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 38.Chimusa E.R., Zaitlen N., Daya M., Möller M., van Helden P.D., Mulder N.J., Price A.L., Hoal E.G. Genome-wide association study of ancestry-specific TB risk in the South African Coloured population. Hum. Mol. Genet. 2014;23:796–809. doi: 10.1093/hmg/ddt462. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 39.Smith E.N., Bloss C.S., Badner J.A., Barrett T., Belmonte P.L., Berrettini W., Byerley W., Coryell W., Craig D., Edenberg H.J., et al. Genome-wide association study of bipolar disorder in European American and African American individuals. Mol. Psychiatry. 2009;14:755–763. doi: 10.1038/mp.2009.43. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 40.Marnetto D., Pärna K., Läll K., Molinaro L., Montinaro F., Haller T., Metspalu M., Mägi R., Fischer K., Pagani L. Ancestry deconvolution and partial polygenic score can improve susceptibility predictions in recently admixed individuals. Nat. Commun. 2020;11:1628. doi: 10.1038/s41467-020-15464-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 41.Hoggart C.J., Choi S.W., García-González J., Souaiaia T., Preuss M., O’Reilly P.F. BridgePRS leverages shared genetic effects across ancestries to increase polygenic risk score portability. Nat. Genet. 2024;56:180–186. doi: 10.1038/s41588-023-01583-9. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 42.Sun Q., Rowland B.T., Chen J., Mikhaylova A.V., Avery C., Peters U., Lundin J., Matise T., Buyske S., Tao R., et al. Improving polygenic risk prediction in admixed populations by explicitly modeling ancestral-differential effects via GAUDI. Nat. Commun. 2024;15:1016. doi: 10.1038/s41467-024-45135-z. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 43.Shang L., Smith J.A., Zhao W., Kho M., Turner S.T., Mosley T.H., Kardia S.L.R., Zhou X. Genetic Architecture of Gene Expression in European and African Americans: An eQTL Mapping Study in GENOA. Am. J. Hum. Genet. 2020;106:496–512. doi: 10.1016/j.ajhg.2020.03.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 44.Liu Y., Chen S., Li Z., Morrison A.C., Boerwinkle E., Lin X. ACAT: A Fast and Powerful p Value Combination Method for Rare-Variant Analysis in Sequencing Studies. Am. J. Hum. Genet. 2019;104:410–421. doi: 10.1016/j.ajhg.2019.01.002. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 45.Purcell S.M., Stone J.L., Stone J.L., Wray N.R., Sullivan P.F., Sullivan P.F., Sklar P., Stone J.L., O'Donovan M.C., et al. International Schizophrenia Consortium Common polygenic variation contributes to risk of schizophrenia and bipolar disorder. Nature. 2009;460:748–752. doi: 10.1038/nature08185. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 46.Chang C.C., Chow C.C., Tellier L.C., Vattikuti S., Purcell S.M., Lee J.J. Second-generation PLINK: rising to the challenge of larger and richer datasets. GigaScience. 2015;4:7. doi: 10.1186/s13742-015-0047-8. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 47.Tibshirani R. Regression Shrinkage and Selection via the Lasso. J. R. Stat. Soc. Ser. B Methodol. 1996;58:267–288. [Google Scholar]
- 48.Mak T.S.H., Porsch R.M., Choi S.W., Zhou X., Sham P.C. Polygenic scores via penalized regression on summary statistics. Genet. Epidemiol. 2017;41:469–480. doi: 10.1002/gepi.22050. [DOI] [PubMed] [Google Scholar]
- 49.Berisa T., Pickrell J.K. Approximately independent linkage disequilibrium blocks in human populations. Bioinforma. Oxf. Engl. 2016;32:283–285. doi: 10.1093/bioinformatics/btv546. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 50.Auton A., Abecasis G.R., Altshuler D.M., Durbin R.M., Abecasis G.R., Bentley D.R., Chakravarti A., Clark A.G., Donnelly P., Eichler E.E., et al. A global reference for human genetic variation. Nature. 2015;526:68–74. doi: 10.1038/nature15393. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 51.Browning S.R., Waples R.K., Browning B.L. Fast, accurate local ancestry inference with FLARE. Am. J. Hum. Genet. 2023;110:326–335. doi: 10.1016/j.ajhg.2022.12.010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 52.Maples B.K., Gravel S., Kenny E.E., Bustamante C.D. RFMix: A Discriminative Modeling Approach for Rapid and Robust Local-Ancestry Inference. Am. J. Hum. Genet. 2013;93:278–288. doi: 10.1016/j.ajhg.2013.06.020. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 53.Guan Y. Detecting Structure of Haplotypes and Local Ancestry. Genetics (Austin, Tex.) 2014;196:625–642. doi: 10.1534/genetics.113.160697. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 54.Salter-Townshend M., Myers S. Fine-Scale Inference of Ancestry Segments Without Prior Knowledge of Admixing Groups. Genetics (Austin, Tex.) 2019;212:869–889. doi: 10.1534/genetics.119.302139. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 55.Hou K., Gogarten S., Kim J., Hua X., Dias J.-A., Sun Q., Wang Y., Tan T., Atkinson E.G., et al. Polygenic Risk Methods in Diverse Populations (PRIMED) Consortium Methods Working Group Admix-kit: An integrated toolkit and pipeline for genetic analyses of admixed populations. bioRxiv. 2023 doi: 10.1101/2023.09.30.560263. Preprint at. [DOI] [Google Scholar]
- 56.Su Z., Marchini J., Donnelly P. HAPGEN2: simulation of multiple disease SNPs. Bioinformatics. 2011;27:2304–2305. doi: 10.1093/bioinformatics/btr341. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 57.Massarat A.R., Lamkin M., Reeve C., Williams A.L., D’Antonio M., Gymrek M. Haptools: a toolkit for admixture and haplotype analysis. Bioinformatics. 2023;39:btad104. doi: 10.1093/bioinformatics/btad104. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 58.Taliun D., Harris D.N., Kessler M.D., Carlson J., Szpiech Z.A., Torres R., Taliun S.A.G., Corvelo A., Gogarten S.M., Kang H.M., et al. Sequencing of 53,831 diverse genomes from the NHLBI TOPMed Program. Nature. 2021;590:290–299. doi: 10.1038/s41586-021-03205-y. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 59.Das S., Forer L., Schönherr S., Sidore C., Locke A.E., Kwong A., Vrieze S.I., Chew E.Y., Levy S., McGue M., et al. Next-generation genotype imputation service and methods. Nat. Genet. 2016;48:1284–1287. doi: 10.1038/ng.3656. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 60.Fuchsberger C., Abecasis G.R., Hinds D.A. minimac2: faster genotype imputation. Bioinformatics. 2015;31:782–784. doi: 10.1093/bioinformatics/btu704. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 61.Browning B.L., Zhou Y., Browning S.R. A One-Penny Imputed Genome from Next-Generation Reference Panels. Am. J. Hum. Genet. 2018;103:338–348. doi: 10.1016/j.ajhg.2018.07.015. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 62.Robinson M.D., Oshlack A. A scaling normalization method for differential expression analysis of RNA-seq data. Genome Biol. 2010;11:R25. doi: 10.1186/gb-2010-11-3-r25. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 63.Ongen H., Buil A., Brown A.A., Dermitzakis E.T., Delaneau O. Fast and efficient QTL mapper for thousands of molecular phenotypes. Bioinformatics. 2016;32:1479–1485. doi: 10.1093/bioinformatics/btv722. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 64.Alexander D.H., Novembre J., Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Res. 2009;19:1655–1664. doi: 10.1101/gr.094052.109. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 65.Ge T., Chen C.-Y., Ni Y., Feng Y.-C.A., Smoller J.W. Polygenic prediction via Bayesian regression and continuous shrinkage priors. Nat. Commun. 2019;10:1776. doi: 10.1038/s41467-019-09718-5. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 66.Uren C., Hoal E.G., Möller M. Putting RFMix and ADMIXTURE to the test in a complex admixed population. BMC Genet. 2020;21:40. doi: 10.1186/s12863-020-00845-3. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 67.Sollis E., Mosaku A., Abid A., Buniello A., Cerezo M., Gil L., Groza T., Güneş O., Hall P., Hayhurst J., et al. The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Res. 2023;51:D977–D985. doi: 10.1093/nar/gkac1010. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 68.Qi Y., Wang X., Li W., Chen D., Meng H., An S. Pseudogenes in Cardiovascular Disease. Front. Mol. Biosci. 2020;7 doi: 10.3389/fmolb.2020.622540. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 69.Mortlock S., Kendarsari R.I., Fung J.N., Gibson G., Yang F., Restuadi R., Girling J.E., Holdsworth-Carson S.J., Teh W.T., Lukowski S.W., et al. Tissue specific regulation of transcription in endometrium and association with disease. Hum. Reprod. 2020;35:377–393. doi: 10.1093/humrep/dez279. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 70.Dinsdale N.L., Crespi B.J. Endometriosis and polycystic ovary syndrome are diametric disorders. Evol. Appl. 2021;14:1693–1715. doi: 10.1111/eva.13244. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 71.Ferreira P.G., Muñoz-Aguirre M., Reverter F., Sá Godinho C.P., Sousa A., Amadoz A., Sodaei R., Hidalgo M.R., Pervouchine D., Carbonell-Caballero J., et al. The effects of death and post-mortem cold ischemia on human tissue transcriptomes. Nat. Commun. 2018;9:490. doi: 10.1038/s41467-017-02772-x. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 72.Lepik K., Annilo T., Kukuškina V., eQTLGen Consortium. Kisand K., Kutalik Z., Peterson P., Peterson H. C-reactive protein upregulates the whole blood expression of CD59 - an integrative analysis. PLoS Comput. Biol. 2017;13 doi: 10.1371/journal.pcbi.1005766. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 73.Kerimov N., Hayhurst J.D., Peikova K., Manning J.R., Walter P., Kolberg L., Samoviča M., Sakthivel M.P., Kuzmin I., Trevanion S.J., et al. A compendium of uniformly processed human gene expression and splicing quantitative trait loci. Nat. Genet. 2021;53:1290–1299. doi: 10.1038/s41588-021-00924-w. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 74.Pasaniuc B., Zaitlen N., Shi H., Bhatia G., Gusev A., Pickrell J., Hirschhorn J., Strachan D.P., Patterson N., Price A.L. Fast and accurate imputation of summary statistics enhances evidence of functional enrichment. Bioinformatics. 2014;30:2906–2914. doi: 10.1093/bioinformatics/btu416. [DOI] [PMC free article] [PubMed] [Google Scholar]
- 75.Wu Y., Eskin E., Sankararaman S. A Unifying Framework for Imputing Summary Statistics in Genome-Wide Association Studies. J. Comput. Biol. 2020;27:418–428. doi: 10.1089/cmb.2019.0449. [DOI] [PMC free article] [PubMed] [Google Scholar]
Associated Data
This section collects any data citations, data availability statements, or supplementary materials included in this article.
Supplementary Materials
Data Availability Statement
The GTEx V8 cis-eQTL summary statistics in Europeans used for the applied analysis of UKB subjects are publicly available and can be accessed at https://www.gtexportal.org/home/downloads/adult-gtex/qtl. The cis-eQTL summary statistics from high-African ancestry African Americans by Kachuri et al.17 also used in our applied UKB analysis can be accessed at https://zenodo.org/records/7735723. The eQTL summary statistics from the Estonian Biobank generated by Lepik et al.72 are available on the eQTL Catalog73 at https://www.ebi.ac.uk/eqtl/. All scripts and code used to generate figures shown in this paper are available from https://github.com/staylorhead/CADET.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.