Positive selection affects the human ADLH2 gene expression: genetic adaptation to alcohol consumption

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract ALDH2 is a key enzyme in alcohol metabolism that protects cells from acetaldehyde toxicity. Using iHS and FST statistics, we identified regulatory acting variants affecting ALDH2 gene expression under positive selection in populations of European ancestry. Several SNPs (rs3184504, rs4766578, rs10774625, rs597808, rs653178, rs847892, rs2013002) that function as eQTLs for ALDH2 in various tissues showed evidence of positive selection. Very large pairwise FST values indicated high genetic differentiation at these loci between populations of European ancestry and populations of other global ancestries. Estimating the timing of positive selection on the beneficial alleles suggests that these variants were recently adapted approximately 3000 to 3700 years ago. The derived beneficial alleles are in complete linkage disequilibrium with the derived ALDH2 promoter variant rs886205, which is associated with higher transcriptional activity. The SNPs rs4766578 and rs847892 are located in binding sequences for the transcription factor HNF4A, which is an important regulatory element of ALDH2 gene expression. In contrast to the missense variant ALDH2 rs671 (ALDH2*2), which is common only in East Asian populations and is associated with greatly reduced enzyme activity and alcohol intolerance, the beneficial alleles of the regulatory variants identified in this study are associated with increased expression of ALDH2. This suggests adaptation of Europeans to higher alcohol consumption.
Full text 179,030 characters · extracted from preprint-html · click to expand
Positive selection affects the human ADLH2 gene expression: genetic adaptation to alcohol consumption | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Positive selection affects the human ADLH2 gene expression: genetic adaptation to alcohol consumption Helmut Schaschl, Tobias Göllner, David L Morris This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-891422/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract ALDH2 is a key enzyme in alcohol metabolism that protects cells from acetaldehyde toxicity. Using iHS and F ST statistics, we identified regulatory acting variants affecting ALDH2 gene expression under positive selection in populations of European ancestry. Several SNPs (rs3184504, rs4766578, rs10774625, rs597808, rs653178, rs847892, rs2013002) that function as eQTLs for ALDH2 in various tissues showed evidence of positive selection. Very large pairwise F ST values indicated high genetic differentiation at these loci between populations of European ancestry and populations of other global ancestries. Estimating the timing of positive selection on the beneficial alleles suggests that these variants were recently adapted approximately 3000 to 3700 years ago. The derived beneficial alleles are in complete linkage disequilibrium with the derived ALDH2 promoter variant rs886205, which is associated with higher transcriptional activity. The SNPs rs4766578 and rs847892 are located in binding sequences for the transcription factor HNF4A , which is an important regulatory element of ALDH2 gene expression. In contrast to the missense variant ALDH2 rs671 ( ALDH2*2 ), which is common only in East Asian populations and is associated with greatly reduced enzyme activity and alcohol intolerance, the beneficial alleles of the regulatory variants identified in this study are associated with increased expression of ALDH2 . This suggests adaptation of Europeans to higher alcohol consumption. Molecular Genetics Evolutionary Biology Anthropology ALDH2 alcohol metabolism eQTLs positive selection genetic adaptation Figures Figure 1 1. Introduction The Neolithic transition from a hunter-gatherer lifestyle to an agriculturist one, about 9,000–13,000 years ago, included substantial changes in food processing and dietary habits associated with plant and animal domestication (Ye and Gu 2011). One of the key questions in biological anthropology is whether these changes resulted in selective pressure, influencing the expression of genes in the human genome. Identifying such loci has the potential to identify the underlying genetic variants contributing to the risk for various human diseases such as autoimmune diseases, cancer or cardiovascular disease (Luca et al. 2010; Valente et al. 2015; Ye and Gu 2011). Alcohol consumption and culture-related drinking behavior is probably one of the major changes in human dietary habits and lifestyle over the last 10,000 years. Production of larger amounts of alcoholic beverages had probably begun by the early Neolithic. A recent study reports archaeological evidence for cereal-based beer brewing by the semi-nomadic Natufians at Raqefet Cave (Mount Carmel in the north of Israel) dating back 11,700–13,700 years ago (Liu et al. 2018). Today, large amounts of alcohol are consumed in many societies. Recent data from the World Health Organization (WHO) show that worldwide about 3 million deaths and 132.6 million disability-adjusted life years are attributable to the harmful use of alcohol (World Health Organization 2018). In particular, Europe stands out in the WHO data as the region with the highest alcohol consumption and the highest burden of alcohol-related diseases. Large amounts of episodic drinking (binge drinking) as well as chronic alcohol consumption are associated with several very harmful effects such as alcoholic liver disease, intestinal inflammation, cancer, hypertension, brain damage including adverse behavioural changes, and decreased fertility (Liu et al. 2016; Ricci et al. 2017; Rocco et al. 2014; Roerecke et al. 2019). While heavy alcohol consumption can cause complex negative physiological effects, positive effects of moderate alcohol consumption have also been reported. Light to moderate alcohol consumption has been associated with a reduced risk of some forms of cardiovascular disease and autoimmune diseases (Aguet et al. 2017; Barbhaiya and Costenbader 2016; Fernandez-Sola 2015; Lu et al. 2014; Ronksley et al. 2011). The most commonly ingested alcohol is ethanol (EtOH, CH 3 CH 2 OH) which is absorbed from the gastrointestinal tract by passive diffusion. EtOH is oxidized in the first step, mainly in the liver, to acetaldehyde by alcohol dehydrogenases (ADH). The genes ADH4 , ADH1A , ADH1B and ADH1C , which are located in an array in the region of chromosome 4q23, encode closely related proteins and carry out most of the EtOH oxidation in liver. Cytochrome P450 2E1 ( CYP2E1 ) and the enzyme catalase ( CAT ) also participate in this metabolic pathway, albeit to a lesser extent. In the second step of EtOH metabolism, acetaldehyde, which is a chemically reactive and toxic compound, is oxidized by aldehyde dehydrogenases (ALDHs) to acetate (Cederbaum 2012; Edenberg and McClintick 2018). Several studies provide evidence that recent positive selection acts on the ADH1B locus in Asian, European and African populations (Galinsky et al. 2016; Gu et al. 2018; Han et al. 2007; Johnson and Voight 2018; Li et al. 2007; Peng et al. 2010). At this locus, two missense substitutions play a role at the SNPs rs1229984 (G > A; p.Arg48His) and rs2066702 (C > T; p.Arg370Cys). They express three different isoforms. The ADH1B*1 isoform with arginine at both codon positions is the most common allele globally, except in populations of East Asian ancestry. In East Asia, the derived allele ADH1B*2 (rs1229984) presents the common allele (with a frequency of about 0.70). The ADH1B*3 allele (rs2066702) occurs only in individuals of African ancestry, with allele frequencies ranging from 0.09 to 0.28. The two derived isoforms ADH1B*2 and ADH1B*3 metabolise EtOH at about 11 and 3 times the rate of ADH1B*1 , respectively, (Edenberg and McClintick 2018). Several studies report that rs1229984 in the ADH1B locus is associated with a reduced risk of alcoholism in different ancestries (Bierut et al. 2012; Craddock et al. 2010; Gelernter et al. 2014; Jorgenson et al. 2017; Thompson et al. 2020; Toth et al. 2010; Way et al. 2015; Xu et al. 2015). The positive selection on the derived allele was estimated to have occurred about 7,000 to 15,000 years ago (Peng et al. 2010; Peter et al. 2012; Smith et al. 2018), which overlaps with the time-frame of the origin and expansion of Neolithic agriculture in East Asia. Nonetheless, it remains unclear whether the driving selective force acting on this genetic polymorphism emanates from the protective effect against alcohol dependence or from the higher efficiency of this polymorphism in metabolizing EtOH. The mitochondrial enzyme ALDH2 plays the key role in the second step of EtOH metabolism by converting acetaldehyde into acetate. ALDH2 is not only a major detoxification enzyme for EtOH-derived acetaldehyde, but is also involved in detoxifying reactive aldehydes derived from reactive oxygen species (ROS) (Budas et al. 2009). Aldehydes are toxic molecules that can form genotoxic DNA- and protein-adducts in cells (Rodriguez-Zavala et al. 2019). Accumulation of high levels of acetaldehyde can be mutagenic, carcinogenic (Chen et al. 2014; Lee et al. 2017; Liu et al. 2016) and may negatively affects the immune system (Ceni et al. 2014). In contrast to the ADH genes, ADLH2 is expressed in most human tissues, with high levels in the liver, heart, kidney, and muscle tissues (Stewart et al. 1998). In the coding region of ALDH2 , the missense variant rs671 (G > A; p.Glu504Lys) expresses the isoforms ALDH2*1 and ALDH2*2 . The ALDH2*2 variant is found only in individuals of East Asian ancestry, reaching frequencies of up to 40% in some East Asian populations such as Han Chinese and Japanese (Edenberg and McClintick 2018; Eng et al. 2007; Li et al. 2009; Oota et al. 2004). This allele significantly affects alcohol metabolism because it results in an inactive enzyme and thus an excess of the toxic acetaldehyde in cells, even with moderate alcohol consumption. The symptoms are severe facial flushing, nausea, headache and tachycardia (Macgregor et al. 2009). East Asians homozygous for ALDH2*2 have a very low risk for alcohol dependency (Jorgenson et al. 2017; Macgregor et al. 2009; Quillen et al. 2014). The ALDH2 enzyme plays a central role in protecting cells from EtOH toxicity by metabolizing acetaldehyde (and other endogenous aldehyde products), is anti-inflammatory (Pan et al. 2016), and functions in myocardial protection (Ma et al. 2011; Panisello-Rosello et al. 2018; Zhang et al. 2012). Accordingly, this gene is of great biomedical interest. ALDH2 is located at the human chromosomal region 12q24.12. Several genome-wide association studies (GWAS) have found this genomic region to be associated with multiple human diseases such as rheumatoid arthritis (Coenen et al. 2009), systemic lupus erythematosus (Bentham et al. 2015), type 1 diabetes (Auburger et al. 2014), hypertension (Yasukochi et al. 2017) and coronary artery disease (Wild et al. 2017). This region, approximately 0.6 Mbp in size, encompasses in addition to ALDH2 the genes CUX2, FAM109A, SH2B2, ATXN2, BRAP, ACAD10 and MAPKAPK5 , as well as the uncharacterized transcript ENST00000546840.3 (UniProt F8VP50 - Aldedh domain-containing protein), which partially overlaps with the genes ACAD10 and ADLH2 . High F ST values at linked sites at the ALDH2 locus point to some form of selection for this genomic region (Oota et al. 2004). A recent study analysing rare singletons in the Japanese population identified the SNP rs3782886, which is in linkage disequilibrium (LD) with the missense SNP rs671 in the 12q24.12 region, as under recent positive selection (Okada et al. 2018). We have previously found (Oberreiter et al. 2021) a signature of positive selection at the chromosomal region 12q24.12 in European populations. In this study, we applied population genetic models of natural selection and included functional genetic data to identify the targets of positive selection in this genomic region. Several lines of evidence indicate that recent positive selection is acting on regulatory variants that influence ALDH2 gene expression in populations of European ancestry. 2. Materials And Methods 2.1. Genomic data We downloaded the phased genomic datasets from the 1000 Gnomes projects (phase 3; ftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/ ) (Auton et al. 2015). We obtained SNP data from 12 human populations: three representative populations each from African ancestry (AFR), European ancestry (EUR), South Asian ancestry (SAS) and East Asian ancestry (EAS) (populations names in accordance to the 1000 Genomes project – see Supplementary Table 1). We excluded related individuals and did not include the admixed populations from the datasets because of the underlying statistical principle of the method used to detect positive selection. We used the software program PLINK 1.9 (Chang et al. 2015) and VCFtools v0.1.14 (Danecek et al. 2011) to process the variant call format (VCF) files. We used the following filter parameters in VCFtools: --maf 0.05 (include only sites with a Minor Allele Frequency (MAF) greater than 0.05), --minQ 30 (include only sites with quality value above this threshold) and --remove-indels (exclude sites that contain an indel). Furthermore, we excluded all SNPs that deviated from Hardy-Weinberg equilibrium (with p -value < 1e-6) using PLINK --hwe midp threshold filter. We further excluded potential duplicated SNPS using bcftools version 1.10.2, ( https://github.com/samtools/bcftools/ ) using the parameter norm --Ov --check-ref w --fasta-ref human_g1k_v37.fasta ( ftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/ ). SNP positions are in accordance to the human genome version GRCh37/hg19 ( https://genome-euro.ucsc.edu/ ) (Kent et al. 2002). 2.2. Population genetic analyses To detect positive selection in phased genomic population data, we used the integrated Haplotype Score (iHS) approach (Voight et al. 2006), which is implemented in the software programme selscan version 1.2.0a (Szpiech and Hernandez 2014). All scans, with default selscan model parameters, were run on phased whole chromosome data (except the Y-chromosome) with a genetic map from HapMap phase II b37 (Altshuler et al. 2010). The iHS approach compares extended haplotype homozygosity (EHH) values between alleles at a given SNP. It is based on the model of a selective sweep, in which a de novo adaptive mutation arises on a haplotype that is rapidly fixed in the population, thereby reducing genetic diversity around that locus (Voight et al. 2006). The unstandardized iHS scores were normalized in default frequency bins across the entire genome using the script ‘norm’ provided by the selscan programme. Negative iHS values (iHS score 2.0) are associated with long haplotypes carrying the ancestral allele (Voight et al. 2006). We used the Ensembl Variant Effect Predictor programme package ( https://github.com/Ensembl/ensembl-vep ) (McLaren et al. 2016) to map genetic information such as gene symbol and biotype to the analysed SNPs. We calculated empirical p -values for the obtained iHS scores across all chromosomes using the R programme version 4.1.0 (R Core Team 2021). In this study we report only results for the human chromosomal region 12q24.12, the genomic location of the ALDH2 gene. We considered statistically significant ( p 2.4 or < -2.4; however, we also applied Bonferroni correction, which yields p -values p 4.2 or < -4.2). We used the script 'colormap.plotting.R' provided by the selscan package to display the EHH plots for the SNPs that are under positive selection. Pairwise F ST were calculated using Weir & Cockerham F ST calculation (Weir and Cockerham 1984) implemented in VCFtool (Danecek et al. 2011). Negative F ST values were set to zero. We calculated empirical p -values for the F ST values (across all chromosomes) to obtain the significant threshold ( p < 0.05) of outlier loci. In addition, locus-specific F ST values and standard deviations (sd) across all analysed populations were calculated for SNPs that were detected to be under positive selection with the Genetix programme version 4.05 (Belkhir et al. 2004) applying the jackknife resampling procedure. We used the R package ggplot2 (Wickham 2009) to plot iHS and F ST values. Allele frequency data, SNP information and ancestral/derived allele state were obtained from the Ensembl genome browser ( https://www.ensembl.org/index.html ) (Howe et al. 2021). We used LDlink, a web-based application ( https://analysistools.cancer.gov/LDlink/?tab=home ) (Machiela and Chanock 2015), to explore population-specific linkage disequilibrium (LD); we report D’ and goodness-of-fit statistics (χ 2 statistics). 2.3. Estimating the timing of positive selection We estimated the timing of selection on a beneficial allele using the R package Startmrca (Smith et al. 2018). The method applies a Markov chain Monte Carlo simulation (MCMC) that samples over the unknown ancestral haplotype to generate a sample of the posterior distribution for the time to the most recent common ancestor (TMRCA). The model takes advantage of both the length of the ancestral haplotype on each chromosome and the accumulation of derived mutations on the ancestral haplotype to generate a sample of the posterior distribution for the TMRCA. The model requires a sample (panel) containing the haplotypes with the selected allele and a reference panel of haplotypes without the selected allele. In this study, populations of European ancestry were used as samples, and populations of the other analysed genetic ancestries were used as reference panels. Because the calculated TMRCA estimates for the populations of European ancestry were very similar regardless of the reference panels used, we report in this study only the TMRCA estimates calculated for the European populations, using the European populations both as sample panel and as reference panel. We also estimated TMRCA for the East Asian-specific functional variants rs671 ( ALDH2 locus) and rs3782886 ( BRAP locus) (Okada et al. 2018) using the East Asian population CHB as sample and reference panel (see Supplementary Table 1 for the corresponding population names). After normalising the TMRCA data we calculated 95% credible intervals (CI = 0.95) for the timing estimates using the equal-tailed interval method implemented in the R package bayestestR (Makowski et al. 2019). We used recombination rates from the sex-averaged recombination map from deCODE to model recombination rate variation across the human genome. We analysed 1 Mb regions up- and downstream of the genetic variants under selection with an assumed mutation rate of 1.6 x 10 − 8 . We ran three independent MCMC chains, each with 25,000 iterations. We discarded the first 9000 iterations (burn–in), retaining the remaining iterations. We assumed 25 years as generation time. 2.4. GTEx and RegulomeDB functional data We utilized expression quantitative trait loci (eQTLs) (accessed between May and July 2021 (dbGaP Accession phs000424.v8.p2) from GTEx Portal V8 Release ( https://www.gtexportal.org/home/ ) (Ardlie et al. 2015) to test whether any of the potential SNPs that are under positive selection function as eQTL. We included cis -eQTL variants within a 1 Mb window of analysed genes. The RegulomeDB database ( https://regulomedb.org/ ) (Boyle et al. 2012) was used to obtained chromatin states; this database comprises known classes of genomic elements such as promoters, enhancers, transcription start sites, and transcription factor (TF) binding motifs. Additionally, mapped phenotype data were obtained from the NHGRI-EBI GWAS catalogue ( https://www.ebi.ac.uk/gwas/ ) (Buniello et al. 2019) (accessed between May and July 2021). 3. Results 3.1 Positive selection in populations of European ancestry The iHS analysis shows evidence that the human chromosomal region 12.q24.12 is under positive selection in populations of European ancestry. Figure 1 (a) plots the iHS scores in the European population GBR; Fig. 1 (b) shows the pairwise F ST values for GBR vs . the African population LWK across 12.q24.12. The red and green lines indicate significant ( p < 0.01 and after Bonferroni correction p 0.3). Table 1 SNPs under positive selection at the human chromosomal region 12q24.12 in populations with European ancestry (GBR, TSI, FIN). Given are iHS scores and the calculated (-log) p -values (in bold Bonferroni correction with p < 1x10 − 5 ), the timing ( t ) of positive selection on the derived beneficial allele in thousand years ago (kya) and 95% credible interval (CI) (rounded to one decimal figure), average allele frequency in % for the derived beneficial allele/ancestral allele in the different ancestries and global locus-specific F ST values (sd = standard deviation) calculated across all analysed populations. Chr Beneficial allele/ ancestral allele Location iHS iHS -logp t (kya) 95% CI Allele frequency in % Locus-specific F ST (sd) GBR TSI FIN AFR EUR SAS EAS 12q24.12 rs3184504-T/C Exon, SH2B3 -3.2 -3.3 -2.9 3.2; 3.3; 2.7 3.7 3.2–4.3 0.2/99.8 46/54 7/93 0.2/99.8 0.351 (0.060) rs4766578-T/A Intron, ATXN2 -3.8 -3.1 -3.1 4.1; 3.0; 3.0 3.5 3.0–4.0 0.2/99.8 48/52 7/93 0.2/99.8 0.366 (0.063) rs10774625-A/G Intron, ATXN2 -3.8 -3.1 -3.0 4.1; 3.0; 2.9 3.0 2.7–3.4 0.2/99.8 48/52 7/93 0.2/99.8 0.366 (0.062) rs597808-A/G Intron, ATXN2 -3.8 -4.2 -3.1 4.1; 4.8; 3.0 3.5 3.0–4.1 0.2/99.8 47/53 7/93 0.2/99.8 0.352 (0.064) rs653178-C/T Intron, ATXN2 -3.5 -4.4 -3.3 3.7; 5.2 ; 3.3 3.1 2.6–3.7 0.3/99.7 47/53 7/93 0/100 0.356 (0.063) rs847892-G/A Intron, ACAD10 -2.7 -2.7 -2.7 2.5; 2.5; 2.5 6.0 5.1–7.0 0.4/99.6 69/31 30/70 7/93 0.405 (0.080) rs2013002-T/C Intron, ENST 00000546840.3 -2.7 -2.1 -2.4 2.5; 1.8; 2.1 3.1 2.8–3.6 0.3/99.7 41/59 6/94 0.3/99.7 0.315 (0.053) 3.2 Positive selection acts on regulatory variants of ALDH2 From the GTEx database, we obtained in total 1591 cis -QTLs that influence ALDH2 gene expression (Supplementary Table 2); of these cis -QTLs, we identified 204, 217 and 53 eQTLs that had significant ( p < 0.01) iHS scores in the European samples GBR, TSI and FIN, respectively (Supplementary Tables 3–5). We also obtained cis -eQTLs (1970 in total) for the other protein-coding genes located in this genomic region ( CUX2, FAM109A, SH2B2, ATXN2, BRAP, ACAD10, MAPKAPK5 ) (Supplementary Table 6). In contrast to the eQTLs for ALDH2 , we did not obtain significant iHS values for these SNP eQTLs, except for SNPs that also function as eQTLs for ALDH2 . We further identified seven SNPs (rs3184504, rs4766578, rs10774625, rs597808, rs653178, rs847892, rs2013002) that are under positive selection in European populations that have very large global locus-specific F ST values > 0.3, i.e. are outlier loci (Table 1 ). The corresponding EHH plots and pairwise F ST values of these SNPs can be found in Supplementary Fig. 1 and Supplementary Table 7, respectively. The pairwise F ST values for these SNPs comparing populations of European ancestries vs . African, East Asian and South Asian ancestries ranged from 0.253 to 0.691. These SNPs function as eQTLs for ALDH2 , and the beneficial alleles are associated with increased ALDH2 gene expression in various human tissues (in accordance to the set of tissues represented in GTEx) such as esophagus – mucosa, skin, muscle – skeletal, brain – nucleus accumbens, artery – tibial, artery – aorta, and thyroid. The average allele frequencies for these SNPs are given in Table 1 ; for most of these SNPs the frequency of the derived beneficial alleles reaches almost 50% in the European populations. In contrast, the derived alleles are very rare (< 0.3%) in African and East Asian ancestries and at low frequency in populations of South Asian ancestry (< 7% with the exception of rs847892). We also compared the allele frequency at these loci with ancient Eurasians, including ancient hunter-gatherers (8.2–7.5 kya) from the study of Mathieson et al. (2015). The allele frequency data from the latter study show for the SNPs rs3184504, rs4766578, rs10774625, and rs653178 that the ancestral alleles were fixed in ancient European hunter-gatherers. As expected, the Neanderthal and Denisovan data on the UCSC Genome Browser also show only ancestral alleles at these loci. In contrast, early European farmers (8.4–4.2 kya) and individuals with steppe ancestry (5.4–3.6 kya) had frequencies between 8% and 25% of the derived alleles at these loci. In the European sample GBR we estimated the timing of positive selection of the derived beneficial alleles to be from about 3.0 to 3.7 kya with the exception of SNP rs847892, which we date at 6.0 kya (Table 1 ). This range of estimates are very similar to the TMRCA estimates calculated for the other European samples (TSI and FIN) (Supplementary Table 8). We also calculated the TMRCA for the derived allele of the East Asian-specific polymorphism (missense variants) at rs671-A/G and rs3782886-C/T in the East Asian population CHB, which yielded an estimation of 5.8 kya (CI: 4.8–6.7) and 5.4 kya (CI: 3.3–6.5), respectively. We then further included in the analysis the ALDH2 promoter variant rs886205-A/G, which is located − 360 bp from the ATG start codon of the ALDH2 gene (Chou et al. 1999). This promoter variant shows very large genetic differentiation with global locus-specific F ST = 0.378 (s.d. = 0.055). In the 1000 Genome data the derived allele A is the common allele in European and South Asian populations with average frequencies of about 83% and 71%, respectively. In contrast, in populations of African and East Asian ancestry the common allele is the ancestral allele G with frequencies of about 78% and 84%, respectively. For the ALDH2 promoter variant, a study showed (in vivo and in vitro experiments) that the − 360G (ancestral) allele has a significantly lower basal transcriptional activity than the − 360A (derived) allele (Kimura et al. 2009). Our LD analysis revealed that the positively selected SNPs are in complete LD ( D’ = 1) with the ALDH2 promoter variant rs886205 (Table 2 ). The chromatin state data from RegulomeDB showed that the identified SNPs are associated with active transcription start site (TSS), enhancers and strong transcription in different tissues (Table 3 ). Importantly, the positively selected SNPs rs4766578 and rs847892 are located in the binding motif for the transcription factor hepatocyte nuclear factor 4 alpha ( HNF4A ). This transcription factor is an important regulatory element of ALDH2 . The mapped phenotypes (Table 3 ) show that the positively selected SNPs are associated with various traits and diseases, in particular with blood pressure, cardiovascular disease, cholesterol level and autoimmune diseases. The variants rs597808 and rs2013002 are also associated with alcohol drinking and physiological traits such as blood pressure. We pooled related traits (Supplementary Table 9) into four main trait category namely autoimmune diseases (AIS), blood pressure (BP), cardiovascular disease (CDS) and cancer to test the null hypothesis that the traits and the allele state are independent. We found a significant (χ 2 = 28.828, df = 3, p -value = 2.4e-06) relationship between the allele state and trait; the derived beneficial alleles are positively associated with AIS, BP and CDS whereas the ancestral alleles with cancer. Table 2 Pairwise LD ( D’ ) of SNPs under positive selection in populations of European ancestry (GBR, TSI, FIN) and the ALDH2 promoter (*) variant rs886205; all calculated D’ values with p -value < 0.0001 (χ 2 statistics). Chr:pos SNP LD ( D’ ) rs3184504 rs4766578 rs10774625 rs597808 rs653178 rs847892 rs2013002 chr12:111884608 rs3184504 - chr12:111904371 rs4766578 1.0 - chr12:111910219 rs10774625 1.0 1.0 - chr12:111973358 rs597808 0.98 0.986 0.986 - chr12:112007756 rs653178 0.98 0.979 0.979 0.859 chr12:112141570 rs847892 0.846 0.851 0.851 0.846 1.0 - chr12:112200150 rs2013002 0.956 0.970 0.97 1.0 1.0 1.0 - chr12:112204427 *rs886205 1.0 1.0 1.0 1.0 1.0 1.0 1.0 Table 3 GTEx and RegulomeDB data on SNPs under positive selection in European populations (GBR, TSI, FIN). Given is also a summary of reported traits from the NHGRI-EBI GWAS catalogue. GTEx eQTLs–eGene interaction with p < 0.0001. RegulomeDB rank: 2b: TF binding + any motif + DNase Footprint + DNase peak; 3a: TF binding + any motif + DNase peak; 4–5: TF binding + DNase peak; 6: motif hit. The RegulomeDB probability score ranges from 0 to 1, with 1 being most likely to be a regulatory variant (for further details see {Boyle, 2012} and {Dong, 2019}). Transcription factor HNF4A , an important regulatory element of the ALDH2 gene expression, is given in bold. GTEx RegulomeDB GWAS reported traits eQTL eGene rank score chromatin state motif rs3184504 ALDH2, LINC01405, TMEM116 3a 0.67022 strong transcription; enhancers MTF1 Cardiovascular disease, blood pressure, ischemic stroke, glaucoma, rheumatoid arthritis, cancer, celiac disease, type I diabetes mellitus, parental longevity, inflammatory bowel disease, multiple sclerosis, blood cell count, hypothyroidism, haemoglobin measurement. rs4766578 ALDH2, TMEM116 2b 0.63936 strong transcription; enhancers ESRRA, ESRRB , HNF4A , NR6A1 Sjögren's syndrome, reticulocyte fraction of red cells, arthritis, vitiligo, HDL cholesterol, smoking status, coronary artery disease. rs10774625 ALDH2, ADAM1B, TMEM116 5 0 strong transcription; enhancers FOXJ2, FOXQ1 Hypertension, myocardial infarction, coronary artery disease, asthma, cholesterol levels, systemic lupus erythematosus, urate measurement, life span, systolic blood pressure, hypothyroidism, glomerular filtration rate. rs597808 ALDH2, LINC01405, ADAM1B 5 0.13454 strong transcription - Systolic blood pressure, alcohol drinking, diastolic blood pressure, cholesterol levels, apolipoprotein B levels, colorectal cancer, allergic diseases, haematocrit, systemic lupus erythematosus, allergy. rs653178 ALDH2, LINC01405 4 0.60906 active TSS; strong transcription; enhancers - Allergic disease, asthma, celiac disease, cholesterol level, eczema, Crohn's disease, chronic kidney disease, blood pressure, eosinophil counts, inflammatory bowel disease, type 1 diabetes, urate level. rs847892 ALDH2, TMEM116, NAA25 6 0.20016 active TSS; strong transcription; enhancers HNF4A No data. rs2013002 ALDH2, ADAM1B 6 0.55195 active TSS; enhancers MAFB, MAFK Alcohol drinking and blood pressure. 4. Discussion 4.1 Evidence of position selection acting on regulatory variants of ALDH2 This study provides evidence of positive selection across the human chromosomal region 12q24.12. This finding is in line with two previous studies (Akbari et al. 2018; Barreiro and Quintana-Murci 2010). We further found that this genomic region is enriched in eQTLs that influence ALDH2 gene expression. A high number of these SNP eQTLs had significant iHS scores in the populations of European ancestry. In contrast, cis -eQTLs of the other genes located at chr12q24.12 showed no significant iHS values. This indicates that the target of positive selection are regulatory acting variants that influence ALDH2 gene expression. We identified seven SNPs that are under positive selection and show very large global locus-specific F ST values (> 0.3), indicating high genetic differentiation between populations of European ancestry and populations from other global ancestries. The GTEx data show that these SNPs function primarily as eQTLs for the ALDH2 gene. The derived beneficial alleles at these SNP eQTLs are associated with increased expression of ALDH2 in multiple human tissues. Importantly, the two variants rs4766578 and rs847892 are located in binding sequences for transcription factor HNF4 . That transcription factor is considered to be a master regulator of liver-specific gene expression (Bolotin et al. 2010) and is an important regulatory element of ALDH2 gene expression (Stewart et al. 1998; You et al. 2002). Furthermore, these SNP eQTLs are in complete LD with the ALDH2 promoter variant rs847892. This promoter polymorphism influences individual differences in acetaldehyde elimination. The ancestral allele G, the common allele in populations of African and East Asian ancestry, has a lower basal transcriptional activity than the derived allele A, the common allele in populations of European and South Asian ancestry (Kimura et al. 2009). These results suggest that higher transcriptional activity and increased ALDH2 expression in individuals of European ancestry represent a form of genetic adaptation to increased alcohol consumption, possibly enabling faster detoxification of acetaldehyde. The derived beneficial alleles of these loci reach almost 50% in the European population, whereas in African and East Asian populations the frequencies are very low (< 0.003). The ancestral alleles at these positively selected loci appear to be fixed in ancient European hunter-gatherers, but in early farmers and individuals with steppe ancestry the frequencies of the derived alleles already range between 8% and 25% (Mathieson et al. 2015). The estimated timing of positive selection on the beneficial alleles in the European population GBR ranges from about 3.0 kya to 3.7 kya (except for rs847892 for which TMRCA was estimated to 6.0 kya). We also calculated the TMRCA for the East Asian-specific derived alleles rs671-A and rs3782886-C, yielding an estimation of 5.8 kya (CI: 4.8–6.7) and 5.4 kya (CI: 4.3–6.5), respectively. Rs3782886, which is in LD with rs671, shows signals of very recent selection for the past 2000–3000 years in the Japanese population as reported in a recent study (Okada et al. 2018). Noteworthy, rs671-A and rs1229984-A ( ADH1B locus) were found in a subsequent study to be significantly associated with better survival in the Japanese population (Sakaue et al. 2020). The estimated TMRCA for the derived alleles in our study suggests that these alleles spread in East Asia at a much earlier time than the beneficial alleles in populations of European ancestry. Archaeological evidence indicates early production of fermented alcohol in China (McGovern et al. 2004). Analysis of starch granules, phytoliths and fungi in food residues adhering to 8000–7000 year-old alcohol-making pottery vessels suggests that, in East Asia in the early Neolithic, alcoholic beverages were already being produced (Liu et al. 2019). For Europe, archaeologically recognizable brewing material in Central European lakeside settlements show that alcoholic beverages were being produced in this region in the late Neolithic period about 6000 years ago (Heiss et al. 2020). Later, in Greek-Roman antiquity, a richly developed viticulture with high wine production was achieved and, in this period, wine became part of the daily diet of many people (Retief and Cilliers 2015). Alcohol consumption has apparently increased steadily since then in Europe, especially in the 19th century. In Germany, for example, the high level of consumption, in particular of strong spirits, in the early 19th century was – in analogy to the plague – referred to as Branntweinpest (brandy plague). Since the rs671-A allele leads to an inactive enzyme and thus to an excess of toxic acetaldehyde in cells with its negative physiological effects, we suggest that this allele may explain the differences in the signature of positive selection between populations of European and East Asian ancestry. 4.2 Wider biomedical context The ALDH2 enzyme plays a critical role both in the detoxification of both acetaldehyde and ROS-generated aldehyde adducts such as 4-hydroxy-2-nonenal and malondialdehyde. This enzyme thus has cytoprotective effects reducing oxidative stress (Budas et al. 2009; Guo et al. 2013). In particular, the ALDH2*2 variant (rs671), which is common only in individuals of East Asian ancestry, has been intensively studied in East Asians. While individuals with the ALDH2*2 allele have a reduced risk of developing alcoholism, it increases their cancer risk (Crabb et al. 2004; Zhang and Fu 2021). Nevertheless, this allele was found to be associated in the Japanese population with better survival (Sakaue et al. 2020). In European populations this allele is virtually absent. In our study, however, the identified variants that are under recent positive selection in European populations act as regulatory variants and are associated with increased ADLH2 gene expression in various human tissues. This suggests that individuals carrying these beneficial alleles should be more quickly able to detoxify the body from higher amounts of acetaldehyde and ROS-generated aldehyde adducts. However, a recent study reports higher methylation in alcohol-dependent patients compared to controls in the ALDH2 promoter region (Pathak et al. 2017). Furthermore, that study suggests that positive and negative regulatory elements interact at the ALDH2 promoter to induce genotype-mediated epigenetic changes, leading to differential transcriptional activity of this gene. We therefore suggest that individuals carrying the beneficial alleles may be able to consume more alcohol (over longer time periods), but may also have a higher likelihood of becoming heavy drinkers and alcohol dependent. This, then, could lead to increased methylation of the ALDH2 promoter, resulting in decreased ALDH2 gene expression. Accordingly, the protective effects of ALDH2 against oxidative damage through acetaldehyde would be lost, resulting in increased risk of numerous oxidative stress-related diseases such as cancer, diabetes, inflammatory disorders and cardiovascular conditions such as hypertension and stroke. Indeed, we found in this study that the derived beneficial alleles are positively associated with AIS, BP and CDS whereas the ancestral alleles with cancer. 5. Conclusions We found that positive selection acts on regulatory variants affecting ALDH2 gene expression in populations of European ancestry. In contrast to the known functional consequence of the ALDH2*2 variant (rs671) in East Asians, which is associated with alcohol intolerance, in Europeans the beneficial derived alleles are associated with increased ALDH2 gene expression. This suggests local adaptation to higher alcohol consumption in Europeans. Estimation of the timing of positive selection on the beneficial alleles suggests that these variants were recently adapted, approximately 3000 to 3700 years ago. We hypothesize that the beneficial effects of higher ALDH2 expression leads to an increased detoxification capacity for acetaldehyde, but possibly also to increased likelihood of chronic alcohol abuse, leading to decreased ALDH2 expression and thus increased cell toxicity from EtOH-derived acetaldehyde as well as from ROS-generated aldehydes. Declarations Funding: Not applicable. Conflicts of interest: All authors declare that the research was conducted without commercial or financial relationships and therefore no potential conflicts of interest occurred. Availability of data and material : The gnomic data can be obtained from 1000 Genomes database. The generated iHS and Fst dataset are available from the corresponding author on reasonable request. Code availability: Not applicable. Data Availability Statement: Not applicable. Authors' contributions: HS conceived and designed this study; HS, TG and DM performed statistical analyses. HS wrote the draft manuscript; all authors contributed to the results, edited, read and approved the final manuscript. Ethics approval: Not applicable. Consent to participate: Not applicable. The 1000 Genomes data are publicly available. Consent for publication: Not applicable. Acknowledgments: We thank Michael Stachowitsch from the Department of Evolutionary Anthropology, University of Vienna and Franz Suchentrunk from the University of Veterinary Medicine Vienna for valuable comments on the manuscript. References Aguet F, Brown AA, Castel SE, Davis JR, He Y, Jo B, Mohammadi P, Park Y, Parsana P, Segre AV, Strober BJ, Zappala Z, Cummings BB, Gelfand ET, Hadley K, Huang KH, Lek M, Li X, Nedzel JL, Nguyen DY, Noble MS, Sullivan TJ, Tukiainen T, MacArthur DG, Getz G, Management NP, Addington A, Guan P, Koester S, Little AR, Lockhart NC, Moore HM, Rao A, Struewing JP, Volpi S, Collection B, Brigham LE, Hasz R, Hunter M, Johns C, Johnson M, Kopen G, Leinweber WF, Lonsdale JT, McDonald A, Mestichelli B, Myer K, Roe B, Salvatore M, Shad S, Thomas JA, Walters G, Washington M, Wheeler J, Bridge J, Foster BA, Gillard BM, Karasik E, Kumar R, Miklos M, Moser MT, Jewell SD, Montroy RG, Rohrer DC, Valley D, Mash DC, Davis DA, Sobin L, Barcus ME, Branton PA, Grp EMW, Abell NS, Balliu B, Delaneau O, Fresard L, Gamazon ER, Garrido-Martin D, Gewirtz ADH, Gliner G, Gloudemans MJ, Han B, He AZ, Hormozdiari F, Liu B, Kang EY, McDowell IC, Ongen H, Palowitch JJ, Peterson CB, Quon G, Ripke S, Saha A, Shabalin AA, Shimko TC, Sul JH, Teran NA, Tsang EK, Zhang H, Zhou YH, Bustamante CD, et al. (2017) Genetic effects on gene expression across human tissues. Nature 550: 204-+. Akbari A, Vitti JJ, Iranmehr A, Bakhtiari M, Sabeti PC, Mirarab S, Bafna V (2018) Identifying the favored mutation in a positive selective sweep. Nature Methods 15: 279-+. Altshuler DM, Gibbs RA, Peltonen L, Dermitzakis E, Schaffner SF, Yu FL, Bonnen PE, de Bakker PIW, Deloukas P, Gabriel SB, Gwilliam R, Hunt S, Inouye M, Jia XM, Palotie A, Parkin M, Whittaker P, Chang K, Hawes A, Lewis LR, Ren YR, Wheeler D, Muzny DM, Barnes C, Darvishi K, Hurles M, Korn JM, Kristiansson K, Lee C, McCarroll SA, Nemesh J, Keinan A, Montgomery SB, Pollack S, Price AL, Soranzo N, Gonzaga-Jauregui C, Anttila V, Brodeur W, Daly MJ, Leslie S, McVean G, Moutsianas L, Nguyen H, Zhang QR, Ghori MJR, McGinnis R, McLaren W, Takeuchi F, Grossman SR, Shlyakhter I, Hostetter EB, Sabeti PC, Adebamowo CA, Foster MW, Gordon DR, Licinio J, Manca MC, Marshall PA, Matsuda I, Ngare D, Wang VO, Reddy D, Rotimi CN, Royal CD, Sharp RR, Zeng CQ, Brooks LD, McEwen JE, Int HapMap C (2010) Integrating common and rare genetic variation in diverse human populations. Nature 467: 52-58. Ardlie KG, DeLuca DS, Segre AV, Sullivan TJ, Young TR, Gelfand ET, Trowbridge CA, Maller JB, Tukiainen T, Lek M, Ward LD, Kheradpour P, Iriarte B, Meng Y, Palmer CD, Esko T, Winckler W, Hirschhorn JN, Kellis M, MacArthur DG, Getz G, Shabalin AA, Li G, Zhou YH, Nobel AB, Rusyn I, Wright FA, Lappalainen T, Ferreira PG, Ongen H, Rivas MA, Battle A, Mostafavi S, Monlong J, Sammeth M, Mele M, Reverter F, Goldmann JM, Koller D, Guigo R, McCarthy MI, Dermitzakis ET, Gamazon ER, Im HK, Konkashbaev A, Nicolae DL, Cox NJ, Flutre T, Wen XQ, Stephens M, Pritchard JK, Tu ZD, Zhang B, Huang T, Long Q, Lin L, Yang JL, Zhu J, Liu J, Brown A, Mestichelli B, Tidwell D, Lo E, Salvatore M, Shad S, Thomas JA, Lonsdale JT, Moser MT, Gillard BM, Karasik E, Ramsey K, Choi C, Foster BA, Syron J, Fleming J, Magazine H, Hasz R, Walters GD, Bridge JP, Miklos M, Sullivan S, Barker LK, Traino HM, Mosavel M, Siminoff LA, Valley DR, Rohrer DC, Jewell SD, Branton PA, Sobin LH, Barcus M, Qi LQ, McLean J, Hariharan P, Um KS, Wu SP, Tabor D, Shive C, Smith AM, Buia SA, et al. (2015) The Genotype-Tissue Expression (GTEx) pilot analysis: Multitissue gene regulation in humans. Science 348: 648-660. Auburger G, Gispert S, Lahut S, Omur O, Damrath E, Heck M, Basak N (2014) 12q24 locus association with type 1 diabetes: SH2B3 or ATXN2 ? World Journal of Diabetes 5: 316-327. Auton A, Abecasis GR, Altshuler DM, Durbin RM, Bentley DR, Chakravarti A, Clark AG, Donnelly P, Eichler EE, Flicek P, Gabriel SB, Gibbs RA, Green ED, Hurles ME, Knoppers BM, Korbel JO, Lander ES, Lee C, Lehrach H, Mardis ER, Marth GT, McVean GA, Nickerson DA, Schmidt JP, Sherry ST, Wang J, Wilson RK, Boerwinkle E, Doddapaneni H, Han Y, Korchina V, Kovar C, Lee S, Muzny D, Reid JG, Zhu Y, Chang Y, Feng Q, Fang X, Guo X, Jian M, Jiang H, Jin X, Lan T, Li G, Li J, Li Y, Liu S, Liu X, Lu Y, Ma X, Tang M, Wang B, Wang G, Wu H, Wu R, Xu X, Yin Y, Zhang D, Zhang W, Zhao J, Zhao M, Zheng X, Gupta N, Gharani N, Toji LH, Gerry NP, Resch AM, Barker J, Clarke L, Gil L, Hunt SE, Kelman G, Kulesha E, Leinonen R, McLaren WM, Radhakrishnan R, Roa A, Smirnov D, Smith RE, Streeter I, Thormann A, Toneva I, Vaughan B, Zheng-Bradley X, Grocock R, Humphray S, James T, Kingsbury Z, Sudbrak R, Albrecht MW, Amstislavskiy VS, Borodina TA, Lienhard M, Mertes F, Sultan M, Timmermann B, Yaspo M-L, Fulton L, Fulton R, et al. (2015) A global reference for human genetic variation. Nature 526: 68-74. Barbhaiya M, Costenbader KH (2016) Environmental exposures and the development of systemic lupus erythematosus. Current Opinion in Rheumatology 28: 497-505. Barreiro LB, Quintana-Murci L (2010) From evolutionary genetics to human immunology: how selection shapes host defence genes. Nature Reviews Genetics 11: 17-30. Belkhir K, Borsa P, Chikhi L, Raufaste N, Bonhomme F (2004) GENETIX4. 05, logiciel sous Windows TM pour la génétiquedes populations. Laboratoire génome, populations, interactions, CNRS UMR 5000: 1996-2004. Bentham J, Morris DL, Graham DSC, Pinder CL, Tombleson P, Behrens TW, Martin J, Fairfax BP, Knight JC, Chen L, Replogle J, Syvanen A-C, Ronnblom L, Graham RR, Wither JE, Rioux JD, Alarcon-Riquelme ME, Vyse TJ (2015) Genetic association analyses implicate aberrant regulation of innate and adaptive immunity genes in the pathogenesis of systemic lupus erythematosus. Nature Genetics 47: 1457-+. Bierut LJ, Goate AM, Breslau N, Johnson EO, Bertelsen S, Fox L, Agrawal A, Bucholz KK, Grucza R, Hesselbrock V, Kramer J, Kuperman S, Nurnberger J, Porjesz B, Saccone NL, Schuckit M, Tischfield J, Wang JC, Foroud T, Rice JP, Edenberg HJ (2012) ADH1B is associated with alcohol dependence and alcohol consumption in populations of European and African ancestry. Molecular Psychiatry 17: 445-450. Bolotin E, Liao HL, Ta TC, Yang CH, Hwang-Verslues W, Evans JR, Jiang T, Sladek FM (2010) Integrated Approach for the Identification of Human Hepatocyte Nuclear Factor 4 alpha Target Genes Using Protein Binding Microarrays. Hepatology 51: 642-653. Boyle AP, Hong EL, Hariharan M, Cheng Y, Schaub MA, Kasowski M, Karczewski KJ, Park J, Hitz BC, Weng S, Cherry JM, Snyder M (2012) Annotation of functional variation in personal genomes using RegulomeDB. Genome Research 22: 1790-1797. Budas GR, Disatnik MH, Mochly-Rosen D (2009) Aldehyde Dehydrogenase 2 in Cardiac Protection: A New Therapeutic Target? Trends in Cardiovascular Medicine 19: 158-164. Buniello A, MacArthur JAL, Cerezo M, Harris LW, Hayhurst J, Malangone C, McMahon A, Morales J, Mountjoy E, Sollis E, Suveges D, Vrousgou O, Whetzel PL, Amode R, Guillen JA, Riat HS, Trevanion SJ, Hall P, Junkins H, Flicek P, Burdett T, Hindorff LA, Cunningham F, Parkinson H (2019) The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Research 47: D1005-D1012. Cederbaum AI (2012) Alcohol Metabolism. Clinics in Liver Disease 16: 667-+. Ceni E, Mello T, Galli A (2014) Pathogenesis of alcoholic liver disease: Role of oxidative metabolism. World Journal of Gastroenterology 20: 17756-17772. Chang CC, Chow CC, Tellier L, Vattikuti S, Purcell SM, Lee JJ (2015) Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience 4. Chen CH, Ferreira JCB, Gross ER, Mochly-Rosen D (2014) TARGETING ALDEHYDE DEHYDROGENASE 2: NEW THERAPEUTIC OPPORTUNITIES. Physiological Reviews 94: 1-34. Chou WY, Stewart MJ, Carr LG, Zheng D, Stewart TR, Williams A, Pinaire J, Crabb DW (1999) An A/G polymorphism in the promoter of mitochondrial aldehyde dehydrogenase (ALDH2): Effects of the sequence variant on transcription factor binding and promoter strength. Alcoholism-Clinical and Experimental Research 23: 963-968. Coenen MJH, Trynka G, Heskamp S, Franke B, van Diemen CC, Smolonska J, van Leeuwen M, Brouwer E, Boezen MH, Postma DS, Platteel M, Zanen P, Lammers J, Groen HJM, Mali W, Mulder CJ, Tack GJ, Verbeek WHM, Wolters VM, Houwen RHJ, Mearin ML, van Heel DA, Radstake T, van Riel P, Wijmenga C, Barrera P, Zhernakova A (2009) Common and different genetic background for rheumatoid arthritis and coeliac disease. Human Molecular Genetics 18: 4195-4203. Crabb DW, Matsumoto M, Chang D, You M (2004) Overview of the role of alcohol dehydrogenase and aldehyde dehydrogenase and their variants in the genesis of alcohol-related pathology. Proceedings of the Nutrition Society 63: 49-63. Craddock N, Hurles ME, Cardin N, Pearson RD, Plagnol V, Robson S, Vukcevic D, Barnes C, Conrad DF, Giannoulatou E, Holmes C, Marchini JL, Stirrups K, Tobin MD, Wain LV, Yau C, Aerts J, Ahmad T, Andrews TD, Arbury H, Attwood A, Auton A, Ball SG, Balmforth AJ, Barrett JC, Barroso I, Barton A, Bennett AJ, Bhaskar S, Blaszczyk K, Bowes J, Brand OJ, Braund PS, Bredin F, Breen G, Brown MJ, Bruce IN, Bull J, Burren OS, Burton J, Byrnes J, Caesar S, Clee CM, Coffey AJ, Connell JMC, Cooper JD, Dominiczak AF, Downes K, Drummond HE, Dudakia D, Dunham A, Ebbs B, Eccles D, Edkins S, Edwards C, Elliot A, Emery P, Evans DM, Evans G, Eyre S, Farmer A, Ferrier IN, Feuk L, Fitzgerald T, Flynn E, Forbes A, Forty L, Franklyn JA, Freathy RM, Gibbs P, Gilbert P, Gokumen O, Gordon-Smith K, Gray E, Green E, Groves CJ, Grozeva D, Gwilliam R, Hall A, Hammond N, Hardy M, Harrison P, Hassanali N, Hebaishi H, Hines S, Hinks A, Hitman GA, Hocking L, Howard E, Howard P, Howson JMM, Hughes D, Hunt S, Isaacs JD, Jain M, Jewell DP, Johnson T, Jolley JD, Jones IR, Jones LA, et al. (2010) Genome-wide association study of CNVs in 16,000 cases of eight common diseases and 3,000 shared controls. Nature 464: 713-U86. Danecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA, Handsaker RE, Lunter G, Marth GT, Sherry ST, McVean G, Durbin R, Genomes Project Anal G (2011) The variant call format and VCFtools. Bioinformatics 27: 2156-2158. Edenberg HJ, McClintick JN (2018) Alcohol Dehydrogenases, Aldehyde Dehydrogenases, and Alcohol Use Disorders: A Critical Review. Alcoholism-Clinical and Experimental Research 42: 2281-2297. Eng MY, Luczak SE, Wall TL (2007) ALDH2, ADH1B, and ADH1C genotypes in Asians: A literature review. Alcohol Research & Health 30: 22-27. Fernandez-Sola J (2015) Cardiovascular risks and benefits of moderate and heavy alcohol consumption. Nature Reviews Cardiology 12: 576-587. Galinsky KJ, Bhatia G, Loh PR, Georgiev S, Mukherjee S, Patterson NJ, Price AL (2016) Fast Principal-Component Analysis Reveals Convergent Evolution of ADH1B in Europe and East Asia. American Journal of Human Genetics 98: 456-472. Gelernter J, Kranzler HR, Sherva R, Almasy L, Koesterer R, Smith AH, Anton R, Preuss UW, Ridinger M, Rujescu D, Wodarz N, Zill P, Zhao H, Farrer LA (2014) Genome-wide association study of alcohol dependence: significant findings in African-and European-Americans including novel risk loci. Molecular Psychiatry 19: 41-49. Guo JM, Liu AJ, Zang P, Dong WZ, Ying L, Wang W, Xu P, Song XR, Cai J, Zhang SQ, Duan JL, Mehta JL, Su DF (2013) ALDH2 protects against stroke by clearing 4-HNE. Cell Research 23: 915-930. Gu S, Li H, Pakstis AJ, Speed WC, Gurwitz D, Kidd JR, Kidd KK (2018) Recent Selection on a Class I ADH Locus Distinguishes Southwest Asian Populations Including Ashkenazi Jews. Genes 9. Han Y, Gu S, Oota H, Osier MV, Pakstis AJ, Speed WC, Kidd JR, Kidd KK (2007) Evidence of positive selection on a class I ADH locus. American Journal of Human Genetics 80: 441-456. Heiss AG, Azorin MB, Antolin F, Kubiak-Martens L, Marinova E, Arendt EK, Biliaderis CG, Kretschmer H, Lazaridou A, Stika HP, Zarnkow M, Baba M, Bleicher N, Cialowicz KM, Chlodnicki M, Matuschik I, Schlichtherle H, Valamoti SM (2020) Mashes to Mashes, Crust to Crust. Presenting a novel microstructural marker for malting in the archaeological record. Plos One 15. Howe KL, Achuthan P, Allen J, Alvarez-Jarreta J, Amode MR, Armean IM, Azov AG, Bennett R, Bhai J, Billis K, Boddu S, Charkhchi M, Cummins C, Fioretto LR, Davidson C, Dodiya K, El Houdaigui B, Fatima R, Gall A, Giron CG, Grego T, Guijarro-Clarke C, Haggerty L, Hemrom A, Hourlier T, Izuogu OG, Juettemann T, Kaikala V, Kay M, Lavidas I, Le T, Lemos D, Martinez JG, Marugan JC, Maurel T, McMahon AC, Mohanan S, Moore B, Muffato M, Oheh DN, Paraschas D, Parker A, Parton A, Prosovetskaia I, Sakthivel MP, Salam AIA, Schmitt BM, Schuilenburg H, Sheppard N, Steed E, Szpak M, Szuba M, Taylor K, Thormann A, Threadgold G, Walts B, Winterbottom A, Chakiachvili M, Chaubal A, De Silva N, Flint B, Frankish A, Hunt SE, Iisley GR, Langridge N, Loveland JE, Martin FJ, Mudge JM, Morales J, Perry E, Ruffier M, Tate J, Thybert D, Trevanion SJ, Cunningham F, Yates AD, Zerbino DR, Flicek P (2021) Ensembl 2021. Nucleic Acids Research 49: D884-D891. Johnson KE, Voight BF (2018) Patterns of shared signatures of recent positive selection across human populations. Nature Ecology & Evolution 2: 713-720. Jorgenson E, Thai KK, Hoffmann TJ, Sakoda LC, Kvale MN, Banda Y, Schaefer C, Risch N, Mertens J, Weisner C, Choquet H (2017) Genetic contributors to variation in alcohol consumption vary by race/ethnicity in a large multi-ethnic genome-wide association study. Molecular Psychiatry 22: 1359-1367. Kent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D (2002) The human genome browser at UCSC. Genome Research 12: 996-1006. Kimura Y, Nishimura FT, Abe S, Fukunaga T, Tanii H, Saijoh K (2009) A Promoter Polymorphism in the ALDH2 Gene Affects Its Basal and Acetaldehyde/Ethanol-Induced Gene Expression in Human Peripheral Blood Leukocytes and HepG2 Cells. Alcohol and Alcoholism 44: 261-266. Lee DJ, Lee HM, Kim JH, Park IS, Rho YS (2017) Heavy alcohol drinking downregulates ALDH2 gene expression but heavy smoking up-regulates SOD2 gene expression in head and neck squamous cell carcinoma. World Journal of Surgical Oncology 15. Li H, Borinskaya S, Yoshimura K, Kal'ina N, Marusin A, Stepanov VA, Qin ZD, Khaliq S, Lee MY, Yang YJ, Mohyuddin A, Gurwitz D, Mehdi SQ, Rogaev E, Jin L, Yankovsky NK, Kidd JR, Kidd KK (2009) Refined Geographic Distribution of the Oriental ALDH2*504Lys (nee 487Lys) Variant. Annals of Human Genetics 73: 335-345. Li H, Mukherjee N, Soundararajan U, Tarnok Z, Barta C, Khaliq S, Mohyuddin A, Kajuna SLB, Mehdi SQ, Kidd JR, Kidd KK (2007) Geographically separate increases in the frequency of the derived ADH1B*47His allele in eastern and western Asia. American Journal of Human Genetics 81: 842-846. Liu L, Wang JJ, Levin MJ, Sinnott-Armstrong N, Zhao H, Zhao YA, Shao J, Di N, Zhang TE (2019) The origins of specialized pottery and diverse alcohol fermentation techniques in Early Neolithic China. Proceedings of the National Academy of Sciences of the United States of America 116: 12767-12774. Liu L, Wang JJ, Rosenberg D, Zhao H, Lengyel G, Nadel D (2018) Fermented beverage and food storage in 13,000 y-old stone mortars at Raqefet Cave, Israel: Investigating Natufian ritual feasting. Journal of Archaeological Science-Reports 21: 783-793. Liu X, Hu C, Bao MH, Li J, Liu XY, Tan XR, Zhou Y, Chen YQ, Wu SL, Chen SH, Zhang R, Jiang F, Jia WP, Wang XY, Yang XC, Cai J (2016) Genome Wide Association Study Identifies L3MBTL4 as a Novel Susceptibility Gene for Hypertension. Scientific Reports 6. Lu B, Solomon DH, Costenbader KH, Karlson EW (2014) Alcohol Consumption and Risk of Incident Rheumatoid Arthritis in Women A Prospective Study. Arthritis & Rheumatology 66: 1998-2005. Luca F, Perry GH, Di Rienzo A (2010) Evolutionary Adaptations to Dietary Changes. Annual Review of Nutrition, Vol 30 30: 291-314. Macgregor S, Lind PA, Bucholz KK, Hansell NK, Madden PAF, Richter MM, Montgomery GW, Martin NG, Heath AC, Whitfield JB (2009) Associations of ADH and ALDH2 gene variation with self report alcohol reactions, consumption and dependence: an integrated analysis. Human Molecular Genetics 18: 580-593. Machiela MJ, Chanock SJ (2015) LDlink: a web-based application for exploring population-specific haplotype structure and linking correlated alleles of possible functional variants. Bioinformatics 31: 3555-3557. Ma H, Guo R, Yu L, Zhang YM, Ren J (2011) Aldehyde dehydrogenase 2 (ALDH2) rescues myocardial ischaemia/reperfusion injury: role of autophagy paradox and toxic aldehyde. European Heart Journal 32: 1025-1038. Makowski D, Ben-Shachar MS, Chen SHA, Ludecke D (2019) Indices of Effect Existence and Significance in the Bayesian Framework. Frontiers in Psychology 10. Mathieson I, Lazaridis I, Rohland N, Mallick S, Patterson N, Roodenberg SA, Harney E, Stewardson K, Fernandes D, Novak M, Sirak K, Gamba C, Jones ER, Llamas B, Dryomov S, Pickrell J, Arsuaga JL, de Castro JMB, Carbonell E, Gerritsen F, Khokhlov A, Kuznetsov P, Lozano M, Meller H, Mochalov O, Moiseyev V, Guerra MAR, Roodenberg J, Verges JM, Krause J, Cooper A, Alt KW, Brown D, Anthony D, Lalueza-Fox C, Haak W, Pinhasi R, Reich D (2015) Genome-wide patterns of selection in 230 ancient Eurasians. Nature 528: 499-+. McGovern PE, Zhang JH, Tang JG, Zhang ZQ, Hall GR, Moreau RA, Nunez A, Butrym ED, Richards MP, Wang CS, Cheng GS, Zhao ZJ (2004) Fermented beverages of pre- and proto-historic China. Proceedings of the National Academy of Sciences of the United States of America 101: 17593-17598. McLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, Flicek P, Cunningham F (2016) The Ensembl Variant Effect Predictor. Genome Biology 17. Oberreiter V, Goellner T, Morris LD, Schaschl H (2021) Positive Selection Affects the Expression of Systemic Lupus Erythematosus Associated Loci in Human Populations. BMC Genomics. Okada Y, Momozawa Y, Sakaue S, Kanai M, Ishigaki K, Akiyama M, Kishikawa T, Arai Y, Sasaki T, Kosaki K, Suematsu M, Matsuda K, Yamamoto K, Kubo M, Hirose N, Kamatani Y (2018) Deep whole-genome sequencing reveals recent selection signatures linked to evolution and disease risk of Japanese. Nature Communications 9. Oota H, Pakstis AJ, Bonne-Tamir B, Goldman D, Grigorenko E, Kajuna SLB, Karoma NJ, Kungulilo S, Lu RB, Odunsi K, Okonofua F, Zhukova OV, Kidd JR, Kidd KK (2004) The evolution and population genetics of the ALDH2 locus: random genetic drift, selection, and low levels of recombination. Annals of Human Genetics 68: 93-109. World Health Organization (2018) Global status report on alcohol and health 2018. Geneva Pan C, Xing JH, Zhang C, Zhang YM, Zhang LT, Wei SJ, Zhang MX, Wang XP, Yuan QH, Xue L, Wang JL, Cui ZQ, Zhang Y, Xu F, Chen YG (2016) Aldehyde dehydrogenase 2 inhibits inflammatory response and regulates atherosclerotic plaque. Oncotarget 7: 35562-35576. Panisello-Rosello A, Lopez A, Folch-Puy E, Carbonell T, Rolo A, Palmeira C, Adam R, Net M, Rosello-Catafau J (2018) Role of aldehyde dehydrogenase 2 in ischemia reperfusion injury: An update. World Journal of Gastroenterology 24: 2984-2994. Pathak H, Frieling H, Bleich S, Glahn A, Heberlein A, Nassab MH, Hillemacher T, Burkert A, Rhein M (2017) Promoter Polymorphism rs886205 Genotype Interacts With DNA Methylation of the ALDH2 Regulatory Region in Alcohol Dependence. Alcohol and Alcoholism 52: 269-276. Peng Y, Shi H, Qi XB, Xiao CJ, Zhong H, Ma RLZ, Su B (2010) The ADH1B Arg47His polymorphism in East Asian populations and expansion of rice domestication in history. Bmc Evolutionary Biology 10. Peter BM, Huerta-Sanchez E, Nielsen R (2012) Distinguishing between Selective Sweeps from Standing Variation and from a De Novo Mutation. Plos Genetics 8. Quillen EE, Chen XD, Almasy L, Yang F, He H, Li X, Wang XY, Liu TQ, Hao W, Deng HW, Kranzler HR, Gelernter J (2014) ALDH2 Is Associated to Alcohol Dependence and Is the Major Genetic Determinant of "Daily Maximum Drinks" in a GWAS Study of an Isolated Rural Chinese Sample. American Journal of Medical Genetics Part B-Neuropsychiatric Genetics 165: 103-110. R Core Team (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https://www.R-project.org/. Retief F, Cilliers L Wine in Graeco-Roman Antiquity with Emphasis on Its Effect on Health 2015 Ricci E, Al Beitawi S, Cipriani S, Candiani M, Chiaffarino F, Vigano P, Noli S, Parazzini F (2017) Semen quality and alcohol intake: a systematic review and meta-analysis. Reproductive Biomedicine Online 34: 38-47. Rocco A, Compare D, Angrisani D, Zamparelli MS, Nardone G (2014) Alcoholic disease: Liver and beyond. World Journal of Gastroenterology 20: 14652-14659. Rodriguez-Zavala JS, Calleja LF, Moreno-Sanchez R, Yoval-Sanchez B (2019) Role of Aldehyde Dehydrogenases in Physiopathological Processes. Chemical Research in Toxicology 32: 405-420. Roerecke M, Vafaei A, Hasan OSM, Chrystoja BR, Cruz M, Lee R, Neuman MG, Rehm J (2019) Alcohol Consumption and Risk of Liver Cirrhosis: A Systematic Review and Meta-Analysis. American Journal of Gastroenterology 114: 1574-1586. Ronksley PE, Brien SE, Turner BJ, Mukamal KJ, Ghali WA (2011) Association of alcohol consumption with selected cardiovascular disease outcomes: a systematic review and meta-analysis. Bmj-British Medical Journal 342. Sakaue S, Akiyama M, Hirata M, Matsuda K, Murakami Y, Kubo M, Kamatani Y, Okada Y (2020) Functional variants in ADH1B and ALDH2 are non-additively associated with all-cause mortality in Japanese population. European Journal of Human Genetics 28: 378-382. Smith J, Coop G, Stephens M, Novembre J (2018) Estimating Time to the Common Ancestor for a Beneficial Allele. Molecular Biology and Evolution 35: 1003-1017. Stewart MJ, Dipple KM, Estonius M, Nakshatri H, Everett LM, Crabb DW (1998) Binding and activation of the human aldehyde dehydrogenase 2 promoter by hepatocyte nuclear factor 4. Biochimica Et Biophysica Acta-Gene Structure and Expression 1399: 181-186. Szpiech ZA, Hernandez RD (2014) selscan: An Efficient Multithreaded Program to Perform EHH-Based Scans for Positive Selection. Molecular Biology and Evolution 31: 2824-2827. Team RC (2021) R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https://www.R-project.org/. Vienna, Austria Thompson A, Cook J, Choquet H, Jorgenson E, Yin J, Kinnunen T, Barclay J, Morris AP, Pirmohamed M (2020) Functional validity, role, and implications of heavy alcohol consumption genetic loci. Science Advances 6. Toth R, Pocsai Z, Fiatal S, Szeles G, Kardos L, Petrovski B, McKee M, Adany R (2010) ADH1B*2 allele is protective against alcoholism but not chronic liver disease in the Hungarian population. Addiction 105: 891-896. Valente C, Alvarez L, Marks SJ, Lopez-Parra AM, Parson W, Oosthuizen O, Oosthuizen E, Amorim A, Capelli C, Arroyo-Pardo E, Gusmao L, Prata MJ (2015) Exploring the relationship between lifestyles, diets and genetic adaptations in humans. Bmc Genetics 16. Voight BF, Kudaravalli S, Wen XQ, Pritchard JK (2006) A map of recent positive selection in the human genome (vol 4, pg 154, 2006). Plos Biology 4: 659-659. Way M, McQuillin A, Saini J, Ruparelia K, Lydall GJ, Guerrini I, Ball D, Smith I, Quadri G, Thomson AD, Kasiakogia-Worlley K, Cherian R, Gunwardena P, Rao H, Kottalgi G, Patel S, Hillman A, Douglas E, Qureshi SY, Reynolds G, Jauhar S, O'Kane A, Dedman A, Sharp S, Kandaswamy R, Dar K, Curtis D, Morgan MY, Gurling HMD (2015) Genetic variants in or near ADH1B and ADH1C affect susceptibility to alcohol dependence in a British and Irish population. Addiction Biology 20: 594-604. Weir BS, Cockerham CC (1984) ESTIMATING F-STATISTICS FOR THE ANALYSIS OF POPULATION-STRUCTURE. Evolution 38: 1358-1370. Wickham H (2009) ggplot2: Elegant Graphics for Data Analysis. Ggplot2: Elegant Graphics for Data Analysis: 1-212. Wild PS, Felix JF, Schillert A, Teumer A, Chen MH, Leening MJG, Volker U, Grossmann V, Brody JA, Irvin MR, Shah SJ, Pramana S, Lieb W, Schmidt R, Stanton AV, Malzahn D, Smith AV, Sundstrom J, Minelli C, Ruggiero D, Lyytikainen LP, Tiller D, Smith JG, Monnereau C, Di Tullio MR, Musani SK, Morrison AC, Pers TH, Morley M, Kleber ME, Aragam J, Benjamin EJ, Bis JC, Bisping E, Broeckel U, Cheng S, Deckers JW, Del Greco MF, Edelmann F, Fornage M, Franke L, Friedrich N, Harris TB, Hofer E, Hofman A, Huang J, Hughes AD, Kahonen M, Kruppa J, Lackner KJ, Lannfelt L, Laskowski R, Launer LJ, Leosdottir M, Lin HH, Lindgren CM, Loley C, MacRae CA, Mascalzoni D, Mayet J, Medenwald D, Morris AP, Muller C, Muller-Nurasyid M, Nappo S, Nilsson PM, Nuding S, Nutile T, Peters A, Pfeufer A, Pietzner D, Pramstaller PP, Raitakari OT, Rice KM, Rivadeneira F, Rotter JI, Ruohonen ST, Sacco RL, Samdarshi TE, Schmidt H, Sharp ASP, Shields DC, Sorice R, Sotoodehnia N, Stricker BH, Surendran P, Thom S, Toglhofer AM, Uitterlinden AG, Wachter R, Volzke H, Ziegler A, Munzel T, Marz W, Cappola TP, Hirschhorn JN, Mitchell GF, Smith NL, Fox ER, Dueker ND, et al. (2017) Large-scale genome-wide analysis identifies genetic variants associated with cardiac structure and function. Journal of Clinical Investigation 127: 1798-1812. Xu K, Kranzler HR, Sherva R, Sartor CE, Almasy L, Koesterer R, Zhao HY, Farrer LA, Gelernter J (2015) Genomewide Association Study for Maximum Number of Alcoholic Drinks in European Americans and African Americans. Alcoholism-Clinical and Experimental Research 39: 1137-1147. Yasukochi Y, Sakuma J, Takeuchi I, Kato K, Oguri M, Fujimaki T, Horibe H, Yamada Y (2017) Longitudinal exome-wide association study to identify genetic susceptibility loci for hypertension in a Japanese population. Experimental and Molecular Medicine 49. Ye KX, Gu ZL (2011) Recent Advances in Understanding the Role of Nutrition in Human Genome Evolution. Advances in Nutrition 2: 486-496. You M, Fischer M, Cho WK, Crabb D (2002) Transcriptional control of the human aldehyde dehydrogenase 2 promoter by hepatocyte nuclear factor 4: Inhibition by cyclic AMP and COUP transcription factors. Archives of Biochemistry and Biophysics 398: 79-86. Zhang H, Fu LW (2021) The role of ALDH2 in tumorigenesis and tumor progression: Targeting ALDH2 as a potential cancer treatment. Acta Pharmaceutica Sinica B 11: 1400-1411. Zhang H, Gong DX, Zhang YJ, Li SJ, Hu SS (2012) Effect of mitochondrial aldehyde dehydrogenase-2 genotype on cardioprotection in patients with congenital heart disease. European Heart Journal 33: 1606-1614. Supplementary Files Supplementaryfile1.docx Supplementaryfile2.xlsx Supplementaryfile3.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-891422","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":51064102,"identity":"526fda39-3999-4470-96be-c923f0c6d311","order_by":0,"name":"Helmut Schaschl","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABH0lEQVRIie3QMUvDQBTA8XcE7BLN+opgv8IFoSpE8lXuEMxSB5fi4JASvC4W1xP9EAddHA2FuOQDRFowxcHFIaVLu4jX1KHDWRwd7g8Jx/F+vBAAm+0fRsGJ1yfUD4mPD1ZvYFuJHqgn1gQPNwmJ/0L45pSRHDV6N8PdBYTeQ/JeLZ8wUuNEYHkFLS92PkoDOblNxdhlwOUkazcHOV6oSSqQ5eDL5x3f+GEFrwmjyNoOEZoUPKFcAFEAZvI2rUlIMZrPNYloTb4gVNCYmbeQmhCFHbqvCdOkV+qfwBW45i05F6+P58hl0ek2BwL9+xVhGZ7JkXspTeRllBWfQRB6MhpWSxG09oqoTBfXweldv68qA/kJTTfO7/M2m81m29436CdqGqRUslEAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-2826-8773","institution":"University of Vienna: Universitat Wien","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Helmut","middleName":"","lastName":"Schaschl","suffix":""},{"id":51064103,"identity":"5d291348-8bd8-4b98-af2c-0d4078924754","order_by":1,"name":"Tobias Göllner","email":"","orcid":"","institution":"University of Vienna: Universitat Wien","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Tobias","middleName":"","lastName":"Göllner","suffix":""},{"id":51064104,"identity":"097695f8-efc8-4165-82ca-68db980430da","order_by":2,"name":"David L Morris","email":"","orcid":"","institution":"King's College London","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"David","middleName":"L","lastName":"Morris","suffix":""}],"badges":[],"createdAt":"2021-09-09 16:52:01","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-891422/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-891422/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":13318637,"identity":"d5853d59-42cd-42b4-ad23-f2ee12f84661","added_by":"auto","created_at":"2021-09-13 15:06:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":164721,"visible":true,"origin":"","legend":"(a.) iHS p-values plotted across the human chromosomal region 12.q24.12 for the population GBR (European genetic ancestry); red/green lines: threshold for significant (p \u003c 0.01; Bonferroni correction p \u003c 1x10-5) iHS scores; (b.) pairwise FST (GBR–LWK); red line: significant outlier loci with FST \u003e 0.3. Bottom: position of genes and SNPs from Table 1.","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-891422/v1/4e7302c56b71e68183bdea41.png"},{"id":14565192,"identity":"69b335af-1ead-42cd-9899-c6e115e1a270","added_by":"auto","created_at":"2021-10-15 21:46:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":631609,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-891422/v1/34fc805c-0846-4548-96ba-54bee5ba8f1d.pdf"},{"id":13318969,"identity":"01ef741a-9537-4b4c-88b2-032b29f1bac1","added_by":"auto","created_at":"2021-09-13 15:09:04","extension":"docx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":242015,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfile1.docx","url":"https://assets-eu.researchsquare.com/files/rs-891422/v1/c0a138e385a138a0564e8aa0.docx"},{"id":13318968,"identity":"8063ee33-8481-4bfb-83ad-098cc3a2380c","added_by":"auto","created_at":"2021-09-13 15:09:04","extension":"xlsx","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":119228,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-891422/v1/33ad6e59f92c76c48dc98569.xlsx"},{"id":13318639,"identity":"629175ff-40cf-4076-b7b9-1ebd24f9d23e","added_by":"auto","created_at":"2021-09-13 15:06:04","extension":"xlsx","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":14589,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfile3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-891422/v1/77a17af25bb7047bd334b1f0.xlsx"}],"financialInterests":"","formattedTitle":"Positive selection affects the human ADLH2 gene expression: genetic adaptation to alcohol consumption","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe Neolithic transition from a hunter-gatherer lifestyle to an agriculturist one, about 9,000\u0026ndash;13,000 years ago, included substantial changes in food processing and dietary habits associated with plant and animal domestication (Ye and Gu 2011). One of the key questions in biological anthropology is whether these changes resulted in selective pressure, influencing the expression of genes in the human genome. Identifying such loci has the potential to identify the underlying genetic variants contributing to the risk for various human diseases such as autoimmune diseases, cancer or cardiovascular disease (Luca et al. 2010; Valente et al. 2015; Ye and Gu 2011). Alcohol consumption and culture-related drinking behavior is probably one of the major changes in human dietary habits and lifestyle over the last 10,000 years. Production of larger amounts of alcoholic beverages had probably begun by the early Neolithic. A recent study reports archaeological evidence for cereal-based beer brewing by the semi-nomadic Natufians at Raqefet Cave (Mount Carmel in the north of Israel) dating back 11,700\u0026ndash;13,700 years ago (Liu et al. 2018). Today, large amounts of alcohol are consumed in many societies. Recent data from the World Health Organization (WHO) show that worldwide about 3\u0026nbsp;million deaths and 132.6\u0026nbsp;million disability-adjusted life years are attributable to the harmful use of alcohol (World Health Organization 2018). In particular, Europe stands out in the WHO data as the region with the highest alcohol consumption and the highest burden of alcohol-related diseases. Large amounts of episodic drinking (binge drinking) as well as chronic alcohol consumption are associated with several very harmful effects such as alcoholic liver disease, intestinal inflammation, cancer, hypertension, brain damage including adverse behavioural changes, and decreased fertility (Liu et al. 2016; Ricci et al. 2017; Rocco et al. 2014; Roerecke et al. 2019). While heavy alcohol consumption can cause complex negative physiological effects, positive effects of moderate alcohol consumption have also been reported. Light to moderate alcohol consumption has been associated with a reduced risk of some forms of cardiovascular disease and autoimmune diseases (Aguet et al. 2017; Barbhaiya and Costenbader 2016; Fernandez-Sola 2015; Lu et al. 2014; Ronksley et al. 2011). The most commonly ingested alcohol is ethanol (EtOH, CH\u003csub\u003e3\u003c/sub\u003eCH\u003csub\u003e2\u003c/sub\u003eOH) which is absorbed from the gastrointestinal tract by passive diffusion. EtOH is oxidized in the first step, mainly in the liver, to acetaldehyde by alcohol dehydrogenases (ADH). The genes \u003cem\u003eADH4\u003c/em\u003e, \u003cem\u003eADH1A\u003c/em\u003e, \u003cem\u003eADH1B\u003c/em\u003e and \u003cem\u003eADH1C\u003c/em\u003e, which are located in an array in the region of chromosome 4q23, encode closely related proteins and carry out most of the EtOH oxidation in liver. Cytochrome P450 2E1 (\u003cem\u003eCYP2E1\u003c/em\u003e) and the enzyme catalase (\u003cem\u003eCAT\u003c/em\u003e) also participate in this metabolic pathway, albeit to a lesser extent. In the second step of EtOH metabolism, acetaldehyde, which is a chemically reactive and toxic compound, is oxidized by aldehyde dehydrogenases (ALDHs) to acetate (Cederbaum 2012; Edenberg and McClintick 2018).\u003c/p\u003e \u003cp\u003eSeveral studies provide evidence that recent positive selection acts on the \u003cem\u003eADH1B\u003c/em\u003e locus in Asian, European and African populations (Galinsky et al. 2016; Gu et al. 2018; Han et al. 2007; Johnson and Voight 2018; Li et al. 2007; Peng et al. 2010). At this locus, two missense substitutions play a role at the SNPs rs1229984 (G\u0026thinsp;\u0026gt;\u0026thinsp;A; p.Arg48His) and rs2066702 (C\u0026thinsp;\u0026gt;\u0026thinsp;T; p.Arg370Cys). They express three different isoforms. The \u003cem\u003eADH1B*1\u003c/em\u003e isoform with arginine at both codon positions is the most common allele globally, except in populations of East Asian ancestry. In East Asia, the derived allele \u003cem\u003eADH1B*2\u003c/em\u003e (rs1229984) presents the common allele (with a frequency of about 0.70). The \u003cem\u003eADH1B*3\u003c/em\u003e allele (rs2066702) occurs only in individuals of African ancestry, with allele frequencies ranging from 0.09 to 0.28. The two derived isoforms \u003cem\u003eADH1B*2\u003c/em\u003e and \u003cem\u003eADH1B*3\u003c/em\u003e metabolise EtOH at about 11 and 3 times the rate of \u003cem\u003eADH1B*1\u003c/em\u003e, respectively, (Edenberg and McClintick 2018). Several studies report that rs1229984 in the \u003cem\u003eADH1B\u003c/em\u003e locus is associated with a reduced risk of alcoholism in different ancestries (Bierut et al. 2012; Craddock et al. 2010; Gelernter et al. 2014; Jorgenson et al. 2017; Thompson et al. 2020; Toth et al. 2010; Way et al. 2015; Xu et al. 2015). The positive selection on the derived allele was estimated to have occurred about 7,000 to 15,000 years ago (Peng et al. 2010; Peter et al. 2012; Smith et al. 2018), which overlaps with the time-frame of the origin and expansion of Neolithic agriculture in East Asia. Nonetheless, it remains unclear whether the driving selective force acting on this genetic polymorphism emanates from the protective effect against alcohol dependence or from the higher efficiency of this polymorphism in metabolizing EtOH.\u003c/p\u003e \u003cp\u003eThe mitochondrial enzyme ALDH2 plays the key role in the second step of EtOH metabolism by converting acetaldehyde into acetate. ALDH2 is not only a major detoxification enzyme for EtOH-derived acetaldehyde, but is also involved in detoxifying reactive aldehydes derived from reactive oxygen species (ROS) (Budas et al. 2009). Aldehydes are toxic molecules that can form genotoxic DNA- and protein-adducts in cells (Rodriguez-Zavala et al. 2019). Accumulation of high levels of acetaldehyde can be mutagenic, carcinogenic (Chen et al. 2014; Lee et al. 2017; Liu et al. 2016) and may negatively affects the immune system (Ceni et al. 2014). In contrast to the \u003cem\u003eADH\u003c/em\u003e genes, \u003cem\u003eADLH2\u003c/em\u003e is expressed in most human tissues, with high levels in the liver, heart, kidney, and muscle tissues (Stewart et al. 1998). In the coding region of \u003cem\u003eALDH2\u003c/em\u003e, the missense variant rs671 (G\u0026thinsp;\u0026gt;\u0026thinsp;A; p.Glu504Lys) expresses the isoforms \u003cem\u003eALDH2*1\u003c/em\u003e and \u003cem\u003eALDH2*2\u003c/em\u003e. The \u003cem\u003eALDH2*2\u003c/em\u003e variant is found only in individuals of East Asian ancestry, reaching frequencies of up to 40% in some East Asian populations such as Han Chinese and Japanese (Edenberg and McClintick 2018; Eng et al. 2007; Li et al. 2009; Oota et al. 2004). This allele significantly affects alcohol metabolism because it results in an inactive enzyme and thus an excess of the toxic acetaldehyde in cells, even with moderate alcohol consumption. The symptoms are severe facial flushing, nausea, headache and tachycardia (Macgregor et al. 2009). East Asians homozygous for \u003cem\u003eALDH2*2\u003c/em\u003e have a very low risk for alcohol dependency (Jorgenson et al. 2017; Macgregor et al. 2009; Quillen et al. 2014). The ALDH2 enzyme plays a central role in protecting cells from EtOH toxicity by metabolizing acetaldehyde (and other endogenous aldehyde products), is anti-inflammatory (Pan et al. 2016), and functions in myocardial protection (Ma et al. 2011; Panisello-Rosello et al. 2018; Zhang et al. 2012). Accordingly, this gene is of great biomedical interest. \u003cem\u003eALDH2\u003c/em\u003e is located at the human chromosomal region 12q24.12. Several genome-wide association studies (GWAS) have found this genomic region to be associated with multiple human diseases such as rheumatoid arthritis (Coenen et al. 2009), systemic lupus erythematosus (Bentham et al. 2015), type 1 diabetes (Auburger et al. 2014), hypertension (Yasukochi et al. 2017) and coronary artery disease (Wild et al. 2017). This region, approximately 0.6 Mbp in size, encompasses in addition to \u003cem\u003eALDH2\u003c/em\u003e the genes \u003cem\u003eCUX2, FAM109A, SH2B2, ATXN2, BRAP, ACAD10\u003c/em\u003e and \u003cem\u003eMAPKAPK5\u003c/em\u003e, as well as the uncharacterized transcript ENST00000546840.3 (UniProt F8VP50 - Aldedh domain-containing protein), which partially overlaps with the genes \u003cem\u003eACAD10\u003c/em\u003e and \u003cem\u003eADLH2\u003c/em\u003e. High \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values at linked sites at the \u003cem\u003eALDH2\u003c/em\u003e locus point to some form of selection for this genomic region (Oota et al. 2004). A recent study analysing rare singletons in the Japanese population identified the SNP rs3782886, which is in linkage disequilibrium (LD) with the missense SNP rs671 in the 12q24.12 region, as under recent positive selection (Okada et al. 2018). We have previously found (Oberreiter et al. 2021) a signature of positive selection at the chromosomal region 12q24.12 in European populations. In this study, we applied population genetic models of natural selection and included functional genetic data to identify the targets of positive selection in this genomic region. Several lines of evidence indicate that recent positive selection is acting on regulatory variants that influence \u003cem\u003eALDH2\u003c/em\u003e gene expression in populations of European ancestry.\u003c/p\u003e"},{"header":"2. Materials And Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1. Genomic data\u003c/h2\u003e \u003cp\u003eWe downloaded the phased genomic datasets from the 1000 Gnomes projects (phase 3; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003eftp://ftp.1000genomes.ebi.ac.uk/vol1/ftp/release/20130502/\u003c/span\u003e\u003c/span\u003e) (Auton et al. 2015). We obtained SNP data from 12 human populations: three representative populations each from African ancestry (AFR), European ancestry (EUR), South Asian ancestry (SAS) and East Asian ancestry (EAS) (populations names in accordance to the 1000 Genomes project \u0026ndash; see Supplementary Table\u0026nbsp;1). We excluded related individuals and did not include the admixed populations from the datasets because of the underlying statistical principle of the method used to detect positive selection. We used the software program PLINK 1.9 (Chang et al. 2015) and VCFtools v0.1.14 (Danecek et al. 2011) to process the variant call format (VCF) files. We used the following filter parameters in VCFtools: \u003cem\u003e--maf 0.05\u003c/em\u003e (include only sites with a Minor Allele Frequency (MAF) greater than 0.05), \u003cem\u003e--minQ 30\u003c/em\u003e (include only sites with quality value above this threshold) and \u003cem\u003e--remove-indels\u003c/em\u003e (exclude sites that contain an indel). Furthermore, we excluded all SNPs that deviated from Hardy-Weinberg equilibrium (with \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;1e-6) using PLINK \u003cem\u003e--hwe midp\u003c/em\u003e threshold filter. We further excluded potential duplicated SNPS using bcftools version 1.10.2, (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/samtools/bcftools/\u003c/span\u003e\u003c/span\u003e) using the parameter norm \u003cem\u003e--Ov --check-ref w --fasta-ref human_g1k_v37.fasta\u003c/em\u003e (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003eftp.1000genomes.ebi.ac.uk/vol1/ftp/technical/reference/\u003c/span\u003e\u003c/span\u003e). SNP positions are in accordance to the human genome version GRCh37/hg19 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://genome-euro.ucsc.edu/\u003c/span\u003e\u003c/span\u003e) (Kent et al. 2002).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2. Population genetic analyses\u003c/h2\u003e \u003cp\u003eTo detect positive selection in phased genomic population data, we used the integrated Haplotype Score (iHS) approach (Voight et al. 2006), which is implemented in the software programme selscan version 1.2.0a (Szpiech and Hernandez 2014). All scans, with default selscan model parameters, were run on phased whole chromosome data (except the Y-chromosome) with a genetic map from HapMap phase II b37 (Altshuler et al. 2010). The iHS approach compares extended haplotype homozygosity (EHH) values between alleles at a given SNP. It is based on the model of a selective sweep, in which a \u003cem\u003ede novo\u003c/em\u003e adaptive mutation arises on a haplotype that is rapidly fixed in the population, thereby reducing genetic diversity around that locus (Voight et al. 2006). The unstandardized iHS scores were normalized in default frequency bins across the entire genome using the script \u0026lsquo;norm\u0026rsquo; provided by the selscan programme. Negative iHS values (iHS score \u0026lt; -2.0) indicate unusually long haplotypes carrying the derived allele, and significant positive values (iHS score\u0026thinsp;\u0026gt;\u0026thinsp;2.0) are associated with long haplotypes carrying the ancestral allele (Voight et al. 2006). We used the Ensembl Variant Effect Predictor programme package (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/Ensembl/ensembl-vep\u003c/span\u003e\u003c/span\u003e) (McLaren et al. 2016) to map genetic information such as gene symbol and biotype to the analysed SNPs. We calculated empirical \u003cem\u003ep\u003c/em\u003e-values for the obtained iHS scores across all chromosomes using the R programme version 4.1.0 (R Core Team 2021). In this study we report only results for the human chromosomal region 12q24.12, the genomic location of the \u003cem\u003eALDH2\u003c/em\u003e gene. We considered statistically significant (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01) iHS scores\u0026thinsp;\u0026gt;\u0026thinsp;2.4 or \u0026lt; -2.4; however, we also applied Bonferroni correction, which yields \u003cem\u003ep\u003c/em\u003e-values \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;1x10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e (=\u0026thinsp;iHS scores\u0026thinsp;\u0026gt;\u0026thinsp;4.2 or \u0026lt; -4.2). We used the script 'colormap.plotting.R' provided by the selscan package to display the EHH plots for the SNPs that are under positive selection. Pairwise \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e were calculated using Weir \u0026amp; Cockerham \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e calculation (Weir and Cockerham 1984) implemented in VCFtool (Danecek et al. 2011). Negative \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values were set to zero. We calculated empirical \u003cem\u003ep\u003c/em\u003e-values for the \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values (across all chromosomes) to obtain the significant threshold (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.05) of outlier loci. In addition, locus-specific \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values and standard deviations (sd) across all analysed populations were calculated for SNPs that were detected to be under positive selection with the Genetix programme version 4.05 (Belkhir et al. 2004) applying the jackknife resampling procedure. We used the R package ggplot2 (Wickham 2009) to plot iHS and \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values. Allele frequency data, SNP information and ancestral/derived allele state were obtained from the Ensembl genome browser (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ensembl.org/index.html\u003c/span\u003e\u003c/span\u003e) (Howe et al. 2021). We used LDlink, a web-based application (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://analysistools.cancer.gov/LDlink/?tab=home\u003c/span\u003e\u003c/span\u003e) (Machiela and Chanock 2015), to explore population-specific linkage disequilibrium (LD); we report \u003cem\u003eD\u0026rsquo;\u003c/em\u003e and goodness-of-fit statistics (χ\u003csup\u003e2\u003c/sup\u003e statistics).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3. Estimating the timing of positive selection\u003c/h2\u003e \u003cp\u003eWe estimated the timing of selection on a beneficial allele using the R package Startmrca (Smith et al. 2018). The method applies a Markov chain Monte Carlo simulation (MCMC) that samples over the unknown ancestral haplotype to generate a sample of the posterior distribution for the time to the most recent common ancestor (TMRCA). The model takes advantage of both the length of the ancestral haplotype on each chromosome and the accumulation of derived mutations on the ancestral haplotype to generate a sample of the posterior distribution for the TMRCA. The model requires a sample (panel) containing the haplotypes with the selected allele and a reference panel of haplotypes without the selected allele. In this study, populations of European ancestry were used as samples, and populations of the other analysed genetic ancestries were used as reference panels. Because the calculated TMRCA estimates for the populations of European ancestry were very similar regardless of the reference panels used, we report in this study only the TMRCA estimates calculated for the European populations, using the European populations both as sample panel and as reference panel. We also estimated TMRCA for the East Asian-specific functional variants rs671 (\u003cem\u003eALDH2\u003c/em\u003e locus) and rs3782886 (\u003cem\u003eBRAP\u003c/em\u003e locus) (Okada et al. 2018) using the East Asian population CHB as sample and reference panel (see Supplementary Table\u0026nbsp;1 for the corresponding population names). After normalising the TMRCA data we calculated 95% credible intervals (CI\u0026thinsp;=\u0026thinsp;0.95) for the timing estimates using the equal-tailed interval method implemented in the R package bayestestR (Makowski et al. 2019). We used recombination rates from the sex-averaged recombination map from deCODE to model recombination rate variation across the human genome. We analysed 1 Mb regions up- and downstream of the genetic variants under selection with an assumed mutation rate of 1.6 x 10\u003csup\u003e\u0026minus;\u0026thinsp;8\u003c/sup\u003e. We ran three independent MCMC chains, each with 25,000 iterations. We discarded the first 9000 iterations (burn\u0026ndash;in), retaining the remaining iterations. We assumed 25 years as generation time.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4. GTEx and RegulomeDB functional data\u003c/h2\u003e \u003cp\u003eWe utilized expression quantitative trait loci (eQTLs) (accessed between May and July 2021 (dbGaP Accession phs000424.v8.p2) from GTEx Portal V8 Release (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.gtexportal.org/home/\u003c/span\u003e\u003c/span\u003e) (Ardlie et al. 2015) to test whether any of the potential SNPs that are under positive selection function as eQTL. We included \u003cem\u003ecis\u003c/em\u003e-eQTL variants within a 1 Mb window of analysed genes. The RegulomeDB database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://regulomedb.org/\u003c/span\u003e\u003c/span\u003e) (Boyle et al. 2012) was used to obtained chromatin states; this database comprises known classes of genomic elements such as promoters, enhancers, transcription start sites, and transcription factor (TF) binding motifs. Additionally, mapped phenotype data were obtained from the NHGRI-EBI GWAS catalogue (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ebi.ac.uk/gwas/\u003c/span\u003e\u003c/span\u003e) (Buniello et al. 2019) (accessed between May and July 2021).\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Positive selection in populations of European ancestry\u003c/h2\u003e \u003cp\u003eThe iHS analysis shows evidence that the human chromosomal region 12.q24.12 is under positive selection in populations of European ancestry. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e (a) plots the iHS scores in the European population GBR; Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e (b) shows the pairwise \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values for GBR \u003cem\u003evs\u003c/em\u003e. the African population LWK across 12.q24.12. The red and green lines indicate significant (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01 and after Bonferroni correction \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;1x10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e, respectively) iHS scores and the genome-wide threshold (95% confidence level) for \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e outlier loci (\u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e \u0026gt;0.3).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSNPs under positive selection at the human chromosomal region 12q24.12 in populations with European ancestry (GBR, TSI, FIN). Given are iHS scores and the calculated (-log) \u003cem\u003ep\u003c/em\u003e-values (in bold Bonferroni correction with \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;1x10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e), the timing (\u003cem\u003et\u003c/em\u003e) of positive selection on the derived beneficial allele in thousand years ago (kya) and 95% credible interval (CI) (rounded to one decimal figure), average allele frequency in % for the derived beneficial allele/ancestral allele in the different ancestries and global locus-specific \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values (sd\u0026thinsp;=\u0026thinsp;standard deviation) calculated across all analysed populations.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"14\"\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eChr\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eBeneficial allele/ ancestral allele\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLocation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"3\" nameend=\"c6\" namest=\"c4\"\u003e \u003cp\u003eiHS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eiHS\u003c/p\u003e \u003cp\u003e-logp\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e\u003cem\u003et\u003c/em\u003e (kya)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003e95%\u003c/p\u003e \u003cp\u003eCI\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colspan=\"4\" nameend=\"c13\" namest=\"c10\"\u003e \u003cp\u003eAllele frequency in %\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLocus-specific \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e (sd)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003eGBR\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003eTSI\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003eFIN\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e\u003cb\u003eAFR\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e\u003cb\u003eEUR\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e\u003cb\u003eSAS\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e\u003cb\u003eEAS\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"6\" rowspan=\"7\"\u003e \u003cp\u003e\u003cb\u003e12q24.12\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers3184504-T/C\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eExon, \u003cem\u003eSH2B3\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-3.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-3.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-2.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e3.2; 3.3; 2.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e3.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e3.2\u0026ndash;4.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e46/54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e7/93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.351 (0.060)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers4766578-T/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIntron, \u003cem\u003eATXN2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-3.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.1; 3.0; 3.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e3.0\u0026ndash;4.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e48/52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e7/93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.366 (0.063)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers10774625-A/G\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIntron, \u003cem\u003eATXN2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-3.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-3.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.1; 3.0; 2.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e3.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e2.7\u0026ndash;3.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e48/52\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e7/93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.366 (0.062)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers597808-A/G\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIntron, \u003cem\u003eATXN2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-3.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-4.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.1; 4.8; 3.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e3.0\u0026ndash;4.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e47/53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e7/93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0.2/99.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.352 (0.064)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers653178-C/T\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIntron, \u003cem\u003eATXN2\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-3.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-4.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-3.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e3.7; \u003cb\u003e5.2\u003c/b\u003e; 3.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e2.6\u0026ndash;3.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.3/99.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e47/53\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e7/93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0/100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.356 (0.063)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers847892-G/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIntron, \u003cem\u003eACAD10\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-2.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-2.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-2.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.5; 2.5; 2.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e6.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e5.1\u0026ndash;7.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.4/99.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e69/31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e30/70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e7/93\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.405 (0.080)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers2013002-T/C\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eIntron, \u003cem\u003eENST 00000546840.3\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-2.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-2.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-2.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e2.5; 1.8; 2.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e2.8\u0026ndash;3.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c10\"\u003e \u003cp\u003e0.3/99.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c11\"\u003e \u003cp\u003e41/59\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c12\"\u003e \u003cp\u003e6/94\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c13\"\u003e \u003cp\u003e0.3/99.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c14\"\u003e \u003cp\u003e0.315 (0.053)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Positive selection acts on regulatory variants of ALDH2\u003c/h2\u003e \u003cp\u003eFrom the GTEx database, we obtained in total 1591 \u003cem\u003ecis\u003c/em\u003e-QTLs that influence \u003cem\u003eALDH2\u003c/em\u003e gene expression (Supplementary Table\u0026nbsp;2); of these \u003cem\u003ecis\u003c/em\u003e-QTLs, we identified 204, 217 and 53 eQTLs that had significant (\u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.01) iHS scores in the European samples GBR, TSI and FIN, respectively (Supplementary Tables\u0026nbsp;3\u0026ndash;5). We also obtained \u003cem\u003ecis\u003c/em\u003e-eQTLs (1970 in total) for the other protein-coding genes located in this genomic region (\u003cem\u003eCUX2, FAM109A, SH2B2, ATXN2, BRAP, ACAD10, MAPKAPK5\u003c/em\u003e) (Supplementary Table\u0026nbsp;6). In contrast to the eQTLs for \u003cem\u003eALDH2\u003c/em\u003e, we did not obtain significant iHS values for these SNP eQTLs, except for SNPs that also function as eQTLs for \u003cem\u003eALDH2\u003c/em\u003e. We further identified seven SNPs (rs3184504, rs4766578, rs10774625, rs597808, rs653178, rs847892, rs2013002) that are under positive selection in European populations that have very large global locus-specific \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values\u0026thinsp;\u0026gt;\u0026thinsp;0.3, i.e. are outlier loci (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The corresponding EHH plots and pairwise \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values of these SNPs can be found in Supplementary Fig.\u0026nbsp;1 and Supplementary Table\u0026nbsp;7, respectively. The pairwise \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values for these SNPs comparing populations of European ancestries \u003cem\u003evs\u003c/em\u003e. African, East Asian and South Asian ancestries ranged from 0.253 to 0.691. These SNPs function as eQTLs for \u003cem\u003eALDH2\u003c/em\u003e, and the beneficial alleles are associated with increased \u003cem\u003eALDH2\u003c/em\u003e gene expression in various human tissues (in accordance to the set of tissues represented in GTEx) such as esophagus \u0026ndash; mucosa, skin, muscle \u0026ndash; skeletal, brain \u0026ndash; nucleus accumbens, artery \u0026ndash; tibial, artery \u0026ndash; aorta, and thyroid. The average allele frequencies for these SNPs are given in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e; for most of these SNPs the frequency of the derived beneficial alleles reaches almost 50% in the European populations. In contrast, the derived alleles are very rare (\u0026lt;\u0026thinsp;0.3%) in African and East Asian ancestries and at low frequency in populations of South Asian ancestry (\u0026lt;\u0026thinsp;7% with the exception of rs847892). We also compared the allele frequency at these loci with ancient Eurasians, including ancient hunter-gatherers (8.2\u0026ndash;7.5 kya) from the study of Mathieson et al. (2015). The allele frequency data from the latter study show for the SNPs rs3184504, rs4766578, rs10774625, and rs653178 that the ancestral alleles were fixed in ancient European hunter-gatherers. As expected, the Neanderthal and Denisovan data on the UCSC Genome Browser also show only ancestral alleles at these loci. In contrast, early European farmers (8.4\u0026ndash;4.2 kya) and individuals with steppe ancestry (5.4\u0026ndash;3.6 kya) had frequencies between 8% and 25% of the derived alleles at these loci.\u003c/p\u003e \u003cp\u003eIn the European sample GBR we estimated the timing of positive selection of the derived beneficial alleles to be from about 3.0 to 3.7 kya with the exception of SNP rs847892, which we date at 6.0 kya (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). This range of estimates are very similar to the TMRCA estimates calculated for the other European samples (TSI and FIN) (Supplementary Table\u0026nbsp;8). We also calculated the TMRCA for the derived allele of the East Asian-specific polymorphism (missense variants) at rs671-A/G and rs3782886-C/T in the East Asian population CHB, which yielded an estimation of 5.8 kya (CI: 4.8\u0026ndash;6.7) and 5.4 kya (CI: 3.3\u0026ndash;6.5), respectively.\u003c/p\u003e \u003cp\u003eWe then further included in the analysis the \u003cem\u003eALDH2\u003c/em\u003e promoter variant rs886205-A/G, which is located \u0026minus;\u0026thinsp;360 bp from the ATG start codon of the \u003cem\u003eALDH2\u003c/em\u003e gene (Chou et al. 1999). This promoter variant shows very large genetic differentiation with global locus-specific \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e = 0.378 (s.d. = 0.055). In the 1000 Genome data the derived allele A is the common allele in European and South Asian populations with average frequencies of about 83% and 71%, respectively. In contrast, in populations of African and East Asian ancestry the common allele is the ancestral allele G with frequencies of about 78% and 84%, respectively. For the \u003cem\u003eALDH2\u003c/em\u003e promoter variant, a study showed (in vivo and in vitro experiments) that the \u0026minus;\u0026thinsp;360G (ancestral) allele has a significantly lower basal transcriptional activity than the \u0026minus;\u0026thinsp;360A (derived) allele (Kimura et al. 2009). Our LD analysis revealed that the positively selected SNPs are in complete LD (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e = 1) with the \u003cem\u003eALDH2\u003c/em\u003e promoter variant rs886205 (Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). The chromatin state data from RegulomeDB showed that the identified SNPs are associated with active transcription start site (TSS), enhancers and strong transcription in different tissues (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Importantly, the positively selected SNPs rs4766578 and rs847892 are located in the binding motif for the transcription factor \u003cem\u003ehepatocyte nuclear factor 4 alpha\u003c/em\u003e (\u003cem\u003eHNF4A\u003c/em\u003e). This transcription factor is an important regulatory element of \u003cem\u003eALDH2\u003c/em\u003e. The mapped phenotypes (Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) show that the positively selected SNPs are associated with various traits and diseases, in particular with blood pressure, cardiovascular disease, cholesterol level and autoimmune diseases. The variants rs597808 and rs2013002 are also associated with alcohol drinking and physiological traits such as blood pressure. We pooled related traits (Supplementary Table\u0026nbsp;9) into four main trait category namely autoimmune diseases (AIS), blood pressure (BP), cardiovascular disease (CDS) and cancer to test the null hypothesis that the traits and the allele state are independent. We found a significant (χ\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;28.828, df\u0026thinsp;=\u0026thinsp;3, \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;=\u0026thinsp;2.4e-06) relationship between the allele state and trait; the derived beneficial alleles are positively associated with AIS, BP and CDS whereas the ancestral alleles with cancer.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003ePairwise LD (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e) of SNPs under positive selection in populations of European ancestry (GBR, TSI, FIN) and the \u003cem\u003eALDH2\u003c/em\u003e promoter (*) variant rs886205; all calculated \u003cem\u003eD\u0026rsquo;\u003c/em\u003e values with \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.0001 (χ\u003csup\u003e2\u003c/sup\u003e statistics).\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eChr:pos\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSNP\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"7\" nameend=\"c9\" namest=\"c3\"\u003e \u003cp\u003eLD (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003ers3184504\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003ers4766578\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ers10774625\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003ers597808\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003ers653178\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003ers847892\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003ers2013002\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:111884608\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers3184504\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:111904371\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers4766578\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:111910219\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers10774625\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:111973358\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers597808\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.986\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.986\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:112007756\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers653178\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.979\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.979\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.859\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:112141570\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers847892\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.846\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.846\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:112200150\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ers2013002\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.956\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.970\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.97\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003echr12:112204427\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e*rs886205\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c9\"\u003e \u003cp\u003e1.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eGTEx and RegulomeDB data on SNPs under positive selection in European populations (GBR, TSI, FIN). Given is also a summary of reported traits from the NHGRI-EBI GWAS catalogue. GTEx eQTLs\u0026ndash;eGene interaction with \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001. RegulomeDB rank: 2b: TF binding\u0026thinsp;+\u0026thinsp;any motif\u0026thinsp;+\u0026thinsp;DNase Footprint\u0026thinsp;+\u0026thinsp;DNase peak; 3a: TF binding\u0026thinsp;+\u0026thinsp;any motif\u0026thinsp;+\u0026thinsp;DNase peak; 4\u0026ndash;5: TF binding\u0026thinsp;+\u0026thinsp;DNase peak; 6: motif hit. The RegulomeDB probability score ranges from 0 to 1, with 1 being most likely to be a regulatory variant (for further details see {Boyle, 2012} and {Dong, 2019}). Transcription factor \u003cem\u003eHNF4A\u003c/em\u003e, an important regulatory element of the \u003cem\u003eALDH2\u003c/em\u003e gene expression, is given in bold.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c2\" namest=\"c1\"\u003e \u003cp\u003eGTEx\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colspan=\"4\" nameend=\"c7\" namest=\"c4\"\u003e \u003cp\u003eRegulomeDB\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eGWAS\u003c/p\u003e \u003cp\u003ereported traits\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eeQTL\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003eeGene\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003erank\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003escore\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003echromatin state\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003emotif\u003c/b\u003e\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers3184504\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, LINC01405, TMEM116\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e3a\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.67022\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003estrong transcription; enhancers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003eMTF1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eCardiovascular disease, blood pressure, ischemic stroke, glaucoma, rheumatoid arthritis, cancer, celiac disease, type I diabetes mellitus, parental longevity, inflammatory bowel disease, multiple sclerosis, blood cell count, hypothyroidism, haemoglobin measurement.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers4766578\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, TMEM116\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e2b\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.63936\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003estrong transcription; enhancers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003eESRRA, ESRRB\u003c/em\u003e, \u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eHNF4A\u003c/span\u003e,\u003c/p\u003e \u003cp\u003e\u003cem\u003eNR6A1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSj\u0026ouml;gren's syndrome, reticulocyte fraction of red cells, arthritis, vitiligo, HDL cholesterol, smoking status, coronary artery disease.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers10774625\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, ADAM1B, TMEM116\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003estrong transcription; enhancers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003eFOXJ2, FOXQ1\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eHypertension, myocardial infarction, coronary artery disease, asthma, cholesterol levels, systemic lupus erythematosus, urate measurement, life span, systolic blood pressure, hypothyroidism, glomerular filtration rate.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers597808\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, LINC01405, ADAM1B\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.13454\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003estrong transcription\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003e-\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eSystolic blood pressure, alcohol drinking, diastolic blood pressure, cholesterol levels, apolipoprotein B levels, colorectal cancer, allergic diseases, haematocrit, systemic lupus erythematosus, allergy.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers653178\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, LINC01405\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.60906\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eactive TSS; strong transcription; enhancers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003e-\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eAllergic disease, asthma, celiac disease, cholesterol level, eczema, Crohn's disease, chronic kidney disease, blood pressure, eosinophil counts, inflammatory bowel disease, type 1 diabetes, urate level.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers847892\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, TMEM116, NAA25\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.20016\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eactive TSS; strong transcription; enhancers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cspan type=\"BoldItalic\" class=\"BoldItalic\" name=\"Emphasis\"\u003eHNF4A\u003c/span\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eNo data.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ers2013002\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e\u003cem\u003eALDH2, ADAM1B\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.55195\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003eactive TSS; enhancers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e\u003cem\u003eMAFB, MAFK\u003c/em\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c8\"\u003e \u003cp\u003eAlcohol drinking and blood pressure.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e4.1 Evidence of position selection acting on regulatory variants of ALDH2\u003c/h2\u003e \u003cp\u003eThis study provides evidence of positive selection across the human chromosomal region 12q24.12. This finding is in line with two previous studies (Akbari et al. 2018; Barreiro and Quintana-Murci 2010). We further found that this genomic region is enriched in eQTLs that influence \u003cem\u003eALDH2\u003c/em\u003e gene expression. A high number of these SNP eQTLs had significant iHS scores in the populations of European ancestry. In contrast, \u003cem\u003ecis\u003c/em\u003e-eQTLs of the other genes located at chr12q24.12 showed no significant iHS values. This indicates that the target of positive selection are regulatory acting variants that influence \u003cem\u003eALDH2\u003c/em\u003e gene expression. We identified seven SNPs that are under positive selection and show very large global locus-specific \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values (\u0026gt;\u0026thinsp;0.3), indicating high genetic differentiation between populations of European ancestry and populations from other global ancestries. The GTEx data show that these SNPs function primarily as eQTLs for the \u003cem\u003eALDH2\u003c/em\u003e gene. The derived beneficial alleles at these SNP eQTLs are associated with increased expression of \u003cem\u003eALDH2\u003c/em\u003e in multiple human tissues. Importantly, the two variants rs4766578 and rs847892 are located in binding sequences for transcription factor \u003cem\u003eHNF4\u003c/em\u003e. That transcription factor is considered to be a master regulator of liver-specific gene expression (Bolotin et al. 2010) and is an important regulatory element of \u003cem\u003eALDH2\u003c/em\u003e gene expression (Stewart et al. 1998; You et al. 2002). Furthermore, these SNP eQTLs are in complete LD with the \u003cem\u003eALDH2\u003c/em\u003e promoter variant rs847892. This promoter polymorphism influences individual differences in acetaldehyde elimination. The ancestral allele G, the common allele in populations of African and East Asian ancestry, has a lower basal transcriptional activity than the derived allele A, the common allele in populations of European and South Asian ancestry (Kimura et al. 2009). These results suggest that higher transcriptional activity and increased \u003cem\u003eALDH2\u003c/em\u003e expression in individuals of European ancestry represent a form of genetic adaptation to increased alcohol consumption, possibly enabling faster detoxification of acetaldehyde.\u003c/p\u003e \u003cp\u003eThe derived beneficial alleles of these loci reach almost 50% in the European population, whereas in African and East Asian populations the frequencies are very low (\u0026lt;\u0026thinsp;0.003). The ancestral alleles at these positively selected loci appear to be fixed in ancient European hunter-gatherers, but in early farmers and individuals with steppe ancestry the frequencies of the derived alleles already range between 8% and 25% (Mathieson et al. 2015). The estimated timing of positive selection on the beneficial alleles in the European population GBR ranges from about 3.0 kya to 3.7 kya (except for rs847892 for which TMRCA was estimated to 6.0 kya). We also calculated the TMRCA for the East Asian-specific derived alleles rs671-A and rs3782886-C, yielding an estimation of 5.8 kya (CI: 4.8\u0026ndash;6.7) and 5.4 kya (CI: 4.3\u0026ndash;6.5), respectively. Rs3782886, which is in LD with rs671, shows signals of very recent selection for the past 2000\u0026ndash;3000 years in the Japanese population as reported in a recent study (Okada et al. 2018). Noteworthy, rs671-A and rs1229984-A (\u003cem\u003eADH1B\u003c/em\u003e locus) were found in a subsequent study to be significantly associated with better survival in the Japanese population (Sakaue et al. 2020). The estimated TMRCA for the derived alleles in our study suggests that these alleles spread in East Asia at a much earlier time than the beneficial alleles in populations of European ancestry. Archaeological evidence indicates early production of fermented alcohol in China (McGovern et al. 2004). Analysis of starch granules, phytoliths and fungi in food residues adhering to 8000\u0026ndash;7000 year-old alcohol-making pottery vessels suggests that, in East Asia in the early Neolithic, alcoholic beverages were already being produced (Liu et al. 2019). For Europe, archaeologically recognizable brewing material in Central European lakeside settlements show that alcoholic beverages were being produced in this region in the late Neolithic period about 6000 years ago (Heiss et al. 2020). Later, in Greek-Roman antiquity, a richly developed viticulture with high wine production was achieved and, in this period, wine became part of the daily diet of many people (Retief and Cilliers 2015). Alcohol consumption has apparently increased steadily since then in Europe, especially in the 19th century. In Germany, for example, the high level of consumption, in particular of strong spirits, in the early 19th century was \u0026ndash; in analogy to the plague \u0026ndash; referred to as Branntweinpest (brandy plague). Since the rs671-A allele leads to an inactive enzyme and thus to an excess of toxic acetaldehyde in cells with its negative physiological effects, we suggest that this allele may explain the differences in the signature of positive selection between populations of European and East Asian ancestry.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e4.2 Wider biomedical context\u003c/h2\u003e \u003cp\u003eThe ALDH2 enzyme plays a critical role both in the detoxification of both acetaldehyde and ROS-generated aldehyde adducts such as 4-hydroxy-2-nonenal and malondialdehyde. This enzyme thus has cytoprotective effects reducing oxidative stress (Budas et al. 2009; Guo et al. 2013). In particular, the \u003cem\u003eALDH2*2\u003c/em\u003e variant (rs671), which is common only in individuals of East Asian ancestry, has been intensively studied in East Asians. While individuals with the \u003cem\u003eALDH2*2\u003c/em\u003e allele have a reduced risk of developing alcoholism, it increases their cancer risk (Crabb et al. 2004; Zhang and Fu 2021). Nevertheless, this allele was found to be associated in the Japanese population with better survival (Sakaue et al. 2020). In European populations this allele is virtually absent. In our study, however, the identified variants that are under recent positive selection in European populations act as regulatory variants and are associated with increased \u003cem\u003eADLH2\u003c/em\u003e gene expression in various human tissues. This suggests that individuals carrying these beneficial alleles should be more quickly able to detoxify the body from higher amounts of acetaldehyde and ROS-generated aldehyde adducts. However, a recent study reports higher methylation in alcohol-dependent patients compared to controls in the \u003cem\u003eALDH2\u003c/em\u003e promoter region (Pathak et al. 2017). Furthermore, that study suggests that positive and negative regulatory elements interact at the \u003cem\u003eALDH2\u003c/em\u003e promoter to induce genotype-mediated epigenetic changes, leading to differential transcriptional activity of this gene. We therefore suggest that individuals carrying the beneficial alleles may be able to consume more alcohol (over longer time periods), but may also have a higher likelihood of becoming heavy drinkers and alcohol dependent. This, then, could lead to increased methylation of the \u003cem\u003eALDH2\u003c/em\u003e promoter, resulting in decreased \u003cem\u003eALDH2\u003c/em\u003e gene expression. Accordingly, the protective effects of \u003cem\u003eALDH2\u003c/em\u003e against oxidative damage through acetaldehyde would be lost, resulting in increased risk of numerous oxidative stress-related diseases such as cancer, diabetes, inflammatory disorders and cardiovascular conditions such as hypertension and stroke. Indeed, we found in this study that the derived beneficial alleles are positively associated with AIS, BP and CDS whereas the ancestral alleles with cancer.\u003c/p\u003e \u003c/div\u003e"},{"header":"5. Conclusions","content":"\u003cp\u003eWe found that positive selection acts on regulatory variants affecting \u003cem\u003eALDH2\u003c/em\u003e gene expression in populations of European ancestry. In contrast to the known functional consequence of the \u003cem\u003eALDH2*2\u003c/em\u003e variant (rs671) in East Asians, which is associated with alcohol intolerance, in Europeans the beneficial derived alleles are associated with increased \u003cem\u003eALDH2\u003c/em\u003e gene expression. This suggests local adaptation to higher alcohol consumption in Europeans. Estimation of the timing of positive selection on the beneficial alleles suggests that these variants were recently adapted, approximately 3000 to 3700 years ago. We hypothesize that the beneficial effects of higher \u003cem\u003eALDH2\u003c/em\u003e expression leads to an increased detoxification capacity for acetaldehyde, but possibly also to increased likelihood of chronic alcohol abuse, leading to decreased \u003cem\u003eALDH2\u003c/em\u003e expression and thus increased cell toxicity from EtOH-derived acetaldehyde as well as from ROS-generated aldehydes.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflicts of interest:\u003c/strong\u003e All authors declare that the research was conducted without commercial or financial relationships and therefore no potential conflicts of interest occurred.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and material\u003c/strong\u003e:\u0026nbsp;The gnomic data can be obtained from 1000 Genomes database. The generated iHS and Fst dataset are available from the corresponding author on reasonable request.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCode availability:\u003c/strong\u003e Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability Statement:\u003c/strong\u003e Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eHS conceived and designed this study; HS, TG and DM performed statistical analyses. HS wrote the draft manuscript; all authors contributed to the results, edited, read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent to participate:\u0026nbsp;\u003c/strong\u003eNot applicable. The 1000 Genomes data are publicly available.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication:\u0026nbsp;\u003c/strong\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgments:\u003c/strong\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWe thank Michael Stachowitsch from the Department of Evolutionary Anthropology, University of Vienna and Franz Suchentrunk from the University of Veterinary Medicine Vienna for valuable comments on the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003cp\u003eAguet F, Brown AA, Castel SE, Davis JR, He Y, Jo B, Mohammadi P, Park Y, Parsana P, Segre AV, Strober BJ, Zappala Z, Cummings BB, Gelfand ET, Hadley K, Huang KH, Lek M, Li X, Nedzel JL, Nguyen DY, Noble MS, Sullivan TJ, Tukiainen T, MacArthur DG, Getz G, Management NP, Addington A, Guan P, Koester S, Little AR, Lockhart NC, Moore HM, Rao A, Struewing JP, Volpi S, Collection B, Brigham LE, Hasz R, Hunter M, Johns C, Johnson M, Kopen G, Leinweber WF, Lonsdale JT, McDonald A, Mestichelli B, Myer K, Roe B, Salvatore M, Shad S, Thomas JA, Walters G, Washington M, Wheeler J, Bridge J, Foster BA, Gillard BM, Karasik E, Kumar R, Miklos M, Moser MT, Jewell SD, Montroy RG, Rohrer DC, Valley D, Mash DC, Davis DA, Sobin L, Barcus ME, Branton PA, Grp EMW, Abell NS, Balliu B, Delaneau O, Fresard L, Gamazon ER, Garrido-Martin D, Gewirtz ADH, Gliner G, Gloudemans MJ, Han B, He AZ, Hormozdiari F, Liu B, Kang EY, McDowell IC, Ongen H, Palowitch JJ, Peterson CB, Quon G, Ripke S, Saha A, Shabalin AA, Shimko TC, Sul JH, Teran NA, Tsang EK, Zhang H, Zhou YH, Bustamante CD, et al. (2017) Genetic effects on gene expression across human tissues. Nature 550: 204-+.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAkbari A, Vitti JJ, Iranmehr A, Bakhtiari M, Sabeti PC, Mirarab S, Bafna V (2018) Identifying the favored mutation in a positive selective sweep. Nature Methods 15: 279-+.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAltshuler DM, Gibbs RA, Peltonen L, Dermitzakis E, Schaffner SF, Yu FL, Bonnen PE, de Bakker PIW, Deloukas P, Gabriel SB, Gwilliam R, Hunt S, Inouye M, Jia XM, Palotie A, Parkin M, Whittaker P, Chang K, Hawes A, Lewis LR, Ren YR, Wheeler D, Muzny DM, Barnes C, Darvishi K, Hurles M, Korn JM, Kristiansson K, Lee C, McCarroll SA, Nemesh J, Keinan A, Montgomery SB, Pollack S, Price AL, Soranzo N, Gonzaga-Jauregui C, Anttila V, Brodeur W, Daly MJ, Leslie S, McVean G, Moutsianas L, Nguyen H, Zhang QR, Ghori MJR, McGinnis R, McLaren W, Takeuchi F, Grossman SR, Shlyakhter I, Hostetter EB, Sabeti PC, Adebamowo CA, Foster MW, Gordon DR, Licinio J, Manca MC, Marshall PA, Matsuda I, Ngare D, Wang VO, Reddy D, Rotimi CN, Royal CD, Sharp RR, Zeng CQ, Brooks LD, McEwen JE, Int HapMap C (2010) Integrating common and rare genetic variation in diverse human populations. Nature 467: 52-58.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eArdlie KG, DeLuca DS, Segre AV, Sullivan TJ, Young TR, Gelfand ET, Trowbridge CA, Maller JB, Tukiainen T, Lek M, Ward LD, Kheradpour P, Iriarte B, Meng Y, Palmer CD, Esko T, Winckler W, Hirschhorn JN, Kellis M, MacArthur DG, Getz G, Shabalin AA, Li G, Zhou YH, Nobel AB, Rusyn I, Wright FA, Lappalainen T, Ferreira PG, Ongen H, Rivas MA, Battle A, Mostafavi S, Monlong J, Sammeth M, Mele M, Reverter F, Goldmann JM, Koller D, Guigo R, McCarthy MI, Dermitzakis ET, Gamazon ER, Im HK, Konkashbaev A, Nicolae DL, Cox NJ, Flutre T, Wen XQ, Stephens M, Pritchard JK, Tu ZD, Zhang B, Huang T, Long Q, Lin L, Yang JL, Zhu J, Liu J, Brown A, Mestichelli B, Tidwell D, Lo E, Salvatore M, Shad S, Thomas JA, Lonsdale JT, Moser MT, Gillard BM, Karasik E, Ramsey K, Choi C, Foster BA, Syron J, Fleming J, Magazine H, Hasz R, Walters GD, Bridge JP, Miklos M, Sullivan S, Barker LK, Traino HM, Mosavel M, Siminoff LA, Valley DR, Rohrer DC, Jewell SD, Branton PA, Sobin LH, Barcus M, Qi LQ, McLean J, Hariharan P, Um KS, Wu SP, Tabor D, Shive C, Smith AM, Buia SA, et al. (2015) The Genotype-Tissue Expression (GTEx) pilot analysis: Multitissue gene regulation in humans. Science 348: 648-660.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAuburger G, Gispert S, Lahut S, Omur O, Damrath E, Heck M, Basak N (2014) 12q24 locus association with type 1 diabetes: SH2B3 or ATXN2 ? World Journal of Diabetes 5: 316-327.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAuton A, Abecasis GR, Altshuler DM, Durbin RM, Bentley DR, Chakravarti A, Clark AG, Donnelly P, Eichler EE, Flicek P, Gabriel SB, Gibbs RA, Green ED, Hurles ME, Knoppers BM, Korbel JO, Lander ES, Lee C, Lehrach H, Mardis ER, Marth GT, McVean GA, Nickerson DA, Schmidt JP, Sherry ST, Wang J, Wilson RK, Boerwinkle E, Doddapaneni H, Han Y, Korchina V, Kovar C, Lee S, Muzny D, Reid JG, Zhu Y, Chang Y, Feng Q, Fang X, Guo X, Jian M, Jiang H, Jin X, Lan T, Li G, Li J, Li Y, Liu S, Liu X, Lu Y, Ma X, Tang M, Wang B, Wang G, Wu H, Wu R, Xu X, Yin Y, Zhang D, Zhang W, Zhao J, Zhao M, Zheng X, Gupta N, Gharani N, Toji LH, Gerry NP, Resch AM, Barker J, Clarke L, Gil L, Hunt SE, Kelman G, Kulesha E, Leinonen R, McLaren WM, Radhakrishnan R, Roa A, Smirnov D, Smith RE, Streeter I, Thormann A, Toneva I, Vaughan B, Zheng-Bradley X, Grocock R, Humphray S, James T, Kingsbury Z, Sudbrak R, Albrecht MW, Amstislavskiy VS, Borodina TA, Lienhard M, Mertes F, Sultan M, Timmermann B, Yaspo M-L, Fulton L, Fulton R, et al. (2015) A global reference for human genetic variation. Nature 526: 68-74.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBarbhaiya M, Costenbader KH (2016) Environmental exposures and the development of systemic lupus erythematosus. Current Opinion in Rheumatology 28: 497-505.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBarreiro LB, Quintana-Murci L (2010) From evolutionary genetics to human immunology: how selection shapes host defence genes. Nature Reviews Genetics 11: 17-30.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBelkhir K, Borsa P, Chikhi L, Raufaste N, Bonhomme F (2004) GENETIX4. 05, logiciel sous Windows TM pour la g\u0026eacute;n\u0026eacute;tiquedes populations. Laboratoire g\u0026eacute;nome, populations, interactions, CNRS UMR 5000: 1996-2004.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBentham J, Morris DL, Graham DSC, Pinder CL, Tombleson P, Behrens TW, Martin J, Fairfax BP, Knight JC, Chen L, Replogle J, Syvanen A-C, Ronnblom L, Graham RR, Wither JE, Rioux JD, Alarcon-Riquelme ME, Vyse TJ (2015) Genetic association analyses implicate aberrant regulation of innate and adaptive immunity genes in the pathogenesis of systemic lupus erythematosus. Nature Genetics 47: 1457-+.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBierut LJ, Goate AM, Breslau N, Johnson EO, Bertelsen S, Fox L, Agrawal A, Bucholz KK, Grucza R, Hesselbrock V, Kramer J, Kuperman S, Nurnberger J, Porjesz B, Saccone NL, Schuckit M, Tischfield J, Wang JC, Foroud T, Rice JP, Edenberg HJ (2012) ADH1B is associated with alcohol dependence and alcohol consumption in populations of European and African ancestry. Molecular Psychiatry 17: 445-450.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBolotin E, Liao HL, Ta TC, Yang CH, Hwang-Verslues W, Evans JR, Jiang T, Sladek FM (2010) Integrated Approach for the Identification of Human Hepatocyte Nuclear Factor 4 alpha Target Genes Using Protein Binding Microarrays. Hepatology 51: 642-653.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBoyle AP, Hong EL, Hariharan M, Cheng Y, Schaub MA, Kasowski M, Karczewski KJ, Park J, Hitz BC, Weng S, Cherry JM, Snyder M (2012) Annotation of functional variation in personal genomes using RegulomeDB. Genome Research 22: 1790-1797.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBudas GR, Disatnik MH, Mochly-Rosen D (2009) Aldehyde Dehydrogenase 2 in Cardiac Protection: A New Therapeutic Target? Trends in Cardiovascular Medicine 19: 158-164.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBuniello A, MacArthur JAL, Cerezo M, Harris LW, Hayhurst J, Malangone C, McMahon A, Morales J, Mountjoy E, Sollis E, Suveges D, Vrousgou O, Whetzel PL, Amode R, Guillen JA, Riat HS, Trevanion SJ, Hall P, Junkins H, Flicek P, Burdett T, Hindorff LA, Cunningham F, Parkinson H (2019) The NHGRI-EBI GWAS Catalog of published genome-wide association studies, targeted arrays and summary statistics 2019. Nucleic Acids Research 47: D1005-D1012.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCederbaum AI (2012) Alcohol Metabolism. Clinics in Liver Disease 16: 667-+.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCeni E, Mello T, Galli A (2014) Pathogenesis of alcoholic liver disease: Role of oxidative metabolism. World Journal of Gastroenterology 20: 17756-17772.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eChang CC, Chow CC, Tellier L, Vattikuti S, Purcell SM, Lee JJ (2015) Second-generation PLINK: rising to the challenge of larger and richer datasets. Gigascience 4.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eChen CH, Ferreira JCB, Gross ER, Mochly-Rosen D (2014) TARGETING ALDEHYDE DEHYDROGENASE 2: NEW THERAPEUTIC OPPORTUNITIES. Physiological Reviews 94: 1-34.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eChou WY, Stewart MJ, Carr LG, Zheng D, Stewart TR, Williams A, Pinaire J, Crabb DW (1999) An A/G polymorphism in the promoter of mitochondrial aldehyde dehydrogenase (ALDH2): Effects of the sequence variant on transcription factor binding and promoter strength. Alcoholism-Clinical and Experimental Research 23: 963-968.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCoenen MJH, Trynka G, Heskamp S, Franke B, van Diemen CC, Smolonska J, van Leeuwen M, Brouwer E, Boezen MH, Postma DS, Platteel M, Zanen P, Lammers J, Groen HJM, Mali W, Mulder CJ, Tack GJ, Verbeek WHM, Wolters VM, Houwen RHJ, Mearin ML, van Heel DA, Radstake T, van Riel P, Wijmenga C, Barrera P, Zhernakova A (2009) Common and different genetic background for rheumatoid arthritis and coeliac disease. Human Molecular Genetics 18: 4195-4203.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCrabb DW, Matsumoto M, Chang D, You M (2004) Overview of the role of alcohol dehydrogenase and aldehyde dehydrogenase and their variants in the genesis of alcohol-related pathology. Proceedings of the Nutrition Society 63: 49-63.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eCraddock N, Hurles ME, Cardin N, Pearson RD, Plagnol V, Robson S, Vukcevic D, Barnes C, Conrad DF, Giannoulatou E, Holmes C, Marchini JL, Stirrups K, Tobin MD, Wain LV, Yau C, Aerts J, Ahmad T, Andrews TD, Arbury H, Attwood A, Auton A, Ball SG, Balmforth AJ, Barrett JC, Barroso I, Barton A, Bennett AJ, Bhaskar S, Blaszczyk K, Bowes J, Brand OJ, Braund PS, Bredin F, Breen G, Brown MJ, Bruce IN, Bull J, Burren OS, Burton J, Byrnes J, Caesar S, Clee CM, Coffey AJ, Connell JMC, Cooper JD, Dominiczak AF, Downes K, Drummond HE, Dudakia D, Dunham A, Ebbs B, Eccles D, Edkins S, Edwards C, Elliot A, Emery P, Evans DM, Evans G, Eyre S, Farmer A, Ferrier IN, Feuk L, Fitzgerald T, Flynn E, Forbes A, Forty L, Franklyn JA, Freathy RM, Gibbs P, Gilbert P, Gokumen O, Gordon-Smith K, Gray E, Green E, Groves CJ, Grozeva D, Gwilliam R, Hall A, Hammond N, Hardy M, Harrison P, Hassanali N, Hebaishi H, Hines S, Hinks A, Hitman GA, Hocking L, Howard E, Howard P, Howson JMM, Hughes D, Hunt S, Isaacs JD, Jain M, Jewell DP, Johnson T, Jolley JD, Jones IR, Jones LA, et al. (2010) Genome-wide association study of CNVs in 16,000 cases of eight common diseases and 3,000 shared controls. Nature 464: 713-U86.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDanecek P, Auton A, Abecasis G, Albers CA, Banks E, DePristo MA, Handsaker RE, Lunter G, Marth GT, Sherry ST, McVean G, Durbin R, Genomes Project Anal G (2011) The variant call format and VCFtools. Bioinformatics 27: 2156-2158.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eEdenberg HJ, McClintick JN (2018) Alcohol Dehydrogenases, Aldehyde Dehydrogenases, and Alcohol Use Disorders: A Critical Review. Alcoholism-Clinical and Experimental Research 42: 2281-2297.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eEng MY, Luczak SE, Wall TL (2007) ALDH2, ADH1B, and ADH1C genotypes in Asians: A literature review. Alcohol Research \u0026amp; Health 30: 22-27.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eFernandez-Sola J (2015) Cardiovascular risks and benefits of moderate and heavy alcohol consumption. Nature Reviews Cardiology 12: 576-587.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGalinsky KJ, Bhatia G, Loh PR, Georgiev S, Mukherjee S, Patterson NJ, Price AL (2016) Fast Principal-Component Analysis Reveals Convergent Evolution of ADH1B in Europe and East Asia. American Journal of Human Genetics 98: 456-472.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGelernter J, Kranzler HR, Sherva R, Almasy L, Koesterer R, Smith AH, Anton R, Preuss UW, Ridinger M, Rujescu D, Wodarz N, Zill P, Zhao H, Farrer LA (2014) Genome-wide association study of alcohol dependence: significant findings in African-and European-Americans including novel risk loci. Molecular Psychiatry 19: 41-49.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGuo JM, Liu AJ, Zang P, Dong WZ, Ying L, Wang W, Xu P, Song XR, Cai J, Zhang SQ, Duan JL, Mehta JL, Su DF (2013) ALDH2 protects against stroke by clearing 4-HNE. Cell Research 23: 915-930.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGu S, Li H, Pakstis AJ, Speed WC, Gurwitz D, Kidd JR, Kidd KK (2018) Recent Selection on a Class I ADH Locus Distinguishes Southwest Asian Populations Including Ashkenazi Jews. Genes 9.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHan Y, Gu S, Oota H, Osier MV, Pakstis AJ, Speed WC, Kidd JR, Kidd KK (2007) Evidence of positive selection on a class I ADH locus. American Journal of Human Genetics 80: 441-456.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHeiss AG, Azorin MB, Antolin F, Kubiak-Martens L, Marinova E, Arendt EK, Biliaderis CG, Kretschmer H, Lazaridou A, Stika HP, Zarnkow M, Baba M, Bleicher N, Cialowicz KM, Chlodnicki M, Matuschik I, Schlichtherle H, Valamoti SM (2020) Mashes to Mashes, Crust to Crust. Presenting a novel microstructural marker for malting in the archaeological record. Plos One 15.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHowe KL, Achuthan P, Allen J, Alvarez-Jarreta J, Amode MR, Armean IM, Azov AG, Bennett R, Bhai J, Billis K, Boddu S, Charkhchi M, Cummins C, Fioretto LR, Davidson C, Dodiya K, El Houdaigui B, Fatima R, Gall A, Giron CG, Grego T, Guijarro-Clarke C, Haggerty L, Hemrom A, Hourlier T, Izuogu OG, Juettemann T, Kaikala V, Kay M, Lavidas I, Le T, Lemos D, Martinez JG, Marugan JC, Maurel T, McMahon AC, Mohanan S, Moore B, Muffato M, Oheh DN, Paraschas D, Parker A, Parton A, Prosovetskaia I, Sakthivel MP, Salam AIA, Schmitt BM, Schuilenburg H, Sheppard N, Steed E, Szpak M, Szuba M, Taylor K, Thormann A, Threadgold G, Walts B, Winterbottom A, Chakiachvili M, Chaubal A, De Silva N, Flint B, Frankish A, Hunt SE, Iisley GR, Langridge N, Loveland JE, Martin FJ, Mudge JM, Morales J, Perry E, Ruffier M, Tate J, Thybert D, Trevanion SJ, Cunningham F, Yates AD, Zerbino DR, Flicek P (2021) Ensembl 2021. Nucleic Acids Research 49: D884-D891.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eJohnson KE, Voight BF (2018) Patterns of shared signatures of recent positive selection across human populations. Nature Ecology \u0026amp; Evolution 2: 713-720.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eJorgenson E, Thai KK, Hoffmann TJ, Sakoda LC, Kvale MN, Banda Y, Schaefer C, Risch N, Mertens J, Weisner C, Choquet H (2017) Genetic contributors to variation in alcohol consumption vary by race/ethnicity in a large multi-ethnic genome-wide association study. Molecular Psychiatry 22: 1359-1367.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eKent WJ, Sugnet CW, Furey TS, Roskin KM, Pringle TH, Zahler AM, Haussler D (2002) The human genome browser at UCSC. Genome Research 12: 996-1006.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eKimura Y, Nishimura FT, Abe S, Fukunaga T, Tanii H, Saijoh K (2009) A Promoter Polymorphism in the ALDH2 Gene Affects Its Basal and Acetaldehyde/Ethanol-Induced Gene Expression in Human Peripheral Blood Leukocytes and HepG2 Cells. Alcohol and Alcoholism 44: 261-266.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLee DJ, Lee HM, Kim JH, Park IS, Rho YS (2017) Heavy alcohol drinking downregulates ALDH2 gene expression but heavy smoking up-regulates SOD2 gene expression in head and neck squamous cell carcinoma. World Journal of Surgical Oncology 15.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLi H, Borinskaya S, Yoshimura K, Kal\u0026apos;ina N, Marusin A, Stepanov VA, Qin ZD, Khaliq S, Lee MY, Yang YJ, Mohyuddin A, Gurwitz D, Mehdi SQ, Rogaev E, Jin L, Yankovsky NK, Kidd JR, Kidd KK (2009) Refined Geographic Distribution of the Oriental ALDH2*504Lys (nee 487Lys) Variant. Annals of Human Genetics 73: 335-345.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLi H, Mukherjee N, Soundararajan U, Tarnok Z, Barta C, Khaliq S, Mohyuddin A, Kajuna SLB, Mehdi SQ, Kidd JR, Kidd KK (2007) Geographically separate increases in the frequency of the derived ADH1B*47His allele in eastern and western Asia. American Journal of Human Genetics 81: 842-846.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLiu L, Wang JJ, Levin MJ, Sinnott-Armstrong N, Zhao H, Zhao YA, Shao J, Di N, Zhang TE (2019) The origins of specialized pottery and diverse alcohol fermentation techniques in Early Neolithic China. Proceedings of the National Academy of Sciences of the United States of America 116: 12767-12774.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLiu L, Wang JJ, Rosenberg D, Zhao H, Lengyel G, Nadel D (2018) Fermented beverage and food storage in 13,000 y-old stone mortars at Raqefet Cave, Israel: Investigating Natufian ritual feasting. Journal of Archaeological Science-Reports 21: 783-793.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLiu X, Hu C, Bao MH, Li J, Liu XY, Tan XR, Zhou Y, Chen YQ, Wu SL, Chen SH, Zhang R, Jiang F, Jia WP, Wang XY, Yang XC, Cai J (2016) Genome Wide Association Study Identifies L3MBTL4 as a Novel Susceptibility Gene for Hypertension. Scientific Reports 6.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLu B, Solomon DH, Costenbader KH, Karlson EW (2014) Alcohol Consumption and Risk of Incident Rheumatoid Arthritis in Women A Prospective Study. Arthritis \u0026amp; Rheumatology 66: 1998-2005.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eLuca F, Perry GH, Di Rienzo A (2010) Evolutionary Adaptations to Dietary Changes. Annual Review of Nutrition, Vol 30 30: 291-314.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMacgregor S, Lind PA, Bucholz KK, Hansell NK, Madden PAF, Richter MM, Montgomery GW, Martin NG, Heath AC, Whitfield JB (2009) Associations of ADH and ALDH2 gene variation with self report alcohol reactions, consumption and dependence: an integrated analysis. Human Molecular Genetics 18: 580-593.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMachiela MJ, Chanock SJ (2015) LDlink: a web-based application for exploring population-specific haplotype structure and linking correlated alleles of possible functional variants. Bioinformatics 31: 3555-3557.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMa H, Guo R, Yu L, Zhang YM, Ren J (2011) Aldehyde dehydrogenase 2 (ALDH2) rescues myocardial ischaemia/reperfusion injury: role of autophagy paradox and toxic aldehyde. European Heart Journal 32: 1025-1038.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMakowski D, Ben-Shachar MS, Chen SHA, Ludecke D (2019) Indices of Effect Existence and Significance in the Bayesian Framework. Frontiers in Psychology 10.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMathieson I, Lazaridis I, Rohland N, Mallick S, Patterson N, Roodenberg SA, Harney E, Stewardson K, Fernandes D, Novak M, Sirak K, Gamba C, Jones ER, Llamas B, Dryomov S, Pickrell J, Arsuaga JL, de Castro JMB, Carbonell E, Gerritsen F, Khokhlov A, Kuznetsov P, Lozano M, Meller H, Mochalov O, Moiseyev V, Guerra MAR, Roodenberg J, Verges JM, Krause J, Cooper A, Alt KW, Brown D, Anthony D, Lalueza-Fox C, Haak W, Pinhasi R, Reich D (2015) Genome-wide patterns of selection in 230 ancient Eurasians. Nature 528: 499-+.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMcGovern PE, Zhang JH, Tang JG, Zhang ZQ, Hall GR, Moreau RA, Nunez A, Butrym ED, Richards MP, Wang CS, Cheng GS, Zhao ZJ (2004) Fermented beverages of pre- and proto-historic China. Proceedings of the National Academy of Sciences of the United States of America 101: 17593-17598.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMcLaren W, Gil L, Hunt SE, Riat HS, Ritchie GRS, Thormann A, Flicek P, Cunningham F (2016) The Ensembl Variant Effect Predictor. Genome Biology 17.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOberreiter V, Goellner T, Morris LD, Schaschl H (2021) Positive Selection Affects the Expression of Systemic Lupus Erythematosus Associated Loci in Human Populations. BMC Genomics.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOkada Y, Momozawa Y, Sakaue S, Kanai M, Ishigaki K, Akiyama M, Kishikawa T, Arai Y, Sasaki T, Kosaki K, Suematsu M, Matsuda K, Yamamoto K, Kubo M, Hirose N, Kamatani Y (2018) Deep whole-genome sequencing reveals recent selection signatures linked to evolution and disease risk of Japanese. Nature Communications 9.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOota H, Pakstis AJ, Bonne-Tamir B, Goldman D, Grigorenko E, Kajuna SLB, Karoma NJ, Kungulilo S, Lu RB, Odunsi K, Okonofua F, Zhukova OV, Kidd JR, Kidd KK (2004) The evolution and population genetics of the ALDH2 locus: random genetic drift, selection, and low levels of recombination. Annals of Human Genetics 68: 93-109.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWorld Health Organization\u0026nbsp;(2018) Global status report on alcohol and health 2018. Geneva\u003c/p\u003e\n\u003cp\u003ePan C, Xing JH, Zhang C, Zhang YM, Zhang LT, Wei SJ, Zhang MX, Wang XP, Yuan QH, Xue L, Wang JL, Cui ZQ, Zhang Y, Xu F, Chen YG (2016) Aldehyde dehydrogenase 2 inhibits inflammatory response and regulates atherosclerotic plaque. Oncotarget 7: 35562-35576.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePanisello-Rosello A, Lopez A, Folch-Puy E, Carbonell T, Rolo A, Palmeira C, Adam R, Net M, Rosello-Catafau J (2018) Role of aldehyde dehydrogenase 2 in ischemia reperfusion injury: An update. World Journal of Gastroenterology 24: 2984-2994.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePathak H, Frieling H, Bleich S, Glahn A, Heberlein A, Nassab MH, Hillemacher T, Burkert A, Rhein M (2017) Promoter Polymorphism rs886205 Genotype Interacts With DNA Methylation of the ALDH2 Regulatory Region in Alcohol Dependence. Alcohol and Alcoholism 52: 269-276.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePeng Y, Shi H, Qi XB, Xiao CJ, Zhong H, Ma RLZ, Su B (2010) The ADH1B Arg47His polymorphism in East Asian populations and expansion of rice domestication in history. Bmc Evolutionary Biology 10.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003ePeter BM, Huerta-Sanchez E, Nielsen R (2012) Distinguishing between Selective Sweeps from Standing Variation and from a De Novo Mutation. Plos Genetics 8.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eQuillen EE, Chen XD, Almasy L, Yang F, He H, Li X, Wang XY, Liu TQ, Hao W, Deng HW, Kranzler HR, Gelernter J (2014) ALDH2 Is Associated to Alcohol Dependence and Is the Major Genetic Determinant of \u0026quot;Daily Maximum Drinks\u0026quot; in a GWAS Study of an Isolated Rural Chinese Sample. American Journal of Medical Genetics Part B-Neuropsychiatric Genetics 165: 103-110.\u003c/p\u003e\n\u003cp\u003eR Core Team (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https://www.R-project.org/.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRetief F, Cilliers L Wine in Graeco-Roman Antiquity with Emphasis on Its Effect on Health 2015\u003c/p\u003e\n\u003cp\u003eRicci E, Al Beitawi S, Cipriani S, Candiani M, Chiaffarino F, Vigano P, Noli S, Parazzini F (2017) Semen quality and alcohol intake: a systematic review and meta-analysis. Reproductive Biomedicine Online 34: 38-47.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRocco A, Compare D, Angrisani D, Zamparelli MS, Nardone G (2014) Alcoholic disease: Liver and beyond. World Journal of Gastroenterology 20: 14652-14659.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRodriguez-Zavala JS, Calleja LF, Moreno-Sanchez R, Yoval-Sanchez B (2019) Role of Aldehyde Dehydrogenases in Physiopathological Processes. Chemical Research in Toxicology 32: 405-420.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRoerecke M, Vafaei A, Hasan OSM, Chrystoja BR, Cruz M, Lee R, Neuman MG, Rehm J (2019) Alcohol Consumption and Risk of Liver Cirrhosis: A Systematic Review and Meta-Analysis. American Journal of Gastroenterology 114: 1574-1586.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRonksley PE, Brien SE, Turner BJ, Mukamal KJ, Ghali WA (2011) Association of alcohol consumption with selected cardiovascular disease outcomes: a systematic review and meta-analysis. Bmj-British Medical Journal 342.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSakaue S, Akiyama M, Hirata M, Matsuda K, Murakami Y, Kubo M, Kamatani Y, Okada Y (2020) Functional variants in ADH1B and ALDH2 are non-additively associated with all-cause mortality in Japanese population. European Journal of Human Genetics 28: 378-382.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSmith J, Coop G, Stephens M, Novembre J (2018) Estimating Time to the Common Ancestor for a Beneficial Allele. Molecular Biology and Evolution 35: 1003-1017.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eStewart MJ, Dipple KM, Estonius M, Nakshatri H, Everett LM, Crabb DW (1998) Binding and activation of the human aldehyde dehydrogenase 2 promoter by hepatocyte nuclear factor 4. Biochimica Et Biophysica Acta-Gene Structure and Expression 1399: 181-186.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSzpiech ZA, Hernandez RD (2014) selscan: An Efficient Multithreaded Program to Perform EHH-Based Scans for Positive Selection. Molecular Biology and Evolution 31: 2824-2827.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTeam RC (2021) R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. URL https://www.R-project.org/. Vienna, Austria\u003c/p\u003e\n\u003cp\u003eThompson A, Cook J, Choquet H, Jorgenson E, Yin J, Kinnunen T, Barclay J, Morris AP, Pirmohamed M (2020) Functional validity, role, and implications of heavy alcohol consumption genetic loci. Science Advances 6.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eToth R, Pocsai Z, Fiatal S, Szeles G, Kardos L, Petrovski B, McKee M, Adany R (2010) ADH1B*2 allele is protective against alcoholism but not chronic liver disease in the Hungarian population. Addiction 105: 891-896.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eValente C, Alvarez L, Marks SJ, Lopez-Parra AM, Parson W, Oosthuizen O, Oosthuizen E, Amorim A, Capelli C, Arroyo-Pardo E, Gusmao L, Prata MJ (2015) Exploring the relationship between lifestyles, diets and genetic adaptations in humans. Bmc Genetics 16.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eVoight BF, Kudaravalli S, Wen XQ, Pritchard JK (2006) A map of recent positive selection in the human genome (vol 4, pg 154, 2006). Plos Biology 4: 659-659.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWay M, McQuillin A, Saini J, Ruparelia K, Lydall GJ, Guerrini I, Ball D, Smith I, Quadri G, Thomson AD, Kasiakogia-Worlley K, Cherian R, Gunwardena P, Rao H, Kottalgi G, Patel S, Hillman A, Douglas E, Qureshi SY, Reynolds G, Jauhar S, O\u0026apos;Kane A, Dedman A, Sharp S, Kandaswamy R, Dar K, Curtis D, Morgan MY, Gurling HMD (2015) Genetic variants in or near ADH1B and ADH1C affect susceptibility to alcohol dependence in a British and Irish population. Addiction Biology 20: 594-604.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWeir BS, Cockerham CC (1984) ESTIMATING F-STATISTICS FOR THE ANALYSIS OF POPULATION-STRUCTURE. Evolution 38: 1358-1370.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWickham H (2009) ggplot2: Elegant Graphics for Data Analysis. Ggplot2: Elegant Graphics for Data Analysis: 1-212.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eWild PS, Felix JF, Schillert A, Teumer A, Chen MH, Leening MJG, Volker U, Grossmann V, Brody JA, Irvin MR, Shah SJ, Pramana S, Lieb W, Schmidt R, Stanton AV, Malzahn D, Smith AV, Sundstrom J, Minelli C, Ruggiero D, Lyytikainen LP, Tiller D, Smith JG, Monnereau C, Di Tullio MR, Musani SK, Morrison AC, Pers TH, Morley M, Kleber ME, Aragam J, Benjamin EJ, Bis JC, Bisping E, Broeckel U, Cheng S, Deckers JW, Del Greco MF, Edelmann F, Fornage M, Franke L, Friedrich N, Harris TB, Hofer E, Hofman A, Huang J, Hughes AD, Kahonen M, Kruppa J, Lackner KJ, Lannfelt L, Laskowski R, Launer LJ, Leosdottir M, Lin HH, Lindgren CM, Loley C, MacRae CA, Mascalzoni D, Mayet J, Medenwald D, Morris AP, Muller C, Muller-Nurasyid M, Nappo S, Nilsson PM, Nuding S, Nutile T, Peters A, Pfeufer A, Pietzner D, Pramstaller PP, Raitakari OT, Rice KM, Rivadeneira F, Rotter JI, Ruohonen ST, Sacco RL, Samdarshi TE, Schmidt H, Sharp ASP, Shields DC, Sorice R, Sotoodehnia N, Stricker BH, Surendran P, Thom S, Toglhofer AM, Uitterlinden AG, Wachter R, Volzke H, Ziegler A, Munzel T, Marz W, Cappola TP, Hirschhorn JN, Mitchell GF, Smith NL, Fox ER, Dueker ND, et al. (2017) Large-scale genome-wide analysis identifies genetic variants associated with cardiac structure and function. Journal of Clinical Investigation 127: 1798-1812.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eXu K, Kranzler HR, Sherva R, Sartor CE, Almasy L, Koesterer R, Zhao HY, Farrer LA, Gelernter J (2015) Genomewide Association Study for Maximum Number of Alcoholic Drinks in European Americans and African Americans. Alcoholism-Clinical and Experimental Research 39: 1137-1147.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eYasukochi Y, Sakuma J, Takeuchi I, Kato K, Oguri M, Fujimaki T, Horibe H, Yamada Y (2017) Longitudinal exome-wide association study to identify genetic susceptibility loci for hypertension in a Japanese population. Experimental and Molecular Medicine 49.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eYe KX, Gu ZL (2011) Recent Advances in Understanding the Role of Nutrition in Human Genome Evolution. Advances in Nutrition 2: 486-496.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eYou M, Fischer M, Cho WK, Crabb D (2002) Transcriptional control of the human aldehyde dehydrogenase 2 promoter by hepatocyte nuclear factor 4: Inhibition by cyclic AMP and COUP transcription factors. Archives of Biochemistry and Biophysics 398: 79-86.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eZhang H, Fu LW (2021) The role of ALDH2 in tumorigenesis and tumor progression: Targeting ALDH2 as a potential cancer treatment. Acta Pharmaceutica Sinica B 11: 1400-1411.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eZhang H, Gong DX, Zhang YJ, Li SJ, Hu SS (2012) Effect of mitochondrial aldehyde dehydrogenase-2 genotype on cardioprotection in patients with congenital heart disease. European Heart Journal 33: 1606-1614.\u0026nbsp;\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"ALDH2, alcohol metabolism, eQTLs, positive selection, genetic adaptation","lastPublishedDoi":"10.21203/rs.3.rs-891422/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-891422/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eALDH2 is a key enzyme in alcohol metabolism that protects cells from acetaldehyde toxicity. Using iHS and \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e statistics, we identified regulatory acting variants affecting \u003cem\u003eALDH2\u003c/em\u003e gene expression under positive selection in populations of European ancestry. Several SNPs (rs3184504, rs4766578, rs10774625, rs597808, rs653178, rs847892, rs2013002) that function as eQTLs for \u003cem\u003eALDH2\u003c/em\u003e in various tissues showed evidence of positive selection. Very large pairwise \u003cem\u003eF\u003c/em\u003e\u003csub\u003e\u003cem\u003eST\u003c/em\u003e\u003c/sub\u003e values indicated high genetic differentiation at these loci between populations of European ancestry and populations of other global ancestries. Estimating the timing of positive selection on the beneficial alleles suggests that these variants were recently adapted approximately 3000 to 3700 years ago. The derived beneficial alleles are in complete linkage disequilibrium with the derived \u003cem\u003eALDH2\u003c/em\u003e promoter variant rs886205, which is associated with higher transcriptional activity. The SNPs rs4766578 and rs847892 are located in binding sequences for the transcription factor \u003cem\u003eHNF4A\u003c/em\u003e, which is an important regulatory element of \u003cem\u003eALDH2\u003c/em\u003e gene expression. In contrast to the missense variant \u003cem\u003eALDH2\u003c/em\u003e rs671 (\u003cem\u003eALDH2*2\u003c/em\u003e), which is common only in East Asian populations and is associated with greatly reduced enzyme activity and alcohol intolerance, the beneficial alleles of the regulatory variants identified in this study are associated with increased expression of \u003cem\u003eALDH2\u003c/em\u003e. This suggests adaptation of Europeans to higher alcohol consumption.\u003c/p\u003e","manuscriptTitle":"Positive selection affects the human ADLH2 gene expression: genetic adaptation to alcohol consumption","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-09-13 15:06:02","doi":"10.21203/rs.3.rs-891422/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"2f0a6c4b-3f9b-4f54-b9c9-92dae6208fed","owner":[],"postedDate":"September 13th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":7139532,"name":"Molecular Genetics"},{"id":7139533,"name":"Evolutionary Biology"},{"id":7139534,"name":"Anthropology"}],"tags":[],"updatedAt":"2021-10-15T21:46:15+00:00","versionOfRecord":[],"versionCreatedAt":"2021-09-13 15:06:02","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-891422","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-891422","identity":"rs-891422","version":["v1"]},"buildId":"7rjqhiLT3MXkJMwkYKINL","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00