Aberrant One-Carbon Metabolism and Ancestral Genetics Underlie Edematous Severe Acute Malnutrition

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Severe acute malnutrition (SAM) contributes to the death of millions of children under age five annually. SAM is clinically classified as non-edematous SAM (NESAM) or the more severe edematous SAM (ESAM), which is more common in east-central Africa and the Caribbean. The reason some children develop ESAM while others develop NESAM remains unclear; however, recent studies have identified aberrant one-carbon metabolism (OCM) in ESAM relative to NESAM. Here, we assess genetic variants at 103 loci known to influence OCM, and determine their association with ESAM in 711 samples from Jamaica and Malawi. Seven OCM loci showed evidence of association across both populations, including five associated with homocysteine and folate metabolism (MTHFR, AHCYL1, PRICKLE2, GABBR2, and PLD2). Three SNPs in PLD2, PRICKLE2, and GABBR2, genotyped using cell-free DNA from serum metabolomic samples, supported causal effects on ESAM risk through homocysteine-related metabolites. Cumulatively, OCM-related variants showed more association with ESAM than expected by chance (z = 3.06), with differing effect magnitudes in the two populations. By leveraging chromosome-level patterns of intracontinental African admixture, we demonstrate that OCM variant associations with ESAM occur on a shared east-African ancestral genetic background. Finally, using whole genome sequence data from eight African populations, we demonstrate that several OCM loci have outlier signatures of selection in multiple populations, including the ESAM-associated PLD2 locus. These findings strengthen support for aberrant OCM in ESAM pathogenesis, with implications for current interventions, and highlight the potential of cell-free DNA, intra-continental admixture, and population genetics in mapping disease risk.
Full text 206,977 characters · extracted from preprint-html · click to expand
Aberrant One-Carbon Metabolism and Ancestral Genetics Underlie Edematous Severe Acute Malnutrition | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Aberrant One-Carbon Metabolism and Ancestral Genetics Underlie Edematous Severe Acute Malnutrition Neil Hanchard, Natasha Lie, Yixing Han, Qing Li, Aarti Jajoo, and 22 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6890799/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Severe acute malnutrition (SAM) contributes to the death of millions of children under age five annually. SAM is clinically classified as non-edematous SAM (NESAM) or the more severe edematous SAM (ESAM), which is more common in east-central Africa and the Caribbean. The reason some children develop ESAM while others develop NESAM remains unclear; however, recent studies have identified aberrant one-carbon metabolism (OCM) in ESAM relative to NESAM. Here, we assess genetic variants at 103 loci known to influence OCM, and determine their association with ESAM in 711 samples from Jamaica and Malawi. Seven OCM loci showed evidence of association across both populations, including five associated with homocysteine and folate metabolism ( MTHFR , AHCYL1 , PRICKLE2 , GABBR2 , and PLD2 ). Three SNPs in PLD2 , PRICKLE2 , and GABBR2 , genotyped using cell-free DNA from serum metabolomic samples, supported causal effects on ESAM risk through homocysteine-related metabolites. Cumulatively, OCM-related variants showed more association with ESAM than expected by chance (z = 3.06), with differing effect magnitudes in the two populations. By leveraging chromosome-level patterns of intracontinental African admixture, we demonstrate that OCM variant associations with ESAM occur on a shared east-African ancestral genetic background. Finally, using whole genome sequence data from eight African populations, we demonstrate that several OCM loci have outlier signatures of selection in multiple populations, including the ESAM-associated PLD2 locus. These findings strengthen support for aberrant OCM in ESAM pathogenesis, with implications for current interventions, and highlight the potential of cell-free DNA, intra-continental admixture, and population genetics in mapping disease risk. Health sciences/Medical research/Genetics research Health sciences/Diseases/Nutrition disorders/Malnutrition kwashiorkor marasmus genomics 1-carbon metabolism admixture local ancestry Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 INTRODUCTION Severe acute malnutrition (SAM) affects 16.9 million children worldwide, predominantly those under five years old, and either directly or indirectly contributes to more than one million childhood deaths per year 1 , 2 . SAM is clinically characterized by a weight-for-height that is more than three standard deviations below the median, a mid-upper-arm circumference of less than 115 mm, or the presence of nutritional edema 3 . In practice, SAM is recognized to occur in two clinically distinct forms: edematous SAM (ESAM), which includes the syndromes of kwashiorkor and marasmic-kwashiorkor, and non-edematous SAM (NESAM), also known as marasmus 4 . ESAM accounts for a higher proportion of hospitalized SAM in east- and central- Africa and the Caribbean 5 , 6 , and is associated with mortality rates of 10–20% among those who are hospitalized with SAM and medical complications 7 – 9 . ESAM is clinically characterized by bilateral pitting edema (swelling) of the extremities and severe systemic involvement that may include skin and hair changes, fatty liver, and multi-system organ failure 4 , 10 , 11 . By contrast, NESAM is more prevalent in Southeast Asia and northern- and western-African geographies, and is characterized by generalized wasting of body tissues, typically with less severe systemic involvement 4 , 6 . Whilst it is well-recognized that the development of SAM is largely linked to food security and availability, the persistence of adverse geopolitical factors mean that SAM is still highly prevalent globally. ESAM was first described in medical literature in the early part of the 20th century 12 – 14 ; however, the reasons why a severely malnourished child will develop ESAM as opposed to NESAM remain largely unknown 15 . In the past, it was hypothesized that diet or protein intake were primary contributors to ESAM risk; however, decades of epidemiological surveys have failed to show consistent evidence for differences in protein intake or any other single dietary deprivation between children who develop ESAM and those who develop NESAM 16 – 19 . Similarly, there have been no consistently demonstrated differences in infectious or environmental exposures between children who develop ESAM and those who develop NESAM 4 , 20 – 22 . There are, however, numerous pathophysiological differences between the two clinical manifestations, particularly in their metabolism and biochemistry; for example, ESAM has been associated with a slower rate of protein breakdown in starvation, impaired heparin sulfate proteoglycan expression, greater dietary cysteine efficiency, and increased oxidative stress, compared to children with NESAM 23 – 28 . Despite this, it remains unknown whether these observations represent downstream effects that result from the disorder or are primary etiological differences. Perhaps as a result, ESAM and NESAM are still treated in the same manner, using the same nutritional supplements, despite well described differences in pathophysiology and underlying biochemistry. Among biochemical differences between children with ESAM and those with NESAM, the cellular movement of methyl groups known as one-carbon metabolism (OCM) has been a particular recent focus. OCM plays a major role in health and disease, providing substrates for mitochondrial activity and ATP generation, dNTPs for DNA repair, and methyl groups needed for the maintenance of DNA and other cellular methylation events in mitotically active cells 29 . The broad metabolic reach of OCM has been hypothesized as a potential unifying factor for the multisystem dysfunction seen in ESAM, and recent studies lend support to this. For instance, metabolomic studies have demonstrated that children with ESAM have significantly lower serum levels of the essential amino acid methionine 30 with concomitant related effects, including dysregulation of asymmetric dimethyl arginine (ADMA), lower cysteine levels, impaired cysteine synthesis, reduced acylcarnitine production, and impaired fatty acid transport 25 , 31 . Mice fed a maize-vegetable diet similar to that seen among SAM children in Malawi, develop a fatty liver phenotype that is similar to that seen in ESAM, which is abrogated by supplementation of the methyl-group donor choline 32 . Similarly, although there are no overt differences in microbiome composition between NESAM and ESAM children, a metagenomic study of severe malnutrition, conducted without regard to subtype, reported significantly lower OCM metabolite levels among gnotobiotic mice transplanted with stool from children with ESAM 33 . Observations of OCM dysfunction in ESAM are also consistent with stable isotope whole-body flux studies reporting slower methionine flux among children with ESAM 26 . We have previously shown that during the acute nutritional stress, children with ESAM have significantly lower buccal cell DNA methylation than children with NESAM 34 , and this was consistent with observations of methionine turnover. Notably, and consistent with turnover studies conducted in children who had recovered from SAM, methylation differences were not observed among adults who had recovered from having either ESAM or NESAM as children, confirming that OCM dysfunction is a feature of the acute nutritional insult. Significantly hypomethylated loci were associated with genes both directly and indirectly implicated in the ESAM phenotype, and with loci associated with disordered nutrition states (e.g., obesity) and body anthropometry. Thus, collectively, there is a growing body of evidence supporting OCM dysfunction in ESAM. These findings, however, do not directly explain why only some children have the OCM derangements associated with ESAM, whilst others from the same region and environment have the OCM profile associated with NESAM. Although there are no formal heritability estimates for ESAM, children with repeated bouts of SAM are more likely to develop the same form of SAM (i.e., either ESAM or NESAM) upon readmission to hospital 35 , 36 ; this observation, alongside the well-documented lack of demonstrable environmental, infectious, or metagenomic differences between ESAM and NESAM, has focused attention on the potential genetic risk of ESAM, which is consistent with recent interest in the genetics of malnutrition 37 . To date, however, attempts at genetic association have been limited to three single-candidate SNP association studies conducted in sample sizes of less than 200 individuals decades ago 38 – 40 . We thus sought to take a modern genetic approach to understanding the risk of developing ESAM, leveraging advances that have characterized genetic risk association studies of other public health disorders, including nutritional disorders like obesity. OCM is a fundamental metabolic pathway that has been genetically characterized 41 , 42 , with well-curated genes and loci mediating interindividual metabolic variation. Therefore, we started by evaluating the potential impact of genetic variation at OCM loci on the risk of developing ESAM. In this study, we compile a comprehensive list of OCM-associated genetic loci and evaluate genetic association with ESAM at these loci in population samples from Jamaica and Malawi (Fig. 1 ). We then leverage population genetics and both local and global population ancestry, to further explore the association between OCM and ESAM. MATERIALS AND METHODS Study participants Details of participant samples used in the study have been previously published 34 , 43 . All recruitment consisted of samples collected after individuals had received a diagnosis of either ESAM or NESAM. Samples from Jamaica included DNA from blood and buccal samples from children (< 18 years) diagnosed with SAM. DNA was extracted using a phenol/chloroform protocol as described previously 43 . SAM subtype was determined using the Wellcome Classification 44 , where NESAM (marasmus) is defined as < 60% weight-for-age without edema, and ESAM includes marasmic-kwashiorkor- defined as < 60% weight-for-age with edema, and kwashiorkor - defined as 60–80% weight-for-age with edema. Pediatric participants were recruited prospectively from the Tropical Metabolism Research Unit of the Caribbean Institute for Health Research at the University Hospital of the West Indies (UHWI), the major tertiary nutritional center on the island. Recruitment was part of a larger, long-running study of the genetic etiology of SAM. A cohort of adult participants (> 18 years) who formerly had SAM as children were recruited from the same site. The study was approved by the Ethics Committees of the UHWI/University of the West Indies Faculty of Medical Sciences. In Malawi, samples from children with SAM were recruited from 18 rural sites across five districts 34 . Samples were originally collected between 2013–2016. The sub-type of SAM was defined using mid-upper arm circumference (MUAC) and the presence or absence of edema. Buccal samples were collected using the Oragene Discover (OGR-250) DNA collection kit (DNA Genotek Inc., Ottawa, Ontario, Canada). The protocol was modified to collect buccal epithelial cells by swabbing the inside of each cheek ten times. DNA was extracted following the manufacturer's instructions. After collection, to maintain clinical uniformity in SAM designation between Jamaica and Malawi, samples from Malawi were re-classified using the Wellcome Classification as having either ESAM or NESAM, which was highly concordant with the MUAC based designation. The study was approved by the National Health Science Review Committee of the Ministry of Health, Government of Malawi. Written informed consent was obtained from the parents or adult guardians of children included in this study. Permission to use de-identified participant samples for genetic studies at Baylor College of Medicine (BCM) was approved by the Institutional Review Board (IRB) of BCM and confirmed by the National Institutes of Health IRB subsequently. Curation of OCM loci GeneWeaver 45 , AmiGO 2 46,47 , GWAS Catalog 48 , 49 , KEGG 50 , NCBI BioSystems 51 , Pathway Commons 52 , and PubMed were queried using the search terms “one-carbon metabolism”, “cysteine”, “methionine”, and “one-carbon pool by folate” to identify OCM loci to be used in this study. At the genome level, a 50 kilobase (kb) buffer region was included on either side of each identified gene to capture potential cis -acting elements. This search resulted in a list of 103 genes involved in OCM. Overlapping regions were consolidated into single regions, resulting in 95 loci ( Supplementary Table S1 ). Genotyping and quality control Peripheral blood and buccal DNA samples from Jamaican and Malawian individuals were genotyped for this study. Downstream quality control and analyses are shown in Supplementary Figure S1 . Genotyping was performed on the Infinium H3Africa Consortium Array v1.1 ( web resources ), which provides LD coverage of ~ 0.89 in the African superpopulation of 1000 Genomes. Raw idat files were processed in GenomeStudio (Illumina, San Diego, California, USA), including within-sample clustering, to create binary input files for PLINK. Genotype fluorescence projections in GenomeStudio were visually inspected for those SNPs that surpassed test-wide significance in order to confirm genotype calls. Downstream SNP and sample quality control (QC) was performed in PLINK v1.07 53 and included limiting to biallelic autosomal SNPs and excluding SNPs with a missingness > 0.95; removal of SNPs outside the Hardy-Weinberg test distribution; and excluding samples with missingness > 10%. Linkage disequilibrium (LD) pruning was conducted using a window size of 50 bases, a sliding window of ten bases, and r 2 threshold of 0.1. Pairs of samples with a proportion-of- identity-by-descent (IBD) (PI_HAT) score > 0.1 were flagged, and the individual within the pair with the higher missingness rate was removed. Post QC, 1.8 million genotyped SNPs were available for downstream analysis across the 711 remaining samples. Imputation Given the relative lack of representation of African genomes in existing imputation panels, we chose to conduct a two-stage imputation and data merge process. Data from Jamaican individuals (n = 340) were imputed using the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) reference panel 54 , which was generated using data from individuals with predominantly west-African ancestry, including Jamaicans. We used minimac3 v1.0.5 (University of Michigan, Ann Arbor, Michigan, USA) 55 and Eagle v2.3 (Broad Institute, Cambridge, Massachusetts, USA) 56 on the Michigan Imputation Server (University of Michigan, Ann Arbor, Michigan, USA) 55 to perform phasing and imputation on our Jamaican samples, selecting the African reference panel ("AFR") as the reference population. Data from Malawian participants (n = 371) was imputed using the H3Africa reference panel (H3ABioNet) 57 ( web resources ), which is designed to represent a wide array of African ancestries, including east- and central-African groups that neighbor Malawi. Phasing was performed using Eagle 2.4 56 (Broad Institute, Cambridge, MA, USA). Minimac4 Version 1.0.0 (University of Michigan, Ann Arbor, MI, USA) was used for imputation on the H3ABioNet Nextflow chip imputation pipeline 58 , which implements a similar workflow to the Michigan Imputation Service pipeline, and was deployed on the University of Cape Town High-Performance Computing Facility (web resources). After imputation, SNPs with r 2 < 0.3 were excluded. The imputed data from both population groups was merged using an in-house command line script ( web resources ), prior to removing non-autosomal SNPs and SNPs with minor allele frequency (MAF) < 0.05. The final imputed dataset included 8,415,377 SNPs, of which 45,411 SNPs were at OCM loci (+/-50kb of genes). Multidimensional scaling Genome-wide SNPs were used to perform multidimensional scaling (MDS) in PLINK v1.07, and these were used for two different purposes (see Supplementary Figure S1 ): 1. - on the QC-ed genotype dataset used for association analyses; 2. on the genotype-only dataset alongside datasets from 1000 Genomes 59 to put the primary populations in the context of African datasets (See Population ancestry and admixture mapping below). 1000 Genomes datasets consisted of Esan individuals in Nigeria (ESN, n = 99), individuals from Gambia in Western Division, the Gambia (GWD, n = 113), Luhya individuals in Webuye, Kenya (LWK, n = 99), Mende individuals in Sierra Leone (MSL, n = 85), and Yoruba individuals in Ibadan, Nigeria (YRI, n = 108). The genotyped-only dataset was merged with the above 1000 Genomes populations, and LD-pruned with a window size of 50 bases, a window of 10 bases, and r 2 < 0.1. MDS was performed PLINK and visualized using an in-house R 60 script ( Supplementary Figure S2 ). Genetic association analyses Association statistics were primarily derived using Genome-Wide Efficient Mixed Model Association (GEMMA) 61 , which better accounts for genetic ancestry variation. Mixed model association analysis in GEMMA was performed using the -lm 4 flag. Standardized relatedness of the combined genotyped and imputed dataset was also calculated using the flag -gk 2. We used the genetic data to estimate the standardized kinship matrix before applying the generalized linear mixed effect model to assess genetic association between SNPs and the risk of developing ESAM, assuming random effect on individuals and fixed effects due to sex and the first component of MDS (as suggested by scree plots of the proportion of variance explained; Supplementary Figure S3 ). In the regional regression analyses, the MDS component was estimated from the dataset of the specified cohort. QQ plots were generated using the R package qqman 62 . Manhattan plots were generated using an in-house script ( web resources ). Cumulative association We employed a permutation test procedure to assess the cumulative association between OCM loci and SAM status; this is a non-parametric method of assessing the statistical significance of observed data without making assumptions about the underlying probability distribution of the data. We hypothesized that OCM loci are enriched for SNPs with p-values below a certain significance threshold, i.e., the proportion of significant loci within OCM loci is more than expected by chance. We choose p = 0.001 as a representative significance threshold for the number of loci being tested (similar to a Bonferroni correction for multiple testing of 0.1/100). To test the robustness of our findings, we included a ceiling threshold of p = 0.01, and two equidistance points (0.001 and 0.01) as sensitivity points. The test procedure involved taking the genome-wide association statistics for all genotyped SNPs (n = 1,843,153) and calculating the proportion of SNPs with a p-value meeting a pre-specified threshold. We then randomly sampled a subset of SNPs equivalent to the number of genotyped SNPs across OCM loci (n = 9,479) and determined the proportion of those SNPs with a p-value meeting the predefined threshold. That step was repeated 10,000 times to derive an empirical distribution of the proportion of significant SNPs. We then calculated a z-score for the observed proportion based on the empirical distribution. To account for the randomness of genome-wide SNPs relative to a restricted metabolic pathway, we performed the same test using a similar number of SNPs sampled from loci involved in ketone metabolism. Haplotype association and calculation of linkage disequilibrium Genotyped SNPs in GABBR2 (n = 5) and PRICKLE2 (n = 6) reaching region-level significance (Fig. 2) were arranged in 5' to 3' order and phased in PLINK v1.07 to create basic region-level haplotypes. PLINK was also used to determine association using the --hap-assoc flag. PLINK files were also parsed to Haploview 63 in order to generate plots of linkage disequilibrium. Population ancestry and admixture mapping To investigate the ancestral African genetic ancestries in the Jamaica and Malawi datasets, genome-wide genotype data was merged with selected population data from publicly available datasets using PLINK v1.9 ( Supplementary Figure S1 ). Available datasets included the African Genome Variation Project (AGVP) 64 , H3Africa Consortium 65 , Schlebusch et al. 66 , and 1000 Genomes Project (Phase III). Only SNPs common to all datasets were retained. As genetic reference data is not available for every African ethnolinguistic group, we sought to ensure that we could assess contributions from Niger-Congo speaking groups, which includes Bantu speakers - the largest ethnolinguistic grouping across the continent, as well as non-Niger-Congo groups, such Rain Forest Forager, Hunter-Gatherer (e.g. Khoe and San), and Nilo-Saharan groups, which are more populous in east- and central- Africa. Populations representing Niger-Congo African ancestry included Bantu speakers from Zambia (BSZ, n = 39), Jola (n = 79), and YRI (n = 100). Populations representing non-Niger-Congo African ancestry included Amhara (n = 42), Botswana (BOT, n = 48), LWK (n = 74), Khoe and San Hunter Gatherers (n = 18), Zulu (n = 95), and MSL (n = 85). Samples from CEU (n = 95) were included to represent European admixture. The merged dataset was further filtered to remove SNPs with excessive missingness (> 0.1) and then LD-pruned (see MDS above). The final merged genotype dataset included ~ 50,000 autosomal SNPs. The resulting merged dataset was used to perform MDS for a subset of groups using the same parameters and procedures as above. Using the unsupervised clustering algorithm as implemented by ADMIXTURE v1.3 67 , global ancestry proportions were inferred for all populations. Since the algorithm is known to be impacted by the number of individuals that represent a population, 100 random individuals each from Jamaica and Malawi were retained to approximate the sample sizes of the reference populations. ADMIXTURE was performed for K = 2 to K = 6, and the K with the lowest cross-validation error estimates was noted (K = 5). Local ancestry inference Supplementary Figure S1 B illustrates the data processes for local ancestry. Data subset from 1000 Genomes consisting of CEU, MSL, and LWK individuals was used as references for European, West African, and East African ancestry, respectively. The SAM cohort dataset and the 1000 Genomes dataset each underwent independent QC filtering using MAF > 0.05; SNP missingness threshold of 0.05; and removal of duplicate SNPs. The remaining SNPs were merged using an in-house script. The merged dataset was subject to additional QC filtering to remove A > T and G > C SNPs. Datasets were then phased by chromosome using ShapeIt2 68 , using the 1000 Genomes Phase 3 genetic map with default parameters. Local ancestry inference scores were calculated for the genotyped-only dataset using RFMix version 2 69 . Local ancestry for the reference 1000 Genomes dataset was run separately from the SAM cohort dataset to maximize the available data generated on differing platforms, and 1000 Genomes-specific indices were used as reference ancestry groups. The 1000 Genomes Phase 3 genetic map was used for RFMix, which was otherwise run with default parameters. Ancestry-adjusted association analyses In ancestry-adjusted association tests, we employed TRACTOR v1.1.0 70 . SNP association with ESAM disease risk was performed whilst accounting for the ancestral origin of the SNP by first assigning ancestry labels to the two haplotypes of individuals, referred as tracts. Tracts were extracted from the genotyped-only dataset using ExtractTracts.py from TRACTOR, with integers from one to three referring to the three proxy ancestries - CEU, LWK, and MSL - from 1000 Genomes. RunTractor.py was then run to perform logistic regression using NESAM vs ESAM status as the outcome for each individual, with tracts, ancestry-specific allele counts, and sex as regressors. OCM loci were then subset from the regression results and visualized 62 . Cell-free DNA genotyping of serum metabolomics samples Cell-free DNA (cfDNA) was obtained from serum samples of 415 individuals from Malawi. These individuals were independently recruited as part of our published metabolomics study of one-carbon metabolites in SAM 31 . Samples were initially stored at -80°C and subsequently slowly brought up to room temperature to minimize degradation. CfDNA was extracted using the Qiagen QIAamp MinElute Virus Spin Kit (Cat #57704), which is optimized for cfDNA isolation, and yielded a mean concentration of 2.7ng/µL of genomic DNA per sample, with concentrations ranging from 0.1 ng/µL to 52.0 ng/µL. We attempted to amplify six ESAM-associated SNPs from cfDNA samples – two in PRICKLE2 (rs753562 and rs17664202), two in GABBR2 (rs17664203 and rs17664204) and one each in SHMT1 (rs651495) and PLD2 (rs1052748). We were unable to obtain consistent or robust genotype data for rs651495 ( SHMT1 ) and rs17664203 ( GABBR2 ) and results from these assays were not included in downstream analyses. Primers for the remaining SNPs were designed using an in silico primer design tool ( https://www.ncbi.nlm.nih.gov/tools/primer-blast/ ) using the SNP ID (rsID) in GRCh38/hg38 ( Supplementary Table S2 ). Primer specificity was confirmed using BLAT. Primers were synthesized by IDT (Coralville, IA, USA) and prepared as 100 µM stock solutions in TE buffer. PCR amplification was performed using 4ng of DNA under a single thermocycling and amplification protocol ( Supplementary Methods ), with products verified via gel electrophoresis (GelDoc, BioRad laboratories, Hercules, CA, USA). Following PCR, products were purified using Exo-Sap IT (Thermo Fisher Scientific, Waltham MA, USA) to remove excess primers and dNTPs. Dideoxy- (Sanger) sequencing was conducted using M13-tagged primers and the Big Dye v3.1 sequencing kit (Thermo Fisher Scientific, Waltham MA, USA). Sequencing reactions were cleaned using the Qiagen DyeEx kit (Qiagen, Germantown, MD, USA) to remove dye terminators, followed by genotype calling on the SeqStudio (Genetic Analyzer #A33770, Fisher Scientific, Waltham MA, USA). Sequencher (software version 4.10.1) was used to visualize and confirm genotype calls. Mendelian randomization (MR) using serum metabolites Following the same classification criterion used in our primary cohort, we selected 252 samples with diagnoses of ESAM (n = 135) and NESAM (n = 117) for whom we had cfDNA genotypes and metabolomic data for all 16 OCM metabolites - choline, betaine, dimethylglycine (DMG), glycine, sarcosine, 5-methyltetrahydrofolate (MTHF), serine, methionine, S-adenosylmethionine (SAMe), S-adenosylhomocysteine (SAH), homocysteine, cysteine, cystathionine, pyridoxal phosphate (PLP) and asymmetric dimethylglycine (ADMA). The remaining samples were either community controls or had a diagnosis of moderate acute malnutrition. Mendelian randomization (MR) was used to assess the causal effect of individual metabolites using the genotypes as the instrument variables. For this analysis, we performed two-sample MR based on multiple SNPs by combining results from the SNP-metabolite and the SNP-outcome associations to show potential in vivo effects. We applied one instance of the inverse-variance approach (IVW) from MendelianRandomization R package. Under this approach, the causal effect is obtained from a weighted linear regression of the associations with the outcome, on the associations with the metabolite, while fixing the intercept to zero and weights being the inverse-variances of the associations with the outcome 71 . We also exploited the MR-Egger method to detect invalid instrument variables and bias 72 . The serum metabolites were highly correlated (R2). Our cumulative association data suggested that there might be synergistic effects among SNPs; therefore, we assumed an additive genetic model for test alleles, and then tested for the causal effect on ESAM risk of each metabolite, after log10 transformation (to better approximate normality). We selected three genotyped SNPs as instrument variables, one SNP from each locus, with SNP rs753562 in gene PRICKLE2 being omitted because the serum genotype data was slightly out of Hardy-Weinberg equilibrium and the SNP is physically close to the other PRICKLE2 SNP. The estimates of SNP-outcome association are drawn from the GEMMA analysis of the Malawi group in the primary cohort with fixed effects being sex and the first component of MDS. The estimates of SNP-metabolite association are drawn from a linear regression of log 10 (metabolite) with sex and age as covariates. Whole genome sequencing A subset of 151 DNA samples from Malawi were chosen to undergo whole-genome sequencing to assess recent positive selection in OCM-related genomic regions. These samples were selected to have an approximately equal distribution of males and females and an equal number of ESAM and NESAM samples to minimize sampling bias. WGS was performed by first generating PCR-free libraries from one microgram genomic DNA using the TruSeq® DNA PCR-Free HT Sample Preparation Kit (Illumina, San Diego, CA, USA). The median insert sizes were approximately 400 bp. Libraries were tagged with unique dual-index DNA barcodes to allow the pooling of libraries and to minimize the impact of barcode hopping. Libraries were then pooled for sequencing on the NovaSeq 6000 (Illumina) to obtain at least 300 million 151-base read pairs per individual library. The sequence data were then aligned to human_g1k_v37_decoy.fasta via BWA-MEMv0.7.15 73 . GATK Resource Bundle B37 74 was used to sort the data, mark duplicates, perform BQSR, and run the GATK Haplotype Caller to generate gVCF files on each sample. Then individual gVCFs were merged into a multi-sample VCF file using GATK 4.3.0.0. GATK Best Practice 75 for variant calling and filtration was followed to obtain high-quality variant data, i.e., applying GATK VQSR function, followed by harder filters (filter parameters set as DP > = 10, QUAL > = 60, and ExcessHet < = 54.69). Signatures of selection We used the iHS score 76 to identify genomic signatures of selection in the dataset; this analysis is agnostic of phenotype and covariates. To compare signatures of selection at OCM loci between different countries, we created subsets of country-specific signatures of selection generated by the H3Africa consortium 65 , which included selection statistics for SNPs across the OCM loci used in the association study: Benin (n = 45,537 SNP scores), Botswana (BOT, n = 44,277 SNP scores), Cameroon (CAM, n = 44,458 SNP scores), Mali (MAL, n = 43,033 SNP scores), Berom from Nigeria (BRN, n = 44,177 SNP scores), Gur speakers from West Africa (GWR, n = 43,523 SNP scores), and Zambia (ZAM, n = 47,572 SNP scores). These samples were sequenced on Illumina HiSeq 2500 and subsequently processed and mapped to GRCh37. Integrated haplotype scores (iHS) were calculated using selscan 77 . We calculated iHS scores for the Malawi WGS dataset using the same QC and run parameters as the H3Africa datasets. SNPs with iHS scores > |2| were considered outliers, potentially under selection. Selection signatures were visualized using R 60 and ggplot2 78 . RESULTS Our primary analyses utilized DNA samples from 833 individuals from Jamaica and Malawi approximately evenly split between ESAM (case/affected group) and NESAM (reference/control group) at each location (Fig. 1 ; Supplementary Table S3 ). Samples were genotyped on the Illumina H3Africa array, with subsequent quality control (QC) processing resulting in 711 samples ( Supplementary Figure S1 ) and 1,639,325 genotyped SNPs ( Methods ). Among genotyped SNPs, 9,479 SNPs fell within 10kb (+/-) of a curated list of 103 autosomal genes and loci previously associated with OCM ( Methods; Supplementary Table S4 ); this included four gene clusters with two or more genes in tandem. After variant imputation and quality control ( Methods ), a set of 45,411 imputed and 9,479 genotyped SNPs at 95 OCM loci were identified to test for association and perform downstream analyses ( Supplementary Figure S1 ). Variation at OCM loci is associated with ESAM We first assessed the association between imputed SNPs at OCM loci and ESAM using the linear mixed model implemented in GEMMA (see Methods ). The top associated SNP (rs79824961 near GABBR2 ) surpassed our multiple-testing adjustment for the total number of SNPs interrogated (p < 9.73x10 − 7 ) (Fig. 2). A further 56 SNPs at nine OCM loci, surpassed a secondary significance threshold adjusted for the number of loci (n = 95) tested (p < 5.26x10 − 4 ). This latter threshold was supported by an excess of observed p-values without evidence of test-wide inflation (λ = 1.02) on the resulting QQ plot (Fig. 2). To further reduce potential false positives, we characterized candidate loci as those with multiple SNPs surpassing our locus-level cut-off or single variants with evidence of association using a secondary method (n = 3 loci; Methods) . This resulted in a final set of seven ESAM-associated OCM loci ( Supplementary Table S4 ). The strongest associated locus was an intragenic region on chromosome 9q22.33 ( Supplementary Figure S4A ) falling within the first intron of gamma-aminobutyric acid type B receptor subunit two ( GABBR2 ; top SNP rs7038285, p = 6.98x10 − 7 , Odds Ratio (OR) = 0.85), where the minor allele (C) was enriched among individuals with NESAM. SNPs within intron 13 of GABBR2 (~ 260kb upstream of our top SNPs) have been associated with lower plasma homocysteine 41 in a European population. The mechanistic relationship of this association to the GABBR2 gene, which is a member of the G-protein coupled receptor 3- and GABA-B receptor families, are unclear. We also found similar evidence of putative association at PRICKLE2 on chromosome 3p14.1 (top SNP rs11130959, p = 4.15x10 − 5 , OR = 0.88) (Supplementary Figure S4B ), with the minor allele also being enriched among NESAM participants. The canonical function of PRICKLE2 is unknown, although it has been associated with variation in folate levels 79 . Of the seven top candidate loci, five ( MTHFR , AHCYL1 , PRICKLE2 , GABBR2 , and PLD2 ) directly impact either conversion to- or breakdown of- homocysteine (Fig. 2), including a coding variant rs1052748 (p.Thr577Ile) in exon 17 of PLD2 that was enriched among ESAM individuals (Supplementary Figure S4C ). PLD2 encodes phospholipase D2 phosphatidyl, which generates the methyl-group donor choline through the hydrolysis of phosphatidylcholine. The rs1052748 variant is reported as an expression quantitative trait locus (eQTL) for PLD2 in multiple tissues in the Gene-Tissue Expression database (GTEx) 80 . Metabolic flux results from cumulative enzymatic activity at multiple points within a biochemical cycle. Given our observation of multiple putatively associated OCM SNPs, each with variable effects on OCM, we considered the potential for a cumulative effect of OCM-associated SNPs on ESAM susceptibility. To test this, we applied a permutation test ( Methods ) to the association statistics of OCM SNPs and compare this with the association tests of the ~ 1.8 million genotyped SNPs genome wide (for which no SNP surpassed genome-wide significance). At a p-value cut-off of < 0.001, a greater proportion of SNPs at OCM loci were associated with ESAM relative to the randomly sampled dataset ((z = 3.2; Fig. 3 a). At more permissive p-value thresholds, the z-score differentiating our OCM SNPs and the random sampling experiment was even more pronounced (e.g., at a threshold of p = 0.01, the resulting z-score was 50 (Fig. 3 c)). Next, we considered whether our cumulative association might have resulted from comparing SNPs in linkage disequilibrium (LD) at metabolically related loci to a random distribution of largely unlinked SNPs. To do so, we repeated the permutation test, but this time randomly sampling the same number of SNPs (n = 9,479) from the 51,105 SNPs found at loci related to ketone metabolism, which is not known to be associated with SAM ( Methods ). Relative to this ketone metabolism background, the proportion of OCM SNPs showing association was still in the upper tail of the distribution (z = 3.19) at the p < 0.001 threshold (Fig. 3 b). Finally, as a negative control, we performed the same genome-wide comparison as above for SNPs at sphingolipid metabolism loci, which has a similar number of SNPs to OCM, but no a priori evidence for involvement in ESAM. The proportion of sphingolipid metabolism SNPs showing association with ESAM fell well within the randomly generated distribution (z= -0.70). There are few established models for evaluating genetic variation in SAM. Therefore, we sought to provide support for the proposed functional impact of associated variants by assessing the impact of ESAM-associated SNPs on OCM metabolites in the context of ESAM vs NESAM. We first extracted cell-free DNA (cfDNA) from ~ 400 serum samples (M ethods ) from our previous study of OCM metabolites in ESAM/NESAM, which was done in an independent cohort of children from Malawi recruited around the time of their SAM diagnosis 31 . We then genotyped four ESAM-associated SNPs (rs753562 ( PRICKLE2 ), rs17664202 ( PRICKLE2) , rs17664204 ( GABBR2 ), and rs651495 ( PLD2 ) using cfDNA samples ( Methods; Supplementary Methods ). Three SNPs (rs17664202 was out of Hardy-Weinberg equilibrium) were then used as instrumental variables in a two-sample Mendelian Randomization (MR) to assess the causal effect of each of the 16 OCM metabolites on ESAM risk ( Methods ). We first assessed SNP-metabolite and SNP-disease associations separately under regression models and then combined effects using the IVW method assuming additive genetic effects. Although OCM metabolites are highly correlated, we found significant casual effects on cystathionine (log-odds ratio ~ 2.49, p = 1.65x10 − 6 ) and on betaine (log-odds ratio ~ 4.69, p = 0.0007) (Fig. 4 ). With three SNPs, there was substantial genotypic heterogeneity in the causal effect of betaine and cysteine (p < 0.001 by Cochran's Q test for both). We thus reevaluated the model after excluding any SNPs with opposing effects (although such effects could reflect the effect of the major allele), which necessarily limited our analyses to the IVW approach. Using the two SNPs with congruent effects, the causal effect of cysteine was estimated ~ -7.97, (p = 2.24x10 − 6 and Cochran’s Q test p = 0.74) and the estimate for betaine was 5.41, (p = 0.0001 and Cochran’s Q test p = 0.005). Although we did not find strong evidence for cystathionine or betaine under the MR-Egger approach, we found that estimates from both approaches were similar in direction and magnitude. Population ancestry influences the association of OCM loci with ESAM In the absence of an available secondary cohort in which to replicate our findings, we also sought evidence of internal consistency in the magnitude and direction of effect between the two countries represented in our cohort. At our seven candidate loci, MAFs and directions of effect among associated SNPs were similar between the two countries ( Supplementary Table S5 ). Despite similarities in the direction of effect in the two countries, we noted that the magnitude of the effect was generally stronger in Malawi than Jamaica despite comparable sample sizes and minor allele frequencies ( Supplementary Figure S5 ). This difference was also evident when we looked at statistically inferred haplotypes ( Methods ) comprised of associated (genotyped) SNPs at our two top loci – GABBR2 and PRICKLE2 . We observed evidence of association with ESAM in each population group; however, the haplotypes with the strongest implied effect differed between the two population groups despite similar patterns of LD in the two groups at both loci ( Supplementary Figure S6 ). Similarly, when we stratified our cumulative association analyses by population group, there was always a higher proportion of association SNPs in Malawi (Fig. 3 c). To explore this further, we first considered our populations in multidimensional scaling (MDS) space. Jamaican and Malawian samples clustered with other African continental populations on MDS components one and two, although samples from Malawi were tightly clustered alongside other Bantu-speaking Niger Congo ethnolinguistic groups (YRI, MSL). Conversely, Jamaican samples were more dispersed (Fig. 5 a; Supplementary Figure S2 ), likely a result of the diverse population ancestry origins of Jamaicans, which include a predominant contribution of African ancestry from the British slave trade of the 14th to 17th centuries, as well as a subsequent influx of indentured peoples from East- and Southeast- Asia, and Europe, in addition to centuries of European (British and Spanish) colonialization. We then used ADMIXTURE 67 to infer patterns of genome-wide admixture in our cohort in the context of parental proxies of geographic and continental differentiation derived from the African Genome Variation Project (AGVP) 64 , the Human Health and Heredity in Africa (H3Africa) Project 65 , and 1000 Genomes (1KG) Project 59 ( Methods ). In this analysis, we proxied European ancestry through the Ceph from Utah (CEU); West African ancestries through the Yoruba from Nigeria (YRI) and Mende from Sierra Leone (MSL); East African ancestries through the Luhya from Webuye, Kenya, and the Amhara from Ethiopia; and southern African ancestries through individuals from Botswana (BOT), Bantu-speakers from Zambia (BSZ), the Zulu, and Khoe and San Hunter-Gatherer groups (KS). As expected, samples from Jamaica had, on average, a significantly higher proportion of European (~ 13%) and west African ancestry (~ 68%) than Malawians, in whom east-African ancestries (~ 83%) were more common (Welch two-sample t-test, p = 2.2x10 − 16 ) (Fig. 5 b; Supplementary Figure S7 ). Association between OCM loci and ESAM is driven by shared east-African ancestry Given the ancestral genetic differences between the two study populations, concomitant differences in associated haplotype backgrounds, and reported geographical variation in ESAM prevalence, we considered that the underlying OCM-ESAM association might be best represented by an ancestral haplotype that is shared by the two populations but seen more commonly in Malawi. In this shared ancestry model, the presence of additional ancestry backgrounds (admixture) might obscure the underlying association signal, especially if using genome-averaged ancestry to account for population stratification. To evaluate this, we used RFmix 69 to derive ‘local’ (locus-level) patterns of ancestry, using the CEU, MSL, LWK populations to proxy European, West African-, and East African- Bantu-speaking parental ancestries, respectively, in our cohort ( Methods ). We then repeated our OCM-wide association tests, but this time adjusting each SNP for its ancestral haplotype background using TRACTOR. This approach allowed us to account for the potential effects of differing ancestral haplotype backgrounds at putatively associated loci in a way that was agnostic to geography but still ancestrally sensitive. Consistent with genome-wide admixture estimates, Jamaican individuals harbored more blocks of European- and west-African- derived ancestry than Malawi (Fig. 6 a). Across the combined cohort, however, the overall ancestry proportions in ESAM were not significantly different than seen in the reference NESAM population (t-test, lowest p = 0.24; Supplementary Figure S8 ), suggesting that differences in effect size and association in our original association were not solely the consequence of mismatched ancestry between ESAM and NESAM. Using TRACTOR to perform logistic regression with 3-way admixture, we found that when conditioned on shared East African ancestry (proxied by LWK), seven loci surpassed the loci-level threshold for association, including candidates MTHFR1 , PRICKLE2 , and PLD2 , but only one OCM locus, and none of our candidates, met the same significance criterion when adjusting for West African ancestry (Fig. 6 b-c). We also observed that stratifying our random sampling cumulative association by ancestry background demonstrated a higher cumulative association between variants of presumed East African ancestry compared to West African ancestry at a p-value threshold of 0.001 (z = 2.81 and = 0.54, respectively) (Fig. 6 d). Collectively, these findings were consistent with our OCM-ESAM association being driven by variants carried on a shared haplotype background more common among individuals of East African ancestry. Signatures of selection at OCM loci among African populations Periods of famine have been associated with substantial childhood mortality, with survival into adolescence following childhood starvation potentially having a major effect on reproductive fitness. In the absence of modern-day nutritional support, ESAM is associated with higher mortality than NESAM, creating the potential for variants and loci that influence the risk of developing ESAM or NESAM to be subject to relatively strong selection. To explore this further, we generated haplotype-similarity-based scores (iHS) of recent selection at our OCM loci (+/- 50kb) using whole genome sequence data from a subset of 151 individuals from Malawi ( Methods ). We contextualized our results by considering iHS scores for the same OCM loci generated from genome sequencing done through the H3Africa consortium 65 , which includes samples from Benin, Botswana, Cameroon, Mali, Nigeria, Zambia, and Gur speakers from West Africa (Burkino Faso and Ghana). Using an ad hoc, but conservative, locus-selection threshold of more than 10% of SNPs having normalized iHS scores > 2 or <-2, we found that in Malawi, 15 of our 95 OCM loci (15.8%) had some evidence of selection, with AMT (0.429), SDS / SDSL (0.301), MDH2 (0.208) and SHMT1 (0.174) having the most SNPs surpassing our threshold ( Supplementary Table S6 ). Of these loci, however, none overlapped our top ESAM-associated candidate loci, most of which had proportions < 3%. Across the African countries in the H3Africa dataset, however, 2,440 SNPs within 50kb of OCM genes had evidence of selection in at least two countries ( Supplementary Table S6; Supplementary Figure S9 ), with 55 SNPs having strong evidence across all surveyed groups (Fig. 7 ) . In this broader dataset, rs1052748 (ESAM-associated PLD2 coding variant) had the strongest selection signal among ESAM-associated OCM SNPs (Fig. 7 ; Supplementary Table S7 ), as did nearby PLD2 SNPs, particularly in Zambia (proportion = 9.6%), Cameroon (9.3%), and among Gur speakers from West Africa (13.3%). Lastly, we considered that haplotype similarity scores at selected loci might be amplified among ESAM samples, mimicking the cumulative allelic effects seen in our association models (i.e. enriching for multiple copies of a risk/protective allele). Across all OCM loci, the mean iHS score was significantly higher in ESAM (n = 90) than NESAM individuals (n = 61) (normalized iHS- two sample t-test p = 4x10 − 7 ; raw iHS - p = 5x10 − 14 ; Supplementary Figure S10a ). Among variant sites within 10KB of SNPs surpassing our locus-wide cut-off (n = 839 ESAM; n = 753 NESAM), mean iHS scores were also significantly different (two sample t-test p = 0.01; Supplementary Figure S10b ) between the two groups. DISCUSSION We report results of a case-control genetic association study of ESAM (affected/cases) versus NESAM (reference/controls) at loci associated with OCM. To our knowledge, this is the largest genetic study of ESAM to date. Our hypothesis-driven targeted pathway approach allowed us to maximize our modest sample size and, in the absence of an available replication cohort, we find strong internal consistency between the two country cohorts assessed. SAM is an acute, life-threatening disorder, for which, traditionally, recruitment, consent, and sampling for large-scale research are secondary considerations; however, the ability to retrospectively recruit individuals several months to years after surviving SAM holds potential for future, larger, germline genetic studies of ESAM. Here, we report evidence of an association between cumulative and locus-specific genetic variation at OCM-loci and the risk of developing ESAM. This association appears to be driven by variants carried on an East African genetic background and is bolstered by evidence that associated variants causally-mediate the association between OCM metabolites and ESAM, as well as by suggestions of recent selection at ESAM-OCM associated loci. Intronic SNPs in GABBR2 and PRICKLE2 showed the strongest locus-specific association. Both loci include multiple sites of reported open chromatin, suggestive of regulatory effects; however, it is unknown whether these effects would be limited to the most proximal genes versus having more distal effects on other cis- or trans- genes. The latter observation is particularly relevant at GABBR2 where the top SNP (rs70387285) had a much higher minor allele frequency among African populations. The relative lack of transcriptional or gene-regulatory contexts from Africa makes it difficult to fully annotate the regulatory potential of candidate variants identified in studies such as ours. We did, however, observe locus-wide positive association between a GTEx-reported eQTL missense variant in PLD2 (rs1052748) and ESAM. The minor allele of rs1052748 is associated with reduced PLD2 transcription. The functional consequence of variable transcription on phospholipase D2 phosphatidyl activity and production of choline from phosphatidylcholine is uncertain; however, reduced choline availability would be consistent with lower choline concentrations noted in ESAM 31 . These single locus metabolic effects, however, do not occur in isolation - small changes in metabolic flux at multiple points in OCM could result in larger effects on OCM consistent with those reported in ESAM. This was consistent with both our cumulative association model and our Mendelian Randomization models. For instance, most of the loci reaching locus-wide association were associated with homocysteine metabolism, and MR analyses suggested causal effects of SNPs at GABBR2, PRICKLE2 , and PLD2 on metabolites that are either directly derived from homocysteine (Cysteine and Cystathionine) or are involved in its re-conversion to methionine (Betaine). These observations strengthen previous studies implicating aberrant OCM, and particularly deficient re-methylation of homocysteine to methionine, in the pathogenesis of ESAM. These data thus provide support for proposed therapeutic interventions to supplement OCM turnover in SAM using cofactors such as choline (clinicaltrials.gov ID NCT06154174). Despite this, our preliminary estimates suggest whilst genetic variation has a moderately strong effect on the variance in SAM outcomes (estimated heritability = 0.22; standard error (se) = 0.37), variation at OCM loci accounts for a much smaller proportion (estimated heritability = 0.06; se = 0.06). The modest sample size employed means that these heritability estimates have a relatively large standard error as do the effect sizes inferred from our MR analysis. Optimistically, however, there may be several unidentified loci that causally influence ESAM risk. Larger cohorts, particularly from areas where the prevalence of ESAM relative to NESAM remains high, would allow for a well-powered genome-wide approach to identify other ESAM-associated loci and/or pathways for therapeutic and nutritional targeting. We were able to uniquely leverage African intra-continental admixture patterns to aid in our mapping. The deep ancestral tree within Africa harbors a complex and diverse genetic landscape that can be as disparate as inter-continental variation. The higher rates of ESAM in east-central and southern Africa, therefore, presented an opportunity to employ admixture mapping approaches that have been long proposed for diseases with different population prevalences. We observed evidence of ESAM association at OCM loci when conditioned on East African, but not West African, haplotype backgrounds shared between Malawi and Jamaica. We hypothesized that reducing the ‘noise’ of West African haplotypes in the analysis, would have augmented the effect of shared East African ancestry beyond the effect seen in our primary association. The weaker signal observed likely represents limitations in the representation of African haplotypes in current imputation panels and public databases 81 , 82 , making imputation and selection of parental ancestral proxies necessarily deficient (e.g. the LWK are an imperfect proxy for the ancestral East African ancestry of Malawi). None the less, our results provide a conceptual model for leveraging admixture within genetically heterogenous groups to genetically map traits with prevalence differences between groups, provided appropriate parental proxies can be identified and there is sufficient genetic distance between the groups in question. There are several examples of how migration across, into, and out of Africa, has interfaced with varying infectious and non-infectious exposures to shape the African genome through adaptation 83 . More recent studies have implicated multiple novel loci, often with varying effect sizes across the continent 65 . The history of famine and food crises in Africa varies by country and are complex; periods of famine have been documented as far back as the 1600s 84 , with speculation that these were more severe in the 19th century, at least in east and southern African countries such as Zimbabwe and Tanzania 85 , 86 . At the same time, children developing ESAM have been noted to have higher birth weights 87 , despite the putative link between low birth weight and early-life SAM. These observations provided a starting point for considering that differential survival between ESAM and NESAM might be reflected in signatures of recent selection at ESAM-associated OCM loci. We noted several OCM loci with strong evidence of selection, particularly the serine dehydratase (SDS/SDSH) complex on chromosome 12, which had a strong signal in all the populations evaluated. Among our ESAM-associated OCM loci, however, only PLD2 showed evidence of selection. Like our genetic association, we observed that OCM selection scores among ESAM individuals were, on average, significantly higher than observed among participants with NESAM, suggesting that we may be similarly underpowered to observe subtle selection signatures at individual loci that contribute to a pathway-wide metabolic effect. Going forward, integration of selection statistics at scale, particularly in the context of African genomic association studies, could provide valuable biological insights, especially for novel loci and understudied disorders. Genome sequencing and genetic mapping efforts in Africa and among populations of predominantly African ancestry are gaining traction, although the number of reference genomes, particularly for non-West African groups, still vastly under-represents the variation across the continent. Our study provides a starting exemplar for applying a hypothesis-driven, population genetics-informed model to study a disease with high prevalence in specific African population groups. Continued expansion of sequencing efforts in African populations, alongside refinement and deployment of association models that can integrate intra-continental admixture and population genetic statistics, has the potential to maximize the future of human genetics studies on the continent and globally. Declarations ACKNOWLEDGMENTS This work utilized the computational resources of the NIH HPC Biowulf cluster (http://hpc.nih.gov) as well as the University of Cape Town’s ICTS High Performance Computing team (hpc.uct.ac.za). This work represents the brainchild of Prof. Colin McKenzie, who died prior to the publication of the manuscript. The authors would also like to acknowledge the advice of Historian Dr. Christopher Donohue of the National Human Genome Research Institute (NHGRI) regarding famine across Africa in recent centuries. This research was supported by a Clinical Scientist Development Award from the Doris Duke Charitable Foundation (Grant #: 2013096) and USDA, ARS cooperative agreement (58-3092-5-001), both to N.A.H., K.V.S. and N.C.L. were supported by Award Numbers T32GM008307 and 5 T32GM008231, respectively, from the National Institute of General Medical Sciences to Baylor College of Medicine. E.G.A was supported by grant R01HG012869 from NHGRI. Portions of the work reported were supported by grant HG-200412 from the NHGRI of the NIH to N.A.H. The work and views expressed do not reflect the views of BCM or the NIH. AUTHOR CONTRIBUTIONS I.T., M.J.M., C.A.M., M.E.R., C.T-B., and N.A.H. designed the study; O.B., C.T-B., I.T., T.M., and M.J.M. recruited participants, obtained informed consent, and collected samples; S.H., K.M., R.S., N.J.H. and A.H. processed and molecularly characterized samples; S.T. conducted cell-free DNA (cfDNA) extractions and J.R. conducted SNP genotyping from cfDNA. A.J., S.S., N.C.L., Y.H., Q.L., K.V.S., and N.A.H. conducted the primary data analyses; D.S. and A.C conducted population genetic analyses including that on H3Africa populations; E.G.A. facilitated local ancestry and TRACTOR analyses; W.A and M.M. performed imputation on the H3Africa server for data from Malawi. N.C.L., Q.L., E.B., and N.A.H. wrote the paper. All authors read, edited, and approved the manuscript. DECLARATION OF INTERESTS The remaining authors do not have any conflicts or relevant interests to declare. WEB RESOURCES H3Africa array v1.1 - https://h3africa.org/index.php/2019/12/12/h3africa-chip-faq/ H3Africa Array manifest - https://www.illumina.com/content/infinium-h3africa-consortium-array-data-sheet) H3ABioNet Imputation Service - https://www.h3abionet.org/resources/h3abionet-imputation-service; h3abionet/chipimputation. Minimac - https://github.com/statgen/Minimac4 University of Cape Town High-Performance Computing - https://ucthpc.uct.ac.za/ 1000 Genomes Project Resource - https://www.internationalgenome.org/ ‘In house scripts’ - https://github.com/NHGRI/SAM_OCMgenetics DATA AND CODE AVAILABILITY The code generated during this study and summary statistics are available through GitHub (https://github.com/NHGRI/SAM_OCMgenetics). The raw sequencing and genotyping data supporting the current study have not been deposited in a public repository, as Ethical approvals in Malawi and Jamaica accompanying recruitment and sample collection for this study predated current data sharing protocols and do not explicitly include publicly sharing data; however, investigators interested in study-based data access should directly contact the corresponding author. References Levels and trends in child malnutrition: UNICEF/WHO/The World Bank Group joint child malnutrition estimates: key findings of the 2020 edition. (2023). Black, R.E., Victora, C.G., Walker, S.P., Bhutta, Z.A., Christian, P., de Onis, M., Ezzati, M., Grantham-McGregor, S., Katz, J., Martorell, R., et al. (2013). Maternal and child undernutrition and overweight in low-income and middle-income countries. Lancet 382 , 427-451. 10.1016/S0140-6736(13)60937-X. WHO child growth standards and the identification of severe acute malnutrition in infants and children. (2024). Bhutta, Z.A., Berkley, J.A., Bandsma, R.H.J., Kerac, M., Trehan, I., and Briend, A. (2017). Severe childhood malnutrition. Nat Rev Dis Primers 3 , 17067. 10.1038/nrdp.2017.67. Frison, S., Checchi, F., and Kerac, M. (2015). Omitting edema measurement: how much acute malnutrition are we missing? The American journal of clinical nutrition 102 , 1176-1181. 10.3945/ajcn.115.108282. Alvarez, J.L., Dent, N., Browne, L., and Briend, M.M.a.A. (2016). Putting Child Kwashiorkor on the Map. Sturgeon, J.P., Mufukari, W., Tome, J., Dumbura, C., Majo, F.D., Ngosa, D., Chandwe, K., Kapoma, C., Mutasa, K., Nathoo, K.J., et al. (2023). Risk factors for inpatient mortality among children with severe acute malnutrition in Zimbabwe and Zambia. European journal of clinical nutrition 77 , 895-904. 10.1038/s41430-023-01320-9. Childhood Acute, I., and Nutrition, N. (2022). Childhood mortality during and after acute illness in Africa and south Asia: a prospective cohort study. Lancet Glob Health 10 , e673-e684. 10.1016/S2214-109X(22)00118-8. Golden, M.H. (1998). Oedematous malnutrition. Br Med Bull 54 , 433-444. Waterlow, J.C. (1997). Protein-energy malnutrition: the nature and extent of the problem. Clinical Nutrition (Edinburgh, Scotland) 16 Suppl 1 , 3-9. 10.1016/s0261-5614(97)80043-x. McKenzie, C.A., Wakamatsu, K., Hanchard, N.A., Forrester, T., and Ito, S. (2007). Childhood malnutrition is associated with a reduction in the total melanin content of scalp hair. Br J Nutr 98 , 159-164. 10.1017/S0007114507694458. Williams, C.D. (1935). Kwashiorkor. A disease of children associated with a maize diet. Lancet ii , 1151-1152. Williams, C.D. (1933). A nutritional disease of childhood associated with a maize diet. Archives of Disease in Childhood 8 , 423-433. Heikens, G.T., and Manary, M. (2009). 75 years of Kwashiorkor in Africa. Malawi medical journal : the journal of Medical Association of Malawi 21 , 96-98. Manary, M.J., Heikens, G.T., and Golden, M. (2009). Kwashiorkor: more hypothesis testing is needed to understand the aetiology of oedema. Malawi medical journal : the journal of Medical Association of Malawi 21 , 106-107. Kismul, H., Van den Broeck, J., and Lunde, T.M. (2014). Diet and kwashiorkor: a prospective study from rural DR Congo. PeerJ 2 , e350. 10.7717/peerj.350. Lin, C.A., Boslaugh, S., Ciliberto, H.M., Maleta, K., Ashorn, P., Briend, A., and Manary, M.J. (2007). A prospective assessment of food and nutrient intake in a population of Malawian children at risk for kwashiorkor. J Pediatr Gastroenterol Nutr 44 , 487-493. 10.1097/MPG.0b013e31802c6e57. Sullivan, J., Ndekha, M., Maker, D., Hotz, C., and Manary, M.J. (2006). The quality of the diet in Malawian children with kwashiorkor and marasmus. Matern Child Nutr 2 , 114-122. 10.1111/j.1740-8709.2006.00053.x. Golden, M.H. (2015). Nutritional and other types of oedema, albumin, complex carbohydrates and the interstitium - a response to Malcolm Coulthard's hypothesis: Oedema in kwashiorkor is caused by hypo-albuminaemia. Paediatrics and International Child Health 35 , 90-109. 10.1179/2046905515Y.0000000010. Laditan, A.A., and Reeds, P.J. (1976). A study of the age of onset, diet and the importance of infection in the pattern of severe protein-energy malnutrition in Ibadan, Nigeria. Br J Nutr 36 , 411-419. 10.1079/bjn19760096. Kristensen, K.H., Wiese, M., Rytter, M.J., Ozcam, M., Hansen, L.H., Namusoke, H., Friis, H., and Nielsen, D.S. (2016). Gut Microbiota in Children Hospitalized with Oedematous and Non-Oedematous Severe Acute Malnutrition in Uganda. PLoS Negl Trop Dis 10 , e0004369. 10.1371/journal.pntd.0004369. Christie, C.D., Heikens, G.T., and Golden, M.H. (1992). Coagulase-negative staphylococcal bacteremia in severely malnourished Jamaican children. Pediatr Infect Dis J 11 , 1030-1036. 10.1097/00006454-199211120-00008. Jahoor, F., Badaloo, A., Reid, M., and Forrester, T. (2008). Protein metabolism in severe childhood malnutrition. Annals of Tropical Paediatrics 28 , 87-101. 10.1179/146532808X302107. Badaloo, A.V., Forrester, T., Reid, M., and Jahoor, F. (2006). Lipid kinetic differences between children with kwashiorkor and those with marasmus2. The American Journal of Clinical Nutrition 83 , 1283-1288. 10.1093/ajcn/83.6.1283. Badaloo, A., Hsu, J.W., Taylor-Bryan, C., Green, C., Reid, M., Forrester, T., and Jahoor, F. (2012). Dietary cysteine is used more efficiently by children with severe acute malnutrition with edema compared with those without edema. Am J Clin Nutr 95 , 84-90. 10.3945/ajcn.111.024323. Jahoor, F., Badaloo, A., Reid, M., and Forrester, T. (2006). Sulfur amino acid metabolism in children with severe childhood undernutrition: methionine kinetics. The American journal of clinical nutrition 84 , 1400-1405. Manary, M.J., Leeuwenburgh, C., and Heinecke, J.W. (2000). Increased oxidative stress in kwashiorkor. J Pediatr 137 , 421-424. 10.1067/mpd.2000.107512. Amadi, B., Fagbemi, A.O., Kelly, P., Mwiya, M., Torrente, F., Salvestrini, C., Day, R., Golden, M.H., Eklund, E.A., Freeze, H.H., and Murch, S.H. (2009). Reduced production of sulfated glycosaminoglycans occurs in Zambian children with kwashiorkor but not marasmus. Am J Clin Nutr 89 , 592-600. 10.3945/ajcn.2008.27092. Ducker, G.S., and Rabinowitz, J.D. (2017). One-Carbon Metabolism in Health and Disease. Cell metabolism 25 , 27-42. 10.1016/j.cmet.2016.08.009. Jahoor, F., Badaloo, A., Reid, M., and Forrester, T. (2008). Protein metabolism in severe childhood malnutrition. Annals of tropical paediatrics 28 , 87-101. 10.1179/146532808X302107. May, T., de la Haye, B., Nord, G., Klatt, K., Stephenson, K., Adams, S., Bollinger, L., Hanchard, N., Arning, E., Bottiglieri, T., et al. (2022). One-carbon metabolism in children with marasmus and kwashiorkor. EBioMedicine 75 , 103791. 10.1016/j.ebiom.2021.103791. May, T., Klatt, K.C., Smith, J., Castro, E., Manary, M., Caudill, M.A., Jahoor, F., and Fiorotto, M.L. (2018). Choline Supplementation Prevents a Hallmark Disturbance of Kwashiorkor in Weanling Mice Fed a Maize Vegetable Diet: Hepatic Steatosis of Undernutrition. Nutrients 10 . 10.3390/nu10050653. Smith, M.I., Yatsunenko, T., Manary, M.J., Trehan, I., Mkakosya, R., Cheng, J., Kau, A.L., Rich, S.S., Concannon, P., Mychaleckyj, J.C., et al. (2013). Gut microbiomes of Malawian twin pairs discordant for kwashiorkor. Science 339 , 548-554. 10.1126/science.1229000. Schulze, K.V., Swaminathan, S., Howell, S., Jajoo, A., Lie, N.C., Brown, O., Sadat, R., Hall, N., Zhao, L., Marshall, K., et al. (2019). Edematous severe acute malnutrition is characterized by hypomethylation of DNA. Nat Commun 10 , 5791. 10.1038/s41467-019-13433-6. Munthali, T., Jacobs, C., Sitali, L., Dambe, R., and Michelo, C. (2015). Mortality and morbidity patterns in under-five children with severe acute malnutrition (SAM) in Zambia: a five-year retrospective review of hospital-based records (2009–2013). Archives of Public Health 73 , 23. 10.1186/s13690-015-0072-1. Gonzales, G.B., Ngari, M.M., Njunge, J.M., Thitiri, J., Mwalekwa, L., Mturi, N., Mwangome, M.K., Ogwang, C., Nyaguara, A., and Berkley, J.A. (2020). Phenotype is sustained during hospital readmissions following treatment for complicated severe malnutrition among Kenyan children: A retrospective cohort study. Matern Child Nutr 16 , e12913. 10.1111/mcn.12913. Duggal, P., and Petri, W.A., Jr. (2018). Does Malnutrition Have a Genetic Component? Annu Rev Genomics Hum Genet 19 , 247-262. 10.1146/annurev-genom-083117-021340. Marshall, K.G., Howell, S., Reid, M., Badaloo, A., Farrall, M., Forrester, T., and McKenzie, C.A. (2006). Glutathione S-transferase polymorphisms may be associated with risk of oedematous severe childhood malnutrition. Br J Nutr 96 , 243-248. 10.1079/bjn20061825. Marshall, K.G., Swaby, K., Hamilton, K., Howell, S., Landis, R.C., Hambleton, I.R., Reid, M., Fletcher, H., Forrester, T., and McKenzie, C.A. (2011). A preliminary examination of the effects of genetic variants of redox enzymes on susceptibility to oedematous malnutrition and on percentage cytotoxicity in response to oxidative stress in vitro. Ann Trop Paediatr 31 , 27-36. 10.1179/146532811X12925735813805. Marshall, K.G., Howell, S., Badaloo, A.V., Reid, M., Farrall, M., Forrester, T., and McKenzie, C.A. (2006). Polymorphisms in genes involved in folate metabolism as risk factors for oedematous severe childhood malnutrition: a hypothesis-generating study. Ann Trop Paediatr 26 , 107-114. 10.1179/146532806X107449. Hazra, A., Kraft, P., Lazarus, R., Chen, C., Chanock, S.J., Jacques, P., Selhub, J., and Hunter, D.J. (2009). Genome-wide significant predictors of metabolites in the one-carbon metabolism pathway. Hum Mol Genet 18 , 4677-4687. 10.1093/hmg/ddp428. Williams, S.R., Yang, Q., Chen, F., Liu, X., Keene, K.L., Jacques, P., Chen, W.-M., Weinstein, G., Hsu, F.-C., Beiser, A., et al. (2014). Genome-wide meta-analysis of homocysteine and methionine metabolism identifies five one carbon metabolism loci and a novel association of ALDH1L1 with ischemic stroke. PLoS genetics 10 , e1004214. 10.1371/journal.pgen.1004214. Marshall, K.G., Howell, S., Reid, M., Badaloo, A.V., Farrall, M., Forrester, T., and McKenzie, C.A. (2006). Glutathione S-transferase polymorphisms may be associated with risk of oedematous severe childhood malnutrition. The British journal of nutrition 96 , 243-248. Waterlow, J.C. (1972). Classification and definition of protein-calorie malnutrition. Br Med J 3 , 566-569. 10.1136/bmj.3.5826.566. Baker, E.J., Jay, J.J., Bubier, J.A., Langston, M.A., and Chesler, E.J. (2012). GeneWeaver: a web-based system for integrative functional genomics. Nucleic Acids Research 40 , D1067-1076. 10.1093/nar/gkr968. Ashburner, M., Ball, C.A., Blake, J.A., Botstein, D., Butler, H., Cherry, J.M., Davis, A.P., Dolinski, K., Dwight, S.S., Eppig, J.T., et al. (2000). Gene Ontology: tool for the unification of biology. Nature genetics 25 , 25-29. 10.1038/75556. Aleksander, S.A., Balhoff, J., Carbon, S., Cherry, J.M., Drabkin, H.J., Ebert, D., Feuermann, M., Gaudet, P., Harris, N.L., Hill, D.P., et al. (2023). The Gene Ontology knowledgebase in 2023. Genetics 224 , iyad031. 10.1093/genetics/iyad031. Sollis, E., Mosaku, A., Abid, A., Buniello, A., Cerezo, M., Gil, L., Groza, T., Güneş, O., Hall, P., Hayhurst, J., et al. (2023). The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Research 51 , D977-D985. 10.1093/nar/gkac1010. Hindorff, L.A., Junkins, H.A., Hall, P.N., and Manolio, T.A. (2011). A Catalog of Published Genome-Wide Association Studies. Kanehisa, M., and Goto, S. (2000). KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Research 28 , 27-30. 10.1093/nar/28.1.27. Geer, L.Y., Marchler-Bauer, A., Geer, R.C., Han, L., He, J., He, S., Liu, C., Shi, W., and Bryant, S.H. (2010). The NCBI BioSystems database. Nucleic Acids Research 38 , D492-496. 10.1093/nar/gkp858. Cerami, E.G., Gross, B.E., Demir, E., Rodchenkov, I., Babur, Ö., Anwar, N., Schultz, N., Bader, G.D., and Sander, C. (2011). Pathway Commons, a web resource for biological pathway data. Nucleic Acids Research 39 , D685-D690. 10.1093/nar/gkq1039. Purcell, S., Neale, B., Todd-Brown, K., Thomas, L., Ferreira, Manuel A.R., Bender, D., Maller, J., Sklar, P., de Bakker, Paul I.W., Daly, Mark J., and Sham, Pak C. (2007). PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses. American Journal of Human Genetics 81 , 559-575. Johnston, H.R., Hu, Y.J., Gao, J., O'Connor, T.D., Abecasis, G.R., Wojcik, G.L., Gignoux, C.R., Gourraud, P.A., Lizee, A., Hansen, M., et al. (2017). Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome. Sci Rep 7 , 46398. 10.1038/srep46398. Das, S., Forer, L., Schonherr, S., Sidore, C., Locke, A.E., Kwong, A., Vrieze, S.I., Chew, E.Y., Levy, S., McGue, M., et al. (2016). Next-generation genotype imputation service and methods. Nat Genet 48 , 1284-1287. 10.1038/ng.3656. Eagle: multi-locus association mapping on a genome-wide scale made routine | Bioinformatics | Oxford Academic. (2024). Mulder, N.J., Adebiyi, E., Alami, R., Benkahla, A., Brandful, J., Doumbia, S., Everett, D., Fadlelmola, F.M., Gaboun, F., Gaseitsiwe, S., et al. (2016). H3ABioNet, a sustainable pan-African bioinformatics network for human heredity and health in Africa. Genome research 26 , 271-277. 10.1101/gr.196295.115. Baichoo, S., Souilmi, Y., Panji, S., Botha, G., Meintjes, A., Hazelhurst, S., Bendou, H., Beste, E., Mpangase, P.T., Souiai, O., et al. (2018). Developing reproducible bioinformatics analysis workflows for heterogeneous computing environments to support African genomics. BMC Bioinformatics 19 , 457. 10.1186/s12859-018-2446-1. Genomes Project, C., Auton, A., Brooks, L.D., Durbin, R.M., Garrison, E.P., Kang, H.M., Korbel, J.O., Marchini, J.L., McCarthy, S., McVean, G.A., and Abecasis, G.R. (2015). A global reference for human genetic variation. Nature 526 , 68-74. 10.1038/nature15393. Dessau, R.B., and Pipper, C.B. (2008). [''R"--project for statistical computing]. Ugeskrift for Laeger 170 , 328-330. Zhou, X., and Stephens, M. (2012). Genome-wide efficient mixed-model analysis for association studies. Nature Genetics 44 , 821-824. 10.1038/ng.2310. Turner, S.D. (2018). qqman: an R package for visualizing GWAS results using Q-Q and manhattan plots. Journal of Open Source Software 3 , 731. 10.21105/joss.00731. Barrett, J.C., Fry, B., Maller, J., and Daly, M.J. (2005). Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics 21 , 263-265. 10.1093/bioinformatics/bth457bth457 [pii]. Jones, B. (2015). Population genetics: the African Genome Variation Project. Nat Rev Genet 16 , 68-69. 10.1038/nrg3886. Choudhury, A., Aron, S., Botigue, L.R., Sengupta, D., Botha, G., Bensellak, T., Wells, G., Kumuthini, J., Shriner, D., Fakim, Y.J., et al. (2020). High-depth African genomes inform human migration and health. Nature 586 , 741-748. 10.1038/s41586-020-2859-7. Schlebusch, C.M., Skoglund, P., Sjodin, P., Gattepaille, L.M., Hernandez, D., Jay, F., Li, S., De Jongh, M., Singleton, A., Blum, M.G., et al. (2012). Genomic variation in seven Khoe-San groups reveals adaptation and complex African history. Science 338 , 374-379. 10.1126/science.1227721. Alexander, D.H., Novembre, J., and Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Res 19 , 1655-1664. 10.1101/gr.094052.109. Delaneau, O., Marchini, J., and Zagury, J.F. (2012). A linear complexity phasing method for thousands of genomes. Nat Methods 9 , 179-181. 10.1038/nmeth.1785. Maples, B.K., Gravel, S., Kenny, E.E., and Bustamante, C.D. (2013). RFMix: a discriminative modeling approach for rapid and robust local-ancestry inference. American Journal of Human Genetics 93 , 278-288. 10.1016/j.ajhg.2013.06.020. Atkinson, E.G., Maihofer, A.X., Kanai, M., Martin, A.R., Karczewski, K.J., Santoro, M.L., Ulirsch, J.C., Kamatani, Y., Okada, Y., Finucane, H.K., et al. (2021). Tractor uses local ancestry to enable the inclusion of admixed individuals in GWAS and to boost power. Nature Genetics 53 , 195-204. 10.1038/s41588-020-00766-y. Burgess, S., Butterworth, A., and Thompson, S.G. (2013). Mendelian randomization analysis with multiple genetic variants using summarized data. Genet Epidemiol 37 , 658-665. 10.1002/gepi.21758. Burgess, S., and Thompson, S.G. (2017). Interpreting findings from Mendelian randomization using the MR-Egger method. Eur J Epidemiol 32 , 377-389. 10.1007/s10654-017-0255-x. Li, H., and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics 25 , 1754-1760. btp324 [pii] 10.1093/bioinformatics/btp324. McKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., Garimella, K., Altshuler, D., Gabriel, S., Daly, M., and DePristo, M.a. (2010). The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome research 20 , 1297-1303. 10.1101/gr.107524.110. van der Auwera, G., and O'Connor, B.D. (2020). Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (O'Reilly Media, Incorporated). Voight, B.F., Kudaravalli, S., Wen, X., and Pritchard, J.K. (2006). A map of recent positive selection in the human genome. PLoS Biol 4 , e72. 10.1371/journal.pbio.0040072. Szpiech, Z.A., and Hernandez, R.D. (2014). selscan: an efficient multithreaded program to perform EHH-based scans for positive selection. Mol Biol Evol 31 , 2824-2827. 10.1093/molbev/msu211. Wickham, H. (2016). ggplot2 (Springer International Publishing). Tanaka, T., Scheet, P., Giusti, B., Bandinelli, S., Piras, M.G., Usala, G., Lai, S., Mulas, A., Corsi, A.M., Vestrini, A., et al. (2009). Genome-wide association study of vitamin B6, vitamin B12, folate, and homocysteine blood concentrations. American Journal of Human Genetics 84 , 477-482. 10.1016/j.ajhg.2009.02.011. Consortium, G.T., Laboratory, D.A., Coordinating Center -Analysis Working, G., Statistical Methods groups-Analysis Working, G., Enhancing, G.g., Fund, N.I.H.C., Nih/Nci, Nih/Nhgri, Nih/Nimh, Nih/Nida, et al. (2017). Genetic effects on gene expression across human tissues. Nature 550 , 204-213. 10.1038/nature24277. Vergara, C., Parker, M.M., Franco, L., Cho, M.H., Valencia-Duarte, A.V., Beaty, T.H., and Duggal, P. (2018). Genotype Imputation Performance of Three Reference Panels Using African Ancestry Individuals. Human genetics 137 , 281-292. 10.1007/s00439-018-1881-4. Sengupta, D., Botha, G., Meintjes, A., Mbiyavanga, M., Study, A.W.-G., Consortium, H.A., Hazelhurst, S., Mulder, N., Ramsay, M., and Choudhury, A. (2023). Performance and accuracy evaluation of reference panels for genotype imputation in sub-Saharan African populations. Cell Genom 3 , 100332. 10.1016/j.xgen.2023.100332. Fan, S., Hansen, M.E., Lo, Y., and Tishkoff, S.A. (2016). Going global by adapting local: A review of recent human adaptation. Science 354 , 54-59. 10.1126/science.aaf5098. Isichei, E.A. (1997). A history of African societies to 1870 (Cambridge University Press). Cheater, A.P., and Iliffe, J.H. (1990). Famine in Zimbabwe 1890-1960. Africa. Bloxham, D., and Moses, A.D. (2022). Genocide : key themes, First edition. Edition (Oxford University Press). Forrester, T.E., Badaloo, A.V., Boyne, M.S., Osmond, C., Thompson, D., Green, C., Taylor-Bryan, C., Barnett, A., Soares-Wynter, S., Hanson, M.A., et al. (2012). Prenatal factors contribute to the emergence of kwashiorkor or marasmus in severe undernutrition: evidence for the predictive adaptation model. PLoS One 7 , e35907. 10.1371/journal.pone.0035907. Additional Declarations There is NO Competing Interest. Supplementary Files SAMOCMSupplementaryMethodsv2.docx Supplementary Methods SAMOCMSupplementaryTablesv5.xlsx Supplementary Tables SAMOCMSupplementaryfiguresv12.pdf Supplementary Figures Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6890799","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":475807619,"identity":"5147aec6-6b50-4b3d-9166-7daa0c149e44","order_by":0,"name":"Neil Hanchard","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6klEQVRIiWNgGAWjYFACxsYDjA1AWgKIeSoYeAyI0NKApOUMUVoYGBBaeNsYGAhqMWdvBtqywy6Pf3bzsw9v592RMWdgv/iYB48Wy56DQC1nkosl7hwznjl32zMeywaeYmN8WgxuJAK1tDEnbpBIMGbm3XaYx+AAT5rkDHxa7j8EaakHakn/zMw7hxgtN0Ah1nYYqCUHaEsDSAv7MYkP+LScATosse144owbOcWMc44BtRzmYTbAq+X48YcPPrZVJ/bPSN/M8KbmsL3B8faHDxLwaAEDVAXMxMUmCmB/QLKWUTAKRsEoGNYAAO45VIkudYskAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0000-0003-1925-2665","institution":"National Human Genome Research Institute","correspondingAuthor":true,"prefix":"","firstName":"Neil","middleName":"","lastName":"Hanchard","suffix":""},{"id":475807620,"identity":"0ee9746a-fc53-4fde-a352-b3c49f0918db","order_by":1,"name":"Natasha Lie","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Natasha","middleName":"","lastName":"Lie","suffix":""},{"id":475807621,"identity":"fa737bc5-8a9e-471d-8d82-ca62695dbac0","order_by":2,"name":"Yixing Han","email":"","orcid":"","institution":"National Institutes of Health","correspondingAuthor":false,"prefix":"","firstName":"Yixing","middleName":"","lastName":"Han","suffix":""},{"id":475807622,"identity":"edb39a90-96b4-4ae0-a44e-5ce45900d27e","order_by":3,"name":"Qing Li","email":"","orcid":"","institution":"National Human Genome Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Qing","middleName":"","lastName":"Li","suffix":""},{"id":475807623,"identity":"3a0d5cef-34a5-4a86-9558-c6726c12352c","order_by":4,"name":"Aarti Jajoo","email":"","orcid":"","institution":"Harvard Medical School","correspondingAuthor":false,"prefix":"","firstName":"Aarti","middleName":"","lastName":"Jajoo","suffix":""},{"id":475807624,"identity":"500d4e87-a13b-48bc-8952-20dc69b2dee6","order_by":5,"name":"Jared Redmond","email":"","orcid":"","institution":"National Human Genome Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Jared","middleName":"","lastName":"Redmond","suffix":""},{"id":475807625,"identity":"0e8df613-d87d-400e-968f-b929c447715b","order_by":6,"name":"Aparna Haldipur","email":"","orcid":"","institution":"National Human Genome Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Aparna","middleName":"","lastName":"Haldipur","suffix":""},{"id":475807626,"identity":"5a0a589f-9517-4147-a510-2fddc7fb8054","order_by":7,"name":"Solome Zewdu","email":"","orcid":"","institution":"National Human Genome Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Solome","middleName":"","lastName":"Zewdu","suffix":""},{"id":475807627,"identity":"b8b97a99-03f8-468e-914f-c27a59296d95","order_by":8,"name":"Emilyn Banfield","email":"","orcid":"","institution":"National Human Genome Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Emilyn","middleName":"","lastName":"Banfield","suffix":""},{"id":475807628,"identity":"654a2256-f424-43f3-88a8-bc897af38d15","order_by":9,"name":"Shanker Swaminathan","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Shanker","middleName":"","lastName":"Swaminathan","suffix":""},{"id":475807629,"identity":"986a6ae9-52e3-4deb-a0c2-9a4249309add","order_by":10,"name":"Sharon Howell","email":"","orcid":"","institution":"University of the West Indies","correspondingAuthor":false,"prefix":"","firstName":"Sharon","middleName":"","lastName":"Howell","suffix":""},{"id":475807630,"identity":"5381ed4e-416b-4a69-a4c0-c3c24a053b5f","order_by":11,"name":"Orgen Brown","email":"","orcid":"","institution":"University of the West Indies","correspondingAuthor":false,"prefix":"","firstName":"Orgen","middleName":"","lastName":"Brown","suffix":""},{"id":475807631,"identity":"b4304248-c543-4922-b2b4-174e2bfdc23e","order_by":12,"name":"Roa Sadat","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Roa","middleName":"","lastName":"Sadat","suffix":""},{"id":475807632,"identity":"482544cf-7793-4bd9-8c3d-2ef9a748b80b","order_by":13,"name":"Nancy Hall","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Nancy","middleName":"","lastName":"Hall","suffix":""},{"id":475807633,"identity":"4ade716e-02d1-440d-9166-f3e2f826911f","order_by":14,"name":"Katharina Schulze","email":"","orcid":"","institution":"Baylor College Of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Katharina","middleName":"","lastName":"Schulze","suffix":""},{"id":475807634,"identity":"996eb1f0-dea5-4d97-b56f-eed2dd055506","order_by":15,"name":"Thaddaeus May","email":"","orcid":"","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Thaddaeus","middleName":"","lastName":"May","suffix":""},{"id":475807635,"identity":"8495e039-54a0-4377-b682-328db88fc9fc","order_by":16,"name":"Marvin Reid","email":"","orcid":"","institution":"University of the West Indies","correspondingAuthor":false,"prefix":"","firstName":"Marvin","middleName":"","lastName":"Reid","suffix":""},{"id":475807636,"identity":"428e4fcd-f0c9-408e-9d79-444eba72dd70","order_by":17,"name":"Mark Manary","email":"","orcid":"","institution":"Washington University at St. Louis","correspondingAuthor":false,"prefix":"","firstName":"Mark","middleName":"","lastName":"Manary","suffix":""},{"id":475807637,"identity":"0ba0cd2a-72a1-49b5-a27a-805385e57dbc","order_by":18,"name":"Indi Trehan","email":"","orcid":"https://orcid.org/0000-0002-3364-6858","institution":"","correspondingAuthor":false,"prefix":"","firstName":"Indi","middleName":"","lastName":"Trehan","suffix":""},{"id":475807638,"identity":"3f31bb60-ec1d-4a9a-90d9-f9d925afb73f","order_by":19,"name":"Mamana Mbiyavanga","email":"","orcid":"","institution":"University of Cape Town","correspondingAuthor":false,"prefix":"","firstName":"Mamana","middleName":"","lastName":"Mbiyavanga","suffix":""},{"id":475807639,"identity":"8ac13ba3-842e-474e-b577-aa46b8db571c","order_by":20,"name":"Wisdom Akurugu","email":"","orcid":"","institution":"University of Cape Town","correspondingAuthor":false,"prefix":"","firstName":"Wisdom","middleName":"","lastName":"Akurugu","suffix":""},{"id":475807640,"identity":"330ac078-16e0-4cb7-b281-bb712354593e","order_by":21,"name":"Colin McKenzie","email":"","orcid":"","institution":"University of the West Indies","correspondingAuthor":false,"prefix":"","firstName":"Colin","middleName":"","lastName":"McKenzie","suffix":""},{"id":475807641,"identity":"f2e60750-3066-4525-925a-240ef0399f98","order_by":22,"name":"Dhriti Sengupta","email":"","orcid":"https://orcid.org/0000-0001-6315-7804","institution":"University of the Witwatersrand","correspondingAuthor":false,"prefix":"","firstName":"Dhriti","middleName":"","lastName":"Sengupta","suffix":""},{"id":475807642,"identity":"084e774b-1326-4306-b653-60732aebb044","order_by":23,"name":"Elizabeth Atkinson","email":"","orcid":"https://orcid.org/0000-0002-6308-776X","institution":"Baylor College of Medicine","correspondingAuthor":false,"prefix":"","firstName":"Elizabeth","middleName":"","lastName":"Atkinson","suffix":""},{"id":475807643,"identity":"0778512e-481c-44ab-9e91-6bca800c0652","order_by":24,"name":"Ananyo Choudhury","email":"","orcid":"https://orcid.org/0000-0001-8225-9531","institution":"University of the Witwatersrand","correspondingAuthor":false,"prefix":"","firstName":"Ananyo","middleName":"","lastName":"Choudhury","suffix":""},{"id":475807644,"identity":"80282f81-6471-4ef1-ae95-28932a9dfa00","order_by":25,"name":"Kwesi Marshall","email":"","orcid":"","institution":"University of the West Indies","correspondingAuthor":false,"prefix":"","firstName":"Kwesi","middleName":"","lastName":"Marshall","suffix":""},{"id":475807645,"identity":"76abb625-6cac-4d3e-a188-7fa59ea69d3a","order_by":26,"name":"Carolyn Taylor-Bryan","email":"","orcid":"","institution":"University of the West Indies","correspondingAuthor":false,"prefix":"","firstName":"Carolyn","middleName":"","lastName":"Taylor-Bryan","suffix":""}],"badges":[],"createdAt":"2025-06-13 21:40:25","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6890799/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6890799/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":85618327,"identity":"059e0094-8a08-46ed-9080-05f388487b2e","added_by":"auto","created_at":"2025-06-29 14:49:38","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":223276,"visible":true,"origin":"","legend":"\u003cp\u003eSAM samples and study recruitment in Jamaica and Malawi are shown alongside the prevalence (%) of nutritional edema (ESAM) across Africa from “Putting Kwashiorkor on the Map”\u003csup\u003e7\u003c/sup\u003e.\u003c/p\u003e","description":"","filename":"Figure1R.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/f4c315e618d6560b38bd8093.png"},{"id":85618324,"identity":"e84bd129-fbc6-4a78-a054-82cec53cc502","added_by":"auto","created_at":"2025-06-29 14:49:38","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":238301,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eA\u003c/strong\u003e - The one-carbon metabolism (OCM) cycle showing the folate and methionine cycles. In the folate cycle, dietary folate is processed to 5-methyl-tetrahydrofolate, which is then used in the production of methionine from homocysteine. Methionine is processed in its eponymous cycle to produce homocysteine, which contributes methyl groups for DNA and other cellular methylation and can be remethylated to methionine through the folate cycle. \u003cstrong\u003e2B\u003c/strong\u003e - QQ plot and Manhattan plot of SNP (genotyped and imputed) association with ESAM demonstrating candidate loci. The activity of loci reaching the locus threshold for significance (dashed line at –-logP= 3.27) or the SNP threshold for significance (dashed line at –logP= 5.9) in OCM is represented by colored arrows in 2A matching the locus colors in 2B.\u003c/p\u003e","description":"","filename":"Figure2R.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/f2901baa7a55b38855c7ab29.png"},{"id":85618328,"identity":"758cd632-afc8-42be-8a78-fb6eaf905cfe","added_by":"auto","created_at":"2025-06-29 14:49:38","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":131300,"visible":true,"origin":"","legend":"\u003cp\u003eHistogram illustrating the proportion of 10,000 randomly sampled SNPs surpassing a significance threshold of p\u0026lt; 0.001, compared to the number of SNPs in OCM loci passing the same threshold (red line). \u003cstrong\u003e3A\u003c/strong\u003e - SNPs were randomly sampled from the total genotyped pool. \u003cstrong\u003e3B\u003c/strong\u003e - SNPs randomly sampled from a pool of ketone metabolism loci. \u003cstrong\u003e3C\u003c/strong\u003e - Z scores of random sampling experiments for p-value thresholds between locus-level significance threshold and a p value ceiling of p=0.01 (see methods).\u003c/p\u003e","description":"","filename":"Figure3R.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/4cbbd77ee354524868a1107c.png"},{"id":85619336,"identity":"57e3c308-3a7d-4553-8290-ff6dabaec055","added_by":"auto","created_at":"2025-06-29 14:57:38","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":230199,"visible":true,"origin":"","legend":"\u003cp\u003eCausal effects estimated using a Two-Sample Mendelian Randomization (MR) approach to modelling combined effects from three independent SNPs. Estimates of SNP-exposure association are drawn from analyses of serum metabolites adjusting for sex and age; the estimates of SNP-disease association are drawn from GEMMA analyses of the Malawi group in the primary cohort. \u003cstrong\u003e4A\u003c/strong\u003e – Forest plot of causal SNP effects on cystathionine in the risk of developing ESAM vs NESAM. \u003cstrong\u003e4B \u003c/strong\u003e– Forest plot of causal SNP effects of cysteine; using MR-Egger, the log-odds ratio is estimated at 9.24 (p = 0.02). \u003cstrong\u003e4C \u003c/strong\u003e– Forest plot of causal SNP effects on betaine. Using the IVW approach, the estimated log odds ratio is 4.69 (p=0.0007), although Cochran's Q test p\u0026lt;0.001. IV – Instrumental Variable; CI – Confience Interval; OR – Odds Ratio.\u003c/p\u003e","description":"","filename":"Figure4R.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/1bc89f332d0c9c0fef8064a0.png"},{"id":85619339,"identity":"4730281a-d120-4e50-8ac8-b92c290abd40","added_by":"auto","created_at":"2025-06-29 14:57:38","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":569761,"visible":true,"origin":"","legend":"\u003cp\u003eAdmixture and Ancestry in the SAM cohort. \u003cstrong\u003e5A \u003c/strong\u003e– Multidimensional scaling dimensions (C1 vs C2) for Jamaican and Malawian SAM cohort in the context of African (AFR) populations from 1000 genomes (ESN – Esan in Nigeria; GWD – Gambians from Western Division; LWK – Luyha in Webuye, Kenya; MSL - Mende in Sierra Leone; YRI – Yoruba in Ibadan, Nigeria). \u003cstrong\u003e5B\u003c/strong\u003e - ADMIXTURE plots (K=5) for Jamaicans and Malawians in the SAM cohort shown in geographical context of putative ancestral source populations (CEU – Utah Residents (CEPH) with Northern and Western European ancestry; BOT – Botswana residents; BSZ – Bantu Speakers from Zambia from the H3Africa Consortium).\u003c/p\u003e","description":"","filename":"Figure5R2.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/455a85b488c7c70fedef2366.png"},{"id":85619334,"identity":"8658acbe-c02f-432d-a9d3-60295631a9c5","added_by":"auto","created_at":"2025-06-29 14:57:38","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":491774,"visible":true,"origin":"","legend":"\u003cp\u003eLocal ancestry association. \u003cstrong\u003e6A\u003c/strong\u003e - illustrative local ancestry plots for Jamaican and Malawian individuals. \u003cstrong\u003e6B\u003c/strong\u003e– QQ plot and Manhattan plot for TRACTOR-based association conditioning on West African (MSL) ancestry. \u003cstrong\u003e6C\u003c/strong\u003e - QQ plot and Manhattan plot for TRACTOR-based association, conditioned on East African (LWK) ancestry. \u003cstrong\u003e6D \u003c/strong\u003e– Cumulative association (see methods) by ancestry group showing the empirical distributions of the proportion of 5,916 SNP associations in TRACTOR with P\u0026lt;0.001 after resampling 10,000 times. \u0026nbsp;The red line shows the observed proportions for 5,916 SNPs at OCM loci.\u003c/p\u003e","description":"","filename":"Figure6R.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/7c42a7a19af9f4b2bc80c316.png"},{"id":85619967,"identity":"3aea4b89-30ee-4fbb-8743-b754563e0f22","added_by":"auto","created_at":"2025-06-29 15:05:38","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":706379,"visible":true,"origin":"","legend":"\u003cp\u003eStandardized iHS scores for OCM loci for eight countries/regions in Africa. Grey open circles identify iHS scores within 50kb of OCM genetic loci. Blue circles indicate iHS scores in regions within 10kb of ESAM-associated OCM genetic loci. The dashed lines represent thresholds of +2 and -2 iHS scores, which are suggestive of strong selection. There were no included OCM loci on chromosome 13.\u003c/p\u003e","description":"","filename":"Figure7R.png","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/33a95b7b741d99fd38f9eb4e.png"},{"id":85620325,"identity":"1453b865-2f99-425f-b118-5df726a1a55a","added_by":"auto","created_at":"2025-06-29 15:13:40","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3773437,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/ea6c6723-cf6c-4485-a802-e24653940035.pdf"},{"id":85618325,"identity":"2a19e61a-c40e-4a10-bf9e-7cd1734a00a8","added_by":"auto","created_at":"2025-06-29 14:49:38","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":42107,"visible":true,"origin":"","legend":"Supplementary Methods","description":"","filename":"SAMOCMSupplementaryMethodsv2.docx","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/f8e2cbe2ee4095287b7273c8.docx"},{"id":85618326,"identity":"2546f057-42d6-4db0-8111-46e62d4bbc54","added_by":"auto","created_at":"2025-06-29 14:49:38","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":68600,"visible":true,"origin":"","legend":"Supplementary Tables","description":"","filename":"SAMOCMSupplementaryTablesv5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/be90266d407e4f526f76a5ec.xlsx"},{"id":85619337,"identity":"ac40c05b-42cd-4ce2-8b0d-64e63960d8f3","added_by":"auto","created_at":"2025-06-29 14:57:38","extension":"pdf","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":1215716,"visible":true,"origin":"","legend":"Supplementary Figures","description":"","filename":"SAMOCMSupplementaryfiguresv12.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6890799/v1/b25c589c37720bf06bd585ca.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"\u003cp\u003eAberrant One-Carbon Metabolism and Ancestral Genetics Underlie Edematous Severe Acute Malnutrition\u003c/p\u003e","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eSevere acute malnutrition (SAM) affects 16.9\u0026nbsp;million children worldwide, predominantly those under five years old, and either directly or indirectly contributes to more than one million childhood deaths per year\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e,\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. SAM is clinically characterized by a weight-for-height that is more than three standard deviations below the median, a mid-upper-arm circumference of less than 115 mm, or the presence of nutritional edema\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. In practice, SAM is recognized to occur in two clinically distinct forms: edematous SAM (ESAM), which includes the syndromes of kwashiorkor and marasmic-kwashiorkor, and non-edematous SAM (NESAM), also known as marasmus\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. ESAM accounts for a higher proportion of hospitalized SAM in east- and central- Africa and the Caribbean\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e, and is associated with mortality rates of 10\u0026ndash;20% among those who are hospitalized with SAM and medical complications\u003csup\u003e\u003cspan additionalcitationids=\"CR8\" citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e. ESAM is clinically characterized by bilateral pitting edema (swelling) of the extremities and severe systemic involvement that may include skin and hair changes, fatty liver, and multi-system organ failure\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. By contrast, NESAM is more prevalent in Southeast Asia and northern- and western-African geographies, and is characterized by generalized wasting of body tissues, typically with less severe systemic involvement\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. Whilst it is well-recognized that the development of SAM is largely linked to food security and availability, the persistence of adverse geopolitical factors mean that SAM is still highly prevalent globally.\u003c/p\u003e \u003cp\u003eESAM was first described in medical literature in the early part of the 20th century\u003csup\u003e\u003cspan additionalcitationids=\"CR13\" citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e; however, the reasons why a severely malnourished child will develop ESAM as opposed to NESAM remain largely unknown\u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. In the past, it was hypothesized that diet or protein intake were primary contributors to ESAM risk; however, decades of epidemiological surveys have failed to show consistent evidence for differences in protein intake or any other single dietary deprivation between children who develop ESAM and those who develop NESAM\u003csup\u003e\u003cspan additionalcitationids=\"CR17 CR18\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. Similarly, there have been no consistently demonstrated differences in infectious or environmental exposures between children who develop ESAM and those who develop NESAM\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan additionalcitationids=\"CR21\" citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. There are, however, numerous pathophysiological differences between the two clinical manifestations, particularly in their metabolism and biochemistry; for example, ESAM has been associated with a slower rate of protein breakdown in starvation, impaired heparin sulfate proteoglycan expression, greater dietary cysteine efficiency, and increased oxidative stress, compared to children with NESAM\u003csup\u003e\u003cspan additionalcitationids=\"CR24 CR25 CR26 CR27\" citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e. Despite this, it remains unknown whether these observations represent downstream effects that result from the disorder or are primary etiological differences. Perhaps as a result, ESAM and NESAM are still treated in the same manner, using the same nutritional supplements, despite well described differences in pathophysiology and underlying biochemistry.\u003c/p\u003e \u003cp\u003eAmong biochemical differences between children with ESAM and those with NESAM, the cellular movement of methyl groups known as one-carbon metabolism (OCM) has been a particular recent focus. OCM plays a major role in health and disease, providing substrates for mitochondrial activity and ATP generation, dNTPs for DNA repair, and methyl groups needed for the maintenance of DNA and other cellular methylation events in mitotically active cells\u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e. The broad metabolic reach of OCM has been hypothesized as a potential unifying factor for the multisystem dysfunction seen in ESAM, and recent studies lend support to this. For instance, metabolomic studies have demonstrated that children with ESAM have significantly lower serum levels of the essential amino acid methionine\u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e with concomitant related effects, including dysregulation of asymmetric dimethyl arginine (ADMA), lower cysteine levels, impaired cysteine synthesis, reduced acylcarnitine production, and impaired fatty acid transport\u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e,\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Mice fed a maize-vegetable diet similar to that seen among SAM children in Malawi, develop a fatty liver phenotype that is similar to that seen in ESAM, which is abrogated by supplementation of the methyl-group donor choline\u003csup\u003e\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e\u003c/sup\u003e. Similarly, although there are no overt differences in microbiome composition between NESAM and ESAM children, a metagenomic study of severe malnutrition, conducted without regard to subtype, reported significantly lower OCM metabolite levels among gnotobiotic mice transplanted with stool from children with ESAM\u003csup\u003e\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eObservations of OCM dysfunction in ESAM are also consistent with stable isotope whole-body flux studies reporting slower methionine flux among children with ESAM\u003csup\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e. We have previously shown that during the acute nutritional stress, children with ESAM have significantly lower buccal cell DNA methylation than children with NESAM\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e, and this was consistent with observations of methionine turnover. Notably, and consistent with turnover studies conducted in children who had recovered from SAM, methylation differences were not observed among adults who had recovered from having either ESAM or NESAM as children, confirming that OCM dysfunction is a feature of the acute nutritional insult. Significantly hypomethylated loci were associated with genes both directly and indirectly implicated in the ESAM phenotype, and with loci associated with disordered nutrition states (e.g., obesity) and body anthropometry. Thus, collectively, there is a growing body of evidence supporting OCM dysfunction in ESAM. These findings, however, do not directly explain why only some children have the OCM derangements associated with ESAM, whilst others from the same region and environment have the OCM profile associated with NESAM.\u003c/p\u003e \u003cp\u003eAlthough there are no formal heritability estimates for ESAM, children with repeated bouts of SAM are more likely to develop the same form of SAM (i.e., either ESAM or NESAM) upon readmission to hospital\u003csup\u003e\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e,\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e; this observation, alongside the well-documented lack of demonstrable environmental, infectious, or metagenomic differences between ESAM and NESAM, has focused attention on the potential genetic risk of ESAM, which is consistent with recent interest in the genetics of malnutrition\u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. To date, however, attempts at genetic association have been limited to three single-candidate SNP association studies conducted in sample sizes of less than 200 individuals decades ago\u003csup\u003e\u003cspan additionalcitationids=\"CR39\" citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e. We thus sought to take a modern genetic approach to understanding the risk of developing ESAM, leveraging advances that have characterized genetic risk association studies of other public health disorders, including nutritional disorders like obesity. OCM is a fundamental metabolic pathway that has been genetically characterized\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e,\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e, with well-curated genes and loci mediating interindividual metabolic variation. Therefore, we started by evaluating the potential impact of genetic variation at OCM loci on the risk of developing ESAM. In this study, we compile a comprehensive list of OCM-associated genetic loci and evaluate genetic association with ESAM at these loci in population samples from Jamaica and Malawi (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). We then leverage population genetics and both local and global population ancestry, to further explore the association between OCM and ESAM.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"MATERIALS AND METHODS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eStudy participants\u003c/h2\u003e \u003cp\u003eDetails of participant samples used in the study have been previously published\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e,\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. All recruitment consisted of samples collected after individuals had received a diagnosis of either ESAM or NESAM.\u003c/p\u003e \u003cp\u003eSamples from Jamaica included DNA from blood and buccal samples from children (\u0026lt;\u0026thinsp;18 years) diagnosed with SAM. DNA was extracted using a phenol/chloroform protocol as described previously\u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e. SAM subtype was determined using the Wellcome Classification\u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e\u003c/sup\u003e, where NESAM (marasmus) is defined as \u0026lt;\u0026thinsp;60% weight-for-age without edema, and ESAM includes marasmic-kwashiorkor- defined as \u0026lt;\u0026thinsp;60% weight-for-age with edema, and kwashiorkor - defined as 60\u0026ndash;80% weight-for-age with edema. Pediatric participants were recruited prospectively from the Tropical Metabolism Research Unit of the Caribbean Institute for Health Research at the University Hospital of the West Indies (UHWI), the major tertiary nutritional center on the island. Recruitment was part of a larger, long-running study of the genetic etiology of SAM. A cohort of adult participants (\u0026gt;\u0026thinsp;18 years) who formerly had SAM as children were recruited from the same site. The study was approved by the Ethics Committees of the UHWI/University of the West Indies Faculty of Medical Sciences.\u003c/p\u003e \u003cp\u003eIn Malawi, samples from children with SAM were recruited from 18 rural sites across five districts\u003csup\u003e\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. Samples were originally collected between 2013\u0026ndash;2016. The sub-type of SAM was defined using mid-upper arm circumference (MUAC) and the presence or absence of edema. Buccal samples were collected using the Oragene Discover (OGR-250) DNA collection kit (DNA Genotek Inc., Ottawa, Ontario, Canada). The protocol was modified to collect buccal epithelial cells by swabbing the inside of each cheek ten times. DNA was extracted following the manufacturer's instructions. After collection, to maintain clinical uniformity in SAM designation between Jamaica and Malawi, samples from Malawi were re-classified using the Wellcome Classification as having either ESAM or NESAM, which was highly concordant with the MUAC based designation. The study was approved by the National Health Science Review Committee of the Ministry of Health, Government of Malawi.\u003c/p\u003e \u003cp\u003e Written informed consent was obtained from the parents or adult guardians of children included in this study. Permission to use de-identified participant samples for genetic studies at Baylor College of Medicine (BCM) was approved by the Institutional Review Board (IRB) of BCM and confirmed by the National Institutes of Health IRB subsequently.\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eCuration of OCM loci\u003c/h3\u003e\n\u003cp\u003eGeneWeaver\u003csup\u003e\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e, AmiGO 2\u003csup\u003e46,47\u003c/sup\u003e, GWAS Catalog\u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e,\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e, KEGG\u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e, NCBI BioSystems\u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e, Pathway Commons\u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e, and PubMed were queried using the search terms \u0026ldquo;one-carbon metabolism\u0026rdquo;, \u0026ldquo;cysteine\u0026rdquo;, \u0026ldquo;methionine\u0026rdquo;, and \u0026ldquo;one-carbon pool by folate\u0026rdquo; to identify OCM loci to be used in this study. At the genome level, a 50 kilobase (kb) buffer region was included on either side of each identified gene to capture potential \u003cem\u003ecis\u003c/em\u003e-acting elements. This search resulted in a list of 103 genes involved in OCM. Overlapping regions were consolidated into single regions, resulting in 95 loci (\u003cb\u003eSupplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e\n\u003ch3\u003eGenotyping and quality control\u003c/h3\u003e\n\u003cp\u003ePeripheral blood and buccal DNA samples from Jamaican and Malawian individuals were genotyped for this study. Downstream quality control and analyses are shown in \u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e. Genotyping was performed on the Infinium H3Africa Consortium Array v1.1 (\u003cb\u003eweb resources\u003c/b\u003e), which provides LD coverage of ~\u0026thinsp;0.89 in the African superpopulation of 1000 Genomes. Raw idat files were processed in GenomeStudio (Illumina, San Diego, California, USA), including within-sample clustering, to create binary input files for PLINK. Genotype fluorescence projections in GenomeStudio were visually inspected for those SNPs that surpassed test-wide significance in order to confirm genotype calls. Downstream SNP and sample quality control (QC) was performed in PLINK v1.07\u003csup\u003e53\u003c/sup\u003e and included limiting to biallelic autosomal SNPs and excluding SNPs with a missingness\u0026thinsp;\u0026gt;\u0026thinsp;0.95; removal of SNPs outside the Hardy-Weinberg test distribution; and excluding samples with missingness\u0026thinsp;\u0026gt;\u0026thinsp;10%. Linkage disequilibrium (LD) pruning was conducted using a window size of 50 bases, a sliding window of ten bases, and r\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e threshold of 0.1. Pairs of samples with a proportion-of- identity-by-descent (IBD) (PI_HAT) score\u0026thinsp;\u0026gt;\u0026thinsp;0.1 were flagged, and the individual within the pair with the higher missingness rate was removed. Post QC, 1.8\u0026nbsp;million genotyped SNPs were available for downstream analysis across the 711 remaining samples.\u003c/p\u003e\n\u003ch3\u003eImputation\u003c/h3\u003e\n\u003cp\u003eGiven the relative lack of representation of African genomes in existing imputation panels, we chose to conduct a two-stage imputation and data merge process. Data from Jamaican individuals (n\u0026thinsp;=\u0026thinsp;340) were imputed using the Consortium on Asthma among African-ancestry Populations in the Americas (CAAPA) reference panel\u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e, which was generated using data from individuals with predominantly west-African ancestry, including Jamaicans. We used minimac3 v1.0.5 (University of Michigan, Ann Arbor, Michigan, USA)\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e and Eagle v2.3 (Broad Institute, Cambridge, Massachusetts, USA)\u003csup\u003e\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e56\u003c/span\u003e\u003c/sup\u003e on the Michigan Imputation Server (University of Michigan, Ann Arbor, Michigan, USA)\u003csup\u003e\u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e55\u003c/span\u003e\u003c/sup\u003e to perform phasing and imputation on our Jamaican samples, selecting the African reference panel (\"AFR\") as the reference population.\u003c/p\u003e \u003cp\u003eData from Malawian participants (n\u0026thinsp;=\u0026thinsp;371) was imputed using the H3Africa reference panel (H3ABioNet)\u003csup\u003e\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e57\u003c/span\u003e\u003c/sup\u003e (\u003cb\u003eweb resources\u003c/b\u003e), which is designed to represent a wide array of African ancestries, including east- and central-African groups that neighbor Malawi. Phasing was performed using Eagle 2.4\u003csup\u003e56\u003c/sup\u003e (Broad Institute, Cambridge, MA, USA). Minimac4 Version 1.0.0 (University of Michigan, Ann Arbor, MI, USA) was used for imputation on the H3ABioNet Nextflow chip imputation pipeline\u003csup\u003e\u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e58\u003c/span\u003e\u003c/sup\u003e, which implements a similar workflow to the Michigan Imputation Service pipeline, and was deployed on the University of Cape Town High-Performance Computing Facility (web resources).\u003c/p\u003e \u003cp\u003eAfter imputation, SNPs with r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.3 were excluded. The imputed data from both population groups was merged using an in-house command line script (\u003cb\u003eweb resources\u003c/b\u003e), prior to removing non-autosomal SNPs and SNPs with minor allele frequency (MAF)\u0026thinsp;\u0026lt;\u0026thinsp;0.05. The final imputed dataset included 8,415,377 SNPs, of which 45,411 SNPs were at OCM loci (+/-50kb of genes).\u003c/p\u003e\n\u003ch3\u003eMultidimensional scaling\u003c/h3\u003e\n\u003cp\u003eGenome-wide SNPs were used to perform multidimensional scaling (MDS) in PLINK v1.07, and these were used for two different purposes (see \u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e): 1. - on the QC-ed genotype dataset used for association analyses; 2. on the genotype-only dataset alongside datasets from 1000 Genomes\u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e to put the primary populations in the context of African datasets (See Population ancestry and admixture mapping below). 1000 Genomes datasets consisted of Esan individuals in Nigeria (ESN, n\u0026thinsp;=\u0026thinsp;99), individuals from Gambia in Western Division, the Gambia (GWD, n\u0026thinsp;=\u0026thinsp;113), Luhya individuals in Webuye, Kenya (LWK, n\u0026thinsp;=\u0026thinsp;99), Mende individuals in Sierra Leone (MSL, n\u0026thinsp;=\u0026thinsp;85), and Yoruba individuals in Ibadan, Nigeria (YRI, n\u0026thinsp;=\u0026thinsp;108). The genotyped-only dataset was merged with the above 1000 Genomes populations, and LD-pruned with a window size of 50 bases, a window of 10 bases, and r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.1. MDS was performed PLINK and visualized using an in-house R\u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e script (\u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003eGenetic association analyses\u003c/h2\u003e \u003cp\u003eAssociation statistics were primarily derived using Genome-Wide Efficient Mixed Model Association (GEMMA)\u003csup\u003e\u003cspan citationid=\"CR61\" class=\"CitationRef\"\u003e61\u003c/span\u003e\u003c/sup\u003e, which better accounts for genetic ancestry variation. Mixed model association analysis in GEMMA was performed using the -lm 4 flag. Standardized relatedness of the combined genotyped and imputed dataset was also calculated using the flag -gk 2. We used the genetic data to estimate the standardized kinship matrix before applying the generalized linear mixed effect model to assess genetic association between SNPs and the risk of developing ESAM, assuming random effect on individuals and fixed effects due to sex and the first component of MDS (as suggested by scree plots of the proportion of variance explained; \u003cb\u003eSupplementary Figure S3\u003c/b\u003e). In the regional regression analyses, the MDS component was estimated from the dataset of the specified cohort. QQ plots were generated using the R package qqman\u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e. Manhattan plots were generated using an in-house script (\u003cb\u003eweb resources\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eCumulative association\u003c/h3\u003e\n\u003cp\u003eWe employed a permutation test procedure to assess the cumulative association between OCM loci and SAM status; this is a non-parametric method of assessing the statistical significance of observed data without making assumptions about the underlying probability distribution of the data. We hypothesized that OCM loci are enriched for SNPs with p-values below a certain significance threshold, i.e., the proportion of significant loci within OCM loci is more than expected by chance. We choose p\u0026thinsp;=\u0026thinsp;0.001 as a representative significance threshold for the number of loci being tested (similar to a Bonferroni correction for multiple testing of 0.1/100). To test the robustness of our findings, we included a ceiling threshold of p\u0026thinsp;=\u0026thinsp;0.01, and two equidistance points (0.001 and 0.01) as sensitivity points.\u003c/p\u003e \u003cp\u003eThe test procedure involved taking the genome-wide association statistics for all genotyped SNPs (n\u0026thinsp;=\u0026thinsp;1,843,153) and calculating the proportion of SNPs with a p-value meeting a pre-specified threshold. We then randomly sampled a subset of SNPs equivalent to the number of genotyped SNPs across OCM loci (n\u0026thinsp;=\u0026thinsp;9,479) and determined the proportion of those SNPs with a p-value meeting the predefined threshold. That step was repeated 10,000 times to derive an empirical distribution of the proportion of significant SNPs. We then calculated a z-score for the observed proportion based on the empirical distribution. To account for the randomness of genome-wide SNPs relative to a restricted metabolic pathway, we performed the same test using a similar number of SNPs sampled from loci involved in ketone metabolism.\u003c/p\u003e\n\u003ch3\u003eHaplotype association and calculation of linkage disequilibrium\u003c/h3\u003e\n\u003cp\u003eGenotyped SNPs in \u003cem\u003eGABBR2\u003c/em\u003e (n\u0026thinsp;=\u0026thinsp;5) and \u003cem\u003ePRICKLE2\u003c/em\u003e (n\u0026thinsp;=\u0026thinsp;6) reaching region-level significance (Fig.\u0026nbsp;2) were arranged in 5' to 3' order and phased in PLINK v1.07 to create basic region-level haplotypes. PLINK was also used to determine association using the --hap-assoc flag. PLINK files were also parsed to Haploview\u003csup\u003e\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e63\u003c/span\u003e\u003c/sup\u003e in order to generate plots of linkage disequilibrium.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003ePopulation ancestry and admixture mapping\u003c/h2\u003e \u003cp\u003eTo investigate the ancestral African genetic ancestries in the Jamaica and Malawi datasets, genome-wide genotype data was merged with selected population data from publicly available datasets using PLINK v1.9 (\u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e). Available datasets included the African Genome Variation Project (AGVP)\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e, H3Africa Consortium\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e, Schlebusch \u003cem\u003eet al.\u003c/em\u003e\u003csup\u003e\u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e66\u003c/span\u003e\u003c/sup\u003e, and 1000 Genomes Project (Phase III). Only SNPs common to all datasets were retained. As genetic reference data is not available for every African ethnolinguistic group, we sought to ensure that we could assess contributions from Niger-Congo speaking groups, which includes Bantu speakers - the largest ethnolinguistic grouping across the continent, as well as non-Niger-Congo groups, such Rain Forest Forager, Hunter-Gatherer (e.g. Khoe and San), and Nilo-Saharan groups, which are more populous in east- and central- Africa. Populations representing Niger-Congo African ancestry included Bantu speakers from Zambia (BSZ, n\u0026thinsp;=\u0026thinsp;39), Jola (n\u0026thinsp;=\u0026thinsp;79), and YRI (n\u0026thinsp;=\u0026thinsp;100). Populations representing non-Niger-Congo African ancestry included Amhara (n\u0026thinsp;=\u0026thinsp;42), Botswana (BOT, n\u0026thinsp;=\u0026thinsp;48), LWK (n\u0026thinsp;=\u0026thinsp;74), Khoe and San Hunter Gatherers (n\u0026thinsp;=\u0026thinsp;18), Zulu (n\u0026thinsp;=\u0026thinsp;95), and MSL (n\u0026thinsp;=\u0026thinsp;85). Samples from CEU (n\u0026thinsp;=\u0026thinsp;95) were included to represent European admixture. The merged dataset was further filtered to remove SNPs with excessive missingness (\u0026gt;\u0026thinsp;0.1) and then LD-pruned (see MDS above). The final merged genotype dataset included\u0026thinsp;~\u0026thinsp;50,000 autosomal SNPs. The resulting merged dataset was used to perform MDS for a subset of groups using the same parameters and procedures as above.\u003c/p\u003e \u003cp\u003eUsing the unsupervised clustering algorithm as implemented by ADMIXTURE v1.3\u003csup\u003e67\u003c/sup\u003e, global ancestry proportions were inferred for all populations. Since the algorithm is known to be impacted by the number of individuals that represent a population, 100 random individuals each from Jamaica and Malawi were retained to approximate the sample sizes of the reference populations. ADMIXTURE was performed for K\u0026thinsp;=\u0026thinsp;2 to K\u0026thinsp;=\u0026thinsp;6, and the K with the lowest cross-validation error estimates was noted (K\u0026thinsp;=\u0026thinsp;5).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003eLocal ancestry inference\u003c/h2\u003e \u003cp\u003e \u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003eB\u003c/b\u003e illustrates the data processes for local ancestry. Data subset from 1000 Genomes consisting of CEU, MSL, and LWK individuals was used as references for European, West African, and East African ancestry, respectively. The SAM cohort dataset and the 1000 Genomes dataset each underwent independent QC filtering using MAF\u0026thinsp;\u0026gt;\u0026thinsp;0.05; SNP missingness threshold of 0.05; and removal of duplicate SNPs. The remaining SNPs were merged using an in-house script. The merged dataset was subject to additional QC filtering to remove A\u0026thinsp;\u0026gt;\u0026thinsp;T and G\u0026thinsp;\u0026gt;\u0026thinsp;C SNPs. Datasets were then phased by chromosome using ShapeIt2\u003csup\u003e68\u003c/sup\u003e, using the 1000 Genomes Phase 3 genetic map with default parameters.\u003c/p\u003e \u003cp\u003eLocal ancestry inference scores were calculated for the genotyped-only dataset using RFMix version 2\u003csup\u003e69\u003c/sup\u003e. Local ancestry for the reference 1000 Genomes dataset was run separately from the SAM cohort dataset to maximize the available data generated on differing platforms, and 1000 Genomes-specific indices were used as reference ancestry groups. The 1000 Genomes Phase 3 genetic map was used for RFMix, which was otherwise run with default parameters.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003eAncestry-adjusted association analyses\u003c/h2\u003e \u003cp\u003eIn ancestry-adjusted association tests, we employed TRACTOR v1.1.0\u003csup\u003e70\u003c/sup\u003e. SNP association with ESAM disease risk was performed whilst accounting for the ancestral origin of the SNP by first assigning ancestry labels to the two haplotypes of individuals, referred as tracts. Tracts were extracted from the genotyped-only dataset using ExtractTracts.py from TRACTOR, with integers from one to three referring to the three proxy ancestries - CEU, LWK, and MSL - from 1000 Genomes. RunTractor.py was then run to perform logistic regression using NESAM vs ESAM status as the outcome for each individual, with tracts, ancestry-specific allele counts, and sex as regressors. OCM loci were then subset from the regression results and visualized\u003csup\u003e\u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e62\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eCell-free DNA genotyping of serum metabolomics samples\u003c/h2\u003e \u003cp\u003eCell-free DNA (cfDNA) was obtained from serum samples of 415 individuals from Malawi. These individuals were independently recruited as part of our published metabolomics study of one-carbon metabolites in SAM\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. Samples were initially stored at -80\u0026deg;C and subsequently slowly brought up to room temperature to minimize degradation. CfDNA was extracted using the Qiagen QIAamp MinElute Virus Spin Kit (Cat #57704), which is optimized for cfDNA isolation, and yielded a mean concentration of 2.7ng/\u0026micro;L of genomic DNA per sample, with concentrations ranging from 0.1 ng/\u0026micro;L to 52.0 ng/\u0026micro;L.\u003c/p\u003e \u003cp\u003eWe attempted to amplify six ESAM-associated SNPs from cfDNA samples \u0026ndash; two in \u003cem\u003ePRICKLE2\u003c/em\u003e (rs753562 and rs17664202), two in \u003cem\u003eGABBR2\u003c/em\u003e (rs17664203 and rs17664204) and one each in \u003cem\u003eSHMT1\u003c/em\u003e (rs651495) and \u003cem\u003ePLD2\u003c/em\u003e (rs1052748). We were unable to obtain consistent or robust genotype data for rs651495 (\u003cem\u003eSHMT1\u003c/em\u003e) and rs17664203 (\u003cem\u003eGABBR2\u003c/em\u003e) and results from these assays were not included in downstream analyses. Primers for the remaining SNPs were designed using an \u003cem\u003ein silico\u003c/em\u003e primer design tool (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/tools/primer-blast/\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/tools/primer-blast/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) using the SNP ID (rsID) in GRCh38/hg38 (\u003cb\u003eSupplementary Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e). Primer specificity was confirmed using BLAT. Primers were synthesized by IDT (Coralville, IA, USA) and prepared as 100 \u0026micro;M stock solutions in TE buffer. PCR amplification was performed using 4ng of DNA under a single thermocycling and amplification protocol (\u003cb\u003eSupplementary Methods\u003c/b\u003e), with products verified via gel electrophoresis (GelDoc, BioRad laboratories, Hercules, CA, USA). Following PCR, products were purified using Exo-Sap IT (Thermo Fisher Scientific, Waltham MA, USA) to remove excess primers and dNTPs. Dideoxy- (Sanger) sequencing was conducted using M13-tagged primers and the Big Dye v3.1 sequencing kit (Thermo Fisher Scientific, Waltham MA, USA). Sequencing reactions were cleaned using the Qiagen DyeEx kit (Qiagen, Germantown, MD, USA) to remove dye terminators, followed by genotype calling on the SeqStudio (Genetic Analyzer #A33770, Fisher Scientific, Waltham MA, USA). Sequencher (software version 4.10.1) was used to visualize and confirm genotype calls.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eMendelian randomization (MR) using serum metabolites\u003c/h2\u003e \u003cp\u003eFollowing the same classification criterion used in our primary cohort, we selected 252 samples with diagnoses of ESAM (n\u0026thinsp;=\u0026thinsp;135) and NESAM (n\u0026thinsp;=\u0026thinsp;117) for whom we had cfDNA genotypes \u003cem\u003eand\u003c/em\u003e metabolomic data for all 16 OCM metabolites - choline, betaine, dimethylglycine (DMG), glycine, sarcosine, 5-methyltetrahydrofolate (MTHF), serine, methionine, S-adenosylmethionine (SAMe), S-adenosylhomocysteine (SAH), homocysteine, cysteine, cystathionine, pyridoxal phosphate (PLP) and asymmetric dimethylglycine (ADMA). The remaining samples were either community controls or had a diagnosis of moderate acute malnutrition.\u003c/p\u003e \u003cp\u003eMendelian randomization (MR) was used to assess the causal effect of individual metabolites using the genotypes as the instrument variables. For this analysis, we performed two-sample MR based on multiple SNPs by combining results from the SNP-metabolite and the SNP-outcome associations to show potential \u003cem\u003ein vivo\u003c/em\u003e effects. We applied one instance of the inverse-variance approach (IVW) from MendelianRandomization R package. Under this approach, the causal effect is obtained from a weighted linear regression of the associations with the outcome, on the associations with the metabolite, while fixing the intercept to zero and weights being the inverse-variances of the associations with the outcome\u003csup\u003e\u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e71\u003c/span\u003e\u003c/sup\u003e. We also exploited the MR-Egger method to detect invalid instrument variables and bias\u003csup\u003e\u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e72\u003c/span\u003e\u003c/sup\u003e. The serum metabolites were highly correlated (R2). Our cumulative association data suggested that there might be synergistic effects among SNPs; therefore, we assumed an additive genetic model for test alleles, and then tested for the causal effect on ESAM risk of each metabolite, after log10 transformation (to better approximate normality). We selected three genotyped SNPs as instrument variables, one SNP from each locus, with SNP rs753562 in gene \u003cem\u003ePRICKLE2\u003c/em\u003e being omitted because the serum genotype data was slightly out of Hardy-Weinberg equilibrium and the SNP is physically close to the other \u003cem\u003ePRICKLE2\u003c/em\u003e SNP. The estimates of SNP-outcome association are drawn from the GEMMA analysis of the Malawi group in the primary cohort with fixed effects being sex and the first component of MDS. The estimates of SNP-metabolite association are drawn from a linear regression of log\u003csub\u003e10\u003c/sub\u003e(metabolite) with sex and age as covariates.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eWhole genome sequencing\u003c/h2\u003e \u003cp\u003eA subset of 151 DNA samples from Malawi were chosen to undergo whole-genome sequencing to assess recent positive selection in OCM-related genomic regions. These samples were selected to have an approximately equal distribution of males and females and an equal number of ESAM and NESAM samples to minimize sampling bias. WGS was performed by first generating PCR-free libraries from one microgram genomic DNA using the TruSeq\u0026reg; DNA PCR-Free HT Sample Preparation Kit (Illumina, San Diego, CA, USA). The median insert sizes were approximately 400 bp. Libraries were tagged with unique dual-index DNA barcodes to allow the pooling of libraries and to minimize the impact of barcode hopping. Libraries were then pooled for sequencing on the NovaSeq 6000 (Illumina) to obtain at least 300\u0026nbsp;million 151-base read pairs per individual library. The sequence data were then aligned to human_g1k_v37_decoy.fasta via BWA-MEMv0.7.15\u003csup\u003e73\u003c/sup\u003e. GATK Resource Bundle B37\u003csup\u003e74\u003c/sup\u003e was used to sort the data, mark duplicates, perform BQSR, and run the GATK Haplotype Caller to generate gVCF files on each sample. Then individual gVCFs were merged into a multi-sample VCF file using GATK 4.3.0.0. GATK Best Practice\u003csup\u003e\u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e75\u003c/span\u003e\u003c/sup\u003e for variant calling and filtration was followed to obtain high-quality variant data, i.e., applying GATK VQSR function, followed by harder filters (filter parameters set as DP\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;10, QUAL\u0026thinsp;\u0026gt;\u0026thinsp;=\u0026thinsp;60, and ExcessHet\u0026thinsp;\u0026lt;\u0026thinsp;=\u0026thinsp;54.69).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003eSignatures of selection\u003c/h2\u003e \u003cp\u003eWe used the iHS score\u003csup\u003e\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e76\u003c/span\u003e\u003c/sup\u003e to identify genomic signatures of selection in the dataset; this analysis is agnostic of phenotype and covariates. To compare signatures of selection at OCM loci between different countries, we created subsets of country-specific signatures of selection generated by the H3Africa consortium\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e, which included selection statistics for SNPs across the OCM loci used in the association study: Benin (n\u0026thinsp;=\u0026thinsp;45,537 SNP scores), Botswana (BOT, n\u0026thinsp;=\u0026thinsp;44,277 SNP scores), Cameroon (CAM, n\u0026thinsp;=\u0026thinsp;44,458 SNP scores), Mali (MAL, n\u0026thinsp;=\u0026thinsp;43,033 SNP scores), Berom from Nigeria (BRN, n\u0026thinsp;=\u0026thinsp;44,177 SNP scores), Gur speakers from West Africa (GWR, n\u0026thinsp;=\u0026thinsp;43,523 SNP scores), and Zambia (ZAM, n\u0026thinsp;=\u0026thinsp;47,572 SNP scores). These samples were sequenced on Illumina HiSeq 2500 and subsequently processed and mapped to GRCh37. Integrated haplotype scores (iHS) were calculated using selscan\u003csup\u003e\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e77\u003c/span\u003e\u003c/sup\u003e. We calculated iHS scores for the Malawi WGS dataset using the same QC and run parameters as the H3Africa datasets. SNPs with iHS scores \u0026gt; |2| were considered outliers, potentially under selection. Selection signatures were visualized using R\u003csup\u003e\u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e60\u003c/span\u003e\u003c/sup\u003e and ggplot2\u003csup\u003e78\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS","content":"\u003cp\u003eOur primary analyses utilized DNA samples from 833 individuals from Jamaica and Malawi approximately evenly split between ESAM (case/affected group) and NESAM (reference/control group) at each location (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e; \u003cb\u003eSupplementary Table S3\u003c/b\u003e). Samples were genotyped on the Illumina H3Africa array, with subsequent quality control (QC) processing resulting in 711 samples (\u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e) and 1,639,325 genotyped SNPs (\u003cb\u003eMethods\u003c/b\u003e). Among genotyped SNPs, 9,479 SNPs fell within 10kb (+/-) of a curated list of 103 autosomal genes and loci previously associated with OCM (\u003cb\u003eMethods; Supplementary Table S4\u003c/b\u003e); this included four gene clusters with two or more genes in tandem. After variant imputation and quality control (\u003cb\u003eMethods\u003c/b\u003e), a set of 45,411 imputed and 9,479 genotyped SNPs at 95 OCM loci were identified to test for association and perform downstream analyses (\u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e).\u003c/p\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eVariation at OCM loci is associated with ESAM\u003c/h2\u003e \u003cp\u003eWe first assessed the association between imputed SNPs at OCM loci and ESAM using the linear mixed model implemented in GEMMA (see \u003cb\u003eMethods\u003c/b\u003e). The top associated SNP (rs79824961 near \u003cem\u003eGABBR2\u003c/em\u003e) surpassed our multiple-testing adjustment for the total number of SNPs interrogated (p\u0026thinsp;\u0026lt;\u0026thinsp;9.73x10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e) (Fig.\u0026nbsp;2). A further 56 SNPs at nine OCM loci, surpassed a secondary significance threshold adjusted for the number of loci (n\u0026thinsp;=\u0026thinsp;95) tested (p\u0026thinsp;\u0026lt;\u0026thinsp;5.26x10\u003csup\u003e\u0026minus;\u0026thinsp;4\u003c/sup\u003e). This latter threshold was supported by an excess of observed p-values without evidence of test-wide inflation (λ\u0026thinsp;=\u0026thinsp;1.02) on the resulting QQ plot (Fig.\u0026nbsp;2). To further reduce potential false positives, we characterized candidate loci as those with multiple SNPs surpassing our locus-level cut-off or single variants with evidence of association using a secondary method (n\u0026thinsp;=\u0026thinsp;3 loci; \u003cb\u003eMethods)\u003c/b\u003e. This resulted in a final set of seven ESAM-associated OCM loci (\u003cb\u003eSupplementary Table S4\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eThe strongest associated locus was an intragenic region on chromosome 9q22.33 (\u003cb\u003eSupplementary Figure S4A\u003c/b\u003e) falling within the first intron of gamma-aminobutyric acid type B receptor subunit two (\u003cem\u003eGABBR2\u003c/em\u003e; top SNP rs7038285, p\u0026thinsp;=\u0026thinsp;6.98x10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e, Odds Ratio (OR)\u0026thinsp;=\u0026thinsp;0.85), where the minor allele (C) was enriched among individuals with NESAM. SNPs within intron 13 of \u003cem\u003eGABBR2\u003c/em\u003e (~\u0026thinsp;260kb upstream of our top SNPs) have been associated with lower plasma homocysteine\u003csup\u003e\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e\u003c/sup\u003e in a European population. The mechanistic relationship of this association to the \u003cem\u003eGABBR2\u003c/em\u003e gene, which is a member of the G-protein coupled receptor 3- and GABA-B receptor families, are unclear. We also found similar evidence of putative association at \u003cem\u003ePRICKLE2\u003c/em\u003e on chromosome 3p14.1 (top SNP rs11130959, p\u0026thinsp;=\u0026thinsp;4.15x10\u003csup\u003e\u0026minus;\u0026thinsp;5\u003c/sup\u003e, OR\u0026thinsp;=\u0026thinsp;0.88) \u003cb\u003e(Supplementary Figure S4B\u003c/b\u003e), with the minor allele also being enriched among NESAM participants. The canonical function of \u003cem\u003ePRICKLE2\u003c/em\u003e is unknown, although it has been associated with variation in folate levels\u003csup\u003e\u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e79\u003c/span\u003e\u003c/sup\u003e. Of the seven top candidate loci, five (\u003cem\u003eMTHFR\u003c/em\u003e, \u003cem\u003eAHCYL1\u003c/em\u003e, \u003cem\u003ePRICKLE2\u003c/em\u003e, \u003cem\u003eGABBR2\u003c/em\u003e, and \u003cem\u003ePLD2\u003c/em\u003e) directly impact either conversion to- or breakdown of- homocysteine (Fig.\u0026nbsp;2), including a coding variant rs1052748 (p.Thr577Ile) in exon 17 of \u003cem\u003ePLD2\u003c/em\u003e that was enriched among ESAM individuals \u003cb\u003e(Supplementary Figure S4C\u003c/b\u003e). \u003cem\u003ePLD2\u003c/em\u003e encodes phospholipase D2 phosphatidyl, which generates the methyl-group donor choline through the hydrolysis of phosphatidylcholine. The rs1052748 variant is reported as an expression quantitative trait locus (eQTL) for \u003cem\u003ePLD2\u003c/em\u003e in multiple tissues in the Gene-Tissue Expression database (GTEx)\u003csup\u003e\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e80\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eMetabolic flux results from cumulative enzymatic activity at multiple points within a biochemical cycle. Given our observation of multiple putatively associated OCM SNPs, each with variable effects on OCM, we considered the potential for a cumulative effect of OCM-associated SNPs on ESAM susceptibility. To test this, we applied a permutation test (\u003cb\u003eMethods\u003c/b\u003e) to the association statistics of OCM SNPs and compare this with the association tests of the ~\u0026thinsp;1.8\u0026nbsp;million genotyped SNPs genome wide (for which no SNP surpassed genome-wide significance). At a p-value cut-off of \u0026lt;\u0026thinsp;0.001, a greater proportion of SNPs at OCM loci were associated with ESAM relative to the randomly sampled dataset ((z\u0026thinsp;=\u0026thinsp;3.2; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). At more permissive p-value thresholds, the z-score differentiating our OCM SNPs and the random sampling experiment was even more pronounced (e.g., at a threshold of p\u0026thinsp;=\u0026thinsp;0.01, the resulting z-score was 50 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003ec)).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eNext, we considered whether our cumulative association might have resulted from comparing SNPs in linkage disequilibrium (LD) at metabolically related loci to a random distribution of largely unlinked SNPs. To do so, we repeated the permutation test, but this time randomly sampling the same number of SNPs (n\u0026thinsp;=\u0026thinsp;9,479) from the 51,105 SNPs found at loci related to ketone metabolism, which is not known to be associated with SAM (\u003cb\u003eMethods\u003c/b\u003e). Relative to this ketone metabolism background, the proportion of OCM SNPs showing association was still in the upper tail of the distribution (z\u0026thinsp;=\u0026thinsp;3.19) at the p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 threshold (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003eb). Finally, as a negative control, we performed the same genome-wide comparison as above for SNPs at sphingolipid metabolism loci, which has a similar number of SNPs to OCM, but no \u003cem\u003ea priori\u003c/em\u003e evidence for involvement in ESAM. The proportion of sphingolipid metabolism SNPs showing association with ESAM fell well within the randomly generated distribution (z= -0.70).\u003c/p\u003e \u003cp\u003eThere are few established models for evaluating genetic variation in SAM. Therefore, we sought to provide support for the proposed functional impact of associated variants by assessing the impact of ESAM-associated SNPs on OCM metabolites in the context of ESAM vs NESAM. We first extracted cell-free DNA (cfDNA) from ~\u0026thinsp;400 serum samples (M\u003cb\u003eethods\u003c/b\u003e) from our previous study of OCM metabolites in ESAM/NESAM, which was done in an independent cohort of children from Malawi recruited around the time of their SAM diagnosis\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e. We then genotyped four ESAM-associated SNPs (rs753562 (\u003cem\u003ePRICKLE2\u003c/em\u003e), rs17664202 (\u003cem\u003ePRICKLE2)\u003c/em\u003e, rs17664204 (\u003cem\u003eGABBR2\u003c/em\u003e), and rs651495 (\u003cem\u003ePLD2\u003c/em\u003e) using cfDNA samples (\u003cb\u003eMethods; Supplementary Methods\u003c/b\u003e). Three SNPs (rs17664202 was out of Hardy-Weinberg equilibrium) were then used as instrumental variables in a two-sample Mendelian Randomization (MR) to assess the causal effect of each of the 16 OCM metabolites on ESAM risk (\u003cb\u003eMethods\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eWe first assessed SNP-metabolite and SNP-disease associations separately under regression models and then combined effects using the IVW method assuming additive genetic effects. Although OCM metabolites are highly correlated, we found significant casual effects on cystathionine (log-odds ratio\u0026thinsp;~\u0026thinsp;2.49, p\u0026thinsp;=\u0026thinsp;1.65x10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e) and on betaine (log-odds ratio\u0026thinsp;~\u0026thinsp;4.69, p\u0026thinsp;=\u0026thinsp;0.0007) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e4\u003c/span\u003e). With three SNPs, there was substantial genotypic heterogeneity in the causal effect of betaine and cysteine (p\u0026thinsp;\u0026lt;\u0026thinsp;0.001 by Cochran's Q test for both). We thus reevaluated the model after excluding any SNPs with opposing effects (although such effects could reflect the effect of the major allele), which necessarily limited our analyses to the IVW approach. Using the two SNPs with congruent effects, the causal effect of cysteine was estimated ~ -7.97, (p\u0026thinsp;=\u0026thinsp;2.24x10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e and Cochran\u0026rsquo;s Q test p\u0026thinsp;=\u0026thinsp;0.74) and the estimate for betaine was 5.41, (p\u0026thinsp;=\u0026thinsp;0.0001 and Cochran\u0026rsquo;s Q test p\u0026thinsp;=\u0026thinsp;0.005). Although we did not find strong evidence for cystathionine or betaine under the MR-Egger approach, we found that estimates from both approaches were similar in direction and magnitude.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003ePopulation ancestry influences the association of OCM loci with ESAM\u003c/h2\u003e \u003cp\u003eIn the absence of an available secondary cohort in which to replicate our findings, we also sought evidence of internal consistency in the magnitude and direction of effect between the two countries represented in our cohort. At our seven candidate loci, MAFs and directions of effect among associated SNPs were similar between the two countries (\u003cb\u003eSupplementary Table S5\u003c/b\u003e). Despite similarities in the direction of effect in the two countries, we noted that the magnitude of the effect was generally stronger in Malawi than Jamaica despite comparable sample sizes and minor allele frequencies (\u003cb\u003eSupplementary Figure S5\u003c/b\u003e). This difference was also evident when we looked at statistically inferred haplotypes (\u003cb\u003eMethods\u003c/b\u003e) comprised of associated (genotyped) SNPs at our two top loci \u0026ndash; \u003cem\u003eGABBR2\u003c/em\u003e and \u003cem\u003ePRICKLE2\u003c/em\u003e. We observed evidence of association with ESAM in each population group; however, the haplotypes with the strongest implied effect differed between the two population groups despite similar patterns of LD in the two groups at both loci (\u003cb\u003eSupplementary Figure S6\u003c/b\u003e). Similarly, when we stratified our cumulative association analyses by population group, there was always a higher proportion of association SNPs in Malawi (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e3\u003c/span\u003ec).\u003c/p\u003e \u003cp\u003eTo explore this further, we first considered our populations in multidimensional scaling (MDS) space. Jamaican and Malawian samples clustered with other African continental populations on MDS components one and two, although samples from Malawi were tightly clustered alongside other Bantu-speaking Niger Congo ethnolinguistic groups (YRI, MSL). Conversely, Jamaican samples were more dispersed (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003ea; \u003cb\u003eSupplementary Figure \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e\u003c/b\u003e), likely a result of the diverse population ancestry origins of Jamaicans, which include a predominant contribution of African ancestry from the British slave trade of the 14th to 17th centuries, as well as a subsequent influx of indentured peoples from East- and Southeast- Asia, and Europe, in addition to centuries of European (British and Spanish) colonialization.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eWe then used ADMIXTURE\u003csup\u003e\u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e67\u003c/span\u003e\u003c/sup\u003e to infer patterns of genome-wide admixture in our cohort in the context of parental proxies of geographic and continental differentiation derived from the African Genome Variation Project (AGVP)\u003csup\u003e\u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e64\u003c/span\u003e\u003c/sup\u003e, the Human Health and Heredity in Africa (H3Africa) Project\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e, and 1000 Genomes (1KG) Project\u003csup\u003e\u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e59\u003c/span\u003e\u003c/sup\u003e (\u003cb\u003eMethods\u003c/b\u003e). In this analysis, we proxied European ancestry through the Ceph from Utah (CEU); West African ancestries through the Yoruba from Nigeria (YRI) and Mende from Sierra Leone (MSL); East African ancestries through the Luhya from Webuye, Kenya, and the Amhara from Ethiopia; and southern African ancestries through individuals from Botswana (BOT), Bantu-speakers from Zambia (BSZ), the Zulu, and Khoe and San Hunter-Gatherer groups (KS). As expected, samples from Jamaica had, on average, a significantly higher proportion of European (~\u0026thinsp;13%) and west African ancestry (~\u0026thinsp;68%) than Malawians, in whom east-African ancestries (~\u0026thinsp;83%) were more common (Welch two-sample t-test, p\u0026thinsp;=\u0026thinsp;2.2x10\u003csup\u003e\u0026minus;\u0026thinsp;16\u003c/sup\u003e) (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003eb; \u003cb\u003eSupplementary Figure S7\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eAssociation between OCM loci and ESAM is driven by shared east-African ancestry\u003c/h2\u003e \u003cp\u003eGiven the ancestral genetic differences between the two study populations, concomitant differences in associated haplotype backgrounds, and reported geographical variation in ESAM prevalence, we considered that the underlying OCM-ESAM association might be best represented by an ancestral haplotype that is shared by the two populations but seen more commonly in Malawi. In this shared ancestry model, the presence of additional ancestry backgrounds (admixture) might obscure the underlying association signal, especially if using genome-averaged ancestry to account for population stratification. To evaluate this, we used RFmix\u003csup\u003e\u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e69\u003c/span\u003e\u003c/sup\u003e to derive \u0026lsquo;local\u0026rsquo; (locus-level) patterns of ancestry, using the CEU, MSL, LWK populations to proxy European, West African-, and East African- Bantu-speaking parental ancestries, respectively, in our cohort (\u003cb\u003eMethods\u003c/b\u003e). We then repeated our OCM-wide association tests, but this time adjusting each SNP for its ancestral haplotype background using TRACTOR. This approach allowed us to account for the potential effects of differing ancestral haplotype backgrounds at putatively associated loci in a way that was agnostic to geography but still ancestrally sensitive.\u003c/p\u003e \u003cp\u003eConsistent with genome-wide admixture estimates, Jamaican individuals harbored more blocks of European- and west-African- derived ancestry than Malawi (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003ea). Across the combined cohort, however, the overall ancestry proportions in ESAM were not significantly different than seen in the reference NESAM population (t-test, lowest p\u0026thinsp;=\u0026thinsp;0.24; \u003cb\u003eSupplementary Figure S8\u003c/b\u003e), suggesting that differences in effect size and association in our original association were not solely the consequence of mismatched ancestry between ESAM and NESAM. Using TRACTOR to perform logistic regression with 3-way admixture, we found that when conditioned on shared East African ancestry (proxied by LWK), seven loci surpassed the loci-level threshold for association, including candidates \u003cem\u003eMTHFR1\u003c/em\u003e, \u003cem\u003ePRICKLE2\u003c/em\u003e, and \u003cem\u003ePLD2\u003c/em\u003e, but only one OCM locus, and none of our candidates, met the same significance criterion when adjusting for West African ancestry (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003eb-c). We also observed that stratifying our random sampling cumulative association by ancestry background demonstrated a higher cumulative association between variants of presumed East African ancestry compared to West African ancestry at a p-value threshold of 0.001 (z\u0026thinsp;=\u0026thinsp;2.81 and =\u0026thinsp;0.54, respectively) (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e6\u003c/span\u003ed). Collectively, these findings were consistent with our OCM-ESAM association being driven by variants carried on a shared haplotype background more common among individuals of East African ancestry.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eSignatures of selection at OCM loci among African populations\u003c/h2\u003e \u003cp\u003ePeriods of famine have been associated with substantial childhood mortality, with survival into adolescence following childhood starvation potentially having a major effect on reproductive fitness. In the absence of modern-day nutritional support, ESAM is associated with higher mortality than NESAM, creating the potential for variants and loci that influence the risk of developing ESAM or NESAM to be subject to relatively strong selection. To explore this further, we generated haplotype-similarity-based scores (iHS) of recent selection at our OCM loci (+/- 50kb) using whole genome sequence data from a subset of 151 individuals from Malawi (\u003cb\u003eMethods\u003c/b\u003e). We contextualized our results by considering iHS scores for the same OCM loci generated from genome sequencing done through the H3Africa consortium\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e, which includes samples from Benin, Botswana, Cameroon, Mali, Nigeria, Zambia, and Gur speakers from West Africa (Burkino Faso and Ghana).\u003c/p\u003e \u003cp\u003eUsing an ad hoc, but conservative, locus-selection threshold of more than 10% of SNPs having normalized iHS scores\u0026thinsp;\u0026gt;\u0026thinsp;2 or \u0026lt;-2, we found that in Malawi, 15 of our 95 OCM loci (15.8%) had some evidence of selection, with \u003cem\u003eAMT\u003c/em\u003e (0.429), \u003cem\u003eSDS\u003c/em\u003e/\u003cem\u003eSDSL\u003c/em\u003e (0.301), \u003cem\u003eMDH2\u003c/em\u003e (0.208) and \u003cem\u003eSHMT1\u003c/em\u003e (0.174) having the most SNPs surpassing our threshold (\u003cb\u003eSupplementary Table S6\u003c/b\u003e). Of these loci, however, none overlapped our top ESAM-associated candidate loci, most of which had proportions\u0026thinsp;\u0026lt;\u0026thinsp;3%. Across the African countries in the H3Africa dataset, however, 2,440 SNPs within 50kb of OCM genes had evidence of selection in at least two countries (\u003cb\u003eSupplementary Table S6; Supplementary Figure S9\u003c/b\u003e), with 55 SNPs having strong evidence across all surveyed groups (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e7\u003c/span\u003e\u003cb\u003e)\u003c/b\u003e. In this broader dataset, rs1052748 (ESAM-associated \u003cem\u003ePLD2\u003c/em\u003e coding variant) had the strongest selection signal among ESAM-associated OCM SNPs (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e7\u003c/span\u003e; \u003cb\u003eSupplementary Table S7\u003c/b\u003e), as did nearby \u003cem\u003ePLD2\u003c/em\u003e SNPs, particularly in Zambia (proportion\u0026thinsp;=\u0026thinsp;9.6%), Cameroon (9.3%), and among Gur speakers from West Africa (13.3%).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eLastly, we considered that haplotype similarity scores at selected loci might be amplified among ESAM samples, mimicking the cumulative allelic effects seen in our association models (i.e. enriching for multiple copies of a risk/protective allele). Across all OCM loci, the mean iHS score was significantly higher in ESAM (n\u0026thinsp;=\u0026thinsp;90) than NESAM individuals (n\u0026thinsp;=\u0026thinsp;61) (normalized iHS- two sample t-test p\u0026thinsp;=\u0026thinsp;4x10\u003csup\u003e\u0026minus;\u0026thinsp;7\u003c/sup\u003e; raw iHS - p\u0026thinsp;=\u0026thinsp;5x10\u003csup\u003e\u0026minus;\u0026thinsp;14\u003c/sup\u003e; \u003cb\u003eSupplementary Figure S10a\u003c/b\u003e). Among variant sites within 10KB of SNPs surpassing our locus-wide cut-off (n\u0026thinsp;=\u0026thinsp;839 ESAM; n\u0026thinsp;=\u0026thinsp;753 NESAM), mean iHS scores were also significantly different (two sample t-test p\u0026thinsp;=\u0026thinsp;0.01; \u003cb\u003eSupplementary Figure S10b\u003c/b\u003e) between the two groups.\u003c/p\u003e \u003c/div\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eWe report results of a case-control genetic association study of ESAM (affected/cases) versus NESAM (reference/controls) at loci associated with OCM. To our knowledge, this is the largest genetic study of ESAM to date. Our hypothesis-driven targeted pathway approach allowed us to maximize our modest sample size and, in the absence of an available replication cohort, we find strong internal consistency between the two country cohorts assessed. SAM is an acute, life-threatening disorder, for which, traditionally, recruitment, consent, and sampling for large-scale research are secondary considerations; however, the ability to retrospectively recruit individuals several months to years after surviving SAM holds potential for future, larger, germline genetic studies of ESAM. Here, we report evidence of an association between cumulative and locus-specific genetic variation at OCM-loci and the risk of developing ESAM. This association appears to be driven by variants carried on an East African genetic background and is bolstered by evidence that associated variants causally-mediate the association between OCM metabolites and ESAM, as well as by suggestions of recent selection at ESAM-OCM associated loci.\u003c/p\u003e \u003cp\u003eIntronic SNPs in \u003cem\u003eGABBR2\u003c/em\u003e and \u003cem\u003ePRICKLE2\u003c/em\u003e showed the strongest locus-specific association. Both loci include multiple sites of reported open chromatin, suggestive of regulatory effects; however, it is unknown whether these effects would be limited to the most proximal genes versus having more distal effects on other \u003cem\u003ecis-\u003c/em\u003e or \u003cem\u003etrans-\u003c/em\u003e genes. The latter observation is particularly relevant at \u003cem\u003eGABBR2\u003c/em\u003e where the top SNP (rs70387285) had a much higher minor allele frequency among African populations. The relative lack of transcriptional or gene-regulatory contexts from Africa makes it difficult to fully annotate the regulatory potential of candidate variants identified in studies such as ours. We did, however, observe locus-wide positive association between a GTEx-reported eQTL missense variant in \u003cem\u003ePLD2\u003c/em\u003e (rs1052748) and ESAM. The minor allele of rs1052748 is associated with reduced \u003cem\u003ePLD2\u003c/em\u003e transcription. The functional consequence of variable transcription on phospholipase D2 phosphatidyl activity and production of choline from phosphatidylcholine is uncertain; however, reduced choline availability would be consistent with lower choline concentrations noted in ESAM\u003csup\u003e\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThese single locus metabolic effects, however, do not occur in isolation - small changes in metabolic flux at multiple points in OCM could result in larger effects on OCM consistent with those reported in ESAM. This was consistent with both our cumulative association model and our Mendelian Randomization models. For instance, most of the loci reaching locus-wide association were associated with homocysteine metabolism, and MR analyses suggested causal effects of SNPs at \u003cem\u003eGABBR2, PRICKLE2\u003c/em\u003e, and \u003cem\u003ePLD2\u003c/em\u003e on metabolites that are either directly derived from homocysteine (Cysteine and Cystathionine) or are involved in its re-conversion to methionine (Betaine). These observations strengthen previous studies implicating aberrant OCM, and particularly deficient re-methylation of homocysteine to methionine, in the pathogenesis of ESAM. These data thus provide support for proposed therapeutic interventions to supplement OCM turnover in SAM using cofactors such as choline (clinicaltrials.gov ID NCT06154174). Despite this, our preliminary estimates suggest whilst genetic variation has a moderately strong effect on the variance in SAM outcomes (estimated heritability\u0026thinsp;=\u0026thinsp;0.22; standard error (se)\u0026thinsp;=\u0026thinsp;0.37), variation at OCM loci accounts for a much smaller proportion (estimated heritability\u0026thinsp;=\u0026thinsp;0.06; se\u0026thinsp;=\u0026thinsp;0.06). The modest sample size employed means that these heritability estimates have a relatively large standard error as do the effect sizes inferred from our MR analysis. Optimistically, however, there may be several unidentified loci that causally influence ESAM risk. Larger cohorts, particularly from areas where the prevalence of ESAM relative to NESAM remains high, would allow for a well-powered genome-wide approach to identify other ESAM-associated loci and/or pathways for therapeutic and nutritional targeting.\u003c/p\u003e \u003cp\u003eWe were able to uniquely leverage African intra-continental admixture patterns to aid in our mapping. The deep ancestral tree within Africa harbors a complex and diverse genetic landscape that can be as disparate as inter-continental variation. The higher rates of ESAM in east-central and southern Africa, therefore, presented an opportunity to employ admixture mapping approaches that have been long proposed for diseases with different population prevalences. We observed evidence of ESAM association at OCM loci when conditioned on East African, but not West African, haplotype backgrounds shared between Malawi and Jamaica. We hypothesized that reducing the \u0026lsquo;noise\u0026rsquo; of West African haplotypes in the analysis, would have augmented the effect of shared East African ancestry beyond the effect seen in our primary association. The weaker signal observed likely represents limitations in the representation of African haplotypes in current imputation panels and public databases\u003csup\u003e\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e81\u003c/span\u003e,\u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e82\u003c/span\u003e\u003c/sup\u003e, making imputation and selection of parental ancestral proxies necessarily deficient (e.g. the LWK are an imperfect proxy for the ancestral East African ancestry of Malawi). None the less, our results provide a conceptual model for leveraging admixture within genetically heterogenous groups to genetically map traits with prevalence differences between groups, provided appropriate parental proxies can be identified and there is sufficient genetic distance between the groups in question.\u003c/p\u003e \u003cp\u003eThere are several examples of how migration across, into, and out of Africa, has interfaced with varying infectious and non-infectious exposures to shape the African genome through adaptation\u003csup\u003e\u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e83\u003c/span\u003e\u003c/sup\u003e. More recent studies have implicated multiple novel loci, often with varying effect sizes across the continent\u003csup\u003e\u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e65\u003c/span\u003e\u003c/sup\u003e. The history of famine and food crises in Africa varies by country and are complex; periods of famine have been documented as far back as the 1600s\u003csup\u003e\u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e84\u003c/span\u003e\u003c/sup\u003e, with speculation that these were more severe in the 19th century, at least in east and southern African countries such as Zimbabwe and Tanzania\u003csup\u003e\u003cspan citationid=\"CR86\" class=\"CitationRef\"\u003e85\u003c/span\u003e,\u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e86\u003c/span\u003e\u003c/sup\u003e. At the same time, children developing ESAM have been noted to have higher birth weights\u003csup\u003e\u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e87\u003c/span\u003e\u003c/sup\u003e, despite the putative link between low birth weight and early-life SAM. These observations provided a starting point for considering that differential survival between ESAM and NESAM might be reflected in signatures of recent selection at ESAM-associated OCM loci. We noted several OCM loci with strong evidence of selection, particularly the serine dehydratase (SDS/SDSH) complex on chromosome 12, which had a strong signal in all the populations evaluated. Among our ESAM-associated OCM loci, however, only \u003cem\u003ePLD2\u003c/em\u003e showed evidence of selection. Like our genetic association, we observed that OCM selection scores among ESAM individuals were, on average, significantly higher than observed among participants with NESAM, suggesting that we may be similarly underpowered to observe subtle selection signatures at individual loci that contribute to a pathway-wide metabolic effect. Going forward, integration of selection statistics at scale, particularly in the context of African genomic association studies, could provide valuable biological insights, especially for novel loci and understudied disorders.\u003c/p\u003e \u003cp\u003eGenome sequencing and genetic mapping efforts in Africa and among populations of predominantly African ancestry are gaining traction, although the number of reference genomes, particularly for non-West African groups, still vastly under-represents the variation across the continent. Our study provides a starting exemplar for applying a hypothesis-driven, population genetics-informed model to study a disease with high prevalence in specific African population groups. Continued expansion of sequencing efforts in African populations, alongside refinement and deployment of association models that can integrate intra-continental admixture and population genetic statistics, has the potential to maximize the future of human genetics studies on the continent and globally.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eACKNOWLEDGMENTS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work utilized the computational resources of the NIH HPC Biowulf cluster (http://hpc.nih.gov) as well as the University of Cape Town\u0026rsquo;s ICTS High Performance Computing team (hpc.uct.ac.za). This work represents the brainchild of Prof. Colin McKenzie, who died prior to the publication of the manuscript. The authors would also like to acknowledge the advice of Historian Dr. Christopher Donohue of the National Human Genome Research Institute (NHGRI) regarding famine across Africa in recent centuries. This research was supported by a Clinical Scientist Development Award from the Doris Duke Charitable Foundation (Grant #: 2013096) and USDA, ARS cooperative agreement (58-3092-5-001), both to N.A.H., K.V.S. and N.C.L. were supported by Award Numbers T32GM008307 and 5 T32GM008231, respectively, from the National Institute of General Medical Sciences to Baylor College of Medicine. E.G.A was supported by grant R01HG012869 from NHGRI. Portions of the work reported were supported by grant HG-200412 from the NHGRI of the NIH to N.A.H. The work and views expressed do not reflect the views of BCM or the NIH.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAUTHOR CONTRIBUTIONS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eI.T., M.J.M., C.A.M., M.E.R., C.T-B., and N.A.H. designed the study; O.B., C.T-B.,\u003c/p\u003e\n\u003cp\u003eI.T., T.M., and M.J.M. recruited participants, obtained informed consent, and collected\u003c/p\u003e\n\u003cp\u003esamples; S.H., K.M., R.S., N.J.H. and A.H. processed and molecularly characterized samples; S.T. conducted cell-free DNA (cfDNA) extractions and J.R. conducted SNP genotyping from cfDNA. A.J., S.S., N.C.L., Y.H., Q.L., K.V.S., and N.A.H. conducted the primary data analyses; D.S. and A.C conducted population genetic analyses including that on H3Africa populations; E.G.A. facilitated local ancestry and TRACTOR analyses; W.A and M.M. performed imputation on the H3Africa server for data from Malawi. N.C.L., Q.L., E.B., and N.A.H. wrote the paper. All authors read, edited, and approved the manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDECLARATION OF INTERESTS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe remaining authors do not have any conflicts or relevant interests to declare.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWEB RESOURCES\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eH3Africa array v1.1 - https://h3africa.org/index.php/2019/12/12/h3africa-chip-faq/\u003c/p\u003e\n\u003cp\u003eH3Africa Array manifest - https://www.illumina.com/content/infinium-h3africa-consortium-array-data-sheet)\u003c/p\u003e\n\u003cp\u003eH3ABioNet Imputation Service - https://www.h3abionet.org/resources/h3abionet-imputation-service; h3abionet/chipimputation.\u003c/p\u003e\n\u003cp\u003eMinimac - https://github.com/statgen/Minimac4\u003c/p\u003e\n\u003cp\u003eUniversity of Cape Town High-Performance Computing - https://ucthpc.uct.ac.za/\u003c/p\u003e\n\u003cp\u003e1000 Genomes Project Resource - https://www.internationalgenome.org/\u003c/p\u003e\n\u003cp\u003e\u0026lsquo;In house scripts\u0026rsquo; - https://github.com/NHGRI/SAM_OCMgenetics\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDATA AND CODE AVAILABILITY\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe code generated during this study and summary statistics are available through GitHub (https://github.com/NHGRI/SAM_OCMgenetics). The raw sequencing and genotyping data supporting the current study have not been deposited in a public repository, as Ethical approvals in Malawi and Jamaica accompanying recruitment and sample collection for this study predated current data sharing protocols and do not explicitly include publicly sharing data; however, investigators interested in study-based data access should directly contact the corresponding author. \u003cbr\u003e \u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eLevels and trends in child malnutrition: UNICEF/WHO/The World Bank Group joint child malnutrition estimates: key findings of the 2020 edition. (2023).\u003c/li\u003e\n\u003cli\u003eBlack, R.E., Victora, C.G., Walker, S.P., Bhutta, Z.A., Christian, P., de Onis, M., Ezzati, M., Grantham-McGregor, S., Katz, J., Martorell, R., et al. (2013). Maternal and child undernutrition and overweight in low-income and middle-income countries. Lancet \u003cem\u003e382\u003c/em\u003e, 427-451. 10.1016/S0140-6736(13)60937-X.\u003c/li\u003e\n\u003cli\u003eWHO child growth standards and the identification of severe acute malnutrition in infants and children. (2024).\u003c/li\u003e\n\u003cli\u003eBhutta, Z.A., Berkley, J.A., Bandsma, R.H.J., Kerac, M., Trehan, I., and Briend, A. (2017). Severe childhood malnutrition. Nat Rev Dis Primers \u003cem\u003e3\u003c/em\u003e, 17067. 10.1038/nrdp.2017.67.\u003c/li\u003e\n\u003cli\u003eFrison, S., Checchi, F., and Kerac, M. (2015). Omitting edema measurement: how much acute malnutrition are we missing? The American journal of clinical nutrition \u003cem\u003e102\u003c/em\u003e, 1176-1181. 10.3945/ajcn.115.108282.\u003c/li\u003e\n\u003cli\u003eAlvarez, J.L., Dent, N., Browne, L., and Briend, M.M.a.A. (2016). Putting Child Kwashiorkor on the Map.\u003c/li\u003e\n\u003cli\u003eSturgeon, J.P., Mufukari, W., Tome, J., Dumbura, C., Majo, F.D., Ngosa, D., Chandwe, K., Kapoma, C., Mutasa, K., Nathoo, K.J., et al. (2023). Risk factors for inpatient mortality among children with severe acute malnutrition in Zimbabwe and Zambia. European journal of clinical nutrition \u003cem\u003e77\u003c/em\u003e, 895-904. 10.1038/s41430-023-01320-9.\u003c/li\u003e\n\u003cli\u003eChildhood Acute, I., and Nutrition, N. (2022). Childhood mortality during and after acute illness in Africa and south Asia: a prospective cohort study. Lancet Glob Health \u003cem\u003e10\u003c/em\u003e, e673-e684. 10.1016/S2214-109X(22)00118-8.\u003c/li\u003e\n\u003cli\u003eGolden, M.H. (1998). Oedematous malnutrition. Br Med Bull \u003cem\u003e54\u003c/em\u003e, 433-444.\u003c/li\u003e\n\u003cli\u003eWaterlow, J.C. (1997). Protein-energy malnutrition: the nature and extent of the problem. Clinical Nutrition (Edinburgh, Scotland) \u003cem\u003e16 Suppl 1\u003c/em\u003e, 3-9. 10.1016/s0261-5614(97)80043-x.\u003c/li\u003e\n\u003cli\u003eMcKenzie, C.A., Wakamatsu, K., Hanchard, N.A., Forrester, T., and Ito, S. (2007). Childhood malnutrition is associated with a reduction in the total melanin content of scalp hair. Br J Nutr \u003cem\u003e98\u003c/em\u003e, 159-164. 10.1017/S0007114507694458.\u003c/li\u003e\n\u003cli\u003eWilliams, C.D. (1935). Kwashiorkor. A disease of children associated with a maize diet. Lancet \u003cem\u003eii\u003c/em\u003e, 1151-1152.\u003c/li\u003e\n\u003cli\u003eWilliams, C.D. (1933). A nutritional disease of childhood associated with a maize diet. Archives of Disease in Childhood \u003cem\u003e8\u003c/em\u003e, 423-433.\u003c/li\u003e\n\u003cli\u003eHeikens, G.T., and Manary, M. (2009). 75 years of Kwashiorkor in Africa. Malawi medical journal : the journal of Medical Association of Malawi \u003cem\u003e21\u003c/em\u003e, 96-98.\u003c/li\u003e\n\u003cli\u003eManary, M.J., Heikens, G.T., and Golden, M. (2009). Kwashiorkor: more hypothesis testing is needed to understand the aetiology of oedema. Malawi medical journal : the journal of Medical Association of Malawi \u003cem\u003e21\u003c/em\u003e, 106-107.\u003c/li\u003e\n\u003cli\u003eKismul, H., Van den Broeck, J., and Lunde, T.M. (2014). Diet and kwashiorkor: a prospective study from rural DR Congo. PeerJ \u003cem\u003e2\u003c/em\u003e, e350. 10.7717/peerj.350.\u003c/li\u003e\n\u003cli\u003eLin, C.A., Boslaugh, S., Ciliberto, H.M., Maleta, K., Ashorn, P., Briend, A., and Manary, M.J. (2007). A prospective assessment of food and nutrient intake in a population of Malawian children at risk for kwashiorkor. J Pediatr Gastroenterol Nutr \u003cem\u003e44\u003c/em\u003e, 487-493. 10.1097/MPG.0b013e31802c6e57.\u003c/li\u003e\n\u003cli\u003eSullivan, J., Ndekha, M., Maker, D., Hotz, C., and Manary, M.J. (2006). The quality of the diet in Malawian children with kwashiorkor and marasmus. Matern Child Nutr \u003cem\u003e2\u003c/em\u003e, 114-122. 10.1111/j.1740-8709.2006.00053.x.\u003c/li\u003e\n\u003cli\u003eGolden, M.H. (2015). Nutritional and other types of oedema, albumin, complex carbohydrates and the interstitium - a response to Malcolm Coulthard\u0026apos;s hypothesis: Oedema in kwashiorkor is caused by hypo-albuminaemia. Paediatrics and International Child Health \u003cem\u003e35\u003c/em\u003e, 90-109. 10.1179/2046905515Y.0000000010.\u003c/li\u003e\n\u003cli\u003eLaditan, A.A., and Reeds, P.J. (1976). A study of the age of onset, diet and the importance of infection in the pattern of severe protein-energy malnutrition in Ibadan, Nigeria. Br J Nutr \u003cem\u003e36\u003c/em\u003e, 411-419. 10.1079/bjn19760096.\u003c/li\u003e\n\u003cli\u003eKristensen, K.H., Wiese, M., Rytter, M.J., Ozcam, M., Hansen, L.H., Namusoke, H., Friis, H., and Nielsen, D.S. (2016). Gut Microbiota in Children Hospitalized with Oedematous and Non-Oedematous Severe Acute Malnutrition in Uganda. PLoS Negl Trop Dis \u003cem\u003e10\u003c/em\u003e, e0004369. 10.1371/journal.pntd.0004369.\u003c/li\u003e\n\u003cli\u003eChristie, C.D., Heikens, G.T., and Golden, M.H. (1992). Coagulase-negative staphylococcal bacteremia in severely malnourished Jamaican children. Pediatr Infect Dis J \u003cem\u003e11\u003c/em\u003e, 1030-1036. 10.1097/00006454-199211120-00008.\u003c/li\u003e\n\u003cli\u003eJahoor, F., Badaloo, A., Reid, M., and Forrester, T. (2008). Protein metabolism in severe childhood malnutrition. Annals of Tropical Paediatrics \u003cem\u003e28\u003c/em\u003e, 87-101. 10.1179/146532808X302107.\u003c/li\u003e\n\u003cli\u003eBadaloo, A.V., Forrester, T., Reid, M., and Jahoor, F. (2006). Lipid kinetic differences between children with kwashiorkor and those with marasmus2. The American Journal of Clinical Nutrition \u003cem\u003e83\u003c/em\u003e, 1283-1288. 10.1093/ajcn/83.6.1283.\u003c/li\u003e\n\u003cli\u003eBadaloo, A., Hsu, J.W., Taylor-Bryan, C., Green, C., Reid, M., Forrester, T., and Jahoor, F. (2012). Dietary cysteine is used more efficiently by children with severe acute malnutrition with edema compared with those without edema. Am J Clin Nutr \u003cem\u003e95\u003c/em\u003e, 84-90. 10.3945/ajcn.111.024323.\u003c/li\u003e\n\u003cli\u003eJahoor, F., Badaloo, A., Reid, M., and Forrester, T. (2006). Sulfur amino acid metabolism in children with severe childhood undernutrition: methionine kinetics. The American journal of clinical nutrition \u003cem\u003e84\u003c/em\u003e, 1400-1405.\u003c/li\u003e\n\u003cli\u003eManary, M.J., Leeuwenburgh, C., and Heinecke, J.W. (2000). Increased oxidative stress in kwashiorkor. J Pediatr \u003cem\u003e137\u003c/em\u003e, 421-424. 10.1067/mpd.2000.107512.\u003c/li\u003e\n\u003cli\u003eAmadi, B., Fagbemi, A.O., Kelly, P., Mwiya, M., Torrente, F., Salvestrini, C., Day, R., Golden, M.H., Eklund, E.A., Freeze, H.H., and Murch, S.H. (2009). Reduced production of sulfated glycosaminoglycans occurs in Zambian children with kwashiorkor but not marasmus. Am J Clin Nutr \u003cem\u003e89\u003c/em\u003e, 592-600. 10.3945/ajcn.2008.27092.\u003c/li\u003e\n\u003cli\u003eDucker, G.S., and Rabinowitz, J.D. (2017). One-Carbon Metabolism in Health and Disease. Cell metabolism \u003cem\u003e25\u003c/em\u003e, 27-42. 10.1016/j.cmet.2016.08.009.\u003c/li\u003e\n\u003cli\u003eJahoor, F., Badaloo, A., Reid, M., and Forrester, T. (2008). Protein metabolism in severe childhood malnutrition. Annals of tropical paediatrics \u003cem\u003e28\u003c/em\u003e, 87-101. 10.1179/146532808X302107.\u003c/li\u003e\n\u003cli\u003eMay, T., de la Haye, B., Nord, G., Klatt, K., Stephenson, K., Adams, S., Bollinger, L., Hanchard, N., Arning, E., Bottiglieri, T., et al. (2022). One-carbon metabolism in children with marasmus and kwashiorkor. EBioMedicine \u003cem\u003e75\u003c/em\u003e, 103791. 10.1016/j.ebiom.2021.103791.\u003c/li\u003e\n\u003cli\u003eMay, T., Klatt, K.C., Smith, J., Castro, E., Manary, M., Caudill, M.A., Jahoor, F., and Fiorotto, M.L. (2018). Choline Supplementation Prevents a Hallmark Disturbance of Kwashiorkor in Weanling Mice Fed a Maize Vegetable Diet: Hepatic Steatosis of Undernutrition. Nutrients \u003cem\u003e10\u003c/em\u003e. 10.3390/nu10050653.\u003c/li\u003e\n\u003cli\u003eSmith, M.I., Yatsunenko, T., Manary, M.J., Trehan, I., Mkakosya, R., Cheng, J., Kau, A.L., Rich, S.S., Concannon, P., Mychaleckyj, J.C., et al. (2013). Gut microbiomes of Malawian twin pairs discordant for kwashiorkor. Science \u003cem\u003e339\u003c/em\u003e, 548-554. 10.1126/science.1229000.\u003c/li\u003e\n\u003cli\u003eSchulze, K.V., Swaminathan, S., Howell, S., Jajoo, A., Lie, N.C., Brown, O., Sadat, R., Hall, N., Zhao, L., Marshall, K., et al. (2019). Edematous severe acute malnutrition is characterized by hypomethylation of DNA. Nat Commun \u003cem\u003e10\u003c/em\u003e, 5791. 10.1038/s41467-019-13433-6.\u003c/li\u003e\n\u003cli\u003eMunthali, T., Jacobs, C., Sitali, L., Dambe, R., and Michelo, C. (2015). Mortality and morbidity patterns in under-five children with severe acute malnutrition (SAM) in Zambia: a five-year retrospective review of hospital-based records (2009\u0026ndash;2013). Archives of Public Health \u003cem\u003e73\u003c/em\u003e, 23. 10.1186/s13690-015-0072-1.\u003c/li\u003e\n\u003cli\u003eGonzales, G.B., Ngari, M.M., Njunge, J.M., Thitiri, J., Mwalekwa, L., Mturi, N., Mwangome, M.K., Ogwang, C., Nyaguara, A., and Berkley, J.A. (2020). Phenotype is sustained during hospital readmissions following treatment for complicated severe malnutrition among Kenyan children: A retrospective cohort study. Matern Child Nutr \u003cem\u003e16\u003c/em\u003e, e12913. 10.1111/mcn.12913.\u003c/li\u003e\n\u003cli\u003eDuggal, P., and Petri, W.A., Jr. (2018). Does Malnutrition Have a Genetic Component? Annu Rev Genomics Hum Genet \u003cem\u003e19\u003c/em\u003e, 247-262. 10.1146/annurev-genom-083117-021340.\u003c/li\u003e\n\u003cli\u003eMarshall, K.G., Howell, S., Reid, M., Badaloo, A., Farrall, M., Forrester, T., and McKenzie, C.A. (2006). Glutathione S-transferase polymorphisms may be associated with risk of oedematous severe childhood malnutrition. Br J Nutr \u003cem\u003e96\u003c/em\u003e, 243-248. 10.1079/bjn20061825.\u003c/li\u003e\n\u003cli\u003eMarshall, K.G., Swaby, K., Hamilton, K., Howell, S., Landis, R.C., Hambleton, I.R., Reid, M., Fletcher, H., Forrester, T., and McKenzie, C.A. (2011). A preliminary examination of the effects of genetic variants of redox enzymes on susceptibility to oedematous malnutrition and on percentage cytotoxicity in response to oxidative stress in vitro. Ann Trop Paediatr \u003cem\u003e31\u003c/em\u003e, 27-36. 10.1179/146532811X12925735813805.\u003c/li\u003e\n\u003cli\u003eMarshall, K.G., Howell, S., Badaloo, A.V., Reid, M., Farrall, M., Forrester, T., and McKenzie, C.A. (2006). Polymorphisms in genes involved in folate metabolism as risk factors for oedematous severe childhood malnutrition: a hypothesis-generating study. Ann Trop Paediatr \u003cem\u003e26\u003c/em\u003e, 107-114. 10.1179/146532806X107449.\u003c/li\u003e\n\u003cli\u003eHazra, A., Kraft, P., Lazarus, R., Chen, C., Chanock, S.J., Jacques, P., Selhub, J., and Hunter, D.J. (2009). Genome-wide significant predictors of metabolites in the one-carbon metabolism pathway. Hum Mol Genet \u003cem\u003e18\u003c/em\u003e, 4677-4687. 10.1093/hmg/ddp428.\u003c/li\u003e\n\u003cli\u003eWilliams, S.R., Yang, Q., Chen, F., Liu, X., Keene, K.L., Jacques, P., Chen, W.-M., Weinstein, G., Hsu, F.-C., Beiser, A., et al. (2014). Genome-wide meta-analysis of homocysteine and methionine metabolism identifies five one carbon metabolism loci and a novel association of ALDH1L1 with ischemic stroke. PLoS genetics \u003cem\u003e10\u003c/em\u003e, e1004214. 10.1371/journal.pgen.1004214.\u003c/li\u003e\n\u003cli\u003eMarshall, K.G., Howell, S., Reid, M., Badaloo, A.V., Farrall, M., Forrester, T., and McKenzie, C.A. (2006). Glutathione S-transferase polymorphisms may be associated with risk of oedematous severe childhood malnutrition. The British journal of nutrition \u003cem\u003e96\u003c/em\u003e, 243-248.\u003c/li\u003e\n\u003cli\u003eWaterlow, J.C. (1972). Classification and definition of protein-calorie malnutrition. Br Med J \u003cem\u003e3\u003c/em\u003e, 566-569. 10.1136/bmj.3.5826.566.\u003c/li\u003e\n\u003cli\u003eBaker, E.J., Jay, J.J., Bubier, J.A., Langston, M.A., and Chesler, E.J. (2012). GeneWeaver: a web-based system for integrative functional genomics. Nucleic Acids Research \u003cem\u003e40\u003c/em\u003e, D1067-1076. 10.1093/nar/gkr968.\u003c/li\u003e\n\u003cli\u003eAshburner, M., Ball, C.A., Blake, J.A., Botstein, D., Butler, H., Cherry, J.M., Davis, A.P., Dolinski, K., Dwight, S.S., Eppig, J.T., et al. (2000). Gene Ontology: tool for the unification of biology. Nature genetics \u003cem\u003e25\u003c/em\u003e, 25-29. 10.1038/75556.\u003c/li\u003e\n\u003cli\u003eAleksander, S.A., Balhoff, J., Carbon, S., Cherry, J.M., Drabkin, H.J., Ebert, D., Feuermann, M., Gaudet, P., Harris, N.L., Hill, D.P., et al. (2023). The Gene Ontology knowledgebase in 2023. Genetics \u003cem\u003e224\u003c/em\u003e, iyad031. 10.1093/genetics/iyad031.\u003c/li\u003e\n\u003cli\u003eSollis, E., Mosaku, A., Abid, A., Buniello, A., Cerezo, M., Gil, L., Groza, T., G\u0026uuml;neş, O., Hall, P., Hayhurst, J., et al. (2023). The NHGRI-EBI GWAS Catalog: knowledgebase and deposition resource. Nucleic Acids Research \u003cem\u003e51\u003c/em\u003e, D977-D985. 10.1093/nar/gkac1010.\u003c/li\u003e\n\u003cli\u003eHindorff, L.A., Junkins, H.A., Hall, P.N., and Manolio, T.A. (2011). A Catalog of Published Genome-Wide Association Studies.\u003c/li\u003e\n\u003cli\u003eKanehisa, M., and Goto, S. (2000). KEGG: kyoto encyclopedia of genes and genomes. Nucleic Acids Research \u003cem\u003e28\u003c/em\u003e, 27-30. 10.1093/nar/28.1.27.\u003c/li\u003e\n\u003cli\u003eGeer, L.Y., Marchler-Bauer, A., Geer, R.C., Han, L., He, J., He, S., Liu, C., Shi, W., and Bryant, S.H. (2010). The NCBI BioSystems database. Nucleic Acids Research \u003cem\u003e38\u003c/em\u003e, D492-496. 10.1093/nar/gkp858.\u003c/li\u003e\n\u003cli\u003eCerami, E.G., Gross, B.E., Demir, E., Rodchenkov, I., Babur, \u0026Ouml;., Anwar, N., Schultz, N., Bader, G.D., and Sander, C. (2011). Pathway Commons, a web resource for biological pathway data. Nucleic Acids Research \u003cem\u003e39\u003c/em\u003e, D685-D690. 10.1093/nar/gkq1039.\u003c/li\u003e\n\u003cli\u003ePurcell, S., Neale, B., Todd-Brown, K., Thomas, L., Ferreira, Manuel A.R., Bender, D., Maller, J., Sklar, P., de Bakker, Paul I.W., Daly, Mark J., and Sham, Pak C. (2007). PLINK: A Tool Set for Whole-Genome Association and Population-Based Linkage Analyses. American Journal of Human Genetics \u003cem\u003e81\u003c/em\u003e, 559-575.\u003c/li\u003e\n\u003cli\u003eJohnston, H.R., Hu, Y.J., Gao, J., O\u0026apos;Connor, T.D., Abecasis, G.R., Wojcik, G.L., Gignoux, C.R., Gourraud, P.A., Lizee, A., Hansen, M., et al. (2017). Identifying tagging SNPs for African specific genetic variation from the African Diaspora Genome. Sci Rep \u003cem\u003e7\u003c/em\u003e, 46398. 10.1038/srep46398.\u003c/li\u003e\n\u003cli\u003eDas, S., Forer, L., Schonherr, S., Sidore, C., Locke, A.E., Kwong, A., Vrieze, S.I., Chew, E.Y., Levy, S., McGue, M., et al. (2016). Next-generation genotype imputation service and methods. Nat Genet \u003cem\u003e48\u003c/em\u003e, 1284-1287. 10.1038/ng.3656.\u003c/li\u003e\n\u003cli\u003eEagle: multi-locus association mapping on a genome-wide scale made routine | Bioinformatics | Oxford Academic. (2024).\u003c/li\u003e\n\u003cli\u003eMulder, N.J., Adebiyi, E., Alami, R., Benkahla, A., Brandful, J., Doumbia, S., Everett, D., Fadlelmola, F.M., Gaboun, F., Gaseitsiwe, S., et al. (2016). H3ABioNet, a sustainable pan-African bioinformatics network for human heredity and health in Africa. Genome research \u003cem\u003e26\u003c/em\u003e, 271-277. 10.1101/gr.196295.115.\u003c/li\u003e\n\u003cli\u003eBaichoo, S., Souilmi, Y., Panji, S., Botha, G., Meintjes, A., Hazelhurst, S., Bendou, H., Beste, E., Mpangase, P.T., Souiai, O., et al. (2018). Developing reproducible bioinformatics analysis workflows for heterogeneous computing environments to support African genomics. BMC Bioinformatics \u003cem\u003e19\u003c/em\u003e, 457. 10.1186/s12859-018-2446-1.\u003c/li\u003e\n\u003cli\u003eGenomes Project, C., Auton, A., Brooks, L.D., Durbin, R.M., Garrison, E.P., Kang, H.M., Korbel, J.O., Marchini, J.L., McCarthy, S., McVean, G.A., and Abecasis, G.R. (2015). A global reference for human genetic variation. Nature \u003cem\u003e526\u003c/em\u003e, 68-74. 10.1038/nature15393.\u003c/li\u003e\n\u003cli\u003eDessau, R.B., and Pipper, C.B. (2008). [\u0026apos;\u0026apos;R\u0026quot;--project for statistical computing]. Ugeskrift for Laeger \u003cem\u003e170\u003c/em\u003e, 328-330.\u003c/li\u003e\n\u003cli\u003eZhou, X., and Stephens, M. (2012). Genome-wide efficient mixed-model analysis for association studies. Nature Genetics \u003cem\u003e44\u003c/em\u003e, 821-824. 10.1038/ng.2310.\u003c/li\u003e\n\u003cli\u003eTurner, S.D. (2018). qqman: an R package for visualizing GWAS results using Q-Q and manhattan plots. Journal of Open Source Software \u003cem\u003e3\u003c/em\u003e, 731. 10.21105/joss.00731.\u003c/li\u003e\n\u003cli\u003eBarrett, J.C., Fry, B., Maller, J., and Daly, M.J. (2005). Haploview: analysis and visualization of LD and haplotype maps. Bioinformatics \u003cem\u003e21\u003c/em\u003e, 263-265. 10.1093/bioinformatics/bth457bth457 [pii].\u003c/li\u003e\n\u003cli\u003eJones, B. (2015). Population genetics: the African Genome Variation Project. Nat Rev Genet \u003cem\u003e16\u003c/em\u003e, 68-69. 10.1038/nrg3886.\u003c/li\u003e\n\u003cli\u003eChoudhury, A., Aron, S., Botigue, L.R., Sengupta, D., Botha, G., Bensellak, T., Wells, G., Kumuthini, J., Shriner, D., Fakim, Y.J., et al. (2020). High-depth African genomes inform human migration and health. Nature \u003cem\u003e586\u003c/em\u003e, 741-748. 10.1038/s41586-020-2859-7.\u003c/li\u003e\n\u003cli\u003eSchlebusch, C.M., Skoglund, P., Sjodin, P., Gattepaille, L.M., Hernandez, D., Jay, F., Li, S., De Jongh, M., Singleton, A., Blum, M.G., et al. (2012). Genomic variation in seven Khoe-San groups reveals adaptation and complex African history. Science \u003cem\u003e338\u003c/em\u003e, 374-379. 10.1126/science.1227721.\u003c/li\u003e\n\u003cli\u003eAlexander, D.H., Novembre, J., and Lange, K. (2009). Fast model-based estimation of ancestry in unrelated individuals. Genome Res \u003cem\u003e19\u003c/em\u003e, 1655-1664. 10.1101/gr.094052.109.\u003c/li\u003e\n\u003cli\u003eDelaneau, O., Marchini, J., and Zagury, J.F. (2012). A linear complexity phasing method for thousands of genomes. Nat Methods \u003cem\u003e9\u003c/em\u003e, 179-181. 10.1038/nmeth.1785.\u003c/li\u003e\n\u003cli\u003eMaples, B.K., Gravel, S., Kenny, E.E., and Bustamante, C.D. (2013). RFMix: a discriminative modeling approach for rapid and robust local-ancestry inference. American Journal of Human Genetics \u003cem\u003e93\u003c/em\u003e, 278-288. 10.1016/j.ajhg.2013.06.020.\u003c/li\u003e\n\u003cli\u003eAtkinson, E.G., Maihofer, A.X., Kanai, M., Martin, A.R., Karczewski, K.J., Santoro, M.L., Ulirsch, J.C., Kamatani, Y., Okada, Y., Finucane, H.K., et al. (2021). Tractor uses local ancestry to enable the inclusion of admixed individuals in GWAS and to boost power. Nature Genetics \u003cem\u003e53\u003c/em\u003e, 195-204. 10.1038/s41588-020-00766-y.\u003c/li\u003e\n\u003cli\u003eBurgess, S., Butterworth, A., and Thompson, S.G. (2013). Mendelian randomization analysis with multiple genetic variants using summarized data. Genet Epidemiol \u003cem\u003e37\u003c/em\u003e, 658-665. 10.1002/gepi.21758.\u003c/li\u003e\n\u003cli\u003eBurgess, S., and Thompson, S.G. (2017). Interpreting findings from Mendelian randomization using the MR-Egger method. Eur J Epidemiol \u003cem\u003e32\u003c/em\u003e, 377-389. 10.1007/s10654-017-0255-x.\u003c/li\u003e\n\u003cli\u003eLi, H., and Durbin, R. (2009). Fast and accurate short read alignment with Burrows-Wheeler transform. Bioinformatics \u003cem\u003e25\u003c/em\u003e, 1754-1760. btp324 [pii] 10.1093/bioinformatics/btp324.\u003c/li\u003e\n\u003cli\u003eMcKenna, A., Hanna, M., Banks, E., Sivachenko, A., Cibulskis, K., Kernytsky, A., Garimella, K., Altshuler, D., Gabriel, S., Daly, M., and DePristo, M.a. (2010). The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome research \u003cem\u003e20\u003c/em\u003e, 1297-1303. 10.1101/gr.107524.110.\u003c/li\u003e\n\u003cli\u003evan der Auwera, G., and O\u0026apos;Connor, B.D. (2020). Genomics in the Cloud: Using Docker, GATK, and WDL in Terra (O\u0026apos;Reilly Media, Incorporated).\u003c/li\u003e\n\u003cli\u003eVoight, B.F., Kudaravalli, S., Wen, X., and Pritchard, J.K. (2006). A map of recent positive selection in the human genome. PLoS Biol \u003cem\u003e4\u003c/em\u003e, e72. 10.1371/journal.pbio.0040072.\u003c/li\u003e\n\u003cli\u003eSzpiech, Z.A., and Hernandez, R.D. (2014). selscan: an efficient multithreaded program to perform EHH-based scans for positive selection. Mol Biol Evol \u003cem\u003e31\u003c/em\u003e, 2824-2827. 10.1093/molbev/msu211.\u003c/li\u003e\n\u003cli\u003eWickham, H. (2016). ggplot2 (Springer International Publishing).\u003c/li\u003e\n\u003cli\u003eTanaka, T., Scheet, P., Giusti, B., Bandinelli, S., Piras, M.G., Usala, G., Lai, S., Mulas, A., Corsi, A.M., Vestrini, A., et al. (2009). Genome-wide association study of vitamin B6, vitamin B12, folate, and homocysteine blood concentrations. American Journal of Human Genetics \u003cem\u003e84\u003c/em\u003e, 477-482. 10.1016/j.ajhg.2009.02.011.\u003c/li\u003e\n\u003cli\u003eConsortium, G.T., Laboratory, D.A., Coordinating Center -Analysis Working, G., Statistical Methods groups-Analysis Working, G., Enhancing, G.g., Fund, N.I.H.C., Nih/Nci, Nih/Nhgri, Nih/Nimh, Nih/Nida, et al. (2017). Genetic effects on gene expression across human tissues. Nature \u003cem\u003e550\u003c/em\u003e, 204-213. 10.1038/nature24277.\u003c/li\u003e\n\u003cli\u003eVergara, C., Parker, M.M., Franco, L., Cho, M.H., Valencia-Duarte, A.V., Beaty, T.H., and Duggal, P. (2018). Genotype Imputation Performance of Three Reference Panels Using African Ancestry Individuals. Human genetics \u003cem\u003e137\u003c/em\u003e, 281-292. 10.1007/s00439-018-1881-4.\u003c/li\u003e\n\u003cli\u003eSengupta, D., Botha, G., Meintjes, A., Mbiyavanga, M., Study, A.W.-G., Consortium, H.A., Hazelhurst, S., Mulder, N., Ramsay, M., and Choudhury, A. (2023). Performance and accuracy evaluation of reference panels for genotype imputation in sub-Saharan African populations. Cell Genom \u003cem\u003e3\u003c/em\u003e, 100332. 10.1016/j.xgen.2023.100332.\u003c/li\u003e\n\u003cli\u003eFan, S., Hansen, M.E., Lo, Y., and Tishkoff, S.A. (2016). Going global by adapting local: A review of recent human adaptation. Science \u003cem\u003e354\u003c/em\u003e, 54-59. 10.1126/science.aaf5098.\u003c/li\u003e\n\u003cli\u003eIsichei, E.A. (1997). A history of African societies to 1870 (Cambridge University Press).\u003c/li\u003e\n\u003cli\u003eCheater, A.P., and Iliffe, J.H. (1990). Famine in Zimbabwe 1890-1960. Africa.\u003c/li\u003e\n\u003cli\u003eBloxham, D., and Moses, A.D. (2022). Genocide : key themes, First edition. Edition (Oxford University Press).\u003c/li\u003e\n\u003cli\u003eForrester, T.E., Badaloo, A.V., Boyne, M.S., Osmond, C., Thompson, D., Green, C., Taylor-Bryan, C., Barnett, A., Soares-Wynter, S., Hanson, M.A., et al. (2012). Prenatal factors contribute to the emergence of kwashiorkor or marasmus in severe undernutrition: evidence for the predictive adaptation model. PLoS One \u003cem\u003e7\u003c/em\u003e, e35907. 10.1371/journal.pone.0035907.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"kwashiorkor, marasmus, genomics, 1-carbon metabolism, admixture, local ancestry","lastPublishedDoi":"10.21203/rs.3.rs-6890799/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6890799/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eSevere acute malnutrition (SAM) contributes to the death of millions of children under age five annually. SAM is clinically classified as non-edematous SAM (NESAM) or the more severe edematous SAM (ESAM), which is more common in east-central Africa and the Caribbean. The reason some children develop ESAM while others develop NESAM remains unclear; however, recent studies have identified aberrant one-carbon metabolism (OCM) in ESAM relative to NESAM. Here, we assess genetic variants at 103 loci known to influence OCM, and determine their association with ESAM in 711 samples from Jamaica and Malawi. Seven OCM loci showed evidence of association across both populations, including five associated with homocysteine and folate metabolism (\u003cem\u003eMTHFR\u003c/em\u003e, \u003cem\u003eAHCYL1\u003c/em\u003e, \u003cem\u003ePRICKLE2\u003c/em\u003e, \u003cem\u003eGABBR2\u003c/em\u003e, and \u003cem\u003ePLD2\u003c/em\u003e). Three SNPs in \u003cem\u003ePLD2\u003c/em\u003e, \u003cem\u003ePRICKLE2\u003c/em\u003e, and \u003cem\u003eGABBR2\u003c/em\u003e, genotyped using cell-free DNA from serum metabolomic samples, supported causal effects on ESAM risk through homocysteine-related metabolites. Cumulatively, OCM-related variants showed more association with ESAM than expected by chance (z\u0026thinsp;=\u0026thinsp;3.06), with differing effect magnitudes in the two populations. By leveraging chromosome-level patterns of intracontinental African admixture, we demonstrate that OCM variant associations with ESAM occur on a shared east-African ancestral genetic background. Finally, using whole genome sequence data from eight African populations, we demonstrate that several OCM loci have outlier signatures of selection in multiple populations, including the ESAM-associated \u003cem\u003ePLD2\u003c/em\u003e locus. These findings strengthen support for aberrant OCM in ESAM pathogenesis, with implications for current interventions, and highlight the potential of cell-free DNA, intra-continental admixture, and population genetics in mapping disease risk.\u003c/p\u003e","manuscriptTitle":"Aberrant One-Carbon Metabolism and Ancestral Genetics Underlie Edematous Severe Acute Malnutrition","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-06-29 14:49:33","doi":"10.21203/rs.3.rs-6890799/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"f62b7a5e-528a-40fa-88b5-8b5a673872a1","owner":[],"postedDate":"June 29th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":50555757,"name":"Health sciences/Medical research/Genetics research"},{"id":50555758,"name":"Health sciences/Diseases/Nutrition disorders/Malnutrition"}],"tags":[],"updatedAt":"2026-04-01T13:43:23+00:00","versionOfRecord":[],"versionCreatedAt":"2025-06-29 14:49:33","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6890799","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6890799","identity":"rs-6890799","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0