Increased homozygosity due to endogamy results in fitness consequences in a human population.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by gemini-2.5-flash-lite, 2026-07-29

This study found that higher genome-wide homozygosity due to background identity by descent, not consanguinity, was significantly associated with lower completed fertility in Namibian Himba women.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Recessive alleles have been shown to directly affect both human Mendelian disease phenotypes and complex traits. Pedigree studies also suggest that consanguinity results in increased childhood mortality and adverse health phenotypes, presumably through penetrance of recessive mutations. Here, we test whether the accumulation of homozygous, recessive alleles decreases reproductive success in a human population. We address this question among the Namibian Himba, an endogamous agro-pastoralist population, who until very recently practiced natural fertility. Using a sample of 681 individuals, we show that Himba exhibit elevated levels of "inbreeding," calculated as the fraction of the genome in runs of homozygosity (FROH). Many individuals contain multiple long segments of ROH in their genomes, indicating that their parents had high kinship coefficients. However, we do not find evidence that this is explained by first-cousin consanguinity, despite a reported social preference for cross-cousin marriages. Rather, we show that elevated haplotype sharing in the Himba is due to a bottleneck, likely in the past 60 generations. We test whether increased recessive mutation load results in observed fitness consequences by assessing the effect of FROH on completed fertility in a cohort of postreproductive women (n = 69). We find that higher FROH is significantly associated with lower fertility. Our data suggest a multilocus genetic effect on fitness driven by the expression of deleterious recessive alleles, especially those in long ROH. However, these effects are not the result of consanguinity but rather elevated background identity by descent.
Full text 36,457 characters · extracted from pmc-nxml · 4 sections · click to expand

Results

We find elevated levels of homozygosity in the Himba as measured by F ROH for three different minimum length thresholds: 500 kb, 1,500 kb, and 5,000 kb ( Fig. 1 ). The lengths of ROH segments reflect the time since a common ancestor, where shorter ROH result from events occurring in the more distant past and longer ROH result from more recent events. In an idealized outbred population, the F ROH expectation for the offspring of first cousins would be 0.0625, and F ROH calculated using a minimum threshold of 1,500 kb to call ROH is comparable to a pedigree estimate of inbreeding ( 14 ). The mean F ROH 1,500 of 2.6% observed in the Himba, and first reported in Swinford et al. ( 37 ), is therefore between the values expected for the offspring of second cousins and offspring of first cousins in an outbred population ( 15 ). However, this measure is not applicable to a population like the Himba who have elevated levels of IBD sharing (see below). Therefore, to estimate the F ROH expectation for the offspring of Himba first cousins, we added 0.0625 to the average amount of background relatedness in the population resulting from more distant demographic history. To calculate this background level of inbreeding, we set a low (500 kb) threshold to call ROH and calculated F ROH for all individuals. We then filtered out the individuals with F ROH levels greater than 0.0625 and calculated the average level of F ROH in the remaining individuals (0.0420). This background F ROH level was then summed with 0.0625, resulting in an expectation of F ROH = 0.1045 for the offspring of Himba first cousins. Only one individual met this expectation (individual 14 in Fig. 5 ). Interestingly, this individual’s parents share IBD consistent with that of Himba first cousins; however, they are not genealogical first cousins themselves ( IBD Sharing between Couples ). Although a first-cousin preference in arranged marriages has been reported, we did not find any parental pairs that were first-cousin relatives. Regardless, the presence of long (>5,000 kb) ROH segments in many individuals and the distribution of F ROH values resulting from these long ROH segments alone suggests an effect of recent demographic changes. F ROH distributions for varying thresholds to call ROH and the relationship between the number of ROH segments and the total amount of ROH for all Himba (n = 681). The Top panel displays individuals’ measurements resulting when F ROH is calculated using a minimum threshold of 500 kb to call ROH, and the Middle and Bottom panels are the values calculated using minimum thresholds of 1,500 kb and 5,000 kb, respectively, to call ROH. Elevated levels of homozygosity can also result from a population bottleneck. Based on knowledge of historical events, we hypothesized that Himba experienced a recent bottleneck within the past ~6 generations ( 38 ). These historical events included heavy cattle raiding beginning in the second half of the 1800s, forcing many Himba out of northern Namibia. In 1897, a severe rinderpest epidemic decimated up to 90% of livestock herds. Contagious Bovine Pleuropneumonia also caused devastating cattle losses in the 1930s and continued to impact herds for at least the next 50 y. Several severe droughts have also occurred throughout the 20th century. Additionally, in the first two decades of the 20th century, the Himba experienced the pressures of genocide, harsh taxes, and decreased mobility and isolation ( 38 , 39 ). To test this hypothesis of recent bottleneck, we selected 120 unrelated individuals and estimated the effective population size (N e ) for the last 100 generations using a nonparametric method that uses inferred identical-by-descent (IBD) segments between pairs of individuals ( 40 ). We optimized the parameters used to infer IBD with a pipeline described by Gopalan et al. ( 41 ) ( Materials and Methods and SI Appendix , Methods ). Our estimation of N e through time indicates that a bottleneck has occurred throughout approximately the past 60 generations, reaching a minimum effective population size of approximately 450 individuals 12 generations ago (ga) ( Fig. 2 A ). Himba exhibit elevated levels of total ROH in the genome compared to other African populations, including representatives from eastern, southern, and western regions: Bench, Zulu, Igbo, and Mandinka. Instead, their distribution is more similar to that of the Chabu foragers from Ethiopia. The Chabu, who are currently experiencing a bottleneck that is estimated to have begun slightly more recently than 60 generations ago ( 41 ), have a higher mean value of total ROH but exhibit a range of values similar to those in the Himba ( Fig. 2 B ). Identical by Descent Segments indicate a recent population bottleneck for the Himba. ( A ) Effective population size through time was estimated for a sample of unrelated Himba individuals (n = 120) using IBDNe. The 95% CI shown in blue. ( B ) To understand whether ROH resulting from IBD sharing in the Himba was unusual, we contrasted them with two Ethiopian populations (Bench and Chabu), two western African populations (Igbo and Mandinka), and the Bantu-speaking Zulu from South Africa. The distribution of total amount of ROH in individuals’ genomes for ROH segments 500 kb and longer is depicted as a violin plot. The Himba are comparable to the Chabu, an Ethiopian hunter-gatherer population which has also experienced recent population decline (Gopalan et al.). Strikingly, they have much higher ROH than the Zulu. ( C ) Bottleneck inference results from HapLD for unrelated Himba individuals (n = 120). 95% CI are shown in blue. ( D ) Himba (blue) and Zulu (red) bottleneck inference results from HapLD with 95% CI. In order to validate our inferences of effective population size history, we used two additional methods to estimate changes in effective population size and founder events. HapLD is a recently developed method that relies on patterns of linkage disequilibrium to infer N e through time ( 42 ). The results are similar to those obtained using IBDNe. Although the absolute estimates of N e vary, they display the same overall shape, indicating that a bottleneck began approximately 87 generations ago and reached a minimum 18 generations ago ( Fig. 2 C ). Because the Himba and the Zulu are both southern African populations of the Bantu expansion, we compared N e trajectory results for the two populations. The Zulu share a similar ancestral N e as the Himba until approximately 38 generations ago and then begin rapid growth while the Himba N e continues to decline ( Fig. 2 D ). Our bottleneck estimates are also consistent with results obtained from ASCEND, a method that infers the timing and strength of founder events by looking at the correlation between pairs of SNPs and compared to an outgroup to account for ancestral allele sharing ( 43 ). ASCEND estimates that a founding event occurred in the Himba 16 to 17 generations ago with an intensity (calculated as the duration of the bottleneck divided by twice the effective population size during the bottleneck) between 2.5% and 2.7% ( SI Appendix , Fig. S1 ). This level of intensity is greater than those experienced by the Jewish populations analyzed by Tournebize et al. ( 43 ). The estimated timing of this founder event is thus similar among all three analyses, ranging between 12 and 18 generations ago. Taken together, these results support the occurrence of a recent bottleneck in the Himba, but this event predates recorded epidemics and warfare. We also used simulation to help validate the results of our inferred bottleneck, we simulated two bottlenecks under different scenarios using msprime and compared the IBDNe and F ROH distribution results of the simulations with the actual Himba data. The simulated bottlenecks specified a beginning N e of 4,000 and final N e of 450 to roughly reflect the starting and minimum N e values inferred from the 120 unrelated Himba individuals. In the two scenarios, the starting N e value was set to begin declining at either 60 or 6 generations ago. We find the IBD inferred from the simulated bottleneck beginning 60 ga more closely matches our real data than the one inferred from the simulated bottleneck beginning 6 ga ( SI Appendix , Fig. S2 ). We also calculated the F ROH 1500 distributions for both sets of simulated individuals as well as for our subset of 120 unrelated Himba individuals. We find the F ROH distribution of our 120 unrelated Himba individuals more closely matches the F ROH distribution of our simulated individuals experiencing a bottleneck beginning 60 ga ( SI Appendix , Fig. S3 ). The mean F ROH 1500 value for the simulated individuals is 0.013 and 0.026 for the 6 and 60 ga simulations, respectively, compared to a value of 0.036 in the set of unrelated Himba. However, both distributions of F ROH values for simulated individuals differ from the distribution of values for unrelated Himba (6 ga bottleneck: P = 2.2e-16, 60 ga bottleneck: P = 0.004). To further investigate the cause of ROH, we looked at the relationship between F ROH 1500, which has been shown to be comparable to a pedigree estimate of the inbreeding coefficient ( 14 ), and F IS ( SI Appendix , Fig. S4 ). F IS represents the departure from random mating in the population, where F IS = 0 indicates random mating, F IS 0 indicates consanguinity ( 36 , 44 ). We find that the population average is equal to −0.0035 and not significantly different than zero ( P = 1.343e-10) and 67% of people have a negative F IS . Additionally, F ROH 1500 was greater than F IS for all individuals. In contrast, populations with high rates of consanguinity lie on the diagonal where F ROH = F IS . These findings suggest that, while some consanguinity is still present ( IBD Sharing between Couples ), it is low N e that is more responsible for higher homozygosity. To assess the effects of F ROH on fertility, we identified postreproductive women in our sample, which we defined as all women greater than 47 y old—as this was the oldest age recorded at which a woman gave birth in this community—and performed a regression fit to a Poisson distribution using year of birth (YOB) and number of marriages (NM) as covariates. We used a proxy measure of reproductive success, defined as the number of children who survived to a minimum age of five. We ran two versions of the model, one including all women regardless of parity, and the other excluding nulliparous women. Many biological conditions, unrelated to recessive load, can affect fertility. In addition to untreated venereal disease ( 45 , 46 ), conditions such as polycystic ovarian syndrome and endometriosis can affect a woman’s ability to become pregnant ( 47 , 48 ). F ROH predicts the total number of children surviving to age five for all three minimum thresholds of ROH in models that contain only women who have given birth (n=65, Pr[β 99%) ( Figs. 3 and 4 ). Our sample of all postreproductive women included four seemingly infertile women (i.e., women who have had zero births). Because we lack medical records that could confirm diagnoses for reduced fertility or sterility, we ran the models again with the addition of the four seemingly sterile women to avoid possible bias. When all women are included in the model, the signal is attenuated, but F ROH continues to have a negative effect on completed fertility for all three minimum thresholds of ROH (n = 69, Pr[β 98%) ( Figs. 3 and 4 ). Examination of potential outliers flagged one individual with high Pareto k values (>0.5) in two models; however, removal of this individual did not alter model results nor the influence of F ROH on model outcomes (Pr[β 95%). Effects of predictors for fertility models. Posterior distributions (posterior median with 66% (thick line) and 95% (thin line) credible intervals) for variables used to predict completed fertility. F ROH predicts lower fertility. Posterior predictions of Bayesian Poisson regression models predicting fertility in fertile women are shown with 50, 80, and 95% intervals. The line indicates the posterior median. Raw data are also shown. While cousin marriage is preferred among the Himba, especially for arranged marriages, “love matches” are also common, particularly following the first marriage ( 49 ). Additionally, it is common and socially acceptable for both husbands and wives to take additional partners ( 50 ). Children born outside of marital unions, either through concurrency or out-of-wedlock, are referred to as omoka . When a woman is married to a man, he becomes the social father of all her children, regardless of biological paternity. Despite this, Himba recognize distinctions between social and biological paternity ( 51 ). To assess the effect of different relationships on the distribution of F ROH values, we analyzed the distributions of IBD sharing between different categories of parental couples in a subset of trios (n = 105) for whom we had marriage data, self-reported kinship for married couples, and were able to verify genetic paternity of offspring. Conditional on a couple having confirmed biological offspring together, trios were divided into three categories: A. married couples self-reported to be unrelated (n = 7), and B. married couples self-reported to be related ( Materials and Methods ) (n = 26). C. mothers and boyfriends (n = 72). Of these self-reported related couples in “B”, 15 were described as first cousins, 3 were described as first cousins once removed, 1 was described as avuncular, 1 was described as avuncular once removed, and the other 6 were described as more distantly related. We found no significant difference ( P > 0.085) between the distributions of IBD sharing between couples within each category (average IBD between biological parents of omoka children = 291 cM, average IBD between unrelated married couples = 347 cM, and average IBD between related married couples = 284 cM). Among married couples purported to be related, the majority of relationships were described as first cousins or first cousins once removed. The average IBD sharing between first-cousin pairs (n = 291 pairs) and half-cousin pairs (n = 957) in the Himba, identified using PONDEROSA ( 34 ), is 1,157 cM and 725 cM, respectively. Thus, while we have many examples of first cousins in the dataset, we observe much lower levels of IBD sharing between married couples. Despite reported relationships, individuals do not appear to marry biologically close kin. However, this does not preclude higher IBD sharing in some couples. In fact, the parental couples with the highest IBD sharing in our dataset are extra-pair relationships (boyfriend/girlfriend) in which one couple shares a total of 1,038 cM (individuals 9 and 11 in Fig. 5 ) and the other shares a total of 1,074 cM (individuals 9 and 10 in Fig. 5 ). These couples have IBD consistent with a third-degree relationship but, after reconstructing their pedigree ( Materials and Methods ), they are not first-cousins. Rather, they are related through multiple distant relationships. The women (individuals 10 and 11) are paternal half-siblings. They are both half-cousins with the man (individual 9) through their fathers (individuals 4 and 7). Additionally, woman 11 shares a fourth-degree relationship with the man’s paternal grandfather, and the man shares a fourth-degree relationship with woman 11’s mother. Woman 10’s mother also shares a cryptic relationship with the man’s father ( Fig. 5 ). This type of pedigree, where two individuals are connected through multiple distant relationships or “reticulations,” is common within the Himba. Partial pedigree reconstruction for a couple with high IBD sharing. To parse the relationship between parental couples with high IBD sharing, we reconstructed their pedigrees. Parents 9 and 11, as well as 9 and 10, are half-cousins through their fathers, individuals 4 and 7. Unsampled individuals are shown in gray. The pairwise shared IBD for additional cryptic relationships are represented with colored dashed lines and corresponding values. Additional offspring, siblings, and partners are not shown for the individuals represented here. Individual 9 is represented twice (connected by a black dashed line) for clarity in the representation of the pedigree. To help quantify this, we analyzed close reticulations only. Many Himba relative pairs exhibit high levels of IBD2 sharing that indicate that they are related through both parents. There are several confirmed cases: double cousins (co-co), double half-cousins (hco-hco), half-siblings/half-cousins (hs-hco), half-siblings/cousins (hs-co), cousin/half-cousin (co-hco), and even double half-avuncular pairs (hav-hav). There are 22 of these relationships confirmed ( SI Appendix , Fig. S6 ), but there are likely more as these can only be confirmed in families with four generations of genotyped individuals. Additionally, we only considered close reticulations here, but more distant reticulations have been observed in the data. Focusing on half-siblings, we took simulated hs, hs-co, hs-hco, and half-sibling/second-cousins and used them to train a linear discriminant analysis classifier, which we used to classify real Himba half-siblings as hs only, hs-co, hs-hco, or hs-sco. We classified 34 of the 835 half-siblings as being either hs-co (n = 7) or hs-hco (n = 11) or hs-sco (n = 16) ( SI Appendix , Fig. S7 ). As we only analyzed half-sibling pairs, this suggests that a minimum of 4% of pairs are closely related through both parents.

Discussion

Long runs of homozygosity in human populations can be generated by consanguinity as well as strong founder effects ( 14 , 52 – 54 ). Using a minimum threshold of 1,500 kb to call ROH, the average fraction of the Himba genome in long runs of homozygosity is 2.6%. Although we observe that elevated levels of F ROH 1500 and Himba social norms promote consanguineous marriage ( 55 ), we do not observe any first-cousin parental pairs in our dataset even among married couples who report consanguinity. Rather, our results indicate that the Himba of northern Namibia experienced a population bottleneck that reached its minimum effective population size between approximately 12 and 18 generations ago. Although historical data suggest a more recent bottleneck (within the past six generations), we were unable to detect such a bottleneck here regardless of the method used. This may be due to a lack of actual genetic bottlenecking during this time or the difficulty of the IBDNe program to estimate accurate effective population sizes for the most recent generations. The bottleneck could trace back to the Bantu expansion across southern Africa over the 2,000 y. However, total ROH in the genome is much higher in Himba than in Zulu, a southern African population derived from the Bantu expansion ( 56 ), and the Zulu exhibit a different N e trajectory compared to the Himba. Our results suggest that bottleneck and endogamy are more responsible for elevated F ROH than consanguinity. The average F IS value for the population is negative but near zero, suggesting that the Himba exhibit random mating (F IS = 0) or inbreeding avoidance (F IS < 0). This is consistent with the Himba practice of concurrency which more closely resembles random mating than do practices in other cultures. However, mating among the Himba is not completely random. In addition to any mate preference and sexual selection variables that may be at play, interviews reveal that the Himba do avoid close inbreeding by alerting paternal half-siblings of their relatedness so that they know not to take each other as a girlfriend or boyfriend. The relationship between F ROH and F IS indicates that increased ROH is due to low N e resulting from bottlenecks or founder events rather than from frequent consanguinity, and our bottleneck simulations indicate that bottleneck can produce F ROH distributions similar to what we observe in the Himba. However, some consanguinity is present in the Himba. This is illustrated in the reconstructed pedigree, where individual 9 has extra-pair relationships with two half-cousins [#10, #11] who themselves are half-siblings. We estimate that approximately 4% of half-siblings have an additional recent (second cousin or closer) reticulation. Thus, while consanguinity does occur, our results suggest that bottleneck and endogamy have contributed most to elevated F ROH even in the presence of random mating or inbreeding avoidance. Among married couples in our sample, half reported that their partner was a second- or third-degree relative. However, after reconstructing genetic pedigrees, we observed low levels of IBD sharing among married couples who are purported to be related. There are several possible explanations for the discrepancy between social and biological relatedness among consanguineous couples. First, errors in the demographic interviews could account for some differences, if individuals reported spouses as a first cousin, when in fact they were a more distant relation. However, another likely explanation is that a high rate of extra-pair paternity has led to an untethering of social and biological relatedness. Himba have a strong cultural tradition of sexual concurrency, with most adults having both marital and nonmarital partners ( 50 ), and previous analyses have shown that Himba have an extra-pair paternity rate of 48% ( 33 ). Concurrency is presumed to be a longstanding practice, as it was written about in the first ethnographies of this group in the early 20th century ( 57 – 59 ). This means that if a man marries his cross-cousin (e.g., his mother’s brother’s daughter), there is only about a 50/50 chance she is his biological cousin, assuming that the brother is the mother’s full sibling. There is a less than 50% chance given that the brother may have a different father, and then, he may also not be the biological father of the woman the man is marrying. With successive generations of concurrency and extra-pair paternity, the chance of relatedness reduces even further. Over time, this could lead to the results we show here. This interplay between concurrency and a preference for consanguinity may thus allow Himba to reap the benefits of a densely connected kin network, while minimizing the costs of higher mutation load that might typically come with consanguineous marriage. The presence of long ROH is particularly important in assessing the effects of mutation load on fertility. Increased risk for complex diseases has been shown in populations where IBD sharing is elevated or consanguinity is common, suggesting a causal role for multiple recessive mutations throughout the genome ( 13 ). Furthermore, long ROH have been shown to be enriched for deleterious nonsynonymous homozygous derived genotypes ( 9 , 18 , 52 ). Therefore, it has been suggested that fitness could be reduced via gene knockouts within long ROH ( 52 ). Although our study was limited in its sample size of postreproductive women, we found F ROH to have a significant effect on completed fertility. Our results show a trend of increasing effect size of F ROH on fertility when longer and longer runs of homozygosity are analyzed. This result is consistent with longer tracts carrying greater numbers of deleterious mutations or more damaging mutations such as gene knockouts. Additional variance in phenotypic effect may be caused by varying placement of ROH within the genome. Specific genomic regions may be especially harmful to reproduction if deleterious variants disrupt genes critical to proper reproductive function ( 60 , 61 ). Overall, our results suggest a multilocus effect on fitness driven by the expression of deleterious recessive alleles, especially those harbored in long ROH. Our results are consistent with earlier self-reported pedigrees studies which considered the effect of consanguinity on fertility. Chagnon et al. ( 23 ) analyzed the South American Yanomamö group and reported that children whose parents were more closely related (i.e. children who would be expected to have higher F ROH ) had significantly lower fertility themselves. Similarly, Postma et al. ( 24 ) used town records to reconstruct genealogies for individuals from a small Swiss village and reported inbreeding depression for fertility among women. Our results strengthen these findings by demonstrating this pattern in another human population using molecular genetic measurements rather than pedigree estimates of relatedness. We caution that self-reported pedigrees may not always reflect biological kinship coefficients, as discussed above. Variance in fertility within our cohort of postreproductive women is not solely due to recessive genetic effects; our models explain 16.3 to 18% of the variance indicating fertility is affected by other factors. Secondary sterility resulting from untreated venereal disease, such as gonorrhea, may be a factor in this population ( 45 , 46 , 62 ). Pennington and Harpending ( 45 ) have suggested that, prior to the 1960s, venereal disease caused lower fertility in the Namibian Herero, a population closely related to the Himba. They argue that women born after ~1945 would be at low risk for sterility due to venereal disease since antibiotics would have been available by the time they became sexually active, and women born prior to ~1915 would be at the greatest risk since antibiotics would not have been available until after they reached menopause ( 45 ). In our sample, nine women were born prior to 1945, including one of the four infertile women, and none of the women in our sample were born prior to 1915. However, more recent work by Hazel et al. ( 62 ) has shown the continued prevalence of gonorrhea infection along with low levels of treatment among Kaokaoland pastoralists, including Himba. Furthermore, this study suggests that these infections may be limiting fertility, but it is unclear to what extent, if any, this is occurring as no official measurements or analyses of the effects of gonorrhea on fertility have been made. A lack of medical records in our Himba sample makes it impossible to know whether there are other exogenous medical conditions that could be affecting fertility as well. Our results are important for understanding the architecture of complex traits in human populations by suggesting that recessive mutation load may play a key role in fitness. A recent study by Szpiech et al. ( 63 ) noted that ROHs from different ancestries had different proportions of damaging homozygotes. This is thought to be due to differing population histories, resulting in haplotypes from high heterozygosity populations (such as African populations) containing more strongly deleterious variants than those found on haplotypes from lower-heterozygosity populations, and thus being more severely damaging when found in ROH. The effects of F ROH on fertility and other polygenic traits may then differ between populations. Thus, this work should ideally be replicated in several populations of varying demographic histories. Future work pertaining to the effects of F ROH on fertility should also annotate variants found in ROH to assess their level of deleteriousness. In conclusion, this work is especially important in demonstrating how differences in mutation loads, especially when considered under a recessive model, may affect differences in evolutionary fitness.

Materials|Methods

The Himba are a small Bantu-speaking population that resides in Kunene Region, northwestern Namibia. They are a seminomadic agro-pastoralist group who practice polygyny with a reported first-cousin preference for arranged marriages; however, “love matches” are also common. Additionally, it is common and socially accepted for both husbands and wives to take additional partners ( 32 ). Within the past 200 y, the Himba have experienced many factors that have likely contributed to population decline, including genocide, climate change, severe drought, and rinderpest epidemics that decimated cattle resources ( 38 ). The data in this study were collected following ethical approval granted by the University of California, Los Angeles (IRB-10-000238) and the State University of New York, Stony Brook (IRB-636415-12). The study was also approved by the Namibian Ministry of Home Affairs and supported by the University of Namibia Office of Academic Affairs and Research, and local community approval of the study was granted by Chief Basekama Ngombe. These data were collected as part of the Kunene Rural Health and Demography Project, which has been working in the community since 2010. Community leaders were actively involved in discussions regarding the research, including who could access the genetic data and what it could be used for, prior to data collection (with initial consultation in 2013 by BMH, BAS, followed by subsequent consultation after DNA collection in 2016). All data and samples were collected with informed consent and parental assent for minors. Individuals were genotyped on either the MEGAex or H3Africa SNP array. Quality control and filtering proceeded as documented in dbGaP phs001995.v1.p1 as part of Scelza et al. ( 33 ), which included using PLINK to filter for missingness greater than 5%, a minor allele frequency less than or equal to 1%, and a Hardy–Weinberg equilibrium exact test with a P -value below 0.0001. We obtained genetic data from three more individuals on the H3Africa array and added them to the dataset after initial QC steps by filtering for SNPs common between the original dataset and the three new individuals. Both datasets were filtered to contain autosomes only to avoid X-chromosome interference in F ROH estimates for males. Genotype data from the H3Africa array were thinned using PLINK2/1.9 to match the SNP density of the MEGAex array, therefore ensuring consistency when calling ROH and allowing for higher SNP density than would be present after merging the datasets due to missingness across platforms. The final H3Africa dataset contained 504 individuals and 755,660 SNPs, and the final MEGAex dataset contained 177 individuals and 755,423 SNPs. We performed a test to ensure that randomly thinning the SNP density did not affect the estimation of an individual’s F ROH . To do this, we thinned the H3Africa dataset 10 separate times, identified ROH in each set as before with a minimum segment length threshold of 1,500 kb, and recalculated F ROH in each set. We then calculated the absolute value of the differences between the F ROH values in each new thinned set and in the original thinned set for each individual. We averaged the 10 values for differences between tests for each individual, resulting in an average difference for F ROH calculated from differently thinned SNP array data for each individual, and used these values to calculate a root-mean-square error (RMSE) of 0.00058, confirming no significant difference between tests. We identified ROH and calculated FROH as in Swinford et al. ( 37 ). To calculate F ROH , we first calculated the length of the genome tested in the SNP arrays by summing the lengths between the first and last SNPs genotyped on each chromosome. We then identified ROH in each individual using PLINK2/1.9 with a scanning window of 50 SNPs and allowing for a maximum of 2 missing SNPs and 1 heterozygote. We identified ROH for each of three different minimum length thresholds—500 kb, 1,500 kb, and 5,000 kb—and then calculated F ROH for each of the three thresholds by dividing the total length of the genome found to be in ROH by the total length of the genome tested in the SNP array. These calculations were done separately for both platforms and then combined in R, resulting in F ROH estimates for a total of 681 Himba individuals. We also used PLINK/1.9 to calculate F IS separately for both platforms and then combined the datasets as was done for identifying ROH. We used a one-sample t-test in R to determine whether the average F IS of the population was significantly different from zero. Reproductive success was measured as the number of children a woman had who survived to a minimum age of five years old and these data were collected during interviews. In populations where infant and child mortality are high, survival to age five is used as a proxy for likelihood of survival to adulthood. In addition, still births and infant mortality are sensitive topics, which women were often reticent to report. Thus, a measure of fitness based on the number of births alone is both less reliable and less pertinent to overall reproductive success than number of children surviving to age five. To assess the effects of F ROH on fertility, we performed a Bayesian Poisson regression in R, fitted using the brms() package ( 64 ). We assessed the effects of F ROH on fertility for three different standardized measurements of F ROH (using minimum thresholds of 500 kb, 1,500 kb, and 5,000 kb to call ROH) and used the measure of total number of children surviving to age five as the response variable. Standardized YOB and NM were used as covariates in the modelling process. Regularizing priors were used for all predictors. Pareto-smoothed importance sampling cross-validation (PSIS) was used to evaluate the impact of potential outliers in the data by calculating and flagging Pareto k values ( 65 ). When Pareto k values were above 0.5, the data in question were removed and model rerun. Where necessary, we report the probability of a positive or negative effect for predictors (Pr[β 0]). To determine the effective population size through time, we first identified a subset of unrelated individuals (n = 120) and ran IBDNe ( 40 ) specifying a minimum centimorgan length of 4 cM and limited the number of generations before present to calculate N e for to 100 (gmax 100). The output was graphed in R. Our msprime simulations followed Gopalan et al. ( 41 ) and specified a starting N e of 4,000, a final N e in the present of 450, the generation at which to begin the bottleneck (60 or 6 ga), the number of haploid individuals to simulate (n = 120), and a mutation rate of 1e-8. We simulated 120 individuals and identified IBD in the output using hap-ibd and then ran IBDNe to reconstruct inferred N e through time. Additionally, we ran HapLD on our 120 unrelated Himba and on 99 unrelated Zulu for comparison ( 42 ). We also ran ASCEND to estimate the timing of a founder event in the Himba, using the Zulu as an outgroup ( 43 ). Additional information on methods, procedures, and parameters can be found in supplement. Omoka status was determined and confirmed with genetic paternity analysis as described in Scelza et al. ( 33 ). Children were described as omoka if their biological father was not their mother’s husband at the time of their birth. Information regarding marital unions and the descriptions of married couples’ relatedness were collected in interviews and then translated into a specific category of relatedness. For example, a description of “mother’s brother’s daughter” was translated to the category of first cousins, while descriptions such as “father’s brother’s daughter’s daughter” or “the husband’s mother is from the wife’s paternal uncle” were translated to the category of first cousins once removed. One couple was described to be related in two different ways, as both first cousins once removed and possibly as third cousins once removed as well. However, it was unclear to what extent these two connections were independent or overlapping so they were included in the category of the closer of the two relationships (i.e., first cousins once removed). For pedigree reconstruction, we used PONDEROSA, an algorithm that infers pedigree relationships and is especially suited for populations with elevated IBD sharing ( 34 ) to identify all pedigree relationships in the data. To estimate the prevalence of close reticulations, we used Ped-Sim ( 66 ) to simulate many multirelationship types and then trained a linear discriminant analysis classifier, which we used to classify the types on consanguinity in the pedigrees of Himba half-siblings. Additional information and parameters can be found in supplement.

Supplementary Material

Appendix 01 (PDF) Click here for additional data file.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-08-23T09:30:01.253652+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-ND-4.0