Full text
92,368 characters
· extracted from
preprint-html
· click to expand
Revisiting the evidence for long-lived balancing selection in humans | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Revisiting the evidence for long-lived balancing selection in humans View ORCID Profile Hannah Munby , View ORCID Profile Molly Przeworski doi: https://doi.org/10.1101/2025.11.10.687682 Hannah Munby 1 Dept. of Biological Sciences, Columbia University Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Hannah Munby For correspondence: hmm2183{at}columbia.edu Molly Przeworski 1 Dept. of Biological Sciences, Columbia University 2 Dept. of Systems Biology, Columbia University Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Molly Przeworski Abstract Full Text Info/History Metrics Preview PDF Abstract Balancing selection maintains variation in a population longer than expected under neutrality. In humans, there are dozens of tentative candidate loci for balancing selection, but only a handful of well-characterized examples, which are either evolutionarily recent alleles or ancient variants shared across species identical by descent (“trans-species polymorphisms”). Here, we look for evidence of balancing selection over a range of timescales, by taking an approach that does not rely on a demographic model or assumptions about the specific mode of balancing selection. Analyzing whole genome sequencing data from 2504 humans and 59 chimpanzees, we identify common single nucleotide polymorphisms (SNPs) that are identical in the two species. This set includes recurrent mutations, a subset of which may be maintained by balancing selection in one or both species, as well as potential trans-species polymorphisms. Using allele ages estimated from ancestral recombination graph reconstructions in humans, we show that shared SNPs are enriched for older alleles as compared to matched human SNPs that are not shared with chimpanzees. On this basis, we estimate that balancing selection has maintained over one thousand alleles in humans longer than expected by chance. Moreover, we identify over 50 trans-species polymorphisms, including an intriguing case that includes an eQTL for the gene MUC7 . However, we also estimate a minimum false discovery rate for any allele age cut-off of ∼70%; as we show, even among the trans-species polymorphisms, many may be shared between humans and chimpanzees simply by chance. Thus, while our empirical approach establishes that there are numerous loci under balancing selection yet to be found, the specific targets remain difficult to identify without independent lines of evidence. Introduction Genetic differences among individuals are shaped by both neutral and selective forces. In studying these genetic differences, much attention has been given to the study of directional selection, which favors a single genetic allele over others and acts to reduce variation at linked sites [ 1 , 2 ]. But it is also possible that variation is in itself adaptive, such that multiple alleles at a single locus are maintained in the population over time by natural selection [ 3 ]. There is a long history of theoretical work describing various selective regimes that result in the maintenance of diversity at a locus longer than expected under genetic drift alone, collectively referred to as balancing selection, and their effects on linked neutral diversity [ 4 ]. Cases of balanced polymorphisms include the S-locus, which leads to self-incompatibility in certain families of flowering plants [ 5 ], a structural polymorphism that controls sexual reproduction strategies in walnuts [ 6 ], and balanced inversions that maintain behavioral ecomorphs in animals, such as the three male ruff morphs [ 7 ]. Examples are few, however, and the broader questions of how much variation is actively maintained by selection, the types of the variants that are maintained, why they have been selected and over what timescales they have persisted, all remain largely open [ 8 ]. In humans specifically, despite millions of genomes now sequenced, we have only relatively few well-characterized examples of balancing selection, most of which seem to be due to heterozygote advantage. Perhaps most famously, in geographic regions endemic for P. falciparum malaria, individuals with one sickle allele and one wild-type allele have a higher fitness than either homozygote [ 9 ]. In this case, the selection pressures are known: in the heterozygous state, the sickle cell allele confers resistance to malaria but in the homozygous state, the fitness benefit provided by protective effect of the allele is outweighed by it leading to sickle cell disease. The sickle cell allele is thus maintained by strong balancing selection in these geographic regions. This polymorphism is thought to have arisen very recently in human evolution, however, only around 7000 years ago [ 10 ]. Other examples of heterozygote advantage are again thought to be in response to malaria and are similarly young (e.g. α-thallassemia [ 11 ] and G6PD deficiency [ 12 ]). The recency of these examples is perhaps not surprising. Heterozygote advantage is unlikely to be evolutionary stable, given that in each generation, there is a segregation load of heterozygous individuals having less fit homozygous offspring [ 13 ]. In the long term, this load can be alleviated by a heterogeneous gene duplication that fixes both alleles in the genome and is thus favored over the polymorphism [ 14 ]. Modeling suggests that, in contrast, other modes of balancing selection–including frequency-dependent selection, temporally or spatially-varying selection, and selection driven by interaction with other species–have the potential to maintain alleles for much longer [ 15 – 17 ]. While such long-lived balanced loci are expected to be rarer, they are the result of strong and persistent selective forces and are thus of importance in understanding adaptation, as well as in identifying genomic regions in which variation is of functional importance. The identification of long-lived balanced polymorphisms often relies on the impact that selection has on the underlying genealogy, specifically on the fact that balancing selection leads the time to the most recent common ancestor (TMRCA) at the target locus to be older than expected under neutrality (or purifying selection). This property causes a signal of increased diversity at neutral sites linked to the selected locus, an observable pattern used by many approaches to detect selection in genomic data [ 18 ]. However, the exact distortion of the underlying genealogy can depend on the mode of balancing selection: as an example, the effects of fluctuating selection can be distinct from those of overdominance or frequency-dependent selection [ 19 – 21 ]. Moreover, these models tend to assume a constant size, panmictic population, even though changes in population size and population structure will also shape the genealogy [ 22 – 24 ]. More generally, existing methods to detect balancing selection tend to rely on strong assumptions about demography and the mode of balancing selection, as well as about mutation and recombination rates. For example, likelihood-based methods (e.g., [ 25 – 27 ]) assume a null model of neutrality under a known population history and an alternative model of heterozygote advantage under a constant population size. Moreover, when such methods rely on a model of heterozygote advantage, they model the balanced polymorphism at deterministic equilibrium allele frequencies, which may not capture the dynamics under alternative forms of balancing selection [ 28 , 29 ]. Other model-based methods make fewer assumptions about the mode of selection but assume the data come from an unstructured population (e.g., [ 30 ]) and are sensitive to model misspecification [ 8 ]. To overcome these limitations, one approach is to focus on regions of the genome with the highest inferred TMRCAs, which are plausibly enriched for targets of balancing selection [ 31 ]; however, the lack of a clear null distribution to compare against makes it difficult to draw firm conclusions about whether these tail cases are truly the targets of selection. There is, however, a special case thought to provide unequivocal evidence for balancing selection [ 32 ]: when the polymorphism persists long enough to lead to the maintenance of a variant from an ancestral species through to the present day, and thus to allele sharing between relatively divergent species through identity by descent (IBD). In humans, as well as in other species, such trans-species polymorphisms (TSPs) with chimpanzees or more distant evolutionary relatives have long been held as a gold standard [ 4 , 33 ], because of estimates suggesting that it would be highly unlikely for alleles to be shared IBD between humans and other primate species by genetic drift alone [ 33 ]. Moreover, well-established examples of balancing selection, such as occurs at loci in the Major Histocompatibility Complex (MHC) [ 34 ] and at the ABO gene, show the clustering by allele rather than by species expected from IBD: for example, in ABO , a small genomic region differs more between humans carrying alleles A versus B than it does between humans and gibbons carrying allele A [ 32 ]. Beyond these two cases, only a handful of TSPs are known between humans and other primate species [ 35 – 37 ]. Evidence that many of these examples are under selection due to host-pathogen interactions [ 37 – 39 ], as well as theoretical reasons to expect such interactions to maintain host variation [ 40 ], have led to the hypothesis that host-pathogen immune functions are the main force driving long-term balancing selection in great apes [ 41 ]. A challenge in identifying TSPs in humans is that the region around the site under balancing selection is expected to be very short, on the order of 100-1000 base pairs (bps) for human recombination rates [ 33 ]. Therefore, such examples can be missed if summary statistics or tree reconstructions are based on larger window sizes [ 33 ]. The search for TSPs and other long-lived balanced polymorphisms has further been made difficult by the high rates of false positives, generated by recent recurrent mutations, especially at CpG sites, and from genome misassemblies (e.g., duplicates being collapsed) [ 35 ]. Recent improvements in genome assemblies alleviate some of these issues [ 42 , 43 ]. Moreover, scalable ancestral recombination graph (ARG) reconstruction methods have now been developed that make it possible to study selection by inferring local genealogies along the genome and analyzing those, rather than relying on summary statistics, such as local diversity levels, in a fixed window size [ 44 , 45 ]. ARG reconstruction methods also infer mutation ages, which can be used to compare the relative ages of alleles. We take advantage of these recent methodological developments in order to look for evidence of long-term balancing selection in humans by an empirical approach that does not rely on assumptions about the mode of selection or demographic history–only on the property that balancing selection maintains genetic variation longer than expected under neutrality. As a first step, we identify the set of intermediate frequency variants that are shared between a set of 2504 humans [ 46 ] and 59 chimpanzees (see Methods). Then, using allele age estimates from ARG reconstructions in humans, we ask whether this set of variants is significantly enriched for older alleles relative to matched control SNPs found in humans but not in chimpanzees. This enrichment could be due to either TSPs or to mutations that arose independently in the two species and have been maintained longer than expected in humans. By reconstructing local genealogies at the oldest of these shared variants, we identify which are likely shared IBD, and thus identify a set of novel TSPs between humans and chimpanzees. Motivated by recent findings about human and chimpanzee demographic histories [ 47 ], we also reassess the evidence that these TSPs are targets of long-term balancing selection. Results An empirical approach to uncovering evidence for balancing selection We first obtain a set of polymorphisms that are segregating in both humans and chimpanzees, which we use to test if such shared variants are older than expected under an empirical null. To survey genetic variation in chimpanzees as comprehensively as possible, we compile publicly available, whole genome sequences from population sequencing [ 48 ] and three pedigree studies [ 49 – 51 ]. After adding 10 newly sequenced samples from a three generation pedigree, we have a total of 59 individuals. We remove first and second degree relatives and are left with 43 samples, including individuals assigned to one of three chimpanzee subspecies: 15 western (ssp. verus , n=15), 14 central (ssp. troglodytes, n=14) and 14 eastern (ssp. schweinfurthii ) (see Methods and Supplementary Figure 1). We jointly call and filter genotypes for these samples using the latest chimpanzee reference genome (PanTro6) before lifting variants to the human reference. Specifically, we reciprocally lift variants (as well as the surrounding +/- 100bp regions) between the two reference assemblies to keep only those that map consistently. The final set of chimpanzee variants is then intersected with the biallelic SNPs called in the 2,504 unrelated human samples of the 1000 Genomes Project phase 3 (1000GP) dataset. After removing variants in regions with known copy number variants and segmental duplications (see Methods), we identify 66,650 biallelic SNPs that are polymorphic and present at an appreciable frequency in both humans and chimpanzees: ≥5% minor allele frequency (MAF) in 1000GP Phase 3, and a minor allele count of three or more in our set of 43 unrelated chimpanzees. As expected, this set of shared variants is enriched for mutations that have appeared independently in the two species, as evidenced by the high proportion that are CpG transitions: 55.6% as compared to only 14.6% of non-shared 1000GP SNPs (given a MAF ≥5%). While the set of shared SNPs is likely to be predominantly neutral, it could also include variants that have been maintained by balancing selection. Cases of balancing selection among shared variants could arise from (i) alleles that persisted since the common ancestor of humans and chimpanzees, i.e., trans-species polymorphisms (ii) recurrent mutations maintained in both species (iii) a variant maintained in humans that recurred by chance in chimpanzees and (iv) a variant maintained in chimpanzees that recurred by chance in humans. In the first three cases, we expect the derived alleles to be older in humans than expected under neutrality. Therefore, if balancing selection plays a role in maintaining variation in humans, the set of shared SNPs should be enriched for old alleles relative to what is expected under neutrality. This prediction is what we test in what follows, using a set of matched controls found only in humans. To this end, we exclude the subset of shared alleles that lie in the MHC region, in which ancient balancing selection is already well established (e.g., [ 52 ]), and in which, as expected, we find 405 shared SNPs (see Methods). Evidence for a large number of targets of long-lived balancing selection in humans To examine evidence for balancing selection in the human genome, we match each shared SNP with a set of control SNPs in humans not shared with chimpanzees, then compare their estimated ages. To generate this set of controls, we want to account for factors that can influence allele ages. Most obviously, since higher frequency alleles tend to be older, we want to match the frequencies of the alleles in humans. We also want to control for the extent of purifying selection at linked sites, as sites in regions of stronger background selection will tend to have lower TMRCAs and hence younger allele ages [ 53 ]. With these considerations in mind, we identify all variants in 1000GP that are not found in our sample of chimpanzees but have the same derived allele frequency in humans (to within 1%) and inferred strength of background selection (by decile [ 54 ]) (see Methods). We verify that the test set and control sets are on average subject to very similar background selection, with a mean value of 841 for the shared SNPs and means of 841-842 across the 100 control sets. We also match the shared and control SNPs for mutation rates. We do so because conditioning on a mutation increases the branch length compared to no mutation; the lower the mutation rate, the bigger the effect [ 55 ]. Moreover, the mutation rate influences the probability of a recurrent mutation in chimpanzees. To match mutation rates, we require that the control variant be the same single nucleotide mutation type (e.g., T>G). Because at methylated CpG sites, mutation rates are much higher than at other sites [ 56 ], we also consider CpG status (see Methods). We further match SNPs based on an estimate of the mutation rate (within 2-fold, as estimated by Roulette [ 57 ]; see Methods). Having identified all possible matched controls for each shared SNP, we then sample with replacement to generate 100 unique control datasets, each containing one matched control per focal SNP. Given that we consider only derived and matched control alleles at substantial population frequency (≥5% in the human sample and three or more copies in the chimpanzee sample), we expect the vast majority of the control SNPs to be neutrally-evolving (or possibly under positive selection), rather than under purifying selection [ 58 ]. Accordingly, across our 100 control sets, the vast majority of the control SNPs considered are non-coding (99.2% on average). Moreover, only only 3.47% of sites in each control set are predicted to be conserved on average (using a PhyloP score threshold corresponding to the 90th percentile of conservation scores, as used in [ 59 ]); while this value is higher than in the test set (1.94%), it is substantially lower than what is seen for even intermediate frequency SNPs genome-wide (5.56% of 1000GP SNPs with MAF≥5%). We therefore rely on the control SNPs to provide a neutral null expectation against which we then test our shared SNPs for an excess of older alleles. To estimate the ages of test and control alleles, we rely on two state-of-the-art, genome-wide genealogical reconstruction methods, Relate and SINGER [ 60 ]. Relate scales to thousands of genomes, whereas SINGER is currently restricted to hundreds. In Relate , tree topologies are jointly estimated using the entire 1000GP dataset and branch lengths are estimated for each of the 26 subpopulations of 1000GP [ 44 ]. Therefore, allele ages can be estimated for each of these populations separately. We use estimates for a set of Yoruba sampled in Ibadan, Nigeria (YRI) and an Iberian sample from Spain (IBS); estimates are available for 82% (54,415 SNPs) of our shared SNPs for both population samples. For SINGER , the published genome-wide genealogical reconstruction relies on a subset of 1000GP, consisting of 200 whole genomes selected by drawing 40 genomes randomly from each of the five African subpopulations. Also made available are estimates of the allele age from 100 posterior samples of the ARG for 29% of our shared SNPs [ 45 ]. Although we have information on fewer SNPs using SINGER , we show results for both methods because there is evidence of a downward bias in age estimates for very old alleles in Relate [ 61 ] and because SINGER presents the advantage of reporting sample trees from the posterior distribution rather than a single fixed topology. These inference methods estimate the ages at which lineages start and end in the ancestral recombination graph rather than mutation ages. To account for the uncertainty in the placement of a mutation along a branch, we assume that the age of an allele is uniformly distributed along its estimated branch length [ 62 ]. Although for alleles under selection, this assumption is not exact, it allows us to approximate the expected total number of SNPs older than a given age by summing information across SNPs (see Methods). With increasing age, the proportion of shared SNPs that are recurrent mutations is expected to decrease. Consistent with this expectation, if we considering age estimates in YRI, the proportion of shared SNPs that are CpG sites decreases from 58.3% for the total set to 51.5% when considering only SNPs estimated to be at least 2 Mya and further to 21.1% for those at least 7 Mya, closer to the genome-wide average of 14.6% (Supplementary Figure 2). Comparing shared SNPs versus matched controls, the shared SNPs are enriched for older alleles ( Fig 1 ): for example, there are 6,388 shared SNPs estimated to be older than 2 Mya when only 5,624 of the non-shared matched controls are equally old [5,565-5,657 90% CI]. This excess–in this example, ∼760 SNPs or 12% of the total set–represents the proportion of shared SNPs that were maintained in the human species longer than expected under neutrality (i.e., based on the control set). While there are quantitative differences in the enrichment seen using age estimates from different populations (YRI vs. IBS) and different estimation methods (Relate vs. SINGER), the qualitative finding is robust to both of these choices. Therefore, we find that balancing selection has maintained a substantial number of shared SNPs between humans and chimpanzees. Download figure Open in new tab Figure 1: The expected number of SNPs above a given age for shared variants and matched controls. In the first column (panels A, C, and E), curves are shown for the shared SNP set (in red) and the median of 100 sets of matched controls (in blue), with the 5th and 95th percentile values of the control sets shaded. The first two rows are based on branch estimates obtained using the Relate software [ 44 ] for different population samples. The top row shows estimates using 108 Yoruba living in Nigeria (YRI) and the bottom using a sample of 107 from Iberia, Spain (IBS). The third row is based on an alternative ARG reconstruction method (SINGER) [ 45 ] applied to 200 genomes sampled across five locations in Africa. We note that the x-axis is truncated at 10 Mya, but 2140 shared SNPs exceed this age by this estimator. In the second column (panels B, D and F) are our empirical estimates of the false discovery rate (FDR) for any given age cut-off, calculated as the expected number of controls over the expected number of shared SNPs (see Methods). Dashed lines in each panel of the second column indicate the minimum FDR in the timeframe shown. If we think of this comparison as representing a test for the presence of balancing selection at a SNP, we can use this excess to obtain an empirical false discovery rate for a given age cut-off. Specifically, we do so by calculating the expected number of matched control SNPs older than a specific age, divided by the expected number of shared SNPs older than this age. Thus, if one considers all alleles with ages estimated to be above 2 Mya to be under balancing selection, the chance of a false discovery is 88% ( Figure 1B ). Equivalently, this estimate can be interpreted as indicating that 12% of the shared SNPs older than this age are true positives, maintained by long-term balancing selection. Although the FDR decreases with increasing allele ages, it is above ⅔ for any age cut-off. One implication is that identifying a specific human allele as old provides very weak evidence for balancing selection. Given that we do not use any criteria to exclude targets of balancing selection from our control set beyond requiring that they are not shared with chimpanzees, a subset of our control SNPs–especially the old alleles–may themselves be under balancing selection. As such, our estimates of enrichments are conservative (and our FDR estimates likely over-estimates). Nonetheless, we estimate that an expected 2,137 shared SNPs that arose over 1 Mya (90%-tile central interval [2120,2212]) have been maintained longer than expected in humans. Every one of these variants will not have been a direct target of balancing selection, as some will have been maintained by their tight linkage to selected variants. But this number appears to be on the same order of magnitude as the number of targets: for instance, if we consider the age of a SNP to be the midpoint of the branch on which it sits, 17,809 shared SNPs are estimated to be older than 1 Mya. Clustering these shared SNPs into independent regions based on proximity ( 0.2) yields 14,267 approximately independent regions. Assuming a similar proportion applies to the subset of true positives, then for alleles maintained >1 My, we expect ∼1,700 approximately independent regions. It is therefore clear that there exist many more targets of balancing selection in humans than have been characterized to date. Further, while our method provides an empirical approach to test for balancing selection, by design it is only able to do so those variants that are shared with chimpanzees. Therefore, the numbers reported are likely underestimated. Novel trans-species polymorphisms at the oldest shared variants We find many shared variants that could plausibly be older than the human-chimpanzee split time, so examined whether such variants are shared identical by descent or only by state. Focusing on the oldest estimated variants in humans, we use a combined sample of humans and chimpanzees to reconstruct the local genealogies, using SINGER . Tree reconstruction previously required an arbitrary choice of window size over which to estimate the local tree; in that setting, if the window size is chosen to be too large, it can lead to the signal being missed [ 32 , 63 ]. ARG reconstruction methods overcome this limitation by inferring the location of changes in local tree structure due to ancestral recombination events along the genome. Using SINGER , we draw samples from the posterior distribution of genealogies, which allows us to characterize the statistical support for a particular topology. To identify the oldest shared variants in our set, we rely on a point estimate of the allele age, the midpoint age of the branch on which the derived mutation lies in humans. We use ages from the Relate dataset because it provides estimates for the majority of our shared SNPs; we note that Relate is downwardly biased for old alleles [ 61 ], which is conservative for our purposes here. Specifically, we focus on the 200 SNPs with the oldest age estimates, excluding the MHC locus; in practice, age estimates above 4.82 Mya. For these 200 SNPs, we then reconstruct local genealogies on a combined sample of humans and chimpanzees, mapped to the human reference. Using 100 samples from the posterior distribution of genealogies, we examine whether the resulting local tree topologies support the hypothesis that the focal SNP is shared identical by descent between the two species, namely that the lineages cluster by allele rather than by species (see Supplementary Figure 3 for an illustration of the criteria used; [ 33 ]). Of the 200 SNPs with the oldest age estimates, seven are removed by our data processing (see Methods). Among the remaining 193, 51 meet our criteria for a TSP in >80% of ARG samples and 40 in >90% of samples. All of the TSPs are non-coding (12 intronic and the remaining 39 intergenic) and none map to known regulatory regions (see Methods). Of the 51 TSPs, 46 cluster into six regions with more than one trans-species polymorphisms ( Table 1 ); these clusters span 164-6,533 bp, plausible segment lengths for trans-species polymorphisms between humans and chimpanzees [ 33 ]. View this table: View inline View popup Download powerpoint Table 1: Genomic regions with evidence of trans-species polymorphism A previous scan for trans-species polymorphism by Leffler et al. (2013) [ 35 ], who analyzed 59 humans and 10 chimpanzees sequenced to medium coverage, reported six regions as strong candidates. Two of these regions also meet all our criteria: those assigned by Leffler et al. to genes IGFBP7 and MTRR . At the third, FREM3 , all of the previously reported TSPs pass our filters and are confirmed to be shared between humans and chimpanzees but ages as estimated by Relate narrowly miss our cutoff criterion. Considering that Relate estimates are downward biased for old alleles [ 61 ], these shared SNPs may indeed be trans-specific. The other three cases are filtered out by our approach and may be artifacts: variants assigned to HUS1 and PROKR2 violate rules of Mendelian segregation in the chimpanzee pedigrees while ST3GAL1 variants fall within a segmentally-duplicated region of the latest chimpanzee reference. In addition to the two regions previously identified by Leffler et al. (2013), we find strong evidence that four additional regions carry multiple variants identical by descent between humans and chimpanzees ( Table 1 ). Of these, one is particularly intriguing: it includes seven SNPs between genes CABS1 and SMR3A on chromosome 4, each meeting our criteria for trans-species polymorphism in 94-99% of ARG samples (example 2 in Table 1 ; Figs. 2A and B ). The TSPs in this region are tightly associated with one another and are found at frequencies of around 40-50% in all human continental groups; in chimpanzees, the average frequencies across chimpanzee subspecies are similar but the derived alleles are fixed in the Western sample ( Fig. 2C ). While the variants do not lie within any known regulatory elements, they are fine-mapped to a reported expression quantitative trait locus (eQTL) that affects the expression of the gene Mucin 7 ( MUC7 ) in the skin [ 64 ]. MUC7 is a secreted antimicrobial protein thought to play an important immune role in clearing bacteria and other microbes from the oral cavity [ 65 , 66 ]. While this protein function is of interest given the theme of host-pathogen interactions among known targets of balancing selection, further evidence would be needed to establish a functional link between our TSPs and the immune role of this gene. Similarly, other regions with multiple TSPs are also nearby or associated with immune-related genes (e.g., IGSF5 ; Table 1 ), but we do not find any clear functional links to a direct immune consequence of the shared SNPs. Download figure Open in new tab Figure 2: Example of a novel trans-species polymorphism (TSP). A . Local genealogical reconstruction using SINGER on the combined sample of 50 human and 43 chimpanzee samples for the region chr4:71,210,309 - 71,210,373 (GRCh37/hg19 coordinates). Check marks denote mutations, branches below the TSP mutations are colored pink, and the coloring at the tips of branches indicate whether the haplotype is from a human (blue) or chimpanzee (red). See Methods for details. B . LocusZoom plot of shared SNPs in this region. On the x-axis is the genomic position and on the left-hand y-axis the proportion of samples from the posterior distribution of SINGER tree topologies that support the shared SNP being shared identical by descent (“TSP support”). Points representing each shared SNP are coloured by r2 value with chr4:71210336_T, which has the highest TSP support (meeting our criteria in 99% of trees; see Methods). Recombination rate along the chromosome in HapMap is shown on the right-hand y-axis. Nearby protein-coding gene locations are shown below. C . Allele frequencies for chr4:71210336_T, in different population and subspecies samples. Many trans-species polymorphisms between humans and chimpanzees may be neutral Examples of trans-species polymorphisms between the two species are widely considered the gold standard for evidence of long-term balancing selection [ 33 ]. The reason being that when the split time of two species is old enough, the probability of a neutral TSP is low, and hence the maintenance of an allele is much more likely driven by selection. Specifically, for a neutral polymorphism to persist requires that at least two lineages in each species not have coalesced by the time of the species split (as well as that the first coalescent event be between lineages from different species and that there be a mutation on the right branch). Therefore, a neutral TSP depends on the probability ( e -T/2N1 ) × ( e -T/2N2 ), where T is the split time in generations and N 1 and N 2 are the effective population sizes of species 1 and species 2, respectively [ 67 ]. It follows that if the two species separated sufficiently long ago relative to their effective population sizes, neutral TSPs are highly unlikely. Given the widely accepted population histories of humans and chimpanzees, this condition was thought to hold [ 4 , 33 , 35 ]. New evidence from Cousins et al. (2024) cast doubt on this assumption [ 47 ]. The authors reported that a larger fraction of the human genome is uncoalesced by 5 Mya than had been previously believed to be the case. They further infer coincident peaks in the estimated effective population sizes of humans and chimpanzees >5 Mya, potentially indicative of ongoing gene flow following the initial separation of the two species [ 47 ]. Given these findings and similar mutation rates and nucleotide diversity in the two species [ 46 , 48 , 51 , 68 ], we expect that the chimpanzee genome too may harbor many regions that have not reached their common ancestor by 5 Mya. If so, there may be a non-negligible number of loci that persisted IBD until the present day in both species, in which a mutation in the ancestral population could give rise to a TSP. We therefore sought to explore the implications of these recent findings for the interpretation of TSPs, including those that we report here. To identify regions of the genome that are uncoalesced >5 Mya in chimpanzees, we follow the approach of Cousins et al., who applied the pairwise sequentially Markovian coalescent (PSMC) to infer coalescent times along the genomes of seven humans of recent African ancestry. Specifically, for each individual and each 1 kb segment of the genome, they estimated the probability that the two haplotypes have not coalesced by 5 Mya, i.e., have not yet reached their common ancestor (see Methods). Taking the same approach for nine of our chimpanzee samples (see Methods), three from each of the subspecies, we estimate that on average the most recent common ancestor of a pair of haplotypes has not been reached by >5 Mya in 0.87%-2.57% of the genome, depending on the subspecies. As expected from their diversity levels (see Methods), the fraction of such uncoalesced segments is greatest in Central and Eastern chimpanzees and lowest in Western chimpanzees. Because the PSMC approach is applied to one individual at a time, it only identifies segments where the two haplotypes of a present-day individual have not coalesced (rather than two haplotypes currently carried by different individuals); therefore, we will miss some regions with old coalescent times. Nonetheless, by this approach, we can examine whether genomic regions in which pairs of haplotypes have not coalesced by >5 Mya tend to lie in regions of reduced background selection, using “‘BMAP”’ values estimated for the human genome [ 54 ]. Regions of the genome that evolved under weak background selection (i.e., are assigned a high BMAP) should be more likely to remain uncoalesced at deep timescales [ 53 ]. Conversely, if uncoalesced regions are the result of neutral processes rather than selection, we would expect to find a higher proportion in regions of high BMAP, where the coalescent rates are lower [ 55 ]. Consistent with this neutral expectation, the expected number of regions that are uncoalesced by 5 Mya is substantially higher for regions under less background selection ( Fig. 3A ). Download figure Open in new tab Figure 3: Expected number of 1 kb segments of the genome that remain uncoalesced by 5 Mya, stratified by background selection decile. A . 1 kb regions of the genome are binned into deciles of background selection (BMAP) on the x-axis, 10 being the regions that are inferred to experience the least background selection [ 54 ]. On the y-axis are the expected number of 1kb segments within that decile that are inferred to be uncoalesced 5 Mya in each individual (see Methods). Results are shown for the posterior decoding of PSMC run on seven humans (green triangles) (taken from [ 47 ]) and nine chimpanzees (circles), three from each of three subspecies: Eastern in blue, Western in purple, and Central in red. Mean population diversity levels, θ, used to run PSMC were 0.001 per bp for humans, 0.001674 in Eastern chimpanzees, 0.000818 in Western chimpanzees and 0.001953 in Central chimpanzees. B . Reported are the total number of our high confidence trans-species polymorphisms (defined as having statistical support above 80%; see Methods) within a given BMAP decile. Looking then at the trans-species polymorphisms that we identify through local genealogical reconstruction, we see that the majority lie in regions of low background selection (i.e., high BMAP; see Fig. 3B ). Together with the findings that many regions remain uncoalesced 5 Mya in humans as in chimpanzees, this distribution over BMAP values suggests that many of the TSP may simply reflect mutations that arose in the ancestral population in regions with low levels of purifying selection and by chance persisted until the present in both species. Indeed, this conclusion is consistent with the high FDR of ∼70% that we estimate even for very old SNPs ( Fig. 1 ). The implication is that even for TSPs between humans and chimpanzees, it remains difficult to know which subset are truly targets of selection. Discussion By taking an empirical approach, we find that variants shared between humans and chimpanzees are enriched for older alleles relative to what is expected from a putatively neutral genomic background, indicating that a substantial number are under long-term balancing selection. Specifically, our results suggest that, excluding the MHC region, an estimated 1700 regions harbor variants that are >1 My old because they were maintained by balancing selection. These variants include both trans-species polymorphisms as well as recurrent mutations under balancing selection in humans (and possibly also in chimpanzees). The approach we have taken does not require assumptions about the mode of balancing selection, relying only on the common feature that it maintains variation longer than expected; nor do we need to specify a demographic model. Therefore, our results provide robust evidence for balancing selection in the human genome. While we detect a clear signal of selection overall, the downside of our approach remains the difficulty of pinpointing the exact targets of selection. The FDR is high for all age cutoffs, above ∼70% even at the oldest ages, and therefore we cannot make strong claims about which of our shared SNPs are responsible for the signal of enrichment. Further work that uses independent lines of evidence is needed to accurately pinpoint the specific loci under long-lived balancing selection in humans. By relying on developments in ARG reconstruction methods, we are able to robustly identify novel TSPs, notably near MUC7 , and confirm those previously reported at MTRR and IGFBP7 . However, we also revisit the assumption that the existence of a trans-species polymorphism between humans and chimpanzees is in itself strong evidence of balancing selection. A recent report that the speciation event between the two species was more complex than previously believed, with ongoing gene flow until ∼5 Mya, increases the odds that some TSPs are neutral alleles that persisted by chance [ 47 ]. This possibility is supported by the high false discovery rate that we estimate even for very old ages, and the elevated proportion of TSPs in regions of low background selection, in which coalescent rates are lower. Almost all TSPs identified to date between humans and chimpanzees lay in non-coding regions of the genome. Previously, this observation was taken to be evidence that the ancient balancing selection primarily acts through selection on regulatory variation or that these sites must be linked to unsampled coding variants that are themselves the sites of selection [ 35 ]. Our results indicate that the high proportion of non-coding variants among TSPs may instead be because many of them are neutral. One implication is that finding particular classes of genes, such as immune-related genes, enriched among TSPs may be a consequence of their distribution along the genome rather than of selection on these genes themselves. For example, if immune-related genes tend to be in regions of the genome with less background selection, they would be more likely to harbor neutral TSPs. Nonetheless, the excess of old, shared variants relative to expectation indicates that there are many targets of balancing selection in humans that are yet to be fully characterized. Elucidating these loci and their functions will help to answer whether host-pathogen interactions are indeed the main balancing selection pressure in humans. The finding of neutral trans-species polymorphisms between human and chimpanzee has implications beyond these species. Human demographic history is, for obvious reasons, studied in more detail than almost any other species. The study of trans-species polymorphism between species that are similarly diverged to humans and chimpanzees may also suffer from similar pitfalls if relevant parameters are not well estimated or if, as is increasingly thought to be the case, speciation frequently does not result from a simple, clean split [ 69 ]. It should be noted, however, that trans-species polymorphisms between much more diverged species, such as variants at the ABO locus shared IBD between humans and gibbons [ 32 ] or MHC alleles shared among highly diverged mammalian species [ 70 ], remain extremely unlikely to have arisen under neutrality and thus remain compelling evidence of balancing selection. Understanding the extent to which balancing selection has played a role in human adaptation and in shaping our present day diversity remains an important but tricky task, which requires disentangling balancing selection from both demographic and other selective processes. We have shown that an empirical approach can offer a way forward in helping to characterize the extent of balancing selection in the genome. However, the accurate identification of the specific targets of selection will likely require independent lines of evidence from ever improving functional characterizations of human genetic variation. Methods & Materials Chimpanzee samples We used publicly available whole-genome sequences for chimpanzees, to which we added data obtained by the Przeworski lab in 2020 from blood samples for a pedigree of ten chimpanzees (including five founders), which were provided to the lab at Brown University by the Primate Foundation of Arizona in 2005. DNA libraries were prepared for sequencing by Genewiz using a PCR-free protocol and sequenced on Illumina HiSeq 2500 with 2 x 150 cycles to mean fold-coverage of between 16 and 44x in 2020 (Supplementary Table 1). Sequencing data for these samples are available on NCBI’s SRA under the project accession PRJNA1357362. Chimpanzee datasets and variant calling We downloaded SRA files of chimpanzee whole genome sequences from Venn et al. 2014 [ 49 ], Tatsumoto et. al. 2017 [ 50 ], Besenbacher et al. 2019 [ 51 ] and de Manuel et al. 2016 [ 48 ] and combined them with the 10 whole genome sequences that we had generated (see above). In total, there are 59 samples, of which 43 are unrelated (i.e., no first- and second-degree relatives) (Supplementary Table 2). For each sample, reads were aligned to the PanTro6 assembly with bwa-mem in paired-end mode. Duplicates were marked in the resulting BAM files with GATK’s MarkDuplicates. We then called variants with GATK (version 4.2.3) following GATK best practices. Specifically, for each chromosome, we generated individual gVCFs with HaplotypeCaller before consolidating all samples with GenomicsDBImport and jointly calling variants with GenotypeGVCFs. After removing indels and multi-allelic sites, we were left with 43,481,990 SNPs. Filtering Following GATK best practices, we retained sites with Phred Quality Score (QUAL) > 30 , QualByDepth (QD) > 2, RMSMappingQuality (MQ) > 40, MappingQualityRankSumTest (MQRankSum) < 6, FisherStrand bias (FS) < 60 and ReadPosRankSum < 6. We removed any sites where the total read depth across samples is 1.5-fold the median of the whole-genome average coverage per site. We excluded sites with excess heterozygosity by retaining only sites with ExcessHet < 28.69 (i.e. with p-values corresponding to more than 3 standard deviations away from Hardy-Weinberg equilibrium). Finally, we removed any sites where ≥10% of samples are missing genotypes. After this filtering, we were left with 38,169,920 SNPs. We used the pedigrees in our full dataset to identify sites with Mendelian errors using PLINK (--mendel). Any sites exhibiting one or more Mendelian violations in our total sample were removed (a further 862,705 SNPs). After this analysis, we retained only founders from the pedigrees and wild-born individuals from de Manuel et al. (2016) for downstream analyses. The reference assembly to which we mapped reads may contain regions where paralogs have been collapsed into a single copy. Differences between such copies could potentially generate false heterozygote calls, affecting our identification of shared variants. We therefore ran BISER [ 71 ] on both the PanTro6 and the telomere-to-telomere (T2T) assembly in order to detect segmental duplications that are collapsed in PanTro6 but are resolved in the T2T assembly. A total of 9707 regions, comprising 11.3% of the genome, were identified as duplicated by this approach. We removed the 4,878,666 SNPs (13.1%) that fall in these cryptic segmental duplications, thereby retaining 32,428,549 SNPs. Sub-species identification We retrieved the subspecies identities of all publicly available samples from their respective publications (Supplementary Table 1). In order to identify the subspecies of the founder individuals in our newly sequenced pedigree, we performed principal component analysis (PCA) on the total set of 59 chimpanzee samples. We used PLINK to LD-prune the variant set, i.e., to reduce linkage disequilibrium levels among SNPs (--indep-pairwise 50 10 0.1), and perform the PCA. We found that the first two principal components separate all unrelated samples into three clusters, corresponding to subspecies labels Pan troglodytes verus (Western), Pan troglodytes schweinfurthii (Eastern) and Pan troglodytes troglodytes (Central). Of the five founder individuals in our newly sequenced pedigree, four clustered with the labelled Western samples and the other one clustered with the Eastern samples (see Supplementary Figure 1). Our final set of 43 unrelated samples thus consists of 14 Eastern, 14 Central and 15 Western chimpanzees. Liftover of chimpanzee variants To map the positions onto the human genome, the filtered VCFs were lifted to hg19 using bcftools +liftover plugin [ 72 ]. We removed sites without a unique mapping in the new reference. Sites where neither the reference or alternative allele called in chimpanzee match the hg19 reference became multi-allelic and therefore were also removed. As a further check, we lifted the remaining variants back to the PanTro6 reference and removed any SNPs that did not map back to their original position. Lastly, we required that the region +/-100bp around each chimpanzee SNP maps also uniquely to hg19. Human (1000GP) data processing We downloaded the 1000 Genomes Project dataset of short variant calls in the Phase 3 dataset 2504 individuals. We removed indels and multiallelic sites to generate a set of biallelic SNPs. To remove SNPs in potential paralogs, as we had done for the chimpanzee data, we used two published datasets: the set of segmental duplications identified using the human T2T genome [ 73 ] and copy-number variants called in 1000GP Phase 3 [ 46 ]. Any SNPs falling within these regions were removed for the shared variant analysis (i.e., 16.0% of the initial set). Identification of shared variants between humans and chimpanzees We considered only human SNPs with a minor allele frequency ≥5% in 1000GP and chimpanzee SNPs with a minor allele count of more than three in our unrelated sample. To identify shared variants, we intersected the human and chimpanzee VCFs using bcftools isec . We then removed any sites that are multi-allelic in the resulting output; such sites represent positions where both species harbour a biallelic polymorphism but the segregating alleles differ (e.g. A/T in human vs. A/G in chimpanzee). The resulting final output is a set of 66,650 SNPs at which the same two alleles are segregating in human and chimpanzee samples. We used the Genome Reference Consortium’s definition of the MHC locus (chr6:28,477,797-33,448,354) to identify MHC variants and excluded them from subsequent analyses. Allele age estimates To characterize the distribution of ages of our shared variants, we used two sets of allele age estimates derived from different local genealogical reconstruction methods, Relate [ 44 ] and SINGER [ 45 ] . The authors of Relate used it to infer the topologies of local genealogies along the genome from the complete 1000GP Phase 3 dataset [ 44 ]. They then inferred population size histories and estimated branch lengths for each of the 26 subpopulations in 1000GP individually. This approach results in separate estimates of allele ages for each of these populations. We relied on the allele age estimates from two of these, one African subpopulation of Yoruba sampled in Ibadan, Nigeria (YRI, n=108), and one European subpopulation from Iberian populations in Spain (IBS, n=107). We choose these population samples because they are among the largest, and because Leffler et al. [ 35 ] had relied on a YRI sample, making the results more readily comparable. For a given 1000GP variant, we recorded age estimates for the start and end of the genealogical branch on which the derived allele was inferred to have arisen. In turn, the developers of SINGER applied it to a sample of 200 African haplotypes from 1000GP and generated 100 posterior samples of the ancestral recombination graph (ARG). We considered all 100 posterior estimates for the age estimates of the start and end of the branch on which the derived allele was inferred to have occurred. Selection of human-only matched controls Before matching, we restricted the shared and non-shared variant sets to those for which there are available age estimates, separately for Relate and SINGER . When using Relate age estimates, we restricted our attention to SNPs that have estimates in both the IBS and YRI populations, which left us with 54,415 shared SNPs (82% of the initial set). For the analysis with SINGER , we had age estimates for 19,009 SNPs (29% of the initial set). Before matching, we further removed any variants from the pool of potential controls if they do not map cleanly to the PanTro6 reference (following the same procedure as taken with the shared variants). To generate a null distribution of allele ages, we took a matched-control approach. Specifically, for each of our variants shared between humans and chimpanzees, we searched the 1000GP biallelic SNP set for variants that are not found in chimpanzees but are matched for other factors. We required a control variant to be of the same single nucleotide mutation type (e.g. T>G) and to be segregating at the same frequency. To match mutation types and their corresponding allele frequencies, we first polarized each variant into the ancestral and derived allele. To this end, we used the ancestral states as annotated in 1000GP, which are inferred using the 4-way EPO alignment of chimpanzee, orangutan, and rhesus macaque [ 74 ]. Our matching criteria are that a control site must have the same ancestral and derived alleles as the shared SNP and have a derived allele frequency in 1000GP that is within +/-1% of that of the shared SNP. A small minority of shared SNPs did not have an ancestral state reported for one of three reasons: because the site is not aligned, because there is a lineage specific insertion at the site, or because the ancestral state is uncertain (encoded as “.”, “-” and “N”, respectively). For these variants, comprising 0.63% of shared SNPs for which we have Relate estimates and 1.78% of SNPs for which we have SINGER estimates, we used the major allele frequency in the total 1000GP sample as the ancestral allele. Given that CpG sites are known to have elevated mutation rates, particularly transition rates, we also included whether a site is a CpG or not in our matching criteria. Specifically, we used the MutationalPatterns package to extract the mutation type and neighboring base pairs for all 1000GP variants. Mutations are coded as one of six single nucleotide mutation types based on the pyrimidine base that mutates, regardless of whether it is the reference or non-reference allele, meaning that C>T mutations and G>A mutations in the right context are both coded as C>T. We then recoded the alleles as ancestral or derived, as described above, and annotated all variants that are C>T transitions at CG dinucleotides as CpG mutations. To account for any remaining variation in the mutation rate within the types we have defined, we considered the mutation rate estimated by Roulette [ 57 ] . Roulette provides state-of-the-art estimates of the mutation rate of all possible single-nucleotide mutations in the human genome, determined using a model that incorporates information about the sequence context, methylation levels and gene expression levels. We require that matched controls have mutation rates estimated by Roulette within two-fold of that of the shared variant. Finally, we matched control and shared SNPs for the effects of linked selection by using estimates of the strength of background selection. Specifically, we assigned all SNPs in 1000GP to an individual B value based on the CADD best fit model from Murphy et al. (2022). We then calculated decile boundaries from the resulting distribution and selected matched controls that are in the same decile of B as the shared SNP. We removed any shared SNPs for which we could not find any appropriately-matched control SNPs: 0.14% of the SNPs with Relate estimates and 0.71% of the SNPs with SINGER estimates. Our aim was to generate 100 non-overlapping sets of matched control sets. The simplest way to proceed would be to identify 100 control SNPs for each shared SNP that were not used as controls for any other shared SNP. In practice, we could not always identify that many unique control SNPs that satisfied all our matching criteria: we identified a median of 22 control SNPs for CpGs and 134 for non-CpGs. We therefore generated the 100 matched control sets by bootstrap sampling control sets for each test variant. To check whether the control variants that were selected were neutrally-evolving, we examined their phyloP conservation scores. We downloaded phyloP scores for each site in the hg19 assembly generated from the 36-way eutherian mammal alignment excluding humans [ 59 ]. As our threshold for conservation [ 59 ], we calculated the 90th percentile value across all sites in the genome as 0.854. We then calculated the number of SNPs that exceed this threshold in our shared SNP set and each of our control sets as well as for the set of all 1000GP SNPs. Test for excess of old shared variation as compared to matched controls We used the matched control sets to test whether the ages of our shared variants are older than expected under neutrality. The age estimates from Relate and SINGER are for the beginning and ends of the branches on which our focal mutations are inferred to have arisen, rather than estimates for the mutations themselves. To assign an age to the mutation, we assumed that mutations are uniformly distributed along these branches, as would be the case if the mutations were neutral [ 62 ]. For a set of SNPs, we then calculated the expected number of SNPs older than a specific age, a , by summing the probability that each SNP is older than a. Specifically, for our shared SNP set and each of the 100 matched control sets, we calculated this expected number of SNPs older than a at intervals of 10 generations. We thus obtained a single estimate at each value of a for the shared SNP set and a distribution for the 100 matched control set, from which we report the median and the 90th-tile central interval in Figure 1 . Ages for which the expected number of SNPs older than a are significantly greater in the shared SNP set versus the 100 sets of matched controls are ages for which there is a significant excess of old alleles among shared SNPs. We can estimate the excess of old alleles at a given age from the expected number of SNPs older than age a among shared SNPs divided by the expected number among control SNPs. Local genealogical reconstruction around old shared variants Our shared SNP dataset includes cases that are old enough to plausibly be trans-species polymorphisms. To explore this possibility, we selected the 200 oldest of our shared variants by taking the mean rank of their midpoint age in YRI and IBS populations, as estimated by Relate. For these 200, we inferred local genealogical histories using a joint sample of 43 chimpanzees and 50 humans. To this end, we mapped the chimpanzee reads for each of our samples to the hg19 reference and then called variants using GATK for the regions containing old shared SNPs, extending 250 kb in either direction. Chimpanzee variants were filtered using the same criteria as when mapped to the PanTro6 (detailed above) and then phased using SHAPEIT2 . The chimpanzee VCF was merged with a 1000GP VCF of the same regions using bcftools merge before removing any sites that became multiallelic as result of being polymorphic for different alleles in the two species. We then downsampled the human sample set to 50 individuals, both because SINGER is limited to sample sizes on the order of hundreds of haplotypes and to generate a combined sample with roughly equal numbers of individuals from the two species. We performed this downsampling independently for each region while preserving the allele frequency of the oldest shared SNP in the region. Specifically, we sampled individuals from each of the diploid genotype states at this SNP in the same proportions as are found in the total 1000GP dataset. We then merged the resulting chimpanzee and human VCFs using bcftools merge. We ran SINGER on the merged VCFs for each region using parameters Ne=20,000 [ 75 ] and m=1.25x10 -8 [ 68 , 76 ] for a total of 10,000 iterations with 4,000 as burn-in before sampling every 60 iterations to get 100 samples of the ARG. We converted the resulting ARG sample outputs to the tskit tree sequence format using SINGER convert_to_tskit and performed all further operations and plotting using tskit [ 77 , 78 ]. Following the expectations for trans-species balanced polymorphisms laid out in Gao et al. (2015), we evaluated the support for a specific variant by inspecting the local topology at that site. Specifically, we required that the topology meet the following two requirements with respect to the mutation at site of interest (for sites with multiple mutations, we focused on the SNP with the oldest age estimate) (see Supplementary Figure 3). Firstly, when the mutation arose, the total number of lineages segregating in the tree must be more than one (i.e., the mutation cannot be on the root branch) and no more than three. Secondly, the first coalescent event after the mutation arose must be between one branch whose descendants are all chimpanzee haplotypes and another branch whose descendants are all human haplotypes. For each variant of interest, we counted how many of 100 posterior samples of the ARG met these requirements as a measure of support for the variants being a trans-species polymorphism. Functional annotation & gene assignment We annotated variants with their functional effects in the human genome using snpEff [ 79 ]. For intergenic and intronic variants, the gene(s) to which they should be assigned is often unknown. We attributed variants to the gene with the top score from the results of Open Targets Genetics’ V2G (variant to gene) pipeline. This approach uses a combination of information, including the distance to a gene, as well as results from quantitative trait loci mapping (eQTL and pQTL) and interaction data (PCHi-C), to examine the functional evidence linking a particular variant to genes nearby, outputting a single aggregate score for each variant-gene prediction [ 80 ]. To further assess whether any of our TSPs fall within known regulatory elements, we ran all TSP variants through ENCODE Project’s registry of candidate cis-regulatory elements, SCREEN [ 81 ]. Coalescent inference and decoding To infer regions of the genome that are uncoalesced at the time of the human and chimpanzee split, we applied the approach of Cousins et al. (2024) to the chimpanzee data. The authors used PSMC [ 82 ], as embedded in cobraa ( https://github.com/trevorcousins/cobraa ), to estimate the demographic histories of seven individuals of recent African ancestry, one from each of the African subpopulations in 1000GP. Using these estimates, they decode the underlying hidden Markov model to obtain the posterior probability that an individual’s two haplotypes coalesced during each time interval at every position in the genome. Specifically, for each individual, the authors ran cobraa in decode mode with a step size of 1 kb to obtain the posterior probabilities of coalescence within each of the time windows. The posterior probability that a 1kb segment has not coalesced by a given time boundary was then calculated as 1 minus the sum of the probabilities of coalescence in every more recent time window. Files containing these posterior probabilities of non-coalescence for each 1 kb segment at every time boundary older than 2 Mya were kindly supplied to us by the authors. To replicate this analysis in chimpanzees, we selected nine samples, three from each of the chimpanzee subspecies. Per-sample VCF files were generated from our PanTro6 aligned BAM files using the command bcftools mpileup -q 20 -Q 20 -C 50 and Stephan Schiffels’s bamCaller.py . We then used these VCF files to prepare the MHS input files for PSMC using Aylwyn Scally’s generate_multihetsep.py . Sites with low mappability were excluded from these input files using a mappability mask that we generated for the PanTro6 assembly with SNPable ( http://lh3lh3.users.sourceforge.net/snpable.shtml ) using a k-mer length of 150 bp and stringency of 75%. We then ran PSMC, as embedded in cobraa , using the same parameters as Cousins et al. used for humans (D=64, b=100, spread1=0.005, spread2=50, mu_over_r=1.5), except that we used different values for the scaled mutation rate, θ , to account for the differences in the present day diversity levels of the different species and subpopulations. Specifically, we estimated the population mutation rate, θ , for each chimpanzee sample from the MHS input files by dividing the number of heterozygous sites by the total unmasked sequence length. We then relied on the mean θ across the three samples for each chimpanzee subspecies when running cobraa : 0.001674 per bp in Eastern chimpanzees, 0.000818 in Western and 0.001953 in Central (while Cousins et al. used a θ value of 0.001 for humans). As in Cousins et al. (2024), for each individual, we use the inferred PSMC parameters and the MHS input files to run cobraa’s decode and obtain the posterior probability that the haplotypes coalesced in each time window for each 1 kb window. As described above, the decoding output can be used to calculate the posterior probability that each 1kb segment of the genome has not coalesced by each time interval boundary in the PSMC time discretization. We looked specifically at the time point of 5 Mya, roughly corresponding to the end of the inferred period of gene flow between the two species following their initial split [ 47 ]. We note that the exact time points are determined by the parameters used in PSMC and vary somewhat between the species and subspecies because of the different θ values, generation times, and mutation rates. In practice, the first time points above 5 Mya are at 5,049,745 years in humans, 5,389,799 years in Eastern chimpanzees, 5,408,462 in Western chimpanzees and 5,445,261 years in Central chimpanzees, using a generation time of 29 years for humans and 25 years for chimpanzees [ 83 ] and a mutation rate per bp per generation of 1.25x10 -8 in humans and 1.27x10 -8 in chimpanzees [ 51 ]. At each 1 kb segment of the genome in each individual, we then calculated the posterior probability that the two haplotypes have not coalesced by 5 Mya as 1 minus the sum of the probabilities of coalescence in all more recent time intervals. For each individual, we summed these posterior probabilities to obtain an estimate of the expected number of 1 kb segments in the genome that remain uncoalesced by 5 Mya. We converted this expected number into an expected fraction of the genome that remains uncoalesced by dividing by the total number of 1 kb segments in the genome. To assess how the expected number of uncoalesced segments varies with background selection, we annotated 1 kb segments with their mean BMAP value in humans (see above) and assigned them to deciles of BMAP. We then repeated the exercise of calculating the expected number of 1 kb segments that remain uncoalesced by 5 Mya for each BMAP decile separately. Download figure Open in new tab Supplementary Figure 1: PC1 vs. PC2 from the Principal Component Analysis (PCA) of the chimpanzee samples. The PCA is based on filtered and LD-pruned, whole-genome SNP data for 43 unrelated chimpanzees (see Methods). Shown are the primary axis of variation (PC1) and secondary axis of variation (PC2), with the proportion of variance explained in parenthesis. Samples are colored according to their subspecies labels (Central = red, Eastern = blue, Western = blue). Shapes indicate whether the subspecies label was previously assigned (circles) or generated in this study (triangles). As can be seen, four of the five samples for which we collected genome sequence data cluster with samples previously labelled as Western and one as Eastern. Download figure Open in new tab Supplementary Figure 2: Proportion of shared variants that are at CpG transitions by age, as estimated by Relate. Shown for every million years is the proportion of shared variants that are CpG transitions. Proportions calculated using age estimates from the YRI population are shown in purple and those calculated using the IBS population are shown in green; all age estimates are from Relate [ 44 ]. The bars denote one standard error. The total proportion of all 1000GP SNPs (MAF≥5%) that are CpG transitions is shown with the dashed red line. Download figure Open in new tab Supplementary Figure 3: Criteria applied to local genealogy for trans-species polymorphisms. The example of a reconstructed genealogy of eight haplotypes illustrates the criteria that we impose to consider a variant a trans-species polymorphism, in particular that lineages that are from different species (blue vs. red tips) but both carry the derived (or ancestral) allele (pink branches) coalesce with one another before they coalesce with lineages of their own species carrying the ancestral (derived) allele. Specifically, we require that, at the time of the oldest mutation at the focal site, there are either two or three lineages segregating in the genealogy (grey dashed line). We further consider whether the first coalescent below this mutation is between a lineage whose descendants are all humans and a lineage whose descendants are all chimpanzees (purple circle). On the left is an example of a polymorphism that meets both of these criteria. On the right, we show a possible reconstruction for a mutation that arose independently in the two species; the mutation in humans is very old and meets our first criterion but, because all of its descendants are human haplotypes, it does not meet the second requirement. View this table: View inline View popup Download powerpoint Supplementary Table 1 Newly sequenced chimpanzee pedigree samples. View this table: View inline View popup Download powerpoint Supplementary Table 2 Summary of all chimpanzee samples. Acknowledgments We thank Felix Wu for organizing the sequencing of the chimpanzee samples; we are also grateful to Trevor Cousins, William Milligan as well as Anastasia Stolyarova and other members of the Przeworski lab for helpful discussions. Funder Information Declared NIH , GM153355 References 1. ↵ Nielsen R . Molecular signatures of natural selection . Annu Rev Genet . 2005 ; 39 : 197 – 218 . OpenUrl CrossRef PubMed Web of Science 2. ↵ Sabeti PC , Schaffner SF , Fry B , Lohmueller J , Varilly P , Shamovsky O , et al. Positive natural selection in the human lineage . Science . 2006 ; 312 : 1614 – 1620 . OpenUrl Abstract / FREE Full Text 3. ↵ Lewontin RC . Senior lecture, annual short course on medical and experimental mammalian genetics, Bar Harbor, Maine, 1968 . Int J Epidemiol. 2016 ; 45 : 654 – 664 . OpenUrl CrossRef PubMed 4. ↵ Charlesworth D . Balancing selection and its effects on sequences in nearby genome regions . PLoS Genet . 2006 ; 2 : e64 . OpenUrl CrossRef PubMed 5. ↵ Roux C , Pauwels M , Ruggiero M-V , Charlesworth D , Castric V , Vekemans X . Recent and ancient signature of balancing selection around the S-locus in Arabidopsis halleri and A. lyrata . Mol Biol Evol . 2013 ; 30 : 435 – 447 . OpenUrl CrossRef PubMed Web of Science 6. ↵ Groh JS , Vik DC , Davis M , Monroe JG , Stevens KA , Brown PJ , et al. Ancient structural variants control sex-specific flowering time morphs in walnuts and hickories . Science . 2025 ; 387 : eado5578 . OpenUrl CrossRef PubMed 7. ↵ Lamichhaney S , Fan G , Widemo F , Gunnarsson U , Thalmann DS , Hoeppner MP , et al. Structural genomic changes underlie alternative reproductive strategies in the ruff (Philomachus pugnax) . Nat Genet . 2016 ; 48 : 84 – 88 . OpenUrl CrossRef PubMed 8. ↵ Bitarello BD , Brandt DYC , Meyer D , Andrés AM . Inferring balancing selection from genome-scale data . Genome Biol Evol . 2023 ; 15 : evad032 . OpenUrl CrossRef PubMed 9. ↵ Allison AC . The sickle-cell and haemoglobin C genes in some African populations . Ann Hum Genet . 1956 ; 21 : 67 – 89 . OpenUrl CrossRef PubMed Web of Science 10. ↵ Shriner D , Rotimi CN . Whole-genome-sequence-based haplotypes reveal single origin of the sickle allele during the Holocene Wet Phase . Am J Hum Genet . 2018 ; 102 : 547 – 556 . OpenUrl CrossRef PubMed 11. ↵ Qiu Q-W , Wu D-D , Yu L-H , Yan T-Z , Zhang W , Li Z-T , et al. Evidence of recent natural selection on the Southeast Asian deletion (--(SEA)) causing α-thalassemia in South China . BMC Evol Biol . 2013 ; 13 : 63 . OpenUrl CrossRef PubMed 12. ↵ Verrelli BC , McDonald JH , Argyropoulos G , Destro-Bisol G , Froment A , Drousiotou A , et al. Evidence for balancing selection from nucleotide sequence analyses of human G6PD . Am J Hum Genet . 2002 ; 71 : 1112 – 1128 . OpenUrl CrossRef PubMed Web of Science 13. ↵ Hedrick PW . Heterozygote advantage: the effect of artificial selection in livestock and pets . J Hered . 2015 ; 106 : 141 – 154 . OpenUrl CrossRef PubMed 14. ↵ Milesi P , Weill M , Lenormand T , Labbé P . Heterogeneous gene duplications can be adaptive because they permanently associate overdominant alleles: OVERDOMINANCE AND HETEROGENEOUS GENE DUPLICATIONS . Evol Lett . 2017 ; 1 : 169 – 180 . OpenUrl PubMed 15. ↵ Levene H . Genetic equilibrium when more than one ecological niche is available . Am Nat . 1953 ; 87 : 331 – 333 . OpenUrl CrossRef Web of Science 16. Stahl EA , Dwyer G , Mauricio R , Kreitman M , Bergelson J . Dynamics of disease resistance polymorphism at the Rpm1 locus of Arabidopsis . Nature . 1999 ; 400 : 667 – 671 . OpenUrl CrossRef PubMed Web of Science 17. ↵ Tellier A , Brown JKM . Spatial heterogeneity, frequency-dependent selection and polymorphism in host-parasite interactions . BMC Evol Biol . 2011 ; 11 : 319 . OpenUrl CrossRef PubMed 18. ↵ Hudson RR , Kaplan NL . The coalescent process in models with selection and recombination . Genetics . 1988 ; 120 : 831 – 840 . OpenUrl Abstract / FREE Full Text 19. ↵ Takahata N . A simple genealogical structure of strongly balanced allelic lines and trans-species evolution of polymorphism . Proc Natl Acad Sci U S A . 1990 ; 87 : 2419 – 2423 . OpenUrl Abstract / FREE Full Text 20. Takahata N , Nei M . Allelic genealogy under overdominant and frequency-dependent selection and polymorphism of major histocompatibility complex loci . Genetics . 1990 ; 124 : 967 – 978 . OpenUrl Abstract / FREE Full Text 21. ↵ Taylor JE . The effect of fluctuating selection on the genealogy at a linked site . Theor Popul Biol . 2013 ; 87 : 34 – 50 . OpenUrl CrossRef PubMed 22. ↵ Charlesworth B , Charlesworth D , Barton NH . The effects of genetic and geographic structure on neutral variation . Annu Rev Ecol Evol Syst . 2003 ; 34 : 99 – 125 . OpenUrl CrossRef 23. Lewontin RC , Krakauer J . Distribution of gene frequency as a test of the theory of the selective neutrality of polymorphisms . Genetics . 1973 ; 74 : 175 – 195 . OpenUrl Abstract / FREE Full Text 24. ↵ Soni V , Jensen JD . Temporal challenges in detecting balancing selection from population genomic data . G3 (Bethesda) . 2024 ; 14 : jkae069 . OpenUrl CrossRef 25. ↵ DeGiorgio M , Lohmueller KE , Nielsen R . A model-based approach for identifying signatures of ancient balancing selection in genetic data . PLoS Genet . 2014 ; 10 : e1004561 . OpenUrl CrossRef PubMed 26. Cheng X , DeGiorgio M . Flexible mixture model approaches that accommodate footprint size variability for robust detection of balancing selection . Mol Biol Evol . 2020 ; 37 : 3267 – 3291 . OpenUrl CrossRef PubMed 27. ↵ Cheng X , DeGiorgio M . BalLeRMix +: mixture model approaches for robust joint identification of both positive selection and long-term balancing selection . Bioinformatics . 2022 ; 38 : 861 – 863 . OpenUrl CrossRef PubMed 28. ↵ Felsenstein J . The theoretical population genetics of variable selection and migration . Annu Rev Genet . 1976 ; 10 : 253 – 280 . OpenUrl CrossRef PubMed Web of Science 29. ↵ Bitarello BD , de Filippo C , Teixeira JC , Schmidt JM , Kleinert P , Meyer D , et al. Signatures of long-term balancing selection in human genomes . Genome Biol Evol . 2018 ; 10 : 939 – 955 . OpenUrl CrossRef PubMed 30. ↵ Siewert KM , Voight BF . BetaScan2: Standardized statistics to detect balancing selection utilizing substitution data . Genome Biol Evol . 2020 ; 12 : 3873 – 3877 . OpenUrl CrossRef PubMed 31. ↵ Rasmussen MD , Hubisz MJ , Gronau I , Siepel A . Genome-wide inference of ancestral recombination graphs . PLoS Genet . 2014 ; 10 : e1004342 . OpenUrl CrossRef PubMed 32. ↵ Ségurel L , Thompson EE , Flutre T , Lovstad J , Venkat A , Margulis SW , et al. The ABO blood group is a trans-species polymorphism in primates . Proc Natl Acad Sci U S A . 2012 ; 109 : 18493 – 18498 . OpenUrl Abstract / FREE Full Text 33. ↵ Gao Z , Przeworski M , Sella G . Footprints of ancient-balanced polymorphisms in genetic variation data from closely related species: FOOTPRINTS OF ANCIENT BALANCING SELECTION . Evolution . 2015 ; 69 : 431 – 446 . OpenUrl CrossRef PubMed 34. ↵ Takahata N , Satta Y , Klein J . Polymorphism and balancing selection at major histocompatibility complex loci . Genetics . 1992 ; 130 : 925 – 938 . OpenUrl Abstract / FREE Full Text 35. ↵ Leffler EM , Gao Z , Pfeifer S , Ségurel L , Auton A , Venn O , et al. Multiple instances of ancient balancing selection shared between humans and chimpanzees . Science . 2013 ; 339 : 1578 – 1582 . OpenUrl Abstract / FREE Full Text 36. Teixeira JC , de Filippo C , Weihmann A , Meneu JR , Racimo F , Dannemann M , et al. Long-term balancing selection in LAD1 maintains a missense trans-species polymorphism in humans, chimpanzees, and bonobos . Mol Biol Evol . 2015 ; 32 : 1186 – 1196 . OpenUrl CrossRef PubMed 37. ↵ Cagliani R , Guerini FR , Fumagalli M , Riva S , Agliardi C , Galimberti D , et al. A trans-specific polymorphism in ZC3HAV1 is maintained by long-standing balancing selection and may confer susceptibility to multiple sclerosis . Mol Biol Evol . 2012 ; 29 : 1599 – 1613 . OpenUrl CrossRef PubMed Web of Science 38. Ségurel L , Gao Z , Przeworski M . Ancestry runs deeper than blood: the evolutionary history of ABO points to cryptic variation of functional importance: Insights & Perspective . Bioessays . 2013 ; 35 : 862 – 867 . OpenUrl CrossRef PubMed 39. ↵ Spurgin LG , Richardson DS . How pathogens drive genetic diversity: MHC, mechanisms and misunderstandings . Proc Biol Sci . 2010 ; 277 : 979 – 988 . OpenUrl CrossRef PubMed Web of Science 40. ↵ Mode CJ . A mathematical model for the co-evolution of obligate parasites and their hosts . Evolution . 1958 ; 12 : 158 . OpenUrl CrossRef Web of Science 41. ↵ Azevedo L , Serrano C , Amorim A , Cooper DN . Trans-species polymorphism in humans and the great apes is generally maintained by balancing selection that modulates the host immune response . Hum Genomics . 2015 ; 9 : 21 . OpenUrl CrossRef PubMed 42. ↵ Nurk S , Koren S , Rhie A , Rautiainen M , Bzikadze AV , Mikheenko A , et al. The complete sequence of a human genome . Science . 2022 ; 376 : 44 – 53 . OpenUrl CrossRef PubMed 43. ↵ Yoo D , Rhie A , Hebbar P , Antonacci F , Logsdon GA , Solar SJ , et al. Complete sequencing of ape genomes . Nature . 2025 ; 641 : 401 – 418 . OpenUrl CrossRef PubMed 44. ↵ Speidel L , Forest M , Shi S , Myers SR . A method for genome-wide genealogy estimation for thousands of samples . Nat Genet . 2019 ; 51 : 1321 – 1329 . OpenUrl CrossRef PubMed 45. ↵ Deng Y , Nielsen R , Song YS . Robust and accurate Bayesian inference of genome-wide genealogies for large samples . bioRxiv . 2024 . p. 2024.03.16.585351. doi: 10.1101/2024.03.16.585351 OpenUrl Abstract / FREE Full Text 46. ↵ 1000 Genomes Project Consortium , Auton A , Brooks LD , Durbin RM , Garrison EP , Kang HM , et al. A global reference for human genetic variation . Nature . 2015 ; 526 : 68 – 74 . OpenUrl CrossRef PubMed 47. ↵ Cousins T , Schweiger R , Durbin R . Deep coalescent history of the hominin lineage . bioRxiv . 2024 . Available: https://www.biorxiv.org/content/10.1101/2024.10.17.618932v2.full.pdf 48. ↵ de Manuel M , Kuhlwilm M , Frandsen P , Sousa VC , Desai T , Prado-Martinez J , et al. Chimpanzee genomic diversity reveals ancient admixture with bonobos . Science . 2016 ; 354 : 477 – 481 . OpenUrl Abstract / FREE Full Text 49. ↵ Venn O , Turner I , Mathieson I , de Groot N , Bontrop R , McVean G . Nonhuman genetics. Strong male bias drives germline mutation in chimpanzees . Science . 2014 ; 344 : 1272 – 1275 . OpenUrl Abstract / FREE Full Text 50. ↵ Tatsumoto S , Go Y , Fukuta K , Noguchi H , Hayakawa T , Tomonaga M , et al. Direct estimation of de novo mutation rates in a chimpanzee parent-offspring trio by ultra-deep whole genome sequencing . Sci Rep . 2017 ; 7 : 13561 . OpenUrl CrossRef PubMed 51. ↵ Besenbacher S , Hvilsom C , Marques-Bonet T , Mailund T , Schierup MH . Direct estimation of mutations in great apes reconciles phylogenetic dating . Nat Ecol Evol . 2019 ; 3 : 286 – 292 . OpenUrl PubMed 52. ↵ Black FL , Hedrick PW . Strong balancing selection at HLA loci: evidence from segregation in South Amerindian families . Proc Natl Acad Sci U S A . 1997 ; 94 : 12452 – 12456 . OpenUrl Abstract / FREE Full Text 53. ↵ Hudson RR , Kaplan NL . Deleterious background selection with recombination . Genetics . 1995 ; 141 : 1605 – 1617 . OpenUrl Abstract / FREE Full Text 54. ↵ Murphy DA , Elyashiv E , Amster G , Sella G . Broad-scale variation in human genetic diversity levels is predicted by purifying selection on coding and non-coding elements . Elife . 2023 ; 12 . doi: 10.7554/eLife.76065 OpenUrl CrossRef 55. ↵ Nordborg M . Structured coalescent processes on different time scales . Genetics . 1997 ; 146 : 1501 – 1514 . OpenUrl Abstract / FREE Full Text 56. ↵ Sved J , Bird A . The expected equilibrium of the CpG dinucleotide in vertebrate genomes under a mutation model . Proc Natl Acad Sci U S A . 1990 ; 87 : 4692 – 4696 . OpenUrl Abstract / FREE Full Text 57. ↵ Seplyarskiy V , Koch EM , Lee DJ , Lichtman JS , Luan HH , Sunyaev SR . A mutation rate model at the basepair resolution identifies the mutagenic effect of polymerase III transcription . Nat Genet . 2023 ; 55 : 2235 – 2242 . OpenUrl CrossRef PubMed 58. ↵ Sullivan PF , Meadows JRS , Gazal S , Phan BN , Li X , Genereux DP , et al. Leveraging base-pair mammalian constraint to understand genetic variation and human disease . Science . 2023 ; 380 : eabn2937 . OpenUrl CrossRef PubMed 59. ↵ Fu W , Gittelman RM , Bamshad MJ , Akey JM . Characteristics of neutral and deleterious protein-coding variation among individuals and populations . Am J Hum Genet . 2014 ; 95 : 421 – 436 . OpenUrl CrossRef PubMed 60. ↵ Nielsen R , Vaughn AH , Deng Y . Inference and applications of ancestral recombination graphs . Nat Rev Genet . 2025 ; 26 : 47 – 58 . OpenUrl CrossRef PubMed 61. ↵ Y C Brandt D , Wei X , Deng Y , Vaughn AH , Nielsen R . Evaluation of methods for estimating coalescence times using ancestral recombination graphs . Genetics . 2022 ; 221 . doi: 10.1093/genetics/iyac044 OpenUrl CrossRef PubMed 62. ↵ Hudson RR . Gene genealogies and the coalescent process . Oxford Surveys in Evolutionary Biology . 1990 ; 7 : 1 – 44 . OpenUrl 63. ↵ Schierup MH , Hein J . Consequences of recombination on traditional phylogenetic analysis . Genetics . 2000 ; 156 : 879 – 891 . OpenUrl Abstract / FREE Full Text 64. ↵ GTEx Consortium . The GTEx Consortium atlas of genetic regulatory effects across human tissues . Science . 2020 ; 369 : 1318 – 1330 . OpenUrl Abstract / FREE Full Text 65. ↵ Liu B , Rayment SA , Gyurko C , Oppenheim FG , Offner GD , Troxler RF . The recombinant N-terminal region of human salivary mucin MG2 (MUC7) contains a binding domain for oral Streptococci and exhibits candidacidal activity . Biochem J . 2000 ; 345 Pt 3 : 557 – 564 . OpenUrl Abstract / FREE Full Text 66. ↵ Liu B , Rayment S , Oppenheim FG , Troxler RF . Isolation of human salivary mucin MG2 by a novel method and characterization of its interactions with oral bacteria . Arch Biochem Biophys . 1999 ; 364 : 286 – 293 . OpenUrl CrossRef PubMed Web of Science 67. ↵ Wiuf C , Zhao K , Innan H , Nordborg M . The probability and chromosomal extent of trans-specific polymorphism . Genetics . 2004 ; 168 : 2363 – 2372 . OpenUrl Abstract / FREE Full Text 68. ↵ Jónsson H , Sulem P , Kehr B , Kristmundsdottir S , Zink F , Hjartarson E , et al. Parental influence on human germline de novo mutations in 1,548 trios from Iceland . Nature . 2017 ; 549 : 519 – 522 . OpenUrl CrossRef PubMed 69. ↵ Taylor SA , Larson EL . Insights from genomes into the evolutionary importance and prevalence of hybridization in nature . Nat Ecol Evol . 2019 ; 3 : 170 – 177 . OpenUrl PubMed 70. ↵ Fortier AL , Pritchard JK . Ancient trans-species polymorphism at the Major Histocompatibility Complex in primates . eLife . 2025 . doi: 10.7554/elife.103547.2 OpenUrl CrossRef 71. ↵ Išerić H , Alkan C , Hach F , Numanagić I . Fast characterization of segmental duplication structure in multiple genome assemblies . Algorithms Mol Biol . 2022 ; 17 : 4 . OpenUrl CrossRef PubMed 72. ↵ Genovese G , Rockweiler NB , Gorman BR , Bigdeli TB , Pato MT , Pato CN , et al. BCFtools/liftover: an accurate and comprehensive tool to convert genetic variants across genome assemblies . Bioinformatics . 2024 ; 40 . doi: 10.1093/bioinformatics/btae038 OpenUrl CrossRef PubMed 73. ↵ Vollger MR , Guitart X , Dishuck PC , Mercuri L , Harvey WT , Gershman A , et al. Segmental duplications and their variation in a complete human genome . Science . 2022 ; 376 : eabj6965 . OpenUrl CrossRef PubMed 74. ↵ Herrero J , Muffato M , Beal K , Fitzgerald S , Gordon L , Pignatelli M , et al. Ensembl comparative genomics resources . Database (Oxford ). 2016 ; 2016 : bav096 . OpenUrl CrossRef PubMed 75. ↵ Schiffels S , Durbin R . Inferring human population size and separation history from multiple genome sequences . Nat Genet . 2014 ; 46 : 919 – 925 . OpenUrl CrossRef PubMed 76. ↵ Kong A , Frigge ML , Masson G , Besenbacher S , Sulem P , Magnusson G , et al. Rate of de novo mutations and the importance of father’s age to disease risk . Nature . 2012 ; 488 : 471 – 475 . OpenUrl CrossRef PubMed Web of Science 77. ↵ Kelleher J , Etheridge AM , McVean G . Efficient coalescent simulation and genealogical analysis for large sample sizes . PLoS Comput Biol . 2016 ; 12 : e1004842 . OpenUrl CrossRef PubMed 78. ↵ Wong Y , Ignatieva A , Koskela J , Gorjanc G , Wohns AW , Kelleher J . A general and efficient representation of ancestral recombination graphs . Genetics . 2024 ; 228 : iyae100 . OpenUrl CrossRef PubMed 79. ↵ Cingolani P , Platts A , Wang LL , Coon M , Nguyen T , Wang L , et al. A program for annotating and predicting the effects of single nucleotide polymorphisms, SnpEff: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3: SNPs in the genome of Drosophila melanogaster strain w1118; iso-2; iso-3 . Fly (Austin) . 2012 ; 6 : 80 – 92 . OpenUrl PubMed 80. ↵ Buniello A , Suveges D , Cruz-Castillo C , Llinares MB , Cornu H , Lopez I , et al. Open Targets Platform: facilitating therapeutic hypotheses building in drug discovery . Nucleic Acids Res . 2025 ; 53 : D1467 – D1475 . OpenUrl CrossRef PubMed 81. ↵ ENCODE Project Consortium , Moore JE , Purcaro MJ , Pratt HE , Epstein CB , Shoresh N , et al. Expanded encyclopaedias of DNA elements in the human and mouse genomes . Nature . 2020 ; 583 : 699 – 710 . OpenUrl CrossRef PubMed 82. ↵ Li H , Durbin R . Inference of human population history from individual whole-genome sequences . Nature . 2011 ; 475 : 493 – 496 . OpenUrl CrossRef PubMed Web of Science 83. ↵ Langergraber KE , Prüfer K , Rowney C , Boesch C , Crockford C , Fawcett K , et al. Generation times in wild chimpanzees and gorillas suggest earlier divergence times in great ape and human evolution . Proc Natl Acad Sci U S A . 2012 ; 109 : 15716 – 15721 . OpenUrl Abstract / FREE Full Text View the discussion thread. Back to top Previous Next Posted November 11, 2025. Download PDF Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Revisiting the evidence for long-lived balancing selection in humans Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Revisiting the evidence for long-lived balancing selection in humans Hannah Munby , Molly Przeworski bioRxiv 2025.11.10.687682; doi: https://doi.org/10.1101/2025.11.10.687682 Share This Article: Copy Citation Tools Revisiting the evidence for long-lived balancing selection in humans Hannah Munby , Molly Przeworski bioRxiv 2025.11.10.687682; doi: https://doi.org/10.1101/2025.11.10.687682 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Evolutionary Biology Subject Areas All Articles Animal Behavior and Cognition (7643) Biochemistry (17717) Bioengineering (13910) Bioinformatics (42015) Biophysics (21477) Cancer Biology (18626) Cell Biology (25536) Clinical Trials (138) Developmental Biology (13392) Ecology (19935) Epidemiology (2067) Evolutionary Biology (24356) Genetics (15617) Genomics (22530) Immunology (17755) Microbiology (40437) Molecular Biology (17200) Neuroscience (88703) Paleontology (667) Pathology (2840) Pharmacology and Toxicology (4832) Physiology (7657) Plant Biology (15171) Scientific Communication and Education (2046) Synthetic Biology (4304) Systems Biology (9828) Zoology (2272)
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.