{"paper_id":"4b4f4694-0ed7-4d53-b4a2-a1fb2426cc8e","body_text":"The performance of genetic-constraint metrics varies \nsignificantly across the human noncoding genome \nPeter McHale, Michael E. Goldberg, and Aaron R. Quinlan* \nDepartment of Human Genetics and Utah Center for Genetic Discovery, University of Utah, Salt \nLake City, UT 84112, USA \n*To whom correspondence should be addressed (aquinlan@genetics.utah.edu). \nAbstract \nA longstanding goal in human genetics is to prioritize noncoding loci that, when disrupted, lead \nto developmental disorders and other Mendelian traits. In pursuit of this goal, multiple metrics \nhave been developed to distinguish neutrally evolving sequences from those subjected to \npurifying selection. These metrics are commonly evaluated genome-wide, e.g., by computing a \nprecision-recall curve on windows tiling the entire noncoding genome. Here, we identify parts of \nthe noncoding genome where these metrics significantly underperform relative to their \ngenome-wide performance due to “bias” in the underlying models of neutral genetic variation \nand/or a low “signal-to-noise ratio” in the genetic data. The most extreme effects are found for \nGnocchi (Chen et al. 2024), the performance of which declines as GC content increases. We \nsuggest annotating constraint scores of noncoding genomic intervals with robust measures of \nthe bias of the corresponding model, allowing users to gauge confidence in those scores.  \nIntroduction  \nComputational approaches designed to predict noncoding variant pathogenicity must be \nimproved if they are to have a meaningful impact upon clinical research (Liu et al. 2019). Many \nof these approaches are based on a simple, but powerful, idea: over the hundreds of millions of \nyears separating species, a single base is likely to have mutated, so the only plausible way to \nexplain why a genomic site is identical between a human and a fish, say, is that whatever \nmutations occurred there were harmful and eliminated by selection. This deep-time filter is so \nrigorous that it can detect even weakly constrained elements. Methods like phyloP (Pollard et al. \n2010) and GERP (Davydov et al. 2010; Cooper et al. 2005) formalize this idea in a number of \nways. (1) They go beyond the vague assertion that a base is \"likely to have mutated\" by \nquantifying the expected (neutral) rate of nucleotide substitution between species. (2) They \nextend the analysis from the comparison of two species to the entire tree of life by accounting \nfor the fact that conservation between, e.g., human and chimp (closely related) is weaker \nevidence of selection than conservation between human and fish (distantly related). (3) They go \nbeyond a simple binary classification of conservation to answer the question \"Is the substitution \npattern we see at this site extremely unlikely to have occurred under the neutral model?”  \nSelection signals derived from sequence conservation are extraordinarily strong and \nunambiguous but have an Achilles heel: they cannot find genomic elements that are newly \n1 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nconstrained in humans, limiting their ability to correctly prioritize noncoding variants. If an \nelement (like an enhancer) becomes functional in the human lineage after the split from \nchimpanzees, it will not be conserved in chimps, mice, or fish. In scenarios such as these, \nconservation methods are underpowered (though they can capture elements that have gained \nfunction in closely related species, e.g., primates (Boffelli et al. 2003)). The blind spot of \nconservation-based methods is not small. Rands et al. estimated that only about 25% of the \nelements under purifying selection in humans are also constrained in mouse (Rands et al. \n2014), while, more recently, Huber et al. estimated that most of the noncoding sequence that is \nfunctional in human has experienced changes in selection coefficients throughout mammalian \nevolution (Huber et al. 2020). This process, called “turnover”, had long been suspected from the \nobservation that, as the divergence between mammalian species increases, the predicted \namount of pairwise shared functional sequence decreases dramatically (Smith et al. 2004; \nMeader et al. 2010). Furthermore, most transcription-factor binding events are species-specific, \nindicative of gain and loss of binding along each lineage (reviewed in (Ponting and Hardison \n2011)).  \nThe only way to discover human-specific genetic constraints in the noncoding genome are \npopulation-based methods. Pleasingly, early deployment of these methods discovered that \nregulatory elements located near neural genes gained function in the human lineage (Ward and \nKellis 2012; Schrider and Kern 2015), and we expect these methods to become increasingly \nsensitive for the discovery of gain- and loss-of-function discovery as more human genomes are \nsequenced. These approaches start with the construction of a model of neutral genetic diversity \n(e.g., number of SNVs in each of a set of 1kb genomic windows) observed in a healthy human \ncohort, using features that are known to co-vary with SNV count. Of these, the most well-known \nfeature is the local sequence composition of a given site, whose probability of being \npolymorphic can increase >10-fold if the sequence is a CpG dinucleotide (Sved and Bird 1990; \nBulmer 1986; Schraiber et al. 2024). With a model of this type in hand, one may identify putative \nsignals of negative selection by searching for anomalous regions in the genome where the \nobserved diversity is significantly lower than the neutral model predicts (Chen et al. 2024; \nHalldorsson et al. 2022; di Iulio et al. 2018). Alternatively, one may construct a model of all \ngenetic diversity, including observed diversity impacted by negative selection. In this approach, \nthe model includes a selection parameter, the value of which is estimated for each genomic \nwindow (Dukler et al. 2022). The two approaches are intimately related: as Dukler et al. show, \nthe maximum-likelihood estimate of the selection coefficient in the second approach is the ratio \nof the observed diversity and that expected under neutrality, the comparison of which is the \nbasis of the first approach.  \nThe success of either approach depends upon how well the underlying model of neutrality \ntracks the observed SNV density across the noncoding genome: if one fails to or inadequately \nmodels the neutral mechanisms that, like purifying selection, deplete SNV density, the result is a \n“biased” model whose predictions of genetic constraint will be flawed. Even if a neutral model \ncan perfectly recapitulate the expected distribution of SNV counts in a locus, one's ability to \npredict that the locus is under purifying selection is still limited by how large the “signal” (the \ndepletion of SNV counts due to purifying selection) is relative to “noise” (unavoidable statistical \nfluctuations in observed SNV counts). Motivated by these concerns, we investigate how existing \n2 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nmodels’ ability to predict genetic constraint is impacted by variations in model bias and \nsignal-to-noise ratio.  \nResults  \nConsider a recent model of neutral genetic diversity in humans (Chen et al. 2024), which \npredicts the distribution of the number of SNVs in a given genomic window, aggregated over \n76,156 humans, as a function of local sequence context, nucleotide-level DNA methylation, and \na comprehensive set of regional genomic features, including GC content and recombination \nrates. This model yields a score called Gnocchi, denoted  below, which evaluates how typical \nor unusual an observed SNV count is in the context of the window containing the SNVs,  \n \nwhere the residual, , is defined by  \n,  \nand  is the observed SNV count, and  and  are the mean and standard deviation, \nrespectively, of the predicted distribution of SNV counts for that window. A value of Gnocchi \nnear zero indicates that the observed SNV count is close to what one would expect from the \nneutral model, . Gnocchi is a standardized residual: the residual, , is reported relative to the \nexpected background noise, . Accordingly,  indicates the SNV count  is not only \nsmaller than , but also smaller than expected statistical fluctuations below , raising the \npossibility (but not guaranteeing) that negative selection is operative in that window.  \nStandardized residuals provide information about how the underlying model of neutral diversity \nis biased. To illustrate how, we developed a simulation consisting of the following sequence of \nsteps. (i) We simulated mutation by random generation of SNVs in a set of mock genomic \nwindows as a function of a mock genomic feature (an example of a real “genomic feature” is \n“nucleotide content”) distributed normally over windows (Figure 1A). The “generative process” \nis characterized by a “true rate” at which mutations are generated as a function of the genomic \nfeature (black line in Figure 1B-D), and sampling from a Poisson distribution with that rate (grey \ndots in Figure 1B-D). (ii) We simulated the development of the model of neutral diversity \nunderlying Gnocchi by estimating the parameters of various models that map the feature value \nof each window to its SNV count. Like the generative process, the estimated model has its own \nrate, which we call the “learned rate” (orange line in Figure 1B-D) (iii) We simulated Gnocchi by \ncombining the results of (i) and (ii) to compute a standardized residual (grey dots in Figure \n1E-G).  \nThe difference between the estimated model and the generative process is quantified by “global \nmodel bias”, defined as the squared difference between the learned rate and the true rate, \naveraged over all windows (Bishop 2006). We started by using the generated SNVs to estimate \na model of the same functional form as the generative process (the “Quadratic model”; Figure \n1B). Consequently, after training, the estimated model recapitulated the simulation's generative \n3 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nprocess, resulting in zero global model bias (Figure 1B). Corresponding to this lack of bias, the \ndistribution of the standardized residual is randomly scattered around zero independent of \ngenomic feature, indicating that observed and predicted SNV counts agree, on average (Figure \n1E). If the estimated model has a simpler functional form than the generative process, such as \nwhen the estimated model is linear but the generative process is quadratic, then the global bias \nbecomes non-zero (Figure 1C). Associated with the global bias is a “feature-specific bias”, \ndefined as the mean of the residuals as a function of feature value (orange line in Figure 1F). \nThe feature-specific bias shifts away from zero, particularly in the tails of the distribution of \nfeature values, indicating that the model systematically over- or under-estimates observed \ngenetic diversity. When the model is an even poorer approximation of the generative process, \nthe global bias is even larger (Figure 1D), and the feature-specific bias deviates significantly \nfrom zero even in the center of the feature distribution (Figure 1G).  \n \nFigure 1 Biased models over- or under-estimate SNV counts, on average, in the tails of the feature distribution. (Top \nrow) Given a mock set of genomic intervals, each with a genomic feature (distributed as shown in A) and a \ncorresponding SNV count (y-axis in B-D), we used Poisson regression to fit three different models mapping the \nfeature to the SNV count (B-D). The quadratic model fits the data perfectly as SNV counts were generated by \nsampling from a Poisson distribution with a rate parameter that is a quadratic function of the genomic feature (note \nthe concordance of the black line and orange line in B). The linear and constant models of the rate parameter are \nprogressively less complex than the true rate parameter. In these cases, minimization of the sum of negative log \nlikelihoods over all the windows during training forces the learned models to pass through the observed data where \nthey are dense at the expense of deviating systematically from them where they are sparse: compare the grey data \npoints and the orange line in panels C and D. (Bottom row) The degree to which the models over- or under-estimate \nSNV count can be quantified by the standardized residuals, shown in grey. The average of these data points (orange \nline) show the feature-dependent biases that emerge when learned models are incomplete. The simulation can be \nfound at the following URL: \nhttp://github.com/quinlan-lab/constraint-tools/blob/main/papers/neutral_models_are_biased/9.regression/simulation.4.\n2.1.interplay-of-SNR-and-bias.ipynb. \n4 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nMotivated by these results, we sought to visualize the bias of the Chen model by plotting the \ndistribution of Gnocchi on a set of   693,270 putatively neutral 1kb intervals (Methods) as a \nfunction of three features (Figure 2A-C): a measure of nucleotide composition (GC content); a \nmeasure of background selection (BGS) computed by (Murphy et al. 2023); and a measure of \nGC-biased gene conversion (gBGC) computed by (Glémin et al. 2015). We chose GC content \nas single-nucleotide substitution probabilities have long been known to decline as GC content \nincreases (Fryxell and Moon 2005; Arndt et al. 2005; Mugal and Ellegren 2011). Additionally, \nthere is evidence that GC content at kilobase scale affects the probability of a CpG site \nmutating, as the deamination process requires local DNA melting, which in turn requires more \nthermodynamic energy for bonds between G and C nucleotides (Elango et al. 2008). Finally, \neven the GC content of the local sequence context of a CpG site may correlate with its mutation \nrate (Chandra and Gao 2025).  \nThough BGS and gBGC metrics are not explicitly included in the Chen model, they are neutral \nmechanisms well-known to affect diversity in the human genome (McVicker et al. 2009; Pouyet \net al. 2018; Murphy et al. 2023), and have been modeled when inferring natural selection in the \ncontext of population genetics, e.g., (Telis et al. 2020). BGS, as a process, reduces variation in \nregions linked to sequences constrained by purifying selection  (for a review see (Charlesworth \n2013)). GC-biased gene conversion (reviewed in (Duret and Galtier 2009)) describes how \nmismatch repair of heteroduplexes formed during meiosis is biased towards GC alleles, and \nthus against AT alleles, altering their probability of fixation. Therefore allele frequency is also \naltered in a manner that mimics natural selection though is neutral (Duret and Galtier 2009). As \nsuch, like GC content, both BGS and gBGC are neutral processes that lead to reduced genetic \ndiversity which could masquerade as selection to models that fail to account for them. \n5 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\n \nFigure 2 Visualizing the bias of the Chen model of neutral genetic diversity. (A-C) The standardized residuals of the \nChen model (Gnocchi) deviate from zero, on average, when the genomic windows they are assessed on exhibit \nextreme values of GC content, background selection (BGS) or GC-biased gene conversion (gBGC). Since we report \nthe (standardized) rank of Gnocchi on the y-axes, the marginal distribution of Gnocchi (the distribution of Gnocchi \nover all windows, irrespective of feature value) is uniform with an average value of 0.5 (horizontal black lines). The \ndark-grey lines are computed by conditioning on each feature bin before computing the mean Gnocchi rank. Note that \nthe slope of the dark-grey lines denote the marginal (not independent) effects of each feature on bias. For example, a \nposition on the dark-grey line in (B) equals the value of Gnocchi conditioned on a specific value of BGS, but \nintegrated over all possible values of GC content and gBGC. The dark-grey lines deviate from the expected value of \n0.5, particularly in the tails of the feature distributions (where the heat map is dark). However the dark-grey lines lie at \na height of approximately 0.5 when conditioned upon the mean of the feature distributions (vertical black lines). For \n(B), this implies that the Chen model captures the expected reduction in SNV counts arising from the weak levels of \nBGS present in the majority of windows, even though the model does not explicitly include BGS, presumably because \nit covaries with features that the model does include. The co-variation of features with one another allows the linear \nregression (light-grey line) to track the observed nonlinear behavior (dark-grey line) in (B). (D) Even in the vicinity of \nmean GC content, the standardized residuals of the Chen model (Gnocchi) exhibit greater variance than expected \n(from a standard normal distribution), indicating the existence of neutral diversity not captured by the model. GC \ncontent was chosen to be within 0.1 * std(GC_content) of the mean GC content, where “std” refers to “standard \ndeviation”. The curve marked “Expected” was constructed by scaling a standard normal distribution according to the \nbin width of the histogram and the number of intervals. An explanation of why a standard normal distribution is \nexpected using a simplification of the Chen model can be found here: \nhttps://github.com/quinlan-lab/constraint-tools/blob/main/papers/neutral_models_are_biased/10.residuals-are-over-di\nspersed.ipynb. \n6 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nWe observe bias (represented by the deviation of the dark-grey line from y = 0.5 in Figure \n2A-C) in the Chen model in rare windows that harbor extreme levels of background selection, or \nGC-biased gene conversion, but particularly those windows with extreme levels of GC content. \nNote, however, that the lack of bias in the Chen model at intermediate feature levels does not \nimply that the model cannot be improved there. In fact, we observe that the variance of Gnocchi \nin that vicinity is larger than what would be expected (Figure 2D), implying that there remains \nvariation in SNV count that the Chen model is unable to capture even at genomic windows \nwhose feature values are observed frequently across the genome.  \nWe next sought to measure how each feature independently contributed to the Chen model’s \nbias. The slopes of the dark-grey lines in Figure 2A-C do not measure these contributions. \nInstead, these slopes represent not only the direct effect of any given feature (e.g, GC content) \non model bias, but also the indirect effect of the complementary features (BGS metric and \ngBGC) due to correlation among these features. To capture the independent effect of each \nfeature on model bias, we fit a linear model of all features to Gnocchi (Methods). The estimated \nparameters of the linear model represent the direct effects of the corresponding feature on the \nbias of the Chen model (Table 1). We find, for example, that the direct effect of the BGS metric \non Gnocchi is negative, implying that stronger background selection (smaller values of the BGS \nmetric) is associated with larger values of Gnocchi. This accords with our naive expectation that \na model of neutral diversity that does not explicitly account for BGS (as is the case with the \nChen model) ought to overestimate diversity where background selection is strong (assuming \nthat BGS does not significantly co-vary with features that the Chen model does include). \nInterestingly, though a positive value of Gnocchi erroneously predicts constraint in the neutral \nwindows subject to strong BGS, it correctly predicts constraint nearby, by definition of BGS. \nSince we standardized all features prior to fitting our linear model to Gnocchi, we may compare \nthe values of the corresponding regression coefficients, revealing that GC content has the \nlargest effect (in magnitude) on model bias: GC content explains almost as much variance in \nGnocchi as explained by the three features jointly (compare  for one feature vs three features, \nrespectively, in Table 2).  \nTable 1. The direct effect of each feature on the bias of the neutral models underlying each constraint metric. The \nregression coefficients may be compared not only within a given model, but also among models, since the features \nare standardized prior to fitting, and the target variables (ranks) in the fitting process are uniformly distributed \nbetween 0 and 1. \n7 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\n \nTable 2. The proportion of variance in each constraint score that is explained, ,  by all three features considered in \nthis study, and by the single feature with the largest direct effect on bias. \n \nTo assess the generality of these results, we repeated our analysis for three other constraint \nmetrics:  (Dukler et al. 2022), Depletion Rank (Halldorsson et al. 2022) and CDTS (di Iulio et λ𝑠\nal. 2018) (Methods). The results, presented in Tables 1 and 2, and Supporting Figure 1, are \nsimilar to those obtained for the Chen model in the sense that all models develop biases as one \nmoves into the tails of the feature distributions, though the details differ. For example, whereas \nhigh GC content biases Gnocchi above zero (rank > 0.5), it biases  below zero (rank < 0.5) λ𝑠\n(Supporting Figure 1). These high-GC-content intervals may be enriched for promoter regions, \nwhere the authors of the   metric observed a similar bias, which they attributed to λ𝑠\nmisspecification of the underlying neutral model there (Dukler et al. 2022). Another difference \n8 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nbetween the Chen model and the others is that BGS, not GC content, explains most of the bias \nobserved in Depletion Rank and CDTS (e.g., Table 2). It is unclear why, but part of the answer \nmay lie in the fact that the Chen model includes GC content as a factor, but the models \nunderlying Depletion Rank and CDTS do not.  \nThe procedure we adopted to identify the feature that drives model bias is closely related to \nGradient Boosting (Friedman 2001), a technique for improving a model by regressing its \nresiduals onto its features, and then adding the resulting “residual model” to the predictions of \nthe original model. Viewed this way, the  values reported in Table 2 measure how much an \nimproved modeling of the corresponding features (e.g., via Gradient Boosting) is expected to \nimprove the underlying model of neutral diversity. Comparing the  values among features and \nmodels leads to two interpretations. First, the largest improvement to the Chen model is \nexpected to come from more complex representations of sequence composition (for which GC \ncontent is a surrogate), whereas the largest improvements to the models underlying Depletion \nRank and CDTS are expected to come from better representations of BGS. Second, of all the \nmodels, the Chen model stands to benefit most from an improved modeling of the three features \nconsidered in this study.  \nThe purpose of constraint metrics, such as Gnocchi, is to discover “genetic constraint”, i.e., to \ndistinguish genomic windows that are neutral from those that are under purifying selection. To \nunderstand how well this approach is likely to work under various circumstances, we returned to \nsimulation. The standard approach to modeling purifying selection is to use forward-time \npopulation genetic simulations, e.g., via stdpopsim (Adrion et al. 2020; Lauterbur et al. 2023; \nGower et al. 2025). Though stdpopsim incorporates a host of neutral processes, each of them is \nexpected to contribute to “noise” (unavoidable statistical fluctuations in observed SNV counts in \ngenomic windows), obscuring the patterns we wish to illustrate with our simulations. To minimize \nthis noise, and therefore maximize our ability to recover “signal” (the depletion of SNV counts \ndue solely to purifying selection), we decided instead to return to our simple simulation of \nneutral diversity (Figure 1A-D), which is stripped of most of the neutral processes included in \nstdpopsim. We extended the simple simulation to model purifying selection by simply reducing \nthe SNV counts of a random subset of our simulated windows, labeling them “constrained”, and \nlabeling the rest “neutral”. We then built a constraint predictor by thresholding residuals \ncorresponding to each of the models of neutrality depicted in Figure 1B-D: windows with \nresiduals greater than the threshold were predicted to be constrained, whereas those with \nresiduals lower than the threshold were predicted to be neutral. Finally, we assessed the \nperformance of those models by comparing predicted labels with actual labels.   \nEven if a model is unbiased, its performance is limited by the signal-to-noise ratio (SNR): the \nratio of the “signal” we are trying to recover (the depletion of SNV counts due to purifying \nselection) and the unwanted background “noise” we must contend with (unavoidable statistical \nfluctuations in observed SNV counts). To illustrate this problem and the effect of variations in \nSNR along the genome, we conditioned on windows with three different values of the genomic \nfeature (Figure 3A-C), and in each case computed the SNR (Figure 3D-F; Methods). \nRegardless of model bias, performance declines (Figure 3G-I) as SNR declines (reflected in an \n9 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nincreasing overlap of the residual distributions of the neutral and constrained windows; Figure \n3D-F).  \nPerformance is impacted not only by low SNR, but also by high model bias. To deconvolve \nthese effects, we introduced, for each of the three cases in Figure 3A-C, enough variability in \nfeature values to expose the impact of bias on the performance of the three models, yet not \nenough to alter the SNR in each case (compare narrow with wide red strips in Figure 3A-C). \nWith this change, not only does performance decline due to declining SNR as found previously \n(notice the similarity of the green lines in Figure 3J-L with the corresponding green lines in \nFigure 3G-I), but also because of increasing bias for any given SNR (for each panel of Figure \n3J-L, compare the three colored lines).  \n \nFigure 3 SNR and model bias are both expected to significantly affect the ability to call constraint. The narrow red \nstrips in panels A-C define the subsets of simulated genomic windows that were used to compute panels D-I. Each \nsubset contains neutral (orange) and constrained (blue) windows with residuals distributed as shown in panels D-F \n(with respect to the unbiased quadratic model). Also shown in those panels are the corresponding values of SNR, \ncomputed from the the standard deviation of the neutral residuals and the depletion  of the simulated SNV counts \nin the constrained windows (Methods). The SNV-count depletion ( ) and class-imbalance (relative numbers of \n“constrained” and “neutral” windows) were chosen to be representative of real data (Supporting Figure 2). For each \nwindow, and each model, we predict that the window is constrained if the residual exceeds a threshold; by varying the \nthreshold we sweep out a PR curve for each model, showing the independent effect of SNR on model performance \n(panels G-I). To see the independent effect of model bias, compare the PR curves for each model in panels J-L, \n10 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nwhich correspond to the wide red strips in panels A-C, respectively. “Local bias” is defined similarly to the definition of \n“Global bias” in Figure 1, but this time using only those genomic windows that comprise the corresponding red strips \nin panels A-C. The distributions of the constrained windows in panels D-F peak at , which measures the \ndepletion of SNV count that we implemented in the simulation, which can be found here. Supporting Figure 3 shows \nsimilar results when the class distributions are computed using standardized residuals, instead of raw residuals, as \nshown in panels D-F. Supporting Figure 4 shows similar results when SNV counts are depleted by an amount \nproportional to the feature value (i.e.,  is now feature-dependent), except that now SNR increases (instead of \ndecreasing) as one moves into the right tail of the feature distribution.  \nGiven that model bias and SNR are expected to jointly affect the ability to predict constraint in a \nway that is impossible to predict without knowledge of that which we are seeking (genetic \nconstraint), we set out to empirically measure performance of Gnocchi as a function of GC \ncontent, BGS and gBGC. We designed a simple classifier that predicts that a window is \nconstrained if Gnocchi exceeds a threshold. Since the main purpose of, and main challenge \nfacing, scores such as Gnocchi is to identify constraint in the noncoding genome, an example of \nwhich may be a conserved enhancer, we constructed a set of noncoding windows and labeled \nthose that overlap a GeneHancer enhancer (Fishilevich et al. 2017) “constrained”, and those \nthat don’t “neutral” (Methods). Measured against this truth set, Gnocchi clearly does better than \nrandom guessing (compare black line and dashed line in Figure 4A), but its performance in \nGC-rich windows is significantly reduced (red lines); conversely, in GC-poor windows, \nperformance is better than expected (blue lines). Since Gnocchi is biased at low GC content \n(Figure 2A), and in light of the results in Figure 3, we conjecture that the increased \nperformance at predicting constraint in GC-poor regions is due to an increased SNR. Similar \nvariability in performance was observed as a function of BGS and gBGC (Supporting Figure \n5A, B). These results illustrate how model bias and/or SNR jointly generate significant variation \nin ability to discover constraint under the Chen model, depending upon the nature of the \nwindows being interrogated.  \n11 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\n \nFigure 4 The ability of models of neutral diversity to predict constraint varies with the nature of the genomic windows. \n(A) The performance of the Chen model at predicting constraint (on a lax truth set defined by whether windows \noverlap GeneHancer enhancers) can be made to vary significantly by changing the nature of the windows upon which \nperformance is assessed. We varied the Gnocchi threshold to trace out the precision-recall curves. As a benchmark, \nthe performance of a classifier that randomly guesses whether a window is constrained or not has a precision pegged \nto the fraction of windows that are “positive”, i.e., labeled “constrained” (dashed line; see Methods). To compare \nperformance across GC-content bins, which vary in the proportion of windows labeled “constrained”, we \ndownsampled the bins so that each had an equal fraction of “constrained” windows. (B) Performance varies with GC \ncontent in a largely similar way for all constraint metrics, at least for low to intermediate GC content. Performance \nwas defined to be the area under the precision-recall curve, auPRC, normalized by the fraction of examples that are \npositive, which varies among the lax truth sets. (C) We downsampled the lax truth set so that its size matched that of \na stringent truth set, defined by essential genes (Methods), and used 1000 bootstraps of the truth sets to compute a \nmean and standard deviation of normalized auPRC for each constraint metric and each truth set (without stratifying \nby genomic features). (D) For each bootstrap sample, we computed the difference in normalized auPRC between \nbins of high and low values of GC content ((0.4, 0.7) and (0.2, 0.375), respectively), and reported the mean and \nstandard deviation of those differences for each constraint metric relative to the stringent truth set. The results show \nthat GC content is associated with changes in the performance of all constraint metrics relative to the stringent truth \nset. Code to reproduce C and D can be found here: \nhttps://github.com/quinlan-lab/constraint-tools/blob/main/papers/neutral_models_are_biased/11.compare-lax-with-stri\nngent-truth-set.ipynb. In all panels, GC content was assessed in 1kb windows.  \nSimilar trends in performance are observed for the other constraint metrics, with the notable \nexception of CDTS and Depletion Rank, which increase in performance in GC-rich and/or \ngBGC-rich regions (Figure 4B, and Supporting Figure 5C and D, respectively). Interestingly, \nthese data show that, of all the models, the Chen model has the best average performance, i.e., \nwhen evaluated at typical values of the features (vertical black lines in Figure 2A-C), even \nthough it has the greatest bias (see  in Table 2). Trading significant modeling bias on rare \n12 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nwindows for high model performance on common windows has been observed elsewhere \n(Kathail et al. 2024a).  \nThe number of windows in the truth set that we used to measure performance is large, enabling \nus to detect performance anomalies that occur deep in the tails of the feature distributions. But \nthat truth set is also lax, as not all windows that overlap a Genehancer enhancer are expected \nto be highly constrained. For example, 18.4% of the noncoding genome is covered by \nGenehancer enhancers (Martini et al. 2023), but others have estimated that only 4.51% of the \nnoncoding genome is under human-specific selection (Huber et al. 2020). We therefore \nidentified noncoding windows that we conjectured would be under strong negative selection \nbecause they regulate the expression of essential genes (Methods). We combined them with an \nequal number of noncoding windows not overlapping any enhancer to form a small but stringent \ntruth set. We used it to re-assess the performance of the various constraint metrics at predicting \nconstraint. We observed that overall performance increases on moving from the lax to the \nstringent truth set, particularly for the Gnocchi and   metrics (Figure 4C and Supporting λ𝑠\nFigure 5E, F), providing a post-hoc validation of stringency. Even though the stringent truth set \nis small–making estimates of performance in the tails of the feature distributions noisy–we were \nnevertheless able to detect associations between performance and features (Figure 4D and \nSupporting Figure 5G,H), as we observed for the lax truth set (Figure 4B and Supporting \nFigure 5C,D).  \nDiscussion \nSequence conservation leverages the power of deep evolutionary time to provide strong signals \nof purifying selection, but this approach is limited to the discovery of selection common to a \ngroup of species. This forces us to look at human-specific diversity-based methods in order to \nfind human-specific selection. The trade-off is that inference of selection via methods that rely \non segregating sites is confounded in a number of ways. (1) Though rare, the fixation of an \nadvantageous allele can drag linked, neutral DNA with it, sweeping away existing variation in \nthat genomic region (Smith and Haigh 1974; Stephan 2019), and naively leading to \nfalse-positive calls of purifying selection. (2) During a population bottleneck, rare alleles tend to \nbe lost quickly, before heterozygosity has been eroded, resulting in a transient imbalance \nbetween heterozygosity and allele number (Luikart et al. 1998; Amos and Hoffman 2010) that \none could falsely interpret as the active removal of deleterious alleles in that interval by purifying \nselection, if demography is not accounted for. (3) The demographic history of a population can \nreduce the ability of natural selection to remove weakly deleterious variants (Lohmueller 2014), \nand so diversity-based methods, which rely on the removal of deleterious variants, will not \ndiscover constrained regions harboring those variants. Here, we identify additional factors that \nconfound our interpretation of diversity-based methods for genetic-constraint calling.  \nWe show that current models of neutral genetic diversity are biased, compromising their ability \nto uncover constraint in the human noncoding genome in rare but important contexts, e.g., high \nGC content (enriched for promoters) or strong background selection (enriched for proximity to \n13 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\ngenes), depending upon the model. This mirrors similar findings in the coding genome for a \nvariety of organisms (Yap et al. 2010; Venkat et al. 2018; Ratnakumar et al. 2010).  \nSimilar to the difficulty in capturing edge cases when training machine-learning models to drive \ncars or answer questions in natural language, models of genetic diversity tend to capture the \naverage behavior seen in a putatively neutral training set, at the expense of incurring large \nprediction errors on classes of genomic intervals that are rare in that training set. Capturing only \naverage behavior is a blessing when rare events are true anomalies (e.g., constrained windows \nthat are erroneously part of the training set), but a curse when they are not (e.g., truly neutral \nwindows with high GC content).   \nWe find that the model with the highest performance on both the lax and stringent truth set—the \nChen model—is also the model with the largest bias, suggesting that bias is a natural tradeoff of \nimproved performance. Interestingly, the direct effect of GC content on bias is larger than that of \ngBGC and BGS, even though the former was explicitly included in the model whereas the latter \nwere not. Since performance of the Chen model declines with GC content (Figure 4A), and \nsince GC content is positively correlated with Gnocchi (Supporting Figure 6A,B), we expect \nthat Gnocchi is most unreliable precisely where it predicts the most constraint. Moreover, given \nthat GC content is auto correlated across the human genome, forming “isochores” (Supporting \nFigure 6B), we expect these false calls of constraint under the Chen model to cluster in \ngenomic space.The relatively high performance of the Chen model (marginalized over all \ngenomic feature values) may result from how comprehensive its set of features is, but it is likely \nthat the way these features, particularly GC content, are included in the model is still not flexible \nenough to comprehensively capture neutral genetic diversity.  \nDeep-learning models, such as those used to predict the density of noncoding SNVs (Fang et \nal. 2022), or their effects on fitness (Benegas et al. 2024; Trevino et al. 2021), do not suffer from \nsuch limitations on model complexity. Nevertheless, a deep-learning model mapping genomic \nfeatures onto SNV density may still fall foul to the type of biases we illustrated here, as recently \nillustrated for other genomic prediction tasks (Kathail et al. 2024b; Bajwa et al. 2024). To the \nextent that this failure mode is due to insufficient drive to improve the fit in sparsely populated \nparts of feature space, a possible solution is to up-weight training examples there, e.g., by \nweighting a genomic interval in inverse proportion to the number of intervals similar to it (Morcos \net al. 2011).  \nIn the meantime, we recommend annotating each genomic interval not only with its constraint \nscore, but also a measure of how biased the score is expected to be in that interval, as a way to \ngauge confidence in the score. This can be done by grouping genomic intervals by the similarity \nof their feature vectors (with elements comprising GC content, etc), computing the distribution of \nconstraint scores in each group, and assigning to each window the mean of the constraint-score \ndistribution for the corresponding group.  \n14 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nMethods \nProvenance of constraint scores \nGnocchi scores for a set of nonoverlapping 1000bp windows (“Chen windows”) were obtained \nfrom here (Supplementary Data #2). Windows on the X and Y chromosomes were omitted. Drs \nNurdan Kuru and Adam Siepel ran ExtRaINSIGHT (Dukler et al. 2022) on the remaining Chen \nwindows in order to compute , excluding Chen windows that did not pass the ExtRaINSIGHT λ𝑠\nfilters (private communication). A set of overlapping 500bp windows (“Halldorsson windows”) \nwith Depletion Rank scores (Halldorsson et al. 2022) was obtained from here. A set of \noverlapping 550bp windows with CDTS scores (“CDTS windows”) (di Iulio et al. 2018) was \nobtained by private communication with Dr Amalio Telenti (Vir Biotechnology), and their \ncoordinates were lifted over from hg19 to hg38.  \nAssignment of genomic feature values to genomic windows \nWe computed GC content in 1kb intervals centered on the Chen, Halldorsson and CDTS \nwindows using “bedtools nuc”, e.g., as described here and here. We obtained BGS values \ncomputed by (Murphy et al. 2023) from here, which we converted to bed format and lifted over \nfrom hg19 to hg38 (see here). We obtained the gBGC coefficients on a set of genomic intervals \nfor the European population computed by (Glémin et al. 2015) from here, and lifted over the \ngenomic coordinates from hg18 to hg38, as described here. We assigned BGS values and \ngBGC coefficients (B_M1star.EUR) to the Chen, Halldorsson and CDTS windows by \nintersecting them with the bed files of BGS and gBGC, e.g., as shown here and here.  \nConstruction of the window sets to assess model bias, and the lax truth sets to assess \nconstraint-prediction performance \nWe filtered out untrustworthy windows, defined as windows that overlap any of the following \nintervals: gaps in the hg38 genome assembly, Encode “exclude regions” (Amemiya et al. 2019), \nand regions with insufficient read coverage in Gnomad version 3 (Chen et al. 2024). Exons were \ncomputed as shown here and merged using “bedtools merge”. Noncoding windows were \ndefined to be Chen, Halldorsson and CDTS windows that don’t significantly overlap merged \nexons. Of the noncoding windows, those that don’t significantly overlap Genehancer enhancers \n(Fishilevich et al. 2017) were labeled “neutral”, and were used to assess model bias. Noncoding \nwindows that do significantly overlap Genehancer enhancers were labeled “constrained”. \nTogether the “neutral” and “constrained” windows comprised the lax truth set used to assess the \nperformance of each of the four constraint metrics.  \nModeling the mapping from genomic features to constraint scores \nThe code to compute the linear regressions of the constraint scores shown in the heat maps, \nincluding the regression coefficients and the proportion of variance explained, reported in the \ntables, can be found here.  \n15 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nComputation of window residuals under the Chen model  \nEach window has a residual, , equal to the difference between the expected SNV count under \nthe Chen model of neutrality, , and the observed SNV count, . We computed  from \nGnomad Version 3, using the bash command “constraint-tools predict-germline-model-Nonly \n--windows ${chen_windows} REQUIRED_ARGUMENTS”, where REQUIRED_ARGUMENTS \nare specified here. We then used , in conjunction with Gnocchi, , to compute  using the \nfollowing algorithm:  \n  a = 1  \n  b = -(2*S + G**2) \n  c = S**2 \n  s = sqrt(b**2 - 4*a*c)  \n  sign = 1 if G > 0 else -1 \n   = (-b + sign*s)/(2*a)  \nFinally: \n.  \nComputation of the PR curve for the random classifier \nGiven a truth set of size , let  represent the proportion of windows that are constrained. Let \n represent the probability that a random classifier predicts that a given window is constrained. \nThen  \n \nand  \n \nwhere TP (average number of True Positives) is the average number of windows that are \nlabeled “constrained” in the truth set and are predicted to be constrained by the classifier. \nSimilar definitions exist for FP (average number of False Positives) and FN (average number of \nFalse Negatives). The formulae for Precision and Recall describe a horizontal line on the \nPrecision-Recall curve (e.g., the dashed line in Figure 4A, traced out as the parameter  is \nvaried between 0 and 1) and represents a baseline level of performance that any meaningful \nconstraint predictor ought to exceed.  \nComputation of SNR \nThe signal-to-noise ratio (SNR) is defined as  \n, \n16 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nwhere  is the random residual of a constrained window,  is the random residual of a neutral \nwindow, and  represents the expectation operator. The numerator is the amplitude of the \nmeaningful signal we are trying to recover (a dip in SNV count due to purifying selection); the \ndenominator is the amplitude of the unwanted background noise we must contend with \n(unavoidable Poisson fluctuations in SNV count). Squaring the relation \n , \nwhere  is the reduction in SNV count induced by purifying selection, we arrive at:  \n .  \nApplying the expectation operator, , we get:  \n , \nwhere we used the facts that  is constant (if  is feature-dependent, then specialize this \nanalysis to a particular bin of that feature), , and . \nSubstituting these expressions for  and  into the definition of SNR, we get \n . \nWhen , SNR is approximately , indicating that the amplitudes of the signal and noise \nare approximately the same, making it impossible to recover the signal. This is reflected in a flat \nPrecision-Recall (PR) curve (e.g., the dashed line in Figure 4A), which represents a baseline \nlevel of performance that any meaningful constraint predictor ought to exceed. We used the \nSNR formula to compute the numerical values of SNR reported in Figure 3 (and in Supporting \nFigure 4, where  is feature-dependent).  \nWhy PR curves for all models overlap when there is little variability in the mock genomic \nfeature (Figure 3G-I)  \nConsider a set of simulated windows each with approximately the same value of the mock \ngenomic feature. Under a given model, the predicted SNV count will be approximately the same \nfor all these windows, because these models depend upon only that genomic feature. Now \nsubset the windows into those that are constrained, and those that are not (neutral). The \nresidual distributions for the two classes of windows (constrained and neutral) are just the \nobserved SNV count distributions for the two classes, offset by a constant–the common \npredicted SNV count. Therefore the overlap of the residual distributions of the two classes is \napproximately the same for all models, and so all models perform approximately equally. \nConstruction of the stringent truth set to assess performance of constraint scores \nChen et al (Chen et al. 2024) obtained a set of highly constrained genes (ClinGen \nhaploinsufficient genes, MGI essential genes, OMIM autosomal dominant genes, LOEUF \n17 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nfirst-decile genes), linked enhancers to each gene using correlated patterns of histone \nmodifications and gene expression, and assigned the most constrained enhancer to each gene, \nyielding a set of enhancers that we downloaded from here (Supplementary Data #6). We \nassigned genomic feature values, and constraint scores, to each of these “essential” enhancers \nsimilarly to how we did the assignments for the Chen, Halldorsson and CDTS windows. Finally, \nwe constructed the stringent truth set by combining these “constrained” windows with an equal \nnumber of “neutral” windows sampled from the Chen lax truth set. See here. \nFunding \nThis work was supported by a National Institutes of Health grant to ARQ (R01HG012252). \nAcknowledgements \nWe graciously thank Nurdan Kuru and Adam Siepel for sharing their ExtRaINSIGHT data. We \nare grateful to Andrew Kern for expert feedback that significantly improved on an early draft of \nthe manuscript.  \nReferences \nAdrion, Jeffrey R., Christopher B. Cole, Noah Dukler, et al. 2020. “A Community-Maintained \nStandard Library of Population Genetic Models.” eLife 9 (e54967). \nhttps://doi.org/10.7554/eLife.54967. \nAmemiya, Haley M., Anshul Kundaje, and Alan P. Boyle. 2019. “The ENCODE Blacklist: \nIdentification of Problematic Regions of the Genome.” Scientific Reports 9 (1): 9354. \nAmos, W., and J. I. Hoffman. 2010. “Evidence That Two Main Bottleneck Events Shaped \nModern Human Genetic Diversity.” Proceedings. Biological Sciences 277 (1678): 131–137. \nArndt, Peter F., Terence Hwa, and Dmitri A. Petrov. 2005. “Substantial Regional Variation in \nSubstitution Rates in the Human Genome: Importance of GC Content, Gene Density, and \nTelomere-Specific Effects.” Journal of Molecular Evolution 60 (6): 748–763. \nBajwa, Ayesha, Ruchir Rastogi, Pooja Kathail, Richard W. Shuai, and Nilah Ioannidis. 2024. \n“Characterizing Uncertainty in Predictions of Genomic Sequence-to-Activity Models.” \nMachine Learning in Computational Biology, March 15, 279–297. \nBenegas, Gonzalo, Carlos Albors, Alan J. Aw, Chengzhong Ye, and Yun S. Song. 2024. \n“GPN-MSA: An Alignment-Based DNA Language Model for Genome-Wide Variant Effect \nPrediction.” bioRxiv.org: The Preprint Server for Biology, ahead of print, April 6. \nhttps://doi.org/10.1101/2023.10.10.561776. \nBishop, Christopher M. 2006. Pattern Recognition and Machine Learning. 1 (August): 740. \nBoffelli, Dario, Jon McAuliffe, Dmitriy Ovcharenko, et al. 2003. “Phylogenetic Shadowing of \nPrimate Sequences to Find Functional Regions of the Human Genome.” Science (New \nYork, N.Y.) 299 (5611): 1391–1394. \n18 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nBulmer, M. 1986. “Neighboring Base Effects on Substitution Rates in Pseudogenes.” Molecular \nBiology and Evolution 3 (4): 322–329. \nChandra, Sheel, and Ziyue Gao. 2025. “Sequence Context and Methylation Interact to Shape \nGermline Mutation Rate Variation at CpG Sites.” In bioRxiv. November 13. \nhttps://doi.org/10.1101/2025.11.13.688199. \nCharlesworth, Brian. 2013. “Background Selection 20 Years on: The Wilhelmine E. Key 2012 \nInvitational Lecture.” The Journal of Heredity 104 (2): 161–171. \nChen, Siwei, Laurent C. Francioli, Julia K. Goodrich, et al. 2024. “A Genomic Mutational \nConstraint Map Using Variation in 76,156 Human Genomes.” Nature 625 (7993): 92–100. \nCooper, Gregory M., Eric A. Stone, George Asimenos, et al. 2005. “Distribution and Intensity of \nConstraint in Mammalian Genomic Sequence.” Genome Research 15 (7): 901–913. \nDavydov, Eugene V., David L. Goode, Marina Sirota, Gregory M. Cooper, Arend Sidow, and \nSerafim Batzoglou. 2010. “Identifying a High Fraction of the Human Genome to Be under \nSelective Constraint Using GERP++.” PLoS Computational Biology 6 (12): e1001025. \nDukler, Noah, Mehreen R. Mughal, Ritika Ramani, Yi-Fei Huang, and Adam Siepel. 2022. \n“Extreme Purifying Selection against Point Mutations in the Human Genome.” Nature \nCommunications 13 (1): 4312. \nDuret, Laurent, and Nicolas Galtier. 2009. “Biased Gene Conversion and the Evolution of \nMammalian Genomic Landscapes.” Annual Review of Genomics and Human Genetics 10: \n285–311. \nElango, Navin, Seong-Ho Kim, Eric Vigoda, and Soojin V. Yi. 2008. “Mutations of Different \nMolecular Origins Exhibit Contrasting Patterns of Regional Substitution Rate Variation.” \nPLoS Computational Biology 4 (2): e1000015. \nFang, Yiyuan, Shuyi Deng, and Cai Li. 2022. “A Generalizable Deep Learning Framework for \nInferring Fine-Scale Germline Mutation Rate Maps.” Nature Machine Intelligence 4 (12): \n1209–1223. \nFishilevich, Simon, Ron Nudel, Noa Rappaport, et al. 2017. “GeneHancer: Genome-Wide \nIntegration of Enhancers and Target Genes in GeneCards.” Database: The Journal of \nBiological Databases and Curation 2017 (January): bax028. \nFriedman, Jerome H. 2001. “Greedy Function Approximation: A Gradient Boosting Machine.” \nAnnals of Statistics 29 (5): 1189–1232. \nFryxell, Karl J., and Won-Jong Moon. 2005. “CpG Mutation Rates in the Human Genome Are \nHighly Dependent on Local GC Content.” Molecular Biology and Evolution 22 (3): 650–658. \nGlémin, Sylvain, Peter F. Arndt, Philipp W. Messer, Dmitri Petrov, Nicolas Galtier, and Laurent \nDuret. 2015. “Quantification of GC-Biased Gene Conversion in the Human Genome.” \nGenome Research 25 (8): 1215–1228. \nGower, Graham, Nathaniel S. Pope, Murillo F. Rodrigues, et al. 2025. “Accessible, Realistic \nGenome Simulation with Selection Using Stdpopsim.” In bioRxivorg. August 12. \nhttps://doi.org/10.1101/2025.03.23.644823. \n19 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nHalldorsson, Bjarni V., Hannes P. Eggertsson, Kristjan H. S. Moore, et al. 2022. “The \nSequences of 150,119 Genomes in the UK Biobank.” Nature 607 (7920): 732–740. \nHuber, Christian D., Bernard Y. Kim, and Kirk E. Lohmueller. 2020. “Population Genetic Models \nof GERP Scores Suggest Pervasive Turnover of Constrained Sites across Mammalian \nEvolution.” PLoS Genetics 16 (5): e1008827. \nIulio, Julia di, Istvan Bartha, Emily H. M. Wong, et al. 2018. “The Human Noncoding Genome \nDefined by Genetic Diversity.” Nature Genetics 50 (3): 333–337. \nKathail, Pooja, Richard W. Shuai, Ryan Chung, Chun Jimmie Ye, Gabriel B. Loeb, and Nilah M. \nIoannidis. 2024a. “Current Genomic Deep Learning Models Display Decreased \nPerformance in Cell Type Specific Accessible Regions.” bioRxiv.org: The Preprint Server for \nBiology, July 10, 2024.07.05.602265. \nKathail, Pooja, Richard W. Shuai, Ryan Chung, Chun Jimmie Ye, Gabriel B. Loeb, and Nilah M. \nIoannidis. 2024b. “Current Genomic Deep Learning Models Display Decreased \nPerformance in Cell Type-Specific Accessible Regions.” Genome Biology 25 (1): 202. \nLauterbur, M. Elise, Maria Izabel A. Cavassim, Ariella L. Gladstein, et al. 2023. “Expanding the \nStdpopsim Species Catalog, and Lessons Learned for Realistic Genome Simulations.” May \n23. https://doi.org/10.7554/elife.84874. \nLiu, Li, Maxwell D. Sanderford, Ravi Patel, Pramod Chandrashekar, Greg Gibson, and Sudhir \nKumar. 2019. “Biological Relevance of Computationally Predicted Pathogenicity of \nNoncoding Variants.” Nature Communications 10 (1): 330. \nLohmueller, Kirk E. 2014. “The Distribution of Deleterious Genetic Variation in Human \nPopulations.” Current Opinion in Genetics & Development 29 (December): 139–146. \nLuikart, G., F. W. Allendorf, J. M. Cornuet, and W. B. Sherwin. 1998. “Distortion of Allele \nFrequency Distributions Provides a Test for Recent Population Bottlenecks.” The Journal of \nHeredity 89 (3): 238–247. \nMartini, Lorenzo, Roberta Bardini, Alessandro Savino, and Stefano Di Carlo. 2023. “GRAIGH: \nGene Regulation Accessibility Integrating GeneHancer Database.” 2023 IEEE International \nConference on Bioinformatics and Biomedicine (BIBM), December 5, 343–348. \nMcVicker, Graham, David Gordon, Colleen Davis, and Phil Green. 2009. “Widespread Genomic \nSignatures of Natural Selection in Hominid Evolution.” PLoS Genetics 5 (5): e1000471. \nMeader, Stephen, Chris P. Ponting, and Gerton Lunter. 2010. “Massive Turnover of Functional \nSequence in Human and Other Mammalian Genomes.” Genome Research 20 (10): \n1335–1343. \nMorcos, Faruck, Andrea Pagnani, Bryan Lunt, et al. 2011. “Direct-Coupling Analysis of Residue \nCoevolution Captures Native Contacts across Many Protein Families.” Proceedings of the \nNational Academy of Sciences of the United States of America 108 (49): E1293–301. \nMugal, Carina F., and Hans Ellegren. 2011. “Substitution Rate Variation at Human CpG Sites \nCorrelates with Non-CpG Divergence, Methylation Level and GC Content.” Genome \nBiology 12 (6): R58. \n20 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nMurphy, David A., Eyal Elyashiv, Guy Amster, and Guy Sella. 2023. “Broad-Scale Variation in \nHuman Genetic Diversity Levels Is Predicted by Purifying Selection on Coding and \nNon-Coding Elements.” eLife 12 (June). https://doi.org/10.7554/eLife.76065. \nPollard, Katherine S., Melissa J. Hubisz, Kate R. Rosenbloom, and Adam Siepel. 2010. \n“Detection of Nonneutral Substitution Rates on Mammalian Phylogenies.” Genome \nResearch 20 (1): 110–121. \nPonting, Chris P., and Ross C. Hardison. 2011. “What Fraction of the Human Genome Is \nFunctional?” Genome Research 21 (11): 1769–1776. \nPouyet, Fanny, Simon Aeschbacher, Alexandre Thiéry, and Laurent Excoffier. 2018. \n“Background Selection and Biased Gene Conversion Affect More than 95% of the Human \nGenome and Bias Demographic Inferences.” eLife 7 (August). \nhttps://doi.org/10.7554/eLife.36317. \nRands, Chris M., Stephen Meader, Chris P. Ponting, and Gerton Lunter. 2014. “8.2% of the \nHuman Genome Is Constrained: Variation in Rates of Turnover across Functional Element \nClasses in the Human Lineage.” PLoS Genetics 10 (7): e1004525. \nRatnakumar, Abhirami, Sylvain Mousset, Sylvain Glémin, et al. 2010. “Detecting Positive \nSelection within Genomes: The Problem of Biased Gene Conversion.” Philosophical \nTransactions of the Royal Society of London. Series B, Biological Sciences 365 (1552): \n2571–2580. \nSchraiber, Joshua G., Jeffrey P. Spence, and Michael D. Edge. 2024. “Estimation of \nDemography and Mutation Rates from One Million Haploid Genomes.” In Genomics, No. \nBiorxiv;2024.09.18.613708v1. BioRxiv, September 22. \nhttps://www.biorxiv.org/content/10.1101/2024.09.18.613708v1.full.pdf. \nSchrider, Daniel R., and Andrew D. Kern. 2015. “Inferring Selective Constraint from Population \nGenomic Data Suggests Recent Regulatory Turnover in the Human Brain.” Genome \nBiology and Evolution 7 (12): 3511–3528. \nSmith, J. M., and J. Haigh. 1974. “The Hitch-Hiking Effect of a Favourable Gene.” Genetical \nResearch 23 (1): 23–35. \nSmith, Nick G. C., Mikael Brandström, and Hans Ellegren. 2004. “Evidence for Turnover of \nFunctional Noncoding DNA in Mammalian Genome Evolution.” Genomics 84 (5): 806–813. \nStephan, Wolfgang. 2019. “Selective Sweeps.” Genetics 211 (1): 5–13. \nSved, J., and A. Bird. 1990. “The Expected Equilibrium of the CpG Dinucleotide in Vertebrate \nGenomes under a Mutation Model.” Proceedings of the National Academy of Sciences of \nthe United States of America 87 (12): 4692–4696. \nTelis, Natalie, Robin Aguilar, and Kelley Harris. 2020. “Selection against Archaic Hominin \nGenetic Variation in Regulatory Regions.” Nature Ecology & Evolution 4 (11): 1558–1566. \nTrevino, Alexandro E., Fabian Müller, Jimena Andersen, et al. 2021. “Chromatin and \nGene-Regulatory Dynamics of the Developing Human Cerebral Cortex at Single-Cell \nResolution.” Cell 184 (19): 5053–5069.e23. \n21 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint \n\nVenkat, Aarti, Matthew W. Hahn, and Joseph W. Thornton. 2018. “Multinucleotide Mutations \nCause False Inferences of Lineage-Specific Positive Selection.” Nature Ecology & Evolution \n2 (8): 1280–1288. \nWard, Lucas D., and Manolis Kellis. 2012. “Evidence of Abundant Purifying Selection in Humans \nfor Recently Acquired Regulatory Functions.” Science (New York, N.Y.) 337 (6102): \n1675–1678. \nYap, Von Bing, Helen Lindsay, Simon Easteal, and Gavin Huttley. 2010. “Estimates of the Effect \nof Natural Selection on Protein-Coding Content.” Molecular Biology and Evolution 27 (3): \n726–734. \n \n22 \n.CC-BY 4.0 International licenseavailable under a \n(which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in perpetuity. It is made \nThe copyright holder for this preprintthis version posted January 28, 2026. ; https://doi.org/10.64898/2026.01.28.701168doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}