Intrahost dynamics, together with genetic and phenotypic traits predict the success of viral mutations | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Intrahost dynamics, together with genetic and phenotypic traits predict the success of viral mutations Cedric Tan, Marina Escalera-Zamudio, Alexei Yavlinksy, Lucy van Dorp, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5298116/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted You are reading this latest preprint version Abstract Predicting the fitness of mutations in the evolution of pathogens is a long-standing and important, yet largely unsolved problem. In this study, we used SARS-CoV-2 as a model system to explore whether the intrahost diversity of viral infections could provide clues on the relative fitness of single amino acid variants (SAVs). To do so, we analysed ~15 million complete genomes and nearly ~8000 sequencing libraries generated from SARS-CoV-2 infections, which were collected at various timepoints during the COVID-19 pandemic. Across timepoints, we found that many successful SAVs were detected in the intrahost diversity of samples collected prior, with a median of 6-40 months between the initial collection dates of samples and the highest frequency seen for these SAVs. Additionally, we found that the co-occurrence of intrahost SAVs significantly captures genetic linkage patterns observed at the interhost level (Pearson’s r =0.28-0.45, all p<0.0001). Further, we show that machine learning models can learn highly generalisable intrahost, physiochemical and phenotypic patterns to forecast the future fitness of intrahost SAVs ( r 2 =0.48-0.63). Most of these models performed significantly better when considering genetic linkage ( r 2 =0.53-0.68). Overall, our results document the evolutionary forces shaping the fitness of mutations, which may offer potential to forecast the emergence of future variants and ultimately inform the design of vaccine targets. Biological sciences/Evolution/Molecular evolution Health sciences/Diseases/Infectious diseases/Viral infection Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Since its emergence in late 2019 1 , SARS-CoV-2, the agent of the COVID-19 pandemic, spread globally and caused over seven million deaths worldwide as of 21 July 2024 ( https://ourworldindata.org/grapher/cumulative-covid-deaths-region ). On 30th January 2020, the World Health Organisation (WHO) declared the pandemic a Public Health Emergency of International Concern, which was then declared over on 5th May 2023 following falling death rates, the deployment of effective vaccines, and increased immunity in the global population ( https://www.who.int/emergencies/diseases/novel-coronavirus-2019/interactive-timeline ). Throughout the COVID-19 pandemic, SARS-CoV-2 has exhibited a repeated pattern of lineage replacement, where multiple waves of genetically distinct lineages, termed ‘variants’, emerged and replaced the previous dominant variant in circulation 2 . These variants, including Delta, Omicron, and BA.2.86 (an Omicron sub-lineage sometimes referred to as “Pirola”), carry constellations of co-occurring mutations, a phenomenon known as genetic linkage. Some of these mutations confer a fitness advantage to the virus such as increased transmissibility, replication or immune evasion 3 . Additionally, some linked mutations have been shown to act synergistically 4 , 5 , conferring greater fitness advantages than the sum of their effects in isolation. This process, termed synergistic epistasis, may explain some of the patterns of genetic linkage observed. The fitness differentials provided by these clusters of co-adapted mutations allow viruses carrying them to outcompete the preceding variant in circulation. SARS-CoV-2 accumulates about 15 mutations per year with an overall evolutionary rate of ~ 1x10 − 3 substitutions per site per year. The accrual of genetic diversity at the genome level is driven by multiple evolutionary forces at different levels. Within a single infection, genome replication errors and the action of host immune RNA editing may generate de novo mutations in the viral population that are subjected to within-host selective pressures. Together, these processes are estimated to result in the acquisition of 1.3x10 − 6 mutations per nucleotide per replication cycle 6 . The intrahost viral population expands and reaches its peak population size around 2–5 days after infection (coinciding with symptom onset) 7 . The viral population then undergoes a narrow transmission bottleneck where only a small fraction of the intrahost viral population is passed on to the next individual in the transmission chain, with estimates as low as one to eight virions per transmission event within households 8 . Selective forces may play an important role during transmission, favouring alleles that improve transmission and gradually purging those that do not. As a result of these processes, the genetic diversity observed within hosts is expected to be substantially higher than that observed at the interhost level. However, since SARS-CoV-2 typically causes acute infections that last about 10–15 days post symptom onset 7 , with samples often collected shortly after symptoms manifest, only a few intrahost mutations have been observed during this short window 8 , 9 . The relative contributions of these evolutionary forces - acting both at the intrahost level and during transmission - on the mutations observed in SARS-CoV-2 lineages remain unclear. Given the short evolutionary timescale and the narrow population bottleneck during transmission, the effects of genetic drift may interfere with the action of natural selection. Neutral or weakly deleterious alleles may be transmitted before they are purged, allowing them to be observed at considerable frequencies amongst sampled genomes, whereas there may be a time lag before weakly advantageous mutations reach high frequencies. As such, we hypothesised that by analysing intrahost allele frequencies, we may be able to predict the relative success of mutations arising in the evolution of SARS-CoV-2. In addition, we hypothesised that the linkage patterns observed at the genomic level may be recapitulated by those found within intrahost viral populations. Addressing these questions may provide insights into the trajectory of mutations in SARS-CoV-2. SARS-CoV-2 is arguably one of the best-studied viruses, owing to concerted, global and multidisciplinary research efforts over the COVID-19 pandemic. One important outcome is the remarkable amount of genomic data that has been generated. As of July 2024, nearly 15 million complete genomes and over seven million sequencing libraries of SARS-CoV-2 have been deposited on the data repositories, GISAID 10 , 11 and NCBI Sequence Read Archive (SRA) 12 , respectively. These datasets and their accompanying metadata, including sample collection dates, sequence quality, lineage designations and viral mutations, provide an unprecedented ‘Big Data’ resource for testing this hypothesis and addressing other important questions relating to viral evolution. In this study, we leverage large-scale SARS-CoV-2 genomic datasets to investigate the relationships between the genetic diversity observed within hosts and the relative fitness of SARS-CoV-2 mutations. We first provide a detailed overview of the landscape of mutations observed in SARS-CoV-2, exploring broad patterns in the physiochemical and phenotypic effects of all non-synonymous mutations observed over the COVID-19 pandemic. Using a curated dataset of 7,862 quality-controlled sequencing libraries generated from SARS-CoV-2 infections, we then test the extent to which intrahost diversity, together with other genetic and phenotypic traits, can predict the relative success of SARS-CoV-2 mutations. These insights are invaluable for understanding the evolutionary forces shaping SARS-CoV-2, and may offer scope to inform the selection of vaccine targets as well as responding to other current and future viral threats. Results The landscape of mutations in SARS-CoV-2 To explore the patterns of mutations that have arisen over the COVID-19 pandemic, we analysed the metadata provided by GISAID 10 , 11 for the first submitted SARS-CoV-2 sequence up to the 8th July 2024, comprising nearly 15 million complete SARS-CoV-2 genomes. Within this dataset, we identified a total of 101,484 single amino acid variants (SAVs). This represents 52% of the 194,660 SAVs that could theoretically be observed considering that SARS-CoV-2 coding regions encode 9,733 amino acids, excluding start and stop codons. The number of SAVs observed over time increased rapidly in the initial stages of the pandemic though rapidly tapered off, with half of all theoretical SAVs already observed by December 2022 (Fig. 1 a). Of all the SAVs observed in our dataset, 101,263 (99.8%) never exceeded a monthly global frequency of 10%, 100 (0.1%) reached frequencies between 10–90%, and 121 (0.1%) reached greater than 90% frequency (i.e., fixation). There were no obvious protein-specific patterns in the number of mutations observed for the vast majority of SAVs that failed to reached fixation ( Extended Data Fig. 1 ). In contrast, SAVs that did were largely concentrated in the spike protein ( Extended Data Fig. 1 ), consistent with its high antigenicity and crucial role in cellular entry. The evolution of SARS-CoV-2 is largely characterised by the repeated emergence of divergent lineages that replace the preceding dominant lineage in circulation (Fig. 1 b). Many mutations that reached fixation at some point over the course of the pandemic exhibited frequency patterns largely resembling those of these dominant lineages (Fig. 1 b). However, there were some exceptions. For instance, Spike_D614G and NSP12_P323L rose in frequency during the initial stages of the pandemic, reaching fixation around mid-2020 and subsequently remained fixed in the SARS-CoV-2 population; Spike_G142D gradually rose in frequency during the Delta wave and remained fixed while NSP4_L438F rose in frequency slightly after the initial Omicron lineages had replaced the preceding Delta variant and quickly dropped in frequency before reaching fixation again coinciding with the emergence of the XBB/EG.5 Omicron sub-lineages (Fig. 1 b). We next explored whether the phenotypic impacts of SAVs observed were potential drivers of their relative fitness, which we estimated as the number of genomes submitted to GISAID carrying each SAV. We first calculated BLOSUM62 scores for all SAVs using the BLOSUM62 matrix, which evaluates amino acid changes based on their likelihood of occurring by chance, as derived from empirical protein alignments 13 . Under this scoring system, non-conservative mutations ( i.e. , rare mutations) are assigned more negative scores 13 , 14 . While this scoring system is typically used for scoring amino acid sequence alignments, it is thought to capture the effects of differing selective forces on amino acid substitutions, where stronger purifying selection acts on less conservative mutations 14 , 15 . Since ‘conservativeness’ is a statistical abstraction of the phenotypic effects of mutation and the consequences of natural selection, we further considered other metrics that capture these effects directly. In particular, we computed the predicted effects of each SAV on the physiochemical properties of proteins (i.e., change in charge, molecular weight and hydropathy). We also investigated the effects of SAVs found within the receptor-binding domain (RBD) on protein expression, host-receptor binding and antibody escape by leveraging direct measurements from previously published Deep Mutational Scanning (DMS) experiments 4 , 16 , 17 (see Methods ). Across SAVs, the score distributions for the three physiochemical properties (i.e., change in charge, molecular weight and hydropathy) appear symmetrically centred around zero (Fig. 1 c), suggesting that most observed SAVs tend to preserve the physiochemical properties of the proteins they are found in. Separately, based on the DMS measurements, 83% and 70% of SAVs in the RBD were detrimental for host-receptor binding and protein expression, respectively, and most conferred minimal antibody escape (Fig. 1 d). These results suggest that most SAVs observed are neutral or deleterious for these phenotypes. Notably, we found a significant positive correlation between fitness and the BLOSUM62 score of all SAVs (Spearman’s ρ = 0.33, p < 0.0001; Fig. 1 e), consistent with the idea that non-conservative mutations are generally deleterious. Further, SAVs associated with larger absolute changes in physiochemical phenotypes had significantly lower fitness values ( ρ =-0.13,-0.040,-0.23, p < 0.0001 for molecular weight, charge and hydropathy, respectively). This relationship, however, does not hold for the small subset of SAVs that have reached fixation over the course of the pandemic ( ρ =-0.068, p = 0.462; Fig. 1 e), which are likely to confer the highest fitness gains. These findings suggest that while non-conservative mutations are generally deleterious, they are also more likely to confer the highest fitness gains to the virus. This is further corroborated by the relationships observed between the physiochemical effects of SAVs and fitness. Indeed, the absolute changes in physiochemical scores were significantly and negatively correlated with fitness ( ρ =-0.040,-0.13,-0.23, respectively; all p < 0.0001). Meanwhile, SAVs that reached fixation were associated with significantly higher absolute changes in charge compared to non-fixed SAVs (two-sided Mann-Whitney U test, p = 0.0076). These results highlight that while more significant biochemical alterations are often deleterious, they can occasionally lead to the highest fitness gains. Finally, SAVs that are associated with increased host-receptor binding and RBD expression had significantly higher fitness values (both ρ = 0.26, p < 0.0001), which was expected given the key role of the RBD in mediating viral entry into host cells. Meanwhile, antibody escape was weakly and negatively correlated with fitness (ρ=-0.084, p = 0.0005). While SAVs that improve antibody escape would likely be beneficial for the virus, and therefore be inherently fitter, this negative correlation may point to a trade-off between the different phenotypes. Indeed, antibody escape was significantly negatively correlated with both binding (ρ=-0.36, p < 0.0001) and expression (ρ=-0.11, p < 0.0001), indicating that, for many SAVs, the fitness advantage due to increased antibody escape may be counteracted by decreased binding and/or expression, thus resulting in a lower overall fitness. Intrahost dynamics predicts the future fitness of mutations We next investigated whether the intrahost diversity of SARS-CoV-2 infections could be used to forecast the relative fitness of SAVs at the population level. To do this, we curated and downloaded seven sequencing datasets (total n = 7862) comprising sequencing libraries of SARS-CoV-2 infections, which were sampled at various timepoints across the pandemic. These datasets correspond to seven time periods of sampling, henceforth timeframes, termed ‘Early’, ‘Alpha’, ‘Delta’, ‘BA.1, ‘BA.5’, ‘XBB’ and ‘BA.2.86’ (Fig. 1 b). The Early timeframe corresponds to the first four months of the pandemic, while the others span the transition points between the major variant waves in our dataset. We then estimated the intrahost frequencies of SAVs in these libraries, retaining only SNPs/SAVs that were supported by at least 100 de-duplicated reads and that were observed at > 3% frequency. We selected these thresholds as they have been previously applied to reduce noise arising from sequencing errors 8 . Additionally, as some libraries may be associated with serial passaging experiments, laboratory contamination, an excess of sequencing artefacts, or erroneous collection dates, we further omitted libraries whose consensus genomes were outliers when assessing the number of single nucleotide polymorphisms relative to the Wuhan-Hu-1 reference (GenBank: MN908947.3) compared to other consensus genomes that were assigned the same lineage ( Extended Data Fig. 2 ; See Methods ). Our final quality-controlled dataset comprised 5885 sequencing libraries, representing seven snapshots of the lineage composition at various points during the COVID-19 pandemic ( Extended Data Fig. 3 ). Within each of the seven datasets, we identified between 2872–9606 intrahost SAVs across the sequencing libraries ( Supplementary Table 1 ), with an increase in the median number of SAVs detected per library by the chronological order of the datasets (Early = 25; BA.2.86 = 128). Within each timeframe, 9-112 (0.1–3.9%) of these intrahost SAVs were already at > 10% monthly frequency at the interhost level. Excluding these SAVs, we found that 4–73 (0.15–1.1%) SAVs reached at least a monthly frequency of > 10% after the timeframe of each dataset (Fig. 2 a). This pattern held for all except the BA.2.86 timeframe, where we observed a substantially (1.8–5.1 fold) higher proportion of high frequency SAVs than the overall expectation of 221/101,795 (0.22%) SAVs when considering the entirety of the GISAID data (Fig. 2 a). Additionally, the intrahost SAVs that were observed at > 10% frequency after the dataset timeframes only reached peak frequency after a median of 6–40 months. In other words, many of these relatively fitter SAVs could already be observed in the intrahost diversity of infections that were sampled much earlier in the pandemic. To explore whether we could forecast the relative fitness of SAVs observed at the intrahost level, we estimated the fitness of each SAV as the number of submitted consensus genomes a SAV was observed in after the timeframe of each dataset (henceforth, ‘future fitness’). We additionally generated 13 predictors of future fitness for each mutation, namely: the number of libraries each SAV was detected in, median and maximum intrahost frequency, BLOSUM62 score, the estimated raw and absolute change in charge, hydropathy or molecular weight, and for RBD mutations, the change in binding and expression, and antibody escape fraction, as directly measured via DMS 4 , 16 , 17 . Most predictors were significantly correlated with the future fitness of SAVs (Spearman’s correlation test, all adjusted p < 0.05; Fig. 2 b) and the direction of the correlations were largely consistent across datasets. Maximum and median intrahost frequency, BLOSUM62 score, hydrophobicity and molecular weight were positively correlated with future fitness, whereas absolute changes to charge, molecular weight and hydropathy were negatively correlated with future fitness (Fig. 2 b). For each dataset, we trained supervised XGBoost 18 regression models – that can capture complex non-linear relationships – to predict the future fitness of SAVs based on predictors generated from the intrahost SAVs detected. We optimised the models and assessed their performance using a cross-validation procedure (see Methods ), where different subsets of the data are withheld to test whether the models can make robust predictions on unseen data. We found that all models were able to predict the future fitness of intrahost SAVs, yielding mean r 2 values between 48–63% for each dataset. Additionally, the mean absolute error of the predictions by each model was 0.45–0.68, indicating that, on average, the models were able to predict the future fitness of intrahost SAVs to within the same order of magnitude. These findings indicate that our models could reliably forecast the success of SAVs based on the intrahost diversity of SARS-CoV-2 infections. To understand how these complex models were making their predictions, we further used the model interpretation framework, SHAP 19 , to visualise the relationships between the predictors and mutation fitness. The core of this framework is the SHAP value, which here represents, for each SAV, the estimated contribution of a single predictor to the predicted fitness value for that SAV. Across the datasets, the median absolute SHAP value for maximum intrahost frequency was the highest (median = 0.18) indicating that this predictor was the most important contributor to the model predictions (Fig. 2 c). Meanwhile, RBD binding, expression, and mean antibody escape were the least important, though noting that these directly measured traits are only available for SAVs in the RBD. We additionally explored the relationships between the predictors and the future fitness of SAVs by comparing the predictors with their corresponding SHAP values (Fig. 2 d-g). Maximum intrahost frequency had a monotonic positive relationship with predicted future fitness (Fig. 2 d), indicating that the selective forces governing the intrahost frequencies of SAVs explain, in part, mutational success at the global level. Separately, there was an apparent parabolic relationship between change in hydropathy and fitness, where SAVs involving relatively larger increases or decreases to hydrophobicity are generally associated with lower mutational fitness (Fig. 2 e), concordant with the idea that significant physiochemical changes are generally negatively associated with fitness. Notably, some SAVs associated with moderate absolute hydropathy or molecular weight changes had the most positive SHAP values ( Extended Data Fig. 4 ), suggesting that whilst most physiochemical changes are deleterious, they can also lead to positive fitness gains. Similarly, BLOSUM62 score was positively correlated with predicted fitness, with some negative scores associated with positive SHAP values (Fig. 2 f). These findings recapitulate the previously observed pattern that larger physiochemical alterations are typically associated with lower fitness, but can, in some cases, also be linked to higher fitness. Finally, RBD binding (Fig. 2 g; Extended Data Fig. 4 ) and expression ( Extended Data Fig. 4 ) were positively correlated with higher predicted fitness values. These findings indicate that, as expected, SAVs associated with increases in binding and expression are fitter. All patterns were largely concordant across the different timeframes considered ( Extended Data Fig. 4 ), highlighting that our models learned biologically relevant patterns to predict the future fitness of mutations. To further test the generalisability of our models, we trained an additional XGBoost regression model on the Alpha dataset, comprising libraries collected in February 2021, to predict the future fitness of intrahost SAVs up to before the first Omicron wave (i.e., BA.1; November 2021). We then tested the trained model on the BA.1 dataset (libraries collected in December 2021), to predict the fitness of intrahost SAVs from BA.1 to the end of our dataset (July 2024). Overall, the model trained on the Alpha dataset and pre-Omicron fitness data was able to reliably predict the future fitness of Omicron SAVs with the BA.1 dataset (Spearman’s ρ = 0.72; r 2 = 0.52; Fig. 2 h), indicating that this model is indeed generalisable. Notably, while some of the intrahost SAVs that reach a future frequency > 10% are predicted to have the highest future fitness, many of the high frequency SAVs had markedly lower predicted fitness values than observed. Indeed, the studentised residuals for high frequency SAVs – in a linear model of predicted versus observed fitness – were considerably more negative than those for low frequency SAVs (median=-1.66 and − 0.066, respectively; two-side Mann-Whitney U test, U = 372,623, p < 0.0001). This prediction error may be a result of other evolutionary processes not captured by the model including epistatic interactions between high frequency mutations. Intrahost evolution shapes global linkage patterns We next explored whether genetic linkage at either the intrahost or interhost level could explain why some of these SAVs are much fitter than predicted by our initial models. To estimate intrahost linkage, we calculated Pearson’s r using the intrahost SAV frequencies of libraries within each dataset. At the interhost level, we calculated the D’ statistic 20 using the SAV frequencies estimated from the GISAID consensus genomes collected prior to the sampling timeframe of each dataset. To reduce noise arising from low frequency SAVs, we considered only SAVs that were observed in more than 1000 genomes. For biallelic loci (i.e., only two alleles or amino acid variants are observed at each locus), linkage is often reported as |D’| or r 2 since negative values indicate linkage between the alternate allele at one locus and the wild type at the other. In our context, there are 20 possible SAVs at each locus so negative D’ or r values are not as meaningful and can only be interpreted as ‘not positively linked’. In total, we calculated D’ for 143,889-7,714,482 and r for 93,528-1,326,006 SAV pairs for each dataset, of which 0.1–0.9% and 0.01–0.6% are strongly linked ( D’ or r > 0.9), respectively. Intrahost linkage (i.e., r ) was significantly correlated with that observed at the interhost level (i.e., D’ ) (Pearson’s r = 0.28–0.45, all adjusted p < 0.0001; Fig. 3 a; Extended Data Fig. 5 ), indicating that the evolutionary processes shaping linkage patterns at the interhost level are also seen within hosts. At both the interhost and intrahost level, within-protein linkage associations were significantly more common in strongly and positively linked ( D’ or r > 0.9) SAVs than those with weaker linkage associations ( D’ or r ≤ 0.9) (Fig. 3 b; odds ratio = 1.7–42, Fisher’s exact test, all adjusted p < 0.0001). This suggests that our linkage statistics likely capture both the signatures of recombination (i.e., linkage decay) and epistasis (i.e., short-range interactions are more common). The SAV pairs that appear strongly linked at both the intrahost and interhost level included SAVs that were highly successful, including N_R203K-N_G204R ( r = 0.98 and D’=1.00 for the Alpha timeframe; Fig. 3 c). The close proximity of these two SAVs and the near-perfect linkage patterns indicate likely synergistic interactions. Separately, Spike_D614G-NSP12_P323L, one of the first few pairs of SAVs whose monthly frequencies co-rose rapidly in the early stages of the pandemic (Fig. 1 b), had a slightly weaker intrahost correlation (‘Early’ timeframe, intrahost r = 0.84; Fig. 3 c). Their occurrence on different proteins and the weaker linkage pattern may point to co-selection of two highly beneficial mutations 21 , 22 in a largely wild-type background, rather than epistasis. Notably, 73–98% of strongly linked pairs involved only SAVs that never exceeded an intrahost frequency of 50% nor were observed in more than 1000 consensus genomes across the 56 months for all timeframes except the BA.2.86 timeframe (Fig. 3 d). The low consensus prevalence and intrahost frequencies of these SAVs jointly indicate that the majority of linked SAVs do not confer a significant fitness advantage and are a product of genetic drift. Additionally, if most mutations are deleterious, then these linkage patterns may represent compensatory epistasis. Overall, these findings indicate that linkage patterns observed within hosts may provide clues on the selective forces shaping those observed at the interhost level. Genetic linkage improves fitness predictions To incorporate genetic linkage at both the intrahost and interhost levels into our models, we generated six further predictors, in addition to the thirteen used previously, namely, whether the SAV is strongly linked to at least one other SAV ( D’ or r > 0.9), the number of strongly linked SAVs ( D’ or r > 0.9), and the maximum D’ or r across all pairs involving the SAV. In addition, D’ statistics were only calculated for SAVs observed in > 1000 consensus genomes prior to the timeframe of each dataset so, given that prior fitness is highly autocorrelated with future fitness (Spearman’s ρ = 0.88–0.92), models may ‘cheat’ by leveraging this autocorrelation to improve performance without learning any additional biological patterns. As such, we included an additional binary predictor, indicating whether a SAV was observed in > 1000 genomes (henceforth ‘prior fitness’), to partition this effect in our models. Inclusion of linkage predictors significantly improved the performance of all models (Mann-Whitney U test, maximum p = 0.018), except for that trained on the Alpha timeframe (p = 0.22). The most prominent improvement was seen for the model trained on the BA.2.86 timeframe (mean r 2 = 0.49 to r 2 = 0.67) (Fig. 3 a). All models trained on one timeframe and tested on another (as for Fig. 2 h) also performed better with the inclusion of linkage predictors ( Extended Data Fig. 6 ). This indicates that the relationships between linkage and future fitness learned by the regression models were generalisable across datasets. Since SHAP values are additive 19 , we compared the sum of the absolute SHAP values for the different types of predictors (i.e., physiochemical, intrahost, intrahost linkage, interhost linkage and prior fitness) to assess their relative importance. Interhost linkage predictors were significantly more important than intrahost linkage predictors (paired t -test, t = 16.0-91.2, d.f.=2761–9560, all p < 0.0001) and the prior fitness predictor for predicting future fitness, across all timeframes ( t = 23.8–95.1, d.f.=2761–9560, all p < 0.0001) (Fig. 3 b). These findings indicate that linkage patterns at the interhost level are largely responsible for the performance boost. Since the fittest SAVs were associated with significantly higher prediction errors than the less successful SAVs in our initial models (Fig. 2 h), we hypothesised that genetic linkage may explain, at least in part, this pattern. To test this, we compared the changes in prediction error for our models with or without linkage predictors for SAVs where D’ could be calculated. Following the inclusion of linkage predictors, the changes in prediction error were more negative for SAVs that reached a future frequency of > 10% compared to those that did not (Fig. 3 c). This effect was statistically significant for the Alpha, Delta and BA.1 timeframes (Mann-Whitney U test, all p < 0.001), but not for the XBB and BA.2.86 timeframes (p = 0.804 and 0.063, respectively). This could be due to the smaller number of high frequency SAVs considered in these timeframes. Additionally, we aggregated the contributions of all linkage predictors by summing their SHAP values and found that the linkage predictors resulted in significantly higher predicted fitness values for SAVs that reached a high future frequency (all p < 0.05; Fig. 3 d). Furthermore, the summed SHAP values of linkage predictors were all negatively correlated with the changes in prediction error across all datasets (Spearman’s ρ =-0.686 to -0.029). This implies that by using information about genetic linkage, these models were predicting higher future fitness values for the high frequency SAVs, thus resulting in lower prediction errors. Overall, these findings suggest that the success of high frequency SAVs may, in part, be governed by genetic linkage. Discussion Predicting which mutations will emerge and reach high frequency in the future is a long-standing and important, yet largely unsolved problem in the study of pathogen dynamics. In this study, we investigated whether the intrahost diversity of SARS-CoV-2 infections could provide clues on the relative fitness of SAVs by leveraging the unparalleled amount of sequencing data available. Our key assumption is that the evolutionary forces shaping the interhost frequencies of mutations in the viral population are also acting within hosts. Using various machine learning techniques, we showed that this assumption holds, that is, that the intrahost dynamics of SARS-CoV-2 infections can predict the future fitness of mutations remarkably well. In addition, our model interpretation analyses revealed highly generalisable patterns of intrahost dynamics and phenotypic traits that correlate with the success of mutations. Our findings lay the groundwork for developing better predictive tools that may be invaluable for pre-empting the emergence of novel lineages and supporting the design of vaccine targets. By analysing the distribution of SAVs across all complete genomes on GISAID (as of July 2024), we found that more than half of all possible SAVs have already emerged in the evolution of SARS-CoV-2, and that the rate of emergence of previously unseen mutations has slowed ( Fig. 1a ). Most SAVs observed thus far tend to preserve the physiochemical properties of the proteins they are found in, suggesting that major phenotypic alterations to proteins are generally deleterious ( Fig. 1c ). This is corroborated by the observation that most SAVs in the RBD are estimated to be phenotypically neutral or deleterious ( Fig. 1d ). Further, only an extremely small subset of SAVs have reached a high frequency at the interhost level thus far. Collectively, these findings point to strong functional and evolutionary constraints on the mutational space that SARS-CoV-2 can explore. Since the rate of emergence of mutations has slowed considerably, one can speculate that future variant lineages may be more likely to harbour novel combinations of moderately or highly fit mutations that have already been observed so far rather than a set of previously unseen mutations. This, however, does not preclude the possibility of unseen mutations emerging in the future. In this study, we used the number of GISAID sequences carrying each SAV across various time windows as a proxy for estimating SAV fitness. While it may be possible to get more direct fitness estimates through animal challenge experiments 23–25 , it is not tractable to do so at the scale of SAVs considered in this study. Separately, our fitness estimates are likely affected by the spatiotemporal biases in genomic surveillance efforts. For example, the estimated fitness of a moderately successful SAV that circulated only in countries with lower surveillance would be lower than a less successful SAV that spread in countries with more extensive surveillance. However, these discrepancies are likely random and manifest as noise, rather than a systematic bias, in the fitness estimates. As such, these sampling biases are not likely to skew the patterns observed in our study. Our modelling approach differs from previous efforts to understand viral fitness, which have largely focused on the relative fitness of variant lineages 26–31 , whereas we consider individual mutations as targets of selection, each associated with its own fitness values. The theoretical number of combinations of SAVs that SARS-CoV-2 can have is practically infinite (i.e., 20 9733 ), which entails that the fitness of previously unseen lineages is difficult to predict, especially for models based on epidemiological data 27,29 . Further, the fitness landscape of lineages is likely confounded by the immunological history of the population and the extant competing lineages in circulation. Focusing on the fitness of mutations is lineage agnostic, circumventing these limitations. Additionally, our models can learn the inherent genetic and phenotypic traits of mutations that underly their success, enabling richer biological interpretations. For example, our models reveal that the SAVs associated with more aggressive physiochemical changes are generally under stronger purifying selection, and expectedly, that the SAVs that negatively impact RBD binding and expression are less fit. Importantly, our modelling approach leverages information about the intrahost dynamics of mutations, viewing each infection as an independent playing field where mutations compete for dominance, with fitter mutations being observed more frequently than less fit ones. Our observation that maximum intrahost frequency is consistently the most important predictor of future fitness supports this, and suggests that within-host evolutionary processes are important components that shape the fitness landscape of mutations at the interhost level. Such processes may include positive selection for mutations that improve viral infectivity and/or replication rate, and purifying selection acting on highly deleterious mutations that drastically alter protein function. Notably, our models performed relatively poorly on making predictions for the fittest mutations ( Fig. 2h ). Incorporating genetic linkage into the models partially reduced this effect, indicating that the co-occurrence of SAVs, which may or may not involve epistasis, plays a role in modulating the fitness of mutations in the evolution of SARS-CoV-2. However, there remains considerable variation in the fitness of SAVs that is not captured by our models, suggesting that some components of mutational fitness are still unaccounted for. Incorporation of further functional, structural or immunological predictors may reduce this gap further, but this may be challenging since existing experimental data for SARS-CoV-2 remain biased towards the spike protein. Additionally, selective forces may act at the point of transmission, with mutations improving transmissibility being more successful, but this is not captured by our models. Alternatively, these discrepancies may indicate that some of these highly successful mutations are not inherently the fittest, but represent a random subset of moderately fit mutations that rose to a high frequency due to genetic hitchhiking 32 or drift. Overall, our results highlight key evolutionary forces governing the success of mutations in the evolution of SARS-CoV-2, and show that the relative fitness of mutations is largely predictable when considering the intrahost dynamics of mutations and their genetic and phenotypic properties. However, the higher prediction error for the most successful mutations indicates that forecasting the exact set of mutations that will reach a high frequency remains out of reach, for now. Nevertheless, the generalisable patterns inferred by our models enables us to narrow down the list of mutations that are likely to be successful in the evolution of SARS-CoV-2. These could be combined with insights gleaned from other models of mutation or lineage fitness, and from immunological assays, to potentially inform on candidate mutations for future iterations of vaccines. Furthermore, we foresee that this approach could be easily ported to other comparable model systems for which large-scale genomic data is available, such as H3N2 Influenza A, which causes seasonal outbreaks globally. There may also be potential to leverage information derived from intrahost genetic variation in very different pathogens, for example to predict the future emergence of antimicrobial resistance in bacteria exposed to drugs. Methods Data acquisition We compiled a comprehensive dataset of SARS-CoV-2 genome assemblies. To do so, the metadata, including the list of SAVs, for all SARS-CoV-2 consensus genomes (n=16,798,863) deposited on GISAID 10,11 was downloaded on 8 th July 2024 in the TSV file format. We retained only genome entries were marked as complete (i.e., >29,000nt), were not low coverage (i.e., >5% Ns), were not isolated from the host genera Manis or Rhinolophus , were not associated with cell cultures, and had complete collection dates (i.e., year, month and day). For collection months with more than 200,000 genomes deposited, we randomly subsampled to 200,000 entries. The metadata of all BioSamples hosted on NCBI were downloaded on 21 st March 2024 from the FTP site (https://ftp.ncbi.nlm.nih.gov/biosample/) in XML file format. This was converted to TSV format with a custom Python script that uses the ElementTree XML API v1.3.0. All BioSamples whose taxonomy name, BioSample title, isolate or strain fields contained the terms ‘ Severe acute respiratory syndrome coronavirus 2 ’, ‘ cov-2 ’ or ‘ cov2 ’ were retained. BioSamples were linked to their associated sequencing libraries where possible using NCBI’s entrez-direct command-line API v21.6. Sequencing libraries were downloaded in the SRA file format and converted to FASTQ format using the prefetch and fastq-dump commands, respectively, as part of NCBI’s sra-tools API v3.11. For the Early dataset, we considered only BioSamples whose collection year and month were known and between 1 st October 2019 and 1 st March 2020. For the rest of the datasets, we considered only BioSamples with complete collection dates that corresponded to the following collection months: Alpha (February 2021), Delta (June 2021), BA.1 (December 2021), BA.5 (June 2022), XBB (February 2023), BA.2.86 (December 2023) ( Fig. 1b ). Bioinformatic pipeline for sequencing data To ensure the homogeneity of sequencing data, we analysed only paired-end Illumina libraries. If BioSamples were associated to multiple sequencing libraries, the sequencing reads were combined prior to further processing. We removed known sequencing adapters, trimmed read ends (<Q20), and removed low quality read pairs (mean base quality of <Q20) using the BBDuk.sh utility of the BBMap v39.06 package (sourceforge.net/projects/bbmap/) 33 . We then aligned the quality-filtered read pairs to the CHM13 human reference genome (assembly accession: GCF_009914755.1) using Bowtie2 v2.5.1 34 , retaining read pairs where both members were unmapped. The unmapped read pairs were then re-aligned to the Wuhan-Hu-1 SARS-CoV-2 reference genome (GenBank accession: MN908947.3), retaining only read pairs where both members were mapped. Duplicate reads were removed using the markdup utility from Samtools v1.20 35 , to minimise the impact of PCR duplicates on our intrahost diversity analysis. SAV frequencies were estimated from the de-duplicated, high quality reads using the call codonvar utility in Quasitools v0.7.0 36 , specifying an sequencing error rate corresponding to Q20 (‘ --error_rate 0.01 ’). Separately, haploid variant calling was performed using the multi-allelic caller (‘ -m ’) in Bcftools v1.20 37 , specifying a minimum mapping quality and base quality of Q30 (‘ --min-MQ 30 ’ and ‘ --min-BQ 30 ’, respectively). Consensus genomes were then generated from the variant calls using a custom R script, masking positions where the coverage depth is less than 10, and ignoring indels. Sites that were masked in >10% of genomes were also replaced with Ns. Finally, the re-assembled consensus genomes were assigned PANGO lineages 38 and variant names using Pangolin v4.3 39 . Quality control of sequencing libraries The diversity of SAVs detected in the sequencing data could be confounded by whether libraries were associated with serial passaging experiments, laboratory contamination, an excess of sequencing artefacts, or erroneous collection dates. In all these cases, we would expect a larger or smaller than expected number of mutations relative to other libraries at each timepoint. For each dataset, we removed sequencing libraries whose consensus genomes had fewer than 1.5 times the lower quartile, or greater than 1.5 times the upper quartile, of the number of SNPs for that dataset. For the Early dataset, we additionally removed libraries that were classified as either Alpha or Beta variant, as these variant lineages are known to have emerged at later stages in the pandemic so were likely mislabelled. BioSamples in the BA.1 and XBB datasets with no lineage assignments were also excluded as they had a much lower number of SNPs compared to other lineages in circulation within the timeframes of these datasets ( Extended Data Fig. 2 ). We additionally removed BioSamples whose re-assembled consensus genomes had >10% Ns. The fitness of mutations The response variable for all models was the future fitness of a SAV after the sampling timeframe of a particular dataset. We define this as the number of genomes in the GISAID metadata carrying said SAV after the timeframe of the dataset with a one month buffer. Similarly, prior fitness is defined as the number of GISAD genomes carrying said SAV before the dataset timeframe with a one month buffer. For example, the Alpha dataset comprises samples collected in February 2021 and we estimated future and prior fitness as the number of GISAID genomes carrying each SAV between April 2021 and July 2024, and between December 2019 to December 2020, respectively. Physiochemical properties and DMS phenotypes For each SAV, we generated the physiochemical predictors using the Bio.SeqUtils.ProtParam module in BioPython v1.84. Changes and absolute changes in physiochemical properties were respectively, where P represents the physiochemical score of an amino acid. Amino acid charge was calculated at pH 7. Hydropathy was estimated using the Gravy scale 40 , for which more positive scores indicate greater hydrophobicity. BLOSUM62 scores were calculated using the Biostrings v2.70.2 package in R. Changes in RBD binding (Δlog 10 K D,app ) and expression (Δlog 10 MFI) were obtained directly from the raw data of previous DMS experiments (https://github.com/jbloomlab/SARS-CoV-2-RBD_DMS_variants/blob/main/results/final_variant_scores/final_variant_scores.csv) 4,17 , and represent the log-ratio of the apparent dissociation constant ( K D,app ) and mean fluorescence intensity (MFI) of the SARS-CoV-2 RBD with a SAV relative to the wildtype, respectively. Mean antibody escape was also calculated from previous experiments (https://github.com/jbloomlab/SARS-CoV-2-RBD_MAP_Crowe_antibodies/blob/master/results/supp_data/MAP_paper_antibodies_raw_data.csv) 16 . This score represents the escape fraction of the RBD with an SAV relative to the wildtype estimated under global epistasis models 16 , which we averaged across the 10 human monoclonal antibodies assayed. These predictors were assigned a value of -100 for SAVs that did not have any associated DMS estimates, which was the case for all non-RBD mutations. This predictor therefore also provides models with information on whether an SAV is found within the RBD or not. Genetic linkage Linkage at the intrahost level was calculated as the Pearson’s correlation coefficient, r , of the intrahost frequencies of two SAVs across all libraries within a dataset. Pearson’s r was only calculated for SAV pairs where there were at least five libraries with non-zero SAV frequencies. Linkage at the interhost level was estimated using the D’ statistic, as follows: where p A and p B , represent the proportion of consensus genomes carrying an SAV at locus A and locus B, respectively, and p AB represents the proportion carrying both SAVs at locus A and B. For these calculations, we only considered genome entries before the timeframe of each dataset, with a one month buffer (as with the calculation of prior fitness). For example, for the Alpha dataset (February 2021) we consider only genomes collected between December 2019 and December 2020. The D’ statistic was only calculated for SAVs with a prior fitness of >1000 genomes. Three types of linkage predictors were generated based on the intrahost and interhost linkage data for each SAV: whether a SAV is strongly linked to at least one other SAV ( D’ or r >0.9) is a binary variable; the number of strongly linked SAVs ( D’ or r >0.9) is an integer variable ranging from zero to n – 1, where n is the total number of SAVs considered in the dataset; maximum D’ or r across all pairs involving the SAV is a continuous variable ranging from zero to 1 (negative D’ or r values are set to zero). All linkage predictors that could not be calculated were set to -1 to indicate missingness. Model training, hyperparameter optimisation and evaluation For each dataset, we trained gradient-boosted regressors as part of the XGBoost v2.0.3 18 module in Python. We optimised the hyperparameters for our models and assessed their test error using a nested 10x10 cross-validation procedure. Briefly, for each iteration of the outer cross-validation loop, a tenth of the data is held out. The remaining nine-tenths of the data is used for the inner cross-validation loop, where an exhaustive search (i.e., grid search) for the best performing combination of hyperparameters is performed. As part of hyperparameter optimisation, 10-fold cross validation is performed to evaluate the performance of models using each combination of hyperparameters, and the best performing model is selected. The held-out tenth of the data in the outer loop is then used to assess the performance of the optimised model, and this process is then repeated for all 10 iterations of the outer loop. We optimised three hyperparameters for our models, particularly the number of trees in the ensemble (n_estimators=100, 300, 500, 700, 900), the maximum depth of each tree (max_depth=1, 2, 3, …, 9), and the proportion of random predictors used to grow each tree in the ensemble (colsample_bytree=0.1, 0.2, …, 1.0). Three main metrics were used for evaluating model performance, the coefficient of determination, r 2 , mean absolute error (MAE), and Spearman’s rank correlation coefficient, ρ . These metrics are defined by: Model interpretation To interpret our XGBoost regression models, we used SHAP v0.45.0 19 . Each predictor of a sample is assigned a SHAP value, which corresponds to the change in predicted future fitness given the information of that predictor. SHAP values therefore enable the decomposition of each future fitness prediction into the sum of contributions from each predictor. The relative importance of our predictors were assessed by their mean absolute SHAP values, with higher values indicating more important to the model predictions. The relative importance of groups of predictors (i.e., physiochemical, intrahost diversity, DMS, intrahost linkage and interhost linkage), were assessed by summing the absolute SHAP values for all predictors within a group and dividing this by the number of SAVs considered. Statistical analysis and visualisation All statistical analysis and visualisations were performed using the stats and ggplot packages in R v4.3.2. Where applicable, p -values where corrected for multiple testing using the Benjamini-Hochberg procedure. Declarations Acknowledgments C.C.S.T. is funded by the National Science Scholarship from the Agency for Science, Technology and Research (A*STAR), Singapore. M.E.Z, F.B. and L.v.D. are funded by the European Commission (Horizon 2021-2024, END-VOC Project). L.v.D. is additionally funded by the UKRI Future Leaders Fellowship (MR/X034828/1). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Health and Digital Executive Agency. For the purpose of open access, the corresponding author has applied a ‘Creative Commons Attribution’ (CC BY) licence to any Author Accepted Manuscript version arising. The authors acknowledge the use of the UCL Computer Science cluster and associated support services, in the completion of this work. Author contributions C.C.S.T and F.B. conceptualised and designed the study. C.C.S.T. performed all analyses with intellectual inputs from all co-authors. C.C.S.T wrote the manuscript with contributions and edits from all co-authors. Declaration of competing interest The authors declare that to current knowledge, there are no legal, financial or personal competing interests. Data and code availability All custom code used to perform the analyses reported here are hosted on GitHub (https://github.com/cednotsed/early_SC2_trajectory). All data used to train and evaluate the models are provided in Supplementary Table 1 . References Van Dorp, L. et al. Emergence of genomic diversity and recurrent mutations in SARS-CoV-2. Infection, Genetics and Evolution 83 , 104351 (2020). Balloux, F. et al. The past, current and future epidemiological dynamic of SARS-CoV-2. Oxford Open Immunology 3 , iqac003 (2022). Carabelli, A. M. et al. SARS-CoV-2 variant biology: immune escape, transmission and fitness. Nature Reviews Microbiology 21 , 162–177 (2023). Starr, T. N. et al. Shifting mutational constraints in the SARS-CoV-2 receptor-binding domain during viral evolution. Science 377 , 420–424 (2022). Witte, L. et al. Epistasis lowers the genetic barrier to SARS-CoV-2 neutralizing antibody escape. Nature Communications 14 , 302 (2023). Amicone, M. et al. Mutation rate of SARS-CoV-2 and emergence of mutators during experimental evolution. Evolution, medicine, and public health 10 , 142–155 (2022). Markov, P. V. et al. The evolution of SARS-CoV-2. Nature Reviews Microbiology 21 , 361–379 (2023). Lythgoe, K. A. et al. SARS-CoV-2 within-host diversity and transmission. Science 372 , eabg0821 (2021). Gu, H. et al. Within-host genetic diversity of SARS-CoV-2 lineages in unvaccinated and vaccinated individuals. Nature Communications 14 , 1793 (2023). Shu, Y. & McCauley, J. GISAID: Global initiative on sharing all influenza data–from vision to reality. Eurosurveillance 22 , 30494 (2017). Elbe, S. & Buckland‐Merrett, G. Data, disease and diplomacy: GISAID’s innovative contribution to global health. Global Challenges 1 , 33–46 (2017). Leinonen, R., Sugawara, H., Shumway, M., & International Nucleotide Sequence Database Collaboration. The sequence read archive. Nucleic acids research 39 , D19–D21 (2010). Henikoff, S. & Henikoff, J. G. Amino acid substitution matrices from protein blocks. Proceedings of the National Academy of Sciences 89 , 10915–10919 (1992). Eddy, S. R. Where did the BLOSUM62 alignment score matrix come from? Nature biotechnology 22 , 1035–1036 (2004). Cargill, M. et al. Characterization of single-nucleotide polymorphisms in coding regions of human genes. Nature Genetics 22 , 231–238 (1999). Greaney, A. J. et al. Complete mapping of mutations to the SARS-CoV-2 spike receptor-binding domain that escape antibody recognition. Cell host & microbe 29 , 44–57 (2021). Starr, T. N. et al. Deep mutational scanning of SARS-CoV-2 receptor binding domain reveals constraints on folding and ACE2 binding. cell 182 , 1295–1310 (2020). Chen, T. & Guestrin, C. Xgboost: A scalable tree boosting system. in 785–794 (2016). Lundberg, S. M. & Lee, S.-I. A unified approach to interpreting model predictions. Advances in neural information processing systems 30 , (2017). Lewontin, R. C. The interaction of selection and linkage. I. General considerations; heterotic models. Genetics 49 , 49 (1964). Plante, J. A. et al. Spike mutation D614G alters SARS-CoV-2 fitness. Nature 592 , 116–121 (2021). Goldswain, H. et al. The P323L substitution in the SARS-CoV-2 polymerase (NSP12) confers a selective advantage during infection. Genome biology 24 , 47 (2023). Ulrich, L. et al. Enhanced fitness of SARS-CoV-2 variant of concern Alpha but not Beta. Nature 602 , 307–313 (2022). Sun, X. et al. Enhanced fitness of SARS-CoV-2 B. 1.617. 2 Delta variant in ferrets. Virology 582 , 57–61 (2023). Wu, H. et al. Nucleocapsid mutations R203K/G204R increase the infectivity, fitness, and virulence of SARS-CoV-2. Cell host & microbe 29 , 1788–1801 (2021). Meijers, M., Ruchnewitz, D., Eberhardt, J., Łuksza, M. & Lässig, M. Population immunity predicts evolutionary trajectories of SARS-CoV-2. Cell 186 , 5151–5164 (2023). Obermeyer, F. et al. Analysis of 6.4 million SARS-CoV-2 genomes identifies mutations associated with fitness. Science 376 , 1327–1332 (2022). Pucci, F. & Rooman, M. Prediction and evolution of the molecular fitness of SARS-CoV-2 variants: introducing SpikePro. Viruses 13 , 935 (2021). Abousamra, E., Figgins, M. & Bedford, T. Fitness models provide accurate short-term forecasts of SARS-CoV-2 variant frequency. PLOS Computational Biology 20 , e1012443 (2024). Ito, J. et al. A Protein Language Model for Exploring Viral Fitness Landscapes. bioRxiv 2024–03 (2024). Dadonaite, B. et al. Spike deep mutational scanning helps predict success of SARS-CoV-2 clades. Nature 631 , 617–626 (2024). Smith, J. M. & Haigh, J. The hitch-hiking effect of a favourable gene. Genetics Research 23 , 23–35 (1974). Bushnell, B. BBMap: a fast, accurate, splice-aware aligner. (2014). Langmead, B. & Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. Nature methods 9 , 357–359 (2012). Li, H. et al. The sequence alignment/map format and SAMtools. bioinformatics 25 , 2078–2079 (2009). Marinier, E. et al. Quasitools: a collection of tools for viral quasispecies analysis. BioRxiv 733238 (2019). Danecek, P. et al. Twelve years of SAMtools and BCFtools. Gigascience 10 , giab008 (2021). Rambaut, A. et al. A dynamic nomenclature proposal for SARS-CoV-2 lineages to assist genomic epidemiology. Nature microbiology 5 , 1403–1407 (2020). O’Toole, Á. et al. Assignment of epidemiological lineages in an emerging pandemic using the pangolin tool. Virus evolution 7 , veab064 (2021). Kyte, J. & Doolittle, R. F. A simple method for displaying the hydropathic character of a protein. Journal of molecular biology 157 , 105–132 (1982). Additional Declarations There is NO Competing Interest. Supplementary Files earlyintrahostdiversitysupplementary161024.docx Extended Data Figures SupplementaryTable1.xlsx Supplementary Table 1 Cite Share Download PDF Status: Under Review Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5298116","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":372185552,"identity":"022c311e-b323-438f-bde5-3687472b9721","order_by":0,"name":"Cedric Tan","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA4ElEQVRIiWNgGAWjYHACA4YECIPxAZDg4SNaCw8DA7MBiGYjSgsDRAubBIhBUAt/e/O2Dw9q7Ozt2c8eq/yaYyfDxsD88NENPFokzhwrnpFwLDmxhycv7bbstmSgw9iMjXPwuUoix5ghseFAAg9DjtltyW3MQC08bNJ4tci/AWux5+F/Y1Ysua2eCC0SPGAtjD0SOWaMH7cdJqxF4kxaMQPYLzfeGEszbjvOw8ZMwC/87Yc3M/4Ahhh7f47hx5/bqu352ZsfPsanBQUw84BJYpWDAOMPUlSPglEwCkbBiAEAncc/lctsnMoAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-3536-8465","institution":"UCL Genetics Institute, University College London","correspondingAuthor":true,"prefix":"","firstName":"Cedric","middleName":"","lastName":"Tan","suffix":""},{"id":372185553,"identity":"85fd1029-8506-455e-a4af-41ecfb0ed8f7","order_by":1,"name":"Marina Escalera-Zamudio","email":"","orcid":"https://orcid.org/0000-0002-4773-2773","institution":"UCL","correspondingAuthor":false,"prefix":"","firstName":"Marina","middleName":"","lastName":"Escalera-Zamudio","suffix":""},{"id":372185554,"identity":"4f19c1a8-9667-4a7b-9755-90170aeec8df","order_by":2,"name":"Alexei Yavlinksy","email":"","orcid":"https://orcid.org/0000-0002-4100-4257","institution":"Institute of Health Informatics, University College London","correspondingAuthor":false,"prefix":"","firstName":"Alexei","middleName":"","lastName":"Yavlinksy","suffix":""},{"id":372185555,"identity":"d5d8b305-1bd5-43c9-abe1-178d4f3fcc7f","order_by":3,"name":"Lucy van Dorp","email":"","orcid":"https://orcid.org/0000-0002-6211-2310","institution":"University College London","correspondingAuthor":false,"prefix":"","firstName":"Lucy","middleName":"van","lastName":"Dorp","suffix":""},{"id":372185556,"identity":"763c410f-e7b0-46a0-bb85-32bbebf1788c","order_by":4,"name":"Francois Balloux","email":"","orcid":"https://orcid.org/0000-0003-1978-7715","institution":"University College London","correspondingAuthor":false,"prefix":"","firstName":"Francois","middleName":"","lastName":"Balloux","suffix":""}],"badges":[],"createdAt":"2024-10-20 11:25:08","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5298116/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5298116/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":69073630,"identity":"622bafc3-aaf4-435d-b0dd-e7de28254b4a","added_by":"auto","created_at":"2024-11-15 10:36:43","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":447628,"visible":true,"origin":"","legend":"\u003cp\u003eThe phenotypic landscape of mutations in SARS-CoV-2. (a) Number of distinct single amino acid variants (SAVs) observed at various time points over the COVID-19 pandemic, normalised by the number of SAVs that could theoretically be observed (n=194,660). (b) Monthly frequencies of different lineages of interest (top) and a representative set of SAVs (bottom). Distributions of the (c) physiochemical effects and (d) phenotypic effects, as measured by deep mutational scanning experiments\u003csup\u003e16,17\u003c/sup\u003e, of all SAVs observed in our dataset. (e) Boxplot of the log10-transformed genome count as a proxy for fitness stratified by BLOSUM62 score, for all SAVs that reached fixation (\u0026gt;90% monthly frequency) and those that did not.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/757c691635c1ad3bc62ed19d.png"},{"id":69073638,"identity":"50583422-8f9c-499b-9527-b3277e2690d6","added_by":"auto","created_at":"2024-11-15 10:36:49","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":352109,"visible":true,"origin":"","legend":"\u003cp\u003eIntrahost diversity predicts the success of mutations. (a) Number of SAVs identified within the intrahost diversity of SARS-CoV-2 infections that reach a high frequency at the interhost level after the sampling timeframe of each dataset. (b) Spearman’s correlation values of the different predictors with the fitness of SAVs estimated after the sampling window of each dataset. The statistical significance of each correlation estimate was determined using a two-sided asymptotic t-test and corrected for multiple testing using the Benjamini-Hochberg procedure. (c) Relative importance of predictors for all XGBoost regression models trained, as assessed using SHAP values, where higher absolute SHAP values indicate higher contributions to model predictions. Scatter plots showing the relationships between future mutational fitness predicted by our XGBoost regression models and the (d) maximum intrahost frequency, (e) change in hydropathy, (f) BLOSUM62 score, and (g) change in RBD binding for each SAVs. The data shown in (d)-(g) were generated using the Alpha dataset, but are representative of the relationships observed in the other datasets considered. Blue lines represent LOESS (Locally Estimated Scatterplot Smoothing) regression smooths for each scatter plot. (h) Density plot showing the high correlation between the predicted future fitness of mutations and the actual observed fitness values for the BA.1 dataset. Predictions were made using an XGBoost regression model trained on the Alpha dataset. SAVs that reached a high future frequency (i.e., \u0026gt;10%) are shown as discrete points. The blue line indicates a linear regression smooth of the predicted and observed SAV fitness values.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/e67d0af4267d0552930fa0cd.png"},{"id":69073633,"identity":"40e72210-693f-472a-ab02-b0241cf93df4","added_by":"auto","created_at":"2024-11-15 10:36:44","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":176855,"visible":true,"origin":"","legend":"\u003cp\u003eIntrahost evolution shapes global linkage patterns. (a) Proportion of SAV pairs involving SAVs that are within the same protein at the (top) interhost or (bottom) intrahost level. (b) Strong positive relationship between the linkage of SAVs at the intrahost level and the interhost level estimated using the Delta dataset. The solid red line represents the linear regression smooth. (c) Scatter plots showing the correlations between the intrahost frequencies of (left) Spike_D614G and (NSP12_P323L) and (right) N_R203K and N_G204R. (d) Proportion of SAV pairs that involve only low fitness SAVs, defined as those that never exceeded an intrahost frequency of 50% nor were observed in more than 1000 consensus genomes across the 56 months.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/ff76a29bfc2c23e548a271ec.png"},{"id":69073628,"identity":"26328e4c-e6d5-4547-8ede-dba4190e09e8","added_by":"auto","created_at":"2024-11-15 10:36:43","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":275125,"visible":true,"origin":"","legend":"\u003cp\u003eConsideration of genetic linkage improves fitness predictions. (a) Nested cross-validation performance (assessed via r2) of XGBoost regression models with and without the inclusion of linkage predictors. (b) Distributions of the summed absolute SHAP values for each predictor category reflecting the aggregated impact of each predictor type. Predictor types are ranked from top to bottom by decreasing importance, as determined via their median absolute SHAP values. Differences in the relative importance of each predictor category were tested using two-sided paired t-tests for each dataset. The tests for each comparison and for each dataset were highly significant (all p\u0026lt;0.0001). (c) The changes in prediction error (i.e., |εlinkage |- \u0026nbsp;|εno linkage|) for the Alpha, Delta, BA.1, BA.5, XBB and BA.2.86 datasets after including linkage predictors, stratified by whether a SAV was observed at \u0026gt;10% monthly frequency after the timeframe of the dataset. (d) Impact of linkage predictors (i.e., ∑(SHAP)) on the predicted fitness of SAVs, stratified in the same way. Boxplot elements are defined as follows: centre line, median; box limits, upper and lower quartiles; whiskers, 1.5x interquartile range. Differences in distributions were tested using two-sided Mann-Whitney U tests and the corresponding p-values are annotated. \u0026nbsp;\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/336a06bbcca3bf36f77fa340.png"},{"id":69074432,"identity":"edc86118-8dc5-4b5d-b760-128320e767ae","added_by":"auto","created_at":"2024-11-15 10:52:45","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1953727,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/2417da90-e03b-4ecf-8142-c52e41279e01.pdf"},{"id":69074233,"identity":"a3539b00-414f-4817-ac0d-8c1d522680a6","added_by":"auto","created_at":"2024-11-15 10:44:43","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":1071532,"visible":true,"origin":"","legend":"Extended Data Figures","description":"","filename":"earlyintrahostdiversitysupplementary161024.docx","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/8625a2161fbaa1177e02a53d.docx"},{"id":69073631,"identity":"78bffc5c-b40b-4ef0-81b8-3387b750c776","added_by":"auto","created_at":"2024-11-15 10:36:44","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":7294346,"visible":true,"origin":"","legend":"Supplementary Table 1","description":"","filename":"SupplementaryTable1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5298116/v1/bcc741633e93de09945acd01.xlsx"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Intrahost dynamics, together with genetic and phenotypic traits predict the success of viral mutations","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSince its emergence in late 2019\u003csup\u003e1\u003c/sup\u003e, SARS-CoV-2, the agent of the COVID-19 pandemic, spread globally and caused over seven million deaths worldwide as of 21 July 2024 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://ourworldindata.org/grapher/cumulative-covid-deaths-region\u003c/span\u003e\u003cspan address=\"https://ourworldindata.org/grapher/cumulative-covid-deaths-region\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). On 30th January 2020, the World Health Organisation (WHO) declared the pandemic a Public Health Emergency of International Concern, which was then declared over on 5th May 2023 following falling death rates, the deployment of effective vaccines, and increased immunity in the global population (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.who.int/emergencies/diseases/novel-coronavirus-2019/interactive-timeline\u003c/span\u003e\u003cspan address=\"https://www.who.int/emergencies/diseases/novel-coronavirus-2019/interactive-timeline\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThroughout the COVID-19 pandemic, SARS-CoV-2 has exhibited a repeated pattern of lineage replacement, where multiple waves of genetically distinct lineages, termed \u0026lsquo;variants\u0026rsquo;, emerged and replaced the previous dominant variant in circulation\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. These variants, including Delta, Omicron, and BA.2.86 (an Omicron sub-lineage sometimes referred to as \u0026ldquo;Pirola\u0026rdquo;), carry constellations of co-occurring mutations, a phenomenon known as genetic linkage. Some of these mutations confer a fitness advantage to the virus such as increased transmissibility, replication or immune evasion\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Additionally, some linked mutations have been shown to act synergistically\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e, conferring greater fitness advantages than the sum of their effects in isolation. This process, termed synergistic epistasis, may explain some of the patterns of genetic linkage observed. The fitness differentials provided by these clusters of co-adapted mutations allow viruses carrying them to outcompete the preceding variant in circulation.\u003c/p\u003e \u003cp\u003eSARS-CoV-2 accumulates about 15 mutations per year with an overall evolutionary rate of ~\u0026thinsp;1x10\u003csup\u003e\u0026minus;\u0026thinsp;3\u003c/sup\u003e substitutions per site per year. The accrual of genetic diversity at the genome level is driven by multiple evolutionary forces at different levels. Within a single infection, genome replication errors and the action of host immune RNA editing may generate \u003cem\u003ede novo\u003c/em\u003e mutations in the viral population that are subjected to within-host selective pressures. Together, these processes are estimated to result in the acquisition of 1.3x10\u003csup\u003e\u0026minus;\u0026thinsp;6\u003c/sup\u003e mutations per nucleotide per replication cycle\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e. The intrahost viral population expands and reaches its peak population size around 2\u0026ndash;5 days after infection (coinciding with symptom onset)\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. The viral population then undergoes a narrow transmission bottleneck where only a small fraction of the intrahost viral population is passed on to the next individual in the transmission chain, with estimates as low as one to eight virions per transmission event within households\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Selective forces may play an important role during transmission, favouring alleles that improve transmission and gradually purging those that do not. As a result of these processes, the genetic diversity observed within hosts is expected to be substantially higher than that observed at the interhost level. However, since SARS-CoV-2 typically causes acute infections that last about 10\u0026ndash;15 days post symptom onset\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e, with samples often collected shortly after symptoms manifest, only a few intrahost mutations have been observed during this short window\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e,\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe relative contributions of these evolutionary forces - acting both at the intrahost level and during transmission - on the mutations observed in SARS-CoV-2 lineages remain unclear. Given the short evolutionary timescale and the narrow population bottleneck during transmission, the effects of genetic drift may interfere with the action of natural selection. Neutral or weakly deleterious alleles may be transmitted before they are purged, allowing them to be observed at considerable frequencies amongst sampled genomes, whereas there may be a time lag before weakly advantageous mutations reach high frequencies. As such, we hypothesised that by analysing intrahost allele frequencies, we may be able to predict the relative success of mutations arising in the evolution of SARS-CoV-2. In addition, we hypothesised that the linkage patterns observed at the genomic level may be recapitulated by those found within intrahost viral populations. Addressing these questions may provide insights into the trajectory of mutations in SARS-CoV-2.\u003c/p\u003e \u003cp\u003eSARS-CoV-2 is arguably one of the best-studied viruses, owing to concerted, global and multidisciplinary research efforts over the COVID-19 pandemic. One important outcome is the remarkable amount of genomic data that has been generated. As of July 2024, nearly 15\u0026nbsp;million complete genomes and over seven million sequencing libraries of SARS-CoV-2 have been deposited on the data repositories, GISAID\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e and NCBI Sequence Read Archive (SRA)\u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u003c/sup\u003e, respectively. These datasets and their accompanying metadata, including sample collection dates, sequence quality, lineage designations and viral mutations, provide an unprecedented \u0026lsquo;Big Data\u0026rsquo; resource for testing this hypothesis and addressing other important questions relating to viral evolution.\u003c/p\u003e \u003cp\u003eIn this study, we leverage large-scale SARS-CoV-2 genomic datasets to investigate the relationships between the genetic diversity observed within hosts and the relative fitness of SARS-CoV-2 mutations. We first provide a detailed overview of the landscape of mutations observed in SARS-CoV-2, exploring broad patterns in the physiochemical and phenotypic effects of all non-synonymous mutations observed over the COVID-19 pandemic. Using a curated dataset of 7,862 quality-controlled sequencing libraries generated from SARS-CoV-2 infections, we then test the extent to which intrahost diversity, together with other genetic and phenotypic traits, can predict the relative success of SARS-CoV-2 mutations. These insights are invaluable for understanding the evolutionary forces shaping SARS-CoV-2, and may offer scope to inform the selection of vaccine targets as well as responding to other current and future viral threats.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eThe landscape of mutations in SARS-CoV-2\u003c/h2\u003e \u003cp\u003eTo explore the patterns of mutations that have arisen over the COVID-19 pandemic, we analysed the metadata provided by GISAID\u003csup\u003e\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e,\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e for the first submitted SARS-CoV-2 sequence up to the 8th July 2024, comprising nearly 15\u0026nbsp;million complete SARS-CoV-2 genomes. Within this dataset, we identified a total of 101,484 single amino acid variants (SAVs). This represents 52% of the 194,660 SAVs that could theoretically be observed considering that SARS-CoV-2 coding regions encode 9,733 amino acids, excluding start and stop codons. The number of SAVs observed over time increased rapidly in the initial stages of the pandemic though rapidly tapered off, with half of all theoretical SAVs already observed by December 2022 (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ea). Of all the SAVs observed in our dataset, 101,263 (99.8%) never exceeded a monthly global frequency of 10%, 100 (0.1%) reached frequencies between 10\u0026ndash;90%, and 121 (0.1%) reached greater than 90% frequency (i.e., fixation). There were no obvious protein-specific patterns in the number of mutations observed for the vast majority of SAVs that failed to reached fixation (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). In contrast, SAVs that did were largely concentrated in the spike protein (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), consistent with its high antigenicity and crucial role in cellular entry.\u003c/p\u003e \u003cp\u003eThe evolution of SARS-CoV-2 is largely characterised by the repeated emergence of divergent lineages that replace the preceding dominant lineage in circulation (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). Many mutations that reached fixation at some point over the course of the pandemic exhibited frequency patterns largely resembling those of these dominant lineages (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). However, there were some exceptions. For instance, Spike_D614G and NSP12_P323L rose in frequency during the initial stages of the pandemic, reaching fixation around mid-2020 and subsequently remained fixed in the SARS-CoV-2 population; Spike_G142D gradually rose in frequency during the Delta wave and remained fixed while NSP4_L438F rose in frequency slightly after the initial Omicron lineages had replaced the preceding Delta variant and quickly dropped in frequency before reaching fixation again coinciding with the emergence of the XBB/EG.5 Omicron sub-lineages (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003eWe next explored whether the phenotypic impacts of SAVs observed were potential drivers of their relative fitness, which we estimated as the number of genomes submitted to GISAID carrying each SAV. We first calculated BLOSUM62 scores for all SAVs using the BLOSUM62 matrix, which evaluates amino acid changes based on their likelihood of occurring by chance, as derived from empirical protein alignments\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. Under this scoring system, non-conservative mutations (\u003cem\u003ei.e.\u003c/em\u003e, rare mutations) are assigned more negative scores\u003csup\u003e\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e,\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e. While this scoring system is typically used for scoring amino acid sequence alignments, it is thought to capture the effects of differing selective forces on amino acid substitutions, where stronger purifying selection acts on less conservative mutations\u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e,\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e. Since \u0026lsquo;conservativeness\u0026rsquo; is a statistical abstraction of the phenotypic effects of mutation and the consequences of natural selection, we further considered other metrics that capture these effects directly. In particular, we computed the predicted effects of each SAV on the physiochemical properties of proteins (i.e., change in charge, molecular weight and hydropathy). We also investigated the effects of SAVs found within the receptor-binding domain (RBD) on protein expression, host-receptor binding and antibody escape by leveraging direct measurements from previously published Deep Mutational Scanning (DMS) experiments\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e (see \u003cb\u003eMethods\u003c/b\u003e).\u003c/p\u003e \u003cp\u003eAcross SAVs, the score distributions for the three physiochemical properties (i.e., change in charge, molecular weight and hydropathy) appear symmetrically centred around zero (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ec), suggesting that most observed SAVs tend to preserve the physiochemical properties of the proteins they are found in. Separately, based on the DMS measurements, 83% and 70% of SAVs in the RBD were detrimental for host-receptor binding and protein expression, respectively, and most conferred minimal antibody escape (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ed). These results suggest that most SAVs observed are neutral or deleterious for these phenotypes.\u003c/p\u003e \u003cp\u003eNotably, we found a significant positive correlation between fitness and the BLOSUM62 score of all SAVs (Spearman\u0026rsquo;s \u003cem\u003eρ\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.33, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001; Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ee), consistent with the idea that non-conservative mutations are generally deleterious. Further, SAVs associated with larger absolute changes in physiochemical phenotypes had significantly lower fitness values (\u003cem\u003eρ\u003c/em\u003e=-0.13,-0.040,-0.23, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001 for molecular weight, charge and hydropathy, respectively). This relationship, however, does not hold for the small subset of SAVs that have reached fixation over the course of the pandemic (\u003cem\u003eρ\u003c/em\u003e=-0.068, \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.462; Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003ee), which are likely to confer the highest fitness gains. These findings suggest that while non-conservative mutations are generally deleterious, they are also more likely to confer the highest fitness gains to the virus. This is further corroborated by the relationships observed between the physiochemical effects of SAVs and fitness. Indeed, the absolute changes in physiochemical scores were significantly and negatively correlated with fitness (\u003cem\u003eρ\u003c/em\u003e=-0.040,-0.13,-0.23, respectively; all \u003cem\u003ep\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). Meanwhile, SAVs that reached fixation were associated with significantly higher absolute changes in charge compared to non-fixed SAVs (two-sided Mann-Whitney U test, p\u0026thinsp;=\u0026thinsp;0.0076). These results highlight that while more significant biochemical alterations are often deleterious, they can occasionally lead to the highest fitness gains.\u003c/p\u003e \u003cp\u003eFinally, SAVs that are associated with increased host-receptor binding and RBD expression had significantly higher fitness values (both ρ\u0026thinsp;=\u0026thinsp;0.26, p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001), which was expected given the key role of the RBD in mediating viral entry into host cells. Meanwhile, antibody escape was weakly and negatively correlated with fitness (ρ=-0.084, p\u0026thinsp;=\u0026thinsp;0.0005). While SAVs that improve antibody escape would likely be beneficial for the virus, and therefore be inherently fitter, this negative correlation may point to a trade-off between the different phenotypes. Indeed, antibody escape was significantly negatively correlated with both binding (ρ=-0.36, p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001) and expression (ρ=-0.11, p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001), indicating that, for many SAVs, the fitness advantage due to increased antibody escape may be counteracted by decreased binding and/or expression, thus resulting in a lower overall fitness.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e\n\u003ch3\u003eIntrahost dynamics predicts the future fitness of mutations\u003c/h3\u003e\n\u003cp\u003eWe next investigated whether the intrahost diversity of SARS-CoV-2 infections could be used to forecast the relative fitness of SAVs at the population level. To do this, we curated and downloaded seven sequencing datasets (total n\u0026thinsp;=\u0026thinsp;7862) comprising sequencing libraries of SARS-CoV-2 infections, which were sampled at various timepoints across the pandemic. These datasets correspond to seven time periods of sampling, henceforth timeframes, termed \u0026lsquo;Early\u0026rsquo;, \u0026lsquo;Alpha\u0026rsquo;, \u0026lsquo;Delta\u0026rsquo;, \u0026lsquo;BA.1, \u0026lsquo;BA.5\u0026rsquo;, \u0026lsquo;XBB\u0026rsquo; and \u0026lsquo;BA.2.86\u0026rsquo; (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb). The Early timeframe corresponds to the first four months of the pandemic, while the others span the transition points between the major variant waves in our dataset. We then estimated the intrahost frequencies of SAVs in these libraries, retaining only SNPs/SAVs that were supported by at least 100 de-duplicated reads and that were observed at \u0026gt;\u0026thinsp;3% frequency. We selected these thresholds as they have been previously applied to reduce noise arising from sequencing errors\u003csup\u003e\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. Additionally, as some libraries may be associated with serial passaging experiments, laboratory contamination, an excess of sequencing artefacts, or erroneous collection dates, we further omitted libraries whose consensus genomes were outliers when assessing the number of single nucleotide polymorphisms relative to the Wuhan-Hu-1 reference (GenBank: MN908947.3) compared to other consensus genomes that were assigned the same lineage (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e; See \u003cb\u003eMethods\u003c/b\u003e). Our final quality-controlled dataset comprised 5885 sequencing libraries, representing seven snapshots of the lineage composition at various points during the COVID-19 pandemic (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWithin each of the seven datasets, we identified between 2872\u0026ndash;9606 intrahost SAVs across the sequencing libraries (\u003cb\u003eSupplementary Table\u0026nbsp;1\u003c/b\u003e), with an increase in the median number of SAVs detected per library by the chronological order of the datasets (Early\u0026thinsp;=\u0026thinsp;25; BA.2.86\u0026thinsp;=\u0026thinsp;128). Within each timeframe, 9-112 (0.1\u0026ndash;3.9%) of these intrahost SAVs were already at \u0026gt;\u0026thinsp;10% monthly frequency at the interhost level. Excluding these SAVs, we found that 4\u0026ndash;73 (0.15\u0026ndash;1.1%) SAVs reached at least a monthly frequency of \u0026gt;\u0026thinsp;10% after the timeframe of each dataset (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). This pattern held for all except the BA.2.86 timeframe, where we observed a substantially (1.8\u0026ndash;5.1 fold) higher proportion of high frequency SAVs than the overall expectation of 221/101,795 (0.22%) SAVs when considering the entirety of the GISAID data (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ea). Additionally, the intrahost SAVs that were observed at \u0026gt;\u0026thinsp;10% frequency after the dataset timeframes only reached peak frequency after a median of 6\u0026ndash;40 months. In other words, many of these relatively fitter SAVs could already be observed in the intrahost diversity of infections that were sampled much earlier in the pandemic.\u003c/p\u003e \u003cp\u003eTo explore whether we could forecast the relative fitness of SAVs observed at the intrahost level, we estimated the fitness of each SAV as the number of submitted consensus genomes a SAV was observed in after the timeframe of each dataset (henceforth, \u0026lsquo;future fitness\u0026rsquo;). We additionally generated 13 predictors of future fitness for each mutation, namely: the number of libraries each SAV was detected in, median and maximum intrahost frequency, BLOSUM62 score, the estimated raw and absolute change in charge, hydropathy or molecular weight, and for RBD mutations, the change in binding and expression, and antibody escape fraction, as directly measured via DMS\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e,\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. Most predictors were significantly correlated with the future fitness of SAVs (Spearman\u0026rsquo;s correlation test, all adjusted p\u0026thinsp;\u0026lt;\u0026thinsp;0.05; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb) and the direction of the correlations were largely consistent across datasets. Maximum and median intrahost frequency, BLOSUM62 score, hydrophobicity and molecular weight were positively correlated with future fitness, whereas absolute changes to charge, molecular weight and hydropathy were negatively correlated with future fitness (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eb).\u003c/p\u003e \u003cp\u003eFor each dataset, we trained supervised XGBoost\u003csup\u003e\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e regression models \u0026ndash; that can capture complex non-linear relationships \u0026ndash; to predict the future fitness of SAVs based on predictors generated from the intrahost SAVs detected. We optimised the models and assessed their performance using a cross-validation procedure (see \u003cb\u003eMethods\u003c/b\u003e), where different subsets of the data are withheld to test whether the models can make robust predictions on unseen data. We found that all models were able to predict the future fitness of intrahost SAVs, yielding mean \u003cem\u003er\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/em\u003e\u003c/sup\u003e values between 48\u0026ndash;63% for each dataset. Additionally, the mean absolute error of the predictions by each model was 0.45\u0026ndash;0.68, indicating that, on average, the models were able to predict the future fitness of intrahost SAVs to within the same order of magnitude. These findings indicate that our models could reliably forecast the success of SAVs based on the intrahost diversity of SARS-CoV-2 infections.\u003c/p\u003e \u003cp\u003eTo understand how these complex models were making their predictions, we further used the model interpretation framework, SHAP\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, to visualise the relationships between the predictors and mutation fitness. The core of this framework is the SHAP value, which here represents, for each SAV, the estimated contribution of a single predictor to the predicted fitness value for that SAV. Across the datasets, the median absolute SHAP value for maximum intrahost frequency was the highest (median\u0026thinsp;=\u0026thinsp;0.18) indicating that this predictor was the most important contributor to the model predictions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ec). Meanwhile, RBD binding, expression, and mean antibody escape were the least important, though noting that these directly measured traits are only available for SAVs in the RBD. We additionally explored the relationships between the predictors and the future fitness of SAVs by comparing the predictors with their corresponding SHAP values (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed-g). Maximum intrahost frequency had a monotonic positive relationship with predicted future fitness (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ed), indicating that the selective forces governing the intrahost frequencies of SAVs explain, in part, mutational success at the global level. Separately, there was an apparent parabolic relationship between change in hydropathy and fitness, where SAVs involving relatively larger increases or decreases to hydrophobicity are generally associated with lower mutational fitness (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ee), concordant with the idea that significant physiochemical changes are generally negatively associated with fitness. Notably, some SAVs associated with moderate absolute hydropathy or molecular weight changes had the most positive SHAP values (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e), suggesting that whilst most physiochemical changes are deleterious, they can also lead to positive fitness gains. Similarly, BLOSUM62 score was positively correlated with predicted fitness, with some negative scores associated with positive SHAP values (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003ef). These findings recapitulate the previously observed pattern that larger physiochemical alterations are typically associated with lower fitness, but can, in some cases, also be linked to higher fitness. Finally, RBD binding (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eg; \u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) and expression (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e) were positively correlated with higher predicted fitness values. These findings indicate that, as expected, SAVs associated with increases in binding and expression are fitter. All patterns were largely concordant across the different timeframes considered (\u003cb\u003eExtended Data\u003c/b\u003e Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e), highlighting that our models learned biologically relevant patterns to predict the future fitness of mutations.\u003c/p\u003e \u003cp\u003eTo further test the generalisability of our models, we trained an additional XGBoost regression model on the Alpha dataset, comprising libraries collected in February 2021, to predict the future fitness of intrahost SAVs up to before the first Omicron wave (i.e., BA.1; November 2021). We then tested the trained model on the BA.1 dataset (libraries collected in December 2021), to predict the fitness of intrahost SAVs from BA.1 to the end of our dataset (July 2024). Overall, the model trained on the Alpha dataset and pre-Omicron fitness data was able to reliably predict the future fitness of Omicron SAVs with the BA.1 dataset (Spearman\u0026rsquo;s ρ\u0026thinsp;=\u0026thinsp;0.72; r\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.52; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eh), indicating that this model is indeed generalisable. Notably, while some of the intrahost SAVs that reach a future frequency\u0026thinsp;\u0026gt;\u0026thinsp;10% are predicted to have the highest future fitness, many of the high frequency SAVs had markedly lower predicted fitness values than observed. Indeed, the studentised residuals for high frequency SAVs \u0026ndash; in a linear model of predicted versus observed fitness \u0026ndash; were considerably more negative than those for low frequency SAVs (median=-1.66 and \u0026minus;\u0026thinsp;0.066, respectively; two-side Mann-Whitney U test, U\u0026thinsp;=\u0026thinsp;372,623, p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). This prediction error may be a result of other evolutionary processes not captured by the model including epistatic interactions between high frequency mutations.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eIntrahost evolution shapes global linkage patterns\u003c/h3\u003e\n\u003cp\u003eWe next explored whether genetic linkage at either the intrahost or interhost level could explain why some of these SAVs are much fitter than predicted by our initial models. To estimate intrahost linkage, we calculated Pearson\u0026rsquo;s \u003cem\u003er\u003c/em\u003e using the intrahost SAV frequencies of libraries within each dataset. At the interhost level, we calculated the \u003cem\u003eD\u0026rsquo;\u003c/em\u003e statistic\u003csup\u003e\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e using the SAV frequencies estimated from the GISAID consensus genomes collected prior to the sampling timeframe of each dataset. To reduce noise arising from low frequency SAVs, we considered only SAVs that were observed in more than 1000 genomes. For biallelic loci (i.e., only two alleles or amino acid variants are observed at each locus), linkage is often reported as \u003cem\u003e|D\u0026rsquo;|\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u003csup\u003e\u003cem\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/em\u003e\u003c/sup\u003e since negative values indicate linkage between the alternate allele at one locus and the wild type at the other. In our context, there are 20 possible SAVs at each locus so negative \u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e values are not as meaningful and can only be interpreted as \u0026lsquo;not positively linked\u0026rsquo;.\u003c/p\u003e \u003cp\u003eIn total, we calculated \u003cem\u003eD\u0026rsquo;\u003c/em\u003e for 143,889-7,714,482 and \u003cem\u003er\u003c/em\u003e for 93,528-1,326,006 SAV pairs for each dataset, of which 0.1\u0026ndash;0.9% and 0.01\u0026ndash;0.6% are strongly linked (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.9), respectively. Intrahost linkage (i.e., \u003cem\u003er\u003c/em\u003e) was significantly correlated with that observed at the interhost level (i.e., \u003cem\u003eD\u0026rsquo;\u003c/em\u003e) (Pearson\u0026rsquo;s \u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.28\u0026ndash;0.45, all adjusted p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001; Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea; \u003cb\u003eExtended Data Fig.\u0026nbsp;5\u003c/b\u003e), indicating that the evolutionary processes shaping linkage patterns at the interhost level are also seen within hosts. At both the interhost and intrahost level, within-protein linkage associations were significantly more common in strongly and positively linked (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.9) SAVs than those with weaker linkage associations (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u0026thinsp;\u0026le;\u003c/em\u003e\u0026thinsp;0.9) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb; odds ratio\u0026thinsp;=\u0026thinsp;1.7\u0026ndash;42, Fisher\u0026rsquo;s exact test, all adjusted p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001). This suggests that our linkage statistics likely capture both the signatures of recombination (i.e., linkage decay) and epistasis (i.e., short-range interactions are more common).\u003c/p\u003e \u003cp\u003eThe SAV pairs that appear strongly linked at both the intrahost and interhost level included SAVs that were highly successful, including N_R203K-N_G204R (\u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.98 and D\u0026rsquo;=1.00 for the Alpha timeframe; Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). The close proximity of these two SAVs and the near-perfect linkage patterns indicate likely synergistic interactions. Separately, Spike_D614G-NSP12_P323L, one of the first few pairs of SAVs whose monthly frequencies co-rose rapidly in the early stages of the pandemic (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eb), had a slightly weaker intrahost correlation (\u0026lsquo;Early\u0026rsquo; timeframe, intrahost \u003cem\u003er\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.84; Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). Their occurrence on different proteins and the weaker linkage pattern may point to co-selection of two highly beneficial mutations\u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e in a largely wild-type background, rather than epistasis. Notably, 73\u0026ndash;98% of strongly linked pairs involved only SAVs that never exceeded an intrahost frequency of 50% nor were observed in more than 1000 consensus genomes across the 56 months for all timeframes except the BA.2.86 timeframe (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ed). The low consensus prevalence and intrahost frequencies of these SAVs jointly indicate that the majority of linked SAVs do not confer a significant fitness advantage and are a product of genetic drift. Additionally, if most mutations are deleterious, then these linkage patterns may represent compensatory epistasis. Overall, these findings indicate that linkage patterns observed within hosts may provide clues on the selective forces shaping those observed at the interhost level.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e\n\u003ch3\u003eGenetic linkage improves fitness predictions\u003c/h3\u003e\n\u003cp\u003eTo incorporate genetic linkage at both the intrahost and interhost levels into our models, we generated six further predictors, in addition to the thirteen used previously, namely, whether the SAV is strongly linked to at least one other SAV (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.9), the number of strongly linked SAVs (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.9), and the maximum \u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e across all pairs involving the SAV. In addition, \u003cem\u003eD\u0026rsquo;\u003c/em\u003e statistics were only calculated for SAVs observed in \u0026gt;\u0026thinsp;1000 consensus genomes prior to the timeframe of each dataset so, given that prior fitness is highly autocorrelated with future fitness (Spearman\u0026rsquo;s \u003cem\u003eρ\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.88\u0026ndash;0.92), models may \u0026lsquo;cheat\u0026rsquo; by leveraging this autocorrelation to improve performance without learning any additional biological patterns. As such, we included an additional binary predictor, indicating whether a SAV was observed in \u0026gt;\u0026thinsp;1000 genomes (henceforth \u0026lsquo;prior fitness\u0026rsquo;), to partition this effect in our models.\u003c/p\u003e \u003cp\u003eInclusion of linkage predictors significantly improved the performance of all models (Mann-Whitney U test, maximum \u003cem\u003ep\u003c/em\u003e\u0026thinsp;=\u0026thinsp;0.018), except for that trained on the Alpha timeframe (p\u0026thinsp;=\u0026thinsp;0.22). The most prominent improvement was seen for the model trained on the BA.2.86 timeframe (mean \u003cem\u003er\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.49 to \u003cem\u003er\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.67) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ea). All models trained on one timeframe and tested on another (as for Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eh) also performed better with the inclusion of linkage predictors (\u003cb\u003eExtended Data Fig.\u0026nbsp;6\u003c/b\u003e). This indicates that the relationships between linkage and future fitness learned by the regression models were generalisable across datasets. Since SHAP values are additive\u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e, we compared the sum of the absolute SHAP values for the different types of predictors (i.e., physiochemical, intrahost, intrahost linkage, interhost linkage and prior fitness) to assess their relative importance. Interhost linkage predictors were significantly more important than intrahost linkage predictors (paired \u003cem\u003et\u003c/em\u003e-test, \u003cem\u003et\u003c/em\u003e\u0026thinsp;=\u0026thinsp;16.0-91.2, d.f.=2761\u0026ndash;9560, all p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001) and the prior fitness predictor for predicting future fitness, across all timeframes (\u003cem\u003et\u003c/em\u003e\u0026thinsp;=\u0026thinsp;23.8\u0026ndash;95.1, d.f.=2761\u0026ndash;9560, all p\u0026thinsp;\u0026lt;\u0026thinsp;0.0001) (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eb). These findings indicate that linkage patterns at the interhost level are largely responsible for the performance boost.\u003c/p\u003e \u003cp\u003eSince the fittest SAVs were associated with significantly higher prediction errors than the less successful SAVs in our initial models (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eh), we hypothesised that genetic linkage may explain, at least in part, this pattern. To test this, we compared the changes in prediction error for our models with or without linkage predictors for SAVs where \u003cem\u003eD\u0026rsquo;\u003c/em\u003e could be calculated. Following the inclusion of linkage predictors, the changes in prediction error were more negative for SAVs that reached a future frequency of \u0026gt;\u0026thinsp;10% compared to those that did not (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ec). This effect was statistically significant for the Alpha, Delta and BA.1 timeframes (Mann-Whitney U test, all p\u0026thinsp;\u0026lt;\u0026thinsp;0.001), but not for the XBB and BA.2.86 timeframes (p\u0026thinsp;=\u0026thinsp;0.804 and 0.063, respectively). This could be due to the smaller number of high frequency SAVs considered in these timeframes. Additionally, we aggregated the contributions of all linkage predictors by summing their SHAP values and found that the linkage predictors resulted in significantly higher predicted fitness values for SAVs that reached a high future frequency (all p\u0026thinsp;\u0026lt;\u0026thinsp;0.05; Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003ed). Furthermore, the summed SHAP values of linkage predictors were all negatively correlated with the changes in prediction error across all datasets (Spearman\u0026rsquo;s \u003cem\u003eρ\u003c/em\u003e=-0.686 to -0.029). This implies that by using information about genetic linkage, these models were predicting higher future fitness values for the high frequency SAVs, thus resulting in lower prediction errors. Overall, these findings suggest that the success of high frequency SAVs may, in part, be governed by genetic linkage.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003ePredicting which mutations will emerge and reach high frequency in the future is a long-standing and important, yet largely unsolved problem in the study of pathogen dynamics. In this study, we investigated whether the intrahost diversity of SARS-CoV-2 infections could provide clues on the relative fitness of SAVs by leveraging the unparalleled amount of sequencing data available. Our key assumption is that the evolutionary forces shaping the interhost frequencies of mutations in the viral population are also acting within hosts. Using various machine learning techniques, we showed that this assumption holds, that is, that the intrahost dynamics of SARS-CoV-2 infections can predict the future fitness of mutations remarkably well. In addition, our model interpretation analyses revealed highly generalisable patterns of intrahost dynamics and phenotypic traits that correlate with the success of mutations. Our findings lay the groundwork for developing better predictive tools that may be invaluable for pre-empting the emergence of novel lineages and supporting the design of vaccine targets. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBy analysing the distribution of SAVs across all complete genomes on GISAID (as of July 2024), we found that more than half of all possible SAVs have already emerged in the evolution of SARS-CoV-2, and that the rate of emergence of previously unseen mutations has slowed (\u003cstrong\u003eFig. 1a\u003c/strong\u003e). Most SAVs observed thus far tend to preserve the physiochemical properties of the proteins they are found in, suggesting that major phenotypic alterations to proteins are generally deleterious (\u003cstrong\u003eFig. 1c\u003c/strong\u003e). This is corroborated by the observation that most SAVs in the RBD are estimated to be phenotypically neutral or deleterious (\u003cstrong\u003eFig. 1d\u003c/strong\u003e). Further, only an extremely small subset of SAVs have reached a high frequency at the interhost level thus far. Collectively, these findings point to strong functional and evolutionary constraints on the mutational space that SARS-CoV-2 can explore. Since the rate of emergence of mutations has slowed considerably, one can speculate that future variant lineages may be more likely to harbour novel combinations of moderately or highly fit mutations that have already been observed so far rather than a set of previously unseen mutations. This, however, does not preclude the possibility of unseen mutations emerging in the future.\u003c/p\u003e\n\u003cp\u003eIn this study, we used the number of GISAID sequences carrying each SAV across various time windows as a proxy for estimating SAV fitness. While it may be possible to get more direct fitness estimates through animal challenge experiments\u003csup\u003e23–25\u003c/sup\u003e, it is not tractable to do so at the scale of SAVs considered in this study. Separately, our fitness estimates are likely affected by the spatiotemporal biases in genomic surveillance efforts. For example, the estimated fitness of a moderately successful SAV that circulated only in countries with lower surveillance would be lower than a less successful SAV that spread in countries with more extensive surveillance. However, these discrepancies are likely random and manifest as noise, rather than a systematic bias, in the fitness estimates. As such, these sampling biases are not likely to skew the patterns observed in our study.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOur modelling approach differs from previous efforts to understand viral fitness, which have largely focused on the relative fitness of variant lineages\u003csup\u003e26–31\u003c/sup\u003e, whereas we consider individual mutations as targets of selection, each associated with its own fitness values. The theoretical number of combinations of SAVs that SARS-CoV-2 can have is practically infinite (i.e., 20\u003csup\u003e9733\u003c/sup\u003e), which entails that the fitness of previously unseen lineages is difficult to predict, especially for models based on epidemiological data\u003csup\u003e27,29\u003c/sup\u003e. Further, the fitness landscape of lineages is likely confounded by the immunological history of the population and the extant competing lineages in circulation. Focusing on the fitness of mutations is lineage agnostic, circumventing these limitations. Additionally, our models can learn the inherent genetic and phenotypic traits of mutations that underly their success, enabling richer biological interpretations. For example, our models reveal that the SAVs associated with more aggressive physiochemical changes are generally under stronger purifying selection, and expectedly, that the SAVs that negatively impact RBD binding and expression are less fit.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eImportantly, our modelling approach leverages information about the intrahost dynamics of mutations, viewing each infection as an independent playing field where mutations compete for dominance, with fitter mutations being observed more frequently than less fit ones. Our observation that maximum intrahost frequency is consistently the most important predictor of future fitness supports this, and suggests that within-host evolutionary processes are important components that shape the fitness landscape of mutations at the interhost level. Such processes may include positive selection for mutations that improve viral infectivity and/or replication rate, and purifying selection acting on highly deleterious mutations that drastically alter protein function.\u003c/p\u003e\n\u003cp\u003eNotably, our models performed relatively poorly on making predictions for the fittest mutations (\u003cstrong\u003eFig. 2h\u003c/strong\u003e). Incorporating genetic linkage into the models partially reduced this effect, indicating that the co-occurrence of SAVs, which may or may not involve epistasis, plays a role in modulating the fitness of mutations in the evolution of SARS-CoV-2. However, there remains considerable variation in the fitness of SAVs that is not captured by our models, suggesting that some components of mutational fitness are still unaccounted for. Incorporation of further functional, structural or immunological predictors may reduce this gap further, but this may be challenging since existing experimental data for SARS-CoV-2 remain biased towards the spike protein. Additionally, selective forces may act at the point of transmission, with mutations improving transmissibility being more successful, but this is not captured by our models. Alternatively, these discrepancies may indicate that some of these highly successful mutations are not inherently the fittest, but represent a random subset of moderately fit mutations that rose to a high frequency due to genetic hitchhiking\u003csup\u003e32\u003c/sup\u003e or drift.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eOverall, our results highlight key evolutionary forces governing the success of mutations in the evolution of SARS-CoV-2, and show that the relative fitness of mutations is largely predictable when considering the intrahost dynamics of mutations and their genetic and phenotypic properties. However, the higher prediction error for the most successful mutations indicates that forecasting the exact set of mutations that will reach a high frequency remains out of reach, for now. Nevertheless, the generalisable patterns inferred by our models enables us to narrow down the list of mutations that are likely to be successful in the evolution of SARS-CoV-2. These could be combined with insights gleaned from other models of mutation or lineage fitness, and from immunological assays, to potentially inform on candidate mutations for future iterations of vaccines. Furthermore, we foresee that this approach could be easily ported to other comparable model systems for which large-scale genomic data is available, such as H3N2 Influenza A, which causes seasonal outbreaks globally. There may also be potential to leverage information derived from intrahost genetic variation in very different pathogens, for example to predict the future emergence of antimicrobial resistance in bacteria exposed to drugs.\u003cbr\u003e\u0026nbsp;\u003c/p\u003e"},{"header":"Methods","content":"\u003ch2\u003eData acquisition\u003c/h2\u003e\n\u003cp\u003eWe compiled a comprehensive dataset of SARS-CoV-2 genome assemblies. To do so, the metadata, including the list of SAVs, for all SARS-CoV-2 consensus genomes (n=16,798,863) deposited on GISAID\u003csup\u003e10,11\u003c/sup\u003e was downloaded on 8\u003csup\u003eth\u003c/sup\u003e July 2024 in the \u003cem\u003eTSV\u003c/em\u003e file format. We retained only genome entries were marked as complete (i.e., \u0026gt;29,000nt), were not low coverage (i.e., \u0026gt;5% Ns), were not isolated from the host genera \u003cem\u003eManis\u003c/em\u003e or \u003cem\u003eRhinolophus\u003c/em\u003e, were not associated with cell cultures, and had complete collection dates (i.e., year, month and day). For collection months with more than 200,000 genomes deposited, we randomly subsampled to 200,000 entries.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eThe metadata of all BioSamples hosted on NCBI were downloaded on 21\u003csup\u003est\u003c/sup\u003e March 2024 from the FTP site (https://ftp.ncbi.nlm.nih.gov/biosample/) in \u003cem\u003eXML\u003c/em\u003e file format. This was converted to \u003cem\u003eTSV\u003c/em\u003e format with a custom Python script that uses the \u003cem\u003eElementTree XML API\u003c/em\u003e v1.3.0. All BioSamples whose taxonomy name, BioSample title, isolate or strain fields contained the terms \u0026lsquo;\u003cem\u003eSevere acute respiratory syndrome coronavirus 2\u003c/em\u003e\u0026rsquo;, \u0026lsquo;\u003cem\u003ecov-2\u003c/em\u003e\u0026rsquo; or \u0026lsquo;\u003cem\u003ecov2\u003c/em\u003e\u0026rsquo; were retained. BioSamples were linked to their associated sequencing libraries where possible using NCBI\u0026rsquo;s \u003cem\u003eentrez-direct\u0026nbsp;\u003c/em\u003ecommand-line API v21.6. Sequencing libraries were downloaded in the \u003cem\u003eSRA\u003c/em\u003e file format and converted to \u003cem\u003eFASTQ\u003c/em\u003e format using the \u003cem\u003eprefetch\u003c/em\u003e and \u003cem\u003efastq-dump\u003c/em\u003e commands, respectively, as part of NCBI\u0026rsquo;s \u003cem\u003esra-tools\u003c/em\u003e API v3.11.\u0026nbsp;For the Early dataset, we considered only BioSamples whose collection year and month were known and between 1\u003csup\u003est\u003c/sup\u003e October 2019 and 1\u003csup\u003est\u003c/sup\u003e March 2020. For the rest of the datasets, we considered only BioSamples with complete collection dates that corresponded to the following collection months: Alpha (February 2021), Delta (June 2021), BA.1 (December 2021), BA.5 (June 2022), XBB (February 2023), BA.2.86 (December 2023) (\u003cstrong\u003eFig. 1b\u003c/strong\u003e).\u003c/p\u003e\n\u003ch2\u003eBioinformatic pipeline for sequencing data\u003c/h2\u003e\n\u003cp\u003eTo ensure the homogeneity of sequencing data, we analysed only paired-end Illumina libraries. If BioSamples were associated to multiple sequencing libraries, the sequencing reads were combined prior to further processing. We removed known sequencing adapters, trimmed read ends (\u0026lt;Q20), and removed low quality read pairs (mean base quality of \u0026lt;Q20) using the \u003cem\u003eBBDuk.sh\u003c/em\u003e utility of the BBMap v39.06 package (sourceforge.net/projects/bbmap/)\u003csup\u003e33\u003c/sup\u003e. We then aligned the quality-filtered read pairs to the CHM13 human reference genome (assembly accession: GCF_009914755.1) using Bowtie2 v2.5.1\u003csup\u003e34\u003c/sup\u003e, retaining read pairs where both members were unmapped. The unmapped read pairs were then re-aligned to the Wuhan-Hu-1 SARS-CoV-2 reference genome (GenBank accession: MN908947.3), retaining only read pairs where both members were mapped. Duplicate reads were removed using the \u003cem\u003emarkdup\u003c/em\u003e utility from Samtools v1.20\u003csup\u003e35\u003c/sup\u003e, to minimise the impact of PCR duplicates on our intrahost diversity analysis.\u0026nbsp;SAV frequencies were estimated from the de-duplicated, high quality reads using the \u003cem\u003ecall codonvar\u0026nbsp;\u003c/em\u003eutility in Quasitools v0.7.0\u003csup\u003e36\u003c/sup\u003e, specifying an sequencing error rate corresponding to Q20 (\u0026lsquo;\u003cem\u003e--error_rate 0.01\u003c/em\u003e\u0026rsquo;).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eSeparately, haploid variant calling was performed using the multi-allelic caller (\u0026lsquo;\u003cem\u003e-m\u003c/em\u003e\u0026rsquo;) in Bcftools v1.20\u003csup\u003e37\u003c/sup\u003e, specifying a minimum mapping quality and base quality of Q30 (\u0026lsquo;\u003cem\u003e--min-MQ 30\u003c/em\u003e\u0026rsquo; and \u0026lsquo;\u003cem\u003e--min-BQ 30\u003c/em\u003e\u0026rsquo;, respectively). Consensus genomes were then generated from the variant calls using a custom R script, masking positions where the coverage depth is less than 10, and ignoring indels. Sites that were masked in \u0026gt;10% of genomes were also replaced with Ns. Finally, the re-assembled consensus genomes were assigned PANGO lineages\u003csup\u003e38\u003c/sup\u003e and variant names using Pangolin v4.3\u003csup\u003e39\u003c/sup\u003e.\u003c/p\u003e\n\u003ch2\u003eQuality control of sequencing libraries\u003c/h2\u003e\n\u003cp\u003eThe diversity of SAVs detected in the sequencing data could be confounded by whether libraries were associated with serial passaging experiments, laboratory contamination, an excess of sequencing artefacts, or erroneous collection dates. In all these cases, we would expect a larger or smaller than expected number of mutations relative to other libraries at each timepoint. For each dataset, we removed sequencing libraries whose consensus genomes had fewer than 1.5 times the lower quartile, or greater than 1.5 times the upper quartile, of the number of SNPs for that dataset. For the Early dataset, we additionally removed libraries that were classified as either Alpha or Beta variant, as these variant lineages are known to have emerged at later stages in the pandemic so were likely mislabelled. BioSamples in the BA.1 and XBB datasets with no lineage assignments were also excluded as they had a much lower number of SNPs compared to other lineages in circulation within the timeframes of these datasets (\u003cstrong\u003eExtended Data Fig. 2\u003c/strong\u003e). We additionally removed BioSamples whose re-assembled consensus genomes had \u0026gt;10% Ns.\u003c/p\u003e\n\u003ch2\u003eThe fitness of mutations\u003c/h2\u003e\n\u003cp\u003eThe response variable for all models was the future fitness of a SAV after the sampling timeframe of a particular dataset. We define this as the number of genomes in the GISAID metadata carrying said SAV after the timeframe of the dataset with a one month buffer. Similarly, prior fitness is defined as the number of GISAD genomes carrying said SAV before the dataset timeframe with a one month buffer. For example, the Alpha dataset comprises samples collected in February 2021 and we estimated future and prior fitness as the number of GISAID genomes carrying each SAV between April 2021 and July 2024, and between December 2019 to December 2020, respectively.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003ePhysiochemical properties and DMS phenotypes\u003c/h2\u003e\n\u003cp\u003eFor each SAV, we generated the physiochemical predictors using the \u003cem\u003eBio.SeqUtils.ProtParam\u003c/em\u003e module in BioPython v1.84. Changes and absolute changes in physiochemical properties were \u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAUUAAAAXCAYAAACcRV0VAAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAAAoBSURBVHhe7Z3bqw5fGMeHf8DxEgkXXIicklBcIEqEkJIiQinkzB1CuVGOUS4IUSRbDsXFxoVDceUGuXC5nf+A/ZvPs9f3tfYy78zs9323d/b7m0+tZmbNWjPP86y1nnWad+9+nTFRSUlJSYnR3x1LSkpKSmJKp1hSUmAWLFgQnThxwl2VNIo0u5ZOsaSkpMSjdIolJSUlHg1zis+ePYv69etnx0bDUJfQU5DlwIED7qp+GG6jY0lJPfRmWympn5YeKR4+fDh68+aNu2o96ChoXH4YPHhwtGXLlujr168uVbFpBR2agTroMKxatSp69+6dS1VsiqpDOX3uw1y9etWO8+fPj/iyinDmzJno7Nmz0aFDh+xe0WkFHZrB7t27oylTpti57Nbe3h69fv06WrZsmcUXnaLqYE6RYbzfY3PuD+3psem5x4wZU0kzderU6P79+y5FMuSjN6DnJw9HrjUCqDYdrRbvQxpkkDzIRpwg7uHDhxY41z3JJF2QiSl2OCqhp5JNSHv9+nV3Jx9Z8kFo9zw29RkyZIgd586da0egl8XBfPr0ycUUmyLqkKfsiKesqDuq30nlp7bDfdW1RoHtsJOYOXNmtHHjxujjx48upvgUUYf+FOKsWbPsoq2tzQIFSZyGsDRcnALTUTz5tWvXLH7RokWpw1zyHTt2LNq7d6/l47hnz55o69atLkVtULF4zooVK+y5yEyPQ9z58+ctDfHEEThfvny5xa9Zs8Zk2rVrl8VrVIKsAp3mzJljduDZp06dik6ePBk9efLEpUgnj3w0+MWLF0cDBw60NITRo0ebTfM6RnVcM2bMsKMIHXyRKZoOecpObNu2zZZnrly5YumA8pMzRwe1HdrM3bt3rbHTjhoBHb7fmcDPnz/dWd+gkDrEDbGT4NPR0cGvXDrPnTvX+fbt285BgwZ1xoXq7nYRVxhLc/z48W7XHIH0XIf59u/fb+/jHeQlTUgYH/ckFgT5N2/e7K7+QB4/XZhPMklGIdkl68qVK01nZBScE5ckb0ge+bAt13EjsWvBO2TTLJLsR3kRh537AkXTIW/d4jpsN2GbUBmH9Y3nJMUnQdqk+qB3+c+gjiJT7MRdTLFppg7V7Ar96bkYrvowpI3vRZs2bYomTJgQffv2zaY0Pgxz07hz544dw3xHjhyJPnz4UJk21QL5GeGFxIq6s2QkEzBCURC6f+PGDZPbl5HzUJdq5JFv+PDhdowbfreRIbZmrSUPjFz9Z/Ic1mIY2ezYscPFFpui6dCTuhV3nu6si7BN3L5920b/Yfy6devcWe28ePHCjno2sxtmQdSfixcvWlzRKaoOtqYYTl2SQGCmAayt4BxYZ0njx48fmU6qHpii0ICQhzUb1nMYiqeBTMDSQBiA+5r6jBw50o4+SXHVyJJv4cKF5hBxwEy5WHPCrj1Zu+R5BPISmM7NmzcvevDgQV2dThZ0JHpnVvCXJZJolg5p5K1bAwYMcGfJ4GCT2smwYcPcWe1oKUd2Y7mHpZinT5/aQKY30TvzhDSaqUMauXafqRgTJ06MVq9ebYqwYNyodZFawHHQA+NMbt68aT3Lhg0bbHSRBWkYBScFGmIjyCsfo2ZG6vE0y0YdOEhsnGcxXiPceMpfkV+jnLzORBsE/mg5D/Tsvt3SQppNG6FDFnTm2jTJctBQT936l+Ck42l+xW7Iiex5nQkdMDYJN5DyoHfmCWnUq0MWrOlKT0JezClqGOtDRcIZsrjMRkQ8/zbBqeRU2nHjxrmUyeDxk3pXemAaI5W1GupBkiAfjgMn0tHREb169coMyVQ/z44V2/1pi/ijRo2y4+fPn+3okxQX0lP5eB/3SEN6RtdHjx51d6ujMps2bZoda4GFf/Cnd5Q35f4vaIQOWTAVx6m1t7dHkydPdrHJ1Fu3Qhgl4uRDvnz54s5qQ53J7Nmz7VgL6AX+LBH9s2aAjaIROmShT7rUrvLSn17xwoUL7rILpg84D6aL2glav369HUWSw/NZsmSJHWV8ce/ePTv6UwhNWQGH9fLlS3f1N79//7bj2rVru40mMPL379/dVTJay7l06ZIdBZXB7zXpvZDbd56ch7okkVc+nA/v9N9BejnlLOg4KLu86ZP49evXXyMg1sF6skxQD43QIQ06YEYfODUcPyPzNOqpW0lQ33CmcgDi8uXL7qw2GtGZqN75gxva9L9yir3dIaIfg7l9+/ZZWfZoFtjW1mY7QHHvaLtAXMcNxXaAYg9b2bHVfQK7c9qJ1Q4O8VxzFDxHO9fEk5Y07MqBdhlJx3v17nCXN/bylZ2/uJJ1y6PnkoegdIDMxJGOfEA+8pOHvMimvErDketQLvL5ciWRVz7ik9IRl7Xr6tstDcqPsiKt9OEdgvf47/J1JPhpG00eHVQ2yEh66iR50AuIo4yJIx12FMrbE116UreqPTOM51nkVRuQvASus+Cd/vPQPbRDNSQ7aXmOL7/qtMDGkkvpe4s8Oqi8kZE0yEN6lTFxSX4I/LpSTRfiwvLTO6yFUzi8XA+h4OQggMx6CUIgDPfJI8PyDO77BY3gGFuC8ww5RIGSejZH3kXgWiCsr5gKlDQEOWwVrAztG0cG4F6oT6gv+A2OgM56fhZ55SNdpSDiIP2zUHoCz64GMiA37yOQHr0E92msQk5B8vUmWTogA3UF+bALtkM+ygt0Lntx368joDLrCXnLjvOksgrjSe83XmSs9qlOEqT3n+fXF9kiCeRFD9Vrzv22x33k8sHO6N/b5NEBWamrpEFO9EAH1V/OKRvsi8yk86HekKYaoV2hIpe7Lmkx5EyE71AEFUCNBrIqUjNIcnZAQ1G8HFnY4f6rRt6bJDXeLCjTsGyxhd8h8txmdYh5oeyos6FMyK14DV7CjjWrQ6xmV+LL3z63KHxzyact4tatW92uWeeKK1a39Tzy8EuOIvHo0aNo6dKl7uoPrO+yBsa6LL842rlzp60dCtap44YejR071sX8f3j8+HG3tVrKmrVVf1cX2/nrieSJHUK3tdRm8/z5c6uzoUzUU9Z4hw4daptokyZNik6fPu3udkG9GT9+vLvKB3sL7GeUTrGFoSEAmzpsplF5cCY4DBa6WeSmIujna2yujRgxwu4rrpmwWI5M06dPdzF/4AuGeCTBTCc6ePCg/QzT37SicodO//8G9sAhbt++vVLWbD5p44fNTn3+xT1sRZ5aPtPpDXBsfLsYQtnHo0Arez5jY8Pu/fv37u6fDjHrC5kQOgr7cN/GjCUtB+tVTDEITDeYKsTlXplScJ9rpguannAvjGsmmj4loSkU8mptyYfpUzil6otUm+algS2YLqu8w7ImYDt/So2tSRNOs5sFMiJPkizEST9/nVGobqSRZtfyv/mVtCR8WsIf/fCn1H0RPjjnDybk/elnSdePTfQheDWq2zWK/gM3NSSZG9Gm4QAAAABJRU5ErkJggg==\"\u003e\u003c/p\u003e\n\u003cp\u003erespectively, where P represents the physiochemical score of an amino acid. Amino acid charge was calculated at pH 7. Hydropathy was estimated using the Gravy scale\u003csup\u003e40\u003c/sup\u003e, for which more positive scores indicate greater hydrophobicity. BLOSUM62 scores were calculated using the Biostrings v2.70.2 package in R.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eChanges in RBD binding (\u0026Delta;log\u003csub\u003e10\u003c/sub\u003e\u003cem\u003eK\u003csub\u003eD,app\u003c/sub\u003e\u003c/em\u003e) and expression (\u0026Delta;log\u003csub\u003e10\u003c/sub\u003eMFI) were obtained directly from the raw data of previous DMS experiments (https://github.com/jbloomlab/SARS-CoV-2-RBD_DMS_variants/blob/main/results/final_variant_scores/final_variant_scores.csv)\u003csup\u003e4,17\u003c/sup\u003e, and represent the log-ratio of the apparent dissociation constant (\u003cem\u003eK\u003csub\u003eD,app\u003c/sub\u003e\u003c/em\u003e) and mean fluorescence intensity (MFI) of the SARS-CoV-2 RBD with a SAV relative to the wildtype, respectively. Mean antibody escape was also calculated from previous experiments (https://github.com/jbloomlab/SARS-CoV-2-RBD_MAP_Crowe_antibodies/blob/master/results/supp_data/MAP_paper_antibodies_raw_data.csv)\u003csup\u003e16\u003c/sup\u003e. This score represents the escape fraction of the RBD with an SAV relative to the wildtype estimated under global epistasis models\u003csup\u003e16\u003c/sup\u003e, which we averaged across the 10 human monoclonal antibodies assayed. These predictors were assigned a value of -100 for SAVs that did not have any associated DMS estimates, which was the case for all non-RBD mutations. This predictor therefore also provides models with information on whether an SAV is found within the RBD or not.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eGenetic linkage\u003c/h2\u003e\n\u003cp\u003eLinkage at the intrahost level was calculated as the Pearson\u0026rsquo;s correlation coefficient, \u003cem\u003er\u003c/em\u003e, of the intrahost frequencies of two SAVs across all libraries within a dataset. Pearson\u0026rsquo;s \u003cem\u003er\u003c/em\u003e was only calculated for SAV pairs where there were at least five libraries with non-zero SAV frequencies. Linkage at the interhost level was estimated using the D\u0026rsquo; statistic, as follows:\u003c/p\u003e\n\u003cp\u003e\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAgQAAACuCAYAAABJGhU3AAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAACGmSURBVHhe7d3rixRX/sfx4+95NGoe7YMQvEDEDQavwZiAghdCWCLrLbIPhASNrs9kNTu6hICbRPNgQdDEsAEfZFfdGJQFjRdQUCPeMagYiCOyiI/GW/wD5tfvM/XVM2equqt7+lbdnxfU1NTp6rqcqjrnW6cuPaK/xImIiEhX+7+kLyIiIl1MAYGIiIgoIBAREREFBCIiIlKigEBEREQUEIiIiIgCAhERESlRQCAiIiIKCEREREQBgYiIiJQoIBAREREFBCIiIqKAQEREREoUEIiIiIgCAhEREVFAICIiIiUKCEREREQBgYiIiCggEBERkRIFBCIiIqKAQERERBQQiIiISIkCAhEREVFAICIiIgoIREREpEQBgYiIiCggEBEREQUEIiIiUqKAQERERBQQiIiIiAICERERKVFAICIiIgoIRERERAGBiIiIlCggEBEREQUEIiIiooBAREREShQQiIiIiAICERERUUAgIiIiJQoIRERERAGBiIiIKCAQERGREgUEIiIiooBAREREFBCIdK3t27e7ESNGDOrGjBnj1q1b5+7cuZOMJSLdQgGBSJfauHGjmz59uv+/v7/fd9999527dOmST//555/9ZyLSHRQQiHSxsWPHuoULFyZDzr3zzjtu//797tGjR+7zzz9PUkWkGyggEOlix44dc/PmzUuGBowbN84HCQQGItI9FBCIdCm7JDB79mzfF5HupoBApEudP3/e9ydNmuT7ItLdFBCIdKlTp075mwe5jyB28eLFQfcWiEjnU0Ag0qWOHz/uFixYkAw9d/bsWX9T4eLFi5MUEekGCghEuhDvGaDSf+2115KU57Zu3epGjx7tlixZkqSISDdQQCDShbgkgJkzZ/q+2b17t3/yYNeuXamXEkSkcykgEOlCp0+f9q0APGIILhOsWLHCffTRR27v3r3+fxHpLiP6eT2ZiHSNffv2uffffz8ZGjB+/Hh/P8HKlSvdnDlzklQR6SYKCESGibNpXuKjQ0lEikyXDEQiixYtGvKjPxMmTPA/BvTgwYNkrOcuX77s1q5dmwyJiBSTAgKRyL/+9S/f5zl8zvr7+vr8nfdffPGFDxbCoIC79Xt7e927776bpIiIFJMCApGI3V1v7/hnmMsC3HlPa8CBAwd8On755Rd//Z0fBRIRKTIFBCIR7rhH/I5/u/P+4MGDvo8bN2641atXJ0MiIsWlgEAkcvPmTd/Pc7f91atX9QIfEekICghEIvaO/zx4hM+e5S+HVof4RsWsjvsURESaTQGBSCTrHf/2c8F2b0E1aG3gBsU83dGjR5NvlZcWTLRTJyLFooBAJEClzzv+33zzzSTlOV7pi3b5FcC0YKKdOhEpFgUEIoFbt275/quvvur7hkcNeeyQYGDKlClJan66ZCAi7U5vKhQJrFu3zl8yuH37dpIygCcMSOf+gloCAhGRdqcWApEAlb7dUEirwJEjR9yMGTMUDIhIx1MLgUiC1oGvvvoqGRpAcMANhh988EGupwlERIpKAYFIB+DegxjBzIYNG/RTxiKSiy4ZiHSAw4cP+/62bdv8Hf78vsLSpUv9zxzT8iEiUokCApEOMHLkSN+31y1zeWPjxo3+Vxi5DGLvUBARyaKAQKQDnDt3zvfj1y2vXLnS9+0dCiIiWRQQiHQAflOhXV6YJCLFpIBApAPwWOS0adOSIRGR6ikgECm4cq9btl9ujH/KWUQkpoBApODOnz/v+7NmzfL90MGDB93o0aNz/ZSziHQ3BQQiBUcLwfjx493YsWOTlAH8fgI3E3788cdJiohINgUEIgXH/QPxzzXz2uVVq1b5lxPxlkXs27fPjRkzxm3evNkHERMmTPAvNOL1zLt37/af0d25c8ePD8bjx5YYj/HpCDTAeJbG/GyaehGSSDEpIBApMCplXkJkv7HAMBX/xIkT/aWCo0eP+pYDKuwnT564Xbt2uf3797v//Oc/7sKFC/7JhE8++cR/99dff/X3Ity/f98PM625c+e6xYsX+5cd7dixwz18+PDZ5YeLFy/6aeDEiRPu888/dz/88IObOnWqTxORYtGri0UKigqbtxDG7xigkqcSX7JkyZDLCLQO0KJw6dIlP8zZ/7x58/xLjAgaXnrpJdfX1+e/Z284JIgAgcahQ4d8P8Q0CTIIDuL5iUhxqIVApKB4GyEtAMT0YUfamjVrUitngoEPP/wwGRp4YRGBAzjLJ5iw71Hxv/322/5/7Nmzx7cYxK5cueJbIxQMiBSbAgKRLkELwOXLl90bb7zhh7l3gIrcfsWRs39aCwgE7D4BLjPwPVoBuEQwefJkt337dv9dxiN9y5Ytfrq0WHAvgogUkwICkS5Bkz4BgN1v8NNPPw26GZEnFTZt2uSuX7/u7xPgnoCPPvrIP864bNkyN3PmTPeHP/zBvfzyy+7GjRv+kgLvPmDcnp4e//1Ro0YlUxORotE9BCIiIqIWAhEREVFAICIiIiUKCEREREQBgYiIiCggEBERkRIFBCIiIqKAQERERBQQiIiISIkCAhEREVFAICIiIgoIREREpEQBgYiIiCggEBEREQUEIiIiUqKAQArv7NmzbsSIEUO6FStW+M9ERKQyBQRSeHPmzHE9PT3+/zNnzrj+/n537do1P/zWW2+5ffv2+f9FRCTbiFLh2Z/8L1JY27dvd5s2bfLBQGjChAnu4cOHvhMRkWxqIZCOcPLkSbdw4cJk6LnVq1e7R48eteTSAfMkIBkzZoz//8GDB27RokX+cgZ9htvV7t27/XKz/Cznzz//7P9n2Tdv3pyMVVmR80Ck2yggkI5w8eJFN2/evGSo9e7cueP+/e9/uwsXLviA5Mcff3R//vOf3a5du/xljWPHjrkDBw4kY7cXu8Ty3//+1/X29vrlJEBgXbZt2+Y+++wzHyBUUuQ8EOlGCgik8KicqHB+//vfJynV4czVbkSs1HFWm8e4ceN8xTd27Fg/vH//frdz506fzj0PePLkie+3G27GXLNmTTLk3KlTp56ty+zZs33a06dPfb+cIueBSDdSQCCFd/78ed+fNWuW74euXr3q+1YBpeEz7j3I0x09ejT5Vj52Jr1169ZnFWNR3Lx50/dpERiOIueBSDdRQCCFxxns9OnTh1Q2XJ8+fvy4W758eZLSfLdu3fJ9zroNTekYNWqU7zeCXadP6/K2clCRc18GZ/Tm3r17vv/CCy/4fh6tygMRqY4CAim8y5cvuxkzZiRDz3377bf+UsL69euTlHSNuGRgTp8+PeRmxxMnTvj+/PnzfT905MgRP5/4Gn28HNykx5MVWWjJSGvhoMvbykEwFd+XcejQITd+/Hg3ZcqUJOU5bjZkuWJ58iBtG3AzIvkhIk1SKiBECqu3t5fnDPv37t2bpAy4du2aT1+7dm2S0hrTp0/vL1WGydCAUoU6JA19fX3+M7ozZ84kqQO2bdvmp2W+/vprv36NwrIwfeZrLK/DNMPysnxpy5Q3Dxju6elJhvr7ly9fnppPItIYaiGQQuPpAkyaNMn3ObPmzPn11193pWDA39TWSrRe3L5921++oFu3bp1/J0LacnEH/l/+8hd/Zvzbb78lqQO4F+LDDz/0/9Pczno28lKINfPbPRjMk/lxaWbjxo0+zbBeq1atcv/85z+fDYfy5gHbctmyZf5/1o8nHBYvXuyHRaTxFBBIYdHMvGXLFv8/AQDNzHPnzvWV2N69e1seDLB84F0IEydOdC+99JKvCKkgw+vy4LG+x48fP7u7/8aNG75vaL7/6KOP/DrSZI9GvoHx3LlzvvLnEgDz5P8FCxakXm4gkGEd7TKCBRPImwcEAFzese1If8OGDYOedhCRBktaCkSkzmhaHz16dDKUjcsbjEcfNJOHzfJ2+cPQPM8wTfiNwjLQZF8Jl2po/ucSA1iu8HJH3jzgEkh4eYDvMV0RaR61EIg0CC0VM2fOTIaycSkgPDvmhT13795NPnV+uFRZJkPPL4/cv3/f9xuB5vupU6cmQ+m4jEDzP037nPmz7LAnEZA3Dw4ePDjoBkbeKcF0RaR5FBCINAjN4tOmTUuG0lGhEgyUzrBpAvAdP9R06dKlZAznvv/++2eVJZXw3/72N1c6634WGNQb86j0oifuBeCeAi4j2HLTcWnh+vXryVj58oBpEfTYS4+4fPDJJ5+kvopaRBqodBCLSJ3ZXfo0l8dPDBiaxRmHzp6SsMsBdHwejkPH9GjKt8sLjXD48GE/L54OyLosQfO+LZONEy4r65EnDxBOi45LBTwdYpchRKQ59GuHIiIioksGIiIiooBAREREShQQiIiIiAICERERUUAgIiIiJQoIRERERAGBiIiIKCAQERGREgUEIiIiooBAREREFBCIiIhIiQICERERUUAgIiIiCgikSvxW/YgRI9y+ffuSlNps377dLVq0KBlqnGbNp1lYnxUrVvjtMFydlDf1zJdymjWfGPOcMGFCMjTU2bNn3YwZM/yxSbd79+7kE5H8FBB0ICscKLzq7datW74/adIk36/WgwcP/LI9efLEHT16NEmtv0rzoUCnMiSvimTjxo1u/fr1bu7cuTUHZZ2YN/XIlzyaNZ/Y5cuX3fTp05OhwdhOb731ltuwYYPj1+wZb/LkycmnQ1EuWOBg3ZgxY9y6devcnTt3krGah+Vh/iwH++WRI0eST6TpSjuQdJC9e/f2s1m//vrrJKW9LFy4sL+npycZapys+fT19fWvXbvW5xHdmTNnkk+K5dq1a3756Verk/NmOPlSjWbNJw+2ZykISIbyYXyW3xw+fNinjR49uuI6Ubawn/T29iYptWM/tHmG+1875Gs3UkDQQayQatdggGCFg58Dv5HKzWf58uW+oqMjr4oaEIDClMqgGt2QN7XkSy2aNZ9ybFtt27YtScmH5Y6XnQqeabEflMO+Qxkzfvx4X4HXup8wnbi8smVgutJ8Cgg6CAcRB2kWDjYKATsIOSAp1BjmewQU4Tj0w4qDNDq+g/D7VohYhUPHWUeIM5C0Az1eLsTLVY2s+YSaWelVyif+r4UVntXkT7vlTb23PdLyhXVheuQ3/7NNsvbzvGrJ/2pR0TMPungZbfnDLm9gkDWuTTMv9l3yle9Vux8zPvOK9zP2UaYpzaeAoINQ2JUr7KmMrACmszNCCjQOTApi0hmHyjw+WC3NDnzG53MrtEi3Qp2DmmkZKzzjIAHxcrEOVmizTlaJ5lFuPiGmH69fo8T5RB5ZPrG+rGOtKDjzVgLtmDf13PahMF+YPtNlmqwX28Pma+tq26Na1eR/rcpVkFapVhPQ2PGetn3ZBnxWLZaD77Kc/J9neex4YDuEal0GGT7dVNghuBno0aNH7pVXXklShuImqHHjxvn/L1686G+OmjNnjpsyZYpP++qrr9yuXbv8OCNHjvRpIUubOXOm7//973/330ep8Hb/+9//3Jo1a/zw2LFjfd/cv3/f99OmGy/XypUr/XSZBvN6/Pix/yyPcvNplTifuJnP8mnevHl+u9WKO89PnjyZDJXXjnlTz20fCvOF6bNf2z65f/9+t3PnTp9u24VtUotq8r9WLPeCBQuSocFOnz7tSpXwkOOtnPPnz/t+rTcGp+EpCG5Q3bNnj+8mTpzobxbMc5OibX9pPQUEHcIK+zwoeDmArTC0g/bjjz9+dnDeu3fP98NC49y5c75Ciw/gq1ev+v4HH3zg+2Ae06ZNS4acu3nzpu/bPNPEy1WLPPOpFnfcx3dlZ3Xl7swnn7ibmjvVW6EReZOWB1ldOfXY9pXYo4Jbt26tqgJttWPHjj0L2mOXLl3KfPogy6lTp/x30vKA7VA6Q0+Gqsf2IzCgLNm0aZM7cOBA8okUgQKCLkOhyBnpu+++m6Q498svv/j+kiVLfB+cecSFBhWatQ6Ejh8/7tauXftsXGutePPNN/0wKp2B2Xc4Qwzdvn3bvfjii8nQc4xP5Ro/b13rmV45FHD9A5fXKnblKjQeHVu9enUyNODu3bs+yDJpwQePg/Go4HA1Im/S8iCry1LttgePpxFAVMMemQ2/Z8HwqFGjfJ+ALs5/WgFa9SicBTFpjxGyT7BPTZ06NUnJh+M1rcWBdWc7LF68OEmpHtNgH/7iiy/ctm3bBpUpWWwdpfUUEHQZKxRfffVV38dPP/005Mw/rdAgjSbukBXmr732WpIycJaBWbNm+T6swM1i3wkrVAqK3t5e99577yUpz1FJ0lQaV3KV5tMq5BPrMnv27CRlAE3mYQW1ZcsW3+/r6/OV6LVr1/ylHNtuw9GueVPttrf3a1R7OYEgNz77PXHihO/Pnz/f91kGxunp6XkWyBAY79ixw3/ebNa8nxZo2j4R71PlpB2vhpYTyoE8lXiM/ZhAYNWqVb779ddffUtYucsBtj8+ffrU90PVtnpIfSgg6BC/+93vfN+a77Ncv359SOV/5cqVQWf+VnmFhYa1LMSFj7UuhJcW0loX7Awnq0md5Ypx9s+yWmFtSOfMcenSpUMCgkrzaRXLpxCFaHxmTLP+8uXLfd5xBkiFQOBTj+u97Zo31Wx79sNvvvnGffrpp1W3mtC8Hvvyyy99ABAeDwQoy5Yt8/9bYDKcs+bhoHk/qwmfS3go1yoVs+Arbukjv7k0Ed5rkQf7MC0o3DdAIECrDgFunmm88cYbvv/jjz/6PtimLEfWPRPSYKUIWDoEm7NUESdD6fg8vnOb74V3Stud19yNzh3DdJZGnzuySQN3bJcKbv+/KRVgfj7caWzzKhWq/vv2vRjf4XN7hIt52DKE+Jz5MW2Wme+FKs3HMF3Ga/Qd4sbuqLa72clH1oP8C5FfjGed5WMWphGug80nTbvmTd5tTz5wFzt5Z/tjljhfwPh8n+nQ8eQB45Evxu7AD7ty+ZU2n3oib+hYxvi4ZbjS8R6zdTbko+1zlfYLQ96xjchLeyqkVqwby2Pb3pav3D4vjdMRAUF8ANNxoOTdwTuFFazl8HlYgFkBGBa+HIzkX1jYkUYBEBeAVmCFyHemSbod6GCaHPBpGJ/KkXnwP+PGBY0tl1WqLEc8b5SbD/Pgc+ZhHcNxxVxvLCcFL33myXraeoTIX1tvKgHGy9qPbduFeUyekJalHfOG+VTa9mC5rVLkc8ZNk5YvNj75Qx7zP9MKgwGwTcJ9ivFZrjRp86k3C47iYwmsR9a2TGPHZdixbkwjLb+zsEx8J867WnBMMy1bnrT1lObpiIAgPqNhR7WCsZoDpugoVFnnag7uZqJAohCLo38rrCsVBLZ+YZdWWGfNp5VY1rQAIJRWyVE52n4dY98OKy8wXK4Cb7e8ybvt7RiPu7RKKS1fLBCohO+F+W3zTZM2n2ZhvVmuSvuUSDU6IiCwQiWuCC3y7JaI084A4nxoJ2kVlp0FlWPrFm5LK6zTKrdKFWMzscx5tgvLaxUM62T5kvY9m2aYH+RRnnVup7zJs+1ZRyrzsPIjf/he2LKFtHyBtc6UY9O0/GYaBGRp38uaT7PY8VCPs3QR0xEBgbUGxDiwSc86w+o0tr6VKp5WotCNz3oJ3EjLYutFx7jG0tIK7LT5tIoV3ixTlnAdw/WKKzwwLhUk061FO+VNpW0Py4+wNYi8sXRTLl/4bqUgKJymzY/li7fbcPN/uFgeWzaReuqIgCAr+rdCVgFB+2Gb2DajQmC5G9H8Gc6nVaiIWD+7/j0crA/TqceZaTvkTb22fbl8oQJlHlTiwz026pn/tbBjPC1QERmuEfwp7WCFNmbMGP9iHF4RG+LxKn4nvHQQt+ztcDFedJJXtZtm8+bN/pn1hw8fJikiIiL5FP49BPZ8fPhWPGOvaq3mxR2NRiWft6sGwY/9FoGIiEi1Ch8Q2Ju8wrfimYMHD/qXm1Tz4o4ispYQ3h9e7etcRURE0BEtBONTfu2LSpI3XlFJtpP4PenlurwIeM6cOePfH86bw0RERKpV+HsIeG0mr7kMm8p5/SUtBrQO8MM01byKs8h4xzu/MNbX19c16ywiIvVR6BYCe+e+/TQow5wh81vccTBAOjcfcuMdrQoEEpyF8ytmvMebz+iYhmE8+/U5xqez98AznqURgNg0W9lkb/dK1OOHcJqJvCPfOqF1g32BH15iP+sE7O/8smDeX/vrpG2JSuvPutbr1yjzoryiTArLqpAts7U0Mr5ILrQQFBEv5OCRKVYh7EjjEabwkRz+J43nhu15ZNIYl8eebHy+b48lMX0eU7LHoXgenGHDtPgO0+N/exSJx5JaxR5JGu6jVVkaMX3bJq16jKueWAfbH9KwD/F50XAscJxUeu690ra0R/aKptL6U0ZQNjRrH2Y5wrIoZMeo7YMsd7nj1caPO7ZTPY/zvOwYYRnot7I87UaFDQhqQSDAAWIICGyHs4CAPjjowgLAKv0Y02THte+1EgUT69CoA9kKj3pNn+lQsFHgFh3bn3VJe5EQFUUYvBaRBb8WIMfKbcuwkCcfiqjS+pPeDuWAneRUw96TYcc1+ytlHWlZwW01yBPK2UoBJftJOE8rz7LyXOqvqwICDpRw52JnswKMnTAsrCjcwoOBz9J2zFoOwEbhgAvXqd1RgHbKGQB5n1bZsS0oXOmzrmyfoqLA5rhIq/SytiXHEOl8h/xJy6OiKLf+oBxo5f5sAXu1y5C1X7JNWd9asc9zXDAdgo5K5RL5F5908V06aY6uCQg4iNnprVnPDm7DjsiBQQFmZzsEAHyPnZlh0hmH7zIe6XYQsrO3OpLlgKp0AHNwWaHBetgwBTXrmpYG8o1pk25nEuQBw3SMZwEJ44XBVBrmw7hxIcEw8+Uzy087g2G5qmmW5fssC99j+fiurRvTrBemzTQrrXNWwdsI9czHEN+N9/OsbRljeeiagX2UZbXjlm1k+UHf9utqpa2/YftXOv6Gw8oaOvI8ZOsWduxveWRtF9tf7XjPi/EpC8gLppEnr9l30pbZ9tlK+5bUR9cEBBxA4cHKjhZGo7bjWUXBQc+wFZ4cMHyfg56dlv/toLTvVqoQGo1lSDuwQ1Z5ha0lLLd9N04LCz8r+MMCwtadAoCDNix4y2H8tJYVO5tmGnSMZwU6eR6fQWRh+Vl2K0T5n2kxHSvoaq0UY5ZXlQo+m28z1CsfY0yH6YWytmXMlqXRWG/b1uS3Heukh/tDLdLW3zB9ph0eH/VmxxvziuXdD2PsD3FFDNtf864P8ydvLGiqZjlsuzCNULXLIMPTNQFBp6Ny48CxgKYcxqNgC5EWF+rxeHbQhgenHbBhAUVhEI8Xq1Q58BkFVTiNSt9JY8scVoBp6zEceSv6vOPVU73y0aStQ97pDWe+tWJZqaDCyok01qMWlbbhcKadB9NmfdJwrGZ9lsXKjbjFAbau5Y4T8pVKnPnSxRV6XlnHZJ5lkPop/IuJZMDTp099f9SoUb5fySuvvJL891zauwuyHm2KjRs3LvnPucmTJyf/Zbt48aKbN29eMjQUn/NI13DfMmmvr/7ss898vxGuXr3qShVdMlQf9shYnq6ceuVjOZW2ZbV4n0baeqZ1jJuFRyCxdevWjnkvx8mTJ10pcE+GBrt06VLmZ1nKvemV/Rrl9h0ecX7//ff9I9e3b9/2+5oUlwICaQl+fyILQQifr1y5MkkZQIHz4osvJkOD8dx1WmFEpUBlHQYs9+7d8/0XXnjB99MqIN4/YRVKJY8fP07+q59SsJ67y1JtPjI+7+Ko9rn1ctuyFvwQWdp6pnXlfrTM3scR7hcW4FrgzDP78bancsv73oVmI/iaOnVqMvQc70G4fPly6mflnDp1ygcRccDE9I4fP+6WL1+epKSbNGmS27t3r9+nyLfhvn/CAnhpDQUEHcIqt6IYPXp08t9QFHoIz0yonHt7e917772XpDxnZ4lpFTOFWnz2eujQIf+6a3uhlVUqhw8f9pVMX1+fL+B49XUeWUFKq1Wbj7xgh3x58uRJkpJPuW3ZSqdPnx7ScnPixAnfnz9/vu+TN4zT09PzLMiggtyxY4f/vJ2w7Qi+0n6szYKfan/IjSCCYDr27bff+nmtX78+SUlHIEHAxfFCS8yePXt8YEBQWc3Lmqz8ytr3CDyk8RQQdAir3KotzFtl5syZvvkzzfXr15P/nqOAoeKxgtxQSH7zzTfu008/HVIAMUzlF+IMcf/+/W716tVJysA0mPY777zjhy9cuOD7eQtXzsryBg/NVE0+kk5gs3Tp0qr3oXLbspVoQo99+eWXQ1qMCJyWLVvm/7eAafHixX64nVjzfloT/rlz53y/XPN+jGOBdX377beTlAHkAa9AX7t2bVXTIzDg7bAEBUyDN8YSrOcJDCi/2C+///77JGUAw2ktGNIYCgg6CM17VHbl2MF59+5d34elhQduWtpvv/3m+9bkDqs8rCkW9rmNn4YCmbOKNFeuXPF9ChVQWfHTzt99992ggoFl++Mf/+gLoJEjR/qznZCdNdm1UJaRPKKACZuaKWg5G7Im4z/96U/uhx9+GFQY2mWFNC+//LLvh3mQxvLKXn/daHnzkc//+te/up07d/ph+16atPslym3LENuLyjfcpxqJ/YHlYn50tIA8fPhw0O+esO5s+9dff91vX/obNmxwa9asScYYrNz9IpbPjfq5ddt/bF3C/Y3lqvb+AWtBsrNvlp/9nDwgGKj1p9Q5bvgu+U85w/0JvMq70vHBD9HxHbvsQJ/hf/zjH35YmqBfOobdqZt2x7DhTmDGobNHr7gD39LsqYI4jbuRSxH8szSGuaPYhvkMdteypTGcxh5hLJ2hJCnPkc7TEraspYIu9S5jlovlhK17iDuU+S7j2fIw3fhxKKZhd4bzGcOWD8budk7Dd/gs6w7rUgUyKN/pSKv1juy8mE+lfGTZSbd9gfVk2bKQh/Eje+W2Jduf6YX7Dv+TlrVv1IPtD6yPzZvtGi8j6xKuL+OTX1nS1t+QzueNwjoxfZYvPsZJj/fZcphWvE8yDfKo3vsl+xj5mmf5wu2Vtp7SWAoIOgwHMwdUoyubeuCApwAIWUFeqbKwSijuwgKfgt4ChnL4XlhRUonGlSLDpGehsIu/00p585F1CvOPLqtCtH0rDqiQti1bySqWSthm4XLbfpWm3PqjVXnAPs8yZwUqInkpIOhAFGqc9bVTAZ3GznjCSpxCLatANlRyfC8sAO0sPTyjYJxKeRBXAAzH36MiKBcMgPnzvXD+rZQnH1kvxgmDBsuPuNJjmAovq9JJ25atRCBYKUCzfcaCQfKB4ybte5XWn3Q+j/OtGWw7tkveS3EpIJCWojCjELZKiTNthsuh8KOjADYU4pYOO2uqVEHbd6wbTiDFOrBM7RAUVMpHa0GgC5tyLS2sFMlLplWpyTfelq3EdqgUxIX7DB3fYR3jSr3S+hMMEAy1Yr0tUKm0bUTyUEAgLUdBatcuKXgpnLPOxPKyM12m18wzJwpoCudKlVGj1SsfCRyYVt4gJ9yWrcI2YN2ppO3sv1aV1p91TQsimsGCulbNXzrPCP6UdioRERHpYnrsUERERBQQiIiIiAICERERcc79P+XI1XAXYGBBAAAAAElFTkSuQmCC\"\u003e\u003c/p\u003e\n\u003cp\u003ewhere \u003cem\u003ep\u003csub\u003eA\u003c/sub\u003e\u003c/em\u003e and \u003cem\u003ep\u003csub\u003eB\u003c/sub\u003e\u003c/em\u003e, represent the proportion of consensus genomes carrying an SAV at locus A and locus B, respectively, and \u003cem\u003ep\u003csub\u003eAB\u003c/sub\u003e\u003c/em\u003e represents the proportion carrying both SAVs at locus A and B. For these calculations, we only considered genome entries before the timeframe of each dataset, with a one month buffer (as with the calculation of prior fitness). For example, for the Alpha dataset (February 2021) we consider only genomes collected between December 2019 and December 2020. The \u003cem\u003eD\u0026rsquo;\u0026nbsp;\u003c/em\u003estatistic was only calculated for SAVs with a prior fitness of \u0026gt;1000 genomes.\u003c/p\u003e\n\u003cp\u003eThree types of linkage predictors were generated based on the intrahost and interhost linkage data for each SAV: whether a SAV is strongly linked to at least one other SAV (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u0026gt;0.9) is a binary variable; the number of strongly linked SAVs (\u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u003c/em\u003e\u0026gt;0.9) is an integer variable ranging from zero to \u003cem\u003en\u003c/em\u003e \u0026ndash; 1, where \u003cem\u003en\u003c/em\u003e is the total number of SAVs considered in the dataset; maximum \u003cem\u003eD\u0026rsquo;\u0026nbsp;\u003c/em\u003eor \u003cem\u003er\u003c/em\u003e across all pairs involving the SAV is a continuous variable ranging from zero to 1 (negative \u003cem\u003eD\u0026rsquo;\u003c/em\u003e or \u003cem\u003er\u0026nbsp;\u003c/em\u003evalues are set to zero). All linkage predictors that could not be calculated were set to -1 to indicate missingness.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eModel training, hyperparameter optimisation and evaluation\u003c/h2\u003e\n\u003cp\u003eFor each dataset, we trained gradient-boosted regressors as part of the XGBoost v2.0.3\u003csup\u003e18\u003c/sup\u003e module in Python. We optimised the hyperparameters for our models and assessed their test error using a nested 10x10 cross-validation procedure. Briefly, for each iteration of the outer cross-validation loop, a tenth of the data is held out. The remaining nine-tenths of the data is used for the inner cross-validation loop, where an exhaustive search (i.e., grid search) for the best performing combination of hyperparameters is performed. As part of hyperparameter optimisation, 10-fold cross validation is performed to evaluate the performance of models using each combination of hyperparameters, and the best performing model is selected. The held-out tenth of the data in the outer loop is then used to assess the performance of the optimised model, and this process is then repeated for all 10 iterations of the outer loop. We optimised three hyperparameters for our models, particularly the number of trees in the ensemble (n_estimators=100, 300, 500, 700, 900), the maximum depth of each tree (max_depth=1, 2, 3, \u0026hellip;, 9), and the proportion of random predictors used to grow each tree in the ensemble (colsample_bytree=0.1, 0.2, \u0026hellip;, 1.0). Three main metrics were used for evaluating model performance, the coefficient of determination, \u003cem\u003er\u003csup\u003e2\u003c/sup\u003e\u003c/em\u003e, mean absolute error (MAE), and Spearman\u0026rsquo;s rank correlation coefficient, \u003cem\u003e\u0026rho;\u003c/em\u003e. These metrics are defined by:\u003c/p\u003e\n\u003cp\u003e\u003cimg src=\"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAxkAAAD2CAYAAAC6GUo1AAAAAXNSR0IArs4c6QAAAARnQU1BAACxjwv8YQUAAAAJcEhZcwAADsMAAA7DAcdvqGQAAEO3SURBVHhe7d37qx7Vvfjxyff3mibmJ09bxEQwRFE0F/FSMGBio0ilWmMtIihqtAgnbb0kKQfxEhOPHpA2F1GQgzF6jmKRxsYUIhgVTaIYqqQcjRSR/hTvf0C++70yn2TtlXmeZ/bez97Z+9nvF8yembXnmVmzZs1lzaw1M+PIkEqtrVy5snrxxRer5cuXV9u2batOPfXU+j+SJEmSYCFDkiRJUl/9v7ovSZIkSX1hIUOSJElSX1nIkCRJktRXFjIkSZIk9ZWFDEmSJEl9ZSFDkiRJUl9ZyJAkSZLUVxYyJEmSJPWVhQxJkiRJfWUhQ5IkSVJfWciQJEmS1FcWMiRJ09pbb71VrVy5spoxY0Y1e/bsau3atfV/JEmjNePIkHpYkqRph8LFnj17qksuuaTaunVrdccdd1Qffvhhde6559ZTSJJGyicZkjQAuDjmYrlNt3HjxvpXwx04cCDd0X/hhRfqkMmLdWBdFi1aVH355Zd1aDPW6c477+w4HffaKGDg2muvTf0f/OAHqS9JGh0LGZI0AG6//fZq4cKFaXjLli3pwjnvDh8+XK1atSr9vwkFi1/84hfV/fffny7KJysKCsRv9+7d1aFDh6rrrruuOvPMM1MBqRPWjacSvabDM888U61Zs6Y644wz6hBJ0mhYXUqSBgQX0Oedd141a9asav/+/SdcKHOBPmfOnGrDhg3VPffcU4cebZNw9dVXN/5msqGAcdNNN1UrVqyoQ46u9/r166tHHnmka/x52vPYY49V7733XnXqqafWocdRGHnzzTerTZs21SGSpNGykCFJA4RGy1xsL1++vNq5c2cdehwX6eeff/6wQsa8efOq2267bVjYoKJ6FU8/ynW1gCFJ/WV1KUkaIKtXr67mzp1bvf766+nOfYmL6fwC+7XXXkvVjqItQo5pedtS2e7hiiuuSO0h6E8kClAsl/YVuV5tTXK//e1vq0cffbQeO6osYJBuPN2RJI2ehQxJGiBUA3r22WfTMO0rPvvsszTcyV/+8pfUlqOsZkQVpL///e/Vq6++mqpR/e1vf6v/U1Xbtm1L/WuuuSb1JwIX/hdffHGq6rV58+Zh67Vjx47U5+lNL4sXL66+/vrrY4UICk8UWphnFFZ4u5QkaWwsZEjSgOFNSTTy5mK6vOtf4mK9qX0CDaUffvjhNC/aeFDgCN9++20Ka3r6MV5o2E47jChI/OMf/0h9fPfddym8zStnozD1zjvvpD7r/tVXX53QUD7eNiVJGh0LGZI0gB588MFUEKDaFNWBOtm7d2+1dOnSeqwZd//ff//9euzobyjENBVOOomnBG26bqIg8dFHH6U+qOp0991312OSpMnAQoYkDSAKAPfdd18anj9/fuo34WlHLxRCKFiEP//5z9Utt9xSjx2tWkXhoNN3KFA+KejW9cJTiw8++CANs0yexuRvm6JqFe1IJEknj4UMSRpAXHzTwJlvZnSrRsTTjl7OPvvsY4UR2jLwdqq8DQfzp3AwkicbY3HBBRdU33zzTRrmuxbr1q1Lw4GqVfv27avHJEkng4UMSRpAN954Y6rmxAV3N0zDh+26+bd/+7fUp4Dxxz/+cdhTjPjydps3O/XLT37yk1QNjIIUTzTy9hPx5ivfDiVJJ5eFDEkaMFQXonpTvAWqG55IfPrpp/VYs3gS8vzzz6cP4eVPLOJ1uBdddFHqT4QFCxak/hNPPJG+CZKLpxqdGm7HF78nMr6SNB1ZyJA05XFHmzvpfFSOu9jUx4+LyemG9eYVrM8991yr6ktXXXVV+k5Gr1fd8prbsu0D4ncT+TamaGMyc+bME169+8UXX3R9le27776bqoj59ihJGl8WMiRNedTLp9rMe++9Vx0+fDi1H1i/fn393+nl1ltvTW9+KgsDgQ/a5R/RYzo+3vfSSy/VIZ01PRnhiUmb71P0E6/QZZlNXyjnTVPd3pb12GOPHWsQL0kaPxYyJE15XGzymlbu3NNdf/316UvVJxtPWPhOxUS1V6AAQQGL19c24akDH50rL8L5eB+NxDs9zWC+DzzwQOOTkbioJ/0nqh0EaRpf5y7t2rUrVYUizqR/Lr6AnrcpkSSNDwsZkgYKF5Yvvvhi9bvf/a4OOTkoWJx55pnpon4icIFP+wSqPs2ZMydVGys7nlg0vbKWqkNctFM4i2pmK1euTPPkwpyG1p2ejFD16N57700f6xvPKkgUGqPQxle/y2pSgfhcffXV1TnnnDOsUMR68AX0l19+uVU1MknS2FjIkDQwuAjlrUpcSHa6CJ0I3EXnwpwvSU+Uhx56qB4aHQoVTz/9dKpmxlOJ/fv3p4t12j10e0MVXwXn9bX0xwtPWCgcLVmyJMWl2yt5eXUt6c76BIYpPH3yySddfytJ6p8ZQyeH3l8+kqSThLvpN998c7pwfPXVV6vTTjutWrNmTXpawZ33+Jp1FDC4yz2ZLiR5gkCcmtoPSJI0qHySIWnS4g42r02lQTd3svlGA9Vl/vSnP6WGv5dddlmarixgRLUaSZJ0cljIkDRpUeWJtgJRh54qPLzhiPGdO3ceq8bD26X4ONt5552XnhzQJqFTI+bAE5K8vUK3Ln8bkyRJ6s1ChqRJLxoj0+6gqdEuVZGo+Zl3FEK6oZFy+ZtOXa959UN8qbpN1/QWp6bp7MbeSZJGx0KGpEmPD6jh8ssvT/1BREGmqYDT1DW9xalpOruxd5Kk0bGQIWnS40kGX5zu56tHrS4lSdL4sZAhadLjA2uLFi2qx/pjslWXkiRpkFjIkDSp8aYoPjA3Fb9vQNyxe/fu1Nfkwbbhg4nz5s1LT6soxEbbH0nS2FnIkDSpHTx4MPUXLFiQ+lMB3+6gihVvuQJvvuJidqKqXfF167LKV6eOC+0mhE/lamJUh6Pg8Nprr9Uhw/FGsg8++CC9Hvnw4cPpFcl8iFCS1B8WMiRNalGtqamx82TFF6bLhtyffvrphFW74tW+tGHBli1bhsWDjovqVatWpf+XuMPPxfm33347qauJNRWYoqNwRH7hg43/8R//0fjNFN5IRmGQdj50fNhx9uzZ9X8lSWPlF78laQBR9YfvhsyaNSt9X4RvjuQoTPCkpfwaORfoF1xwQfXwww/XIVMb67lkyZLq97///bHvqpRiGtr+lOkkSRodn2RI0gCiDcuaNWtSNaCmO/lx9z7Hnf29e/dWq1evrkOmPtbzySefrO6///5UmCgRxtfiX375ZQsYktRHFjIkaUBRWJg7d25qE0I7jRKFivwpxuOPP56qepWvCuapCG1KqE6UN46m3UZUUZpII43PihUr0rQvvfRSHXJUFDB4mjMVXywgSZOZhQxJGlAUFp599tk0zJ38zz77LA034X9Uq7rqqqvqkKO4EKeAQlUi5IUVCii0/SifiIyn0cZn2bJl1SuvvFKPnVjAYLzpiY8kaXQsZEjSAKMBNI28O1WbCv/6179S/5RTTkn9QEFl06ZNqSoRF+r79u2r/3MU8/3Nb35Tj42/0cbn9NNPT090Am+XYpx2Kzz5oH1Kt0KYJGlkLGRI0oB78MEHUwNwLqqpItXk448/Tv1ub/E6//zz09OOwEU58x3Jm7/yKk29uk6v1w1jiQ9PPfI3btH50UVJ6h8LGZI04Lj7f99996Xh+fPnp36JV9b2ctFFF6V+tIOgjcMDDzyQhgNtHzp9mwJNF/edury9SJM28ZEknRwWMiRpwNHe4NFHH03fzOjUwHnmzJn1UGdRQPn+++/TPPmYHY2qc1999dUJYeOlTXwkSSeHhQxJGnA0cF68eHHH70QgvqjOl7I74YkI1ZHeeeed1KYhb/vA76jiNJFfCe8WH0nSyWUhQ5IGGG9f4tsX27Ztq0OanXbaaan/xRdfpH4nFFZ4YoC87QPDy5cvr5YuXVqHTIxO8SkxDfGTJE0MCxmSNKBoq3DHHXdUzz33XLrr3w1va+L1r2+++WYd0oyvgfP62FtuuaUOOe7TTz891k5ionSLT45prrnmmnpMkjTeLGRI0oC69dZb0+trO7VTWLt27bDqTb/97W/T26do39AJr4h99dVXTyi08GanQ4cOjehNU/3QKT65eKPWtddem/qSpPFnIUOSBhAFCC7AeX1tEwoFmzdvHla9ia99U/3oiSeeqEOG461RnV4RS5UsqiPx9CT/QN546hafQIFp3bp11fr163s+zZEk9c+MI7wnUJI0MGiEfemll9Zj3fHF6/xVsVyU83TjuuuuS+EUGHi9LdWgnn/++fQhvCZc8F955ZWpoEH7j/G6oG8bH1CQ4uvfixYt6jqdJKn/fJIhSQPmoYceqodGjsJBfEWbwgZPJu69997qr3/9a9cLdapkcc+KD9qN5xODtvGhoEUBg+9mWMCQpInnkwxJkiRJfeWTDEmSJEl9ZSFDkiRJUl9ZyJAkSZLUVxYyJEmSJPWVhQxJkiRJfWUhQ5IkSVJfWciQJEmS1FcWMiRJkiT1lYUMSZIkSX1lIUOSJElSX1nIkCRJktRXFjIkSZIk9ZWFDEkaEFu3bq1mzJjRqtu4cWP9q6q64oorho1PdsSVdVi0aFH15Zdf1qHdTbV1lKSpzkKGJA2I22+/vVq4cGEa3rJlS3XkyJFh3eHDh6tVq1al/09FFChWrlxZ7d69uzp06FB13XXXVWeeeWZ14MCBegpJ0mRhIUOSBsjTTz+d+vfff3/12WefpeFw6qmnVg8++GA9NvXcdddd1U033VTt3LmzOuOMM6p77rmneuONN6r169efsK6SpJPLQoYkDZBzzz23WrNmTfX1119Xd955Zx16HAWN66+/vh6bWl544YVqxYoV9dhRrC/hFDokSZOHhQxJGjCrV6+u5s6dW73++uupnUaJi3KeAnSzdu3a1O6hLKg0temYCJMtPpKk7ixkSNKA4WnFs88+m4abqk31QsHk4osvrjZs2FBt3rx52O937NiR+suXL0/9iTDZ4iNJ6s1ChiQNoEsuuSQ18u5UbaobGpBTLSku3P/xj3+kPr777rsUTjWliTLZ4iNJ6s1ChiQNKBp5z5o1K1WboorUSMWF+0cffZT6ePPNN6u77767HmsnqjS16brpV3wkSePPQoYkDSiqTd13331peP78+ak/Ujwl+OCDD9Iwr5ClqlLe+JqqTHyvopvyVbrdul76ER9J0vizkCFJA4qL8EcffTR9M2O01YkuuOCC6ptvvknDzzzzTLVu3bo0HKjKtG/fvnps/E22+EiSmlnIkKQBdeONN1aLFy9OF96j9ZOf/CRVt6LAwhME2noEvqJNFae33nqrDhl/ky0+kqRmFjIkaQBRbWjv3r3Vtm3b6pDRWbBgQeo/8cQT1SOPPJKGQzxFyC/0x9tki48kqZmFDEkaMAcOHKjuuOOO6rnnnkvtMsYi2nLMnDnzhA/effHFF8fe+DRRJlt8JEnNLGRI0oC59dZb0+try69jBz5sR9WiNr799tt04d708T7e7LR06dJ6bGJMtvhIkppZyJA0cPJXojbVz6c+/+zZs9P/V65cWYc24xsTTNck2gB06k5G2wAKEHwbg9fXNuFtTHzQru3FOOu/adOmemy4Xbt2VRdddFFaJmk6ESZbfCRJzSxkSBo48RVo8LG2EvX5uRDnGxLdvh9BIaHb/6O9A3fW89ewxvJH+9rY0SK+tFM4dOhQNWfOnMaCz9y5c9O6d0MBjIt0Luj5ynZZLSmQfldffXV1zjnnjLlaVjeTLT6SpN4sZEgaOKeccsqxuvn5h9vAnfwXX3yxWrhwYbVs2bI6tNm///u/H/vOBO0cSnEhWz4VWLJkSbqYn+gL3YceeqgeGj3Sh0II68Bbqbq9+pZXxX711Vc9nwaNxWSLjySpnRlH2nz9SJKmkI0bN6b+U089Vd12223D6u9Txemaa65JDaP5fkSn17vyBIM6/r/61a+qSy+9tNqzZ88Jby3iyUGn/0mSNJ35JEPSwNm9e3d19tlnV/PmzUvDIdpI/PjHP079eB1qiao5vA61U7uG8PHHH6d+FDBee+01vzYtSdIQCxmSBg4fa6N6zQ9/+MM65Kibb745FR7efvvtNN7p6QNtNqh61Ku60xtvvJH60d7hyiuv7FkFK/RqNJ53J6MBuSRJY2EhQ9JA4YI82kOcf/756YN04ON0FAAoWPAWok7fU6ANwPvvv3+sXn803n7nnXdSP8d81qxZc6zBN6+N/dnPflb/t7udO3cOayzerWtbFaupgGJ3YidJGn8WMiQNFAoD8TSBD7bRaJjqT4899liq/sTw/v37O77ClTcY8SQkLkh5S1MTGoIz74svvrgOqdKrVU9m24ymAordiZ0kafxZyJA0UGiD8dOf/jQNR5uLu+66KzUA5+nGwYMHUxhtNkq0qUCbi9KYz1lnnZX6I2V1KUnSILOQIWmg8BRi8eLFafgHP/hB6vPkIt4wFdWeaLNRuvvuu1ObjSZ8aTrHm6eollV+s4EPwVE1q5fxqC4lSdJkYSFD0sCIi3uqSSG+qfDkk0+mPlWleK1tE157y0fsygv6+D4G7TRytMfgWxs5ls/H8C688MI6RJKk6cnvZEgaCPHNCnDxz4fZSlRR4kkHaPjN0wRQOOC7GSi/nUF1pbB9+/bUIJx2G5s3b65Dh+Or03wQTpKk6cxChiRJkqS+srqUJEmSpL6ykCFJkiSpryxkSJL6gg8Zxqt5oxE+b9tifN68een/kqTpwUKGJKkv+Po5HySkUf0rr7ySChh8Af3w4cPpzV0vvfRSPaUkadBZyJAk9cULL7xw7Lshe/fuTQUMXgnMRxAlSdOLhQxJUl9RwOCDiPHNkfjWSNNX1iVJg8lChiSpbyhQfP311+nr6eHgwYOp3/SVdUnSYLKQIUnqmyhQnHXWWamPN998M30g0WpTkjR9WMiQJPXN3//+9/TV82ibgV27dlXLli2rxyRJ04GFDElS3zQVKHizFKhKxRunJEmDz0KGJKlv9u/fX51//vn12FGrVq2qHnnkkfTtjNWrV9ehkqRBNuPIkHpYkiRJksbMJxmSJEmS+spChiRJkqS+spAhSZIkqa8sZEiSJEnqKwsZkiRJkvrKQoYkSZKkvrKQIUmSJKmvLGRIkiRJ6isLGZIkSZL6ykKGJEmSpL6ykCFJkiSpryxkSJImvS+//LLauHFjNW/evGrGjBnVokWLqgMHDtT/lSRNNhYyJEmT3jPPPFN98MEH1XvvvVcdPny4+vrrr6v169fX/5UkTTYzjgyphyVJmhLWrl2bChqbNm2qQyRJk4mFDEnSlELVqSVLllS7du2qzjjjjDpUkjSZWF1KktTT1q1bU1uINh1tJ9p66623jv2OZfRCAePGG2+sXn75ZQsYkjSJWciQJFUvvPBCtXLlynSx/9prr9Whx91+++3VwoUL0/CWLVsqHoLnHe0kVq1alf4/Epdcckm1Y8eONLxgwYLU7yQKGBs2bKjOPffcOlSSNBlZyJCkaYwLdwoX69atq37+85+nwsKKFSvq/w739NNPp/79999fffbZZ2k4nHrqqdWDDz5Yj40OBY5OygIG43feeWf9X0nSZGMhQ5KmMS7cDx06lN7aRGGDwkInXNyvWbMmNbhuusDnt9dff3091t7bb7997ClJJ7xd6vXXX6/OO++89LRlzpw5JxR0JEmThw2/JWmaoorUDTfcUO3Zs6frU4QcTxBodE3BhGpTVKMaCX7/xBNPVJs3b06FFapY7du3r1q2bFn18MMP11NJkqY6n2RI0jRFFSmeILQtYICnFc8++2wabqo21csVV1yR3gr16quvprYc2L9/f3XOOeekYUnSYLCQIUnTEF/L5mkEX86m6tPs2bNTNSS+qN3U8DtHoYQnEJ2qTXXCW6coUNC2Iwo2p59+euovXrw49SVJg8FChiRNQ99//33qU1Xpqquuqr766qvU6JsnG7/+9a97PqGgkfesWbNSOwmqXbXx1FNPpTYb5ZuhmI+vo5WkwWIhQ5Kmsf/6r/869jYpqkJRBYonFHv37k1hnTDtfffdl4bnz5+f+t3Ek5PLLrusDjlq9+7dqT1GiUbofNVbkjQ1WciQJB0TTxk+//zz1O+EBtyPPvpoavzd5psV8eQk/xYGT0t4EnL++efXIcfxdMSG4JI0dVnIkKRp6LTTTkv9v/71r6kfeOKAs88+O/U74dW3tKMY6dulcv/5n/+Z+hdddFHqh/gCuCRp6rKQIUnTEG0g+OYFr5J96623UhhPFm699dbULqPTB/mwdevWVJ1q27ZtdUhvUah55513Up8G49HoG8wz2nbwwb3ly5enYUnS1GQhQ5KmqdWrV6e3RF199dXpyQGFC942tXPnznqKE/Gk44477qiee+65rh/uK1GoofBw7733pjdY0dj8lltuqebOnZuWD9ph4J///Ge1dOnSNCxJmpr8GJ8kqTUKIXSbNm2qQ4ajsfb777/ftaDSC4UQvsUxku93SJImF59kSNI44tsQcYd+qqMAwZuneH1tE6pbUf1qLE8hmAdvoeKNVb5dSpKmLgsZkjQO+KAdd+SpHvTNN9/UoVMX7TYeeeSRVACYM2fOscbZeUfVJwohY8V3M5YsWVL98pe/rEMkSVONhQxJ6jMaMH/00UfVe++9NzANmB966KF6aHzRdoMPA3766aetXo0rSZqcbJMhaSDxHYc//OEPqfoO+NL0n/70pxE1Vu6HK664IvXH0kZBkqSpxicZkgYOBYwzzzwzDR8+fDh1+/fvr+66664UJkmSxpeFDEkDhycYs2fPTm9A4skFHU8yXnzxxXqKZk3tDDp1kiSpMwsZkgYKTzGoIvX73/++Djlq5syZ9VBn1B5t2/UbDaubCjNNXVTBKjVNa3diJ0kafxYyJA0UGlvj8ssvT/2we/fu9PajyYpvQjQVZpq6Tu07mqa1O7GTJI0/CxmSBgpvdeIVqLylKPB04/XXX6+WLVtWhzRruuvdqZMkSZ1ZyJA0UHhisXjx4nrsqJdeein1f/e736V+J013vTt1kiSpMwsZkgbK3r170zcW+HI0+GbF/fffX23YsGHY042JwlMU4kR/0LBOfNGcjw7ydGfRokXVgQMH6v9KkqYzCxmSBgYXuHxxeuHChalqFBe+69atq5577rnqnnvuqacaf8SDxtm84YpX5xInXqlL2GS9CKew0KlBeSfPPPNM9cEHH6R2MLwmmPVcv359/d+Th0b0FHj46rok6eTwY3ySBgZPLW644QarM40ATyMoXFAoe/jhh+vQ0Vm7dm0qaPDq4IlC4YiCDts+x5MsXltMYWMi4yNJOsonGZIGxt///vf0FEPt3XjjjX0pYFBY4Tskvdq99AtPKaimde+991bffPNNHXocVeN4C9euXbuqrVu31qGSpIliIUPSwOCCkjvXaoe7/7QXWb16dR0yOhQwKKy8/PLLE9LuhXjzFjGqaS1fvrwOPREfYXzyySdTm5xBbBMjSZOZhQxJA4P2D+eee249Nv3QFoG7+7QFYTiqQtE2pam9xeOPP16tXLkyXYznuPPPPJgX86AdSTTupnpSLgoYNKyfqLQnzrSxKePdZMWKFWld4g1jkqSJYSFD0sCgLcbtt99ej00vtEF4/vnn09192kX89a9/rf7whz+k9gg7duxI3wmh4BGYnkLZVVddVYccFW0bXn311erQoUPp4vx//ud/0nzXrFmTqieFsoDB+J133ln/d/KgOtgrr7xSj0mSJoKFDEkaAFRTokARd/dpH/Hggw+m8FNOOSWF5f71r3+lfvk/nhLkBbU33ngjtddgvjNnzqxDj+LtUhRezjvvvPSUY86cOcdeHTyZnH766SmekqSJYyFDkgZIvCL3tttuO1bg+Pjjj1N//vz5qY8Iu+SSS1K/FP//zW9+k/rgLU55w3qqLPH0KO9obN0JT1IojLTpRvo6XUnS5GIhQ5IGyLvvvpv6t9xyS+qDggeFg7wNw7ffflsPNeMJBr/JCyFUr6Lq0Wgxr7JQ0qnrVliRJE1+FjIkaYBQoJg7d+6wAgVv3SoLB2XVpxIFivxNXVSDoo3GxRdfXIdIktSZhQxJGiD79u0bVqWpU+FgwYIFqZ83Bg804OY3+duieNUtlixZkvqSJHVjIUOSBgSFA55AXHbZZXXI8Qbe4M1R8fao0047LfW/+OKL1M/xJilceOGFqY/PP/+8HqrSG6QmSwNv1pkCEP1OaEvS7XsakqT+s5AhSQPi4MGDqR9PKUBjb55s/PrXv04FBd4eBd46Rfibb76ZxnN86A75kwwu0mfNmpWeZPDa24n46F4nVAmjYTjfv6BQxSt7zzzzzBQWDd9zVBe75ppr6jFJ0kSYcYQWdpKkaYenGjyV+OSTT4a14Rgk02EdJWky8kmGJE1TPNVYvHhx9cQTT9Qhg4UqVOvWravWr19vAUOSJpiFDEmaxrZt25aqE23cuLEOGQy0GaH6FG/Vmq5fgZekk8lChiRNY9zh541UGJQP4PHGrOuvv7564IEH0lfQJUkTzzYZkiRJkvrKJxmSJEmS+spChiRJkqS+spAhSZIkqa8sZEiSJEnqKwsZkiRJkvrKQoYkSZKkvrKQIUmSJKmvLGRIkiRJ6isLGZIkSZL6ykKGJEmSpL6ykCFJkiSpryxkSJIkSeorCxmSJEmS+qovhYwZM2ZUGzdurMek/rriiitSN1L8hrw5e/bsOkT99sILL1Rbt26tx5p99tln1cqVK+uxoybDMeOtt95K8aA/6Ehr1rWb1157rZo3b16aLtIlT58vv/wyzWc6pNd4K49ppPNI94c2+95ItMkjTMPxlOmIf/mbpn1dJ8+BAweOnQfpYp/O81q/85FGf80yiHyS0Wdx0F20aFE6KfdLHCSaOjNzs507d1YLFy6sFi9eXIeo32644Ybq22+/rceavfTSS9WLL75Yj2myuvvuu1N/x44d1Z49e9Jw7uDBg9W9995bj+lka7Pv9RMFCLY/x1Pyx4YNG+r/HOe+PrlQeHj99der7du3p202f/78+j/HTXQ+0vRiIaNPKFBwB2f37t3VoUOHquuuu64688wz052Efjhy5EjHjotpdbZ06dJ6SFInHLd4krFixYrqkksuSR3HF/oaX6TzPffcU49NTv/6179Sn+MpeeLcc89NcSbumpwoGIJrE7bZqaeeOiXymgbHsEIGd9/pcpSEuVueP14D0+WPRSkJr1279tijVP7P4/dcPG6PR/JMy2/yO/78nzvzsVymifnwqI9lEh7LmCyP7u+6667qpptuShf8Z5xxRtqJ33jjjWr9+vXHdvTJhPQjrePxaXQjxbbJH8eyve68887Gbcqy2GadpkNeZYNpx1JIY9779++vzj777Dpk9LrlS7Yv60IY/yP+5eNnfhu/j+ma9hHCeXwd6cRv0GbfYZiwfJqmNG6zHxFO3Drt05FvwN1Nhst5gPjH3W+mifVBv44ZnZDWZVo07Ysff/xxilc+3WjStW0+YH0i/ZmG6Vn3EmnBNOQHtE2LvIoE08bvO8m3JXc9GWY5EU6f8UsvvTRNQz+mJ5xlMU3k2U7pwzS98h3jEfeYpswTpGksi455jvQ4wXL4LWnTbdvHdPRju7LOaLO9QfxjmxHvprjm8w3ELX5Hx7JYZsQJ5b7XNk4jzSP59s+XSXjEhfk17ev0mS7fH+mX6wviUW7bcp/ttf3Zfm32114inVl+nm9jXvn2oV/mU37fK78zH+YX84np8nl1ike53iWmYX+O4dgeDJP2MV/k23Qky2uT32I++TTltmeZkR/pyjRoq21eI7wp/+XhndKh7fYH84r17rROLCeff1M+ITzfN2JbTglDpdpjNmzYwC2JI4cPH65Djhy5/vrrUxj9wP8J27JlSxpnmG758uVHduzYkbqFCxemsEOHDqVpwP9nzZqVfrdnz54j27dvT+NMG4hDhDEfpmF5TB/xIJzxiBvjnXz44YdH5s6dm+bJcIh1pRsL4sZ6MR/iGiK+dAyPxZo1a9J8Vq1aVYccFfNnXUaCbRK/LTvWpS3Snd+wHVhHuohrnl+IX2wDhpku0j9fJ9Ivnx/j/I680CZe/CbyHb+LuOT5ebSIb1O+pItwwvJ1Y/kh8gjTxnQRluffmCbSKf7HtIR323dYXj7/pmkIZxm99iPGSUOWSzhdpC35J98n2YYMN6Uz+xz/ZzqmiX2QcbpO8w9t1rtJLJM0id9FHoz5R/zpIs+xnHL+bdJ1JPmA6fI8zjLL9UasQ6Rrm7QgfSMsT1d+y7ya5NuSaRkmLhEW4xHPWD5Yxzb7dsyrW75jGXnaNE1Dn/F8/yDO/K5Mv24iPhEnxpu2fUzHOpLeLI80bru9+V++DMaZV2yTwDT8PjAPwmLfim3Nb//v//7vWLzi/8SnbZxGk0fy7Z8vM+aPTvs68yXekYb8L7Yr4yGPK9M0bVvCmIZpO03D7xmPZUXaMd1I8FuWxW9jeRFH1qlMP6YjTfLfsp78n/EyLyN+l8eVMKaL9It5kX5N+aETpot5MRzzY5z1IK4x73ybRliv5bXNb/yfLtKhnKbNft9W27zGOPEo5eH8lnHiRlge917bP/YlwiIeEZavUyyjVz6JeUU88v9NdsOOKLHC+cZgxaIL/J/pYqdmuMzsZGjCY4PFb1hGrlxmbMQyEZk/G6kUmaoJG5ydJDIxwzkyBht0LGL+TfGIg91YcGAnLSJdIs1BOGFx8GgrtgVdvj2If6e0bMK6N00fO1OIuOf5CuV0pFW5jSN/lOGlSAuWxXYnnRgfyfp00ylfkgbEOw4uodxesa5l/i/TnGnIl7m2+05THif/MP+IB8NNaUlYGY98HLE81i2U400iLXJN8x/tMaNUzidEnog0ivmUx4BY7kjSdaT5IJ8u4sU8c/ly26YF61LGg2HCmK4b/p/njZh3LLMcR6xfuS1iPUObfBcXsZFWgbjHtmT9mvIN05T7ZjexLmWcym0f0+UXTWi7vZmmXEbTshmPdYz8UOY51o/5Rfrnv0HbOI02j0S882XGvEM5DtYzX34gLNIg1jmfNyJekRZttn8+fSj31zZifct5MZ9yfeL8E9uGacrtDsIi/hHvct8p07lTPCKtY5lNIu1z+bxRjrddHv8n/t3yWxyLy3VkO0b6tNnv24r1LedFWL49GG+adx7eKR3abP+IB+uf47ex/dEmn4B5ldcFU8Ww6lLU2RvasNWbb76Zxnk09vXXX6cqP/TjURn/H0qAVC0oDGWaeugo6mvm/vznP9dDRx8PRRfy/2PJkiX10NF4DG3QtLz8t3SE8b+mx4bUP9y0aVOaZtmyZdW+ffvq/xzFOv3mN7+px0Yn5n/NNdekeOSPZJn/fffdV4+Nzu23357qSA9lujT+j3/8I/Xx3XffpfAyrXuJ7Tu0Qwyrb80jv5Fg3T/99NN67LhObSAuv/zyeuiofDq2Jel188031yFHET/yWjek+a9//euUB6mmxnaPvMl276c8X4JHmMTv4MGDw/LlzJkz0///9re/pT6GDhLD0hu33XbbCfm3jHPbfYd4EB8eFfOIF+QfthHpMdL9qNyny7iPVb+PGSGqCFx77bWpH1hHlrl58+Y65CiqOeZ4dI3YT3qlK0aaD8ijgXkQ9vTTT9chR6vYsD9cddVVabxtWtDolvjn82c41mm8dNu32+a7H//4x2n6oQv6YdUKvvrqq2N1yEljpqc6TORV8g3TcJwcqWjsHiKdyrx18cUX10NHtdnejI/mmBZ55Ve/+lXqB9aP9ey0H7bNgycjj5C/Y18JcU5DxI245nFnXVgn1g1ttn9M321/HYnYBwPnyXJ9TjnllHqofX6PeJfp3mn7lvG46KKL6qHx0Wt5bfLbj370ozT8+OOPp+nj+ojhaEvaZr8fiV55baRGuv0D05TntDjfky/b5pPQ72uZiXJCw29WZNeuXWmYzEOhgx00xkEGKS8QImN18s0336Q+9TrLDvH/kB8Av//++9Tn4qD8bVwwxDSdnH/++al+fmAjs279unBasGBB6kcagXS85ZZb6rGxicz60UcfpT64CCpPkm0QLw4O5bpzcTbSjMxBg52Cuocc+KkrGPVyS/k2LX3xxRepHwelXK/CDwczTub3339/HXJcuZOPVbkOLJd8VebLO+64I/0/f2tH0/rHgTvPv+W+1Hbf4SJ19uzZadlsX9KNbRIn2pHuR7326bHq9zEjfPDBB6lfnmjAcaDU6QQR6dYrXTHWfHDrrbem38c8//KXv6TjU1w4tUmL+O3pp5+e+rmmsH7qtm+3zXesKxcaXARfeeWVqe4xF2Ccb8Lq1avTBcMjjzxSnXfeeWm7UEeaY9BoNG175l/mrXK6Ntt7tMe0yCsjPTe1idPJyiPd8gdinYlrGX/WiXVDm+3fZn8diaY80o/8HriQJI9zDiW/d8obTfEYT72W1ya/kU7bt29PF8y8xWrOnDmpXQEFwChwtNnvR6JXXhupkW7/0DRNtA/lJQojzSfjfT4eLycUMn7+858fK0FxERsXnezYjBNO5vrZz36WwkeCE/eRIye+HYmuzRuSduzY0fhbul4H5LiYi5Ihr9p74IEH0nA/xPJpRAoOehTE+pnh2QZxAcUOykEzLkLa4jdsX0rUuThI//SnP039NkhL3qDFTvHUU09V77//frqDuWrVqnqKicFdRy7I8gJFrM+FF16Y+uOJA2RTnqTrdSeGp1FttNl3WH/u1rGfECdwMs4vmDGW/WiijfWY0Q9t03Us+SCeusQdXU6w5R3OyZAWY9Em3z388MPp+LRly5Z0/OTCg4sTLhLB8ZR1/fDDD9MrVLmIjJN0U6PK0YiLn17Gsr3Hy2SM00gcPny4Me50aLP92+6v461NfqeARGGJPM6bKSkcPfTQQ+l/U0Gb/MZx7JNPPkmFDa4N2L8piOSNl3vt94NsKp2PR+OEQkZ8U+Ddd99Nd7wvu+yyNH7BBRekce52j/YJAKXetgfwJm+//XY9NHLxfmhKhsSBi/XyAp0dfCwnKg5icUfmj3/8Y7rrkmNnG8tOwzaIO2zPPPNMtW7dujQ8Env37k398tHnO++8k/oj+aYEeYM04+DAQZ2DPwcWHnOOVNzti7t/OebdDWlSxvv5559P/fJJBgUj7pSMJR+W4snfaMSTqdNOOy31OxnJvkO+5qAdJ1puClCoDmPZjybaaI4Z8bSi6YIiCum9sNzySUivdB1LPuDiiZMrd2EpIDPveIIceqVFxPef//xn6ueawiZa23zHerDuFLS46Iw71zn2a441VIHl+MM56cknn6z/OzZN275Jr+092mNa3LHM786DbR9vmemkV5wmex5577336qHu2mz/XvvreOuV37mbTwGJwhIXlJw/qYLc9C2LyartMY/jG9c/rB/XBxRO2M/yPN5mvx9P5f42HprO91PpfDwaJxQy2NDcHXjllVfSzht3gqmTyjh3rMs7bG1EvVQujnNx0df0OrFAgSbuWJQnWeLChW6vCxEyOQciLqaJQ1NbjLxe52gQR+5GkFm50GGZOXYeDnqj9ZOf/CQV8lhXLpZGU9D7/PPPU7+8qCXexL/NyTVw0ObCKP8NcRvNxVZs42effbYOOYq0JN+NBHmKtOYgVeLkxAG93DajxZ2Z8mAJTiDk6zycbVde+LI/UTjtlu5t9h3SPaoF5M4666x6qD/70UQa7TEjtnt5QUHac4eMPJuLAmlg24E68W3SFSPJB53wFJl5EB+2U15AbpsWxIO8n29HhrtdmI63tvku0iqfhv003zeYnuoWOf7PPEaj3PaRTmU97FKb7T3aY1q0bynjFk+5mqpfoW0enIx5JJ7klQUF4sX+F3e9e23/mL7X/jqe2ub3uCFZVqmONmWTXZv8Rsf6ljdvzznnnHro+PR5WpX7/XgoC9Vxo7VfyvM968f5nvMT6zbVzsejNnTBdYKhUmZqzZ7/mzcIRFj5pgDCerXUx9DF1LEwWuEzn6EL/9QNHXTTNPyPaUo76tb7zINhfj90sZDCiG8bQxs3/aaMa7xFgP/nOoV3wnyJ39DOV4ccx3zoSp3Cm0R8WN9Ir1ykXbl+OdZlKGPXY8fxuzzebebFfNh2bMfYnhGWr1PMq1SGxzZmG8X8Yl7dtgHxZjryKG9ziO1Mn3RiGJ3WqVN4roxrYP4sO0+HmJa8EIgLYaQP07GuEcZvQqd4tNl3Yn4xDcvgd/k0bfejTvEow5k382I+pH+TeHsI/XjbRtv5t1nvJuQJfkc/fhd5M35HONPQse6MsxzGI8+gTbrSj3j1ygd0TUg/fh/xKbVJi4gH0xLPiCu/o+uG/+dxi/ShD7Yd45GmiHUsleFt8l0sL58m5hPTxDi/5f/5fPgNYj6d0hkxDV2sD/Mm7bqlQWi7vWO9I76xzQjLl8M4vw+sb4Tlv8vzJeMsi/+Td9rGabR5hPmV8Yz5h6Z9nfVs2hZleKxzpFUer9i2sbxu2595Mh5pF/NhnWM/ifnk61KK9aWfa1qfclqWyXikMeERz8jLbCPG83WJcxnhEbdy3qFTeC7SIpfPG5EXmA/5qO3y2uS3OKbl00RejmlivnlaxXwircA4XTdN2wZleGwLlhPLI04RhogX/VzTMspp+T/jnHNYp8iDhMV+gTb5BIxHvKaaxi0WK14mJAlGeHkx0SkBynB+x3jMh41KgsaOD/7P/5rkG4CODcPBrC02Gsss4w/WtVyHyDhlOnQS6RYZLce8m+YT69IG8WbaMp4h0q7T/8H/OZDlYj3Z+UObebGzlNuDeUQ60EfMq9QUzm9iZ2RbMQ3p1m0bkH8iT7FupFOc7Phdnr8IK7dPm3XttA5g/vnJgbiQ1/J8FuvAfGI6xsu4dIpHm32HaYhHTEPHNPlBDSyz135EeFM8yvA4YRDelO9BvFhXpqGPtvNvs96dlL8jbfLfRb5n3SN+TFfGq226jiQfdMLv+W05b7RNi3K/ZJ7Eg+Fu+H8et0iffLvGfGJexKdpvk3hbfId+39sCzrWtdwejMcxIuYTxxpEvLulc7dtn2+vpjQIbbY32hzT+F+5nsQt39blvJv2vbZxGk0eibTI48lw/huWE+kZ61eua2gKZ53ybcv/y7Rnmd22P3FgfSLt6Mr9NeJdpnmu07ZvinfTtAz3yu8sP9/GxJttyLR06BSPTuE54sk0Ocbz9S7z0UiW1ya/RV6LaWI982na7Pfxv26atg3K8MgjMc/IHwzHcjulQ9MyymljGuYV6814ngcDv+mVTwgv02OqmMGfoRWYFmhkRRUIHlOVeMTKY+3yfzzuo31Fm0fJTEN1pKYGdiybt3c0/W/27Nmt2jHw6I35dGvkSbUJ6vSW9blHo5/zmgxIv6GDF0epOuS48V7XeNzfbdtJ6o9ex23+T2PhoRN84/lAg2/Qzm/TQdtrJU0eJ7TJGFTUCRwqUTaeULj4HCqRN/6PE1U0fu+GeVBIaSpEgHYKTe+1pj5i29fGUsCg4VQnUeevH29U6ue8JgsavS9vaKcxiOsqTWdtj9uanjzmTz0juVbS5DHQhQwyJXcruGvFe+c7NbqOi08aUfKbwIGINxeVH/UKTM8dagoYa9asqbZt21b/Z7goxPDWiLJB2htvvNH4fYdAyZ14UMDYsGFD18ZQfKODnbB8o9Jo9HNekwWvYF66dGm6u0meCIO4rtJ01eu4LXnMn3p6XStpkqK61KCKOnfUEeyG+oBMR525vJ5gL9RlbPM76i1SL496hk118jrhd8x/pL9Ts6hz3Cs/jAfyCJ2kk6+sQy1J6r9p1SZDkiRJ0vibNm0yJEmSJE0MCxmSJEmS+spChiRJkqS+spAhSZIkqa8sZEiSJEnqKwsZkiRJkvrKQoYkSZKkvrKQIUmSJKmvLGRIkiRJ6isLGZIkSZL6ykKGJEmSpL6ykCFJkiSpryxkSJIkSeqrURcyDhw4UF1xxRXVjBkzUvfWW2/V/znRa6+9Nmza2bNnVytXrkzz6OTLL788Nn3TdPye/3322Wd1yIliHnfeeWcdMhhIa9arW5pj48aNabrJoG2cT4axxK1MY/I5XSAPkv+Yho7pR7LvDJrJlCfbmGrxnWzI/6RhnsfHM02ZL/OXJJ18oy5kbN26tXr99der7du3V3v27Knmz59f/2c4DvhXXnll9emnn1Zr1qypNmzYkAoIu3btqs4777xUAGnyt7/9LfVnzZqVllX6+c9/nvoxXZOXXnop9a+66qrUlyYa+XPz5s0p77OfXHvtta33HWmqO3jwYHXvvffWY5Kk6WTUhYx4gkCB4ZJLLqlOPfXUNJ7jLhYnmOXLl6dCxsMPP1zdc8891aZNm6pPPvkkFSDuvvvueurhHn/88WrhwoVp/lykMa8c4fz+6aefrkNOxP/mzp1brVixog4ZDKT3kSNHUl+Ty86dO1MXPv/889T/2c9+lrbXGWec0WrfGVTs/+RdTV/mAUmaHhoLGTzaXrRoUXr0TMfFUP64mzDuxMZwXj0kx10sLF26NPVzXFitWrWqOnTo0AnVobgI279/f7Vs2bJjTyHiqUSO3zNdU5WpmMdtt91Whxxdr7yaCuvY6UlKjnnlVV7mzZtXvfDCC/V/j2Ia0imfdzkNy+bJDneymUfMq3y83yue/J/wfJvkVXCa4hdiXaiyFtOWT4oinrE+TBNpzHzLvBH/C3kVIZazdu3a+j+jQ1zyZTalGeGkEcuKdWvavmOJW5s05v90MRx3cS+99NL0O7pO+06vtM23e6xjpMNItmuv/AfiEtPQMe9yO7fJCyWWxbSBceLFOsW8WAeWV95YKDFtHodIy9jGefzLvBBpSXzz/ZbhpuqZucg/sf1jHnnaN6VpL922L+vEcKwT/yfv5mnEMGH5NGU6kkZ05IFYRpk2oc32jXjFvOgzHuHke0T+B+ExHGmZxxGME87/A+mSbyviRliTpt8H1ov/9cqrkqQxOlLYvn07t5iOXH/99Ud27NiRuuXLl6cwhrFnz54jCxcuTGEMf/jhhym8dPjw4SOzZs06Mnfu3DRdW2vWrEnzjvnGPEr8n+mYvhTzGCrEpHH6zIf1Ii50DOfr1SR+x/JJm/x3W7ZsSdMQD6YhTWKaoQLQCXEjHZlP07wYR5t4EsY4feTLj23GcGy3ENsjj+eGDRsa45nHIeKWT0t4LIdpI51ZRoTl6xhxiTi3FduRZccyIz0i/cE46cpy8jQgvB9xa5vGjNOB30Q+IK7Mny7ixTDToE3aEs40kX+Yht+PZLv2yn+INCfuTMP/mD+/Y1loE98m8bvAOPPldwzncWf53TBN/jvigIhHrCP9SPNIb8IZZ9lN6xnK+Mb2zNOLedOx/Dz+edq3kceJ+cf2BduOuEU+iriy3MDyyvUup4n8mk8XYZF+yNeBaTpt3wiLbRC/I18xHfFlPOKNPE1ZP4aJS45xwiNO/DbmG+kceTePN+PMH/yfuJUIz9NEkjQ+jp89axyU4yIpx4kvP/nGiamXOFnQxQmdsG4XIsShPHny+zhJ5Zguj1cgLF+PONmVy40TZCdxUVH+juVyskKczOICLJS/jTQr50VYxLVNPOOEG+nRtPy48GS6QHya4hkn/TKe+XT8j7AyrWI5zBsR/3JbxTybtmE3kWdKzCvfvoyX+SDSKeI8lri1TWPmlccr0jafdywvtE3bWJ/y4nWk27XMW4RFnCMuZZpzIccyiEPb+DaJOIUYLy8yyzRqwv/Li0UuWolDOb9Iu4hzjJdxjfjE9srjy7QM5/PudJFMfsnzQRudti/zzuMUYvpYdlPak+fZL8o8UM6LaWL/abt9I17luhN/5sX0Ecd8eXmagmnjWBoYj/iA4ab0JCyfLo83eZbxvBBCnAgjXSRJ42tYdSkePX/99dfVzTffXIcc9/vf/z5Vber0eLoTHm/zu6ETT2pDQfuKG264IbWV4LF9WTWBx/bE4dZbb61DquqXv/xl6j///POpn/vtb397QpUr5kFYvh4//vGPU5945FUDvvrqq1RHuBMerQ+dyFJd+ty+ffuOVZd48cUX03qWdetvv/321M+reg1dFJ0wL+YfRhPPpuUzTFiO+JLuVGNjO0Y3c+bM9P+8ET3xzOcX/2Pa/LfMi3lGWrzyyitpvGwv0pSn2qAtD214SnmahaELk3roqDIOY4lb2zQejbZpGy6++OJ66KiRbtdu+S+m/dWvfpX6gXZN5EHSbqTxbePyyy+vh45qqmLZhCqVuXPPPTfFs9wu5TYP5UshLrroonpoONaJY9fQxemwef/oRz9KfdqQMU1U+2E4b5szEuX2/fOf/1wPHT1GRxfi/5H2VIWKqkAcg9iH8m1OHijTg2qlcRxtu31juWVa0/aOZeb7Sjfst+xfkXb0GY/9mTgRN9Yhjw8dYeXxP5BnOef85S9/qUOOH4vL/CZJ6r9hhYyPP/449ePEmVuwYEE9NHKcCDjxcGF++PDhaseOHak9BXXTL7vssmF1Y//7v/879fOTABcOnNyaGoDHdHn9c04qnFzykx8nHC7cOXnxtivq5PL/OGF2QoGn2wVPnNxOP/301M8R71KvE+9I4xlp17T8Mox1oZ0K9aPz7o477kj///bbb1MfZTzjf0xb/p55Mm9wcUGd8FJTnmqLdaTARV1u6lhTFzvaNeTiorqT0cZtJGk8Gm3TNpxyyin10FFj2a6lmLbTRTlGGt822l6Qljptc/ZL9hnyDPtP03ZHmZadcGMEFFRzxHv79u3pQpdp5syZk/Inx6PyWNVWGadvvvkm9cu0pkP8nxdd0CaC7cLxknWmjUZ+fEVTWkfh6vvvv2+9fVluU2F/pOImUhRuoh/hxAkc/8v4EIaYplQeO0mjpoK2JKn/Rv12qdHiBMeFNHenuSvICStOKnEHC5wkucCOjpM4ygbgzI8CS5xImAfDhJUo6DAflht3z7gwGEnj34kwnvGkAHPkSKomd0LX7YlOoJDY9Fu68cC2JC9Q4Prf//3fdJeap1xcKAyasaTtWLfraEx0XmiLgiivx2af2b17d7rwfuihh+r/jg7Hkw0bNqTCbX7RCi5keVsehQ2mY9/lAj0aovcD+b0pneniiQk3NShIcxOH/IBHHnkk/bYsaJS+++67eui4idq+xJs4xpMR+uzz5U0a1qspLnTdnlRxjuEmBWlAIYmn35Kk8TeskBFPK7744ovUz8VTjpHg5EsBodMdPb4ZgLhzFgUILqw5oZcdHnvssdTPUbUjTiQUWBiOu2Al7mBRhYALBU6i3InjRNwJT0S4UCnFW1ziRPjPf/4z9XNNj/DbahvPuCPXtPymML5PMhbvvfdePdSMNOFCp9SUp3oh/bhQJD+QBlFFjXSJQudIjDZuI03j0eqVtt2MdbuGeDJAVZQc+zD7cn6BPZb4jheeIHB3m+MFF59cgHNDY6zfImEeFNa4+KUQUx7TuNnB8Y7pKAhzkc8FbZmOo8W8Oh1HS9zEiSpLXJhzPGx6O1/uo48+Sv3TTjst9dFr+/7whz9sfKLIcZiC3UiOf9w44GYK60if6rmlt99+ux5qj7Rgm/F0O25mWVVKkibGsEIGd4O4qH722WfrkOO4uOd/3apRlOKDeX/4wx9Sv/TEE0+k/tlnn536UYD405/+lE7oZceFNheX5YmbOMWJhLtg3BUr74Jx8VEWeLgw6PXYnAsHTqTlnUBOhFEFg4tgLr7Ki4CowjWSKgWjiWc8ycl/w3B+QQima7rwiWV2uyCKAuGTTz6Z+oHlkA5x15b2DU3bqClP9RJVIG666aaUBoF5c+E0UmOJW9s0Ho22advJWLZrKS7AyvZPcYFG1bKxxnc8xQ2LW265JfVD08XwaJBXyHtxTCNtuaDO20/hnHPOqYfGLtoMPfPMM6kfuIhn+8YrY0n78mnnWWedVQ8d13Q8e+qpp9Jxk+NM2+0bx/dyH4g2ECOpIhn5LtI1LwjE8b2puizHZ9K/DM/F8fmNN95Iw/mxRJI0jo4U4o0hQxfG6a0cdAwTxnCIsF6GLoDSdEMniTS8YcOG1A2d0FI4YYi3tMR4E5bfaRrmOVQISv8v33aCeMsJy2U+jPMbwtYUb3PJDV2Upvnmvxs6UaXfxesl4402rCPLZppY73zepBldKQ9vE8+Yhj7KONJF+tKFmI4u4hnzZvrQKZ4sn2lZf36bL4fhQFi+jEgvuogzfcablhOIL9PEekV8Yx3y3zId/yuV4W3i1qRtGpdpF+mbz5v/579Bm7QlvJwXxrpdy/CIC/NgXsyTeRO30DYvlCJeoRwPncJz/J/pcsSV8IgXHfsi8c+nJ5xx+rkyvCkekWeY5nD9xqU87SO9yrRvWl6uXHYu0pb45MugY/sjlhHTxDZpmiaOVUzTFLfR7uuRXvH2pvy4HvNvSlNEPPJ0Cywz/scw84rtQFwD48w/F3GIeJb4H8uWJPVX41k8P6HQxYkmFyeENjiwx/TRMf/8gB8X5eVySpwcmS5OmoHxmDcn/iasVx4P5lWekJow7zih0RH3Mp6cyMppyhMay246mZXhveLJsgnP41Aun/SMC4Uc65JfdDFvpsvTrFM8wTrleYPpyrRgXvkymIaLDoZj2liHTssJnfJirFvEm+GmbVmGt4lbJ23SmPnl68Syy3nz//w3oVfaRpo1xXMs27UpnDSJfY15lvNCm7xQivQI5XjoFJ7j//m2DYTlcSddSB/iSodOaVmGN8WDeTFfloHIF5H2scwy7ZuWl+sUJzCvcr1YJnEJkbdjGjqmIX4htjXzivh22m5t93XyRsyLZUcBI8Q+QoemNAXLI7z8fWDZ+f5H3MppCWf+pUiTMg+DcNZNktRfM/gzdJCVJhxVTf74xz/2pcqRNNlRzWjoAviEqpwTKao6jfb1ulMVb/yia3odtiRpfEz426WkQAGDVxhLg472E3v37j2pBYzpirSnzVL5TRRJ0viykKGTgoaavGc/GplKg4zG1t5Fn1g8KaVR/C9+8YvUqJ03TUmSJo7VpSRpmphO1aV4u1p8mPDll1/2KZIkTTALGZIkSZL6yupSkiRJkvrKQoYkSZKkvrKQIUmSJKmvLGRIkiRJ6isLGZIkSZL6ykKGJEmSpL6ykCFJkiSpryxkSJIkSeorCxmSJEmS+spChiRJkqS+spAhSZIkqa8sZEiSJEnqKwsZkiRJkvrKQoYkSZKkvrKQIUmSJKmvLGRIkiRJ6isLGZIkSZL6qKr+Pw6J/Kme34UiAAAAAElFTkSuQmCC\"\u003e\u003c/p\u003e\n\u003ch2\u003eModel interpretation\u003c/h2\u003e\n\u003cp\u003eTo interpret our XGBoost regression models, we used SHAP v0.45.0\u003csup\u003e19\u003c/sup\u003e. Each predictor of a sample is assigned a SHAP value, which corresponds to the change in predicted future fitness given the information of that predictor. SHAP values therefore enable the decomposition of each future fitness prediction into the sum of contributions from each predictor. The relative importance of our predictors were assessed by their mean absolute SHAP values, with higher values indicating more important to the model predictions. The relative importance of groups of predictors (i.e., physiochemical, intrahost diversity, DMS, intrahost linkage and interhost linkage), were assessed by summing the absolute SHAP values for all predictors within a group and dividing this by the number of SAVs considered.\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eStatistical analysis and visualisation\u003c/h2\u003e\n\u003cp\u003eAll statistical analysis and visualisations were performed using the stats and ggplot packages in R v4.3.2. Where applicable,\u0026nbsp;\u003cem\u003ep\u003c/em\u003e-values where corrected for multiple testing using the Benjamini-Hochberg procedure.\u0026nbsp;\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eAcknowledgments\u003c/p\u003e\n\u003cp\u003eC.C.S.T. is funded by the National Science Scholarship from the Agency for Science, Technology and Research (A*STAR), Singapore. M.E.Z, F.B. and L.v.D. are funded by the European Commission (Horizon 2021-2024, END-VOC Project). L.v.D. is additionally funded by the UKRI Future Leaders Fellowship (MR/X034828/1). Views and opinions expressed are however those of the authors only and do not necessarily reflect those of the European Union or the European Health and Digital Executive Agency. For the purpose of open access, the corresponding author has applied a \u0026lsquo;Creative Commons Attribution\u0026rsquo; (CC BY) licence to any Author Accepted Manuscript version arising. The authors acknowledge the use of the UCL Computer Science cluster and associated support services, in the completion of this work.\u003c/p\u003e\n\u003cp\u003eAuthor contributions\u003c/p\u003e\n\u003cp\u003eC.C.S.T and F.B. conceptualised and designed the study. C.C.S.T. performed all analyses with intellectual inputs from all co-authors. C.C.S.T wrote the manuscript with contributions and edits from all co-authors.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eDeclaration of competing interest\u003c/p\u003e\n\u003cp\u003eThe authors declare that to current knowledge, there are no legal, financial or personal competing interests.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eData and code availability\u003c/p\u003e\n\u003cp\u003eAll custom code used to perform the analyses reported here are hosted on GitHub (https://github.com/cednotsed/early_SC2_trajectory). All data used to train and evaluate the models are provided in \u003cstrong\u003eSupplementary Table 1\u003c/strong\u003e.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eVan Dorp, L. \u003cem\u003eet al.\u003c/em\u003e Emergence of genomic diversity and recurrent mutations in SARS-CoV-2. \u003cem\u003eInfection, Genetics and Evolution\u003c/em\u003e \u003cstrong\u003e83\u003c/strong\u003e, 104351 (2020).\u003c/li\u003e\n \u003cli\u003eBalloux, F. \u003cem\u003eet al.\u003c/em\u003e The past, current and future epidemiological dynamic of SARS-CoV-2. \u003cem\u003eOxford Open Immunology\u003c/em\u003e \u003cstrong\u003e3\u003c/strong\u003e, iqac003 (2022).\u003c/li\u003e\n \u003cli\u003eCarabelli, A. M. \u003cem\u003eet al.\u003c/em\u003e SARS-CoV-2 variant biology: immune escape, transmission and fitness. \u003cem\u003eNature Reviews Microbiology\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 162\u0026ndash;177 (2023).\u003c/li\u003e\n \u003cli\u003eStarr, T. N. \u003cem\u003eet al.\u003c/em\u003e Shifting mutational constraints in the SARS-CoV-2 receptor-binding domain during viral evolution. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e377\u003c/strong\u003e, 420\u0026ndash;424 (2022).\u003c/li\u003e\n \u003cli\u003eWitte, L. \u003cem\u003eet al.\u003c/em\u003e Epistasis lowers the genetic barrier to SARS-CoV-2 neutralizing antibody escape. \u003cem\u003eNature Communications\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 302 (2023).\u003c/li\u003e\n \u003cli\u003eAmicone, M. \u003cem\u003eet al.\u003c/em\u003e Mutation rate of SARS-CoV-2 and emergence of mutators during experimental evolution. \u003cem\u003eEvolution, medicine, and public health\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 142\u0026ndash;155 (2022).\u003c/li\u003e\n \u003cli\u003eMarkov, P. V. \u003cem\u003eet al.\u003c/em\u003e The evolution of SARS-CoV-2. \u003cem\u003eNature Reviews Microbiology\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 361\u0026ndash;379 (2023).\u003c/li\u003e\n \u003cli\u003eLythgoe, K. A. \u003cem\u003eet al.\u003c/em\u003e SARS-CoV-2 within-host diversity and transmission. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e372\u003c/strong\u003e, eabg0821 (2021).\u003c/li\u003e\n \u003cli\u003eGu, H. \u003cem\u003eet al.\u003c/em\u003e Within-host genetic diversity of SARS-CoV-2 lineages in unvaccinated and vaccinated individuals. \u003cem\u003eNature Communications\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 1793 (2023).\u003c/li\u003e\n \u003cli\u003eShu, Y. \u0026amp; McCauley, J. GISAID: Global initiative on sharing all influenza data\u0026ndash;from vision to reality. \u003cem\u003eEurosurveillance\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 30494 (2017).\u003c/li\u003e\n \u003cli\u003eElbe, S. \u0026amp; Buckland‐Merrett, G. Data, disease and diplomacy: GISAID\u0026rsquo;s innovative contribution to global health. \u003cem\u003eGlobal Challenges\u003c/em\u003e \u003cstrong\u003e1\u003c/strong\u003e, 33\u0026ndash;46 (2017).\u003c/li\u003e\n \u003cli\u003eLeinonen, R., Sugawara, H., Shumway, M., \u0026amp; International Nucleotide Sequence Database Collaboration. The sequence read archive. \u003cem\u003eNucleic acids research\u003c/em\u003e \u003cstrong\u003e39\u003c/strong\u003e, D19\u0026ndash;D21 (2010).\u003c/li\u003e\n \u003cli\u003eHenikoff, S. \u0026amp; Henikoff, J. G. Amino acid substitution matrices from protein blocks. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e \u003cstrong\u003e89\u003c/strong\u003e, 10915\u0026ndash;10919 (1992).\u003c/li\u003e\n \u003cli\u003eEddy, S. R. Where did the BLOSUM62 alignment score matrix come from? \u003cem\u003eNature biotechnology\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 1035\u0026ndash;1036 (2004).\u003c/li\u003e\n \u003cli\u003eCargill, M. \u003cem\u003eet al.\u003c/em\u003e Characterization of single-nucleotide polymorphisms in coding regions of human genes. \u003cem\u003eNature Genetics\u003c/em\u003e \u003cstrong\u003e22\u003c/strong\u003e, 231\u0026ndash;238 (1999).\u003c/li\u003e\n \u003cli\u003eGreaney, A. J. \u003cem\u003eet al.\u003c/em\u003e Complete mapping of mutations to the SARS-CoV-2 spike receptor-binding domain that escape antibody recognition. \u003cem\u003eCell host \u0026amp; microbe\u003c/em\u003e \u003cstrong\u003e29\u003c/strong\u003e, 44\u0026ndash;57 (2021).\u003c/li\u003e\n \u003cli\u003eStarr, T. N. \u003cem\u003eet al.\u003c/em\u003e Deep mutational scanning of SARS-CoV-2 receptor binding domain reveals constraints on folding and ACE2 binding. \u003cem\u003ecell\u003c/em\u003e \u003cstrong\u003e182\u003c/strong\u003e, 1295\u0026ndash;1310 (2020).\u003c/li\u003e\n \u003cli\u003eChen, T. \u0026amp; Guestrin, C. Xgboost: A scalable tree boosting system. in 785\u0026ndash;794 (2016).\u003c/li\u003e\n \u003cli\u003eLundberg, S. M. \u0026amp; Lee, S.-I. A unified approach to interpreting model predictions. \u003cem\u003eAdvances in neural information processing systems\u003c/em\u003e \u003cstrong\u003e30\u003c/strong\u003e, (2017).\u003c/li\u003e\n \u003cli\u003eLewontin, R. C. The interaction of selection and linkage. I. General considerations; heterotic models. \u003cem\u003eGenetics\u003c/em\u003e \u003cstrong\u003e49\u003c/strong\u003e, 49 (1964).\u003c/li\u003e\n \u003cli\u003ePlante, J. A. \u003cem\u003eet al.\u003c/em\u003e Spike mutation D614G alters SARS-CoV-2 fitness. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e592\u003c/strong\u003e, 116\u0026ndash;121 (2021).\u003c/li\u003e\n \u003cli\u003eGoldswain, H. \u003cem\u003eet al.\u003c/em\u003e The P323L substitution in the SARS-CoV-2 polymerase (NSP12) confers a selective advantage during infection. \u003cem\u003eGenome biology\u003c/em\u003e \u003cstrong\u003e24\u003c/strong\u003e, 47 (2023).\u003c/li\u003e\n \u003cli\u003eUlrich, L. \u003cem\u003eet al.\u003c/em\u003e Enhanced fitness of SARS-CoV-2 variant of concern Alpha but not Beta. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e602\u003c/strong\u003e, 307\u0026ndash;313 (2022).\u003c/li\u003e\n \u003cli\u003eSun, X. \u003cem\u003eet al.\u003c/em\u003e Enhanced fitness of SARS-CoV-2 B. 1.617. 2 Delta variant in ferrets. \u003cem\u003eVirology\u003c/em\u003e \u003cstrong\u003e582\u003c/strong\u003e, 57\u0026ndash;61 (2023).\u003c/li\u003e\n \u003cli\u003eWu, H. \u003cem\u003eet al.\u003c/em\u003e Nucleocapsid mutations R203K/G204R increase the infectivity, fitness, and virulence of SARS-CoV-2. \u003cem\u003eCell host \u0026amp; microbe\u003c/em\u003e \u003cstrong\u003e29\u003c/strong\u003e, 1788\u0026ndash;1801 (2021).\u003c/li\u003e\n \u003cli\u003eMeijers, M., Ruchnewitz, D., Eberhardt, J., Łuksza, M. \u0026amp; L\u0026auml;ssig, M. Population immunity predicts evolutionary trajectories of SARS-CoV-2. \u003cem\u003eCell\u003c/em\u003e \u003cstrong\u003e186\u003c/strong\u003e, 5151\u0026ndash;5164 (2023).\u003c/li\u003e\n \u003cli\u003eObermeyer, F. \u003cem\u003eet al.\u003c/em\u003e Analysis of 6.4 million SARS-CoV-2 genomes identifies mutations associated with fitness. \u003cem\u003eScience\u003c/em\u003e \u003cstrong\u003e376\u003c/strong\u003e, 1327\u0026ndash;1332 (2022).\u003c/li\u003e\n \u003cli\u003ePucci, F. \u0026amp; Rooman, M. Prediction and evolution of the molecular fitness of SARS-CoV-2 variants: introducing SpikePro. \u003cem\u003eViruses\u003c/em\u003e \u003cstrong\u003e13\u003c/strong\u003e, 935 (2021).\u003c/li\u003e\n \u003cli\u003eAbousamra, E., Figgins, M. \u0026amp; Bedford, T. Fitness models provide accurate short-term forecasts of SARS-CoV-2 variant frequency. \u003cem\u003ePLOS Computational Biology\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, e1012443 (2024).\u003c/li\u003e\n \u003cli\u003eIto, J. \u003cem\u003eet al.\u003c/em\u003e A Protein Language Model for Exploring Viral Fitness Landscapes. \u003cem\u003ebioRxiv\u003c/em\u003e 2024\u0026ndash;03 (2024).\u003c/li\u003e\n \u003cli\u003eDadonaite, B. \u003cem\u003eet al.\u003c/em\u003e Spike deep mutational scanning helps predict success of SARS-CoV-2 clades. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e631\u003c/strong\u003e, 617\u0026ndash;626 (2024).\u003c/li\u003e\n \u003cli\u003eSmith, J. M. \u0026amp; Haigh, J. The hitch-hiking effect of a favourable gene. \u003cem\u003eGenetics Research\u003c/em\u003e \u003cstrong\u003e23\u003c/strong\u003e, 23\u0026ndash;35 (1974).\u003c/li\u003e\n \u003cli\u003eBushnell, B. BBMap: a fast, accurate, splice-aware aligner. (2014).\u003c/li\u003e\n \u003cli\u003eLangmead, B. \u0026amp; Salzberg, S. L. Fast gapped-read alignment with Bowtie 2. \u003cem\u003eNature methods\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 357\u0026ndash;359 (2012).\u003c/li\u003e\n \u003cli\u003eLi, H. \u003cem\u003eet al.\u003c/em\u003e The sequence alignment/map format and SAMtools. \u003cem\u003ebioinformatics\u003c/em\u003e \u003cstrong\u003e25\u003c/strong\u003e, 2078\u0026ndash;2079 (2009).\u003c/li\u003e\n \u003cli\u003eMarinier, E. \u003cem\u003eet al.\u003c/em\u003e Quasitools: a collection of tools for viral quasispecies analysis. \u003cem\u003eBioRxiv\u003c/em\u003e 733238 (2019).\u003c/li\u003e\n \u003cli\u003eDanecek, P. \u003cem\u003eet al.\u003c/em\u003e Twelve years of SAMtools and BCFtools. \u003cem\u003eGigascience\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, giab008 (2021).\u003c/li\u003e\n \u003cli\u003eRambaut, A. \u003cem\u003eet al.\u003c/em\u003e A dynamic nomenclature proposal for SARS-CoV-2 lineages to assist genomic epidemiology. \u003cem\u003eNature microbiology\u003c/em\u003e \u003cstrong\u003e5\u003c/strong\u003e, 1403\u0026ndash;1407 (2020).\u003c/li\u003e\n \u003cli\u003eO\u0026rsquo;Toole, \u0026Aacute;. \u003cem\u003eet al.\u003c/em\u003e Assignment of epidemiological lineages in an emerging pandemic using the pangolin tool. \u003cem\u003eVirus evolution\u003c/em\u003e \u003cstrong\u003e7\u003c/strong\u003e, veab064 (2021).\u003c/li\u003e\n \u003cli\u003eKyte, J. \u0026amp; Doolittle, R. F. A simple method for displaying the hydropathic character of a protein. \u003cem\u003eJournal of molecular biology\u003c/em\u003e \u003cstrong\u003e157\u003c/strong\u003e, 105\u0026ndash;132 (1982).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5298116/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5298116/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003ePredicting the fitness of mutations in the evolution of pathogens is a long-standing and important, yet largely unsolved problem. In this study, we used SARS-CoV-2 as a model system to explore whether the intrahost diversity of viral infections could provide clues on the relative fitness of single amino acid variants (SAVs). To do so, we analysed ~15 million complete genomes and nearly ~8000 sequencing libraries generated from SARS-CoV-2 infections, which were collected at various timepoints during the COVID-19 pandemic. Across timepoints, we found that many successful SAVs were detected in the intrahost diversity of samples collected prior, with a median of 6-40 months between the initial collection dates of samples and the highest frequency seen for these SAVs. Additionally, we found that the co-occurrence of intrahost SAVs significantly captures genetic linkage patterns observed at the interhost level (Pearson’s \u003cem\u003er\u003c/em\u003e=0.28-0.45, all p\u0026lt;0.0001). Further, we show that machine learning models can learn highly generalisable intrahost, physiochemical and phenotypic patterns to forecast the future fitness of intrahost SAVs (\u003cem\u003er\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e=0.48-0.63). Most of these models performed significantly better when considering genetic linkage (\u003cem\u003er\u003c/em\u003e\u003csup\u003e\u003cem\u003e2\u003c/em\u003e\u003c/sup\u003e=0.53-0.68). Overall, our results document the evolutionary forces shaping the fitness of mutations, which may offer potential to forecast the emergence of future variants and ultimately inform the design of vaccine targets.\u003c/p\u003e","manuscriptTitle":"Intrahost dynamics, together with genetic and phenotypic traits predict the success of viral mutations","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-11-15 10:36:39","doi":"10.21203/rs.3.rs-5298116/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"nature-communications","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"NCOMMS","sideBox":"Learn more about [Nature Communications](http://www.nature.com/ncomms/)","snPcode":"","submissionUrl":"https://mts-ncomms.nature.com/","title":"Nature Communications","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Nature Communications","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"2c96ff30-5e93-4557-8219-ab68e16ff6ae","owner":[],"postedDate":"November 15th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[{"id":39611053,"name":"Biological sciences/Evolution/Molecular evolution"},{"id":39611054,"name":"Health sciences/Diseases/Infectious diseases/Viral infection"}],"tags":[],"updatedAt":"2024-11-15T10:36:39+00:00","versionOfRecord":[],"versionCreatedAt":"2024-11-15 10:36:39","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5298116","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5298116","identity":"rs-5298116","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.