Exposing Structural Variations in SARS-CoV-2 Evolution | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Exposing Structural Variations in SARS-CoV-2 Evolution Jiaan Yang, Peng Zhang, Wen Xiang Cheng, Youyong Lu, Wu Gang, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-800496/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 11 Nov, 2021 Read the published version in Scientific Reports → Version 1 posted 9 You are reading this latest preprint version Abstract The mutation of SARS-CoV-2 influences viral function as residue replacements affect both physiochemical properties and folding conformations. Although a large amount of data on SARS-CoV-2 is available, the investigation of how viral functions change in response to mutations is hampered by a lack of effective structural analysis. Here, we exploit advances in protein structure fingerprint technology to study the folding conformational changes induced by mutations. With the integration of both protein sequences and folding conformations and alignments of SARS-CoV to SARS-CoV-2 , the UK variant and India variant, we found that structural variations in the spike protein at the binding interface interacting with ACE2 play a critical role in coronavirus entry into human cells. Additionally, the structural variations impact vaccine effectiveness and drug function over the course of SARS-CoV-2 evolution. The analysis of structural variations revealed how the coronavirus has gradually evolved in both structure and function and how the SARS-CoV-2 variants have contributed to more severe acute disease worldwide. Infectious Diseases Structural Biology Mutation SARS-CoV-2 spike protein physiochemical properties folding conformation Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) belongs to the coronavirus family, and it is an urgent need to explore the SARS-CoV-2 structure, function and activity. 1 In particular, mutations in SARS-CoV-2 are considered a priority because of their potential to increase transmissibility and virulence while reducing the effectiveness of vaccines and impacting the development of medical treatments. 2 , 3 Mutations that alter the protein sequence, including replacements or deletions of amino acid residues, may affect protein properties and folding conformations and result in changes to the biological functions of the virus. 4 The interaction of the receptor-binding domain (RBD) of the spike protein of SARS-CoV-2 with angiotensin-converting enzyme 2 (ACE2) receptors is key for allowing the virus to enter human cells. 5 , 6 Thus, mutations in the RBD directly influence disease development and epidemic spread. 7 , 8 To date, over 3,000 SARS-CoV-2 sequences and nearly 800 spike protein 3D structural data sources are available in the National Center for Biotechnology Information (NCBI) database and Protein Data Bank (PDB). According to the COVID-19 Genomics UK (COG-UK) Consortium, more than 4,000 mutations have been detected in the spike protein alone, 9 which provides sufficient data to investigate coronavirus mutations to understand changes in its physiochemical properties as well as folding conformations leading to virus evolution over time. With protein sequence alignment, the positions of replaced amino acid residues can be discovered, and the concomitant changes in physiochemical properties can be further probed. 9,10 In addition to physiochemical properties, changes in the protein folding conformation also impact biological viral functions. For proteins with known 3D structures, the conformational differences caused by mutations can be roughly compared by structure superposition with root-mean-square deviation (RMSD) as a measurement. 11 For proteins without known 3D structures, the protein structures first need to be predicted by computational dynamics simulations. However, for mutational differences, the reliability of the predicted protein structure remains a challenge even when using ab initio modeling methods. 12-14 Thus, it is crucial that a new approach overcomes these barriers to studying structural mutations. At this point, we propose using the protein structure fingerprint approach 15 , 16 to analyze the changes in folding conformation caused by mutations. With protein structure fingerprints, the protein folding shape code (PFSC) provides an alphabetical string to completely describe the folding conformation for 3D protein structure. Additionally, according to the protein sequence, the protein folding variation matrix (PFVM) reveals the folding variations along the sequence and generates the possible folding conformations. Thus, the alignment of the protein sequence with the PFSC string can comprehensively expose the variations in both biological functions and folding conformations caused by mutations in SARS-CoV-2. Here, the structural variations in evolved coronavirus strains, from SARS-CoV to SARS-CoV-2, UK variant and India variant, are studied. Methods Structural bioinformation. All protein structural data for SARS-CoV-2 were extracted from public databases. The sequences were obtained from the NCBI and UniProt databases, and protein 3D structures were obtained from the PDB. The cd21477 cluster was obtained from the NCBI Conserved Domain Database, which contains the protein structure with PDB ID 6ACC for the SARS-CoV spike protein released in August 2018 and the protein structure with PDB ID 6VSB for the SARS-CoV-2 spike protein released in February 2020. Then, the mutations between 6ACC and 6VSB were analyzed according to either protein 3D structures or sequences by protein structure fingerprint technology. Information on the UK variant and India variant of SARS-CoV-2 was obtained from Public Health England. Seven mutations in the spike protein were identified, and the variations in physiochemical properties and folding conformations were studied. The complexes of coronavirus with ACE2 were obtained from PDB; PBD ID 6ACG is the complex with SARS-CoV, and PBD ID 7A98 is the complex with SARS-CoV-2. Protein comparison. The sequences of the spike protein between SARS-CoV and SARS-CoV-2 were aligned with the Clustal Omega program through UniProt and then compared according to their physiochemical properties. Discovery Studio (version 4.5) was used to generate 3D images of protein structures, and then the superimposition of protein 3D structures was performed. Furthermore, with protein structure fingerprint technology, the variations in protein folding conformations were exposed in detail. Protein structure fingerprint. First, the complete folding space for a set of 5 successive points was mathematically covered by a set of folding shapes. Second, the possible folds of a fragment of 5 amino acids could be defined by the 27 protein folding shape code (PFSC) with alphabetical letters, as shown in Fig. 4 . Third, any protein sequence has a protein folding variation matrix (PFVM), and any protein with a given 3D structure can be expressed by a PFSC string as a protein structure conformation. In the PFSC string, two PFSC letters next each other overlap by four amino acids; thus, a PFSC string represents the complete folding conformation of the 3D protein structure. It is significant that protein folding conformations as PFSC strings can be aligned for comparison. Therefore, the folding variations of SARS-CoV-2 as well as its mutations may be well analyzed by protein structure fingerprints. Software availability. The protein structure fingerprint can be accessed on Website http://www.micropht.com . Results The changes in both physiochemical properties and folding conformations of SARS-CoV-2 due to mutation are studied based on the protein 3D structures and sequences of the spike protein, and the interaction between coronavirus and ACE2 are a particular focus. The structural analysis covers coronavirus strains from early SARS-CoV to SARS-CoV-2, UK variant and recent India variant. Variations based on 3D structures. Coronavirus spike proteins have an S1 subunit at the N-terminus (~700 amino acids) and anS2 subunit at the C-terminus (~600 amino acids). Analysis of many protein 3D structures confirmed that three S1/S2 heterodimers assembled to form a trimer spike protruding from the viral envelope. 17 The S1 subunit of the spike protein in SARS-CoV-2 is an envelope glycoprotein that plays the most important role in viral attachment, fusion, and entry into host cells, and is a major target for the development of neutralizing antibodies, inhibitors and vaccines. The S1 subunit contains a receptor-binding domain (RBD), and many studies have found that the RBD of the spike protein in SARS-CoV-2 strongly binds to human and bat angiotensin- converting enzyme 2 (ACE2) receptors. 18-20 A set of sequences for the cluster of the SARS-CoV-like_Spike_S1_RBD subfamily (cd21477) that contains the conserved protein domain of the S1 RBD subfamily for SARS-CoV-like and SARS-CoV-2 spike proteins is available in the NCBI database. The sequences of the cd21477 cluster were aligned and are presented in Table 1 , where the red color font indicates highly conserved fragments, blue indicates less conserved fragments and gray indicates unaligned fragments. It is not surprising that the mutations were most frequent on less conserved residues (blue font). Additionally, it is noted that some sequences have the given 3D structures in the PDB. The protein structure of the SARS-CoV spike protein, with PDB ID 6ACC, was released in August 2018; 21 the protein structure of the SARS-CoV-2 spike protein, with PDB ID 6VSB, was released in February 2020. 17 The residues that differ between 6ACC and 6VSB are marked in green. Some changes in the physiochemical properties of the protein based on the differences between 6ACC to 6VSB are summarized in Table 2 , including hydrophobicity, negative or positive charge, polarity, size of side chain, aromatic and etc . The changes in the physiochemical properties are represented by a “+” sign for an increase in the property after mutation and a “-” sign for a decrease in the property after mutation. From the given structures of 6ACC and 6VSB, 3D images of folding conformations are compared and displayed in Fig. 1 . The 3D structures directly provide a visualization to observe the protein structures, and the superposition allows a comparison of the structures. Although more than 30 mutations in the fragment between 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541 for SARS-CoV-2 occurred, representing up to 37.5% residue replacement, the structure superposition showed that the folding conformations of 6ACC and 6VSB were still similar overall. It is difficult to distinguish the folding differences of spike proteins of SARS-CoV and SARS-CoV-2 based on the 3D structure only. With the protein folding shape code (PFSC), however, the differences in folding can be exposed. Any protein 3D structure can be converted into a PFSC description, which is an alphabetical string representing the continuous folding shape of each five-amino-acid in sequence. Thus, the folding conformations of 6ACC for SARS-CoV and 6VSB for SARS-CoV-2 can be compared by PFSC alignment and displayed in Table 3 . In PFSC, generally, the red color indicates typical alpha helix, pink indicates alpha-like helix, blue indicates a typical beta strand, light blue indicates a beta-like strand, and black indicates an irregular fold. According tothe PFSC color notation, it is obvious that the secondary structural fragments are well aligned. For example, the fragments of alpha helices at 324-330 and 352-364 and the beta strands at 349-351 and 378-386 on 6ACC are aligned with the corresponding fragments in 6VSB. Also, the PFSC alignment exposes local folding comparison in detail, which the local folding similarity and differences between PFSC strings are indicated; “|” indicates an identical folding shape, “:” indicates a similar folding shape, and “.” indicates dissimilar folding. For example, the folding letters at 334, 335, 340, 341, 343, 370, 379 and 380 on 6ACC are different from 6VSB. Also, due to mutations, the adjustments of beta strand at 349-351 and 452-454 fragments on 6ACC are exposed. Thus, the PFSC revealed the changes in local folding shapes caused by the mutations. Variations based on sequences. The variations of folding conformations for a protein based on sequence alone can be exposed by the PFVM. The PFVMs for sequences taken directly from the structures of 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541 for SARS-CoV-2 separately are exhibited in Table 4 . The PFSC letters in each column represent the folding variations of 5 successive amino acid residues in sequence while the favored folding shapes are ranked on top, and the numbers are different in each column. The PFVM exhibits the folding variations along the sequence. The number deviations of folding shapes along sequences in PFVM between 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541 for SARS-CoV-2 are shown by the curves in Fig. 2 , where the yellow and green blocks indicate the regions of fluctuation due to mutations. It was apparent that the mutations caused changes in folding flexibility; some fragments have the potential to be more flexible, and other fragments are more rigid. Thus, along the sequence from the N-terminus to the C-terminus, the variations in the folding conformation are well exposed. The most likely conformations for a protein can be predicted from PFVM. Taking one letter from each column, a massive number of PFSC strings can be formed, and each string is one possible folding conformation. Although a large number of folding conformations exist, the letters on top of each column are directly constructed into one of the most likely conformations, which is named PFVM-01. This predicted conformation may be assessed by a given 3D protein structure through PFSC alignment. Two PFSC strings of the first row from Table 4 for 6ACC-A-306-527 and 6VSB-C-319-541 and two PFSC strings of their 3D structures from Table 3 are aligned in Table 5 . The PFSC letters in red and pink colors represent alpha helices, those in blue and light blue represent beta strands and those in black represent irregular folding shapes. Overall, with observation, the secondary fragments are aligned, so the predicted folding conformations of PFVM-01 for SARS-CoV and SARS-CoV-2 are similar to the given 3D structures. Thus, the PFVM-01 as the most likely conformation, which is generated from PFVM, is a reliable prediction. Structure variations with virus evolution. The sequence of the SARS-CoV spike protein (UniProtKB = P59594 (SPIKE_SARS)) was first determined in 2003. 22 The sequence of the SARS-CoV-2 spike protein (UniProtKB = P0DTC2 (SPIKE_SARS2)) was determined in January 2020. 23 After 17 years of evolution from SARS-CoV to SARS-CoV-2, the spike protein sequences are approximately 24% different. The UK variant is mutant of SARS-CoV-2 that was noted in November 2020 from a sample taken in the UK in September. SARS-CoV-2 infections in the UK increased because of one or more mutations in the virus spike protein. The lineage of the UK mutant B.1.1.7 VOC-202012/01 is taken from the Public Health England, 24 which was reported on March 5, 2021, with seven mutations in the spike protein: E484K, N501Y, A570D, P681H, T716I, S982A and D1118H. 25 Similar variants have also emerged in South Africa (lineage B.1.351) and Brazil (lineage P.1). Recently, India variant (lineage B.1.617.2) with mutations K417N, N440K, L452R, T478K and E484Q in the spike protein caused the epidemic to become severe. Thus, it is important to understand the effects of the mutations following virus evolution. In order to study the mutations in SARS-CoV-2, a sequence of QTJ15692 (GenBank) was taken as the background reference, which was deposited in the NCBI database on April 2020 before the UK mutation. The mutations may cause changes in physicochemical functions as well as in folding conformations, which together impact biological functions. The changes in physiochemical properties, including hydrophobicity, negative or positive charge, polarity, residue size and aromaticity, are listed in the top rows of Table 6 . For example, the mutation E484K changed a negative charge to a positive charge; A570D is a change from hydrophobic to negative, from non-charged to charged and from non-polar to polar; P681H is a change from hydrophobic to positive, from non-charged to charged and from non-polar to polar and an increase in the size of the side chain due to an aromatic moiety; T716I and S982A are changes from polar to hydrophobic. The changes in physiochemical properties caused by mutations are indicated in detail. Furthermore, the variations in local folding shapes may be revealed by PFVM because each PFSC letter in PFVM represents the folding shape of 5 successive amino acids in sequence. The PFVMs of seven regions for these related mutations are displayed in Table 7 , which shows the folding variations before and after mutations. To compare each pair of PFVMs, the fluctuations in the number of folding shapes and the contributions of the alpha helix and beta strand are summed and listed in the bottom three rows in Table 6 . It is apparent that the number of folding variations is reduced after mutation for seven regions, which indicates that the flexibilities are reduced. For the E484 mutation, the contribution of the alpha helix increased while that of the beta strand was reduced; for N501Y, the contribution of the alpha helix decreased while that of the beta strand increased; for D1118H, the factor of contribution of the alpha helix decreased while that of the beta strand increased. Therefore, the variations in both physiochemical properties and folding features for UK mutations of SARS-CoV-2 are well exposed. Mutations vs. ACE2 interaction. The RBD of the spike protein of SARS-CoV-2 binds to angiotensin-converting enzyme 2 (ACE2) receptors, serving as the entry point into human cells and causing the global coronavirus disease pandemic. 26 Thus, analysis of mutations in the RBD of the SARS-CoV-2 spike protein is significant, as it helps to explain why SARS-CoV-2 has been more dramatically widespread than SARS-CoV and benefits vaccine and drug development. The structure variation in RBD fragment, as the affinitive interface with ACE2, is focused in this study, especially the evolution from SARS-CoV to SARS-CoV-2, and to the UK variant (lineage B.1.1.7) and India variant (lineage B.1.617.2). The complete sequences of the spike proteins were obtained from the Universal Protein Resource (UniProt) database, 27 with UniProtKB P59594 for SARS-CoV (SPIKE_SARS) and UniProtKB P0DTC2 for SARS-CoV-2 (SPIKE_SARS2). Then, the mutational fragments of RBD sequences interfacing with ACE2 were aligned and displayed in Fig. 3-C and Fig. 3-D , which exposed the evolution of sequences from SARS-CoV to SARS-CoV-2, UK variant and India variant. The residues involving mutations are marked with bold font, in which SARS-CoV is black, SARS-CoV-2 blue, and UK variant and India variant red. Sequence alignment showed that the evolution of the RBD from SARS-CoV to SARS-CoV-2 involved the replacement of 9 residues; to the UK variant, two residues; and to the India variant, five residues. The protein 3D structure of the complex of the spike protein and ACE2 is available in the PDB, and images of the interaction between the spike protein and ACE2 are displayed in Fig. 3 . The protein structure with PDB ID 6ACG is the complex of the SARS-CoV spike protein and ACE2, which was deposited in July 2018; the structure with PDB ID 7A98 is the complex of the SARS-CoV-2 spike protein with ACE2, which was deposited in September 2020. The mutational residues of SARS-CoV-2 as well as the residues of ACE2 on the binding interface are marked by wire mesh to show the interpolated charged surface. It is apparent that most residues on the binding surface of ACE2 are negative charge and polar, except for K26 and K31 are positive charge. The binding surface of SARS-CoV has one residue, D480, with a negative charge facing negative residues on ACE2, and most residues of SARS-CoV are polar and without charge. After evolution from SARS-CoV to SARS-CoV-2, residue E484 near the positive residue K31 of ACE2 becomes negative, and the T501N mutation increases the polarity, which favors the interaction between SARS-CoV-2 and ACE2. All hydrogen bond (H-bond) interactions between the spike protein and ACE2 are listed in Fig. 3-A and Fig. 3-B . The distribution of H-bonds is different between SARS-CoV and SARS-CoV-2. For SARS-CoV, the residues on ACE2 involved in the H-bonds are K353, N330, Q325, Q42, Y41 and D38; for SARS-CoV-2, the residues on ACE2 involved in the H-bonds are Y83, Y41, H34, Q31 and Q24. This result indicated that the distribution of H-bonds shifted toward the N-terminus on SARS-CoV-2 compared with SARS-CoV. The change in the distribution of H-bonds is consistent with the influence of the residue 484 mutation on the spike protein from SARS-CoV to SARS-CoV-2. The evolution from SARS-CoV to SARS-CoV-2, and to the UK variant and India variant enhanced the interaction of the spike protein with ACE2. The arrows in Fig. 3 indicate the mutational residues on the binding interface of the spike protein of SARS-CoV-2. Alteration to charge residues is an important factor in virus evolution. SARS-CoV does not have an effective charge residue on the interface with ACE2. In contrast, SARS-CoV-2 has residue E484 with a negative charge near the positive K31 of ACE2. In the UK variant, the E484K mutation reverses the charge of the residue from negative to positive and triggers a folding change, and K484 interacts with the nearby negative E23 on ACE2. In the recent variant from India, although the E484Q mutation changing from charge to polar, the N440K, L452R and T478K mutations changed to residues with positive charge. N440K changed from polar to positive charge and forwarded to negative residue E329 on ACE2; L452R changed from hydrophobic to positive charge and forwarded to negative residue D38 on ACE2; T478K changed from polar to positive charge and forwarded to negative residue E22 and E23 on ACE2. Also, it is noted that K417N mutation avoided the positive charge repulsion between K417 residue of spike protein and K31 residue of ACE2. These mutations in the Indian variant increase the affinity between the spike protein and ACE2. Overall, structural mutation analysis revealed that the evolution from SARS-CoV to SARS-CoV-2, and to the UK variant and India variant enhanced the spike protein interactions with ACE2, which helps the coronavirus to infiltrate into human cells and spread more easily, leading to the coronavirus disease pandemic. Discussion The combination of sequence and protein structure fingerprint enhanced the study of mutations. In principle, sequence alignment is a useful means for studying mutations. First, it can be used to handle a large amount of data from databases with multiple sequence alignments for residue-by-residue investigation. Second, the protein structure fingerprint provides a complete description of protein folding without any gap, generating a unique sequence for alignment to study changes in folding conformation. With the protein structure fingerprint, the PFSC string as a complete folding description, which is acquired according to either the protein 3D structure or the PFVM, can cover regular secondary fragments and irregular tertiary fragments. Third, the alignment of PFSC strings is able to discover the folding structure difference caused by mutation. In addition, the combination of alignment of sequence with PFSC alphabetic string provides comprehensive analysis for mutation with integration of residue replacement and folding shape change. Moreover, the PFVM as folding variations, which is obtained directly according to protein sequence, reveals the fluctuation of the folding conformation along the sequence. It is significant that the protein structure fingerprint overcomes barriers in the study of the effects of mutations on protein structure when protein 3D structure data are absent. Thus, directly associating the protein sequence with the protein structure fingerprint is better to probe the mutations of SARS-CoV-2, which exhibits the changes on both physiochemical property and folding conformation, and provides more complete information understanding the variant in biological functions caused by mutations. The mutations in fragment at the binding interface of the RBD of the SARS-CoV-2 spike protein interacting with ACE2 are critical for coronavirus epidemic spread. From SARS-CoV to SARS-CoV-2 ( Fig. 3-A and Fig. 3-B ), the interface at the RBD was involved with at least 9 residue mutations. Before SARS-CoV-2, the residues at the interface of SARS-CoV did not have apparent charge features, and most residues were polar. After evolution to SARS-CoV-2, residue E484 with a negative charge appears nearby positive charge residue K31 of ACE2, which is one of the factors making SARS-CoV-2 a more severe disease than SARS-CoV. The UK variant has 7 residue mutations (E484K, N501Y, A570D, P681H, T716I, S982A and D1118H) at the spike protein, but only E484K and N501Y are critical because of their positions at the binding interface of the RBD, which strengthen the interaction between SARS-CoV-2 and ACE2 to increase SARS-CoV-2 infectivity. The India variant has many mutations in the spike protein, but only K417N, N440K, L452R, T478K and E484Q impact the interaction with ACE2. Although the E484Q mutation reduced the charge feature of the residue, the N440K, L452R and T478K mutations generated three positive residues near the negative residues E329, D38, E22 and E23 in ACE2, and K417N mutation reduced repulsion. Thus, the mutations in the India variant enhanced the affinity between the spike protein and ACE2 and increased viral function. As viruses undergo genetic changes, some of these genetic changes can confer evolutionary advantages, and mutations of SARS-CoV-2 at the interface with ACE2 are especially critical. In the process of evolution, many mutations occurred at different positions in the spike protein and even on other proteins of SARS-CoV-2. 28 Of course, some mutations may be neutral because they involve substitution of amino acids with physicochemical similarity; some mutations are missense because of substitution of amino acids with different physicochemical properties that change viral biological function. Understanding the structural variations caused by missense mutations in the protein epitope of SARS-CoV-2 is significant for antibody development. Thus, a protein structure fingerprint approach offers a better means to probe the mutations in SARS-CoV-2 and other viruses. SARS-CoV-2 variant may be more transmissible than previously evolved ones, so understanding structural variations is important for drug development. The structural variations caused by mutations can provide lead information for vaccine and antibody research. With new mRNA vaccine technology, short-lived synthetic fragments of the RNA sequence of a virus is introduced into the human body, where they are taken up by dendritic cells, which use their own internal ribosomes to read the mRNA and produce viral antigen proteins. The synthetic mRNA fragment is a copy of the specific part of the viral RNA that carries the instructions to build the protein spike of SARS-CoV-2. Thus, the structural variations at the binding interface of the RBD of the SARS-CoV-2 spike protein provide an important reference for designing synthetic mRNA fragments. Developing a cocktail with multiple synthetic mRNA fragments according to the mutations in the fragment at the interface of the RBD of the spike protein may be a solution to continuously counter the evolution of SARS-CoV-2. Moreover, antibody engineering requires structural data related to spike protein mutations to design the product. Antibodies contain complementarity determining regions (CDRs) for a particular epitope on an antigen, which allow these two structures to bind together with precision. Mutations in the RBD of the SARS-CoV-2 spike protein provide significant structural information for CDR design and the production of effective antibodies. Thus, understanding the structural variations, particularly at the RBD of the spike protein of SARS-CoV-2, is significant for vaccine and antibody development. Conclusion The alignment of the protein sequence and folding description reveals the structural variations caused by mutations of SARS-CoV-2. The protein structure fingerprint applies an alphabetical string to achieve a complete description of folding, which provides supplemental structural information for mutation study. The integration of changes in both physicochemical properties and folding features at the affinity interface of the RBD of the spike protein revealed how the coronavirus has gradually evolved in both structure and function and why SARS-CoV-2, the UK variant and Indian variant have led to more severe acute disease worldwide. Declarations Acknowledgments Work at the Molecular Foundry was supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231. G. Ren was partially supported by the National Heart, Lung, National Institutes of Health and Blood Institute (NIHLI), the National Institute of Mental Health (NIMH) and the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) of the National Institutes of Health under award numbers of R01HL115153, R01MH077303, and R01DK042667. Author Contributions Data curation, Peng Zhang and Wenxiang Cheng; Investigation, Gang Wu; Methodology, Gang Ren; Resources, Youyong Lu; Writing – original draft, Jiaan Yang. References Huang, Y., Yang, C., Xu, X. F., Xu, W. & Liu, S. W. Structural and functional properties of SARS-CoV-2 spike protein: potential antivirus drug development for COVID-19. Acta Pharmacol Sin , 41 , 1141–1149 https://doi.org/10.1038/s41401-020-0485-4 (2020). Harapan, H. et al. Coronavirus disease 2019 (COVID-19): A literature review. J Infect Public Health , 13 , 667–673 https://doi.org/10.1016/j.jiph.2020.03.019 (2020). Tay, M. Z., Poh, C. M., Renia, L., MacAry, P. A. & Ng, L. F. P. The trinity of COVID-19: immunity, inflammation and intervention. Nat Rev Immunol , 20 , 363–374 https://doi.org/10.1038/s41577-020-0311-8 (2020). Li, Q. et al. The Impact of Mutations in SARS-CoV-2 Spike on Viral Infectivity and Antigenicity. Cell 182 , 1284–1294 e1289, doi: 10.1016/j.cell.2020.07.012 (2020). Ni, W. et al. Role of angiotensin-converting enzyme 2 (ACE2) in COVID-19. Crit Care , 24 , 422 https://doi.org/10.1186/s13054-020-03120-0 (2020). Crackower, M. A. et al. Angiotensin-converting enzyme 2 is an essential regulator of heart function. Nature , 417 , 822–828 https://doi.org/10.1038/nature00786 (2002). Plante, J. A. et al. Spike mutation D614G alters SARS-CoV-2 fitness. Nature , 592 , 116–121 https://doi.org/10.1038/s41586-020-2895-3 (2021). Jungreis, I., Sealfon, R. & Kellis, M. SARS-CoV-2 gene content and COVID-19 mutation impact by comparing 44 Sarbecovirus genomes. Nat Commun , 12 , 2642 https://doi.org/10.1038/s41467-021-22905-7 (2021). Wise, J. Covid-19: New coronavirus variant is identified in UK. BMJ , 371 , m4857 https://doi.org/10.1136/bmj.m4857 (2020). Sievers, F. & Higgins, D. G. Clustal Omega for making accurate alignments of many protein sequences. Protein Sci , 27 , 135–145 https://doi.org/10.1002/pro.3290 (2018). Kuzmanic, A. & Zagrovic, B. Determination of ensemble-average pairwise root mean-square deviation from experimental B-factors. Biophys J , 98 , 861–871 https://doi.org/10.1016/j.bpj.2009.11.011 (2010). Cheung, N. J. & Yu, W. De novo protein structure prediction using ultra-fast molecular dynamics simulation. PloS one , 13 , e0205819 https://doi.org/10.1371/journal.pone.0205819 (2018). Pierce, L. C., Salomon-Ferrer, R., Augusto, F. O. C., McCammon, J. A. & Walker, R. C. Routine Access to Millisecond Time Scale Events with Accelerated Molecular Dynamics. J Chem Theory Comput , 8 , 2997–3002 https://doi.org/10.1021/ct300284c (2012). Zhang, Y. Progress and challenges in protein structure prediction. Curr Opin Struct Biol , 18 , 342–348 https://doi.org/10.1016/j.sbi.2008.02.004 (2008). Yang, J. Protein Structure Fingerprint Technology. J Bioinform, Genomics Proteomics , 3 , 1036 (2018). Yang, J. Comprehensive description of protein structures using protein folding shape code., 71 , 1497–1518 https://doi.org/10.1002/prot.21932 (2008). Wrapp, D. et al. Cryo-EM structure of the 2019-nCoV spike in the prefusion conformation., 367 , 1260–1263 https://doi.org/10.1126/science.abb2507 (2020). Xu, X. et al. Evolution of the novel coronavirus from the ongoing Wuhan outbreak and modeling of its spike protein for risk of human transmission. Sci China Life Sci , 63 , 457–460 https://doi.org/10.1007/s11427-020-1637-5 (2020). Millet, J. K. & Whittaker, G. R. Physiological and molecular triggers for SARS-CoV membrane fusion and entry into host cells., 517 , 3–8 https://doi.org/10.1016/j.virol.2017.12.015 (2018). Wang, H. et al. SARS coronavirus entry into host cells through a novel clathrin- and caveolae-independent endocytic pathway. Cell Res , 18 , 290–301 https://doi.org/10.1038/cr.2008.15 (2008). Song, W., Gui, M., Wang, X. & Xiang, Y. Cryo-EM structure of the SARS coronavirus spike glycoprotein in complex with its host cell receptor ACE2. PLoS Pathog , 14 , e1007236 https://doi.org/10.1371/journal.ppat.1007236 (2018). Rota, P. A. et al. Characterization of a novel coronavirus associated with severe acute respiratory syndrome., 300 , 1394–1399 https://doi.org/10.1126/science.1085952 (2003). Wu, F. et al. A new coronavirus associated with human respiratory disease in China. Nature , 579 , 265–269 https://doi.org/10.1038/s41586-020-2008-3 (2020). Galloway, S. E. et al. Emergence of SARS-CoV-2 B.1.1.7 Lineage - United States, December 29, 2020-January 12, 2021. MMWR Morb Mortal Wkly Rep , 70 , 95–99 https://doi.org/10.15585/mmwr.mm7003e2 (2021). Collier, D. A. et al. Sensitivity of SARS-CoV-2 B.1.1.7 to mRNA vaccine-elicited antibodies. Nature , 593 , 136–141 https://doi.org/10.1038/s41586-021-03412-7 (2021). Zhang, H., Penninger, J. M., Li, Y., Zhong, N. & Slutsky, A. S. Angiotensin-converting enzyme 2 (ACE2) as a SARS-CoV-2 receptor: molecular mechanisms and potential therapeutic target. Intensive Care Med , 46 , 586–590 https://doi.org/10.1007/s00134-020-05985-9 (2020). Bairoch, A. et al. The Universal Protein Resource (UniProt). Nucleic Acids Res , 33 , D154–159 https://doi.org/10.1093/nar/gki070 (2005). Wang, R. et al. Characterizing SARS-CoV-2 mutations in the United States. Res Sq , https://doi.org/10.21203/rs.3.rs-49671/v1 (2020). Tables Due to technical limitations, table 1 to 7 is only available as a download in the Supplemental Files section. Additional Declarations No competing interests reported. Supplementary Files Tables.docx Cite Share Download PDF Status: Published Journal Publication published 11 Nov, 2021 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Major revision 27 Sep, 2021 Reviews received at journal 23 Sep, 2021 Reviewers agreed at journal 20 Sep, 2021 Reviewers agreed at journal 14 Sep, 2021 Reviewers invited by journal 14 Sep, 2021 Editor assigned by journal 14 Sep, 2021 Editor invited by journal 08 Sep, 2021 Submission checks completed at journal 08 Sep, 2021 First submitted to journal 11 Aug, 2021 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-800496","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":50685440,"identity":"2520ef3b-e444-4a8e-bd21-e0ef5083dcfd","order_by":0,"name":"Jiaan Yang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAw0lEQVRIiWNgGAWjYBACxgYQWcHAA6KYSdByBqYlgWir2iA0cVqYZyQ/e/h13jYZ/tmHjz0u/MEgzy92gIAFM9LMjWW33eaROJeWbjwjgcFw5mwCVjHOSDCTlgRqMeDhMZPmSWBIMLhNUEv6N2nJOaRpyTGT/NhAkpaeN2XSDMeAfjnDlm7MkyZB2C+G7enbJH/U3Lbn72E+9pjHxkaeX5qQlgkJDMw8EDYbEEvgVw4C8vwHGBh/ILSMglEwCkbBKMAEABaWO6A+htXQAAAAAElFTkSuQmCC","orcid":"","institution":"Chinese Academy of Sciences","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Jiaan","middleName":"","lastName":"Yang","suffix":""},{"id":50685442,"identity":"cca987a4-c7f1-4d7d-8b8a-538abd972e56","order_by":1,"name":"Peng Zhang","email":"","orcid":"","institution":"Chinese Academy of Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Peng","middleName":"","lastName":"Zhang","suffix":""},{"id":50685445,"identity":"ea0f8e0d-fa9a-4911-a9c5-de749ffd3053","order_by":2,"name":"Wen Xiang Cheng","email":"","orcid":"","institution":"Chinese Academy of Sciences","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wen","middleName":"Xiang","lastName":"Cheng","suffix":""},{"id":50685448,"identity":"5e444182-8f8d-45b7-b837-f1863e3a1bad","order_by":3,"name":"Youyong Lu","email":"","orcid":"","institution":"Peking University Cancer Hospital \u0026 Institute","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Youyong","middleName":"","lastName":"Lu","suffix":""},{"id":50685450,"identity":"c417f22e-b84c-4794-8ab2-bd7ef3448ca9","order_by":4,"name":"Wu Gang","email":"","orcid":"","institution":"Huazhong University of Science and Technology","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Wu","middleName":"","lastName":"Gang","suffix":""},{"id":50685451,"identity":"9bc1438a-60fe-4b1f-913d-5c52e0dcc117","order_by":5,"name":"Gang Ren","email":"","orcid":"","institution":"Lawrence Berkeley National Laboratory","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Gang","middleName":"","lastName":"Ren","suffix":""}],"badges":[],"createdAt":"2021-08-11 08:44:05","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-800496/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-800496/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-021-01650-3","type":"published","date":"2021-11-11T11:58:25+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":13338310,"identity":"367296af-e7a4-4a32-9a88-1eaf980d2aed","added_by":"auto","created_at":"2021-09-13 21:19:17","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":92998,"visible":true,"origin":"","legend":"Comparison 3D structures between 6ACC and 6VSB. PDB ID 6ACC is the SARS-CoV spike protein, and PDB ID 6VSB is the SARS-CoV-2 spike protein. The protein 3D structural trimer, chain and domain fragment are displayed. The superposition of fragments between 6ACC-A-306-527 (blue) and 6VSB-C-319-541 (red) are shown at the bottom.","description":"","filename":"Fig01.png","url":"https://assets-eu.researchsquare.com/files/rs-800496/v1/4a68624c2fd77e357a8a72dc.png"},{"id":13338309,"identity":"89e0c6f6-c21f-4163-a15c-0be05a43aaf1","added_by":"auto","created_at":"2021-09-13 21:19:17","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":34812,"visible":true,"origin":"","legend":"The numbers of folding variations in PFVM between 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541. The horizontal coordinate is the sequence position, and the longitudinal coordinate is the number of folding shapes, i.e., number of PFSC letters. Yellow indicates the ranges with more variation in SARS-CoV, whereas green indicates more variation in SARS-CoV-2.","description":"","filename":"Fig02.png","url":"https://assets-eu.researchsquare.com/files/rs-800496/v1/e84fc7d2f86d414fc9a25286.png"},{"id":13338772,"identity":"528f9f95-e594-4bba-862a-929629e3e90c","added_by":"auto","created_at":"2021-09-13 21:22:17","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":405545,"visible":true,"origin":"","legend":"Mutations impact the interaction between SARS-CoV-2 and ACE2. The 3D images display the binding interface between the RBD of the SARS-CoV-2 spike protein and ACE2. The protein structure shown in brown color is SARS-CoV or SARS-CoV-2. The wire meshes represent the charge surfaces for residues involved in the interaction; red wire mesh indicates negative charge, and blue indicates positive charge. Row A shows the SARS-CoV structure and intermolecular H-bonds (PDB ID 6ACG); row B shows SARS-CoV-2 and H-bonds (PDB ID 7A98). Row C shows the structure of the UK variant, and row D shows the structure of the Indian variant, which were both obtained by computational modeling. The contributions of hydrogen bonds from ACE2 are marked by bold font. The arrows indicate the residues altered in viral evolution. The sequences of SARS-CoV (UniProtKB = P59594 (SPIKE_SARS)), SARS-CoV-2 (UniProtKB = P0DTC2 (SPIKE_SARS2)) and the UK variant and Indian variant are aligned and listed in row C and D, and the mutational residues are shown in colored bold font.","description":"","filename":"Fig03.png","url":"https://assets-eu.researchsquare.com/files/rs-800496/v1/6682ca5b74204ab112689bf2.png"},{"id":13338313,"identity":"7c358f27-386f-42bc-93c9-c160cdc1a962","added_by":"auto","created_at":"2021-09-13 21:19:17","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":128757,"visible":true,"origin":"","legend":"Protein structure fingerprint technology. The set of 27 protein folding shape code (PFSC) is presented in the cubic box. The blue arrows indicate how the complete conformation description with using PFSC is obtained from a protein 3D structure. The red arrows indicate how the comprehensive protein folding variations in the protein folding variation matrix (PFVM) are obtained from gene or protein sequence and expressed in PFSC description. ","description":"","filename":"Fig04.png","url":"https://assets-eu.researchsquare.com/files/rs-800496/v1/bc1c23ee46a70e241e37379b.png"},{"id":15431894,"identity":"8bff263e-cb32-408f-935f-2fc761290057","added_by":"auto","created_at":"2021-11-11 11:58:28","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":976763,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-800496/v1/4b5fb2ef-d81a-4935-97fd-d4020727355d.pdf"},{"id":13338311,"identity":"d8d38ff4-cc0f-432f-9cb8-59c9cf15715e","added_by":"auto","created_at":"2021-09-13 21:19:17","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":2753136,"visible":true,"origin":"","legend":"","description":"","filename":"Tables.docx","url":"https://assets-eu.researchsquare.com/files/rs-800496/v1/7808d8e7262f0f8a4f1abe85.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Exposing Structural Variations in SARS-CoV-2 Evolution","fulltext":[{"header":"Introduction","content":"\u003cp\u003eSevere acute respiratory syndrome coronavirus 2 (SARS-CoV-2) belongs to the coronavirus family, and it is an urgent need to explore the SARS-CoV-2 structure, function and activity.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e In particular, mutations in SARS-CoV-2 are considered a priority because of their potential to increase transmissibility and virulence while reducing the effectiveness of vaccines and impacting the development of medical treatments.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e,\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e Mutations that alter the protein sequence, including replacements or deletions of amino acid residues, may affect protein properties and folding conformations and result in changes to the biological functions of the virus.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e The interaction of the receptor-binding domain (RBD) of the spike protein of SARS-CoV-2 with angiotensin-converting enzyme 2 (ACE2) receptors is key for allowing the virus to enter human cells.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e,\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e Thus, mutations in the RBD directly influence disease development and epidemic spread.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e,\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003e To date, over 3,000 SARS-CoV-2 sequences and nearly 800 spike protein 3D structural data sources are available in the National Center for Biotechnology Information (NCBI) database and Protein Data Bank (PDB). According to the COVID-19 Genomics UK (COG-UK) Consortium, more than 4,000 mutations have been detected in the spike protein alone,\u003csup\u003e\u003cb\u003e9\u003c/b\u003e\u003c/sup\u003e which provides sufficient data to investigate coronavirus mutations to understand changes in its physiochemical properties as well as folding conformations leading to virus evolution over time. With protein sequence alignment, the positions of replaced amino acid residues can be discovered, and the concomitant changes in physiochemical properties can be further probed.\u003csup\u003e\u003cb\u003e9,10\u003c/b\u003e\u003c/sup\u003e In addition to physiochemical properties, changes in the protein folding conformation also impact biological viral functions. For proteins with known 3D structures, the conformational differences caused by mutations can be roughly compared by structure superposition with root-mean-square deviation (RMSD) as a measurement.\u003csup\u003e\u003cb\u003e11\u003c/b\u003e\u003c/sup\u003e For proteins without known 3D structures, the protein structures first need to be predicted by computational dynamics simulations. However, for mutational differences, the reliability of the predicted protein structure remains a challenge even when using \u003cem\u003eab initio\u003c/em\u003e modeling methods.\u003csup\u003e\u003cb\u003e12-14\u003c/b\u003e\u003c/sup\u003e Thus, it is crucial that a new approach overcomes these barriers to studying structural mutations.\u003c/p\u003e \u003cp\u003eAt this point, we propose using the protein structure fingerprint approach\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e,\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e to analyze the changes in folding conformation caused by mutations. With protein structure fingerprints, the protein folding shape code (PFSC) provides an alphabetical string to completely describe the folding conformation for 3D protein structure. Additionally, according to the protein sequence, the protein folding variation matrix (PFVM) reveals the folding variations along the sequence and generates the possible folding conformations. Thus, the alignment of the protein sequence with the PFSC string can comprehensively expose the variations in both biological functions and folding conformations caused by mutations in SARS-CoV-2. Here, the structural variations in evolved coronavirus strains, from SARS-CoV to SARS-CoV-2, UK variant and India variant, are studied.\u003c/p\u003e "},{"header":"Methods","content":"\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003cp\u003e \u003cb\u003eStructural bioinformation.\u003c/b\u003e All protein structural data for SARS-CoV-2 were extracted from public databases. The sequences were obtained from the NCBI and UniProt databases, and protein 3D structures were obtained from the PDB. The cd21477 cluster was obtained from the NCBI Conserved Domain Database, which contains the protein structure with PDB ID 6ACC for the SARS-CoV spike protein released in August 2018 and the protein structure with PDB ID 6VSB for the SARS-CoV-2 spike protein released in February 2020. Then, the mutations between 6ACC and 6VSB were analyzed according to either protein 3D structures or sequences by protein structure fingerprint technology. Information on the UK variant and India variant of SARS-CoV-2 was obtained from Public Health England. Seven mutations in the spike protein were identified, and the variations in physiochemical properties and folding conformations were studied. The complexes of coronavirus with ACE2 were obtained from PDB; PBD ID 6ACG is the complex with SARS-CoV, and PBD ID 7A98 is the complex with SARS-CoV-2.\u003c/p\u003e \u003cp\u003e \u003cb\u003eProtein comparison.\u003c/b\u003e The sequences of the spike protein between SARS-CoV and SARS-CoV-2 were aligned with the Clustal Omega program through UniProt and then compared according to their physiochemical properties. Discovery Studio (version 4.5) was used to generate 3D images of protein structures, and then the superimposition of protein 3D structures was performed. Furthermore, with protein structure fingerprint technology, the variations in protein folding conformations were exposed in detail.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eProtein structure fingerprint.\u003c/b\u003e First, the complete folding space for a set of 5 successive points was mathematically covered by a set of folding shapes. Second, the possible folds of a fragment of 5 amino acids could be defined by the 27 protein folding shape code (PFSC) with alphabetical letters, as shown in \u003cb\u003eFig.\u0026nbsp;4\u003c/b\u003e. Third, any protein sequence has a protein folding variation matrix (PFVM), and any protein with a given 3D structure can be expressed by a PFSC string as a protein structure conformation. In the PFSC string, two PFSC letters next each other overlap by four amino acids; thus, a PFSC string represents the complete folding conformation of the 3D protein structure. It is significant that protein folding conformations as PFSC strings can be aligned for comparison. Therefore, the folding variations of SARS-CoV-2 as well as its mutations may be well analyzed by protein structure fingerprints.\u003c/p\u003e \u003cp\u003e \u003cb\u003eSoftware availability.\u003c/b\u003e The protein structure fingerprint can be accessed on Website \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://www.micropht.com\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e"},{"header":"Results","content":" \u003cp\u003eThe changes in both physiochemical properties and folding conformations of SARS-CoV-2 due to mutation are studied based on the protein 3D structures and sequences of the spike protein, and the interaction between coronavirus and ACE2 are a particular focus. The structural analysis covers coronavirus strains from early SARS-CoV to SARS-CoV-2, UK variant and recent India variant.\u003c/p\u003e \u003cp\u003e \u003cb\u003eVariations based on 3D structures.\u003c/b\u003e Coronavirus spike proteins have an S1 subunit at the N-terminus (~700 amino acids) and anS2 subunit at the C-terminus (~600 amino acids). Analysis of many protein 3D structures confirmed that three S1/S2 heterodimers assembled to form a trimer spike protruding from the viral envelope.\u003csup\u003e\u003cb\u003e17\u003c/b\u003e\u003c/sup\u003eThe S1 subunit of the spike protein in SARS-CoV-2 is an envelope glycoprotein that plays the most important role in viral attachment, fusion, and entry into host cells, and is a major target for the development of neutralizing antibodies, inhibitors and vaccines. The S1 subunit contains a receptor-binding domain (RBD), and many studies have found that the RBD of the spike protein in SARS-CoV-2 strongly binds to human and bat angiotensin- converting enzyme 2 (ACE2) receptors.\u003csup\u003e\u003cb\u003e18-20\u003c/b\u003e\u003c/sup\u003e A set of sequences for the cluster of the SARS-CoV-like_Spike_S1_RBD subfamily (cd21477) that contains the conserved protein domain of the S1 RBD subfamily for SARS-CoV-like and SARS-CoV-2 spike proteins is available in the NCBI database. The sequences of the cd21477 cluster were aligned and are presented in \u003cb\u003eTable 1\u003c/b\u003e, where the red color font indicates highly conserved fragments, blue indicates less conserved fragments and gray indicates unaligned fragments. It is not surprising that the mutations were most frequent on less conserved residues (blue font). Additionally, it is noted that some sequences have the given 3D structures in the PDB. The protein structure of the SARS-CoV spike protein, with PDB ID 6ACC, was released in August 2018;\u003csup\u003e\u003cb\u003e21\u003c/b\u003e\u003c/sup\u003e the protein structure of the SARS-CoV-2 spike protein, with PDB ID 6VSB, was released in February 2020.\u003csup\u003e\u003cb\u003e17\u003c/b\u003e\u003c/sup\u003e The residues that differ between 6ACC and 6VSB are marked in green. Some changes in the physiochemical properties of the protein based on the differences between 6ACC to 6VSB are summarized in \u003cb\u003eTable 2\u003c/b\u003e, including hydrophobicity, negative or positive charge, polarity, size of side chain, aromatic and \u003cem\u003eetc\u003c/em\u003e. The changes in the physiochemical properties are represented by a \u0026ldquo;+\u0026rdquo; sign for an increase in the property after mutation and a \u0026ldquo;-\u0026rdquo; sign for a decrease in the property after mutation.\u003c/p\u003e \u003cp\u003eFrom the given structures of 6ACC and 6VSB, 3D images of folding conformations are compared and displayed in \u003cb\u003eFig. 1\u003c/b\u003e. The 3D structures directly provide a visualization to observe the protein structures, and the superposition allows a comparison of the structures. Although more than 30 mutations in the fragment between 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541 for SARS-CoV-2 occurred, representing up to 37.5% residue replacement, the structure superposition showed that the folding conformations of 6ACC and 6VSB were still similar overall. It is difficult to distinguish the folding differences of spike proteins of SARS-CoV and SARS-CoV-2 based on the 3D structure only. With the protein folding shape code (PFSC), however, the differences in folding can be exposed. Any protein 3D structure can be converted into a PFSC description, which is an alphabetical string representing the continuous folding shape of each five-amino-acid in sequence. Thus, the folding conformations of 6ACC for SARS-CoV and 6VSB for SARS-CoV-2 can be compared by PFSC alignment and displayed in \u003cb\u003eTable 3\u003c/b\u003e. In PFSC, generally, the red color indicates typical alpha helix, pink indicates alpha-like helix, blue indicates a typical beta strand, light blue indicates a beta-like strand, and black indicates an irregular fold. According tothe PFSC color notation, it is obvious that the secondary structural fragments are well aligned. For example, the fragments of alpha helices at 324-330 and 352-364 and the beta strands at 349-351 and 378-386 on 6ACC are aligned with the corresponding fragments in 6VSB. Also, the PFSC alignment exposes local folding comparison in detail, which the local folding similarity and differences between PFSC strings are indicated; \u0026ldquo;|\u0026rdquo; indicates an identical folding shape, \u0026ldquo;:\u0026rdquo; indicates a similar folding shape, and \u0026ldquo;.\u0026rdquo; indicates dissimilar folding. For example, the folding letters at 334, 335, 340, 341, 343, 370, 379 and 380 on 6ACC are different from 6VSB. Also, due to mutations, the adjustments of beta strand at 349-351 and 452-454 fragments on 6ACC are exposed. Thus, the PFSC revealed the changes in local folding shapes caused by the mutations.\u003c/p\u003e \u003cp\u003e \u003cb\u003eVariations based on sequences.\u003c/b\u003e The variations of folding conformations for a protein based on sequence alone can be exposed by the PFVM. The PFVMs for sequences taken directly from the structures of 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541 for SARS-CoV-2 separately are exhibited in \u003cb\u003eTable\u0026nbsp;4\u003c/b\u003e. The PFSC letters in each column represent the folding variations of 5 successive amino acid residues in sequence while the favored folding shapes are ranked on top, and the numbers are different in each column. The PFVM exhibits the folding variations along the sequence. The number deviations of folding shapes along sequences in PFVM between 6ACC-A-306-527 for SARS-CoV and 6VSB-C-319-541 for SARS-CoV-2 are shown by the curves in \u003cb\u003eFig.\u0026nbsp;2\u003c/b\u003e, where the yellow and green blocks indicate the regions of fluctuation due to mutations. It was apparent that the mutations caused changes in folding flexibility; some fragments have the potential to be more flexible, and other fragments are more rigid. Thus, along the sequence from the N-terminus to the C-terminus, the variations in the folding conformation are well exposed.\u003c/p\u003e \u003cp\u003eThe most likely conformations for a protein can be predicted from PFVM. Taking one letter from each column, a massive number of PFSC strings can be formed, and each string is one possible folding conformation. Although a large number of folding conformations exist, the letters on top of each column are directly constructed into one of the most likely conformations, which is named PFVM-01. This predicted conformation may be assessed by a given 3D protein structure through PFSC alignment. Two PFSC strings of the first row from \u003cb\u003eTable\u0026nbsp;4\u003c/b\u003e for 6ACC-A-306-527 and 6VSB-C-319-541 and two PFSC strings of their 3D structures from \u003cb\u003eTable\u0026nbsp;3\u003c/b\u003e are aligned in \u003cb\u003eTable\u0026nbsp;5\u003c/b\u003e. The PFSC letters in red and pink colors represent alpha helices, those in blue and light blue represent beta strands and those in black represent irregular folding shapes. Overall, with observation, the secondary fragments are aligned, so the predicted folding conformations of PFVM-01 for SARS-CoV and SARS-CoV-2 are similar to the given 3D structures. Thus, the PFVM-01 as the most likely conformation, which is generated from PFVM, is a reliable prediction.\u003c/p\u003e \u003cp\u003e \u003cb\u003eStructure variations with virus evolution.\u003c/b\u003e The sequence of the SARS-CoV spike protein (UniProtKB\u0026thinsp;=\u0026thinsp;P59594 (SPIKE_SARS)) was first determined in 2003.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e The sequence of the SARS-CoV-2 spike protein (UniProtKB\u0026thinsp;=\u0026thinsp;P0DTC2 (SPIKE_SARS2)) was determined in January 2020.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e After 17 years of evolution from SARS-CoV to SARS-CoV-2, the spike protein sequences are approximately 24% different. The UK variant is mutant of SARS-CoV-2 that was noted in November 2020 from a sample taken in the UK in September. SARS-CoV-2 infections in the UK increased because of one or more mutations in the virus spike protein. The lineage of the UK mutant B.1.1.7 VOC-202012/01 is taken from the Public Health England,\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e which was reported on March 5, 2021, with seven mutations in the spike protein: E484K, N501Y, A570D, P681H, T716I, S982A and D1118H.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e Similar variants have also emerged in South Africa (lineage B.1.351) and Brazil (lineage P.1). Recently, India variant (lineage B.1.617.2) with mutations K417N, N440K, L452R, T478K and E484Q in the spike protein caused the epidemic to become severe. Thus, it is important to understand the effects of the mutations following virus evolution.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eIn order to study the mutations in SARS-CoV-2, a sequence of QTJ15692 (GenBank) was taken as the background reference, which was deposited in the NCBI database on April 2020 before the UK mutation. The mutations may cause changes in physicochemical functions as well as in folding conformations, which together impact biological functions. The changes in physiochemical properties, including hydrophobicity, negative or positive charge, polarity, residue size and aromaticity, are listed in the top rows of \u003cb\u003eTable\u0026nbsp;6\u003c/b\u003e. For example, the mutation E484K changed a negative charge to a positive charge; A570D is a change from hydrophobic to negative, from non-charged to charged and from non-polar to polar; P681H is a change from hydrophobic to positive, from non-charged to charged and from non-polar to polar and an increase in the size of the side chain due to an aromatic moiety; T716I and S982A are changes from polar to hydrophobic. The changes in physiochemical properties caused by mutations are indicated in detail. Furthermore, the variations in local folding shapes may be revealed by PFVM because each PFSC letter in PFVM represents the folding shape of 5 successive amino acids in sequence. The PFVMs of seven regions for these related mutations are displayed in \u003cb\u003eTable\u0026nbsp;7\u003c/b\u003e, which shows the folding variations before and after mutations. To compare each pair of PFVMs, the fluctuations in the number of folding shapes and the contributions of the alpha helix and beta strand are summed and listed in the bottom three rows in \u003cb\u003eTable\u0026nbsp;6\u003c/b\u003e. It is apparent that the number of folding variations is reduced after mutation for seven regions, which indicates that the flexibilities are reduced. For the E484 mutation, the contribution of the alpha helix increased while that of the beta strand was reduced; for N501Y, the contribution of the alpha helix decreased while that of the beta strand increased; for D1118H, the factor of contribution of the alpha helix decreased while that of the beta strand increased. Therefore, the variations in both physiochemical properties and folding features for UK mutations of SARS-CoV-2 are well exposed.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cb\u003eMutations vs. ACE2 interaction.\u003c/b\u003e The RBD of the spike protein of SARS-CoV-2 binds to angiotensin-converting enzyme 2 (ACE2) receptors, serving as the entry point into human cells and causing the global coronavirus disease pandemic.\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e Thus, analysis of mutations in the RBD of the SARS-CoV-2 spike protein is significant, as it helps to explain why SARS-CoV-2 has been more dramatically widespread than SARS-CoV and benefits vaccine and drug development. The structure variation in RBD fragment, as the affinitive interface with ACE2, is focused in this study, especially the evolution from SARS-CoV to SARS-CoV-2, and to the UK variant (lineage B.1.1.7) and India variant (lineage B.1.617.2). The complete sequences of the spike proteins were obtained from the Universal Protein Resource (UniProt) database,\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e with UniProtKB P59594 for SARS-CoV (SPIKE_SARS) and UniProtKB P0DTC2 for SARS-CoV-2 (SPIKE_SARS2). Then, the mutational fragments of RBD sequences interfacing with ACE2 were aligned and displayed in \u003cb\u003eFig.\u0026nbsp;3-C\u003c/b\u003e and \u003cb\u003eFig.\u0026nbsp;3-D\u003c/b\u003e, which exposed the evolution of sequences from SARS-CoV to SARS-CoV-2, UK variant and India variant. The residues involving mutations are marked with bold font, in which SARS-CoV is black, SARS-CoV-2 blue, and UK variant and India variant red. Sequence alignment showed that the evolution of the RBD from SARS-CoV to SARS-CoV-2 involved the replacement of 9 residues; to the UK variant, two residues; and to the India variant, five residues.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThe protein 3D structure of the complex of the spike protein and ACE2 is available in the PDB, and images of the interaction between the spike protein and ACE2 are displayed in \u003cb\u003eFig.\u0026nbsp;3\u003c/b\u003e. The protein structure with PDB ID 6ACG is the complex of the SARS-CoV spike protein and ACE2, which was deposited in July 2018; the structure with PDB ID 7A98 is the complex of the SARS-CoV-2 spike protein with ACE2, which was deposited in September 2020. The mutational residues of SARS-CoV-2 as well as the residues of ACE2 on the binding interface are marked by wire mesh to show the interpolated charged surface. It is apparent that most residues on the binding surface of ACE2 are negative charge and polar, except for K26 and K31 are positive charge. The binding surface of SARS-CoV has one residue, D480, with a negative charge facing negative residues on ACE2, and most residues of SARS-CoV are polar and without charge. After evolution from SARS-CoV to SARS-CoV-2, residue E484 near the positive residue K31 of ACE2 becomes negative, and the T501N mutation increases the polarity, which favors the interaction between SARS-CoV-2 and ACE2. All hydrogen bond (H-bond) interactions between the spike protein and ACE2 are listed in \u003cb\u003eFig.\u0026nbsp;3-A\u003c/b\u003e and \u003cb\u003eFig.\u0026nbsp;3-B\u003c/b\u003e. The distribution of H-bonds is different between SARS-CoV and SARS-CoV-2. For SARS-CoV, the residues on ACE2 involved in the H-bonds are K353, N330, Q325, Q42, Y41 and D38; for SARS-CoV-2, the residues on ACE2 involved in the H-bonds are Y83, Y41, H34, Q31 and Q24. This result indicated that the distribution of H-bonds shifted toward the N-terminus on SARS-CoV-2 compared with SARS-CoV. The change in the distribution of H-bonds is consistent with the influence of the residue 484 mutation on the spike protein from SARS-CoV to SARS-CoV-2.\u003c/p\u003e \u003cp\u003eThe evolution from SARS-CoV to SARS-CoV-2, and to the UK variant and India variant enhanced the interaction of the spike protein with ACE2. The arrows in \u003cb\u003eFig.\u0026nbsp;3\u003c/b\u003e indicate the mutational residues on the binding interface of the spike protein of SARS-CoV-2. Alteration to charge residues is an important factor in virus evolution. SARS-CoV does not have an effective charge residue on the interface with ACE2. In contrast, SARS-CoV-2 has residue E484 with a negative charge near the positive K31 of ACE2. In the UK variant, the E484K mutation reverses the charge of the residue from negative to positive and triggers a folding change, and K484 interacts with the nearby negative E23 on ACE2. In the recent variant from India, although the E484Q mutation changing from charge to polar, the N440K, L452R and T478K mutations changed to residues with positive charge. N440K changed from polar to positive charge and forwarded to negative residue E329 on ACE2; L452R changed from hydrophobic to positive charge and forwarded to negative residue D38 on ACE2; T478K changed from polar to positive charge and forwarded to negative residue E22 and E23 on ACE2. Also, it is noted that K417N mutation avoided the positive charge repulsion between K417 residue of spike protein and K31 residue of ACE2. These mutations in the Indian variant increase the affinity between the spike protein and ACE2. Overall, structural mutation analysis revealed that the evolution from SARS-CoV to SARS-CoV-2, and to the UK variant and India variant enhanced the spike protein interactions with ACE2, which helps the coronavirus to infiltrate into human cells and spread more easily, leading to the coronavirus disease pandemic.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe combination of sequence and protein structure fingerprint enhanced the study of mutations. In principle, sequence alignment is a useful means for studying mutations. First, it can be used to handle a large amount of data from databases with multiple sequence alignments for residue-by-residue investigation. Second, the protein structure fingerprint provides a complete description of protein folding without any gap, generating a unique sequence for alignment to study changes in folding conformation. With the protein structure fingerprint, the PFSC string as a complete folding description, which is acquired according to either the protein 3D structure or the PFVM, can cover regular secondary fragments and irregular tertiary fragments. Third, the alignment of PFSC strings is able to discover the folding structure difference caused by mutation. In addition, the combination of alignment of sequence with PFSC alphabetic string provides comprehensive analysis for mutation with integration of residue replacement and folding shape change. Moreover, the PFVM as folding variations, which is obtained directly according to protein sequence, reveals the fluctuation of the folding conformation along the sequence. It is significant that the protein structure fingerprint overcomes barriers in the study of the effects of mutations on protein structure when protein 3D structure data are absent. Thus, directly associating the protein sequence with the protein structure fingerprint is better to probe the mutations of SARS-CoV-2, which exhibits the changes on both physiochemical property and folding conformation, and provides more complete information understanding the variant in biological functions caused by mutations.\u003c/p\u003e \u003cp\u003eThe mutations in fragment at the binding interface of the RBD of the SARS-CoV-2 spike protein interacting with ACE2 are critical for coronavirus epidemic spread. From SARS-CoV to SARS-CoV-2 (\u003cb\u003eFig. 3-A\u003c/b\u003e and \u003cb\u003eFig. 3-B\u003c/b\u003e), the interface at the RBD was involved with at least 9 residue mutations. Before SARS-CoV-2, the residues at the interface of SARS-CoV did not have apparent charge features, and most residues were polar. After evolution to SARS-CoV-2, residue E484 with a negative charge appears nearby positive charge residue K31 of ACE2, which is one of the factors making SARS-CoV-2 a more severe disease than SARS-CoV. The UK variant has 7 residue mutations (E484K, N501Y, A570D, P681H, T716I, S982A and D1118H) at the spike protein, but only E484K and N501Y are critical because of their positions at the binding interface of the RBD, which strengthen the interaction between SARS-CoV-2 and ACE2 to increase SARS-CoV-2 infectivity. The India variant has many mutations in the spike protein, but only K417N, N440K, L452R, T478K and E484Q impact the interaction with ACE2. Although the E484Q mutation reduced the charge feature of the residue, the N440K, L452R and T478K mutations generated three positive residues near the negative residues E329, D38, E22 and E23 in ACE2, and K417N mutation reduced repulsion. Thus, the mutations in the India variant enhanced the affinity between the spike protein and ACE2 and increased viral function. As viruses undergo genetic changes, some of these genetic changes can confer evolutionary advantages, and mutations of SARS-CoV-2 at the interface with ACE2 are especially critical. In the process of evolution, many mutations occurred at different positions in the spike protein and even on other proteins of SARS-CoV-2.\u003csup\u003e\u003cb\u003e28\u003c/b\u003e\u003c/sup\u003e Of course, some mutations may be neutral because they involve substitution of amino acids with physicochemical similarity; some mutations are missense because of substitution of amino acids with different physicochemical properties that change viral biological function. Understanding the structural variations caused by missense mutations in the protein epitope of SARS-CoV-2 is significant for antibody development. Thus, a protein structure fingerprint approach offers a better means to probe the mutations in SARS-CoV-2 and other viruses.\u003c/p\u003e \u003cp\u003eSARS-CoV-2 variant may be more transmissible than previously evolved ones, so understanding structural variations is important for drug development. The structural variations caused by mutations can provide lead information for vaccine and antibody research. With new mRNA vaccine technology, short-lived synthetic fragments of the RNA sequence of a virus is introduced into the human body, where they are taken up by dendritic cells, which use their own internal ribosomes to read the mRNA and produce viral antigen proteins. The synthetic mRNA fragment is a copy of the specific part of the viral RNA that carries the instructions to build the protein spike of SARS-CoV-2. Thus, the structural variations at the binding interface of the RBD of the SARS-CoV-2 spike protein provide an important reference for designing synthetic mRNA fragments. Developing a cocktail with multiple synthetic mRNA fragments according to the mutations in the fragment at the interface of the RBD of the spike protein may be a solution to continuously counter the evolution of SARS-CoV-2. Moreover, antibody engineering requires structural data related to spike protein mutations to design the product. Antibodies contain complementarity determining regions (CDRs) for a particular epitope on an antigen, which allow these two structures to bind together with precision. Mutations in the RBD of the SARS-CoV-2 spike protein provide significant structural information for CDR design and the production of effective antibodies. Thus, understanding the structural variations, particularly at the RBD of the spike protein of SARS-CoV-2, is significant for vaccine and antibody development.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThe alignment of the protein sequence and folding description reveals the structural variations caused by mutations of SARS-CoV-2. The protein structure fingerprint applies an alphabetical string to achieve a complete description of folding, which provides supplemental structural information for mutation study. The integration of changes in both physicochemical properties and folding features at the affinity interface of the RBD of the spike protein revealed how the coronavirus has gradually evolved in both structure and function and why SARS-CoV-2, the UK variant and Indian variant have led to more severe acute disease worldwide.\u003c/p\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAcknowledgments\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWork at the Molecular Foundry was supported by the Office of Science of the U.S. Department of Energy under Contract No. DE-AC02-05CH11231. G. Ren was partially supported by the National Heart, Lung, National Institutes of Health and Blood Institute (NIHLI), the National Institute of Mental Health (NIMH) and the National Institute of Diabetes and Digestive and Kidney Diseases (NIDDK) of the National Institutes of Health under award numbers of R01HL115153, R01MH077303, and R01DK042667.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eData curation, Peng Zhang and Wenxiang Cheng; Investigation, Gang Wu; Methodology, Gang Ren; Resources, Youyong Lu; Writing \u0026ndash; original draft, Jiaan Yang.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eHuang, Y., Yang, C., Xu, X. F., Xu, W. \u0026amp; Liu, S. W. Structural and functional properties of SARS-CoV-2 spike protein: potential antivirus drug development for COVID-19. \u003cem\u003eActa Pharmacol Sin\u003c/em\u003e, \u003cb\u003e41\u003c/b\u003e, 1141\u0026ndash;1149 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41401-020-0485-4\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarapan, H. \u003cem\u003eet al.\u003c/em\u003e Coronavirus disease 2019 (COVID-19): A literature review. \u003cem\u003eJ Infect Public Health\u003c/em\u003e, \u003cb\u003e13\u003c/b\u003e, 667\u0026ndash;673 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jiph.2020.03.019\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTay, M. Z., Poh, C. M., Renia, L., MacAry, P. A. \u0026amp; Ng, L. F. P. The trinity of COVID-19: immunity, inflammation and intervention. \u003cem\u003eNat Rev Immunol\u003c/em\u003e, \u003cb\u003e20\u003c/b\u003e, 363\u0026ndash;374 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41577-020-0311-8\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi, Q. \u003cem\u003eet al.\u003c/em\u003e The Impact of Mutations in SARS-CoV-2 Spike on Viral Infectivity and Antigenicity. \u003cem\u003eCell\u003c/em\u003e \u003cb\u003e182\u003c/b\u003e, 1284\u0026ndash;1294 e1289, doi:\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cell.2020.07.012\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNi, W. \u003cem\u003eet al.\u003c/em\u003e Role of angiotensin-converting enzyme 2 (ACE2) in COVID-19. \u003cem\u003eCrit Care\u003c/em\u003e, \u003cb\u003e24\u003c/b\u003e, 422 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1186/s13054-020-03120-0\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCrackower, M. A. \u003cem\u003eet al.\u003c/em\u003e Angiotensin-converting enzyme 2 is an essential regulator of heart function. \u003cem\u003eNature\u003c/em\u003e, \u003cb\u003e417\u003c/b\u003e, 822\u0026ndash;828 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/nature00786\u003c/span\u003e\u003c/span\u003e (2002).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePlante, J. A. \u003cem\u003eet al.\u003c/em\u003e Spike mutation D614G alters SARS-CoV-2 fitness. \u003cem\u003eNature\u003c/em\u003e, \u003cb\u003e592\u003c/b\u003e, 116\u0026ndash;121 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41586-020-2895-3\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJungreis, I., Sealfon, R. \u0026amp; Kellis, M. SARS-CoV-2 gene content and COVID-19 mutation impact by comparing 44 Sarbecovirus genomes. \u003cem\u003eNat Commun\u003c/em\u003e, \u003cb\u003e12\u003c/b\u003e, 2642 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41467-021-22905-7\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWise, J. Covid-19: New coronavirus variant is identified in UK. \u003cem\u003eBMJ\u003c/em\u003e, \u003cb\u003e371\u003c/b\u003e, m4857 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1136/bmj.m4857\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSievers, F. \u0026amp; Higgins, D. G. Clustal Omega for making accurate alignments of many protein sequences. \u003cem\u003eProtein Sci\u003c/em\u003e, \u003cb\u003e27\u003c/b\u003e, 135\u0026ndash;145 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/pro.3290\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKuzmanic, A. \u0026amp; Zagrovic, B. Determination of ensemble-average pairwise root mean-square deviation from experimental B-factors. \u003cem\u003eBiophys J\u003c/em\u003e, \u003cb\u003e98\u003c/b\u003e, 861\u0026ndash;871 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.bpj.2009.11.011\u003c/span\u003e\u003c/span\u003e (2010).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheung, N. J. \u0026amp; Yu, W. De novo protein structure prediction using ultra-fast molecular dynamics simulation. \u003cem\u003ePloS one\u003c/em\u003e, \u003cb\u003e13\u003c/b\u003e, e0205819 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.pone.0205819\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePierce, L. C., Salomon-Ferrer, R., Augusto, F. O. C., McCammon, J. A. \u0026amp; Walker, R. C. Routine Access to Millisecond Time Scale Events with Accelerated Molecular Dynamics. \u003cem\u003eJ Chem Theory Comput\u003c/em\u003e, \u003cb\u003e8\u003c/b\u003e, 2997\u0026ndash;3002 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1021/ct300284c\u003c/span\u003e\u003c/span\u003e (2012).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, Y. Progress and challenges in protein structure prediction. \u003cem\u003eCurr Opin Struct Biol\u003c/em\u003e, \u003cb\u003e18\u003c/b\u003e, 342\u0026ndash;348 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.sbi.2008.02.004\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, J. Protein Structure Fingerprint Technology. \u003cem\u003eJ Bioinform, Genomics Proteomics\u003c/em\u003e, \u003cb\u003e3\u003c/b\u003e, 1036 (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang, J. Comprehensive description of protein structures using protein folding shape code., \u003cb\u003e71\u003c/b\u003e, 1497\u0026ndash;1518 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1002/prot.21932\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWrapp, D. \u003cem\u003eet al.\u003c/em\u003e Cryo-EM structure of the 2019-nCoV spike in the prefusion conformation., \u003cb\u003e367\u003c/b\u003e, 1260\u0026ndash;1263 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.abb2507\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu, X. \u003cem\u003eet al.\u003c/em\u003e Evolution of the novel coronavirus from the ongoing Wuhan outbreak and modeling of its spike protein for risk of human transmission. \u003cem\u003eSci China Life Sci\u003c/em\u003e, \u003cb\u003e63\u003c/b\u003e, 457\u0026ndash;460 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11427-020-1637-5\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMillet, J. K. \u0026amp; Whittaker, G. R. Physiological and molecular triggers for SARS-CoV membrane fusion and entry into host cells., \u003cb\u003e517\u003c/b\u003e, 3\u0026ndash;8 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.virol.2017.12.015\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, H. \u003cem\u003eet al.\u003c/em\u003e SARS coronavirus entry into host cells through a novel clathrin- and caveolae-independent endocytic pathway. \u003cem\u003eCell Res\u003c/em\u003e, \u003cb\u003e18\u003c/b\u003e, 290\u0026ndash;301 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/cr.2008.15\u003c/span\u003e\u003c/span\u003e (2008).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong, W., Gui, M., Wang, X. \u0026amp; Xiang, Y. Cryo-EM structure of the SARS coronavirus spike glycoprotein in complex with its host cell receptor ACE2. \u003cem\u003ePLoS Pathog\u003c/em\u003e, \u003cb\u003e14\u003c/b\u003e, e1007236 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1371/journal.ppat.1007236\u003c/span\u003e\u003c/span\u003e (2018).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRota, P. A. \u003cem\u003eet al.\u003c/em\u003e Characterization of a novel coronavirus associated with severe acute respiratory syndrome., \u003cb\u003e300\u003c/b\u003e, 1394\u0026ndash;1399 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.1085952\u003c/span\u003e\u003c/span\u003e (2003).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWu, F. \u003cem\u003eet al.\u003c/em\u003e A new coronavirus associated with human respiratory disease in China. \u003cem\u003eNature\u003c/em\u003e, \u003cb\u003e579\u003c/b\u003e, 265\u0026ndash;269 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41586-020-2008-3\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGalloway, S. E. \u003cem\u003eet al.\u003c/em\u003e Emergence of SARS-CoV-2 B.1.1.7 Lineage - United States, December 29, 2020-January 12, 2021. \u003cem\u003eMMWR Morb Mortal Wkly Rep\u003c/em\u003e, \u003cb\u003e70\u003c/b\u003e, 95\u0026ndash;99 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.15585/mmwr.mm7003e2\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCollier, D. A. \u003cem\u003eet al.\u003c/em\u003e Sensitivity of SARS-CoV-2 B.1.1.7 to mRNA vaccine-elicited antibodies. \u003cem\u003eNature\u003c/em\u003e, \u003cb\u003e593\u003c/b\u003e, 136\u0026ndash;141 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41586-021-03412-7\u003c/span\u003e\u003c/span\u003e (2021).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhang, H., Penninger, J. M., Li, Y., Zhong, N. \u0026amp; Slutsky, A. S. Angiotensin-converting enzyme 2 (ACE2) as a SARS-CoV-2 receptor: molecular mechanisms and potential therapeutic target. \u003cem\u003eIntensive Care Med\u003c/em\u003e, \u003cb\u003e46\u003c/b\u003e, 586\u0026ndash;590 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00134-020-05985-9\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBairoch, A. \u003cem\u003eet al.\u003c/em\u003e The Universal Protein Resource (UniProt). \u003cem\u003eNucleic Acids Res\u003c/em\u003e, \u003cb\u003e33\u003c/b\u003e, D154\u0026ndash;159 \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1093/nar/gki070\u003c/span\u003e\u003c/span\u003e (2005).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang, R. \u003cem\u003eet al.\u003c/em\u003e Characterizing SARS-CoV-2 mutations in the United States. \u003cem\u003eRes Sq\u003c/em\u003e, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.21203/rs.3.rs-49671/v1\u003c/span\u003e\u003c/span\u003e (2020).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eDue to technical limitations, table 1 to 7 is only available as a download in the Supplemental Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Mutation, SARS-CoV-2, spike protein, physiochemical properties, folding conformation","lastPublishedDoi":"10.21203/rs.3.rs-800496/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-800496/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe \u003ca href=\"https://en.wikipedia.org/wiki/Mutation\" rel=\"noopener noreferrer\" target=\"_blank\"\u003emutation\u003c/a\u003e\u0026nbsp;of \u003ca href=\"https://en.wikipedia.org/wiki/SARS-CoV-2\" rel=\"noopener noreferrer\" target=\"_blank\"\u003eSARS-CoV-2\u003c/a\u003e\u0026nbsp;influences viral function as residue replacements affect both physiochemical properties and folding conformations. Although a large amount of data on \u003ca href=\"https://en.wikipedia.org/wiki/SARS-CoV-2\" rel=\"noopener noreferrer\" target=\"_blank\"\u003eSARS-CoV-2\u003c/a\u003e is available, the investigation of how viral functions change in response to mutations is hampered by a lack of effective structural analysis. Here, we exploit advances in protein structure fingerprint technology to study the folding conformational changes induced by mutations. With the integration of both protein sequences and folding conformations and alignments of \u003ca href=\"https://en.wikipedia.org/wiki/SARS-CoV-2\" rel=\"noopener noreferrer\" target=\"_blank\"\u003eSARS-CoV to SARS-CoV-2\u003c/a\u003e, the UK variant and India variant, we found that structural variations in the spike protein at the binding interface interacting with ACE2 play a critical role in coronavirus entry into human cells. Additionally, the structural variations impact vaccine effectiveness and drug function over the course of \u003ca href=\"https://en.wikipedia.org/wiki/SARS-CoV-2\" rel=\"noopener noreferrer\" target=\"_blank\"\u003eSARS-CoV-2\u003c/a\u003e evolution. The analysis of structural variations revealed how the coronavirus has gradually evolved in both structure and function and how the \u003ca href=\"https://en.wikipedia.org/wiki/SARS-CoV-2\" rel=\"noopener noreferrer\" target=\"_blank\"\u003eSARS-CoV-2\u003c/a\u003e variants have contributed to more severe acute disease worldwide.\u003c/p\u003e","manuscriptTitle":"Exposing Structural Variations in SARS-CoV-2 Evolution","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2021-09-13 21:19:15","doi":"10.21203/rs.3.rs-800496/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revision","date":"2021-09-27T09:23:13+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2021-09-23T04:25:24+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"14dac9a7-fde3-472e-b800-e3e3cd58c568","date":"2021-09-20T12:04:41+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"32437df6-d59b-48a3-8c3d-1b33207897fb","date":"2021-09-14T23:31:55+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2021-09-14T19:10:07+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2021-09-14T13:02:07+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2021-09-08T12:24:19+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2021-09-08T08:16:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2021-08-11T08:34:13+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d77930cd-0afa-406f-9e6e-cb77f5eaca17","owner":[],"postedDate":"September 13th, 2021","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":7156475,"name":"Infectious Diseases"},{"id":7156476,"name":"Structural Biology"}],"tags":[],"updatedAt":"2021-11-11T11:58:25+00:00","versionOfRecord":{"articleIdentity":"rs-800496","link":"https://doi.org/10.1038/s41598-021-01650-3","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2021-11-11 11:58:25","publishedOnDateReadable":"November 11th, 2021"},"versionCreatedAt":"2021-09-13 21:19:15","video":"","vorDoi":"10.1038/s41598-021-01650-3","vorDoiUrl":"https://doi.org/10.1038/s41598-021-01650-3","workflowStages":[]},"version":"v1","identity":"rs-800496","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-800496","identity":"rs-800496","version":["v1"]},"buildId":"rHA-KDH7Qsr4HCuvH75dn","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.