Mass spectrometry-based thermostability profiling of virus-derived MHC peptide complexes serves as an effective predictor of immunogenicity | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Mass spectrometry-based thermostability profiling of virus-derived MHC peptide complexes serves as an effective predictor of immunogenicity Anthony Purcell, Mohammad Shahbazy, Sri Ramarathinam, David Tscharke, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5824434/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract The major histocompatibility complex (MHC) encodes molecules that present peptides on the surface of cells to stimulate T-cell-mediated immune responses. The stability of peptide-MHC class I complexes (pMHCI) has been postulated to influence the immunogenicity of virus-derived epitopes and cancer neoepitopes. Here, we sought to investigate this further by conducting thermostability profiling of thousands of individual pMHCI, including a panel of 110 vaccinia virus (VACV) derived peptides with known CD8 + T cell response profiles. The denaturation profiles of these peptides spanned thermostability (T m ) ranges of 41.2°C to 65.1°C, and we found that thermostability correlated with immunogenicity in VACV-infected mice. We developed two machine learning-based models from these thermostability data to predict peptide immunogenicity and demonstrate the ability of this model to distinguish immunogenic epitopes derived from an unrelated infectious pathogen, influenza A virus in mice. Using such models, we provide evidence that the thermostability of pMHCI allows for improved prediction of immunogenic CD8 + T cell epitopes and conclude that this information is a valuable measurement for selecting optimal targets for T cell-mediated therapies and vaccine design. Biological sciences/Immunology/Adaptive immunity/Cellular immunity/Antigen presentation Biological sciences/Immunology/Vaccines/Peptide vaccines Biological sciences/Immunology/Antigen processing and presentation/Cellular immunity Figures Figure 1 Figure 2 Figure 3 Figure 4 Introduction Major histocompatibility complex (MHC) molecules present peptide antigens on the surface of antigen-presenting cells (APCs) that are recognized by T cell receptors (TCRs) expressed by T cells, thereby controlling immune responses 1 , 2 . A comprehensive understanding of anti-viral immunity can be achieved by studying the presentation of viral peptides by MHC class I molecules (MHCI) and the magnitude and diversity of the responding T cell repertoire 3 , 4 . Firstly, a peptide should bind to an MHCI molecule with appropriate affinity to facilitate the transport of the complex to the cell surface 5 . These pMHCI must remain on the cell surface long enough to be scrutinized by CD8 + T cells 6 , 7 . It has been suggested that pMHCI stability correlates with immunogenicity and is a better predictor than MHC-binding affinity (BA) alone 7 . Although it has been shown that the stability of pMHC complexes can influence the immunogenicity of selected virus-derived epitopes and cancer neoepitopes 7 , 8 , 9 , 10 , 11 , 12 , 13 , 14 , 15 , 16 , we chose to validate these observations further using a deeply studied system where T cell responses to the viral immunopeptidome have been studied systematically. Thus, we investigated the correlation between pMHCI stability and peptide immunogenicity for vaccinia virus (VACV) infection in C57BL/6 mice 3 by generating mass spectrometry-based thermostability measurements for individual viral peptides. VACV is the prototypic orthopoxvirus utilized as a smallpox vaccine and a vector for recombinant vaccines 17 , 18 and is also used as the monkeypox (Mpox) vaccine 19 . VACV infection in mice is also used as a model to research the fundamentals of anti-viral responses and host-pathogen interactions 17 . Recently, several studies have assessed the size and specificity of immune responses and the pMHCI derived from the VACV proteome assessed bioinformatic approaches and using comprehensive immunopeptidomics studies 3 , 10 , 17 , 20 . Here, we used a panel of VACV-derived peptides with known immunogenicity profiles to assess their pMHCI thermostability and sought to correlate these with the magnitude of the immune responses towards these viral determinants. Further, these data allowed the creation of a machine-learning-based model to predict the immunogenicity of CD8 + T cell epitopes. Results MS-thermostability profiling assay for endogenous H-2 b and VACV pMHCI complexes To profile the thermostability of endogenous murine and VACV-peptide ligands, we utilized and refined a quantitative MS-based immunopeptidomics workflow for immunopeptidome-wide thermostability measurements 10 , 21 , 22 . To investigate the thermostability of these peptides, the H-2 b -expressing DC2.4 cell line was pulsed with an exogenous mixture of 119 VACV peptides in order to enrich their presentation on the cell surface 23 , 24 (Fig. 1A). Cells were then lysed, and aliquots of the solubilised pMHCI complexes were subjected to a thermal gradient from 37℃ to 73℃, followed by immunoprecipitation of the heat treated aliquots with conformation-specific monoclonal antibodies to H-2 b MHCI allotypes. Peptides remaining in the thermostable pMHCI were analysed by quantitative mass spectrometry to assess the yield of peptide ligands at each temperature point to generate their thermostability profiles (Fig. 1A). As expected, the thermal treatment resulted in the depletion of precursor-derived MHC-peptide signal intensity (Fig. 1B) (additional analyses in Figure S2 – Supplemental Data). The quantitative analysis and processing of MS data, followed by a refined computational workflow, enabled thermal profiling of pMHCI complexes and determining melting temperature (T m ) values for individual peptide ligands. This immunopeptidomic profiling resulted in the confident detection and quantification of H-2 b MHCI peptides with defined thermostability profiles and T m values. Hierarchical clustering was used to visualize the triplicate data, which showed clear segregation from 37°C to 61°C. The row-wise clustering is representative of the overall thermostability profiles of the MHCI peptides. The latter analysis showed three profiles based on the shape of the denaturation profiles and T m values: non-stable, semi-stable, and stable (Fig. 1C). These ligand-specific profiles were extracted for comparative analysis, demonstrating the difference between the thermostability of the identified H-2 b peptide ligand classes (Figure S3 – Supplemental Data). A sigmoidal-shaped decay was observed in the global precursor MS signals resulting from the thermal dissociation of all pMHCI complexes (the average T m based on the total peak area for all peptides was estimated as 52.68°C) (Fig. 1D). We used this computational approach to determine the thermal denaturation profiles for 110 (out of 119) VACV MHCI peptides. Three examples are shown in Fig. 1E, with these peptides categorized as non-stable (VTKYYINL, T m = 44.76°C), semi-stable (KNYLFNAI, T m = 49.58°C), and stable (TSYKFESV, T m = 53.19°C) (Fig. 1E). TSYKFESV, the most immunogenic VACV MHCI peptide 3 , 25 , 26 , 27 , 28 , 29 , was notable in having a T m value higher than the median value for all endogenous H-2 b and VACV pMHCIs. Machine learning-based predictive models of viral peptide immunogenicity Given the relationship between pMHC thermostability and immunogenicity, we next sought to develop ML-based ANN models to predict peptide immunogenicity for MHCI H-2D b /K b restricted VACV peptides using the descriptors of MHCI peptide ligand residues, binding affinity (as a standard predictor), and the incorporation of the T m data as a new predictor. First, we constructed a multivariate regression model using ML algorithms to predict anti-VACV CD8 + T cell responses. We examined different sets of predictors, as listed above, as the input layer to train the ANN model. Then, we assessed the performance in modelling T-cell responses for the VACV peptides profiled by the MS-thermostability assay (Fig. 3A). In Model I, we utilized logical-based fingerprint sequence encoding (LFSE) 30 as an effective sequence scoring function to convert sequences to binary vectors based on amino acid residues in the peptides. This approach generated a data matrix containing numerically encoded descriptors of the sequences of VACV MHCI peptides as the input layer for the ANN model. We iterated the model training 200 times to examine the robustness of the model and to ensure the choice of an accurate training set throughout the data space. At each iteration, the modelling algorithm randomly selected 33% (36 IDs) of the dataset as a test set for the model validation. We evaluated the model performance independently per iteration using the correlative analysis of the predicted anti-VACV CD8 + T cell responses versus measured immune responses from infected mice 3 . Then, we picked the top 20 models (based on the outcome PCC values) to select the most reliable models for further analysis and benchmarking. In the first model, the PCC values had an average of 0.46 ± 0.03. In Model II, we utilized a substitution matrix index (SMI) method as a peptide sequence score function based on physicochemical property scorers for each amino acid residue 30 . This method generated a data matrix of new predictors to train the ANN model with the same modelling strategy as for Model I and showed an improvement compared with the first model (PCC = 0.60 ± 0.03). In the third model (Model III), we aimed to use the two sets of predictors from the previous models. Since these predictors are different in nature and scale, we used PCA to achieve their linear combinations to export the scores on the first 20 principal components (PCs) as newly transformed predictors for the model training. This data reduction resulted in a data matrix of the predictors: the rows representing VACV peptide sequences and the columns representing the 20 PCs. We trained the ANN model with the same procedure and improved the model performance in the test validation (PCC = 0.65 ± 0.02). Next, for Model IV, we added BA scores predicted by NetMHCpan 4.1 to the predictors of the third model (Model III plus BAs), and the BA predictors slightly improved the model performance in predicting anti-VACV CD8 + responses (PCC = 0.67 ± 0.02). Next, we added the T m values to Model III (Model III plus T m ) as a new predictor to evaluate the impact of the thermostability dimension in modelling peptide immunogenicity. This predictor improved the model performance significantly (P < 0.01), with a PCC of 0.80 ± 0.02. We benchmarked all modelling strategies by the top 20 constructed models per approach. The model that included the thermostability predictor demonstrated significant improvements compared with other models (statistical test by Kruskal–Wallis multiple comparisons, Fig. 3B). For further analysis, the same test set was selected for Model III (without thermostability) and the final model (Model III with thermostability), and we compared the empirical and predicted anti-VACV CD8 + T cell responses for each model, further showing the significant improvement upon inclusion of the thermostability predictor (Fig. 3C-D). In the subsequent modelling, we utilized ANN to build a classification model to recognize immunogenic VACV-derived peptides from non-immunogenic ones. We developed this qualitative model to simplify the immunogenicity model regardless of the immune reactivity level and the magnitude of the CD8 + T cell responses. We first defined the negative and positive immunogenicity classes of each peptide, as well as the minor and major immunogenic VACV peptides (as ranked by Croft et al . 3 ). Of the 110 peptides with T m profiles, 89 are immunogenic, and 21 are non-immunogenic. The endogenous H-2D b and − 2K b peptides with T m profiles are all assumed to be non-immunogenic, and 59 peptides (selected based on spanning the full T m range) were thus assigned to the non-immunogenic class. Therefore, we chose 80 non-immunogenic in total to keep the number of objects in the negative dataset roughly the same as the positive dataset. In the previous section, we organized two sets of predictors using LFSE and SMI sequence scoring functions and then augmented them (Model III). We used this strategy to initially set up predictors as input layers for the ANN algorithm and train classification models with and without including thermostability data. Then, we assessed the performance of the classification models in distinguishing immunogenic from non-immunogenic peptides by using standard evaluation metrics, e.g., accuracy, precision, sensitivity (recall), and specificity. This approach shaped the framework of the ML modelling (Fig. 3E). The model training was iterated 600 times to assess robustness and select accurate training and test sets throughout the data space. At each iteration, the modelling ANN algorithm randomly chose 67% (113 IDs) of the dataset as the training set. After the model training, the model was used to predict immunogenic peptides using the predictors of the remaining 33% of data (56 IDs). The accuracy of the classification model was evaluated by calculating the number of peptides assigned correctly to the immunogenicity classes (i.e., immunogenic and non-immunogenic), i.e. the non-error rate (NER). We selected the top 30 models and compared the accuracy (NER values) for classifying peptide immunogenicity. The model with the thermostability data outperformed other models significantly (Mann-Whitney test statistical test, p -value = 0.0006) with a NER of 0.81 ± 0.01 compared with the model without thermostability (NER = 0.78 ± 0.01) (Fig. 3F). We used a randomly selected set of non-immunogenic peptides (70 endogenous H-2D b and − 2K b MHCI peptides) to evaluate the model in true negative predictions. The model with thermostability was significantly better in accurately predicting non-immunogenic peptides, showing higher true negative rates (Fig. 3G). We used the receiver operating characteristic (ROC) curve as a standard metric of the classification model performance, a trade-off between the true positive and false positive rates. We compared the ROC curves for each model with and without thermostability predictors. This analysis revealed slightly higher area under the ROC curve (AUC) values for the model with thermostability for the non-immunogenic class (0.89 ± 0.01) compared to the model without thermostability (0.86 ± 0.01). For the immunogenic class, the AUC values equal 0.88 ± 0.01 (with thermostability) vs. 0.86 ± 0.01 (without thermostability) (Supplemental Data – Figure S9A-B). We compared the AUC values for each model, demonstrating a significant improvement upon including the thermostability predictor in distinguishing immunogenic peptides ( Mann-Whitney statistical test, p -value of 0.0282, Fig. 3H). The overall AUC values for the classifier without and with thermostability were 0.86 ± 0.01 and 0.89 ± 0.01, respectively (Fig. 3I). Given there may be an imbalance in the size of immunogenic and non-immunogenic class-specific datasets at each iteration, we compared the precision-recall (PR) curve as a trade-off for the true positive rate and the positive predictive value where a higher precision shows a lower false positive rate and higher recall (sensitivity) corresponds with a lower false negative rate. Significantly higher PR area under the curve (AUCPR) values were reported for the model with thermostability (0.89 ± 0.01) compared with the model without (0.75 ± 0.02) (Fig. 3J). Furthermore, upon a 5-fold interval cross-validation strategy for assessing the model with thermostability, good segregation was achieved by drawing cross-validated responses against the transformed scores of the immunogenicity classes (Fig. 3K - Additional metrics are reported in Supplemental Data – Figure S9). In summary, the model with thermostability was better at predicting immunogenic VACV peptides. Expanding the immunogenicity model for vaccinia and influenza A virus-derived MHCI peptides We expanded our model to other viruses to further demonstrate the effectiveness of the thermostability-based predictors. We exported VACV and influenza A virus (IAV) derived peptides with known immunogenicity from previously published data deposited in the immune epitope database (IEDB 31 ). For consistency, we exported all annotated VACV and IAV data in IEDB to generate equivalent and unbiased models for each virus and have a consolidated data pre-processing approach before modelling for all VACV- and IAV-derived MHCI peptides with validated immunogenicity profiles. We organized two discrete and confident H-2 b MHCI peptide datasets with negative (n = 115) and positive (n = 185) T cell reactivity. Notably, this dataset contains 105 VACV-derived peptides that were not thermally profiled by this study and 57 peptides that were not studied by Croft et al. 3 . As the thermostability profiles are unknown for most of these selected peptides, we developed a thermostability model embedded in the modelling workflow to predict the thermostability category for these peptides by training with the experimental T m data acquired for the ~ 3,600 endogenous H-2 b MHCI peptides above. Upon training the thermostability model, we subjected the IEDB-exported peptide data as the test set to predict their thermostability category (predicted thermostability categories; PTCs) as either non-stable or stable so that they could be incorporated into the subsequent immunogenicity model (Fig. 4A). We trained the model with Model III, incorporating predictors derived from LFSE and SMI sequence scoring functions, with and without PTC, as the input layer for training the ANN algorithm. The model training was iterated 600 times by bootstrapping to select different training and test sets from throughout the data space and to evaluate the robustness of the model. At each iteration, the ANN modelling algorithm selected 70% (210 IDs) of the dataset as the training set and the rest of the peptides (30%, 90 IDs) as the test set. Then, the trained model was utilized to distinguish immunogenic peptides (Fig. 4A). The model with the PTC predictors significantly outperformed the model without PTCs (Mann-Whitney test statistical test, p -value < 0.0001, Fig. 4B). However, considerable variation was observed in the model accuracy; therefore, we chose the 30 most accurate models for further comparison, wherein models with the PTC data were more accurate in predicting immunogenic peptides compared with the models without PTCs (Mann-Whitney statistical test, p -value = 0.0066) (Fig. 4C). We examined the models without and with the PTCs across multiple model performance evaluation measures. We assessed the AUC and AUCPR for the models, showing again a significant improvement by including the PTC predictor in distinguishing immunogenic peptides (Fig. 4D-E). We showed the minimum and maximum of the obtained ROC curves for the classifier without and with the PTCs predictor (Fig. 4F). We performed a paired-wise comparison, demonstrating a clear difference and improvement in the model with the PTCs. The ROC curves are shown for the non-immunogenic and immunogenic peptide classes alongside the overall ROC to show the performance of the classifier in the best model with the involvement of the PTC predictor (AUC = 0.87, Fig. 4G), compared favourably to the ROC curves for the best model without PTCs (AUC = 0.80, Supplemental Data – Figure S11A). Since the length of the MHC peptides may impact the model performance with the change in the number of variables in the sequence descriptors, we first analyzed the length distribution of the peptide data in both negative and positive datasets (Fig. 4H). Subsequently, at each iteration of the bootstrapping process for selecting the training (70%, 210 IDs) and test (30%, 90 IDs) sets, we applied these proportions to achieve the same length distribution to all the models. Furthermore, we performed a paired-wise t-test per iteration. Once again, the PTC predictor significantly improved model accuracy and AUC metrics (paired t-test, p -value < 0.0001, Fig. 4I-J). In the regression models, we demonstrated that the involvement of experimentally calculated T m could improve the model in predicting the magnitude of the CD8 + T cell responses better than BA. Here, we assessed the models with PTCs and compared them with the models with the inclusion of predicted binding affinities (from NetMHCpan 4.1 for the H-2D b /-2K b alleles). We show that PTCs can significantly improve the peptide immunogenicity models more than BA (ordinary one-way ANOVA, p -value of 0.0278, Fig. 4K). As multicollinearity (occurring when two independent predictors show correlation) has been shown above for thermostability and BA, an algebra-based linear combination can be a potential input for the ANN model to boost the immunogenicity model further. Thus, we included both BA and PTC predictors in the model, and the resultant model could significantly improve prediction compared with the models with only PTC (p-value of 0.0096) or BA ( p -value < 0.0001) (Fig. 4K). Thus, we demonstrated that predicted thermostability profiles can be considered a better and synergistic predictor than BA alone to interpret peptide immunogenicity. Discussion To stimulate antigen-specific T-cell responses, an immunogenic peptide should meet the prerequisites to first bind with sufficient affinity to MHC molecules and for the complex to be stable enough to allow for prolonged presentation on the cell surface for circulating T-cell immunosurveillance 4 , 7 . Therefore, it has been hypothesized that the stability of pMHC can impact the immunogenicity of virus-derived epitopes and cancer neoepitopes presented to CD8 + T cells 7 , 8 , 9 , 10 , 11 . Although the impact of pMHC thermostability on peptide immunogenicity has been shown in earlier studies 9 , 10 , few have assessed this in the context of robust immunogenicity models. In this study, we sought to investigate the extent to which the thermostability profiles of viral peptides correlated with their immunogenicity. We utilized an optimized MS-thermostability screening protocol to generate T m profiles for endogenous H-2 b peptides presented on DC2.4 cells and for a subset of VACV peptides that spanned a range of known CD8 + responses. Notably, we show that more stable VACV-derived peptide MHCI complexes are more immunogenic with higher measured anti-VACV CD8 + T cell responses and generally have higher measured binding affinities. Importantly, we show that despite a strong correlation between binding affinity and thermostability, these parameters are complementary predictors, and their combination into AI-aided models predicts immunogenicity with improved accuracy. We developed and validated ML-based models to predict virus-derived MHCI peptide immunogenicity. Using a non-linear multivariate regression ANN algorithm, a quantitative model predicted anti-VACV CD8 + T cell responses for viral peptides, with significant improvements when thermostability data was included. Overall, according to the observed robustness of these regression models, we can further improve the model if we train the models with an appropriate dataset sampled to have diversity in immunogenicity levels. A classification model that effectively distinguishes immunogenic from non-immunogenic and endogenous self-peptides was also developed that too was improved when incorporating a pMHCI thermostability predictor. Collectively, these models can confer a promising approach to predict anti-viral CD8 + T cell responses and distinguish immunogenic from non-immunogenic peptides. Although the models suffer from limitations in the training data size and the lack of large data for model validation (internal and external), this study underpins approaches for developing the next generation of immunogenicity models by implementing potentially impactful features of pMHCI complex stability. In addition, uncovering different features of the structure and interaction of the pMHC-TCR can be another important feature for decoding the peptide immunogenicity and leveraging these features as predictors for the next generation of AI-aided models for predicting immunogenicity 32 , 33 . To acquire a comprehensive snapshot of immune specificity, it is also important to simultaneously study the interaction between the pMHCI complex and CD8 + T cells 3 , 4 , 34 . Indeed, TCR repertoire is the third feature governing peptide immunogenicity. The predictors describing the interactions of TCR-pMHC alongside high-throughput TCR sequencing workflows may further improve the next generation of predictive immunogenicity models 32 , 35 , 36 . Altogether, these findings demonstrate an important correlation between thermostability and immunogenicity for pMHCI complexes as a superior predictive matrix for next-generation peptide immunogenicity models. This may be predictive of prolonged cell surface dwell times, enhancing the likelihood of T cell interaction and induction of T cell responses. Thus, including pMHC stability predictors is an effective strategy in choosing targets for T cell immunotherapy. The prediction of peptide immunogenicity may thus be strengthened by implementing this additional thermostability dimension as a new predictor for AI-aided models for immunotherapy. The thermostability screening can assist in shortlisting MHC peptides derived from tissues, biopsies, and patient-derived cells to conduct more efficient discovery of neoepitopes for cancer vaccination and immunotherapy. Methods 4.1. Experimental Design and Statistical Rationale We implemented well-established immunoaffinity purification (IP) protocols to prepare sample fractions and biological replicates by small-scale IP for data acquisition through DIA and MRM HR techniques for MS-thermostability measurement assay from lysed cells 21 , 22 . We used a large-scale IP protocol to acquire DDA data to generate spectral libraries of endogenous MHCI H-2D b and H-2K b and exogenous VACV peptides. The DIA and MRM HR (targeted) datasets were acquired on MHCI peptides isolated and purified from smaller samples using an optimized small-scale IP and peptide elution protocol. These resulting datasets were utilized to develop an MS-thermostability measurement assay to screen pMHCI complexes. The DIA dataset contains endogenous MHCI H-2D b and − 2K b peptides purified and isolated from DC2.4 cells. Both DIA and MRM HR datasets are expected to have VACV-derived MHCI peptides. The standard sequence analysis approaches were used to assess the results derived from MS-thermostability assay, e.g., pMHCI binding prediction, cluster analysis, Venn overlap graphs, and sequence motif analysis. We employed computational MS workflows for data visualization and quantitative analysis of the MS datasets for thermostability screening of pMHCI complexes. We applied the multiple pairwise comparisons derived from one-way and two-way ANOVA analysis and standard t -test analyses as canonical statistical tests to assess the measurements of different features of pMHCI complexes. Machine learning algorithms (ANN and SVM) were used to construct regression and classification models to predict anti-VACV CD8 + T cell responses and immunogenic peptides. 4.2. Cell culture We cultured the DC2.4 cell line as immortalized murine dendritic cells manufactured through transduction of bone marrow isolates derived from C57BL/6 mice by retrovirus vectors that can express MHC class I and II H-2 b alleles to present endogenous self-peptides 37 . The DC2.4 were harvested and maintained in RF-10 media of RPMI-1640 medium (Gibco™, Thermo Fisher Scientific, Waltham, MA). The media was supplemented with 10% fetal calf serum (FCS; Sigma-Aldrich), 50 µM of β-mercaptoethanol (Sigma-Aldrich), 1% (v/v) mM non-essential amino acids (Gibco™), 5 mM HEPES (Sigma-Aldrich, St Louis, MO), 2 mM L-glutamine (MP Biomedicals), 50 µg mL-1 Penicillin-Streptomycin (Pen-Strep, Gibco™), and 0.3 mg mL-1 hygromycin (Invitrogen) (to support a stable expression of the transfected MHC-I allotypes). The cells were expanded and grown in flasks and roller bottles incubated at 37°C and a 5% CO2 atmosphere. The cells were counted, harvested by centrifugation (3724 × g at 4°C for 15 min), washed with phosphate-buffered saline (PBS), pelleted, snap-frozen in liquid nitrogen, 4.3. Synthetic VACV peptides The synthetic peptides matched with the identified VACV sequences (by Croft et al. 3 ) were purchased from Mimotopes (VIC, Australia) and Genscript (Piscataway, NJ) with HPLC purity from 90–99%. The original stocks were dissolved in 100% DMSO (LC-MS grade), diluted to have 5 mM per peptide, and stored at -20℃. The peptides were diluted into water (Optima™ LC/MS Grade, Fisher Scientific®) containing 0.1% (v/v) formic acid to analyze by LC-MS/MS in DDA and MRM HR modes. 4.4. Pulsing DC2.4 cells with vaccinia virus peptides DC2.4 cells were loaded with a mixture of 122 exogenous viral peptides. First, we detached adherent cells (by 2.5 mM of EDTA in PBS) and suspended them in the RF10 media. Then, the cells were counted, and the required amount was transferred to a separate 50 mL tube (4×10 8 to 5×10 8 cells per pellet). After spinning at 1,500 rpm for 5 mins at RT, the cells were washed with PBS twice. The cells were resuspended in RPMI-1640 medium (Gibco™, Thermo Fisher Scientific, Waltham, MA) at a cell density of 2×10 7 cells mL − 1 . Afterwards, we added the mixture of synthetic VACV peptides so that the final concentration of each peptide was 1 µM. The pulsed cells were incubated at 37℃ for 1 hr on a roller while we gently resuspended the cells every 10 mins. Once the incubation was complete, we extensively washed them with PBS three times (spun between the washes at 3,724 × g , for 10 mins, at 4℃). The cells were pelleted and snap-frozen via liquid nitrogen and then stored at -80°C. 4.5. Large-scale isolation and elution of MHCI peptides A well-established large-scale immunoprecipitation protocol was utilized to isolate and purify peptides bound to MHCI H-2D b /-2K b expressed individually on the surface of the DC2.4 cells using monoclonal antibodies. A frozen pellet with 5×10 8 cells was pulverized by liquid nitrogen-cooled cryo-milling (Retsch® MM 400 mixer mill) and then resuspended in lysis buffer (50 mM Tris pH 8 (Sigma-Aldrich), 150 mM NaCl (Merck-Millipore, Darmstadt, Germany), 0.5% IGEPAL 630 (Sigma-Aldrich), and cOmplete™ Protease Inhibitor Cocktail (Roche Applied Science®)), and then incubated on a roller for 60 min at 4℃. The resulting lysate was pre-cleared by centrifugation at 2,000 × g , 4°C for 10 min, and the supernatant was ultracentrifuged at 100,000 × g, 4°C for 45 min. MHCI peptides were isolated and purified from the lysate immunoaffinity chromatography. We utilized the prepared immunoaffinity columns of monoclonal antibodies of choice (purified from hybridomas in-house), i.e., Y-3 (anti-H-2K b ; ATCC® HB-176™) and 28-14-8S (anti-H-2D b ; ATCC® HB-27™), cross-linked to Protein A sepharose (each antibody/Protein A; 10 mg mL − 1 ) for immunoprecipitation (as detailed 21 ), separately per each antibody. The captured pMHC complexes throughout the immunoaffinity column were dissociated by 10% acetic acid. After pMHC immunoaffinity purification, the eluate (a composition of MHCI heavy chains, β 2 -microglobulin (β 2 m), and peptide ligands) was fractionated through C18 reversed-phase (RP) end-capped high-performance liquid chromatography (HPLC) column (4.6 mm internal diameter × 100 mm long; Chromolith® SpeedROD; Merck-Millipore) on ÄKTAmicro™ HPLC System (GE Healthcare, UK controlled by the UNICORN™ ver. 5.11 software) to separate peptides from the MHC heavy chains and β 2 m molecules. A mobile phase constituted of buffer A (0.1% trifluoroacetic acid (TFA) [Thermo Fisher Scientific, San Jose, CA]) and buffer B (80% acetonitrile (ACN) [Thermo Fisher Scientific, Waltham, MA], and 0.1% TFA). The separated fractions were mixed into nine pools of peptides and then concentrated by vacuum centrifugation (CentriVap Benchtop Vacuum Concentrator [LABCONCO®]). The concentrated peptides were resuspended in 15 µL of 0.1% v/v of formic acid (FA) in water (Optima™ LC/MS Grade, Fisher Scientific®). 200 fmoles of iRT peptides were spiked into each sample as an internal retention time (iRT) standard for the RT normalization 38 . The samples were stored at -80℃ until LC-MS/MS analysis. 4.6. LC-MS/MS analysis to identify MHC-bound peptides DC2.4 and VACV DDA data – A SCIEX ZenoTOF 7600 LC-MS/MS system - high-resolution mass spectrometer (SCIEX™) coupled to ACQUITY™ UPLC M-Class (Microscale) LC System (Waters™, Milford, MA), Waters™ ACQUITY UPLC Console, was used to analyse isolated MHC-bound peptides and acquire MS/MS spectra (operated in the DDA mode). The UPLC system was equipped with a Waters™ nanoACQUITY Sample Manager and a Waters™ ACQUITY Auxiliary Solvent Manager as the autosampler and LC pump, respectively. 4 µL of sample fractions were loaded with a flow rate of 5 µL min − 1 onto a Kinetex® 00F-4496-AC - XB-C18 LC column (0.3 mm internal diameter × 150 mm length, Particle Size 2.6 µm, Pore Size 100 Å [Phenomenex®, Torrance, CA]). The peptides were eluted at 5 µL min − 1 flow rate over a gradient of buffer A (0.1% FA) and B (98% ACN, 0.1% FA): 0–1 min 3% B, 1–46 min 30% B, 46–47 min 80% B, 47–49 min 80% B, 49–50 min, then down to 3% B, and 50–55 min re-equilibration at 3% B. The mass spectrometer was operated in data-dependent acquisition (DDA) mode with collision-induced dissociation (CID) fragmentation. The other parameters were set to the following settings: precursor charge state of 2 + to 5+, full-scan MS1 range 400–1,500 m/z, collision energy of 10 V, MS2 scan range of 100–1,800 m/z, dynamic exclusion of 5 s, switch criteria of intensity greater than 100 counts per second, mass tolerance of 50 mDa, precursors selected for MS/MS of 45 per cycle time, declustering potential of 80 V, and accumulation times of 150 ms and 20 ms for MS1 and MS/MS, respectively. 4.7. DDA database search to discover immunopeptidomes and generation of a spectral library of H-2b and VACV-derived pMHCIs MS/MS spectra from DDA data matched with MHCI-bound peptide sequences - MS/MS spectra (PSMs) were searched against a combined database of the mouse ( Mus musculus ; UP000000589) and vaccinia virus strain Western Reserve (VACV-WR; UP000000344) proteomes (UniProtKB/SwissProt v.19102022; 17,560 entries) to identify and sequence MHC-associated peptides for generating an extensive spectral library. The database search provided a sizeable peptide repertoire containing pMHC targets for spectral library-based DIA data searches. An MSBooster-assisted database search using MSFragger v. 3.7 (over a computational platform - through a GUI pipeline; FragPipe v. 19.1) 39 , 40 was used to process and search LC − MS/MS DDA data against the mentioned database alongside a database of 17,368 decoys (50.0% - the same size as the target proteome database) and 11 standard iRT peptide sequences. A spectral library of the immunopeptidome-derived spectra was generated from this DDA dataset validated by the target-decoy algorithm with a false discovery rate (FDR) threshold of 5% for confident peptide identification by Philosopher v. 4.8.0 41 , which is a customized algorithm to filter MS/MS spectra and an FDR estimator via a multi-level target-decoy algorithm. MSBooster 42 was used in tandem with MSFragger and Percolator 43 to rescore PSMs through deep learning-based predictions of a set of features (i.e., LC retention time and MS/MS spectra). EasyPQP (version 0.1.35) was used to generate the consensus spectral library. In the MSFragger search, the peak matching parameters were set to mass calibration and parameter optimization, precursor and fragment mass error tolerances were set to 20 ppm, and isotope error was set to 0/1/2. In the protein digestion setting, cleavage was set to “nonspecific” and enzyme specificity to “nonspecific” with up to 2 missed cleavages. Peptide length was set from 7 to 25, and peptide mass range was set from 200 to 5,000 Da. Variable post-translational modifications (PTMs) were set to oxidation (M [+ 15.99]) and deamidation (NQ [+ 0.98]) with a maximum number of three variable PTMs for each peptide. A wide-scale spectral library of the immunopeptidome-derived MS/MS spectra was generated from this DDA dataset. 4.8. Small-scale immunoprecipitation of pMHC complexes to develop MS thermostability assay We implemented a small-scale immunoprecipitation protocol 22 , modified from the formerly developed protocol 21 , to purify and isolate pMHCs. The pellets of DC2.4 cells (pulsed with VACV peptides) were lysed through resuspending in the lysis buffer. The resultant lysates were incubated on a roller for 45 mins at 4°C and then cleared by centrifuge at 3700 × g for 10 mins at 4°C. The supernatant was ultracentrifuged at 100,000 × g and 4°C for 45 min. Then, the collected supernatant was run over pre-column (with 1 mL of Protein A resin) and flow lysate through. After pre-column, cleared lysates were split into aliquots of 5×10 7 cell equivalents, each triplicate incubated at 37°C, 41°C, 45°C, 49°C, 53°C, 57°C, 61°C, 65°C, 69°C, and 73°C, for 10 mins ( n = 3; where n denotes the number of biological replicates per datapoint with specific thermal treatment). Therefore, for ten temperature treatments, we prepared 30 sample replicates. The resultant thermally-treated samples were incubated overnight at 4°C with 500 µg of a mixture of purified 28-14-8S (anti-H-2D b ) and Y-3 (anti-H-2K b ) antibodies (with a ratio of 1:1–250 µg of each antibody) bound to protein A sepharose for immunoprecipitation. The pMHC complexes captured and bound to the Protein A resin-antibody were settled down and filtered through MobiSpin columns (with 10 µm pore size filters [MoBiTec GmbH, Germany]) followed by extensive washes by PBS. 400 µL of 10% acetic acid (in Optima™ LC-MS grade water) was added to release and elute the bound pMHC complexes through the Mobispin columns. Eluates constituting class-I heavy chains, β 2 m, and antibodies were filtered (from class I HC and β 2 m molecules) using pre-washed 5 kDa MWCO (molecular weight cut-off) centrifugal filter units (Ultrafree®-MC-PLHCC, Merck Millipore, Germany) through centrifuging for 60 min at 16,000 × g to collect filtered samples in new Eppendorf tubes. Excess 200 µL of 10% acetic acid was added to wash filters to elute the remaining peptides. Eluted peptides were desalted and eluted using C18-Omix 100 µL Pipette Zip tips (Agilent©, OMIX A57003100) by a buffer of 40% ACN and 0.1% FA (in Optima™ water LC-MS grade). Cleaned-up samples were concentrated by vacuum centrifugation by running Labconco Vacuum Concentrator System. MHC peptides were reconstituted in 15 µL of 0.1% v/v FA buffer in Optima water per replicate. 200 fmoles of the standard iRT peptides were spiked into the samples for retention time prediction and peak normalization for LC-MS/MS analysis. The samples were sonicated for 10 min and centrifuged at 21,000 × g for 10 min before transferring to MS vials to prepare for MS analyses. 4.9. Zeno SWATH-DIA LC-MS/MS data independent acquisition A SCIEX ZenoTOF 7600 system - high-resolution mass spectrometer (SCIEX™) coupled to ACQUITY™ UPLC M-Class (Microscale) LC System (Waters™, Milford, MA), controlled by Waters™ ACQUITY UPLC Console, was operated for LC-MS/MS DIA analysis of sample biological replicates prepared by small-scale immunoprecipitation as described above after the thermal treatments. The injection volume was 4 µL to load the sample onto a Kinetex® 00F-4496-AC - XB-C18 LC column (0.3 mm internal diameter × 150 mm length, Particle Size 2.6 µm, Pore Size 100 Å [Phenomenex®, Torrance, CA]) to elute through. The analytical column and LC gradient method were identical to the DDA data acquisition. DIA parameters were as follows: 75 DIA scans with the fixed isolation precursor windows of 8 Da (m/z) ranging from 375 to 975 m/z, MS1 scans across 400-1,500 m/z, collision energy of 10, MS2 scan range of 140–1,800 m/z, MS/MS collision energy of 17, and accumulation times of 100 ms and 15 ms for MS1 and MS/MS scans, respectively. 4.10. Library search-based DIA data analysis We employed two “peptide-centric” DIA software tools (i.e., Skyline ver. 22.2 44, 45 , and DIA-NN ver. 1.8.0 46 to analyze the DIA dataset at 5% FDR (at peptide level) using an extensive DDA spectral library previously generated. The settings are described in Supplementary Material. Perseus 47 (version 2.0.7.0) was used for post-proteome and peptidome data analysis. 4.11. Computational workflow for drawing denaturation profiles and T m value estimation The computational workflow comprises DIA data processing followed by the MS data post-analysis through in-house code and the Perseus pipeline 47 . Through the spectral library-based DIA data processing, the DIA software of choice 48 (a combination of Skyline 44 , 45 , and DIA-NN 46 ) matches the acquired composite thermostability data (peptide-specific spectra) with individual peptide MS/MS within the spectral library. After the library search and identifying specific peptide ligands, the raw chromatographic data for an eluted peptide is directly exported, and the accumulated peak areas (total peak area) for the spectral transitions can be utilized to quantify a peptide across replicates. To draw denaturation profiles for individual pMHCs, we extracted the total chromatographic peak area from each temperature-specific data point. All precursor peak areas (quantities) were normalized to the reference temperature data point (corresponding peak area) at 37°C to have fold-changes and draw denaturation profiles. This temperature was chosen as this is the condition where we expect a maximal peak intensity. We calculated the peak area under the curve (at different MS levels, i.e., MS1, MS2, and total area) per integrated chromatographic profile (assigned to a precursor (peptide) and its fragment ion peaks) to measure the relative quantity of peptides in each replicate. By the customized computational MS workflow, a sigmoidal curve was fitted to the denaturation profiles derived from SWATH-MS data (for the precursors identified at least in two biological replicates) to estimate the T m value. 4.12. MRM HR LC-MS/MS data acquisition Thermally-treated samples were analyzed for the quantification of targeted VACV peptides by MRM HR MS technique, acquired on SCIEX ZenoTOF 7600 system - high-resolution mass spectrometer (SCIEX™) coupled to ACQUITY™ UPLC M-Class (Microscale) LC System (Waters™, Milford, MA), controlled by Waters™ ACQUITY UPLC Console utilizing SCIEX OS software (v. 3.1) for data acquisition. The injection volume was 4 µL to load the sample onto the C18 column and eluted through. The analytical column and LC gradient method were identical to the DDA data acquisition. We exported collision energy (CE) values from the Zeno SWATH results for VACV peptides to set up the acquisition method for the MRM HR analysis. The LC-MS/MS system was operated in the scheduled MRM HR mode with CID fragmentation. The acquisition parameters were set as follows: CE from 16 to 47 V (dependent on the targeted precursor), precursor charges of + 1 to + 3, MS1 scan range of 210–1,250 m/z, declustering potential of 80 V, MS2 scan range of 100–2,000 m/z, and accumulation times of 200 ms and 5 ms for MS1 and MS/MS scans, respectively. MRM HR transitions (159 singly and doubly-charged precursors corresponding with 122 unique peptide sequences were tracked as a targeted list) are listed in Supplementary Data. The MRM HR dataset was processed and analyzed in Skyline 44 , 45 (v. 22.2; MacCoss Laboratory, University of Washington, Seattle, WA). We validated the detected VACV-derived precursors using two approaches. First, we examined the RT values of MRM HR precursors by comparing them with the Zeno SWATH data using the iRT predictor. Next, we matched MS/MS spectra with the spectral library of the VACV peptides derived from synthetic peptides. 4.13. Statistics We used standard statistical tests, e.g., ordinary one-/two-way ANOVA, Mann-Whitney test, Kruskal–Wallis, and Tukey’s multiple pairwise comparisons, to assess the differences between data. Z-score was used for clustering analysis calculated per peptide by subtracting the mean of intensities of each precursor from individual precursor signals in each replicate divided by the standard deviation of each precursor signal across temperatures. A p -value of ≤ 0.05 was considered the statistically significant cut-off for all statistical analyses. 4.14. Software tools for peptide sequence and statistical data analysis All statistical data analyses were executed by GraphPad Prism v. 9.0.0 and MATLAB programming code routines (The MathWorks Inc., Natick, MA) v. R2021a. The Classification toolbox was used to construct machine-learning models and evaluate their performance 49 . NetMHCpan (v. 4.1) 50 was used to determine allelic specificity according to peptide binding rank. The peptides were also segregated based on their sequence features using GibbsCluster (v. 2.0) 51 . Seq2Logo 52 was used to produce sequence motif analysis of the identified HLA-bound peptides. We used BioVenn 53 and InteractiVenn 54 to generate Venn overlap graphs. BioRender (BioRender.com) and Microsoft PowerPoint were used to design and create schematic figures and experimental workflows. 4.15. ANN algorithm setting and network structure We used a feedforward fully connected neural network with two hidden layers for the ANN-based regression model. The first layer is fully connected with the input layer (predictors), with a size of 30 neurons. In the first layer, the input data are multiplied by a weight matrix (in each fully connected neuron of the layer) and then corrected by adding a bias vector. The size of the second hidden layer was set to 10, which contains the activation function after the first hidden layer. The activation function was set to the rectified linear unit (ReLU), which applies a threshold operation to the data to set negative values to zero. The final fully connected layer is named the output layer and generates the final output of the network, which is the predicted response. For the ANN classifier, the size of the hidden layers was set to 20 and 10 for the first and second fully connected networks, respectively. 4.16. Parameters to assess the models Some parameters used to assess the classification model in recognizing immunogenic from non-immunogenic peptides are defined as follows: Precision (also known as the positive predictive value) is the ratio of the number of true positives to the sum of the true positives and false positives, explaining how well a model can predict the positive class. Sensitivity (or recall) is the ratio of the number of true positives to the sum of the true positives and the false negatives. Specificity is the ratio of the number of true negatives to the sum of the number of true negatives and false positives. 4.17. Sequence encoding functions The LFSE method was used as an effective sequence scoring function that can convert sequence strings to binary vectors logically by a numerical substitution based on the amino acid residues to provide a unique sequence descriptor for individual MHCI peptides 30 . We used the SMI scoring function to encode physicochemical characteristics-based residue groups. In the SMI method, amino acids are grouped according to the characteristics of the side chains. In the first categorical strategy, per sequence, the code substitutes the aliphatic amino acids, i.e., alanine, glycine, isoleucine, leucine, proline, and valine, with one. This function replaces aromatic (phenylalanine, tryptophan, and tyrosine), acidic (aspartic acid and glutamic acid), basic (arginine, histidine, and lysine), hydroxylic (serine and threonine), sulfur-containing (cysteine and methionine), and amidic residues (asparagine and glutamine), with two to seven, respectively. 4.18. Parameters for the IEDB search and post refinements of data We used the following setting as the search parameters for the initial search in the IEDB database’s website: epitope structure as only linear peptide sequences, no B cell assays, no MHC assays, only T cell assays, only positive assays (for the positive dataset), only negative assays (for the negative dataset), MHC restriction type of class I, and host as mouse. The post-exportation filters to make the data ready for ML modelling were the following refinements: 1) Mus musculus (C57BL/6) as the host; 2) VACV- and IAV-derived MHCI peptides; 3) CD8 + T cell responses as either positive or negative; 4) IFNγ release as the measured T cell assay response; 5) Removing the joint peptides between negative and positive datasets; 6) Subjecting the data to NetMHCpan (v. 4.1) 50 and removing non-H-2D b /-2K b binders; 7) Sequence length range was 8-to-11-mer for the selected H-2 b MHCI ligands. Declarations Acknowledgements The R@CMon/Monash Node of the NeCTAR Research Cloud, Australia’s national research cloud specifically designed for research computing, supported computational resources. We appreciate Rochelle Ayala's technical help and lab resources. Data availability All LC-MS/MS immunopeptidomics data, MSFragger DDA database search, SWATH-DIA library search results, MRM HR results, peptide quantification matrices, and thermostability results have been organized and uploaded to the ProteomeXchange Consortium via the PRIDE partner repository 55, 56 . For DDA data acquired from the DC2.4 dendritic cell line and synthetic VACV by DDA LC-MS/MS (n = 27), the accession code is PXD057866 (Username: [email protected] - Password: kYZxYytIptVf). DC2.4 cells pulsed by VACV peptides Zeno SWATH-DIA LC-MS/MS (n = 30) and LC-MRM HR (n = 30) data were deposited under the dataset identifier PXD057919 (Username: [email protected] - Password: wBb44QKOBrA0) and PXD058188 (Username: [email protected] - Password: EKCLoDDnM01e), respectively. All search and results summaries were deposited (also provided as Supplementary Data) with the corresponding datasets. Supplementary information This article contains supplemental data. Author Contributions M.S. wrote the original manuscript, performed experiments, MS-assay development, data acquisition, data analysis, wrote code to develop computational workflow for thermostability profiling, machine learning models; M.S., P.F., S.H.R., N.P.C., and A.W.P. design of experimental and computational workflows, conceptualization, methodology; P.F., D.C.T., C.L., S.H.R., N.P.C., and A.W.P. investigation; P.F., S.H.R., N.P.C., and A.W.P. supervision; N.P.C. and A.W.P. model evaluation, editing the manuscript, leading project; A.W.P. funding acquisition. Conflict of interest AWP is a scientific advisor for Bioinformatics Solutions Inc (Canada), a shareholder and scientific advisor for Evaxion Biotech (Denmark), and a co-founder of Resseptor Therapeutics (Australia). They had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results. There are no other conflicts of interest declared by the authors. Funding This work was funded by grants from the National Health and Medical Research Council of Australia (NHMRC) APP1084283 (to D.C.T., A.W.P., and N.P.C.) and APP2016596 (to A.W.P.). A.W.P is supported by a NHMRC Investigator Fellowship (APP2016596). DCT was supported by a NHMRC Investigator Fellowship (APP2008990). C.L. was supported by an Australian Research Council (ARC) Future Fellowship (FT240100798) and an NHMRC Ideas Grant (2024/GNT2037597). P.F. was supported by the Victorian Department of Health and Human Services acting through the Victorian Cancer Agency, grant 2022/GNT2019729 awarded through the National Health and Medical Research Council (NHMRC) and grant NCRI000108 awarded through the Medical Research Future Fund (MRFF). Postgraduate research scholarship from Monash University, including Monash Graduate Scholarship and Monash International Tuition Support (to M.S.). References Calis JJA et al (2013) Properties of MHC Class I Presented Peptides That Enhance Immunogenicity. PLoS Comput Biol 9:e1003266 Rothbard JB, Gefter ML (1991) Interactions between Immunogenic Peptides and MHC Proteins. Annu Rev Immunol 9:527–565 Croft NP et al (2019) Most viral peptides displayed by class I MHC on infected cells are immunogenic. Proceedings of the National Academy of Sciences 116, 3112 Croft NP (2020) Peptide Presentation to T Cells: Solving the Immunogenic Puzzle. BioEssays 42:1900200 Neefjes J, Jongsma MLM, Paul P, Bakke O (2011) Towards a systems understanding of MHC class I and MHC class II antigen presentation. Nat Rev Immunol 11:823 Rasmussen M et al (2016) Pan-Specific Prediction of Peptide–MHC Class I Complex Stability, a Correlate of T Cell Immunogenicity. J Immunol 197:1517–1524 Harndahl M et al (2012) Peptide-MHC class I stability is a better predictor than peptide affinity of CTL immunogenicity. Eur J Immunol 42:1405–1416 Hellman LM et al (2016) Differential scanning fluorimetry based assessments of the thermal and kinetic stability of peptide–MHC complexes. J Immunol Methods 432:95–101 Blaha DT et al (2019) High-Throughput Stability Screening of Neoantigen/HLA Complexes Improves Immunogenicity Predictions. Cancer Immunol Res 7:50–61 Jappe EC et al (2020) Thermostability profiling of MHC-bound peptides: a new dimension in immunopeptidomics and aid for immunotherapy design. Nat Commun 11:6305 van der Burg SH, Visseren MJ, Brandt RM, Kast WM, Melief CJ (1996) Immunogenicity of peptides bound to MHC class I molecules depends on the MHC-peptide complex stability. J Immunol 156:3308–3314 Nicholls S et al (2009) Secondary anchor polymorphism in the HA-1 minor histocompatibility antigen critically affects MHC stability and TCR recognition. Proceedings of the National Academy of Sciences 106, 3889–3894 Micheletti F et al (1999) Selective amino acid substitutions of a subdominant Epstein-Barr virus LMP2-derived epitope increase HLA/peptide complex stability and immunogenicity: implications for immunotherapy of Epstein-Barr virus-associated malignancies. Eur J Immunol 29:2579–2589 Spierings E et al (2009) Steric Hindrance and Fast Dissociation Explain the Lack of Immunogenicity of the Minor Histocompatibility HA-1Arg Null Allele 1. J Immunol 182:4809–4816 Lipford GB, Bauer S, Wagner H, Heeg K (1995) In vivo CTL induction with point-substituted ovalbumin peptides: immunogenicity correlates with peptide-induced MHC class I stability. Vaccine 13:313–320 van Stipdonk MJB et al (2009) Design of Agonistic Altered Peptides for the Robust Induction of CTL Directed towards H-2Db in Complex with the Melanoma-Associated Epitope gp100. Cancer Res 69:7784–7792 Yuen Tracy J et al (2010) Analysis of A47, an Immunoprevalent Protein of Vaccinia Virus, Leads to a Reevaluation of the Total Antiviral CD8 + T Cell Response. J Virol 84:10220–10229 Bamford D, Zuckerman M (2021) Encyclopedia of virology. Academic Chiuppesi F et al (2024) Synthetic modified vaccinia Ankara vaccines confer cross-reactive and protective immunity against mpox virus. Commun Med 4:19 Assarsson E et al (2007) A Quantitative Analysis of the Variables Affecting the Repertoire of T Cell Specificities Recognized after Vaccinia Virus Infection. J Immunol 178:7890 Purcell AW, Ramarathinam SH, Ternette N (2019) Mass spectrometry–based identification of MHC-bound peptides for immunopeptidomics. Nat Protoc 14:1687–1707 Pandey K, Ramarathinam SH, Purcell AW (2021) Isolation of HLA Bound Peptides by Immunoaffinity Capture and Identification by Mass Spectrometry. Curr Protocols 1:e92 Jensen PE, Weber DA, Thayer WR, Westerman LE, Dao CT (1999) Peptide exchange in MHC molecules. Immunol Rev 172:229–238 Chefalo PJ, Harding CV (2001) Processing of Exogenous Antigens for Presentation by Class I MHC Molecules Involves Post-Golgi Peptide Exchange Influenced by Peptide-MHC Complex Stability and Acidic pH1. J Immunol 167:1274–1282 York IA, Brehm MA, Zendzian S, Towne CF, Rock KL (2006) Endoplasmic reticulum aminopeptidase 1 (ERAP1) trims MHC class I-presented peptides in vivo and plays an important role in immunodominance. Proceedings of the National Academy of Sciences 103, 9202–9207 Tscharke DC et al (2004) Identification of poxvirus CD8 + T cell determinants to enable rational design and characterization of smallpox vaccines. J Exp Med 201:95–104 Remakus S et al (2018) Cutting Edge: Protection by Antiviral Memory CD8 T Cells Requires Rapidly Produced Antigen in Large Amounts. J Immunol 200:3347–3352 Flesch IEA et al (2009) Altered CD8 + T Cell Immunodominance after Vaccinia Virus Infection and the Naive Repertoire in Inbred and F1 Mice. J Immunol 184:45–55 Xu R-H, Remakus S, Ma X, Roscoe F, Sigal LJ (2010) Direct Presentation Is Sufficient for an Efficient Anti-Viral CD8 + T Cell Response. PLoS Pathog 6:e1000768 Shahbazy M et al (2024) MHCpLogics: an interactive machine learning-based tool for unsupervised data visualization and cluster analysis of immunopeptidomes. Brief Bioinform 25:bbae087 Vita R et al (2019) The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res 47:D339–D343 Bentzen AK et al (2018) T cell receptor fingerprinting enables in-depth characterization of the interactions governing recognition of peptide–MHC complexes. Nat Biotechnol 36:1191–1196 Szeto C, Lobos CA, Nguyen AT, Gras S (2021) TCR Recognition of Peptide–MHC-I: Rule Makers and Breakers. Int J Mol Sci Chong C, Coukos G, Bassani-Sternberg M (2022) Identification of tumor antigens with immunopeptidomics. Nat Biotechnol 40:175–188 La Gruta NL, Gras S, Daley SR, Thomas PG, Rossjohn J (2018) Understanding the drivers of MHC restriction of T cell receptors. Nat Rev Immunol 18:467–478 Dash P et al (2017) Quantifiable predictive features define epitope-specific T cell receptor repertoires. Nature 547:89–93 Shen Z, Reznikoff G, Dranoff G, Rock KL (1997) Cloned dendritic cells can present exogenous antigens on both MHC class I and class II molecules. J Immunol 158:2723–2730 Escher C et al (2012) Using iRT, a normalized retention time for more targeted measurement of peptides. Proteomics 12:1111–1121 Kong AT, Leprevost FV, Avtonomov DM, Mellacheruvu D, Nesvizhskii AI (2017) MSFragger: ultrafast and comprehensive peptide identification in mass spectrometry–based proteomics. Nat Methods 14:513–520 Teo GC, Polasky DA, Yu F, Nesvizhskii AI (2021) Fast Deisotoping Algorithm and Its Implementation in the MSFragger Search Engine. J Proteome Res 20:498–505 da Leprevost V (2020) Philosopher: a versatile toolkit for shotgun proteomics data analysis. Nat Methods 17:869–870 Yang KL et al (2022) MSBooster: Improving Peptide Identification Rates using Deep Learning-Based Features. bioRxiv , 2022.2010.2019.512904 Käll L, Canterbury JD, Weston J, Noble WS, MacCoss MJ (2007) Semi-supervised learning for peptide identification from shotgun proteomics datasets. Nat Methods 4:923–925 MacLean B et al (2010) Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. Bioinformatics 26:966–968 Pino LK et al (2020) The Skyline ecosystem: Informatics for quantitative mass spectrometry proteomics. Mass Spectrom Rev 39:229–244 Demichev V, Messner CB, Vernardis SI, Lilley KS, Ralser M (2020) DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat Methods 17:41–44 Tyanova S et al (2016) The Perseus computational platform for comprehensive analysis of (prote)omics data. Nat Methods 13:731–740 Shahbazy M et al (2023) Benchmarking Bioinformatics Pipelines in Data-Independent Acquisition Mass Spectrometry for Immunopeptidomics. Mol Cell Proteom 22:100515 Ballabio D, Consonni V (2013) Classification tools in chemistry. Part 1: linear models. PLS-DA. Anal Methods 5:3790–3798 Reynisson B, Alvarez B, Paul S, Peters B, Nielsen M (2020) NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Res 48:W449–W454 Andreatta M, Alvarez B, Nielsen M (2017) GibbsCluster: unsupervised clustering and alignment of peptide sequences. Nucleic Acids Res 45:W458–W463 Thomsen MCF, Nielsen M (2012) Seq2Logo: a method for construction and visualization of amino acid binding motifs and sequence profiles including sequence weighting, pseudo counts and two-sided representation of amino acid enrichment and depletion. Nucleic Acids Res 40:W281–W287 Hulsen T, de Vlieg J, Alkema W (2008) BioVenn – a web application for the comparison and visualization of biological lists using area-proportional Venn diagrams. BMC Genomics 9:488 Heberle H, Meirelles GV, da Silva FR, Telles GP, Minghim R (2015) InteractiVenn: a web-based tool for the analysis of sets through Venn diagrams. BMC Bioinformatics 16:169 Deutsch EW et al (2020) The ProteomeXchange consortium in 2020: enabling ‘big data’ approaches in proteomics. Nucleic Acids Res 48:D1145–D1152 Perez-Riverol Y et al (2022) The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res 50:D543–D552 Additional Declarations Yes there is potential Competing Interest. AWP is a scientific advisor for Bioinformatics Solutions Inc (Canada), a shareholder and scientific advisor for Evaxion Biotech (Denmark), and a co-founder of Resseptor Therapeutics (Australia). They had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results. There are no other conflicts of interest declared by the authors. Supplementary Files SupplementaryMaterialNatImmpMHCIVACVthermostability.docx Supplemntary Figures S1DBsearchMSFraggerMSBoosterDC24VACV.xlsx Supplemental Table 1 S2SpectralLibraryMSFraggerDC24VACV.xlsx Supplemental Table 2 S3SkylineZenoSWATHSLMSFrggDC24VACVrawoutput.xlsx Supplemental Table 3 S4DIANNZenoSWATHSLMSFrggDC24VACVrawoutput.xlsx Supplemental Table 4 S5H2DbKbVACVTmBAZenoSWATHdataNetMHCpan.xlsx Supplemental Table 5 S6VACVTmBAZenoSWATHXMRMhr.xlsx Supplemental Table 6 S7SkylineMRMHRSearchResultsVACVrawreport.xlsx Supplemental Table 7 Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5824434","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":409359421,"identity":"7f3f85c5-551f-46dd-8b42-8992c7ecb490","order_by":0,"name":"Anthony Purcell","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABB0lEQVRIiWNgGAWjYHACNiR2gYUcmGZsABIHcOlgRtZiIGFMupbEBkJa+Gf3H3vwccdhBnOJ3IcfPhhIpK9t7z348ecOBjm+GwlYtUjcOcxuOPPMYQbLGenGkjMMJHK3nTmXLM17hsFYEocWhhvJbNK8bYcZDG6ksTHzgLTcyDGQZmxjSNyAQ4s8ipY/QIeZ3cgx/vmzjaEelxYDFC1A7ycAtZhJ8LYxJBjg0GJ4I9lMcmZbOo9lzzNmyR4DCcNtZ86YWfO2SQB9+ACrFrkbic8kPrZZy5mzpzF++FFhI292vMf45s82G3m+4zi8DwHNPAZoIhL4lINAHQO6llEwCkbBKBgFcAAAqgJcoQSUGpgAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0003-0532-8331","institution":"Monash University","correspondingAuthor":true,"prefix":"","firstName":"Anthony","middleName":"","lastName":"Purcell","suffix":""},{"id":409359422,"identity":"a8279e3e-5f0f-459b-997f-4cc70058481f","order_by":1,"name":"Mohammad Shahbazy","email":"","orcid":"https://orcid.org/0000-0003-2684-3092","institution":"Yale University","correspondingAuthor":false,"prefix":"","firstName":"Mohammad","middleName":"","lastName":"Shahbazy","suffix":""},{"id":409359423,"identity":"113a46b4-60a8-4101-a0d3-69a8c525ae24","order_by":2,"name":"Sri Ramarathinam","email":"","orcid":"","institution":"Monash University","correspondingAuthor":false,"prefix":"","firstName":"Sri","middleName":"","lastName":"Ramarathinam","suffix":""},{"id":409359424,"identity":"33c93417-33e3-44f2-b17d-573f14e9124c","order_by":3,"name":"David Tscharke","email":"","orcid":"https://orcid.org/0000-0001-6825-9172","institution":"The Australian National University","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"","lastName":"Tscharke","suffix":""},{"id":409359425,"identity":"29b57472-4bbe-4df4-b565-2bb863cf5b47","order_by":4,"name":"Chen Li","email":"","orcid":"https://orcid.org/0000-0002-1847-754X","institution":"Monash Unversity","correspondingAuthor":false,"prefix":"","firstName":"Chen","middleName":"","lastName":"Li","suffix":""},{"id":409359426,"identity":"2a6dc8d2-270a-48c2-9fc5-100fbf90b5bd","order_by":5,"name":"Pouya Faridi","email":"","orcid":"https://orcid.org/0000-0002-2712-3356","institution":"Monash University","correspondingAuthor":false,"prefix":"","firstName":"Pouya","middleName":"","lastName":"Faridi","suffix":""},{"id":409359427,"identity":"b09e355d-5e1c-46e8-a17d-9cc51f668a2a","order_by":6,"name":"Nathan Croft","email":"","orcid":"https://orcid.org/0000-0002-2128-5127","institution":"Monash University","correspondingAuthor":false,"prefix":"","firstName":"Nathan","middleName":"","lastName":"Croft","suffix":""}],"badges":[],"createdAt":"2025-01-14 06:10:49","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5824434/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5824434/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":75284089,"identity":"4ccb366f-3bac-4207-b672-6e288c04d94f","added_by":"auto","created_at":"2025-02-03 03:59:45","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":813166,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eMS-immunopeptidomics workflow for thermostability profiling of pMHC molecules. (A) \u003c/strong\u003eSchematic of the workflow utilised. As well as capturing endogenously presented H-2b-bound peptides, DC2.4 cells were pulsed with a mixture of 119 synthetic peptides before cells were lysed under mild conditions, with aliquots subjected to a thermal gradient before immunoprecipitation of pMHC and immunopeptidomics. Data were acquired by quantitative mass spectrometry (DIA and MRM analyses) and searched against a spectral library. \u003cstrong\u003e(B)\u003c/strong\u003e Extracted ion chromatograms (XICs) were utilized to measure relative quantities of peptide per data point. \u003cstrong\u003e(C)\u003c/strong\u003e Hierarchical clustering in both data directions to visualize the thermostability of quantified H-2\u003csup\u003eb \u003c/sup\u003eMHCI-bound peptides. This analysis segregated thermostability profiles into non-, semi-, and stable profiles based on the denaturation profile shapes. \u003cstrong\u003e(D)\u003c/strong\u003e Logistic sigmoid functional curve fitted to a peptide thermal denaturation profile, allowing an estimate of the overall T\u003csub\u003em\u003c/sub\u003e value for H-2\u003csup\u003eb\u003c/sup\u003e MHCI peptide (52.68 °C in the example shown). \u003cstrong\u003e(E)\u003c/strong\u003e Peptide spectra (top panels) were validated by matching the eluted spectrum with that from an in-house generated spectral library. Three example denaturation profiles are shown for VACV peptides, representing non-stable (VTKYYINL), semi-stable (KNYLFNAI), and stable peptide (TSYKFESV).\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/3a9c3496cf9b23c7a844ade1.png"},{"id":75284841,"identity":"c83d6f53-85e9-4906-ba16-4ff07735ae04","added_by":"auto","created_at":"2025-02-03 04:15:45","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":365435,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eAnalysis of MS-thermostability assay-derived thermostability profiles for H-2\u003c/strong\u003e\u003csup\u003e\u003cstrong\u003eb\u003c/strong\u003e\u003c/sup\u003e\u003cstrong\u003e MHCI and VACV peptides, and association of T\u003c/strong\u003e\u003csub\u003e\u003cstrong\u003em\u003c/strong\u003e\u003c/sub\u003e\u003cstrong\u003e values with immunogenicity for VACV-derived pMHCs. (A)\u003c/strong\u003e The proportion of the H-2\u003csup\u003eb\u003c/sup\u003e and VACV MHCI peptides thermally profiled by MS-thermostability assay.\u003cstrong\u003e (B) \u003c/strong\u003eThe distribution for the calculated T\u003csub\u003em\u003c/sub\u003e values with a median of 46.48 °C (H-2D\u003csup\u003eb \u003c/sup\u003e– green), 46.62 °C (H-2K\u003csup\u003eb \u003c/sup\u003e– blue), and 48.06 °C (VACV peptides – red). \u003cstrong\u003e(C)\u003c/strong\u003e Venn diagram overlap between the VACV peptide sequences profiled by Zeno SWATH DIA and MRM\u003csup\u003eHR\u003c/sup\u003e approaches. \u003cstrong\u003e(D)\u003c/strong\u003e Examination of the association of the measured T\u003csub\u003em\u003c/sub\u003e values by MS-thermostability assays for 89 commonly-profiled VACV peptides (PCC = 0.844). \u003cstrong\u003e(E)\u003c/strong\u003e Comparison of the Tm values distribution for the commonly profiled VACV peptides by DIA and MRM\u003csup\u003eHR\u003c/sup\u003e MS-thermostability assays. \u003cstrong\u003e(F)\u003c/strong\u003e Comparison of the measured binding affinities (as IC\u003csub\u003e50\u003c/sub\u003e values) \u003csup\u003e3\u003c/sup\u003e for thermostability-ranked VACV-derived MHCI peptides. \u003cstrong\u003e(G)\u003c/strong\u003e Examination of the measured anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses (log\u003csub\u003e2\u003c/sub\u003e) for the thermostability-ranked VACV-derived MHCI epitopes demonstrating VACV-derived pMHCIs with higher stability have significantly higher CD8\u003csup\u003e+\u003c/sup\u003e responses. \u003cstrong\u003e(H) \u003c/strong\u003eComparison between T\u003csub\u003em\u003c/sub\u003e values of the immune reactivity-ranked VACV peptides, demonstrating that the major peptides have significantly higher Tm values than\u0026nbsp;non- and minor-immunogenic pMHCIs. \u003cem\u003eThe thermostability ranks were defined based on the average of T\u003c/em\u003e\u003csub\u003e\u003cem\u003em\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e values (T\u003c/em\u003e\u003csub\u003e\u003cem\u003emean\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e) as follows: non-stable (T\u003c/em\u003e\u003csub\u003e\u003cem\u003em\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e \u0026lt; [T\u003c/em\u003e\u003csub\u003e\u003cem\u003emean\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e – 3°C]); semi-stable ([T\u003c/em\u003e\u003csub\u003e\u003cem\u003emean\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e - 3°C] ≤ T\u003c/em\u003e\u003csub\u003e\u003cem\u003em\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e \u0026lt; [T\u003c/em\u003e\u003csub\u003e\u003cem\u003emean\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e + 3°C]); stable (T\u003c/em\u003e\u003csub\u003e\u003cem\u003em\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e ≥ [T\u003c/em\u003e\u003csub\u003e\u003cem\u003emean\u003c/em\u003e\u003c/sub\u003e\u003cem\u003e + 3°C]). The\u003c/em\u003e \u003cem\u003eKruskal–Wallis (F) and ordinary one-way ANOVA (G, H, and I) multiple comparison tests were used to determine the significance of the pairs. ns, not significant; *p-value \u0026lt; 0.05; **p-value \u0026lt; 0.01; ***p-value \u0026lt; 0.001; and ****p-value \u0026lt; 0.0001.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/98ad14f5fa0dee6554c17162.png"},{"id":75284261,"identity":"ee923ecf-209d-4b2a-88d8-18c459465b39","added_by":"auto","created_at":"2025-02-03 04:07:45","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":620742,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eANN models to predict anti-VACV CD8\u003c/strong\u003e\u003csup\u003e\u003cstrong\u003e+\u003c/strong\u003e\u003c/sup\u003e\u003cstrong\u003e T cell responses and immunogenic VACV peptides.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDifferent predictors were used as the input layer to train the ANN model. Then, the model performance was assessed to predict CD8\u003csup\u003e+\u003c/sup\u003e T cell responses for 110 VACV peptides profiled by the MS-thermostability assay. \u003cstrong\u003e(B)\u003c/strong\u003e Benchmarking the top-20 constructed models per each strategy. The model that included the thermostability predictor demonstrated significant improvements compared with others. \u003cstrong\u003e(C-D)\u003c/strong\u003e For further analysis, the same test set was selected for \u003cstrong\u003e(C)\u003c/strong\u003e Model III (without thermostability) and \u003cstrong\u003e(D)\u003c/strong\u003e the final model (Model III with thermostability). The measured and predicted anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses were compared for each model, showing significant improvement upon including the thermostability predictor. \u003cstrong\u003e(E)\u003c/strong\u003e The framework of the ML model was used to train the ANN algorithm by different sets of predictors to distinguish immunogenic peptides. The models were assessed by predicting immunogenicity classes for the test set using standard metrics of evaluation classification models. \u003cstrong\u003e(F)\u003c/strong\u003e The comparison between the calculated NER values shows ML-modelling accuracy in classifying peptides. The model with thermostability predictors (NER = 0.8134 ± 0.0073) was significantly superior (\u003cem\u003ep\u003c/em\u003e-value = 0.0006) to the model without thermostability (NER = 0.7971 ± 0.0069). \u003cstrong\u003e(G) \u003c/strong\u003eThe comparison between the TN rates shows that the model with thermostability predictors was significantly higher than those without. \u003cstrong\u003e(H) \u003c/strong\u003eThe comparison between the AUC values for the models without and with the involvement of the thermostability profiles showed a significant improvement upon including the thermostability predictor in distinguishing immunogenic peptides (\u003cem\u003ep\u003c/em\u003e-value of 0.0282). \u003cstrong\u003e(I)\u003c/strong\u003e The overall AUC values for the models without and with thermostability were 0.8622 ± 0.0071 and 0.8863 ± 0.0069, respectively. \u003cstrong\u003e(J)\u003c/strong\u003e The assessment of PR curves with the AUCPR values for the models with and without thermostability data. The PR curves showed higher AUCPR values significantly (\u003cem\u003ep\u003c/em\u003e-value \u0026lt; 0.0001) for the model with thermostability (0.8852 ± 0.0108) compared with the model without thermostability (0.7529 ± 0.0150). \u003cstrong\u003e(K)\u003c/strong\u003e A 5-fold interval (cross) validation strategy for assessing the model with thermostability, demonstrating good segregation by drawing cross-validated responses against the transformed scores of the immunogenicity classes. \u003cem\u003eThe\u003c/em\u003e \u003cem\u003estatistical tests of the Kruskal–Wallis multiple-comparison (B) and Mann-Whitney (F, G, H, J) were used to determine the significance.\u003c/em\u003e \u003cem\u003ens, not significant; *p-value \u0026lt; 0.05; **p-value \u0026lt; 0.01; ***p-value \u0026lt; 0.001; and ****p-value \u0026lt; 0.0001. PCC, Pearson correlation coefficient;\u003c/em\u003e \u003cem\u003eNER, \u003c/em\u003eNon-error rate (model accuracy); ROC, Receiver operating characteristic; PR, Precision-recall; AUC, Area under the ROC curve; AUCPR, Area under the PR curve.\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/3c7fd0f7c4684c8b270c336c.png"},{"id":75284109,"identity":"841072a5-e046-41a9-a333-404e32f1d0ee","added_by":"auto","created_at":"2025-02-03 03:59:46","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":3312318,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eClassification ML model to distinguish immunogenic VACV- and IAV-derived MHCI peptides. (A)\u003c/strong\u003e A schematic flowchart illustrates the workflow for exporting, refining, and classifying immunogenic VACV- and IAV-derived MHCI peptides. Peptides with known immunogenicity were exported from the IEDB (\u003ca href=\"http://www.iedb.org\"\u003ewww.iedb.org\u003c/a\u003e), followed by post-filters for VACV and IAV peptides. Two distinct datasets of H-2\u003csup\u003eb\u003c/sup\u003e MHCI peptides were organized based on T cell reactivity (negative (n = 115) and positive (n = 185)). A machine learning model was trained using MS-thermostability data of ~3,600 mouse-derived H-2\u003csup\u003eb\u003c/sup\u003e MHCI peptides to predict stability categories (PTCs) for the selected MHCI peptides with unknown thermostability profiles. Using an ANN algorithm, these PTCs, along with MHCI pocket and physicochemical features of amino acid residues, were integrated into immunogenicity models. Model training involved iterative bootstrapping for robustness evaluation and external validation using a test set. \u003cstrong\u003e(B-C)\u003c/strong\u003e Assessment of the model accuracy with and without the PTC predictors, where a significant improvement was observed for the model with PTC predictor for the external test set for \u003cstrong\u003e(B)\u003c/strong\u003e all 600 iterations and \u003cstrong\u003e(C) \u003c/strong\u003ethe top 5% (n = 30) models. \u003cstrong\u003e(D)\u003c/strong\u003e Assessment of the AUC values for each model shows a significant improvement by including the PTCs predictor. \u003cstrong\u003e(E)\u003c/strong\u003e Comparison of PR curves and AUCPR for the models with and without PTCs predictor; although there is a slightly higher average in the model with PTCs, no significant difference was observed. \u003cstrong\u003e(F)\u003c/strong\u003e A paired-wise comparison between the minimum and maximum of the obtained ROC curves for the classifier -/+ PTCs. \u003cstrong\u003e(G)\u003c/strong\u003e The ROC curves show the classifier performance for the non-immunogenic and immunogenic peptide classes alongside the overall ROC in the best model, including the PTC predictor (AUC = 0.868). \u003cstrong\u003e(H)\u003c/strong\u003e The length distribution of the peptide data in negative and positive datasets. \u003cstrong\u003e(I-J)\u003c/strong\u003eThe models were constructed with the same length distribution in the training and test sets, followed by paired-wise t-test statistics, showing that the inclusion of PTC led to significant improvements in \u003cstrong\u003e(I)\u003c/strong\u003e model accuracy and \u003cstrong\u003e(J)\u003c/strong\u003e AUC for 200 model iterations. \u003cstrong\u003e(K)\u003c/strong\u003e Evaluation of the models with PTCs and predicted BAs (from NetMHCpan 4.1). PTCs significantly improved the immunogenicity models better than BA. A combination of BA and PTC predictors demonstrated further significant improvement. \u003cem\u003eThe Mann-Whitney (B, C, D, and E), paired-wise t-test (I and J), and ordinary one-way ANOVA multiple comparison (K) statistical tests were used to determine the significance. ns, not significant; *p-value \u0026lt; 0.05; **p-value \u0026lt; 0.01; ***p-value \u0026lt; 0.001; and ****p-value \u0026lt; 0.0001.\u003c/em\u003e \u003cem\u003eANN, artificial neural network; VACV, vaccinia virus; IAV, Influenza A virus; PTC, Predicted thermostability category; AUC, Area under the ROC curve; BA, binding affinity.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/0d9434b95e233bf53e743806.png"},{"id":75285486,"identity":"c5c31d44-eabf-49dc-91db-2e1deaaff9ed","added_by":"auto","created_at":"2025-02-03 04:23:49","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6159451,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/23af7bbd-4a30-4e65-8c9e-5aa1937ab178.pdf"},{"id":75284098,"identity":"380e2ac6-ece1-40d5-b951-ce94677bea54","added_by":"auto","created_at":"2025-02-03 03:59:46","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":3253883,"visible":true,"origin":"","legend":"Supplemntary Figures","description":"","filename":"SupplementaryMaterialNatImmpMHCIVACVthermostability.docx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/d45df326ec8c7c19f84ba2f4.docx"},{"id":75284093,"identity":"57dcf58d-fc7d-451a-b445-c5fae92f35e5","added_by":"auto","created_at":"2025-02-03 03:59:45","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":970211,"visible":true,"origin":"","legend":"Supplemental Table 1","description":"","filename":"S1DBsearchMSFraggerMSBoosterDC24VACV.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/09944aa16d3f4bd9300b113e.xlsx"},{"id":75284106,"identity":"85d34b2e-7550-4f7a-bf5d-6f260336eb42","added_by":"auto","created_at":"2025-02-03 03:59:46","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":12505828,"visible":true,"origin":"","legend":"Supplemental Table 2","description":"","filename":"S2SpectralLibraryMSFraggerDC24VACV.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/6777b997c264a5e0124f615b.xlsx"},{"id":75284114,"identity":"27279e0e-8b92-426b-927c-3c1f9be61034","added_by":"auto","created_at":"2025-02-03 03:59:47","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":59966085,"visible":true,"origin":"","legend":"Supplemental Table 3","description":"","filename":"S3SkylineZenoSWATHSLMSFrggDC24VACVrawoutput.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/9823eff257d3357fa1922264.xlsx"},{"id":75284266,"identity":"e15f9aad-2e0d-4f45-bc33-511927b54b56","added_by":"auto","created_at":"2025-02-03 04:07:46","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":16889672,"visible":true,"origin":"","legend":"Supplemental Table 4","description":"","filename":"S4DIANNZenoSWATHSLMSFrggDC24VACVrawoutput.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/ec64e247570b7a30d345f417.xlsx"},{"id":75284107,"identity":"4f4b6492-e9a0-43ee-bf8a-d66e9eebed22","added_by":"auto","created_at":"2025-02-03 03:59:46","extension":"xlsx","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":972133,"visible":true,"origin":"","legend":"\u003cp\u003eSupplemental Table 5\u003c/p\u003e","description":"","filename":"S5H2DbKbVACVTmBAZenoSWATHdataNetMHCpan.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/485162d263efbcfe36d3e18c.xlsx"},{"id":75284095,"identity":"689fa634-9b5f-4d7f-ac36-c1c659954aed","added_by":"auto","created_at":"2025-02-03 03:59:46","extension":"xlsx","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":49607,"visible":true,"origin":"","legend":"Supplemental Table 6","description":"","filename":"S6VACVTmBAZenoSWATHXMRMhr.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/cceaa57c022bcff610c3faea.xlsx"},{"id":75284267,"identity":"aec88a7b-5f7f-47bf-af86-7ce9bb2327dc","added_by":"auto","created_at":"2025-02-03 04:07:46","extension":"xlsx","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":919951,"visible":true,"origin":"","legend":"\u003cp\u003eSupplemental Table 7\u003c/p\u003e","description":"","filename":"S7SkylineMRMHRSearchResultsVACVrawreport.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-5824434/v1/341fd115141e1454cf48dd54.xlsx"}],"financialInterests":"\u003cb\u003eYes\u003c/b\u003e there is potential Competing Interest.\nAWP is a scientific advisor for Bioinformatics Solutions Inc (Canada), a shareholder and scientific advisor for Evaxion Biotech (Denmark), and a co-founder of Resseptor Therapeutics (Australia). They had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results. There are no other conflicts of interest declared by the authors.","formattedTitle":"Mass spectrometry-based thermostability profiling of virus-derived MHC peptide complexes serves as an effective predictor of immunogenicity","fulltext":[{"header":"Introduction","content":"\u003cp\u003eMajor histocompatibility complex (MHC) molecules present peptide antigens on the surface of antigen-presenting cells (APCs) that are recognized by T cell receptors (TCRs) expressed by T cells, thereby controlling immune responses \u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e. A comprehensive understanding of anti-viral immunity can be achieved by studying the presentation of viral peptides by MHC class I molecules (MHCI) and the magnitude and diversity of the responding T cell repertoire \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. Firstly, a peptide should bind to an MHCI molecule with appropriate affinity to facilitate the transport of the complex to the cell surface \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. These pMHCI must remain on the cell surface long enough to be scrutinized by CD8\u003csup\u003e+\u003c/sup\u003e T cells \u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. It has been suggested that pMHCI stability correlates with immunogenicity and is a better predictor than MHC-binding affinity (BA) alone \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Although it has been shown that the stability of pMHC complexes can influence the immunogenicity of selected virus-derived epitopes and cancer neoepitopes \u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e, \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e, we chose to validate these observations further using a deeply studied system where T cell responses to the viral immunopeptidome have been studied systematically. Thus, we investigated the correlation between pMHCI stability and peptide immunogenicity for vaccinia virus (VACV) infection in C57BL/6 mice \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e by generating mass spectrometry-based thermostability measurements for individual viral peptides.\u003c/p\u003e \u003cp\u003eVACV is the prototypic orthopoxvirus utilized as a smallpox vaccine and a vector for recombinant vaccines \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e and is also used as the monkeypox (Mpox) vaccine \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e\u003c/sup\u003e. VACV infection in mice is also used as a model to research the fundamentals of anti-viral responses and host-pathogen interactions \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e. Recently, several studies have assessed the size and specificity of immune responses and the pMHCI derived from the VACV proteome assessed bioinformatic approaches and using comprehensive immunopeptidomics studies \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e, \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e. Here, we used a panel of VACV-derived peptides with known immunogenicity profiles to assess their pMHCI thermostability and sought to correlate these with the magnitude of the immune responses towards these viral determinants. Further, these data allowed the creation of a machine-learning-based model to predict the immunogenicity of CD8\u003csup\u003e+\u003c/sup\u003e T cell epitopes.\u003c/p\u003e"},{"header":"Results","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n\u003ch2\u003eMS-thermostability profiling assay for endogenous H-2\u003csup\u003eb\u003c/sup\u003e and VACV pMHCI complexes\u003c/h2\u003e\n\u003cp\u003eTo profile the thermostability of endogenous murine and VACV-peptide ligands, we utilized and refined a quantitative MS-based immunopeptidomics workflow for immunopeptidome-wide thermostability measurements \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. To investigate the thermostability of these peptides, the H-2\u003csup\u003eb\u003c/sup\u003e-expressing DC2.4 cell line was pulsed with an exogenous mixture of 119 VACV peptides in order to enrich their presentation on the cell surface \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e23\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e (Fig.\u0026nbsp;1A). Cells were then lysed, and aliquots of the solubilised pMHCI complexes were subjected to a thermal gradient from 37℃ to 73℃, followed by immunoprecipitation of the heat treated aliquots with conformation-specific monoclonal antibodies to H-2\u003csup\u003eb\u003c/sup\u003e MHCI allotypes. Peptides remaining in the thermostable pMHCI were analysed by quantitative mass spectrometry to assess the yield of peptide ligands at each temperature point to generate their thermostability profiles (Fig.\u0026nbsp;1A). As expected, the thermal treatment resulted in the depletion of precursor-derived MHC-peptide signal intensity (Fig.\u0026nbsp;1B) (additional analyses in Figure S2 \u0026ndash; Supplemental Data).\u003c/p\u003e\n\u003cp\u003eThe quantitative analysis and processing of MS data, followed by a refined computational workflow, enabled thermal profiling of pMHCI complexes and determining melting temperature (T\u003csub\u003em\u003c/sub\u003e) values for individual peptide ligands. This immunopeptidomic profiling resulted in the confident detection and quantification of H-2\u003csup\u003eb\u003c/sup\u003e MHCI peptides with defined thermostability profiles and T\u003csub\u003em\u003c/sub\u003e values. Hierarchical clustering was used to visualize the triplicate data, which showed clear segregation from 37\u0026deg;C to 61\u0026deg;C. The row-wise clustering is representative of the overall thermostability profiles of the MHCI peptides. The latter analysis showed three profiles based on the shape of the denaturation profiles and T\u003csub\u003em\u003c/sub\u003e values: non-stable, semi-stable, and stable (Fig.\u0026nbsp;1C). These ligand-specific profiles were extracted for comparative analysis, demonstrating the difference between the thermostability of the identified H-2\u003csup\u003eb\u003c/sup\u003e peptide ligand classes (Figure S3 \u0026ndash; Supplemental Data).\u003c/p\u003e\n\u003cp\u003eA sigmoidal-shaped decay was observed in the global precursor MS signals resulting from the thermal dissociation of all pMHCI complexes (the average T\u003csub\u003em\u003c/sub\u003e based on the total peak area for all peptides was estimated as 52.68\u0026deg;C) (Fig.\u0026nbsp;1D). We used this computational approach to determine the thermal denaturation profiles for 110 (out of 119) VACV MHCI peptides. Three examples are shown in Fig.\u0026nbsp;1E, with these peptides categorized as non-stable (VTKYYINL, T\u003csub\u003em\u003c/sub\u003e = 44.76\u0026deg;C), semi-stable (KNYLFNAI, T\u003csub\u003em\u003c/sub\u003e = 49.58\u0026deg;C), and stable (TSYKFESV, T\u003csub\u003em\u003c/sub\u003e = 53.19\u0026deg;C) (Fig.\u0026nbsp;1E). TSYKFESV, the most immunogenic VACV MHCI peptide \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e25\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e27\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e28\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e, was notable in having a T\u003csub\u003em\u003c/sub\u003e value higher than the median value for all endogenous H-2\u003csup\u003eb\u003c/sup\u003e and VACV pMHCIs.\u003c/p\u003e\n\u003c/div\u003e\n\u003ch3\u003eMachine learning-based predictive models of viral peptide immunogenicity\u003c/h3\u003e\n\u003cp\u003eGiven the relationship between pMHC thermostability and immunogenicity, we next sought to develop ML-based ANN models to predict peptide immunogenicity for MHCI H-2D\u003csup\u003eb\u003c/sup\u003e/K\u003csup\u003eb\u003c/sup\u003e restricted VACV peptides using the descriptors of MHCI peptide ligand residues, binding affinity (as a standard predictor), and the incorporation of the T\u003csub\u003em\u003c/sub\u003e data as a new predictor.\u003c/p\u003e\n\u003cp\u003eFirst, we constructed a multivariate regression model using ML algorithms to predict anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses. We examined different sets of predictors, as listed above, as the input layer to train the ANN model. Then, we assessed the performance in modelling T-cell responses for the VACV peptides profiled by the MS-thermostability assay (Fig.\u0026nbsp;3A). In Model I, we utilized logical-based fingerprint sequence encoding (LFSE) \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e as an effective sequence scoring function to convert sequences to binary vectors based on amino acid residues in the peptides. This approach generated a data matrix containing numerically encoded descriptors of the sequences of VACV MHCI peptides as the input layer for the ANN model. We iterated the model training 200 times to examine the robustness of the model and to ensure the choice of an accurate training set throughout the data space. At each iteration, the modelling algorithm randomly selected 33% (36 IDs) of the dataset as a test set for the model validation. We evaluated the model performance independently per iteration using the correlative analysis of the predicted anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses versus measured immune responses from infected mice \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. Then, we picked the top 20 models (based on the outcome PCC values) to select the most reliable models for further analysis and benchmarking. In the first model, the PCC values had an average of 0.46\u0026thinsp;\u0026plusmn;\u0026thinsp;0.03.\u003c/p\u003e\n\u003cp\u003eIn Model II, we utilized a substitution matrix index (SMI) method as a peptide sequence score function based on physicochemical property scorers for each amino acid residue \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. This method generated a data matrix of new predictors to train the ANN model with the same modelling strategy as for Model I and showed an improvement compared with the first model (PCC\u0026thinsp;=\u0026thinsp;0.60\u0026thinsp;\u0026plusmn;\u0026thinsp;0.03). In the third model (Model III), we aimed to use the two sets of predictors from the previous models. Since these predictors are different in nature and scale, we used PCA to achieve their linear combinations to export the scores on the first 20 principal components (PCs) as newly transformed predictors for the model training. This data reduction resulted in a data matrix of the predictors: the rows representing VACV peptide sequences and the columns representing the 20 PCs. We trained the ANN model with the same procedure and improved the model performance in the test validation (PCC\u0026thinsp;=\u0026thinsp;0.65\u0026thinsp;\u0026plusmn;\u0026thinsp;0.02).\u003c/p\u003e\n\u003cp\u003eNext, for Model IV, we added BA scores predicted by NetMHCpan 4.1 to the predictors of the third model (Model III plus BAs), and the BA predictors slightly improved the model performance in predicting anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e responses (PCC\u0026thinsp;=\u0026thinsp;0.67\u0026thinsp;\u0026plusmn;\u0026thinsp;0.02). Next, we added the T\u003csub\u003em\u003c/sub\u003e values to Model III (Model III plus T\u003csub\u003em\u003c/sub\u003e) as a new predictor to evaluate the impact of the thermostability dimension in modelling peptide immunogenicity. This predictor improved the model performance significantly (P\u0026thinsp;\u0026lt;\u0026thinsp;0.01), with a PCC of 0.80\u0026thinsp;\u0026plusmn;\u0026thinsp;0.02. We benchmarked all modelling strategies by the top 20 constructed models per approach. The model that included the thermostability predictor demonstrated significant improvements compared with other models (statistical test by Kruskal\u0026ndash;Wallis multiple comparisons, Fig.\u0026nbsp;3B). For further analysis, the same test set was selected for Model III (without thermostability) and the final model (Model III with thermostability), and we compared the empirical and predicted anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses for each model, further showing the significant improvement upon inclusion of the thermostability predictor (Fig.\u0026nbsp;3C-D).\u003c/p\u003e\n\u003cp\u003eIn the subsequent modelling, we utilized ANN to build a classification model to recognize immunogenic VACV-derived peptides from non-immunogenic ones. We developed this qualitative model to simplify the immunogenicity model regardless of the immune reactivity level and the magnitude of the CD8\u003csup\u003e+\u003c/sup\u003e T cell responses. We first defined the negative and positive immunogenicity classes of each peptide, as well as the minor and major immunogenic VACV peptides (as ranked by Croft \u003cem\u003eet al\u003c/em\u003e.\u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e). Of the 110 peptides with T\u003csub\u003em\u003c/sub\u003e profiles, 89 are immunogenic, and 21 are non-immunogenic. The endogenous H-2D\u003csup\u003eb\u003c/sup\u003e and \u0026minus;\u0026thinsp;2K\u003csup\u003eb\u003c/sup\u003e peptides with T\u003csub\u003em\u003c/sub\u003e profiles are all assumed to be non-immunogenic, and 59 peptides (selected based on spanning the full T\u003csub\u003em\u003c/sub\u003e range) were thus assigned to the non-immunogenic class. Therefore, we chose 80 non-immunogenic in total to keep the number of objects in the negative dataset roughly the same as the positive dataset. In the previous section, we organized two sets of predictors using LFSE and SMI sequence scoring functions and then augmented them (Model III). We used this strategy to initially set up predictors as input layers for the ANN algorithm and train classification models with and without including thermostability data.\u003c/p\u003e\n\u003cp\u003eThen, we assessed the performance of the classification models in distinguishing immunogenic from non-immunogenic peptides by using standard evaluation metrics, e.g., accuracy, precision, sensitivity (recall), and specificity. This approach shaped the framework of the ML modelling (Fig.\u0026nbsp;3E). The model training was iterated 600 times to assess robustness and select accurate training and test sets throughout the data space. At each iteration, the modelling ANN algorithm randomly chose 67% (113 IDs) of the dataset as the training set. After the model training, the model was used to predict immunogenic peptides using the predictors of the remaining 33% of data (56 IDs). The accuracy of the classification model was evaluated by calculating the number of peptides assigned correctly to the immunogenicity classes (i.e., immunogenic and non-immunogenic), i.e. the non-error rate (NER). We selected the top 30 models and compared the accuracy (NER values) for classifying peptide immunogenicity. The model with the thermostability data outperformed other models significantly (Mann-Whitney test statistical test, \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;=\u0026thinsp;0.0006) with a NER of 0.81\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01 compared with the model without thermostability (NER\u0026thinsp;=\u0026thinsp;0.78\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01) (Fig.\u0026nbsp;3F). We used a randomly selected set of non-immunogenic peptides (70 endogenous H-2D\u003csup\u003eb\u003c/sup\u003e and \u0026minus;\u0026thinsp;2K\u003csup\u003eb\u003c/sup\u003e MHCI peptides) to evaluate the model in true negative predictions. The model with thermostability was significantly better in accurately predicting non-immunogenic peptides, showing higher true negative rates (Fig.\u0026nbsp;3G).\u003c/p\u003e\n\u003cp\u003eWe used the receiver operating characteristic (ROC) curve as a standard metric of the classification model performance, a trade-off between the true positive and false positive rates. We compared the ROC curves for each model with and without thermostability predictors. This analysis revealed slightly higher area under the ROC curve (AUC) values for the model with thermostability for the non-immunogenic class (0.89\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01) compared to the model without thermostability (0.86\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01). For the immunogenic class, the AUC values equal 0.88\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01 (with thermostability) vs. 0.86\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01 (without thermostability) (Supplemental Data \u0026ndash; Figure S9A-B). We compared the AUC values for each model, demonstrating a significant improvement upon including the thermostability predictor in distinguishing immunogenic peptides (\u003cem\u003eMann-Whitney\u003c/em\u003e statistical test, \u003cem\u003ep\u003c/em\u003e-value of 0.0282, Fig.\u0026nbsp;3H). The overall AUC values for the classifier without and with thermostability were 0.86\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01 and 0.89\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01, respectively (Fig.\u0026nbsp;3I).\u003c/p\u003e\n\u003cp\u003eGiven there may be an imbalance in the size of immunogenic and non-immunogenic class-specific datasets at each iteration, we compared the precision-recall (PR) curve as a trade-off for the true positive rate and the positive predictive value where a higher precision shows a lower false positive rate and higher recall (sensitivity) corresponds with a lower false negative rate. Significantly higher PR area under the curve (AUCPR) values were reported for the model with thermostability (0.89\u0026thinsp;\u0026plusmn;\u0026thinsp;0.01) compared with the model without (0.75\u0026thinsp;\u0026plusmn;\u0026thinsp;0.02) (Fig.\u0026nbsp;3J). Furthermore, upon a 5-fold interval cross-validation strategy for assessing the model with thermostability, good segregation was achieved by drawing cross-validated responses against the transformed scores of the immunogenicity classes (Fig.\u0026nbsp;3K - Additional metrics are reported in Supplemental Data \u0026ndash; Figure S9). In summary, the model with thermostability was better at predicting immunogenic VACV peptides.\u003c/p\u003e\n\u003ch3\u003eExpanding the immunogenicity model for vaccinia and influenza A virus-derived MHCI peptides\u003c/h3\u003e\n\u003cp\u003eWe expanded our model to other viruses to further demonstrate the effectiveness of the thermostability-based predictors. We exported VACV and influenza A virus (IAV) derived peptides with known immunogenicity from previously published data deposited in the immune epitope database (IEDB \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e31\u003c/span\u003e\u003c/sup\u003e). For consistency, we exported all annotated VACV and IAV data in IEDB to generate equivalent and unbiased models for each virus and have a consolidated data pre-processing approach before modelling for all VACV- and IAV-derived MHCI peptides with validated immunogenicity profiles. We organized two discrete and confident H-2\u003csup\u003eb\u003c/sup\u003e MHCI peptide datasets with negative (n\u0026thinsp;=\u0026thinsp;115) and positive (n\u0026thinsp;=\u0026thinsp;185) T cell reactivity. Notably, this dataset contains 105 VACV-derived peptides that were not thermally profiled by this study and 57 peptides that were not studied by Croft et al. \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. As the thermostability profiles are unknown for most of these selected peptides, we developed a thermostability model embedded in the modelling workflow to predict the thermostability category for these peptides by training with the experimental T\u003csub\u003em\u003c/sub\u003e data acquired for the ~\u0026thinsp;3,600 endogenous H-2\u003csup\u003eb\u003c/sup\u003e MHCI peptides above. Upon training the thermostability model, we subjected the IEDB-exported peptide data as the test set to predict their thermostability category (predicted thermostability categories; PTCs) as either non-stable or stable so that they could be incorporated into the subsequent immunogenicity model (Fig.\u0026nbsp;4A). We trained the model with Model III, incorporating predictors derived from LFSE and SMI sequence scoring functions, with and without PTC, as the input layer for training the ANN algorithm. The model training was iterated 600 times by bootstrapping to select different training and test sets from throughout the data space and to evaluate the robustness of the model. At each iteration, the ANN modelling algorithm selected 70% (210 IDs) of the dataset as the training set and the rest of the peptides (30%, 90 IDs) as the test set. Then, the trained model was utilized to distinguish immunogenic peptides (Fig.\u0026nbsp;4A). The model with the PTC predictors significantly outperformed the model without PTCs (Mann-Whitney test statistical test, \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.0001, Fig.\u0026nbsp;4B). However, considerable variation was observed in the model accuracy; therefore, we chose the 30 most accurate models for further comparison, wherein models with the PTC data were more accurate in predicting immunogenic peptides compared with the models without PTCs (Mann-Whitney statistical test, \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;=\u0026thinsp;0.0066) (Fig.\u0026nbsp;4C).\u003c/p\u003e\n\u003cp\u003eWe examined the models without and with the PTCs across multiple model performance evaluation measures. We assessed the AUC and AUCPR for the models, showing again a significant improvement by including the PTC predictor in distinguishing immunogenic peptides (Fig.\u0026nbsp;4D-E). We showed the minimum and maximum of the obtained ROC curves for the classifier without and with the PTCs predictor (Fig.\u0026nbsp;4F). We performed a paired-wise comparison, demonstrating a clear difference and improvement in the model with the PTCs. The ROC curves are shown for the non-immunogenic and immunogenic peptide classes alongside the overall ROC to show the performance of the classifier in the best model with the involvement of the PTC predictor (AUC\u0026thinsp;=\u0026thinsp;0.87, Fig.\u0026nbsp;4G), compared favourably to the ROC curves for the best model without PTCs (AUC\u0026thinsp;=\u0026thinsp;0.80, Supplemental Data \u0026ndash; Figure S11A).\u003c/p\u003e\n\u003cp\u003eSince the length of the MHC peptides may impact the model performance with the change in the number of variables in the sequence descriptors, we first analyzed the length distribution of the peptide data in both negative and positive datasets (Fig.\u0026nbsp;4H). Subsequently, at each iteration of the bootstrapping process for selecting the training (70%, 210 IDs) and test (30%, 90 IDs) sets, we applied these proportions to achieve the same length distribution to all the models. Furthermore, we performed a paired-wise t-test per iteration. Once again, the PTC predictor significantly improved model accuracy and AUC metrics (paired t-test, \u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.0001, Fig.\u0026nbsp;4I-J).\u003c/p\u003e\n\u003cp\u003eIn the regression models, we demonstrated that the involvement of experimentally calculated T\u003csub\u003em\u003c/sub\u003e could improve the model in predicting the magnitude of the CD8\u003csup\u003e+\u003c/sup\u003e T cell responses better than BA. Here, we assessed the models with PTCs and compared them with the models with the inclusion of predicted binding affinities (from NetMHCpan 4.1 for the H-2D\u003csup\u003eb\u003c/sup\u003e/-2K\u003csup\u003eb\u003c/sup\u003e alleles). We show that PTCs can significantly improve the peptide immunogenicity models more than BA (ordinary one-way ANOVA, \u003cem\u003ep\u003c/em\u003e-value of 0.0278, Fig.\u0026nbsp;4K). As multicollinearity (occurring when two independent predictors show correlation) has been shown above for thermostability and BA, an algebra-based linear combination can be a potential input for the ANN model to boost the immunogenicity model further. Thus, we included both BA and PTC predictors in the model, and the resultant model could significantly improve prediction compared with the models with only PTC (p-value of 0.0096) or BA (\u003cem\u003ep\u003c/em\u003e-value\u0026thinsp;\u0026lt;\u0026thinsp;0.0001) (Fig.\u0026nbsp;4K). Thus, we demonstrated that predicted thermostability profiles can be considered a better and synergistic predictor than BA alone to interpret peptide immunogenicity.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n\u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n\u003c/div\u003e"},{"header":"Discussion","content":"\u003cp\u003eTo stimulate antigen-specific T-cell responses, an immunogenic peptide should meet the prerequisites to first bind with sufficient affinity to MHC molecules and for the complex to be stable enough to allow for prolonged presentation on the cell surface for circulating T-cell immunosurveillance \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e. Therefore, it has been hypothesized that the stability of pMHC can impact the immunogenicity of virus-derived epitopes and cancer neoepitopes presented to CD8\u003csup\u003e+\u003c/sup\u003e T cells \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e7\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e\u003c/sup\u003e. Although the impact of pMHC thermostability on peptide immunogenicity has been shown in earlier studies \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e10\u003c/span\u003e\u003c/sup\u003e, few have assessed this in the context of robust immunogenicity models. In this study, we sought to investigate the extent to which the thermostability profiles of viral peptides correlated with their immunogenicity. We utilized an optimized MS-thermostability screening protocol to generate T\u003csub\u003em\u003c/sub\u003e profiles for endogenous H-2\u003csup\u003eb\u003c/sup\u003e peptides presented on DC2.4 cells and for a subset of VACV peptides that spanned a range of known CD8\u003csup\u003e+\u003c/sup\u003e responses. Notably, we show that more stable VACV-derived peptide MHCI complexes are more immunogenic with higher measured anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses and generally have higher measured binding affinities. Importantly, we show that despite a strong correlation between binding affinity and thermostability, these parameters are complementary predictors, and their combination into AI-aided models predicts immunogenicity with improved accuracy.\u003c/p\u003e\n\u003cp\u003eWe developed and validated ML-based models to predict virus-derived MHCI peptide immunogenicity. Using a non-linear multivariate regression ANN algorithm, a quantitative model predicted anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses for viral peptides, with significant improvements when thermostability data was included. Overall, according to the observed robustness of these regression models, we can further improve the model if we train the models with an appropriate dataset sampled to have diversity in immunogenicity levels. A classification model that effectively distinguishes immunogenic from non-immunogenic and endogenous self-peptides was also developed that too was improved when incorporating a pMHCI thermostability predictor. Collectively, these models can confer a promising approach to predict anti-viral CD8\u003csup\u003e+\u003c/sup\u003e T cell responses and distinguish immunogenic from non-immunogenic peptides. Although the models suffer from limitations in the training data size and the lack of large data for model validation (internal and external), this study underpins approaches for developing the next generation of immunogenicity models by implementing potentially impactful features of pMHCI complex stability.\u003c/p\u003e\n\u003cp\u003eIn addition, uncovering different features of the structure and interaction of the pMHC-TCR can be another important feature for decoding the peptide immunogenicity and leveraging these features as predictors for the next generation of AI-aided models for predicting immunogenicity \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e33\u003c/span\u003e\u003c/sup\u003e. To acquire a comprehensive snapshot of immune specificity, it is also important to simultaneously study the interaction between the pMHCI complex and CD8\u003csup\u003e+\u003c/sup\u003e T cells \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e4\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e34\u003c/span\u003e\u003c/sup\u003e. Indeed, TCR repertoire is the third feature governing peptide immunogenicity. The predictors describing the interactions of TCR-pMHC alongside high-throughput TCR sequencing workflows may further improve the next generation of predictive immunogenicity models \u003csup\u003e\u003cspan class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan class=\"CitationRef\"\u003e36\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eAltogether, these findings demonstrate an important correlation between thermostability and immunogenicity for pMHCI complexes as a superior predictive matrix for next-generation peptide immunogenicity models. This may be predictive of prolonged cell surface dwell times, enhancing the likelihood of T cell interaction and induction of T cell responses. Thus, including pMHC stability predictors is an effective strategy in choosing targets for T cell immunotherapy. The prediction of peptide immunogenicity may thus be strengthened by implementing this additional thermostability dimension as a new predictor for AI-aided models for immunotherapy. The thermostability screening can assist in shortlisting MHC peptides derived from tissues, biopsies, and patient-derived cells to conduct more efficient discovery of neoepitopes for cancer vaccination and immunotherapy.\u003c/p\u003e"},{"header":"Methods","content":"\u003ch2\u003e4.1. Experimental Design and Statistical Rationale\u003c/h2\u003e\u003cp\u003eWe implemented well-established immunoaffinity purification (IP) protocols to prepare sample fractions and biological replicates by small-scale IP for data acquisition through DIA and MRM\u003csup\u003eHR\u003c/sup\u003e techniques for MS-thermostability measurement assay from lysed cells \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e. We used a large-scale IP protocol to acquire DDA data to generate spectral libraries of endogenous MHCI H-2D\u003csup\u003eb\u003c/sup\u003e and H-2K\u003csup\u003eb\u003c/sup\u003e and exogenous VACV peptides. The DIA and MRM\u003csup\u003eHR\u003c/sup\u003e (targeted) datasets were acquired on MHCI peptides isolated and purified from smaller samples using an optimized small-scale IP and peptide elution protocol. These resulting datasets were utilized to develop an MS-thermostability measurement assay to screen pMHCI complexes. The DIA dataset contains endogenous MHCI H-2D\u003csup\u003eb\u003c/sup\u003e and − 2K\u003csup\u003eb\u003c/sup\u003e peptides purified and isolated from DC2.4 cells. Both DIA and MRM\u003csup\u003eHR\u003c/sup\u003e datasets are expected to have VACV-derived MHCI peptides. The standard sequence analysis approaches were used to assess the results derived from MS-thermostability assay, e.g., pMHCI binding prediction, cluster analysis, Venn overlap graphs, and sequence motif analysis. We employed computational MS workflows for data visualization and quantitative analysis of the MS datasets for thermostability screening of pMHCI complexes. We applied the multiple pairwise comparisons derived from one-way and two-way ANOVA analysis and standard \u003cem\u003et\u003c/em\u003e-test analyses as canonical statistical tests to assess the measurements of different features of pMHCI complexes. Machine learning algorithms (ANN and SVM) were used to construct regression and classification models to predict anti-VACV CD8\u003csup\u003e+\u003c/sup\u003e T cell responses and immunogenic peptides.\u003c/p\u003e\u003ch3\u003e4.2. Cell culture\u003c/h3\u003e\u003cp\u003eWe cultured the DC2.4 cell line as immortalized murine dendritic cells manufactured through transduction of bone marrow isolates derived from C57BL/6 mice by retrovirus vectors that can express MHC class I and II H-2\u003csup\u003eb\u003c/sup\u003e alleles to present endogenous self-peptides \u003csup\u003e\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e\u003c/sup\u003e. The DC2.4 were harvested and maintained in RF-10 media of RPMI-1640 medium (Gibco™, Thermo Fisher Scientific, Waltham, MA). The media was supplemented with 10% fetal calf serum (FCS; Sigma-Aldrich), 50 µM of β-mercaptoethanol (Sigma-Aldrich), 1% (v/v) mM non-essential amino acids (Gibco™), 5 mM HEPES (Sigma-Aldrich, St Louis, MO), 2 mM L-glutamine (MP Biomedicals), 50 µg mL-1 Penicillin-Streptomycin (Pen-Strep, Gibco™), and 0.3 mg mL-1 hygromycin (Invitrogen) (to support a stable expression of the transfected MHC-I allotypes). The cells were expanded and grown in flasks and roller bottles incubated at 37°C and a 5% CO2 atmosphere. The cells were counted, harvested by centrifugation (3724 × \u003cem\u003eg\u003c/em\u003e at 4°C for 15 min), washed with phosphate-buffered saline (PBS), pelleted, snap-frozen in liquid nitrogen,\u003c/p\u003e\u003ch2\u003e4.3. Synthetic VACV peptides\u003c/h2\u003e\u003cp\u003eThe synthetic peptides matched with the identified VACV sequences (by Croft et al. \u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e) were purchased from Mimotopes (VIC, Australia) and Genscript (Piscataway, NJ) with HPLC purity from 90–99%. The original stocks were dissolved in 100% DMSO (LC-MS grade), diluted to have 5 mM per peptide, and stored at -20℃. The peptides were diluted into water (Optima™ LC/MS Grade, Fisher Scientific®) containing 0.1% (v/v) formic acid to analyze by LC-MS/MS in DDA and MRM\u003csup\u003eHR\u003c/sup\u003e modes.\u003c/p\u003e\u003ch2\u003e4.4. Pulsing DC2.4 cells with vaccinia virus peptides\u003c/h2\u003e\u003cp\u003eDC2.4 cells were loaded with a mixture of 122 exogenous viral peptides. First, we detached adherent cells (by 2.5 mM of EDTA in PBS) and suspended them in the RF10 media. Then, the cells were counted, and the required amount was transferred to a separate 50 mL tube (4×10\u003csup\u003e8\u003c/sup\u003e to 5×10\u003csup\u003e8\u003c/sup\u003e cells per pellet). After spinning at 1,500 rpm for 5 mins at RT, the cells were washed with PBS twice. The cells were resuspended in RPMI-1640 medium (Gibco™, Thermo Fisher Scientific, Waltham, MA) at a cell density of 2×10\u003csup\u003e7\u003c/sup\u003e cells mL\u003csup\u003e− 1\u003c/sup\u003e. Afterwards, we added the mixture of synthetic VACV peptides so that the final concentration of each peptide was 1 µM. The pulsed cells were incubated at 37℃ for 1 hr on a roller while we gently resuspended the cells every 10 mins. Once the incubation was complete, we extensively washed them with PBS three times (spun between the washes at 3,724 × \u003cem\u003eg\u003c/em\u003e, for 10 mins, at 4℃). The cells were pelleted and snap-frozen via liquid nitrogen and then stored at -80°C.\u003c/p\u003e\u003ch2\u003e4.5. Large-scale isolation and elution of MHCI peptides\u003c/h2\u003e\u003cp\u003eA well-established large-scale immunoprecipitation protocol was utilized to isolate and purify peptides bound to MHCI H-2D\u003csup\u003eb\u003c/sup\u003e/-2K\u003csup\u003eb\u003c/sup\u003e expressed individually on the surface of the DC2.4 cells using monoclonal antibodies. A frozen pellet with 5×10\u003csup\u003e8\u003c/sup\u003e cells was pulverized by liquid nitrogen-cooled cryo-milling (Retsch® MM 400 mixer mill) and then resuspended in lysis buffer (50 mM Tris pH 8 (Sigma-Aldrich), 150 mM NaCl (Merck-Millipore, Darmstadt, Germany), 0.5% IGEPAL 630 (Sigma-Aldrich), and cOmplete™ Protease Inhibitor Cocktail (Roche Applied Science®)), and then incubated on a roller for 60 min at 4℃. The resulting lysate was pre-cleared by centrifugation at 2,000 × \u003cem\u003eg\u003c/em\u003e, 4°C for 10 min, and the supernatant was ultracentrifuged at 100,000 × g, 4°C for 45 min. MHCI peptides were isolated and purified from the lysate immunoaffinity chromatography. We utilized the prepared immunoaffinity columns of monoclonal antibodies of choice (purified from hybridomas in-house), i.e., Y-3 (anti-H-2K\u003csup\u003eb\u003c/sup\u003e; ATCC® HB-176™) and 28-14-8S (anti-H-2D\u003csup\u003eb\u003c/sup\u003e; ATCC® HB-27™), cross-linked to Protein A sepharose (each antibody/Protein A; 10 mg mL\u003csup\u003e− 1\u003c/sup\u003e) for immunoprecipitation (as detailed \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e), separately per each antibody. The captured pMHC complexes throughout the immunoaffinity column were dissociated by 10% acetic acid. After pMHC immunoaffinity purification, the eluate (a composition of MHCI heavy chains, β\u003csub\u003e2\u003c/sub\u003e-microglobulin (β\u003csub\u003e2\u003c/sub\u003em), and peptide ligands) was fractionated through C18 reversed-phase (RP) end-capped high-performance liquid chromatography (HPLC) column (4.6 mm internal diameter × 100 mm long; Chromolith® SpeedROD; Merck-Millipore) on ÄKTAmicro™ HPLC System (GE Healthcare, UK controlled by the UNICORN™ ver. 5.11 software) to separate peptides from the MHC heavy chains and β\u003csub\u003e2\u003c/sub\u003em molecules. A mobile phase constituted of buffer A (0.1% trifluoroacetic acid (TFA) [Thermo Fisher Scientific, San Jose, CA]) and buffer B (80% acetonitrile (ACN) [Thermo Fisher Scientific, Waltham, MA], and 0.1% TFA). The separated fractions were mixed into nine pools of peptides and then concentrated by vacuum centrifugation (CentriVap Benchtop Vacuum Concentrator [LABCONCO®]). The concentrated peptides were resuspended in 15 µL of 0.1% v/v of formic acid (FA) in water (Optima™ LC/MS Grade, Fisher Scientific®). 200 fmoles of iRT peptides were spiked into each sample as an internal retention time (iRT) standard for the RT normalization \u003csup\u003e\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e\u003c/sup\u003e. The samples were stored at -80℃ until LC-MS/MS analysis.\u003c/p\u003e\u003ch2\u003e4.6. LC-MS/MS analysis to identify MHC-bound peptides\u003c/h2\u003e\u003cp\u003e \u003cem\u003eDC2.4 and VACV DDA data –\u003c/em\u003e A SCIEX ZenoTOF 7600 LC-MS/MS system - high-resolution mass spectrometer (SCIEX™) coupled to ACQUITY™ UPLC M-Class (Microscale) LC System (Waters™, Milford, MA), Waters™ ACQUITY UPLC Console, was used to analyse isolated MHC-bound peptides and acquire MS/MS spectra (operated in the DDA mode). The UPLC system was equipped with a Waters™ nanoACQUITY Sample Manager and a Waters™ ACQUITY Auxiliary Solvent Manager as the autosampler and LC pump, respectively. 4 µL of sample fractions were loaded with a flow rate of 5 µL min\u003csup\u003e− 1\u003c/sup\u003e onto a Kinetex® 00F-4496-AC - XB-C18 LC column (0.3 mm internal diameter × 150 mm length, Particle Size 2.6 µm, Pore Size 100 Å [Phenomenex®, Torrance, CA]). The peptides were eluted at 5 µL min\u003csup\u003e− 1\u003c/sup\u003e flow rate over a gradient of buffer A (0.1% FA) and B (98% ACN, 0.1% FA): 0–1 min 3% B, 1–46 min 30% B, 46–47 min 80% B, 47–49 min 80% B, 49–50 min, then down to 3% B, and 50–55 min re-equilibration at 3% B. The mass spectrometer was operated in data-dependent acquisition (DDA) mode with collision-induced dissociation (CID) fragmentation. The other parameters were set to the following settings: precursor charge state of 2 + to 5+, full-scan MS1 range 400–1,500 m/z, collision energy of 10 V, MS2 scan range of 100–1,800 m/z, dynamic exclusion of 5 s, switch criteria of intensity greater than 100 counts per second, mass tolerance of 50 mDa, precursors selected for MS/MS of 45 per cycle time, declustering potential of 80 V, and accumulation times of 150 ms and 20 ms for MS1 and MS/MS, respectively.\u003c/p\u003e\u003cp\u003e \u003cb\u003e4.7. DDA database search to discover immunopeptidomes and generation of a spectral library of H-2b and VACV-derived pMHCIs\u003c/b\u003e \u003c/p\u003e\u003cp\u003e \u003cem\u003eMS/MS spectra from DDA data matched with MHCI-bound peptide sequences\u003c/em\u003e - MS/MS spectra (PSMs) were searched against a combined database of the mouse (\u003cem\u003eMus musculus\u003c/em\u003e; UP000000589) and vaccinia virus strain Western Reserve (VACV-WR; UP000000344) proteomes (UniProtKB/SwissProt v.19102022; 17,560 entries) to identify and sequence MHC-associated peptides for generating an extensive spectral library. The database search provided a sizeable peptide repertoire containing pMHC targets for spectral library-based DIA data searches. An MSBooster-assisted database search using MSFragger v. 3.7 (over a computational platform - through a GUI pipeline; FragPipe v. 19.1) \u003csup\u003e\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u003c/sup\u003e was used to process and search LC − MS/MS DDA data against the mentioned database alongside a database of 17,368 decoys (50.0% - the same size as the target proteome database) and 11 standard iRT peptide sequences. A spectral library of the immunopeptidome-derived spectra was generated from this DDA dataset validated by the target-decoy algorithm with a false discovery rate (FDR) threshold of 5% for confident peptide identification by Philosopher v. 4.8.0 \u003csup\u003e41\u003c/sup\u003e, which is a customized algorithm to filter MS/MS spectra and an FDR estimator via a multi-level target-decoy algorithm. MSBooster \u003csup\u003e\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u003c/sup\u003e was used in tandem with MSFragger and Percolator \u003csup\u003e\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e\u003c/sup\u003e to rescore PSMs through deep learning-based predictions of a set of features (i.e., LC retention time and MS/MS spectra). EasyPQP (version 0.1.35) was used to generate the consensus spectral library. In the MSFragger search, the peak matching parameters were set to mass calibration and parameter optimization, precursor and fragment mass error tolerances were set to 20 ppm, and isotope error was set to 0/1/2. In the protein digestion setting, cleavage was set to “nonspecific” and enzyme specificity to “nonspecific” with up to 2 missed cleavages. Peptide length was set from 7 to 25, and peptide mass range was set from 200 to 5,000 Da. Variable post-translational modifications (PTMs) were set to oxidation (M [+ 15.99]) and deamidation (NQ [+ 0.98]) with a maximum number of three variable PTMs for each peptide. A wide-scale spectral library of the immunopeptidome-derived MS/MS spectra was generated from this DDA dataset.\u003c/p\u003e\u003ch2\u003e4.8. Small-scale immunoprecipitation of pMHC complexes to develop MS thermostability assay\u003c/h2\u003e\u003cp\u003eWe implemented a small-scale immunoprecipitation protocol \u003csup\u003e\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e, modified from the formerly developed protocol \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e\u003c/sup\u003e, to purify and isolate pMHCs. The pellets of DC2.4 cells (pulsed with VACV peptides) were lysed through resuspending in the lysis buffer. The resultant lysates were incubated on a roller for 45 mins at 4°C and then cleared by centrifuge at 3700 × \u003cem\u003eg\u003c/em\u003e for 10 mins at 4°C. The supernatant was ultracentrifuged at 100,000 × \u003cem\u003eg\u003c/em\u003e and 4°C for 45 min. Then, the collected supernatant was run over pre-column (with 1 mL of Protein A resin) and flow lysate through. After pre-column, cleared lysates were split into aliquots of 5×10\u003csup\u003e7\u003c/sup\u003e cell equivalents, each triplicate incubated at 37°C, 41°C, 45°C, 49°C, 53°C, 57°C, 61°C, 65°C, 69°C, and 73°C, for 10 mins (\u003cem\u003en\u003c/em\u003e = 3; where \u003cem\u003en\u003c/em\u003e denotes the number of biological replicates per datapoint with specific thermal treatment). Therefore, for ten temperature treatments, we prepared 30 sample replicates. The resultant thermally-treated samples were incubated overnight at 4°C with 500 µg of a mixture of purified 28-14-8S (anti-H-2D\u003csup\u003eb\u003c/sup\u003e) and Y-3 (anti-H-2K\u003csup\u003eb\u003c/sup\u003e) antibodies (with a ratio of 1:1–250 µg of each antibody) bound to protein A sepharose for immunoprecipitation. The pMHC complexes captured and bound to the Protein A resin-antibody were settled down and filtered through MobiSpin columns (with 10 µm pore size filters [MoBiTec GmbH, Germany]) followed by extensive washes by PBS. 400 µL of 10% acetic acid (in Optima™ LC-MS grade water) was added to release and elute the bound pMHC complexes through the Mobispin columns. Eluates constituting class-I heavy chains, β\u003csub\u003e2\u003c/sub\u003em, and antibodies were filtered (from class I HC and β\u003csub\u003e2\u003c/sub\u003em molecules) using pre-washed 5 kDa MWCO (molecular weight cut-off) centrifugal filter units (Ultrafree®-MC-PLHCC, Merck Millipore, Germany) through centrifuging for 60 min at 16,000 × \u003cem\u003eg\u003c/em\u003e to collect filtered samples in new Eppendorf tubes. Excess 200 µL of 10% acetic acid was added to wash filters to elute the remaining peptides. Eluted peptides were desalted and eluted using C18-Omix 100 µL Pipette Zip tips (Agilent©, OMIX A57003100) by a buffer of 40% ACN and 0.1% FA (in Optima™ water LC-MS grade). Cleaned-up samples were concentrated by vacuum centrifugation by running Labconco Vacuum Concentrator System. MHC peptides were reconstituted in 15 µL of 0.1% v/v FA buffer in Optima water per replicate. 200 fmoles of the standard iRT peptides were spiked into the samples for retention time prediction and peak normalization for LC-MS/MS analysis. The samples were sonicated for 10 min and centrifuged at 21,000 × \u003cem\u003eg\u003c/em\u003e for 10 min before transferring to MS vials to prepare for MS analyses.\u003c/p\u003e\u003ch2\u003e4.9. Zeno SWATH-DIA LC-MS/MS data independent acquisition\u003c/h2\u003e\u003cp\u003eA SCIEX ZenoTOF 7600 system - high-resolution mass spectrometer (SCIEX™) coupled to ACQUITY™ UPLC M-Class (Microscale) LC System (Waters™, Milford, MA), controlled by Waters™ ACQUITY UPLC Console, was operated for LC-MS/MS DIA analysis of sample biological replicates prepared by small-scale immunoprecipitation as described above after the thermal treatments. The injection volume was 4 µL to load the sample onto a Kinetex® 00F-4496-AC - XB-C18 LC column (0.3 mm internal diameter × 150 mm length, Particle Size 2.6 µm, Pore Size 100 Å [Phenomenex®, Torrance, CA]) to elute through. The analytical column and LC gradient method were identical to the DDA data acquisition. DIA parameters were as follows: 75 DIA scans with the fixed isolation precursor windows of 8 Da (m/z) ranging from 375 to 975 m/z, MS1 scans across 400-1,500 m/z, collision energy of 10, MS2 scan range of 140–1,800 m/z, MS/MS collision energy of 17, and accumulation times of 100 ms and 15 ms for MS1 and MS/MS scans, respectively.\u003c/p\u003e\u003ch2\u003e4.10. Library search-based DIA data analysis\u003c/h2\u003e\u003cp\u003eWe employed two “peptide-centric” DIA software tools (i.e., Skyline ver. 22.2 \u003csup\u003e44, 45\u003c/sup\u003e, and DIA-NN ver. 1.8.0 \u003csup\u003e46\u003c/sup\u003e to analyze the DIA dataset at 5% FDR (at peptide level) using an extensive DDA spectral library previously generated. The settings are described in Supplementary Material. Perseus \u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e (version 2.0.7.0) was used for post-proteome and peptidome data analysis.\u003c/p\u003e\u003ch2\u003e4.11. Computational workflow for drawing denaturation profiles and T\u003csub\u003em\u003c/sub\u003e value estimation\u003c/h2\u003e\u003cp\u003eThe computational workflow comprises DIA data processing followed by the MS data post-analysis through \u003cem\u003ein-house\u003c/em\u003e code and the Perseus pipeline \u003csup\u003e\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e47\u003c/span\u003e\u003c/sup\u003e. Through the spectral library-based DIA data processing, the DIA software of choice \u003csup\u003e\u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e48\u003c/span\u003e\u003c/sup\u003e (a combination of \u003cem\u003eSkyline\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e, and \u003cem\u003eDIA-NN\u003c/em\u003e \u003csup\u003e\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e\u003c/sup\u003e) matches the acquired composite thermostability data (peptide-specific spectra) with individual peptide MS/MS within the spectral library. After the library search and identifying specific peptide ligands, the raw chromatographic data for an eluted peptide is directly exported, and the accumulated peak areas (total peak area) for the spectral transitions can be utilized to quantify a peptide across replicates. To draw denaturation profiles for individual pMHCs, we extracted the total chromatographic peak area from each temperature-specific data point. All precursor peak areas (quantities) were normalized to the reference temperature data point (corresponding peak area) at 37°C to have fold-changes and draw denaturation profiles. This temperature was chosen as this is the condition where we expect a maximal peak intensity. We calculated the peak area under the curve (at different MS levels, i.e., MS1, MS2, and total area) per integrated chromatographic profile (assigned to a precursor (peptide) and its fragment ion peaks) to measure the relative quantity of peptides in each replicate. By the customized computational MS workflow, a sigmoidal curve was fitted to the denaturation profiles derived from SWATH-MS data (for the precursors identified at least in two biological replicates) to estimate the T\u003csub\u003em\u003c/sub\u003e value.\u003c/p\u003e\u003ch2\u003e4.12. MRM\u003csup\u003eHR\u003c/sup\u003e LC-MS/MS data acquisition\u003c/h2\u003e\u003cp\u003eThermally-treated samples were analyzed for the quantification of targeted VACV peptides by MRM\u003csup\u003eHR\u003c/sup\u003e MS technique, acquired on SCIEX ZenoTOF 7600 system - high-resolution mass spectrometer (SCIEX™) coupled to ACQUITY™ UPLC M-Class (Microscale) LC System (Waters™, Milford, MA), controlled by Waters™ ACQUITY UPLC Console utilizing SCIEX OS software (v. 3.1) for data acquisition. The injection volume was 4 µL to load the sample onto the C18 column and eluted through. The analytical column and LC gradient method were identical to the DDA data acquisition. We exported collision energy (CE) values from the Zeno SWATH results for VACV peptides to set up the acquisition method for the MRM\u003csup\u003eHR\u003c/sup\u003e analysis. The LC-MS/MS system was operated in the scheduled MRM\u003csup\u003eHR\u003c/sup\u003e mode with CID fragmentation. The acquisition parameters were set as follows: CE from 16 to 47 V (dependent on the targeted precursor), precursor charges of + 1 to + 3, MS1 scan range of 210–1,250 m/z, declustering potential of 80 V, MS2 scan range of 100–2,000 m/z, and accumulation times of 200 ms and 5 ms for MS1 and MS/MS scans, respectively. MRM\u003csup\u003eHR\u003c/sup\u003e transitions (159 singly and doubly-charged precursors corresponding with 122 unique peptide sequences were tracked as a targeted list) are listed in Supplementary Data. The MRM\u003csup\u003eHR\u003c/sup\u003e dataset was processed and analyzed in Skyline \u003csup\u003e\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e\u003c/sup\u003e (v. 22.2; MacCoss Laboratory, University of Washington, Seattle, WA). We validated the detected VACV-derived precursors using two approaches. First, we examined the RT values of MRM\u003csup\u003eHR\u003c/sup\u003e precursors by comparing them with the Zeno SWATH data using the iRT predictor. Next, we matched MS/MS spectra with the spectral library of the VACV peptides derived from synthetic peptides.\u003c/p\u003e\u003ch2\u003e4.13. Statistics\u003c/h2\u003e\u003cp\u003eWe used standard statistical tests, e.g., ordinary one-/two-way ANOVA, Mann-Whitney test, Kruskal–Wallis, and Tukey’s multiple pairwise comparisons, to assess the differences between data. Z-score was used for clustering analysis calculated per peptide by subtracting the mean of intensities of each precursor from individual precursor signals in each replicate divided by the standard deviation of each precursor signal across temperatures. A \u003cem\u003ep\u003c/em\u003e-value of ≤ 0.05 was considered the statistically significant cut-off for all statistical analyses.\u003c/p\u003e\u003ch2\u003e4.14. Software tools for peptide sequence and statistical data analysis\u003c/h2\u003e\u003cp\u003eAll statistical data analyses were executed by GraphPad Prism v. 9.0.0 and MATLAB programming code routines (The MathWorks Inc., Natick, MA) v. R2021a. The Classification toolbox was used to construct machine-learning models and evaluate their performance \u003csup\u003e\u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e49\u003c/span\u003e\u003c/sup\u003e. NetMHCpan (v. 4.1) \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e was used to determine allelic specificity according to peptide binding rank. The peptides were also segregated based on their sequence features using GibbsCluster (v. 2.0) \u003csup\u003e\u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e51\u003c/span\u003e\u003c/sup\u003e. Seq2Logo \u003csup\u003e\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e52\u003c/span\u003e\u003c/sup\u003e was used to produce sequence motif analysis of the identified HLA-bound peptides. We used BioVenn \u003csup\u003e\u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e53\u003c/span\u003e\u003c/sup\u003e and InteractiVenn \u003csup\u003e\u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e54\u003c/span\u003e\u003c/sup\u003e to generate Venn overlap graphs. BioRender (BioRender.com) and Microsoft PowerPoint were used to design and create schematic figures and experimental workflows.\u003c/p\u003e\u003ch2\u003e4.15. ANN algorithm setting and network structure\u003c/h2\u003e\u003cp\u003eWe used a feedforward fully connected neural network with two hidden layers for the ANN-based regression model. The first layer is fully connected with the input layer (predictors), with a size of 30 neurons. In the first layer, the input data are multiplied by a weight matrix (in each fully connected neuron of the layer) and then corrected by adding a bias vector. The size of the second hidden layer was set to 10, which contains the activation function after the first hidden layer. The activation function was set to the rectified linear unit (ReLU), which applies a threshold operation to the data to set negative values to zero. The final fully connected layer is named the output layer and generates the final output of the network, which is the predicted response. For the ANN classifier, the size of the hidden layers was set to 20 and 10 for the first and second fully connected networks, respectively.\u003c/p\u003e\u003ch2\u003e4.16. Parameters to assess the models\u003c/h2\u003e\u003cp\u003eSome parameters used to assess the classification model in recognizing immunogenic from non-immunogenic peptides are defined as follows: Precision (also known as the positive predictive value) is the ratio of the number of true positives to the sum of the true positives and false positives, explaining how well a model can predict the positive class. Sensitivity (or recall) is the ratio of the number of true positives to the sum of the true positives and the false negatives. Specificity is the ratio of the number of true negatives to the sum of the number of true negatives and false positives.\u003c/p\u003e\u003ch2\u003e4.17. Sequence encoding functions\u003c/h2\u003e\u003cp\u003eThe LFSE method was used as an effective sequence scoring function that can convert sequence strings to binary vectors logically by a numerical substitution based on the amino acid residues to provide a unique sequence descriptor for individual MHCI peptides \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e. We used the SMI scoring function to encode physicochemical characteristics-based residue groups. In the SMI method, amino acids are grouped according to the characteristics of the side chains. In the first categorical strategy, per sequence, the code substitutes the aliphatic amino acids, i.e., alanine, glycine, isoleucine, leucine, proline, and valine, with one. This function replaces aromatic (phenylalanine, tryptophan, and tyrosine), acidic (aspartic acid and glutamic acid), basic (arginine, histidine, and lysine), hydroxylic (serine and threonine), sulfur-containing (cysteine and methionine), and amidic residues (asparagine and glutamine), with two to seven, respectively.\u003c/p\u003e\u003ch2\u003e4.18. Parameters for the IEDB search and post refinements of data\u003c/h2\u003e\u003cp\u003eWe used the following setting as the search parameters for the initial search in the IEDB database’s website: epitope structure as only linear peptide sequences, no B cell assays, no MHC assays, only T cell assays, only positive assays (for the positive dataset), only negative assays (for the negative dataset), MHC restriction type of class I, and host as mouse. The post-exportation filters to make the data ready for ML modelling were the following refinements: 1) Mus musculus (C57BL/6) as the host; 2) VACV- and IAV-derived MHCI peptides; 3) CD8\u003csup\u003e+\u003c/sup\u003e T cell responses as either positive or negative; 4) IFNγ release as the measured T cell assay response; 5) Removing the joint peptides between negative and positive datasets; 6) Subjecting the data to NetMHCpan (v. 4.1) \u003csup\u003e\u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e50\u003c/span\u003e\u003c/sup\u003e and removing non-H-2D\u003csup\u003eb\u003c/sup\u003e/-2K\u003csup\u003eb\u003c/sup\u003e binders; 7) Sequence length range was 8-to-11-mer for the selected H-2\u003csup\u003eb\u003c/sup\u003e MHCI ligands.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAcknowledgements\u003c/h2\u003e\n\u003cp\u003eThe R@CMon/Monash Node of the NeCTAR Research Cloud, Australia\u0026rsquo;s national research cloud specifically designed for research computing, supported computational resources. We appreciate Rochelle Ayala\u0026apos;s technical help and lab resources.\u003c/p\u003e\n\u003ch2\u003eData availability\u003c/h2\u003e\n\u003cp\u003eAll LC-MS/MS immunopeptidomics data, MSFragger DDA database search, SWATH-DIA library search results, MRM\u003csup\u003eHR\u003c/sup\u003e results, peptide quantification matrices, and thermostability results have been organized and uploaded to the ProteomeXchange Consortium via the PRIDE partner repository \u003csup\u003e55, 56\u003c/sup\u003e. For DDA data acquired from the DC2.4 dendritic cell line and synthetic VACV by DDA LC-MS/MS (n = 27), the accession code is PXD057866 (Username:
[email protected] - Password: kYZxYytIptVf). DC2.4 cells pulsed by VACV peptides Zeno SWATH-DIA LC-MS/MS (n = 30) and LC-MRM\u003csup\u003eHR\u003c/sup\u003e (n = 30) data were deposited under the dataset identifier PXD057919 (Username:
[email protected] - Password: wBb44QKOBrA0) and PXD058188 (Username:
[email protected] - Password: EKCLoDDnM01e), respectively. All search and results summaries were deposited (also provided as Supplementary Data) with the corresponding datasets.\u003c/p\u003e\n\u003ch2\u003eSupplementary information\u003c/h2\u003e\n\u003cp\u003eThis article contains supplemental data.\u003c/p\u003e\n\u003ch2\u003eAuthor Contributions\u003c/h2\u003e\n\u003cp\u003eM.S. wrote the original manuscript, performed experiments, MS-assay development, data acquisition, data analysis, wrote code to develop computational workflow for thermostability profiling, machine learning models; M.S., P.F., S.H.R., N.P.C., and A.W.P. design of experimental and computational workflows, conceptualization, methodology; P.F., D.C.T., C.L., S.H.R., N.P.C., and A.W.P. investigation; P.F., S.H.R., N.P.C., and A.W.P. supervision; N.P.C. and A.W.P. model evaluation, editing the manuscript, leading project; A.W.P. funding acquisition. \u0026nbsp;\u0026nbsp;\u003c/p\u003e\n\u003ch2\u003eConflict of interest\u003c/h2\u003e\n\u003cp\u003eAWP is a scientific advisor for Bioinformatics Solutions Inc (Canada), a shareholder and scientific advisor for Evaxion Biotech (Denmark), and a co-founder of Resseptor Therapeutics (Australia). They had no role in the design of the study; in the collection, analyses, or interpretation of data; in the writing of the manuscript; or in the decision to publish the results. There are no other conflicts of interest declared by the authors.\u003c/p\u003e\n\u003ch2\u003eFunding\u003c/h2\u003e\n\u003cp\u003eThis work was funded by grants from the National Health and Medical Research Council of Australia (NHMRC) APP1084283 (to D.C.T., A.W.P., and N.P.C.) and APP2016596 (to A.W.P.). A.W.P is supported by a NHMRC Investigator Fellowship (APP2016596). DCT was supported by a NHMRC Investigator Fellowship (APP2008990). C.L. was supported by an Australian Research Council (ARC) Future Fellowship (FT240100798) and an NHMRC Ideas Grant (2024/GNT2037597). P.F. was supported by the Victorian Department of Health and Human Services acting through the Victorian Cancer Agency, grant 2022/GNT2019729 awarded through the National Health and Medical Research Council (NHMRC) and grant NCRI000108 awarded through the Medical Research Future Fund (MRFF). Postgraduate research scholarship from Monash University, including Monash Graduate Scholarship and Monash International Tuition Support (to M.S.).\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eCalis JJA et al (2013) Properties of MHC Class I Presented Peptides That Enhance Immunogenicity. PLoS Comput Biol 9:e1003266\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRothbard JB, Gefter ML (1991) Interactions between Immunogenic Peptides and MHC Proteins. Annu Rev Immunol 9:527\u0026ndash;565\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCroft NP et al (2019) Most viral peptides displayed by class I MHC on infected cells are immunogenic. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e 116, 3112\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCroft NP (2020) Peptide Presentation to T Cells: Solving the Immunogenic Puzzle. BioEssays 42:1900200\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNeefjes J, Jongsma MLM, Paul P, Bakke O (2011) Towards a systems understanding of MHC class I and MHC class II antigen presentation. Nat Rev Immunol 11:823\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRasmussen M et al (2016) Pan-Specific Prediction of Peptide\u0026ndash;MHC Class I Complex Stability, a Correlate of T Cell Immunogenicity. J Immunol 197:1517\u0026ndash;1524\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHarndahl M et al (2012) Peptide-MHC class I stability is a better predictor than peptide affinity of CTL immunogenicity. Eur J Immunol 42:1405\u0026ndash;1416\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHellman LM et al (2016) Differential scanning fluorimetry based assessments of the thermal and kinetic stability of peptide\u0026ndash;MHC complexes. J Immunol Methods 432:95\u0026ndash;101\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBlaha DT et al (2019) High-Throughput Stability Screening of Neoantigen/HLA Complexes Improves Immunogenicity Predictions. Cancer Immunol Res 7:50\u0026ndash;61\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJappe EC et al (2020) Thermostability profiling of MHC-bound peptides: a new dimension in immunopeptidomics and aid for immunotherapy design. Nat Commun 11:6305\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan der Burg SH, Visseren MJ, Brandt RM, Kast WM, Melief CJ (1996) Immunogenicity of peptides bound to MHC class I molecules depends on the MHC-peptide complex stability. J Immunol 156:3308\u0026ndash;3314\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNicholls S et al (2009) Secondary anchor polymorphism in the HA-1 minor histocompatibility antigen critically affects MHC stability and TCR recognition. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e 106, 3889\u0026ndash;3894\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMicheletti F et al (1999) Selective amino acid substitutions of a subdominant Epstein-Barr virus LMP2-derived epitope increase HLA/peptide complex stability and immunogenicity: implications for immunotherapy of Epstein-Barr virus-associated malignancies. Eur J Immunol 29:2579\u0026ndash;2589\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSpierings E et al (2009) Steric Hindrance and Fast Dissociation Explain the Lack of Immunogenicity of the Minor Histocompatibility HA-1Arg Null Allele 1. J Immunol 182:4809\u0026ndash;4816\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLipford GB, Bauer S, Wagner H, Heeg K (1995) In vivo CTL induction with point-substituted ovalbumin peptides: immunogenicity correlates with peptide-induced MHC class I stability. Vaccine 13:313\u0026ndash;320\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003evan Stipdonk MJB et al (2009) Design of Agonistic Altered Peptides for the Robust Induction of CTL Directed towards H-2Db in Complex with the Melanoma-Associated Epitope gp100. Cancer Res 69:7784\u0026ndash;7792\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYuen Tracy J et al (2010) Analysis of A47, an Immunoprevalent Protein of Vaccinia Virus, Leads to a Reevaluation of the Total Antiviral CD8\u0026thinsp;+\u0026thinsp;T Cell Response. J Virol 84:10220\u0026ndash;10229\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBamford D, Zuckerman M (2021) Encyclopedia of virology. Academic\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChiuppesi F et al (2024) Synthetic modified vaccinia Ankara vaccines confer cross-reactive and protective immunity against mpox virus. Commun Med 4:19\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAssarsson E et al (2007) A Quantitative Analysis of the Variables Affecting the Repertoire of T Cell Specificities Recognized after Vaccinia Virus Infection. J Immunol 178:7890\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePurcell AW, Ramarathinam SH, Ternette N (2019) Mass spectrometry\u0026ndash;based identification of MHC-bound peptides for immunopeptidomics. Nat Protoc 14:1687\u0026ndash;1707\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePandey K, Ramarathinam SH, Purcell AW (2021) Isolation of HLA Bound Peptides by Immunoaffinity Capture and Identification by Mass Spectrometry. Curr Protocols 1:e92\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJensen PE, Weber DA, Thayer WR, Westerman LE, Dao CT (1999) Peptide exchange in MHC molecules. Immunol Rev 172:229\u0026ndash;238\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChefalo PJ, Harding CV (2001) Processing of Exogenous Antigens for Presentation by Class I MHC Molecules Involves Post-Golgi Peptide Exchange Influenced by Peptide-MHC Complex Stability and Acidic pH1. J Immunol 167:1274\u0026ndash;1282\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYork IA, Brehm MA, Zendzian S, Towne CF, Rock KL (2006) Endoplasmic reticulum aminopeptidase 1 (ERAP1) trims MHC class I-presented peptides\u0026thinsp;\u0026lt;\u0026thinsp;i\u0026thinsp;\u0026gt;\u0026thinsp;in vivo\u0026thinsp;and plays an important role in immunodominance. \u003cem\u003eProceedings of the National Academy of Sciences\u003c/em\u003e 103, 9202\u0026ndash;9207\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTscharke DC et al (2004) Identification of poxvirus CD8\u0026thinsp;+\u0026thinsp;T cell determinants to enable rational design and characterization of smallpox vaccines. J Exp Med 201:95\u0026ndash;104\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRemakus S et al (2018) Cutting Edge: Protection by Antiviral Memory CD8 T Cells Requires Rapidly Produced Antigen in Large Amounts. J Immunol 200:3347\u0026ndash;3352\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFlesch IEA et al (2009) Altered CD8\u0026thinsp;+\u0026thinsp;T Cell Immunodominance after Vaccinia Virus Infection and the Naive Repertoire in Inbred and F1 Mice. J Immunol 184:45\u0026ndash;55\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu R-H, Remakus S, Ma X, Roscoe F, Sigal LJ (2010) Direct Presentation Is Sufficient for an Efficient Anti-Viral CD8\u0026thinsp;+\u0026thinsp;T Cell Response. PLoS Pathog 6:e1000768\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShahbazy M et al (2024) MHCpLogics: an interactive machine learning-based tool for unsupervised data visualization and cluster analysis of immunopeptidomes. Brief Bioinform 25:bbae087\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVita R et al (2019) The Immune Epitope Database (IEDB): 2018 update. Nucleic Acids Res 47:D339\u0026ndash;D343\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBentzen AK et al (2018) T cell receptor fingerprinting enables in-depth characterization of the interactions governing recognition of peptide\u0026ndash;MHC complexes. Nat Biotechnol 36:1191\u0026ndash;1196\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSzeto C, Lobos CA, Nguyen AT, Gras S (2021) TCR Recognition of Peptide\u0026ndash;MHC-I: Rule Makers and Breakers. Int J Mol Sci\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChong C, Coukos G, Bassani-Sternberg M (2022) Identification of tumor antigens with immunopeptidomics. Nat Biotechnol 40:175\u0026ndash;188\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLa Gruta NL, Gras S, Daley SR, Thomas PG, Rossjohn J (2018) Understanding the drivers of MHC restriction of T cell receptors. Nat Rev Immunol 18:467\u0026ndash;478\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDash P et al (2017) Quantifiable predictive features define epitope-specific T cell receptor repertoires. Nature 547:89\u0026ndash;93\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShen Z, Reznikoff G, Dranoff G, Rock KL (1997) Cloned dendritic cells can present exogenous antigens on both MHC class I and class II molecules. J Immunol 158:2723\u0026ndash;2730\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eEscher C et al (2012) Using iRT, a normalized retention time for more targeted measurement of peptides. Proteomics 12:1111\u0026ndash;1121\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKong AT, Leprevost FV, Avtonomov DM, Mellacheruvu D, Nesvizhskii AI (2017) MSFragger: ultrafast and comprehensive peptide identification in mass spectrometry\u0026ndash;based proteomics. Nat Methods 14:513\u0026ndash;520\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTeo GC, Polasky DA, Yu F, Nesvizhskii AI (2021) Fast Deisotoping Algorithm and Its Implementation in the MSFragger Search Engine. J Proteome Res 20:498\u0026ndash;505\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eda Leprevost V (2020) Philosopher: a versatile toolkit for shotgun proteomics data analysis. Nat Methods 17:869\u0026ndash;870\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYang KL et al (2022) MSBooster: Improving Peptide Identification Rates using Deep Learning-Based Features. \u003cem\u003ebioRxiv\u003c/em\u003e, 2022.2010.2019.512904\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eK\u0026auml;ll L, Canterbury JD, Weston J, Noble WS, MacCoss MJ (2007) Semi-supervised learning for peptide identification from shotgun proteomics datasets. Nat Methods 4:923\u0026ndash;925\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMacLean B et al (2010) Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. Bioinformatics 26:966\u0026ndash;968\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePino LK et al (2020) The Skyline ecosystem: Informatics for quantitative mass spectrometry proteomics. Mass Spectrom Rev 39:229\u0026ndash;244\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDemichev V, Messner CB, Vernardis SI, Lilley KS, Ralser M (2020) DIA-NN: neural networks and interference correction enable deep proteome coverage in high throughput. Nat Methods 17:41\u0026ndash;44\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTyanova S et al (2016) The Perseus computational platform for comprehensive analysis of (prote)omics data. Nat Methods 13:731\u0026ndash;740\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShahbazy M et al (2023) Benchmarking Bioinformatics Pipelines in Data-Independent Acquisition Mass Spectrometry for Immunopeptidomics. Mol Cell Proteom 22:100515\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBallabio D, Consonni V (2013) Classification tools in chemistry. Part 1: linear models. PLS-DA. Anal Methods 5:3790\u0026ndash;3798\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReynisson B, Alvarez B, Paul S, Peters B, Nielsen M (2020) NetMHCpan-4.1 and NetMHCIIpan-4.0: improved predictions of MHC antigen presentation by concurrent motif deconvolution and integration of MS MHC eluted ligand data. Nucleic Acids Res 48:W449\u0026ndash;W454\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAndreatta M, Alvarez B, Nielsen M (2017) GibbsCluster: unsupervised clustering and alignment of peptide sequences. Nucleic Acids Res 45:W458\u0026ndash;W463\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThomsen MCF, Nielsen M (2012) Seq2Logo: a method for construction and visualization of amino acid binding motifs and sequence profiles including sequence weighting, pseudo counts and two-sided representation of amino acid enrichment and depletion. Nucleic Acids Res 40:W281\u0026ndash;W287\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHulsen T, de Vlieg J, Alkema W (2008) BioVenn \u0026ndash; a web application for the comparison and visualization of biological lists using area-proportional Venn diagrams. BMC Genomics 9:488\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHeberle H, Meirelles GV, da Silva FR, Telles GP, Minghim R (2015) InteractiVenn: a web-based tool for the analysis of sets through Venn diagrams. BMC Bioinformatics 16:169\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDeutsch EW et al (2020) The ProteomeXchange consortium in 2020: enabling \u0026lsquo;big data\u0026rsquo; approaches in proteomics. Nucleic Acids Res 48:D1145\u0026ndash;D1152\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePerez-Riverol Y et al (2022) The PRIDE database resources in 2022: a hub for mass spectrometry-based proteomics evidences. Nucleic Acids Res 50:D543\u0026ndash;D552\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-5824434/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5824434/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe major histocompatibility complex (MHC) encodes molecules that present peptides on the surface of cells to stimulate T-cell-mediated immune responses. The stability of peptide-MHC class I complexes (pMHCI) has been postulated to influence the immunogenicity of virus-derived epitopes and cancer neoepitopes. Here, we sought to investigate this further by conducting thermostability profiling of thousands of individual pMHCI, including a panel of 110 vaccinia virus (VACV) derived peptides with known CD8\u003csup\u003e+\u003c/sup\u003e T cell response profiles. The denaturation profiles of these peptides spanned thermostability (T\u003csub\u003em\u003c/sub\u003e) ranges of 41.2\u0026deg;C to 65.1\u0026deg;C, and we found that thermostability correlated with immunogenicity in VACV-infected mice. We developed two machine learning-based models from these thermostability data to predict peptide immunogenicity and demonstrate the ability of this model to distinguish immunogenic epitopes derived from an unrelated infectious pathogen, influenza A virus in mice. Using such models, we provide evidence that the thermostability of pMHCI allows for improved prediction of immunogenic CD8\u003csup\u003e+\u003c/sup\u003e T cell epitopes and conclude that this information is a valuable measurement for selecting optimal targets for T cell-mediated therapies and vaccine design.\u003c/p\u003e","manuscriptTitle":"Mass spectrometry-based thermostability profiling of virus-derived MHC peptide complexes serves as an effective predictor of immunogenicity","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-02-03 03:59:41","doi":"10.21203/rs.3.rs-5824434/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"9386c15a-63dd-439b-ba6a-484181a10df7","owner":[],"postedDate":"February 3rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":43652597,"name":"Biological sciences/Immunology/Adaptive immunity/Cellular immunity/Antigen presentation"},{"id":43652598,"name":"Biological sciences/Immunology/Vaccines/Peptide vaccines"},{"id":43652599,"name":"Biological sciences/Immunology/Antigen processing and presentation/Cellular immunity"}],"tags":[],"updatedAt":"2025-02-03T03:59:41+00:00","versionOfRecord":[],"versionCreatedAt":"2025-02-03 03:59:41","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-5824434","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5824434","identity":"rs-5824434","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.