Machine learning-based exploration of enzyme-substrate networks: SET8-mediated methyllysine and its changing impact within cancer proteomes | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Machine learning-based exploration of enzyme-substrate networks: SET8-mediated methyllysine and its changing impact within cancer proteomes Kyle Biggar, Nashira Ridgeway, Anand Chopra, Valentina Lukinovic, and 4 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3771179/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 07 Nov, 2025 Read the published version in Communications Chemistry → Version 1 posted You are reading this latest preprint version Abstract The exploration of post-translational modifications (PTMs) within the proteome is pivotal for advancing disease and cancer therapeutics. However, identifying genuine PTM sites amid numerous candidates is challenging. Integrating machine learning (ML) models with high-throughput in vitro peptide synthesis has introduced an ML-hybrid search methodology, enhancing enzyme-substrate selection prediction. In this study we have developed a ML-hybrid search methodology to better predict enzyme-substrate selection. This model achieved a 37.4% experimentally validated precision, unveiling 885 SET8 candidate methylation sites in the human proteome—marking a 19-fold accuracy increase over traditional in vitro methods. Mass spectrometry analysis confirmed the methylation status of several sites, responding positively to SET8 overexpression in mammalian cells. This approach to substrate discovery has also shed light on the changing SET8-regulated substrate network in breast cancer, revealing a predicted gain (376) and loss (62) of substrates due to missense mutations. By unraveling enzyme selection features, this approach offers transformative potential, revolutionizing enzyme-substrate discovery across diverse PTMs while capturing crucial biochemical substrate properties. Biological sciences/Biochemistry/Enzymes/Transferases Biological sciences/Computational biology and bioinformatics/Protein function predictions Biological sciences/Computational biology and bioinformatics/Machine learning Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 INTRODUCTION In the current era, in which the human genome 1 has been decoded for nearly two decades, unraveling the functional intricacies of the vast majority of human proteins remains a mystery. This challenge predominantly arises from the influence of post-translational modifications (PTMs), which are reversible chemical alterations with the potential to profoundly shape a function of a modified protein. With more than 500 distinct PTMs identified to date, the functional proteome transcends the approximately 20,000 proteins encoded by the human genome. This dynamic process involves covalently attaching functional chemical groups, such as methyl, phosphate, or acetyl, to specific amino acid (AA) residues within proteins. PTMs, which are mediated by specific modifying enzymes, can lead to substantial alterations in a protein's activity, stability, and folding 2 . Notably, proteins leverage PTMs as a cellular messaging system, responding to various external cues and stressors. This dynamic interplay between proteins and PTMs facilitates cellular adaptation to diverse environments (i.e., ensuring the maintenance of cellular homeostasis). However, the dysregulation of PTMs has been implicated in conditions ranging from cancer to inflammatory and immune disorders, highlighting their critical role in both health and disease 3,4 . Central to understanding PTMs is recognizing the interaction networks between the modifying enzymes and their corresponding PTM-modified substrates (i.e., enzyme-substrate networks). This recognition illuminates the extent to which a modifying enzyme impacts the proteome. Yet, conventional discovery methodologies face significant challenges that have historically hindered the growth of these substrate networks. Peptide arrays and mass spectrometry (MS) analysis, although valuable, have their own set of limitations and biases 5,6 . Peptide arrays offer high-throughput representation of protein segments, but fall short of capturing the full scope of PTM function as new modification sites are identified 7,8 . In contrast, MS analysis provides a comprehensive view of cellular mechanics, but often necessitates challenges to affinity or chemical enrichment steps, particularly for the discovery of lysine methylation and methyltransferase (KMT) substrates 1,9,10 . Amid these challenges, the integration of artificial intelligence has emerged as a novel approach to dissecting protein function and enzyme-substrate selection. Although generalized in silico methods have paved the way for predicting PTMs 11–13 , the application of deep learning, as seen in MusiteDeep 14 , is a transformative leap forward. However, the prediction of specific substrates for PTM-inducing enzymes remains a starkly unexplored frontier 15–17 . The innovative fusion of in vitro experiments with in silico predictions, exemplified by methodologies such as the Bayesian framework employed to characterize the substrates of protein-tyrosine phosphatase, PTP1B, using protein–protein interaction prediction, offer a promising avenue for substrate prediction 18 . However, such methods are limited by their reliance on databases that can be poorly representative of the enzyme of interest or of uncertain quality 19 . Often machine learning (ML)-based PTM prediction methods that are enzyme-specific require details about structure or the metabolic networks of the enzyme 19,20 . In this study, we transcend traditional techniques by adopting a ML-hybrid ensemble approach to enzyme-substrate identification that is generalizable across diverse enzyme classes. We show that this paradigm can successfully identify enzyme-catalyzed PTM sites for lysine methylation (SET8) and (de)acetylation (SIRT1-7) modifying enzymes. SET8, a member of the Su(var)3–9, Enhancer of zeste, Trithorax-homology (SET) family of KMT enzymes, serves as a representative model in this study. The previously identified recognition site of SET8 manifests a notable specificity towards lysines located in unfolded regions of proteins 21 , positioned ± 4 amino acids from the central lysine. This heightened specificity poses challenges in distinguishing SET8 methylation sites solely from peptide arrays, as the involvement of biophysical features beyond a simple sequential representation is anticipated in substrate recognition. Considering these complexities, SET8 emerges as a prime candidate for a systematic machine learning-based approach to substrate identification. SET8 mono-methylates histone H4 lysine 20 (H4K20), an event implicated in DNA damage repair, DNA replication, and cell cycle control 22,23 . SET8 also targets non-histone proteins, including K382 in the C-terminal protein domain of the p53 tumor suppressor, K248 of the proliferating cell nuclear antigen (PCNA), and two sites on the mitosis-associated protein Numb, K158 and K163 21,23,24 . Additional substrates have been proposed, including UHRF1-K385 and α-tubulin-K311 25,26 . SET8 has been shown to be overexpressed in bladder cancer, non-small cell and small cell lung carcinomas, pancreatic cancer, leukemia, and other diseases 23 . To broadly apply our ML-hybrid approach in delineating substrate specificity across diverse enzyme families and evaluating model performance in various enzyme classes, we investigate the substrate networks associated with the sirtuin (SIRT) family of nicotinamide adenine dinucleotide (NAD+)-dependent deacetylases. Comprising seven homologs denoted as SIRT1 to SIRT7, the SIRT family plays a pivotal role in diverse physiological processes, including inflammation, glucose and lipid metabolism, oxidative stress response, cell apoptosis, autophagy, cell proliferation, as well as cell migration and invasion 27 . This family is broadly involved in histone and non-histone deacetylation, with significant variations in subcellular localization and catalytic activity levels observed among its members. In the context of cancer, the SIRT family has been implicated in a wide spectrum of malignancies, autoimmune disorders, cardiovascular diseases, and respiratory disorders 27 . Consequently, the identification of substrates governing SIRT function holds timely and substantial relevance. Here, we improved on conventional in vitro and in silico techniques of substrate discovery by employing a novel ML approach trained on a complete peptide representation of the modified methyl-lysine and acetyl-lysine proteomes. Unlike most ML predictors of PTMs, our “hybrid” approach begins with the experimental generation of enzyme-specific training data, rather than relying purely on an online database, in which just a handful of substrates exist 28 . By chemically synthesizing a representative PTM proteome using peptide arrays, then subjecting them to in vitro enzymatic activity, we can characterize enzymatic PTM activity in a facile way. With the development of a machine learning model, augmented by generalised PTM-specific prediction 11,29 , we created ML-hybrid ensemble models unique to each enzyme that demonstrates that significantly enhanced predictive accuracy in cell models (Fig. 1 ). The application of the ML-hybrid ensemble model to a proteome bearing missense mutations associated with breast cancer uncovers potential novel pathways of SET8-mediated function in cancer cells. Furthermore, the enzyme-substrate networks produced by ML-hybrid ensemble models specific to each SIRT family member reveals novel potential pathways of conserved and enzyme-specific interaction. This pioneering ML-hybrid ensemble method not only outperforms conventional in vitro and in silico approaches, it also demonstrates broad applicability across diverse PTM-inducing enzymes. As we stand at the intersection of artificial intelligence and protein function, this novel paradigm not only sheds light on the intricate world of PTMs, but it also sets a precedent for future investigations into enzyme-substrate networks and protein function modulation. RESULTS SET8 expression and conventional substrate prediction with permutation array analysis A well-characterized and highly active SET8 193-352 construct 28 was applied to the peptide array experiments. The purity ( Figure 2A ) and methyltransferase activity of the construct was assessed for the H4K20 peptide (GGAKRHRKVLRDNIQ) ( Figure 2B ). There are several generalized methods of identifying novel substrates for PTM-inducing enzymes. To accurately assess the ability of the proposed SET8 ML-hybrid ensemble methodology to determine novel substrates, a comparable array-based permutation motif was generated to identify potential substrates 7,21 . The permutation array was created using the histone H4K20 sequence (±4 AAs, each sequence has 15 AAs in total) and was exposed to SET8 activity to identify peptide variants susceptible to methylation ( Figure 2C, Supplementary Table 1 ). Densitometry results processed by PeSA2.0 yielded the following motif: [KPGCHIVD]XH[RVIKYSAHML] K [IVT]L[RDLGI]X ( Supplementary Figure 1 ) 29 . A search of the known methyllysine proteome ( Supplementary Table 2 ) was performed with the scoring matrix, and a normalized score cutoff of 0.5, relative to unmodified H4K20 peptide (assigned a score of 1), yielded 346 hits ( Supplementary Table 3 ). Of these candidate substrate hits, just 26 peptides were validated as being methylated by SET8 in vitro with peptide array, indicating a method precision rate of 7.5% in this enriched methyllysine proteome dataset ( Figure 2D, Supplementary Table 3) . To accurately compare both the SET8 ML-hybrid ensemble model and permutation methods of substrate prediction, we next identified novel SET8 substrates from the dataset of surface-exposed lysine. This approach reveals the applicability of these approaches to the exploration of enzyme substrates beyond those that are currently known to be modified (e.g., missense cancer mutations). Using the scoring matrix from the permutation array ( Figure 2C ), a search of the surface-exposed lysine proteome was performed and yielded 15,961 sites contained within 2,424 proteins ( Supplementary Table 4 ). A randomly selected subset of these positively predicted sites (n=100) identified two positive hits, indicating a precision rate of 2% (precision represents the quantity of validated positive predictions among all positive predictions) ( Figure 2E ). Training set generation: SET8 substrates within the known methyllysine proteome To apply an effective ML model, the initial dataset must provide sufficient samples of the positive case and the negative case, ideally in equal amounts 30,31 . A randomly sampled subset (n=100) of the approximately 600,000 lysine-centric sites within the proteome tested with peptide arrays identified no sites of SET8 methylation. This observation is indicative of SET8’s highly specific recognition of substrates for methylation, and emphasizes the need for an improved approach for generating training data 21 . To efficiently enhance the likelihood of detecting positive sites, a targeted subset of the known methyllysine proteome was obtained from PhosphoSitePlus 27 . This dataset contains modified lysines, including mono-, di-, and tri-methylation. Enzymes such as KMTs often contain conserved catalytic domains and act upon methylatable histone and non-histone substrates, meaning the methyllysine proteome should contain an enriched number of substrates for SET8; in contrast, the methylated status of the full proteome is unknown 24 . Upon analysis of the SET8-exposed peptide arrays that comprise the methyllysine proteome obtained from PhosphoSitePlus 27 , the targeted subset successfully contained peptide substrates of SET8. Specifically, of the 4,593 peptides tested, 213 were deemed to be positive for SET8 methylation ( Supplementary Table 3 ). The 213 positive sites were identified across 179 proteins, indicating that several proteins harbored multiple SET8 methylation sites. To date, the commonly accepted substrates of SET8 are P53-K382, PCNA-K248, Numb-K158, and Numb-K163 24 . The 213 sites identified in vitro with the targeted search alone expands upon these four substrates. SET8 base model fitting and fine-tuning The lysine methylome for SET8 was numerically encoded with the application of MACCS keys, one-hot sequential encoding, and ProtDCal molecular descriptions ( Supplementary Table 5 ) 32,33,34,35,36 . The resulting set contains 483 features. With the stratified K-fold cross-validation method, the SET8 methyllysine proteome dataset was split into training and testing sets to effectively assess model fitting and prevent overfitting 37 . The F-score was selected as the best way to measure the predictive performance of the model on the imbalanced dataset 30,31 . A linear discriminant analysis, along with random oversampling of the positive class (i.e., sites positive for SET8 methylation), attained the highest F-score, which was 0.13 38,39 . An m-threshold analysis was performed ( Figure 2F ), as well as both precision-recall and receiver operating characteristic (ROC) curves were generated ( Figure 2G ). Metrics for the default threshold of 0.5 resulted in an F-score of 0.13, a precision of 0.085, recall of 0.24, and specificity of 0.83. The metrics further demonstrate the benefit of the F-score, defined by the harmonic mean of precision and recall, and describes our positive identification rate, rather than using specificity, which is falsely inflated by the negative identification rate. Feature importance was analyzed for the selected model, hereafter referred to as the base model. Features deemed crucial by the model for identifying positive sites of SET8 methylation were specific one-hot-encoded AA/position combinations, including tryptophan, cysteine, and tyrosine at positions +3, +4, and –6 from the central lysine, respectively. Additionally, the MACCS key corresponding to the aromatic bond between carbon and nitrogen (key 65), found in histidine, proline, and tryptophan, was deemed to be of high importance 40 . Sulphur (key 88), found in cysteine and methionine was also determined to be highly important to the model’s classification of positives 40 . Regarding the classification of negatives, or sites not methylated by SET8, once again, one-hot encoded positions played a crucial role. Specifically, regarding the central lysine (position 0), cysteine at –6, phenylalanine, and methionine at +6, and methionine at –2 scored highly for feature importance in negative classification. One MACCS key was included as well, key 132, which represents AO-CH 2 -A (where A represents any elemental symbol) and likely corresponds to the presence of aspartic acid, glutamic acid, serine, or threonine within the site 40 . SET8 ML-hybrid ensemble model construction The ability of a lysine residue to undergo methylation is a prerequisite for any newly predicted SET8 substrate. To enhance the performance of the SET8 substrate prediction, or base model, a composite or ensemble model was constructed using MethylSight, the current state-of-the-art generalized predictor of lysine methylation 11 . In the inaugural study, MethylSight identified 51 novel sites of histone methylation, and 89% of the sites were confirmed to physically exhibit methylated lysine 11 . Much like the SET8 substrate prediction model previously described, MethylSight uses ProtDCal to characterize the 15-AA-long site surrounding a central lysine 11 . Hence, it is well suited for integration with the SET8 substrate prediction model using stacked ensemble learning. As with the initial model fitting, stratified K-fold cross-validation was applied to assess the performance of each model. Two features were applied: the SET8 substrate prediction score (described above); and the MethylSight score (i.e., the likelihood of methylation). The F-score was optimized with the application of a logistic regression model and SVM SMOTE oversampling 41,42,43 . The simplicity of logistic regression was reflected in the singular hyperparameter of 100 max iterations determined from the tuning process 41 . A much-improved F-score of 0.12 was determined for the ensemble model, along with improved values of 0.25 for precision, 0.08 for recall, and 0.98 for specificity. A comparison of performance metrics with classification threshold is illustrated in Figure 2H . To optimize the performance of the ensemble model, a threshold cutoff of 0.82 was applied. Given the performance increase gained from the integration of methyllysine prediction into the ensemble model (hereafter referred to as the SET8 ML-hybrid ensemble model), the investigation proceeded with this hybrid model ( Supplementary Tables 6 and 7 ). Proteome-wide prediction of SET8 substrates Using our SET8 ML-hybrid ensemble model, experimental validation of the 2,367 predicted positive sites of SET8 methylation was completed by testing each site for in vitro methylation. Of these predictions, 885 sites permitted in vitro SET8 methyltransferase activity, representing a validated precision of 37.4%. The precision of this method is much improved over the 0% validated precision of the random search method and the 2% validated precision determined with the permutation array within the surface-exposed lysine proteome ( Figure 2I ). An analysis of the sequence composition of the predicted sites of SET8 methylation by the ML-hybrid ensemble model demonstrates that the known SET8 substrates differ from the predicted sites, with substantial variation observed in the latter ( Figure 2J ). The SET8 ML-hybrid ensemble model proved to be 100% accurate in identifying a subset (n=362) of predicted negative, lowest-scoring sites, as verified by peptide array experiments with SET8 ( Supplementary Table 8 ). Based on these findings, it is clear our SET8 ML-hybrid ensemble model improves on the traditional substrate identification approach. A total of 2,367 positive SET8 methylation sites were predicted by the SET8 ML-hybrid ensemble model within the surface-exposed lysine dataset, representing sites within 1,203 proteins. To investigate the enriched biological functions of the 1,203 proteins (i.e., predicted SET8 substrate network), clustering analysis with GO annotations was performed using the spatial analysis of functional enrichment (SAFE) approach ( Figure 3 ) 44,45 . The HuRI proteome was selected for protein mapping because of its quality, high-confidence interactions among proteins 46 . A shared theme among the enriched biological processes is involvement in cell homeostasis, regulation, and control of the cell cycle ( Figure 3A ). Given the established involvement of SET8 with these cellular events, mediated through known substrates, the possibility that SET8 might participate in such processes through the methylation of other substrates identified by our SET8 ML-hybrid ensemble model is bolstered 23,24,28,47 . Other affiliated processes include mRNA and RNA polyadenylation. Regulation through the polyadenylation of mRNA has been associated with other SET-domain-containing methyltransferases, specifically SET1 and SET2, through histone methylation 48 . The involvement of SET8 in transcription modulation may also implicate it in polyadenylation regulation; however, a direct connection has not been reported 23 . In summary, the substrates generated by the ML-hybrid ensemble model provide the potential to unveil new functional narratives for SET8 and its role(s) in disease. The efficacy of the SET8 ML-hybrid ensemble model is further demonstrated by the progression from the proteome isolated for surface-exposed lysine residues that contain 145,379 sites to the 2,367 predictions ( Supplementary Tables 9 and 10 ), which resulted in 885 in vitro validated sites, as shown in Figure 3B . Cell-based validation of SET8 substrate candidates To validate SET8-influenced cellular methyllysine events, we used parallel reaction monitoring MS to assess the in vitro SET8 substrates newly identified by our SET8 ML-hybrid ensemble model in a targeted manner. To restrict the number of methylation sites monitored with this approach, we generated an isolation list that was constrained to the primary and secondary interactors of SET8, as described by the STRING database 49 ( Supplementary Table 11 ). The 44 proteins within the network in Figure 4A contained 75 sites of predicted SET8 methylation that were verified by peptide array experiments. Of these 75 sites, it was predicted that 32 sites would create suitable digested peptides in silico; these were targeted for MS monitoring in SET8 overexpressed HCT116 cells ( Figure 4B; Supplementary Table 12 ). Of the 32 monitored sites, only nine were reliability detectable, and elevated levels of mono-methylation were observed in three (33%) of these substrates: SETD1B-K41, KAT6A-K314, and PRDM12-K269 ( Figure 4C–4E; Supplementary Figure 3 ). In conclusion, the ML-hybrid ensemble model is able to identify novel substrates of possible SET8 methylation activity, as confirmed by site-targeted MS monitoring. SET8 substrate discovery in cancer Elevated expression of SET8 is linked to a high mortality rate in patients with breast cancer ( Figure 5A ) 50,51 . However, the behavior of SET8 in cancerous cells remains unclear, and further investigation is required to uncover the functional role(s) SET8 plays in tumorigenesis. As cancer-associated mutations continue to diversify, mutation datasets serve as a valuable resource with which to elucidate the effect mutations have on protein structure and function. In the case of missense mutations, they may cause the gain or loss of methylatable lysine or make changes to neighboring residues, which then dictate the suitability of these sites for SET8 methylation 50 . To explore the possibility of gain or loss of SET8 substrates in breast cancer, missense mutations were downloaded from the COSMIC database (v.96) and applied to the human proteome. Of the initial mutations, 9,438 either occurred within seven AAs of a lysine residue (e.g., any residue) or resulted in the gain or loss of an individual lysine, directly impacting the creation or loss of a potential methylation site ( Supplementary Table 13 ). The corresponding unmutated sites, except sites in which a lysine did not previously exist, were also assembled. Application of the SET8 ML-hybrid ensemble model to normal and breast cancer datasets predicted that most of the mutations (94.6%) would not affect SET8’s methylation behavior toward the site, likely because their structure is not changed dramatically by a single AA mutation. In contrast, 4.0% (376) of mutations resulted in a predicted gain of SET8 methylation, and 0.7% (62) resulted in a loss ( Figure 5B ). Of the 4.0% of sites predicted to gain SET8 substrate status, 46.8% (176) were the result of the mutation introducing a new lysine that is itself predicted to be methylated by SET8 ( Figure 5C ). MCODE clustering analysis of the total set of mutations revealed a directed subset of predicted substrate interactions that were highly interconnected with SET8 ( Supplementary Figure 4 ). Mutations within the subset that resulted in a gain of predicted SET8 methylation were investigated for involvement in pathways implicated in breast cancer. Of particular interest was XPF (encoded by ERCC4), a protein associated with the vital cellular process of DNA damage repair 52 . Specifically, the XPF-S352A mutation led to the prediction of a new SET8 methylation site at XPF-K350. As detailed in Figure 5D , XPF is directly involved in DNA damage repair, including nucleotide excision, double-strand break, and interstrand cross-link repair pathways 53 . Gap filling is then proceeded by PCNA, a known substrate of SET8 24 , further implicating SET8 in the NER pathway 54 . Finally, ligation is performed with DNA ligase and the NER pathway for DNA damage is complete 53 . Interestingly, SET8 has been implicated in DNA repair previously, specifically in 53BPI/BRCA1 double-stranded DNA repair through histone H4K20 mono-methylation 54 . ERCC4 gene mutations have also been determined to affect XPF function within the NER pathway 52 . Beyond breast cancer, our SET8 ML-hybrid ensemble model was also applied to missense mutations present in pancreatic cancer (COSMIC database, v.96) ( Supplementary Figure 5A and 5B; Supplementary Table 14 ). DISCUSSION The introduction of our ML-hybrid ensemble model marks a significant leap forward in the identification of substrates for specific PTM-inducing enzymes. This methodology is successfully generalized across multiple enzyme classes and is inevitably extendable to a broader spectrum of PTMs, including phosphorylation, sumoylation, and ubiquitination. The power of this approach lies in its ability to efficiently characterize PTM-specific subsets of the proteome, a task made feasible by the manageable size; only a few thousand sites can be readily profiled in parallel using peptide arrays (Fig. 1 ). Moreover, this approach ensures greater success in positively identifying new enzyme substrates than established array-based methods 5,57 . Although many predictive models can identify sites of a particular PTM, few possess the capacity to pinpoint the enzyme responsible 11,15,58 . In contrast, our ML-hybrid ensemble methodology is highly specific to the enzyme under investigation and reveals the predicted PTM site. Notably, the ensemble models for both SET8 and SIRT2 demonstrate exceptional performance when experimentally validated in cells overexpressing the target enzyme, a departure from many PTM prediction tools 58,59 . The ML-hybrid ensemble model’s experimentally validated precision of 37–43% stands in stark contrast to the 2.0% validated precision achieved with a traditional permutation array in the in vitro search for new substrates within the human proteome. Further the high precision of our models enables its facile use in directing substrate identification studies using more targeted, and intensive validation strategies. The limited amount of labeled PTM-specific training data imposes restrictions on ML model complexity, but significantly reduces computational load. Deep learning algorithms, including neural networks, are precluded from predicting enzyme-specific substrates because of the dataset’s small size. However, this constraint aligns with the notion that deep learning in PTM prediction is most effective in a more generalized context, as seen in MusiteDeep 58 . Most potential substrates for an enzyme within the proteome are, in fact, not substrates, as a result there is an inherent imbalance that narrows the range of applicable models. Although the abbreviated peptide sequences may not fully represent a full-length enzyme-substrate interaction, this limitation is not exclusive to our approach; conventional permutation array experiments employ similar methodologies when validating results 21 . The unique advantage of our ML-hybrid ensemble method lies in the extensive number of predictions (more than 2,300 for SET8), yielding a substantial number of validated predictions (885 for SET8), thereby expanding the list of possible substrates for MS-monitoring methods compared with conventional methods 21 . The value of any predictive model of substrate selection lies within its ability to prioritize the exploration of new datasets and reveal novel insights. The targeted MS monitoring of predicted sites of SET8 methylation in cells highlighted the further elevation of SETD1B-K41, KAT6A-K314, and PRDM12-K269 mono-methylation levels in SET8 overexpressed HCT116 cells (Fig. 3 ). These findings suggest the potential involvement of SET8 in gene activation and regulatory pathways. The change in SETD1B-K41me1 levels in SET8 overexpressed cells may link SET8 to the COMPASS complex, of which SETD1B is a part 60 . KAT6A, commonly localized to CpG islands, may be influenced by SET8 mono-methylation at K269 61 , potentially impacting gene regulation. Furthermore, the association of both KAT6A and SETD1B with CXXC1 (CpG-binding protein) implicates SET8 in gene activation 61,62 . The regulatory domain binding factor PRDM12 was observed to demonstrate elevated K269me1 levels with SET8 overexpression 63 . With additional investigation, this may suggest SET8’s involvement in gene activation and regulation, potentially through the methylation of PRDM12-K269 63 . To help reveal potential insights into the role(s) SET8 plays in breast cancer, our SET8 ML-hybrid ensemble model was applied to explore the creation and loss of potential substrates to help annotate a predicted breast-cancer-specific SET8 enzyme-substrate network. When applied to cancer mutation datasets, the ML-hybrid ensemble model has specific advantages. The sensitivity of the model is evident when comparing oncogenic and healthy proteome scores, which provide a unique perspective on oncogenic mutations. For example, the predicted gain of SET8 substrate methylation due to a missense mutation sheds light on potential proteins of interest, particularly the association of SET8 with breast cancer within the NER pathway through XPF (Fig. 4 ) 54 . This implicates SET8 in the DNA repair process, employing PCNA, another substrate of SET8, for gap filling 54 . The XPF-S352A mutation and subsequent SET8 methylation at XPF-K350 could be hypothesized to affect the NER pathway, such that DNA damage repair may not be completed. Interestingly, the K350 site falls within the XPF-SLX4 interaction region 64 , an interaction associated with the Fanconi anemia DNA repair pathway and defects in SLX4 function have been linked to breast cancer 65 . The generalizability of the ML-hybrid ensemble approach is demonstrated by the sirtuin family investigation. The rates of identified in vitro peptide substrates for each SIRT range from as low as 2.59% positive (97.4% negative), to 38.4% positive (61.6% negative). With the balancing methodology and the careful tuning of models, each SIRT-based model produced metrics as impressive, if not more so, than the SET8 data. This is exemplified by the MS-validated precision of the SIRT2 model of 43.0%. The conserved and enzyme-specific substrates within each SIRT network further provides a unique insight into potential pathways of activity (Fig. 5 ). Interestingly, SIRT2 and SIRT3 share almost twice the substrates as any other combination of SIRT. This finding is supported by prior investigation that uncovered that the pair may compensate for each other when the other is deficient 66 . The second largest overlap in predicted substrates occurs between SIRT5 and SIRT7. As the most recently discovered member, SIRT7 is underrepresented in literature, however some associations between SIRT5 and SIRT7 exist. The two have been implicated within disease states such as cardiac hypertrophy and inflammatory bowel diseases, as well as the regulation of the NF-κB inflammation signalling pathway 27 . The predicted substrates of the ML-hybrid ensemble model for SIRT7 may prove to be a key to direct future investigations to uncover SIRT7’s cellular function(s). In conclusion, our generalizable ML-hybrid ensemble approach represents a significant advance in the precise identification of enzyme-substrate networks and the features of substrate enzymes that influence selection. The methodology’s proven versatility offers promise for a wide array of PTM-inducing enzymes, thereby illuminating previously uncharted territories of protein function modulation. Although some limitations exist, including dataset size and an unavoidable class imbalance, our approach applies techniques to overcome such constraints to offer a powerful tool for exploring the intricate world of PTMs. The verification of predicted SET8 and SIRT2 substrates in cells underscores the model’s robustness, outperforming traditional methods and paving the way for a more holistic understanding of enzyme-substrate selection and involvement in cellular processes. The application of this model to cancer datasets provides a promising avenue with which to uncover critical insights into oncogenic mutations and the role of enzymes within these modified proteomes. Moreover, our ML-hybrid ensemble approach presents the capability to uncover the enzyme-substrate network of an entire family of enzymes to accurately predict the extent of an enzyme family’s functional role within the cell. With its potential for broader application and its capacity to shed light on unexplored facets of cellular regulation, our ML-hybrid ensemble model stands poised to have a profound impact on the study of enzyme-substrate networks and protein function modulation. METHODS Peptide synthesis Peptide SPOT arrays were synthesized to commercial aminated cellulose membranes (Intavis Inc.) with standard Fmoc (N-(9-fluorenyl)methoxycarbonyl) chemistry automatically using a Multipep synthesizer (Intavis Inc.) 67 . Resulting arrays contained approximately 2 nmol peptide per SPOT, separated from the membrane surface by a flexible C-terminal 6-aminohexanoic acid linker. Post-treatment involved the cleavage of the protective groups from the side-chains with an acidic solution (51% water, 47.5% trifluoroacetic acid, and 1.5% tri-isopropylsilane) 67,68 . Arrays were then washed in ethanol and dried for future use. Dried arrays were stored desiccated at 4°C. Protein expression and purification A bacterial expression construct of human SET8 encoding the catalytic site residues 191–352 were cloned into the parallel expression vector pHIS2 using BamHI and XhoI restriction sites 48 . The construct was transformed in BL21 DE3 Escherichia coli , and after propagation were induced with 0.3 mM isopropyl β-D-1-thiogalacttopyranoside 48 . Following protein extraction, affinity column purification of the soluble lysate was completed with 500 µL HisPur™ Ni-NTA Resin (ThermoFisher Scientific, Cat# 88221). Fractions were eluted using P500 buffer (50 mM NaHPO 4 (pH 7), 500 mM NaCl, 10% glycerol, 0.05% TritonX-100, 1 mM DTT, 500 mM Imidazole) and then dialyzed into a storage buffer (20 mM tris pH 7.5-8, 200 mM NaCl, 10% glycerol, 1 mM DTT). Protein purity was assessed with a 12% SDS PAGE gel ( Supplementary Fig. 1A ) stained with Coomassie Brilliant Blue G-250 stain. Concentration was determined with a Bradford Assay 69 . Prior to storage, activity was confirmed using the Methyltransferase-Glo Assay to assess in vitro activity with H4K20 peptide (GGAKRHRKVLRDNIQ) ( Supplementary Fig. 1B ), according to the manufacturer’s specifications (Promega, Cat# V7601). Recombinant SET8 protein was then snap frozen stored at -80°C. Training dataset generation The annotated Human lysine methylome was obtained from PhosphoSitePlus (accessed February 5th, 2020). This dataset contained sequence information for each lysine methylation site represented as 15 residue peptides, centered on the methylated lysine. (i.e., position 8) 28,70 . The data was then cleaned using Python 3 and the Pandas package for efficient isolation of human protein sites of lysine methylation 71,72 . Briefly, sites that were within 7 Aas from the beginning or end of the sequence were padded with alanine residues, and any duplicate sequences were removed (i.e., methylation events that occur within a conserved 15 AA sequence, but among unique proteins). The resulting dataset yielded a total of 4,593 Human lysine methylation sites ( Supplementary Table 2 ). Peptide arrays were then methylated using 1 µM SET8 191− 325 with 5 µCi/mL S-[methyl- 3 H]-Adenosyl-L-methionine (SAM) in methylation buffer (50 mM Tris pH 8.5, 2 mM MgCl2, 10 µM DTT) overnight at room temperature 73 . After methylation, arrays were washed 6x 3 min in buffer (100 mM NH4HCO3, 1% SDS) and then a 7% 2,5-Diphenyloxazole solution in ethanol was sprayed generously over the array and left to air dry. This procedure was repeated a total of three times. Dried peptide arrays were then exposed to intensifying screens (Dupont, Cronex Lighting Plus) at -80°C for two weeks and imaged using a Typhoon™ FLA 7000 IP phosphorimager (General Electric). Peptide methylation was determined by densitometry using the Protein Array Analyzer (v.1.1.c) toolset for Image J (v.1.53) was applied to obtain SPOT densitometry from the array images. The peptide substrate dataset for the sirtuin family was adapted from Rauh et al., 2013 5 . Peptides which produced a signal ratio score of 1 or greater were classified as positive, and positions with masked AA residues were replaced with alanine ( Supplementary Table 16 ). Biochemical Feature Generation Peptide libraries were represented numerically through one-hot encoding, a common method of vectorizing peptides 34,35 . A matrix describing AA position and identity with a 1 or 0 is applied to the sequence, resulting in 300 features for SET8, and 260 for the sirtuins (Eq. 1). \({x}_{AA\text{, }pos}=\left\{\begin{array}{c}1 if seq\left[pos\right]==AA\\ 0 if else \end{array}\right. \text{, for }pos=\left[1,N\right]\text{, }AA=\{A, C, D, \dots , W\}\) [1] To encode molecular structure, the Molecular Access System (MACCS) Keys of 166 binary fingerprints were applied 36 . Created to describe molecular structure, MACCS Keys include predefined atom symbols, bond types, and atom properties 36 . The RDKit package for Python ( www.rdkit.org ) was applied to generate MACCS Keys from the site sequence, representing the molecular structure for the AAs present. Aggregate molecularly descriptive features were generated for the peptide sequence through ProtDCal (v4.5) 38 . The 17 metrics selected included molecular weight, hydrophobicity, isoelectric point, free energy, and Levitt’s Probability for various protein conformation amongst other descriptive sequence-based metrics 38 ( Supplementary Table 5 ). Combined, the sequential information provided by one-hot encoding, along with the unique molecular elements represented by the MACCS keys and the molecularly descriptive features determined by ProtDCal resulted in 483 features for SET8, and 443 for the sirtuins. Machine Learning – Model Fitting All procedures related to model fitting and data balancing were completed in Python 3 using the Scikit-Learn and Imbalanced-learn packages 71,74,75 . Pandas and Numpy were applied for data handling and storage, and plots were generated with Matplotlib 72,76,77 . Class imbalance was present in all samples, although it ranged in degree for each enzyme studied. The SET8-methylated lysine peptide array data provided a class imbalance of 213 positives within the 4,593 sites tested for SET8 activity. Imbalance, along with the relatively small size of the training data, was taken into careful consideration when selecting ML models 33 . For proper comparison to a baseline model, the dummy classifier which simply randomly classifies data input was first fit to the data. Applying the F-score as the primary selection criteria and a model complexity considerate of the size and imbalance of our dataset as the secondary selection criteria, the selected models were linear discriminant analysis (LDA) and the decision tree classifier 40 . Class imbalance in which the positive case composed less than 30% of the training dataset was addressed with data balancing or sampling methods, which was the case for all datasets except SIRT5 (38.4% positive, 61.6% negative). Both over and undersampling techniques were tested to improve F-score with our models, and the best performing data balancing approach was selected ( Supplementary Table 17 ). In these approaches, a selection of the positive class is replicated within the dataset, increasing the overall percentage of positive values 74 . Alternatively, negative values are removed to reduce the overall percentage of the negative case within the training dataset. Machine Learning – Cross-validation The fit of each model was assessed with repeated stratified K-fold cross-validation 39 . Cross-validation involves the splitting of the labelled training set and related features into n-folds. In this case, the label indicates whether our datapoint (i.e., modification site) is methylated. Due to the imbalanced nature of our data, stratification was applied to ensure the same percentage of positive and negative labels are represented within each fold. Many metrics may be used to assess classification performance; however, the nature of the data must be taken into careful consideration. In a dataset with fewer positive instances than negatives, the preferred metric is the F-score, as it accounts for true positive and false positive counts, as well as false negatives 33,39 (Eq. 2). To assess the fit of each model, stratified 10-fold cross-validation was applied to the lysine methylome with 3 repeats and using F-score. \(F=\frac{2TP}{2TP+FN+FP}\) [2] The best-scoring models were determined to be the linear discriminant analysis (LDA) method and the decision tree classification method. The LDA method determines a linear combination of features that optimally separates the two classes; “modified by our enzyme” and “ not modified by our enzyme”. The LDA method has been applied to solve biochemical problems in the past, including the prediction of protein function and tertiary structure from sequence 78,79 . One investigation employed LDA to predict the site of protein sumoylation, highlighting the benefit of LDA in the prediction of PTMs 80 . The decision tree classifier composes a model through the creation of nodes representing a test on the features provided. With each node, a branch descends to eventually point towards a classification of the input to one of the two classes as defined in the LDA method. Decision tree classification has been applied to the classification of enzymatic and non-enzymatic metal binding sites, as well as the familial classification of an unknown protein 81,82 . Machine Learning – Hyperparameter Tuning The next step of model fitting involves the tuning of the model’s hyperparameters. These model-specific parameters control the method in which the model is applied to the data. To effectively search for the optimal hyperparameters, a grid search was applied 83 . Each combination of hyperparameter was tested with the repeated stratified K-fold cross-validation method, and the combination resulting in the highest F-score was returned 39,83 . The balancing methods are also tuned through hyperparameter optimization to better the model’s performance. Machine Learning – Optimizing Decision Threshold Adjustment of the model’s threshold is required to fine-tune the performance, particularly so with imbalanced training data. When applied to new data, the trained model outputs a score representative of the probability that a data point is a member of the positive class. By default, a decision threshold of 0.5 is applied the probability; however, this decision threshold can be tuned to trade off recall and precision 33 . Recall assesses the model’s identification of true positives; precision measures the proportion of positive predictions that are correct; and accuracy over truly negative instances are quantified with specificity 33 . For significantly imbalanced datasets, the best threshold for the model was determined by identifying the point at which the F-metric was maximized 11,33,39 . Otherwise, a decision threshold of 0.5 was maintained. Machine Learning – Ensemble Learning To improve the overall performance of the models, ensemble learning algorithms were explored. In ensemble algorithms, a secondary ML predictor is applied in parallel. The lysine methylation predictor MethylSight was implemented as the secondary method for the SET8 model, as it generally predicts sites of lysine methylation; a dependent variable for a lysine to be a SET8 substrate 11 . MethylSight uses support vector ML, and was validated experimentally to predict novel sites of methylation, including the identification of novel KDM5B substrate H2B-K43me2 11 . For the sirtuin family the generalized deep learning predictor MusiteDeep was applied as the secondary predictor. The N6-acetyl-lysine predictive function of MusiteDeep was employed to generally identify sites of lysine acetylation, much like MethylSight for SET8. MusiteDeep outperformed other representative ML and deep learning algorithms for N6-acetyl-lysine prediction 29 . The ensemble methods explored included hard voting, soft voting, and stacking. Each method takes in the probability score output by either predictor that a site is methylated. The hard voting method classifies through most votes, or scores above our specified positive cutoff thresholds for the positive class by either model 84 . In soft voting, classification is determined through the mean of probabilities by either model 84 . Finally, stacking utilizes the scores of each model as features of a third model, along with dataset balancing as previously outlined 84 . Each method was tested, and F-scores were generated with repeated stratified K-fold cross-validation. For SET8, a holdout set from the MethylSight investigation was employed to provide an unbiased final validation of each ensemble method 11 . Experimental Dataset Generation and Evaluation The human proteome was obtained from UniProt, and analyzed with NetSurfP2.0 to generate values for relevant surface accessibility 85,86 . Lysine residues with a relevant surface accessibility value of over 0.2 were selected, and the surrounding (±7 and ±6 from central lysine) AAs were isolated from the full protein sequence to form the experimental datasets. As with the generation of the lysine methylome, sites less than 7 or 6 AAs from the beginning or end of a protein sequence were padded with alanine. Additionally, any duplicated sequences or sequences found within the training dataset were dropped. The SET8 ensemble learning model was applied to the resulting set and resulted in 2,367 novel positive predictions of new lysine methylation sites that are also SET8 substrates ( Supplementary Table 9 ). These values were mapped to the human binary protein interactome (HuRI), and abundant biological processes were showcased with the Spatial Analysis of Functional Enrichment (SAFE) package for Python 3 45,47,71 . Additionally, all predicted positive sites were tested individually via peptide SPOT array experiments for SET8 in vitro methylation, as previously described. The resulting seven sirtuin models were applied and predicted various novel sites of deacetylation ( Supplementary Table 18 ). The commonality of predictions between each SIRT was illustrated through Circos and UpSet plots as generated by the pyCircos and UpSetPlot packages for Python 3 71,87,88 . Cell line transfection Cells were cultured in Dulbecco’s modified Eagle’s medium (DMEM, Gibco Cat# 11965092) supplemented with 10% heat-inactivated FBS and 1% penicillin/streptomycin at 37°C, 5% CO 2 . Cell lines were regularly tested for Mycoplasma pneumoniae using polymerase chain reaction (PCR) assay. For SET8 overexpression experiments, cells were transfected with 10 µg pcDNA3.1 HA-SET8 using jetOptimus transfection reagent (Polyplus, Cat# 101000025), following manufacturer instruction. The cells were harvested 48h after transfection and cell pellets were snap frozen and stored at -20°C until further use. Cell pellets were lysed in extraction buffer containing 50 mM tris (pH 8), 150 mM NaCl, 1 mM EDTA, 10% glycerol, 0.5% NP-40, and protease inhibitors (1 mM PMSF, 10 µM E-64, 1 µM Pepstatin, 1 µM Leupeptin, and 1 mM sodium orthovanadate). Protein concentration was determined by Bradford assay and samples stored at -80°C until use. Conditions of SIRT2 overexpression in HCT116 cells are described previously 56 . Mass Spectrometry Protein concentration was determined by Pierce BCA Protein Assay (Thermo Fisher Scientific) and 500 µg of soluble lysate was digested with Arg-C (Roche, 11370529001). Cell lysates were first diluted in Arg-C digestion buffer (100 mM Tris-HCl, 10 mM CaCl 2 , pH 7.6) followed by standard reduction and alkylation at room temperature. Briefly, proteins were first reduced with 3 mM TCEP for 45 minutes, alkylated with 15 mM iodoacetamide (IAA) for 60 minutes, followed by quenching of any unused IAA by incubation with 20 mM DTT for 45 minutes. Next, activation solution and Arg-C were added to a final concentration of 1X and a 1:200 protein-to-protease ratio, respectively. The digestion occurred overnight at 37 o C with end-to-end rotation. Resultant peptides were desalted C18 Spin Columns (Thermo Fisher Scientific, 89870) and C18-ZipTip (Millipore Sigma, ZTC18S096), dried by SpeedVac, and resuspended in 20 µL of MS-grade water + 0.1% formic acid. To assess mono-methylation levels of newly identified in vitro SETD8 substrates, parallel reaction monitoring (PRM)-MS was performed using a Q-Exactive Plus hybrid quadrupole-orbitrap mass spectrometer, at the John L. Holmes Mass Spectrometry Facility at the University of Ottawa, as previously described 89 . PRM-MS scanning was guided by an isolation list built in Skyline Software 90 . The isolation list was generated by first identifying the primary interactors of SET8, isolated from the STRING human interactome dataset using Python 3 and the Pandas package 50,71,72 . The interactors of the primary interactors of SET8, or secondary interactors were also included in this list. Proteins which contained the predicted and validated SET8 in vitro methylation sites were selected from within the primary and secondary interactors of SET8, and the resulting network was mapped with the NetworkX package for Python 91 . The verified sites contained within this subnetwork were then used to generate an isolation list for targeted mass spectrometry analysis. To quantify relative mono-methylation levels of detectable target peptides, Skyline software was used to design an isolation list ( Supplementary Table 15 ) and was used to derive total peak areas for modified target substrates as well as other unmodified peptides along the protein sequence. The total peak area of each modification site of interest was divided by that of all other reliably detected unmodified peptides within a given parental protein. The MS-verified sites for SIRT2 were obtained from the dataset provided in Zhang et. al., 2022 56 . Application to the Cancer Proteome Targeted screen mutant repositories were obtained for breast cancer from v96 of the COSMIC database (cancer.sanger.ac.uk) 92 . In Python 3, the data were cleaned to isolate missense mutations, which were then applied to the full human proteome obtained from UniProt 71,85 . Mutations which occur within ±7 AAs from either ( 1 ) a lysine, or ( 2 ) a position mutated to a lysine, were accepted within the oncoproteome set. An accompanying dataset for the healthy proteome was generated. The ML-hybrid ensemble model for SET8 methylation prediction was applied to both the healthy and cancerous dataset. The resulting scores were compared, yielding several potential states for mutations; ( 1 ) mutations predicted to have a null effect on SET8 methylation (i.e., no change from healthy population), ( 2 ) mutations predicted to remove SET8 methylation activity (i.e., loss of SET8 substrate from healthy population), and ( 3 ) mutations predicted to induce SET8 methylation (i.e., gain of SET8 from health population). The human interactome from the STRING database was applied to the mutated proteins and the probability scores from the SET8 ML-hybrid ensemble model were scaled to resemble STRING scores and included within the interaction list as well. The network was mapped with Cytoscape (v3.9.1), and molecular subcomplexes were isolated using the Molecular Complex Detection cluster generator from the clusterMaker (v2.0) toolset 93–95 . The subcomplex which contained SET8 was further investigated for proteins implicated in pathways associated with the cancer type, and resulting pathways were represented in Adobe Illustrator. Declarations DATA AVAILABILITY The data conveying all results reported may be readily obtained within the github repository https://github.com/nashirag/ML-Hybrid_Ensemble_Method . Source data for all figures is accessible within the same repository, including all SPOT peptide array densitometry figures are available. All mass spectrometry data is available through the PeptideAtlas repository (project ID: PASS05848). CODE AVAILABILITY The Python code implemented to produce all results may be accessed at https://github.com/nashirag/ML-Hybrid_Ensemble_Method_SET8. References Brandi, J., Noberini, R., Bonaldi, T. & Cecconi, D. Advances in enrichment methods for mass spectrometry-based proteomics analysis of post-translational modifications. J. Chromatogr. A 1678 , 463352 (2022). Deribe, Y. L., Pawson, T. & Dikic, I. Post-translational modifications in signal integration. Nat. Struct. Mol. Biol. 17 , 666–672 (2010). Liu, J., Qian, C. & Cao, X. Post-Translational Modification Control of Innate Immunity. Immunity 45 , 15–30 (2016). Qian, M. et al. Targeting post-translational modification of transcription factors as cancer therapy. Drug Discov. Today 25 , 1502–1512 (2020). Rauh, D. et al. An acetylome peptide microarray reveals specificities and deacetylation substrates for all human sirtuin isoforms. Nat. Commun. 4 , 2327 (2013). Merbl, Y. & Kirschner, M. W. Large-scale detection of ubiquitination substrates using cell extracts and protein microarrays. Proc. Natl. Acad. Sci. 106 , 2543–2548 (2009). Moore, K. E. & Gozani, O. An unexpected journey: Lysine methylation across the proteome. Biochim. Biophys. Acta BBA - Gene Regul. Mech. 1839 , 1395–1403 (2014). Polo, S. et al. A single motif responsible for ubiquitin recognition and monoubiquitination in endocytic proteins. Nature 416 , 451–455 (2002). Mitchell, C. J. et al. Unbiased identification of substrates of protein tyrosine phosphatase ptp‐3 in C. elegans. Mol. Oncol. 10 , 910–920 (2016). Yu-Ying, Y., Markus, G. & Howard, H. C. Identification of lysine acetyltransferase p300 substrates using 4-pentynoyl-coenzyme A and bioorthogonal proteomics. Bioorg. Med. Chem. Lett. 21 , 4976–4979 (2011). Biggar, K. K. et al. Proteome-wide Prediction of Lysine Methylation Leads to Identification of H2BK43 Methylation and Outlines the Potential Methyllysine Proteome. Cell Rep. 32 , 107896 (2020). Jamal, S., Ali, W., Nagpal, P., Grover, A. & Grover, S. Predicting phosphorylation sites using machine learning by integrating the sequence, structure, and functional information of proteins. J. Transl. Med. 19 , 218 (2021). Kiemer, L., Bendtsen, J. D. & Blom, N. NetAcet: prediction of N-terminal acetylation sites. Bioinformatics 21 , 1269–1270 (2005). Neely, B. A. et al. Toward an Integrated Machine Learning Model of a Proteomics Experiment. J. Proteome Res. 22 , 681–696 (2023). Deng, W. et al. GPS-PAIL: prediction of lysine acetyltransferase-specific modification sites from protein sequences. Sci. Rep. 6 , 39787 (2016). Wu, Z., Lu, M. & Li, T. Prediction of substrate sites for protein phosphatases 1B, SHP-1, and SHP-2 based on sequence features. Amino Acids 46 , 1919–1928 (2014). Wang, X. et al. UbiBrowser 2.0: a comprehensive resource for proteome-wide known and predicted ubiquitin ligase/deubiquitinase–substrate interactions in eukaryotic species. Nucleic Acids Res. 50 , D719–D728 (2022). Ferrari, E. et al. Identification of New Substrates of the Protein-tyrosine Phosphatase PTP1B by Bayesian Integration of Proteome Evidence. J. Biol. Chem. 286 , 4173–4185 (2011). Smith, K., Rhoads, N. & Chandrasekaran, S. Protocol for CAROM: A machine learning tool to predict post-translational regulation from metabolic signatures. STAR Protoc. 3 , 101799 (2022). Lanouette, S. et al. Discovery of Substrates for a SET Domain Lysine Methyltransferase Predicted by Multistate Computational Protein Design. Structure 23 , 206–215 (2015). Kudithipudi, S., Dhayalan, A., Kebede, A. F. & Jeltsch, A. The SET8 H4K20 protein lysine methyltransferase has a long recognition sequence covering seven amino acid residues. Biochimie 94 , 2212–2218 (2012). Fang, J. et al. Purification and Functional Characterization of SET8, a Nucleosomal Histone H4-Lysine 20-Specific Methyltransferase. Curr. Biol. 12 , 1086–1099 (2002). Milite, C. et al. The emerging role of lysine methyltransferase SETD8 in human diseases. Clin. Epigenetics 8 , 102 (2016). Biggar, K. K., Wang, Z. & Li, S. S.-C. SnapShot: Lysine Methylation beyond Histones. Mol. Cell 68 , 1016-1016.e1 (2017). Zhang, H. et al. SET8 prevents excessive DNA methylation by methylation-mediated degradation of UHRF1 and DNMT1. Nucleic Acids Res. 47 , 9053–9068 (2019). Chin, H. G. et al. The microtubule-associated histone methyltransferase SET8, facilitated by transcription factor LSF, methylates α-tubulin. J. Biol. Chem. 295 , 4748–4759 (2020). Hornbeck, P. V. et al. PhosphoSitePlus, 2014: mutations, PTMs and recalibrations. Nucleic Acids Res. 43 , D512–D520 (2015). Yin, Y. et al. SET8 recognizes the sequence RHRK20VLRDN within the N terminus of histone H4 and mono-methylates lysine 20. J. Biol. Chem. 280 , 30025–30031 (2005). Topcu, E., Ridgeway, N. H. & Biggar, K. K. PeSA 2.0: A software tool for peptide specificity analysis implementing positive and negative motifs and motif-based peptide scoring. Comput. Biol. Chem. 101 , 107753 (2022). Burkov, A. The hundred-page machine learning book . (Andriy Burkov, 2019). Brownlee, J. Imbalanced Classification with Python: Choose Better Metrics, Balance Skewed Classes, and Apply Cost-Sensitive Learning . (Machine Learning Mastery, 2021). Yang, K. K., Wu, Z. & Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. Nat. Methods 16 , 687–694 (2019). Erjavac, I., Kalafatovic, D. & Mauša, G. Coupled encoding methods for antimicrobial peptide prediction: How sensitive is a highly accurate model? Artif. Intell. Life Sci. 2 , 100034 (2022). Durant, J. L., Leland, B. A., Henry, D. R. & Nourse, J. G. Reoptimization of MDL Keys for Use in Drug Discovery. J. Chem. Inf. Comput. Sci. 42 , 1273–1280 (2002). Ruiz-Blanco, Y. B., Paz, W., Green, J. & Marrero-Ponce, Y. ProtDCal: A program to compute general-purpose-numerical descriptors for sequences and 3D-structures of proteins. BMC Bioinformatics 16 , 162 (2015). Romero‐Molina, S., Ruiz‐Blanco, Y. B., Green, J. R. & Sanchez‐Garcia, E. ProtDCal‐Suite: A web server for the numerical codification and functional analysis of proteins. Protein Sci. pro.3673 (2019) doi:10.1002/pro.3673. Szeghalmy, S. & Fazekas, A. A Comparative Study of the Use of Stratified Cross-Validation and Distribution-Balanced Stratified Cross-Validation in Imbalanced Learning. Sensors 23 , 2333 (2023). Izenman, A. J. Linear Discriminant Analysis. in Modern Multivariate Statistical Techniques 237–280 (Springer New York, 2013). doi:10.1007/978-0-387-78189-1_8. Kamalov, F., Leung, H.-H. & Cherukuri, A. K. Keep it simple: random oversampling for imbalanced data. in 2023 Advances in Science and Engineering Technology International Conferences (ASET) 1–4 (IEEE, 2023). doi:10.1109/ASET56582.2023.10180891. Cereto-Massagué, A. et al. Molecular fingerprint similarity search in virtual screening. Methods 71 , 58–63 (2015). Wright, R. E. Logistic regression. Read. Underst. Multivar. Stat. 217–244 (1995). Chawla, N. V., Bowyer, K. W., Hall, L. O. & Kegelmeyer, W. P. SMOTE: Synthetic Minority Over-sampling Technique. J. Artif. Intell. Res. 16 , 321–357 (2002). Nguyen, H. M., Cooper, E. W. & Kamei, K. Borderline over-sampling for imbalanced data classification. Int J Knowl Eng Soft Data Paradig. 3 , 4–21 (2009). Baryshnikova, A. Spatial Analysis of Functional Enrichment (SAFE) in Large Biological Networks. in Computational Cell Biology (eds. von Stechow, L. & Santos Delgado, A.) vol. 1819 249–268 (Springer New York, 2018). The Gene Ontology Consortium et al. The Gene Ontology resource: enriching a GOld mine. Nucleic Acids Res. 49 , D325–D334 (2021). Luck, K. et al. A reference map of the human binary protein interactome. Nature 580 , 402–408 (2020). Couture, J.-F., Collazo, E., Brunzelle, J. S. & Trievel, R. C. Structural and functional analysis of SET8, a histone H4 Lys-20 methyltransferase. Genes Dev. 19 , 1455–1465 (2005). Kaczmarek Michaels, K., Mohd Mostafa, S., Ruiz Capella, J. & Moore, C. L. Regulation of alternative polyadenylation in the yeast Saccharomyces cerevisiae by histone H3K4 and H3K36 methyltransferases. Nucleic Acids Res. 48 , 5407–5425 (2020). Szklarczyk, D. et al. STRING v10: protein–protein interaction networks, integrated over the tree of life. Nucleic Acids Res. 43 , D447–D452 (2015). Liu, B. et al. A functional single nucleotide polymorphism of SET8 is prognostic for breast cancer. Oncotarget 7 , 34277–34287 (2016). Yang, C., Wang, K., Zhou, Y. & Zhang, S.-L. Histone lysine methyltransferase SET8 is a novel therapeutic target for cancer treatment. Drug Discov. Today 26 , 2423–2430 (2021). Bogliolo, M. et al. Mutations in ERCC4, Encoding the DNA-Repair Endonuclease XPF, Cause Fanconi Anemia. Am. J. Hum. Genet. 92 , 800–806 (2013). Faridounnia, M., Folkers, G. & Boelens, R. Function and Interactions of ERCC1-XPF in DNA Damage Response. Molecules 23 , 3205 (2018). Xu, L. et al. Roles for the methyltransferase SETD8 in DNA damage repair. Clin. Epigenetics 14 , 34 (2022). Levy, D. et al. A proteomic approach for the identification of novel lysine methyltransferase substrates. Epigenetics Chromatin 4 , 19 (2011). Meng, L. et al. Mini-review: Recent advances in post-translational modification site prediction based on deep learning. Comput. Struct. Biotechnol. J. 20 , 3522–3532 (2022). Schwartz, D. Prediction of lysine post-translational modifications using bioinformatic tools. Essays Biochem. 52 , 165–177 (2012). Shilatifard, A. The COMPASS Family of Histone H3K4 Methylases: Mechanisms of Regulation in Development and Disease Pathogenesis. Annu. Rev. Biochem. 81 , 65–95 (2012). Weber, L. M. et al. The histone acetyltransferase KAT6A is recruited to unmethylated CpG islands via a DNA binding winged helix domain. Nucleic Acids Res. 51 , 574–594 (2023). Shinsky, S. A., Monteith, K. E., Viggiano, S. & Cosgrove, M. S. Biochemical Reconstitution and Phylogenetic Comparison of Human SET1 Family Core Complexes Involved in Histone Methylation. J. Biol. Chem. 290 , 6361–6375 (2015). Rienzo, M. et al. PRDM12 in Health and Diseases. Int. J. Mol. Sci. 22 , 12030 (2021). Hashimoto, K., Wada, K., Matsumoto, K. & Moriya, M. Physical interaction between SLX4 (FANCP) and XPF (FANCQ) proteins and biological consequences of interaction-defective missense mutations. DNA Repair 35 , 48–54 (2015). Bakker, J. L. et al. Analysis of the Novel Fanconi Anemia Gene SLX4 / FANCP in Familial Breast Cancer Cases. Hum. Mutat. 34 , 70–73 (2013). Chopra, A. et al. A peptide array pipeline for the development of Spike-ACE2 interaction inhibitors. Peptides 158 , 170898 (2022). Hilpert, K., Winkler, D. F. & Hancock, R. E. Cellulose-bound Peptide Arrays: Preparation and Applications. Biotechnol. Genet. Eng. Rev. 24 , 31–106 (2007). Bradford, M. M. A rapid and sensitive method for the quantitation of microgram quantities of protein utilizing the principle of protein-dye binding. Anal. Biochem. 72 , 248–254 (1976). Hornbeck, P. V. et al. PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. Nucleic Acids Res. 40 , D261–D270 (2012). Rossum, G. V. & Drake, F. L. Python 3 Reference Manual . (CreateSpace, 2009). McKinney, W. Data Structures for Statistical Computing in Python. in 56–61 (2010). doi:10.25080/Majora-92bf1922-00a. Rowe, E. M. & Biggar, K. K. An optimized method using peptide arrays for the identification of in vitro substrates of lysine methyltransferase enzymes. MethodsX 5 , 118–124 (2018). Pedregosa, F. et al. Scikit-learn: Machine Learning in Python. ArXiv12010490 Cs (2018). Lemaitre, G., Nogueira, F. & Aridas, C. K. Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. (2016) doi:10.48550/ARXIV.1609.06570. Harris, C. R. et al. Array programming with NumPy. Nature 585 , 357–362 (2020). Hunter, J. D. Matplotlib: A 2D Graphics Environment. Comput. Sci. Eng. 9 , 90–95 (2007). Wang, H., Yan, L., Huang, H. & Ding, C. From Protein Sequence to Protein Function via Multi-Label Linear Discriminant Analysis. IEEE/ACM Trans. Comput. Biol. Bioinform. 14 , 503–513 (2017). Álvarez, Ó., Fernández-Martínez, J. L., Corbeanu, A. C., Fernández-Muñiz, Z. & Kloczkowski, A. Predicting protein tertiary structure and its uncertainty analysis via particle swarm sampling. J. Mol. Model. 25 , 79 (2019). Xu, Y., Ding, Y.-X., Deng, N.-Y. & Liu, L.-M. Prediction of sumoylation sites in proteins using linear discriminant analysis. Gene 576 , 99–104 (2016). Bergstra, J. & Bengio, Y. Random Search for Hyper-Parameter Optimization. J Mach Learn Res 13 , 281–305 (2012). Dietterich, T. G. Ensemble Methods in Machine Learning. in Multiple Classifier Systems vol. 1857 1–15 (Springer Berlin Heidelberg, 2000). The UniProt Consortium et al. UniProt: the Universal Protein Knowledgebase in 2023. Nucleic Acids Res. 51 , D523–D531 (2023). Klausen, M. S. et al. NetSurfP‐2.0: Improved prediction of protein structural features by integrated deep learning. Proteins Struct. Funct. Bioinforma. 87 , 520–527 (2019). Rosario, F. J. et al. Placental Remote Control of Fetal Metabolism: Trophoblast mTOR Signaling Regulates Liver IGFBP-1 Phosphorylation and IGF-1 Bioavailability. Int. J. Mol. Sci. 24 , 7273 (2023). MacLean, B. et al. Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. Bioinformatics 26 , 966–968 (2010). Hagberg, A. A., Schult, D. A. & Swart, P. J. Exploring Network Structure, Dynamics, and Function using NetworkX. in Proceedings of the 7th Python in Science Conference (eds. Varoquaux, G., Vaught, T. & Millman, J.) 11–15 (2008). Tate, J. G. et al. COSMIC: the Catalogue Of Somatic Mutations In Cancer. Nucleic Acids Res. 47 , D941–D947 (2019). Shannon, P. et al. Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks. Genome Res. 13 , 2498–2504 (2003). Morris, J. H. et al. clusterMaker: a multi-algorithm clustering plugin for Cytoscape. BMC Bioinformatics 12 , 436 (2011). Bader, G. D. & Hogue, C. W. An automated method for finding molecular complexes in large protein interaction networks. BMC Bioinformatics 4 , 2 (2003). Additional Declarations There is NO Competing Interest. Supplementary Files SupplementaryTables.xlsx Supplementary Tables SupplementaryFigures.pdf Cite Share Download PDF Status: Published Journal Publication published 07 Nov, 2025 Read the published version in Communications Chemistry → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3771179","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":262074013,"identity":"b36d6988-2a62-439b-b4d2-203b074ad096","order_by":0,"name":"Kyle Biggar","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABEklEQVRIie2PsUrEQBCGZwlstRDLORDzChuEYCFJ6WtsCOw9wDWCwVsJ7DXRewRfQd/g5GCv0QcQQWKTOjaSStzLNTabsxRuP5ZhWOZj5gfweP4hfBUoID1ue9LYcrxXiRWxSj0oAbeFAd2jnMJW0UNP8U9KgvlNQ+qz9P7ipb3sy3cWRpVpoEydyjnmFbdZisfnafJamxmbaDrlYIoxRaPNUsRK0jeggnHDEiQqGDts0RNtlWVLZ/AtWLZT5u74douNj2mEkgZE2y10UNZOJa4/KsxrFBzbYHJ7JxgaKVGYjVPhm+Kp++yvs2gpSdd/iZOwWhvsyiunAkcCwL78YfX7V7gFgHA3mkVqbMrj8XgOmh/D/U7WffLuhwAAAABJRU5ErkJggg==","orcid":"","institution":"Carleton University","correspondingAuthor":true,"prefix":"","firstName":"Kyle","middleName":"","lastName":"Biggar","suffix":""},{"id":262074014,"identity":"ea1b1f0a-70f6-4153-b988-1eb497893a5f","order_by":1,"name":"Nashira Ridgeway","email":"","orcid":"","institution":"Carleton University","correspondingAuthor":false,"prefix":"","firstName":"Nashira","middleName":"","lastName":"Ridgeway","suffix":""},{"id":262074015,"identity":"493f978c-e5a7-46b1-bd2c-615c6871f876","order_by":2,"name":"Anand Chopra","email":"","orcid":"","institution":"Carleton University","correspondingAuthor":false,"prefix":"","firstName":"Anand","middleName":"","lastName":"Chopra","suffix":""},{"id":262074016,"identity":"1349324c-768f-49b3-bfc6-9bab0dc62815","order_by":3,"name":"Valentina Lukinovic","email":"","orcid":"","institution":"Carleton University","correspondingAuthor":false,"prefix":"","firstName":"Valentina","middleName":"","lastName":"Lukinovic","suffix":""},{"id":262074017,"identity":"12f59239-aae4-4b32-994c-98a81de16d6e","order_by":4,"name":"Michal Feldman","email":"","orcid":"","institution":"Ben Gurion University of the Negev","correspondingAuthor":false,"prefix":"","firstName":"Michal","middleName":"","lastName":"Feldman","suffix":""},{"id":262074018,"identity":"a6997f06-c9af-4b8b-8d1d-c18e3055d71b","order_by":5,"name":"Francois Charih","email":"","orcid":"","institution":"Carleton University","correspondingAuthor":false,"prefix":"","firstName":"Francois","middleName":"","lastName":"Charih","suffix":""},{"id":262074019,"identity":"f742022c-fc7b-4b27-b74f-613a9aa19934","order_by":6,"name":"Dan Levy","email":"","orcid":"","institution":"Ben Gurion University of the Negev","correspondingAuthor":false,"prefix":"","firstName":"Dan","middleName":"","lastName":"Levy","suffix":""},{"id":262074020,"identity":"521bde24-7439-41d8-bfd6-b5fa7c07a872","order_by":7,"name":"James Green","email":"","orcid":"https://orcid.org/0000-0002-6039-2355","institution":"Carleton University","correspondingAuthor":false,"prefix":"","firstName":"James","middleName":"","lastName":"Green","suffix":""}],"badges":[],"createdAt":"2023-12-18 09:47:36","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3771179/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3771179/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s42004-025-01717-6","type":"published","date":"2025-11-07T05:00:00+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":51689997,"identity":"fa881534-f7b3-4f6d-8ab5-b7f8eca1cbb1","added_by":"auto","created_at":"2024-02-27 09:25:51","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":624357,"visible":true,"origin":"","legend":"\u003cp\u003eGraphical overview of the hybrid in vitro and in silico\u003cem\u003e \u003c/em\u003emethods employed to generate a ML model for the prediction of PTMs by specific enzymes. The generation of data is completed with peptide arrays that consist of the known sites of the PTM of interest and an enzymatic construct. ML models with data balancing are applied and assessed for performance on the resulting imbalanced dataset with metrics such as F-score, precision, and recall. Finally, the resulting model is applied to the proteome to search for novel substrates of the PTM of interest, shown for the model enzyme SET8.\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/bfb40e655ea352940704ac23.png"},{"id":51689998,"identity":"524679f7-080a-45fc-a336-18020234e349","added_by":"auto","created_at":"2024-02-27 09:25:51","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":1768493,"visible":true,"origin":"","legend":"\u003cp\u003eFoundational experimental outcomes. \u003cstrong\u003eA \u003c/strong\u003eSDS-PAGE of SET8\u003csub\u003e193-352\u003c/sub\u003e-His\u003csub\u003e6\u003c/sub\u003e construct expression to verify and assess purity. \u003cstrong\u003eB \u003c/strong\u003eMTAse-Glo activity assay of SET8\u003csub\u003e193-352\u003c/sub\u003e-His\u003csub\u003e6\u003c/sub\u003e with H4K20 demonstrates lysine methyltransferase capability. \u003cstrong\u003eC \u003c/strong\u003ePhosphorimage of permutation array experiment of H4K20 site (±7 AAs) with SET8, performed in triplicate. \u003cstrong\u003eD \u003c/strong\u003eSites of the SET8 recognition motif that scored above 0.5 validated with peptide array experiments of the lysine methylome, \u003cstrong\u003eE \u003c/strong\u003eand the proteome, isolated for surface-exposed lysine (random subset, n=100). \u003cstrong\u003eF \u003c/strong\u003eThreshold analysis of the basic SET8 ML model, including performance metrics of precision, recall, and specificity. \u003cstrong\u003eG\u003c/strong\u003e The top graph is the precision-recall curve for the ML-hybrid ensemble model, with the receiver-operating-characteristic (ROC) curve on the bottom. \u003cstrong\u003eH\u003c/strong\u003e Threshold analysis of the ML-hybrid ensemble model, like that in F. Selected threshold for the optimization of model performance is indicated at 0.82 in red. \u003cstrong\u003eI \u003c/strong\u003eComparison of methods, including a random search, the SET8 recognition motif generated with permutation arrays, and the ML-hybrid ensemble model applied to the proteome dataset comprising surface-exposed lysine. \u003cstrong\u003eJ \u003c/strong\u003eAmino acid composition comparison for the ML-hybrid ensemble model predictions of SET8 methylation sites.\u0026nbsp;\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/9a5fda7dfcfe013fb71c51b6.png"},{"id":51690001,"identity":"5e337bcf-e3d3-4013-90bc-45a160a1d357","added_by":"auto","created_at":"2024-02-27 09:25:51","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":895938,"visible":true,"origin":"","legend":"\u003cp\u003eML-hybrid ensemble model results of the proteome set that contains surface-exposed lysine. \u003cstrong\u003eA \u003c/strong\u003eSAFE mapping of ML-hybrid ensemble model predictions to the HuRI interactome, clustered for associated biological processes. \u003cstrong\u003eB\u003c/strong\u003e Progression of model predictions and SET8 substrates validated with peptide array experiments.\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/9c8b450f20722a951722bb1a.png"},{"id":51690000,"identity":"d610126e-4803-4686-bcc4-782ed2776ff9","added_by":"auto","created_at":"2024-02-27 09:25:51","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1088267,"visible":true,"origin":"","legend":"\u003cp\u003eMS analysis of the validated predictions of SET8 methyltransferase activity, generated by the ML-hybrid ensemble model. \u003cstrong\u003eA \u003c/strong\u003eWestern blot analysis of the overexpression of SET8 within HCT116 cells, compared with wild type. \u003cstrong\u003eB \u003c/strong\u003ePrimary and secondary interactors of SET8, as represented in the STRING database. Circles represent MS-monitored sites. \u003cstrong\u003eC \u003c/strong\u003eTargeted MS results of SETD1BK41me1, including the peptide sequence described in the isolation list. \u003cstrong\u003eD \u003c/strong\u003eRelative expression of SETD1B-K41me1, KAT6A-K314me1, and PRDM12-K269me1 of SET8 overexpression (SET8 OE), compared with wild type (NT). \u003cstrong\u003eE \u003c/strong\u003eTotal peak intensities of targeted MS analysis for the three substrates shown in D.\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/641946ce10c39644fc3a039c.png"},{"id":51690002,"identity":"a19335c6-8de9-4ab8-909d-4b83415dab04","added_by":"auto","created_at":"2024-02-27 09:25:52","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1020661,"visible":true,"origin":"","legend":"\u003cp\u003ePrediction of SET8 substrates among missense mutations in breast cancer. \u003cstrong\u003eA \u003c/strong\u003eKaplan-Meier plot depicting patient survival with high (green) and low (red) SET8 expression in breast cancer (hazard ratio is 5.90 and \u003cem\u003ep\u003c/em\u003e-value is 0.009). \u003cstrong\u003eB \u003c/strong\u003eComparison of predicted SET8 sites of lysine methyltransferase activity between healthy and mutated cancerous states. Squares at the end of each row correspond to the color code applied in C, which\u003cstrong\u003e \u003c/strong\u003eis a visualization of cancer-mutated and healthy SET8 predicted ML scores for each site analyzed. \u003cstrong\u003eD \u003c/strong\u003eNuclear excision repair pathway initiated by DNA damage. Mutation of ERCC4 involved in the critical excision step is predicted to invoke SET8 methyltransferase activity by the ML-hybrid ensemble model. Gap filling by PCNA, a known substrate of SET8 in normal cells, is included.\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/27823904c5f09af17fadf740.png"},{"id":95430519,"identity":"9168d4bd-2afe-4c00-b083-3a606115102e","added_by":"auto","created_at":"2025-11-08 08:11:14","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":6678795,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/55d52e20-20d8-4911-aea9-49f324393624.pdf"},{"id":51690003,"identity":"4cff06e6-919f-441f-88dd-efb9edd92216","added_by":"auto","created_at":"2024-02-27 09:25:52","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":10075929,"visible":true,"origin":"","legend":"\u003cp\u003eSupplementary Tables\u003c/p\u003e","description":"","filename":"SupplementaryTables.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/e44023e7774e25f5242a9beb.xlsx"},{"id":51689999,"identity":"c40a3be4-2915-4611-afc5-ab1099a0dfe3","added_by":"auto","created_at":"2024-02-27 09:25:51","extension":"pdf","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":756417,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cbr\u003e\u003c/p\u003e","description":"","filename":"SupplementaryFigures.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3771179/v1/562a6759218040fa5842ba0b.pdf"}],"financialInterests":"There is \u003cb\u003eNO\u003c/b\u003e Competing Interest.","formattedTitle":"Machine learning-based exploration of enzyme-substrate networks: SET8-mediated methyllysine and its changing impact within cancer proteomes","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eIn the current era, in which the human genome\u003csup\u003e1\u003c/sup\u003e has been decoded for nearly two decades, unraveling the functional intricacies of the vast majority of human proteins remains a mystery. This challenge predominantly arises from the influence of post-translational modifications (PTMs), which are reversible chemical alterations with the potential to profoundly shape a function of a modified protein. With more than 500 distinct PTMs identified to date, the functional proteome transcends the approximately 20,000 proteins encoded by the human genome. This dynamic process involves covalently attaching functional chemical groups, such as methyl, phosphate, or acetyl, to specific amino acid (AA) residues within proteins. PTMs, which are mediated by specific modifying enzymes, can lead to substantial alterations in a protein's activity, stability, and folding\u003csup\u003e2\u003c/sup\u003e. Notably, proteins leverage PTMs as a cellular messaging system, responding to various external cues and stressors. This dynamic interplay between proteins and PTMs facilitates cellular adaptation to diverse environments (i.e., ensuring the maintenance of cellular homeostasis). However, the dysregulation of PTMs has been implicated in conditions ranging from cancer to inflammatory and immune disorders, highlighting their critical role in both health and disease\u003csup\u003e3,4\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eCentral to understanding PTMs is recognizing the interaction networks between the modifying enzymes and their corresponding PTM-modified substrates (i.e., enzyme-substrate networks). This recognition illuminates the extent to which a modifying enzyme impacts the proteome. Yet, conventional discovery methodologies face significant challenges that have historically hindered the growth of these substrate networks. Peptide arrays and mass spectrometry (MS) analysis, although valuable, have their own set of limitations and biases\u003csup\u003e5,6\u003c/sup\u003e. Peptide arrays offer high-throughput representation of protein segments, but fall short of capturing the full scope of PTM function as new modification sites are identified\u003csup\u003e7,8\u003c/sup\u003e. In contrast, MS analysis provides a comprehensive view of cellular mechanics, but often necessitates challenges to affinity or chemical enrichment steps, particularly for the discovery of lysine methylation and methyltransferase (KMT) substrates\u003csup\u003e1,9,10\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAmid these challenges, the integration of artificial intelligence has emerged as a novel approach to dissecting protein function and enzyme-substrate selection. Although generalized \u003cem\u003ein silico\u003c/em\u003e methods have paved the way for predicting PTMs\u003csup\u003e11\u0026ndash;13\u003c/sup\u003e, the application of deep learning, as seen in MusiteDeep\u003csup\u003e14\u003c/sup\u003e, is a transformative leap forward. However, the prediction of specific substrates for PTM-inducing enzymes remains a starkly unexplored frontier\u003csup\u003e15\u0026ndash;17\u003c/sup\u003e. The innovative fusion of \u003cem\u003ein vitro\u003c/em\u003e experiments with \u003cem\u003ein silico\u003c/em\u003e predictions, exemplified by methodologies such as the Bayesian framework employed to characterize the substrates of protein-tyrosine phosphatase, PTP1B, using protein\u0026ndash;protein interaction prediction, offer a promising avenue for substrate prediction\u003csup\u003e18\u003c/sup\u003e. However, such methods are limited by their reliance on databases that can be poorly representative of the enzyme of interest or of uncertain quality\u003csup\u003e19\u003c/sup\u003e. Often machine learning (ML)-based PTM prediction methods that are enzyme-specific require details about structure or the metabolic networks of the enzyme\u003csup\u003e19,20\u003c/sup\u003e. In this study, we transcend traditional techniques by adopting a ML-hybrid ensemble approach to enzyme-substrate identification that is generalizable across diverse enzyme classes. We show that this paradigm can successfully identify enzyme-catalyzed PTM sites for lysine methylation (SET8) and (de)acetylation (SIRT1-7) modifying enzymes.\u003c/p\u003e \u003cp\u003eSET8, a member of the Su(var)3\u0026ndash;9, Enhancer of zeste, Trithorax-homology (SET) family of KMT enzymes, serves as a representative model in this study. The previously identified recognition site of SET8 manifests a notable specificity towards lysines located in unfolded regions of proteins\u003csup\u003e21\u003c/sup\u003e, positioned\u0026thinsp;\u0026plusmn;\u0026thinsp;4 amino acids from the central lysine. This heightened specificity poses challenges in distinguishing SET8 methylation sites solely from peptide arrays, as the involvement of biophysical features beyond a simple sequential representation is anticipated in substrate recognition. Considering these complexities, SET8 emerges as a prime candidate for a systematic machine learning-based approach to substrate identification. SET8 mono-methylates histone H4 lysine 20 (H4K20), an event implicated in DNA damage repair, DNA replication, and cell cycle control\u003csup\u003e22,23\u003c/sup\u003e. SET8 also targets non-histone proteins, including K382 in the C-terminal protein domain of the p53 tumor suppressor, K248 of the proliferating cell nuclear antigen (PCNA), and two sites on the mitosis-associated protein Numb, K158 and K163\u003csup\u003e21,23,24\u003c/sup\u003e. Additional substrates have been proposed, including UHRF1-K385 and α-tubulin-K311\u003csup\u003e25,26\u003c/sup\u003e. SET8 has been shown to be overexpressed in bladder cancer, non-small cell and small cell lung carcinomas, pancreatic cancer, leukemia, and other diseases\u003csup\u003e23\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTo broadly apply our ML-hybrid approach in delineating substrate specificity across diverse enzyme families and evaluating model performance in various enzyme classes, we investigate the substrate networks associated with the sirtuin (SIRT) family of nicotinamide adenine dinucleotide (NAD+)-dependent deacetylases. Comprising seven homologs denoted as SIRT1 to SIRT7, the SIRT family plays a pivotal role in diverse physiological processes, including inflammation, glucose and lipid metabolism, oxidative stress response, cell apoptosis, autophagy, cell proliferation, as well as cell migration and invasion\u003csup\u003e27\u003c/sup\u003e. This family is broadly involved in histone and non-histone deacetylation, with significant variations in subcellular localization and catalytic activity levels observed among its members. In the context of cancer, the SIRT family has been implicated in a wide spectrum of malignancies, autoimmune disorders, cardiovascular diseases, and respiratory disorders\u003csup\u003e27\u003c/sup\u003e. Consequently, the identification of substrates governing SIRT function holds timely and substantial relevance.\u003c/p\u003e \u003cp\u003eHere, we improved on conventional \u003cem\u003ein vitro\u003c/em\u003e and \u003cem\u003ein silico\u003c/em\u003e techniques of substrate discovery by employing a novel ML approach trained on a complete peptide representation of the modified methyl-lysine and acetyl-lysine proteomes. Unlike most ML predictors of PTMs, our \u0026ldquo;hybrid\u0026rdquo; approach begins with the experimental generation of enzyme-specific training data, rather than relying purely on an online database, in which just a handful of substrates exist\u003csup\u003e28\u003c/sup\u003e. By chemically synthesizing a representative PTM proteome using peptide arrays, then subjecting them to \u003cem\u003ein vitro\u003c/em\u003e enzymatic activity, we can characterize enzymatic PTM activity in a facile way. With the development of a machine learning model, augmented by generalised PTM-specific prediction\u003csup\u003e11,29\u003c/sup\u003e, we created ML-hybrid ensemble models unique to each enzyme that demonstrates that significantly enhanced predictive accuracy in cell models (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The application of the ML-hybrid ensemble model to a proteome bearing missense mutations associated with breast cancer uncovers potential novel pathways of SET8-mediated function in cancer cells. Furthermore, the enzyme-substrate networks produced by ML-hybrid ensemble models specific to each SIRT family member reveals novel potential pathways of conserved and enzyme-specific interaction.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eThis pioneering ML-hybrid ensemble method not only outperforms conventional \u003cem\u003ein vitro\u003c/em\u003e and \u003cem\u003ein silico\u003c/em\u003e approaches, it also demonstrates broad applicability across diverse PTM-inducing enzymes. As we stand at the intersection of artificial intelligence and protein function, this novel paradigm not only sheds light on the intricate world of PTMs, but it also sets a precedent for future investigations into enzyme-substrate networks and protein function modulation.\u003c/p\u003e"},{"header":"RESULTS","content":"\u003cp\u003e\u003cstrong\u003eSET8 expression and conventional substrate prediction with permutation array analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA well-characterized and highly active SET8\u003csub\u003e193-352\u003c/sub\u003e construct\u003csup\u003e28\u003c/sup\u003e was applied to the peptide array\u003cem\u003e\u0026nbsp;\u003c/em\u003eexperiments. The purity (\u003cstrong\u003eFigure 2A\u003c/strong\u003e) and methyltransferase activity of the construct was assessed for the H4K20 peptide (GGAKRHRKVLRDNIQ) (\u003cstrong\u003eFigure 2B\u003c/strong\u003e). There are several generalized methods of identifying novel substrates for PTM-inducing enzymes. To accurately assess the ability of the proposed SET8\u0026nbsp;ML-hybrid ensemble\u0026nbsp;methodology to determine novel substrates, a comparable array-based permutation motif was generated to identify potential substrates\u003csup\u003e7,21\u003c/sup\u003e. The permutation array was created using the histone H4K20 sequence (\u0026plusmn;4 AAs, each sequence has 15 AAs in total) and was exposed to SET8 activity to identify peptide variants susceptible to methylation (\u003cstrong\u003eFigure 2C, Supplementary Table 1\u003c/strong\u003e). Densitometry results processed by PeSA2.0 yielded the following motif: [KPGCHIVD]XH[RVIKYSAHML]\u003cstrong\u003e\u003cu\u003eK\u003c/u\u003e\u003c/strong\u003e[IVT]L[RDLGI]X (\u003cstrong\u003eSupplementary Figure 1\u003c/strong\u003e)\u003csup\u003e29\u003c/sup\u003e. A search of the known methyllysine proteome (\u003cstrong\u003eSupplementary Table 2\u003c/strong\u003e) was performed with the scoring matrix, and a normalized score cutoff of 0.5, relative to unmodified H4K20 peptide (assigned a score of 1), yielded 346 hits (\u003cstrong\u003eSupplementary Table 3\u003c/strong\u003e). Of these candidate substrate hits, just 26 peptides were validated as being methylated by SET8 in vitro\u003cem\u003e\u0026nbsp;\u003c/em\u003ewith peptide array, indicating a method precision rate of 7.5% in this enriched methyllysine proteome dataset (\u003cstrong\u003eFigure 2D, Supplementary Table 3)\u003c/strong\u003e.\u003c/p\u003e\n\u003cp\u003eTo accurately compare both the SET8 ML-hybrid ensemble model and permutation methods of substrate prediction, we next identified novel SET8 substrates from the dataset of surface-exposed lysine. This approach reveals the applicability of these approaches to the exploration of enzyme substrates beyond those that are currently known to be modified (e.g., missense cancer mutations). Using the scoring matrix from the permutation array (\u003cstrong\u003eFigure 2C\u003c/strong\u003e), a search of the surface-exposed lysine proteome was performed and yielded 15,961 sites contained within 2,424 proteins (\u003cstrong\u003eSupplementary Table 4\u003c/strong\u003e). A randomly selected subset of these positively predicted sites (n=100) identified two positive hits, indicating a precision rate of 2% (precision represents the quantity of validated positive predictions among all positive predictions) (\u003cstrong\u003eFigure 2E\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTraining set generation: SET8 substrates within the known methyllysine proteome\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo apply an effective ML model, the initial dataset must provide sufficient samples of the positive case and the negative case, ideally in equal amounts\u003csup\u003e30,31\u003c/sup\u003e. A randomly sampled subset (n=100) of the approximately 600,000 lysine-centric sites within the proteome tested with peptide arrays identified no sites of SET8 methylation. This observation is indicative of SET8\u0026rsquo;s highly specific recognition of substrates for methylation, and emphasizes the need for an improved approach for generating training data\u003csup\u003e21\u003c/sup\u003e. To efficiently enhance the likelihood of detecting positive sites, a targeted subset of the known methyllysine proteome was obtained from PhosphoSitePlus\u003csup\u003e27\u003c/sup\u003e. This dataset contains modified lysines, including mono-, di-, and tri-methylation. Enzymes such as KMTs often contain conserved catalytic domains and act upon methylatable histone and non-histone substrates, meaning the methyllysine proteome should contain an enriched number of substrates for SET8; in contrast, the methylated status of the full proteome is unknown\u003csup\u003e24\u003c/sup\u003e.\u003c/p\u003e\n\u003cp\u003eUpon analysis of the SET8-exposed peptide arrays that comprise the methyllysine proteome obtained from PhosphoSitePlus\u003csup\u003e27\u003c/sup\u003e, the targeted subset successfully contained peptide substrates of SET8. Specifically, of the 4,593 peptides tested, 213 were deemed to be positive for SET8 methylation (\u003cstrong\u003eSupplementary Table 3\u003c/strong\u003e). The 213 positive sites were identified across 179 proteins, indicating that several proteins harbored multiple SET8 methylation sites. To date, the commonly accepted substrates of SET8 are P53-K382, PCNA-K248, Numb-K158, and Numb-K163\u003csup\u003e24\u003c/sup\u003e. The 213 sites identified in vitro with the targeted search alone expands upon these four substrates. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSET8 base model fitting and fine-tuning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe lysine methylome for SET8 was numerically encoded with the application of MACCS keys, one-hot sequential encoding, and ProtDCal molecular descriptions (\u003cstrong\u003eSupplementary Table 5\u003c/strong\u003e)\u003csup\u003e32,33,34,35,36\u003c/sup\u003e. The resulting set contains 483 features. With the stratified K-fold cross-validation method, the SET8 methyllysine proteome dataset was split into training and testing sets to effectively assess model fitting and prevent overfitting\u003csup\u003e37\u003c/sup\u003e. The F-score was selected as the best way to measure the predictive performance of the model on the imbalanced dataset\u003csup\u003e30,31\u003c/sup\u003e. A linear discriminant analysis, along with random oversampling of the positive class (i.e., sites positive for SET8 methylation), attained the highest F-score, which was 0.13\u003csup\u003e38,39\u003c/sup\u003e. An m-threshold analysis was performed (\u003cstrong\u003eFigure 2F\u003c/strong\u003e), as well as both precision-recall and receiver operating characteristic (ROC) curves were generated (\u003cstrong\u003eFigure 2G\u003c/strong\u003e). Metrics for the default threshold of 0.5 resulted in an F-score of 0.13, a precision of 0.085, recall of 0.24, and specificity of 0.83. The metrics further demonstrate the benefit of the F-score, defined by the harmonic mean of precision and recall, and describes our positive identification rate, rather than using specificity, which is falsely inflated by the negative identification rate.\u003c/p\u003e\n\u003cp\u003eFeature importance was analyzed for the selected model, hereafter referred to as the base model. Features deemed crucial by the model for identifying positive sites of SET8 methylation were specific one-hot-encoded AA/position combinations, including tryptophan, cysteine, and tyrosine at positions +3, +4, and \u0026ndash;6 from the central lysine, respectively. Additionally, the MACCS key corresponding to the aromatic bond between carbon and nitrogen (key 65), found in histidine, proline, and tryptophan, was deemed to be of high importance\u003csup\u003e40\u003c/sup\u003e. Sulphur (key 88), found in cysteine and methionine was also determined to be highly important to the model\u0026rsquo;s classification of positives\u003csup\u003e40\u003c/sup\u003e. Regarding the classification of negatives, or sites not methylated by SET8, once again, one-hot encoded positions played a crucial role. Specifically, regarding the central lysine (position 0), cysteine at \u0026ndash;6, phenylalanine, and methionine at +6, and methionine at \u0026ndash;2 scored highly for feature importance in negative classification. One MACCS key was included as well, key 132, which represents AO-CH\u003csub\u003e2\u003c/sub\u003e-A (where A represents any elemental symbol) and likely corresponds to the presence of aspartic acid, glutamic acid, serine, or threonine within the site\u003csup\u003e40\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSET8 ML-hybrid ensemble model construction\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe ability of a lysine residue to undergo methylation is a prerequisite for any newly predicted SET8 substrate. To enhance the performance of the SET8 substrate prediction, or base model, a composite or ensemble model was constructed using MethylSight, the current state-of-the-art generalized predictor of lysine methylation\u003csup\u003e11\u003c/sup\u003e. In the inaugural study, MethylSight identified 51 novel sites of histone methylation, and 89% of the sites were confirmed to physically exhibit methylated lysine\u003csup\u003e11\u003c/sup\u003e. Much like the SET8 substrate prediction model previously described, MethylSight uses ProtDCal to characterize the 15-AA-long site surrounding a central lysine\u003csup\u003e11\u003c/sup\u003e. Hence, it is well suited for integration with the SET8 substrate prediction model using stacked ensemble learning.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAs with the initial model fitting, stratified K-fold cross-validation was applied to assess the performance of each model. Two features were applied: the SET8 substrate prediction score (described above); and the MethylSight score (i.e., the likelihood of methylation). The F-score was optimized with the application of a logistic regression model and SVM SMOTE oversampling\u003csup\u003e41,42,43\u003c/sup\u003e. The simplicity of logistic regression was reflected in the singular hyperparameter of 100 max iterations determined from the tuning process\u003csup\u003e41\u003c/sup\u003e. A much-improved F-score of 0.12 was determined for the ensemble model, along with improved values of 0.25 for precision, 0.08 for recall, and 0.98 for specificity. A comparison of performance metrics with classification threshold is illustrated in \u003cstrong\u003eFigure 2H\u003c/strong\u003e. To optimize the performance of the ensemble model, a threshold cutoff of 0.82 was applied. Given the performance increase gained from the integration of methyllysine prediction into the ensemble model (hereafter referred to as the SET8 ML-hybrid ensemble model), the investigation proceeded with this hybrid model (\u003cstrong\u003eSupplementary Tables 6 and 7\u003c/strong\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eProteome-wide prediction of SET8 substrates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUsing our SET8 ML-hybrid ensemble model, experimental validation of the 2,367 predicted positive sites of SET8 methylation was completed by testing each site for in vitro methylation. Of these predictions, 885 sites permitted in vitro SET8 methyltransferase activity, representing a validated precision of 37.4%. The precision of this method is much improved over the 0% validated precision of the random search method and the 2% validated precision determined with the permutation array within the surface-exposed lysine proteome (\u003cstrong\u003eFigure 2I\u003c/strong\u003e). An analysis of the sequence composition of the predicted sites of SET8 methylation by the ML-hybrid ensemble model demonstrates that the known SET8 substrates differ from the predicted sites, with substantial variation observed in the latter (\u003cstrong\u003eFigure 2J\u003c/strong\u003e). The SET8 ML-hybrid ensemble model proved to be 100% accurate in identifying a subset (n=362) of predicted negative, lowest-scoring sites, as verified by peptide array experiments with SET8 (\u003cstrong\u003eSupplementary Table 8\u003c/strong\u003e). Based on these findings, it is clear our SET8 ML-hybrid ensemble model improves on the traditional substrate identification approach.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eA total of 2,367 positive SET8 methylation sites were predicted by the SET8\u0026nbsp;ML-hybrid ensemble model\u0026nbsp;within the surface-exposed lysine dataset, representing sites within 1,203 proteins. To investigate the enriched biological functions of the 1,203 proteins (i.e., predicted SET8 substrate network), clustering analysis with GO annotations was performed using the spatial analysis of functional enrichment (SAFE) approach (\u003cstrong\u003eFigure 3\u003c/strong\u003e)\u003csup\u003e44,45\u003c/sup\u003e. The HuRI proteome was selected for protein mapping because of its quality, high-confidence interactions among proteins\u003csup\u003e46\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eA shared theme among the enriched biological processes is involvement in cell homeostasis, regulation, and control of the cell cycle (\u003cstrong\u003eFigure 3A\u003c/strong\u003e). Given the established involvement of SET8 with these cellular events, mediated through known substrates, the possibility that SET8 might participate in such processes through the methylation of other substrates identified by our SET8\u0026nbsp;ML-hybrid ensemble\u0026nbsp;model is bolstered\u003csup\u003e23,24,28,47\u003c/sup\u003e. Other affiliated processes include mRNA and RNA polyadenylation. Regulation through the polyadenylation of mRNA has been associated with other SET-domain-containing methyltransferases, specifically SET1 and SET2, through histone methylation\u003csup\u003e48\u003c/sup\u003e. The involvement of SET8 in transcription modulation may also implicate it in polyadenylation regulation; however, a direct connection has not been reported\u003csup\u003e23\u003c/sup\u003e. In summary, the substrates generated by the ML-hybrid ensemble model provide the potential to unveil new functional narratives for SET8 and its role(s) in disease. The efficacy of the SET8 ML-hybrid ensemble model is further demonstrated by the progression from the proteome isolated for surface-exposed lysine residues that contain 145,379 sites to the 2,367 predictions (\u003cstrong\u003eSupplementary Tables 9 and 10\u003c/strong\u003e), which resulted in 885 in vitro validated sites, as shown in \u003cstrong\u003eFigure 3B\u003c/strong\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCell-based validation of SET8 substrate candidates\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo validate SET8-influenced cellular methyllysine events, we used parallel reaction monitoring MS to assess the in vitro SET8 substrates newly identified by our SET8 ML-hybrid ensemble model in a targeted manner. To restrict the number of methylation sites monitored with this approach, we generated an isolation list that was constrained to the primary and secondary interactors of SET8, as described by the STRING database\u003csup\u003e49\u003c/sup\u003e (\u003cstrong\u003eSupplementary Table 11\u003c/strong\u003e). The 44 proteins within the network in \u003cstrong\u003eFigure 4A\u003c/strong\u003e contained 75 sites of predicted SET8 methylation that were verified by peptide array experiments. Of these 75 sites, it was predicted that 32 sites would create suitable digested peptides in silico; these were targeted for MS monitoring in SET8 overexpressed HCT116 cells (\u003cstrong\u003eFigure 4B;\u003c/strong\u003e \u003cstrong\u003eSupplementary Table 12\u003c/strong\u003e). Of the 32 monitored sites, only nine were reliability detectable, and elevated levels of mono-methylation were observed in three (33%) of these substrates: SETD1B-K41, KAT6A-K314, and PRDM12-K269 (\u003cstrong\u003eFigure 4C\u0026ndash;4E; Supplementary Figure 3\u003c/strong\u003e). In conclusion, the ML-hybrid ensemble model is able to identify novel substrates of possible SET8 methylation activity, as confirmed by site-targeted MS monitoring.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSET8 substrate discovery in cancer\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eElevated expression of SET8 is linked to a high mortality rate in patients with breast cancer (\u003cstrong\u003eFigure 5A\u003c/strong\u003e)\u003csup\u003e50,51\u003c/sup\u003e. However, the behavior of SET8 in cancerous cells remains unclear, and further investigation is required to uncover the functional role(s) SET8 plays in tumorigenesis. As cancer-associated mutations continue to diversify, mutation datasets serve as a valuable resource with which to elucidate the effect mutations have on protein structure and function. In the case of missense mutations, they may cause the gain or loss of methylatable lysine or make changes to neighboring residues, which then dictate the suitability of these sites for SET8 methylation\u003csup\u003e50\u003c/sup\u003e.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTo explore the possibility of gain or loss of SET8 substrates in breast cancer, missense mutations were downloaded from the COSMIC database (v.96) and applied to the human proteome. Of the initial mutations, 9,438 either occurred within seven AAs of a lysine residue (e.g., any residue) or resulted in the gain or loss of an individual lysine, directly impacting the creation or loss of a potential methylation site (\u003cstrong\u003eSupplementary Table 13\u003c/strong\u003e). The corresponding unmutated sites, except sites in which a lysine did not previously exist, were also assembled. Application of the SET8\u0026nbsp;ML-hybrid ensemble model\u0026nbsp;to normal and breast cancer datasets predicted that most of the mutations (94.6%) would not affect SET8\u0026rsquo;s methylation behavior toward the site, likely because their structure is not changed dramatically by a single AA mutation. In contrast, 4.0% (376) of mutations resulted in a predicted gain of SET8 methylation, and 0.7% (62) resulted in a loss (\u003cstrong\u003eFigure 5B\u003c/strong\u003e). Of the 4.0% of sites predicted to gain SET8 substrate status, 46.8% (176) were the result of the mutation introducing a new lysine that is itself predicted to be methylated by SET8 (\u003cstrong\u003eFigure 5C\u003c/strong\u003e). MCODE clustering analysis of the total set of mutations revealed a directed subset of predicted substrate interactions that were highly interconnected with SET8 (\u003cstrong\u003eSupplementary Figure 4\u003c/strong\u003e). Mutations within the subset that resulted in a gain of predicted SET8 methylation were investigated for involvement in pathways implicated in breast cancer. Of particular interest was XPF (encoded by ERCC4), a protein associated with the vital cellular process of DNA damage repair\u003csup\u003e52\u003c/sup\u003e. Specifically, the XPF-S352A mutation led to the prediction of a new SET8 methylation site at XPF-K350. As detailed in \u003cstrong\u003eFigure 5D\u003c/strong\u003e, XPF is directly involved in DNA damage repair, including nucleotide excision, double-strand break, and interstrand cross-link repair pathways\u003csup\u003e53\u003c/sup\u003e. Gap filling is then proceeded by PCNA, a known substrate of SET8\u003csup\u003e24\u003c/sup\u003e, further implicating SET8 in the NER pathway\u003csup\u003e54\u003c/sup\u003e. Finally, ligation is performed with DNA ligase and the NER pathway for DNA damage is complete\u003csup\u003e53\u003c/sup\u003e. Interestingly, SET8 has been implicated in DNA repair previously, specifically in 53BPI/BRCA1 double-stranded DNA repair through histone H4K20 mono-methylation\u003csup\u003e54\u003c/sup\u003e. ERCC4 gene mutations have also been determined to affect XPF function within the NER pathway\u003csup\u003e52\u003c/sup\u003e. Beyond breast cancer, our SET8\u0026nbsp;ML-hybrid ensemble model was also applied to missense mutations present in pancreatic cancer\u0026nbsp;(COSMIC database, v.96) (\u003cstrong\u003eSupplementary Figure 5A and 5B;\u0026nbsp;Supplementary Table 14\u003c/strong\u003e).\u003c/p\u003e"},{"header":"DISCUSSION","content":"\u003cp\u003eThe introduction of our ML-hybrid ensemble model marks a significant leap forward in the identification of substrates for specific PTM-inducing enzymes. This methodology is successfully generalized across multiple enzyme classes and is inevitably extendable to a broader spectrum of PTMs, including phosphorylation, sumoylation, and ubiquitination. The power of this approach lies in its ability to efficiently characterize PTM-specific subsets of the proteome, a task made feasible by the manageable size; only a few thousand sites can be readily profiled in parallel using peptide arrays (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Moreover, this approach ensures greater success in positively identifying new enzyme substrates than established array-based methods\u003csup\u003e5,57\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eAlthough many predictive models can identify sites of a particular PTM, few possess the capacity to pinpoint the enzyme responsible\u003csup\u003e11,15,58\u003c/sup\u003e. In contrast, our ML-hybrid ensemble methodology is highly specific to the enzyme under investigation and reveals the predicted PTM site. Notably, the ensemble models for both SET8 and SIRT2 demonstrate exceptional performance when experimentally validated in cells overexpressing the target enzyme, a departure from many PTM prediction tools\u003csup\u003e58,59\u003c/sup\u003e. The ML-hybrid ensemble model\u0026rsquo;s experimentally validated precision of 37\u0026ndash;43% stands in stark contrast to the 2.0% validated precision achieved with a traditional permutation array in the \u003cem\u003ein vitro\u003c/em\u003e search for new substrates within the human proteome. Further the high precision of our models enables its facile use in directing substrate identification studies using more targeted, and intensive validation strategies.\u003c/p\u003e \u003cp\u003eThe limited amount of labeled PTM-specific training data imposes restrictions on ML model complexity, but significantly reduces computational load. Deep learning algorithms, including neural networks, are precluded from predicting enzyme-specific substrates because of the dataset\u0026rsquo;s small size. However, this constraint aligns with the notion that deep learning in PTM prediction is most effective in a more generalized context, as seen in MusiteDeep\u003csup\u003e58\u003c/sup\u003e. Most potential substrates for an enzyme within the proteome are, in fact, not substrates, as a result there is an inherent imbalance that narrows the range of applicable models. Although the abbreviated peptide sequences may not fully represent a full-length enzyme-substrate interaction, this limitation is not exclusive to our approach; conventional permutation array experiments employ similar methodologies when validating results\u003csup\u003e21\u003c/sup\u003e. The unique advantage of our ML-hybrid ensemble method lies in the extensive number of predictions (more than 2,300 for SET8), yielding a substantial number of validated predictions (885 for SET8), thereby expanding the list of possible substrates for MS-monitoring methods compared with conventional methods\u003csup\u003e21\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe value of any predictive model of substrate selection lies within its ability to prioritize the exploration of new datasets and reveal novel insights. The targeted MS monitoring of predicted sites of SET8 methylation in cells highlighted the further elevation of SETD1B-K41, KAT6A-K314, and PRDM12-K269 mono-methylation levels in SET8 overexpressed HCT116 cells (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). These findings suggest the potential involvement of SET8 in gene activation and regulatory pathways. The change in SETD1B-K41me1 levels in SET8 overexpressed cells may link SET8 to the COMPASS complex, of which SETD1B is a part\u003csup\u003e60\u003c/sup\u003e. KAT6A, commonly localized to CpG islands, may be influenced by SET8 mono-methylation at K269\u003csup\u003e61\u003c/sup\u003e, potentially impacting gene regulation. Furthermore, the association of both KAT6A and SETD1B with CXXC1 (CpG-binding protein) implicates SET8 in gene activation\u003csup\u003e61,62\u003c/sup\u003e. The regulatory domain binding factor PRDM12 was observed to demonstrate elevated K269me1 levels with SET8 overexpression\u003csup\u003e63\u003c/sup\u003e. With additional investigation, this may suggest SET8\u0026rsquo;s involvement in gene activation and regulation, potentially through the methylation of PRDM12-K269\u003csup\u003e63\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eTo help reveal potential insights into the role(s) SET8 plays in breast cancer, our SET8 ML-hybrid ensemble model was applied to explore the creation and loss of potential substrates to help annotate a predicted breast-cancer-specific SET8 enzyme-substrate network. When applied to cancer mutation datasets, the ML-hybrid ensemble model has specific advantages. The sensitivity of the model is evident when comparing oncogenic and healthy proteome scores, which provide a unique perspective on oncogenic mutations. For example, the predicted gain of SET8 substrate methylation due to a missense mutation sheds light on potential proteins of interest, particularly the association of SET8 with breast cancer within the NER pathway through XPF (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e)\u003csup\u003e54\u003c/sup\u003e. This implicates SET8 in the DNA repair process, employing PCNA, another substrate of SET8, for gap filling\u003csup\u003e54\u003c/sup\u003e. The XPF-S352A mutation and subsequent SET8 methylation at XPF-K350 could be hypothesized to affect the NER pathway, such that DNA damage repair may not be completed. Interestingly, the K350 site falls within the XPF-SLX4 interaction region\u003csup\u003e64\u003c/sup\u003e, an interaction associated with the Fanconi anemia DNA repair pathway and defects in SLX4 function have been linked to breast cancer\u003csup\u003e65\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe generalizability of the ML-hybrid ensemble approach is demonstrated by the sirtuin family investigation. The rates of identified \u003cem\u003ein vitro\u003c/em\u003e peptide substrates for each SIRT range from as low as 2.59% positive (97.4% negative), to 38.4% positive (61.6% negative). With the balancing methodology and the careful tuning of models, each SIRT-based model produced metrics as impressive, if not more so, than the SET8 data. This is exemplified by the MS-validated precision of the SIRT2 model of 43.0%. The conserved and enzyme-specific substrates within each SIRT network further provides a unique insight into potential pathways of activity (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e). Interestingly, SIRT2 and SIRT3 share almost twice the substrates as any other combination of SIRT. This finding is supported by prior investigation that uncovered that the pair may compensate for each other when the other is deficient\u003csup\u003e66\u003c/sup\u003e. The second largest overlap in predicted substrates occurs between SIRT5 and SIRT7. As the most recently discovered member, SIRT7 is underrepresented in literature, however some associations between SIRT5 and SIRT7 exist. The two have been implicated within disease states such as cardiac hypertrophy and inflammatory bowel diseases, as well as the regulation of the NF-κB inflammation signalling pathway\u003csup\u003e27\u003c/sup\u003e. The predicted substrates of the ML-hybrid ensemble model for SIRT7 may prove to be a key to direct future investigations to uncover SIRT7\u0026rsquo;s cellular function(s).\u003c/p\u003e \u003cp\u003eIn conclusion, our generalizable ML-hybrid ensemble approach represents a significant advance in the precise identification of enzyme-substrate networks and the features of substrate enzymes that influence selection. The methodology\u0026rsquo;s proven versatility offers promise for a wide array of PTM-inducing enzymes, thereby illuminating previously uncharted territories of protein function modulation. Although some limitations exist, including dataset size and an unavoidable class imbalance, our approach applies techniques to overcome such constraints to offer a powerful tool for exploring the intricate world of PTMs. The verification of predicted SET8 and SIRT2 substrates in cells underscores the model\u0026rsquo;s robustness, outperforming traditional methods and paving the way for a more holistic understanding of enzyme-substrate selection and involvement in cellular processes. The application of this model to cancer datasets provides a promising avenue with which to uncover critical insights into oncogenic mutations and the role of enzymes within these modified proteomes. Moreover, our ML-hybrid ensemble approach presents the capability to uncover the enzyme-substrate network of an entire family of enzymes to accurately predict the extent of an enzyme family\u0026rsquo;s functional role within the cell. With its potential for broader application and its capacity to shed light on unexplored facets of cellular regulation, our ML-hybrid ensemble model stands poised to have a profound impact on the study of enzyme-substrate networks and protein function modulation.\u003c/p\u003e "},{"header":"METHODS","content":"\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\u003cdiv id=\"Sec13\" class=\"Section3\"\u003e \u003ch2\u003ePeptide synthesis\u003c/h2\u003e \u003cp\u003ePeptide SPOT arrays were synthesized to commercial aminated cellulose membranes (Intavis Inc.) with standard Fmoc (N-(9-fluorenyl)methoxycarbonyl) chemistry automatically using a Multipep synthesizer (Intavis Inc.)\u003csup\u003e67\u003c/sup\u003e. Resulting arrays contained approximately 2 nmol peptide per SPOT, separated from the membrane surface by a flexible C-terminal 6-aminohexanoic acid linker. Post-treatment involved the cleavage of the protective groups from the side-chains with an acidic solution (51% water, 47.5% trifluoroacetic acid, and 1.5% tri-isopropylsilane)\u003csup\u003e67,68\u003c/sup\u003e. Arrays were then washed in ethanol and dried for future use. Dried arrays were stored desiccated at 4\u0026deg;C.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003eProtein expression and purification\u003c/h2\u003e \u003cp\u003eA bacterial expression construct of human SET8 encoding the catalytic site residues 191\u0026ndash;352 were cloned into the parallel expression vector pHIS2 using BamHI and XhoI restriction sites\u003csup\u003e48\u003c/sup\u003e. The construct was transformed in BL21 DE3 \u003cem\u003eEscherichia coli\u003c/em\u003e, and after propagation were induced with 0.3 mM isopropyl β-D-1-thiogalacttopyranoside\u003csup\u003e48\u003c/sup\u003e. Following protein extraction, affinity column purification of the soluble lysate was completed with 500 \u0026micro;L HisPur\u0026trade; Ni-NTA Resin (ThermoFisher Scientific, Cat# 88221). Fractions were eluted using P500 buffer (50 mM NaHPO\u003csub\u003e4\u003c/sub\u003e (pH 7), 500 mM NaCl, 10% glycerol, 0.05% TritonX-100, 1 mM DTT, 500 mM Imidazole) and then dialyzed into a storage buffer (20 mM tris pH 7.5-8, 200 mM NaCl, 10% glycerol, 1 mM DTT). Protein purity was assessed with a 12% SDS PAGE gel (\u003cb\u003eSupplementary Fig.\u0026nbsp;1A\u003c/b\u003e) stained with Coomassie Brilliant Blue G-250 stain. Concentration was determined with a Bradford Assay\u003csup\u003e69\u003c/sup\u003e. Prior to storage, activity was confirmed using the Methyltransferase-Glo Assay to assess \u003cem\u003ein vitro\u003c/em\u003e activity with H4K20 peptide (GGAKRHRKVLRDNIQ) (\u003cb\u003eSupplementary Fig.\u0026nbsp;1B\u003c/b\u003e), according to the manufacturer\u0026rsquo;s specifications (Promega, Cat# V7601). Recombinant SET8 protein was then snap frozen stored at -80\u0026deg;C.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec15\" class=\"Section2\"\u003e \u003ch2\u003eTraining dataset generation\u003c/h2\u003e \u003cp\u003eThe annotated Human lysine methylome was obtained from PhosphoSitePlus (accessed February 5th, 2020). This dataset contained sequence information for each lysine methylation site represented as 15 residue peptides, centered on the methylated lysine. (i.e., position 8) \u003csup\u003e28,70\u003c/sup\u003e. The data was then cleaned using Python 3 and the Pandas package for efficient isolation of human protein sites of lysine methylation\u003csup\u003e71,72\u003c/sup\u003e. Briefly, sites that were within 7 Aas from the beginning or end of the sequence were padded with alanine residues, and any duplicate sequences were removed (i.e., methylation events that occur within a conserved 15 AA sequence, but among unique proteins). The resulting dataset yielded a total of 4,593 Human lysine methylation sites (\u003cb\u003eSupplementary Table\u0026nbsp;2\u003c/b\u003e).\u003c/p\u003e \u003cp\u003ePeptide arrays were then methylated using 1 \u0026micro;M SET8\u003csub\u003e191\u0026minus;\u0026thinsp;325\u003c/sub\u003e with 5 \u0026micro;Ci/mL S-[methyl-\u003csup\u003e3\u003c/sup\u003eH]-Adenosyl-L-methionine (SAM) in methylation buffer (50 mM Tris pH 8.5, 2 mM MgCl2, 10 \u0026micro;M DTT) overnight at room temperature\u003csup\u003e73\u003c/sup\u003e. After methylation, arrays were washed 6x 3 min in buffer (100 mM NH4HCO3, 1% SDS) and then a 7% 2,5-Diphenyloxazole solution in ethanol was sprayed generously over the array and left to air dry. This procedure was repeated a total of three times. Dried peptide arrays were then exposed to intensifying screens (Dupont, Cronex Lighting Plus) at -80\u0026deg;C for two weeks and imaged using a Typhoon\u0026trade; FLA 7000 IP phosphorimager (General Electric). Peptide methylation was determined by densitometry using the Protein Array Analyzer (v.1.1.c) toolset for Image J (v.1.53) was applied to obtain SPOT densitometry from the array images. The peptide substrate dataset for the sirtuin family was adapted from Rauh et al., 2013\u003csup\u003e5\u003c/sup\u003e. Peptides which produced a signal ratio score of 1 or greater were classified as positive, and positions with masked AA residues were replaced with alanine (\u003cb\u003eSupplementary Table\u0026nbsp;16\u003c/b\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003eBiochemical Feature Generation\u003c/h2\u003e \u003cp\u003ePeptide libraries were represented numerically through one-hot encoding, a common method of vectorizing peptides\u003csup\u003e34,35\u003c/sup\u003e. A matrix describing AA position and identity with a 1 or 0 is applied to the sequence, resulting in 300 features for SET8, and 260 for the sirtuins (Eq.\u0026nbsp;1).\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\({x}_{AA\\text{, }pos}=\\left\\{\\begin{array}{c}1 if seq\\left[pos\\right]==AA\\\\ 0 if else \\end{array}\\right. \\text{, for }pos=\\left[1,N\\right]\\text{, }AA=\\{A, C, D, \\dots , W\\}\\)\u003c/span\u003e \u003c/span\u003e [1]\u003c/p\u003e \u003cp\u003eTo encode molecular structure, the Molecular Access System (MACCS) Keys of 166 binary fingerprints were applied\u003csup\u003e36\u003c/sup\u003e. Created to describe molecular structure, MACCS Keys include predefined atom symbols, bond types, and atom properties\u003csup\u003e36\u003c/sup\u003e. The RDKit package for Python (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e\u003ca href=\"http://www.rdkit.org\" target=\"_blank\"\u003ewww.rdkit.org\u003c/a\u003e\u003c/span\u003e\u003cspan address=\"http://www.rdkit.org\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was applied to generate MACCS Keys from the site sequence, representing the molecular structure for the AAs present.\u003c/p\u003e \u003cp\u003eAggregate molecularly descriptive features were generated for the peptide sequence through ProtDCal (v4.5)\u003csup\u003e38\u003c/sup\u003e. The 17 metrics selected included molecular weight, hydrophobicity, isoelectric point, free energy, and Levitt\u0026rsquo;s Probability for various protein conformation amongst other descriptive sequence-based metrics\u003csup\u003e38\u003c/sup\u003e (\u003cb\u003eSupplementary Table\u0026nbsp;5\u003c/b\u003e). Combined, the sequential information provided by one-hot encoding, along with the unique molecular elements represented by the MACCS keys and the molecularly descriptive features determined by ProtDCal resulted in 483 features for SET8, and 443 for the sirtuins.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e\u003cb\u003eMachine Learning \u0026ndash; Model Fitting\u003c/b\u003e\u003c/h2\u003e \u003cp\u003eAll procedures related to model fitting and data balancing were completed in Python 3 using the Scikit-Learn and Imbalanced-learn packages\u003csup\u003e71,74,75\u003c/sup\u003e. Pandas and Numpy were applied for data handling and storage, and plots were generated with Matplotlib\u003csup\u003e72,76,77\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eClass imbalance was present in all samples, although it ranged in degree for each enzyme studied. The SET8-methylated lysine peptide array data provided a class imbalance of 213 positives within the 4,593 sites tested for SET8 activity. Imbalance, along with the relatively small size of the training data, was taken into careful consideration when selecting ML models\u003csup\u003e33\u003c/sup\u003e. For proper comparison to a baseline model, the dummy classifier which simply randomly classifies data input was first fit to the data. Applying the F-score as the primary selection criteria and a model complexity considerate of the size and imbalance of our dataset as the secondary selection criteria, the selected models were linear discriminant analysis (LDA) and the decision tree classifier\u003csup\u003e40\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eClass imbalance in which the positive case composed less than 30% of the training dataset was addressed with data balancing or sampling methods, which was the case for all datasets except SIRT5 (38.4% positive, 61.6% negative). Both over and undersampling techniques were tested to improve F-score with our models, and the best performing data balancing approach was selected (\u003cb\u003eSupplementary Table\u0026nbsp;17\u003c/b\u003e). In these approaches, a selection of the positive class is replicated within the dataset, increasing the overall percentage of positive values\u003csup\u003e74\u003c/sup\u003e. Alternatively, negative values are removed to reduce the overall percentage of the negative case within the training dataset.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003eMachine Learning \u0026ndash; Cross-validation\u003c/h2\u003e \u003cp\u003eThe fit of each model was assessed with repeated stratified K-fold cross-validation\u003csup\u003e39\u003c/sup\u003e. Cross-validation involves the splitting of the labelled training set and related features into n-folds. In this case, the label indicates whether our datapoint (i.e., modification site) is methylated. Due to the imbalanced nature of our data, stratification was applied to ensure the same percentage of positive and negative labels are represented within each fold.\u003c/p\u003e \u003cp\u003eMany metrics may be used to assess classification performance; however, the nature of the data must be taken into careful consideration. In a dataset with fewer positive instances than negatives, the preferred metric is the F-score, as it accounts for true positive and false positive counts, as well as false negatives\u003csup\u003e33,39\u003c/sup\u003e (Eq.\u0026nbsp;2). To assess the fit of each model, stratified 10-fold cross-validation was applied to the lysine methylome with 3 repeats and using F-score.\u003cdiv class=\"BlockQuote\"\u003e\u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(F=\\frac{2TP}{2TP+FN+FP}\\)\u003c/span\u003e \u003c/span\u003e [2]\u003c/p\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThe best-scoring models were determined to be the linear discriminant analysis (LDA) method and the decision tree classification method. The LDA method determines a linear combination of features that optimally separates the two classes; \u0026ldquo;modified by our enzyme\u0026rdquo; and \u0026ldquo;\u003cem\u003enot\u003c/em\u003e modified by our enzyme\u0026rdquo;. The LDA method has been applied to solve biochemical problems in the past, including the prediction of protein function and tertiary structure from sequence\u003csup\u003e78,79\u003c/sup\u003e. One investigation employed LDA to predict the site of protein sumoylation, highlighting the benefit of LDA in the prediction of PTMs\u003csup\u003e80\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe decision tree classifier composes a model through the creation of nodes representing a test on the features provided. With each node, a branch descends to eventually point towards a classification of the input to one of the two classes as defined in the LDA method. Decision tree classification has been applied to the classification of enzymatic and non-enzymatic metal binding sites, as well as the familial classification of an unknown protein\u003csup\u003e81,82\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003eMachine Learning \u0026ndash; Hyperparameter Tuning\u003c/h2\u003e \u003cp\u003eThe next step of model fitting involves the tuning of the model\u0026rsquo;s hyperparameters. These model-specific parameters control the method in which the model is applied to the data. To effectively search for the optimal hyperparameters, a grid search was applied\u003csup\u003e83\u003c/sup\u003e. Each combination of hyperparameter was tested with the repeated stratified K-fold cross-validation method, and the combination resulting in the highest F-score was returned\u003csup\u003e39,83\u003c/sup\u003e. The balancing methods are also tuned through hyperparameter optimization to better the model\u0026rsquo;s performance.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003eMachine Learning \u0026ndash; Optimizing Decision Threshold\u003c/h2\u003e \u003cp\u003eAdjustment of the model\u0026rsquo;s threshold is required to fine-tune the performance, particularly so with imbalanced training data. When applied to new data, the trained model outputs a score representative of the probability that a data point is a member of the positive class. By default, a decision threshold of 0.5 is applied the probability; however, this decision threshold can be tuned to trade off recall and precision\u003csup\u003e33\u003c/sup\u003e. Recall assesses the model\u0026rsquo;s identification of true positives; precision measures the proportion of positive predictions that are correct; and accuracy over truly negative instances are quantified with specificity\u003csup\u003e33\u003c/sup\u003e. For significantly imbalanced datasets, the best threshold for the model was determined by identifying the point at which the F-metric was maximized\u003csup\u003e11,33,39\u003c/sup\u003e. Otherwise, a decision threshold of 0.5 was maintained.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec21\" class=\"Section2\"\u003e \u003ch2\u003eMachine Learning \u0026ndash; Ensemble Learning\u003c/h2\u003e \u003cp\u003eTo improve the overall performance of the models, ensemble learning algorithms were explored. In ensemble algorithms, a secondary ML predictor is applied in parallel. The lysine methylation predictor MethylSight was implemented as the secondary method for the SET8 model, as it generally predicts sites of lysine methylation; a dependent variable for a lysine to be a SET8 substrate\u003csup\u003e11\u003c/sup\u003e. MethylSight uses support vector ML, and was validated experimentally to predict novel sites of methylation, including the identification of novel KDM5B substrate H2B-K43me2\u003csup\u003e11\u003c/sup\u003e. For the sirtuin family the generalized deep learning predictor MusiteDeep was applied as the secondary predictor. The N6-acetyl-lysine predictive function of MusiteDeep was employed to generally identify sites of lysine acetylation, much like MethylSight for SET8. MusiteDeep outperformed other representative ML and deep learning algorithms for N6-acetyl-lysine prediction\u003csup\u003e29\u003c/sup\u003e.\u003c/p\u003e \u003cp\u003eThe ensemble methods explored included hard voting, soft voting, and stacking. Each method takes in the probability score output by either predictor that a site is methylated. The hard voting method classifies through most votes, or scores above our specified positive cutoff thresholds for the positive class by either model\u003csup\u003e84\u003c/sup\u003e. In soft voting, classification is determined through the mean of probabilities by either model\u003csup\u003e84\u003c/sup\u003e. Finally, stacking utilizes the scores of each model as features of a third model, along with dataset balancing as previously outlined\u003csup\u003e84\u003c/sup\u003e. Each method was tested, and F-scores were generated with repeated stratified K-fold cross-validation. For SET8, a holdout set from the MethylSight investigation was employed to provide an unbiased final validation of each ensemble method\u003csup\u003e11\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec22\" class=\"Section2\"\u003e \u003ch2\u003eExperimental Dataset Generation and Evaluation\u003c/h2\u003e \u003cp\u003eThe human proteome was obtained from UniProt, and analyzed with NetSurfP2.0 to generate values for relevant surface accessibility\u003csup\u003e85,86\u003c/sup\u003e. Lysine residues with a relevant surface accessibility value of over 0.2 were selected, and the surrounding (\u0026plusmn;7 and \u0026plusmn;6 from central lysine) AAs were isolated from the full protein sequence to form the experimental datasets. As with the generation of the lysine methylome, sites less than 7 or 6 AAs from the beginning or end of a protein sequence were padded with alanine. Additionally, any duplicated sequences or sequences found within the training dataset were dropped.\u003c/p\u003e \u003cp\u003eThe SET8 ensemble learning model was applied to the resulting set and resulted in 2,367 novel positive predictions of new lysine methylation sites that are also SET8 substrates (\u003cb\u003eSupplementary Table\u0026nbsp;9\u003c/b\u003e). These values were mapped to the human binary protein interactome (HuRI), and abundant biological processes were showcased with the Spatial Analysis of Functional Enrichment (SAFE) package for Python 3\u003csup\u003e45,47,71\u003c/sup\u003e. Additionally, all predicted positive sites were tested individually via peptide SPOT array experiments for SET8 \u003cem\u003ein vitro\u003c/em\u003e methylation, as previously described. The resulting seven sirtuin models were applied and predicted various novel sites of deacetylation (\u003cb\u003eSupplementary Table\u0026nbsp;18\u003c/b\u003e). The commonality of predictions between each SIRT was illustrated through Circos and UpSet plots as generated by the pyCircos and UpSetPlot packages for Python 3\u003csup\u003e71,87,88\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec23\" class=\"Section3\"\u003e \u003ch2\u003eCell line transfection\u003c/h2\u003e \u003cp\u003eCells were cultured in Dulbecco\u0026rsquo;s modified Eagle\u0026rsquo;s medium (DMEM, Gibco Cat# 11965092) supplemented with 10% heat-inactivated FBS and 1% penicillin/streptomycin at 37\u0026deg;C, 5% CO\u003csub\u003e2\u003c/sub\u003e. Cell lines were regularly tested for \u003cem\u003eMycoplasma pneumoniae\u003c/em\u003e using polymerase chain reaction (PCR) assay. For SET8 overexpression experiments, cells were transfected with 10 \u0026micro;g pcDNA3.1 HA-SET8 using jetOptimus transfection reagent (Polyplus, Cat# 101000025), following manufacturer instruction. The cells were harvested 48h after transfection and cell pellets were snap frozen and stored at -20\u0026deg;C until further use. Cell pellets were lysed in extraction buffer containing 50 mM tris (pH 8), 150 mM NaCl, 1 mM EDTA, 10% glycerol, 0.5% NP-40, and protease inhibitors (1 mM PMSF, 10 \u0026micro;M E-64, 1 \u0026micro;M Pepstatin, 1 \u0026micro;M Leupeptin, and 1 mM sodium orthovanadate). Protein concentration was determined by Bradford assay and samples stored at -80\u0026deg;C until use. Conditions of SIRT2 overexpression in HCT116 cells are described previously\u003csup\u003e56\u003c/sup\u003e.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec24\" class=\"Section2\"\u003e \u003ch2\u003eMass Spectrometry\u003c/h2\u003e \u003cp\u003eProtein concentration was determined by Pierce BCA Protein Assay (Thermo Fisher Scientific) and 500 \u0026micro;g of soluble lysate was digested with Arg-C (Roche, 11370529001). Cell lysates were first diluted in Arg-C digestion buffer (100 mM Tris-HCl, 10 mM CaCl\u003csub\u003e2\u003c/sub\u003e, pH 7.6) followed by standard reduction and alkylation at room temperature. Briefly, proteins were first reduced with 3 mM TCEP for 45 minutes, alkylated with 15 mM iodoacetamide (IAA) for 60 minutes, followed by quenching of any unused IAA by incubation with 20 mM DTT for 45 minutes. Next, activation solution and Arg-C were added to a final concentration of 1X and a 1:200 protein-to-protease ratio, respectively. The digestion occurred overnight at 37 \u003csup\u003eo\u003c/sup\u003eC with end-to-end rotation. Resultant peptides were desalted C18 Spin Columns (Thermo Fisher Scientific, 89870) and C18-ZipTip (Millipore Sigma, ZTC18S096), dried by SpeedVac, and resuspended in 20 \u0026micro;L of MS-grade water\u0026thinsp;+\u0026thinsp;0.1% formic acid.\u003c/p\u003e \u003cp\u003eTo assess mono-methylation levels of newly identified \u003cem\u003ein vitro\u003c/em\u003e SETD8 substrates, parallel reaction monitoring (PRM)-MS was performed using a Q-Exactive Plus hybrid quadrupole-orbitrap mass spectrometer, at the John L. Holmes Mass Spectrometry Facility at the University of Ottawa, as previously described\u003csup\u003e89\u003c/sup\u003e. PRM-MS scanning was guided by an isolation list built in Skyline Software\u003csup\u003e90\u003c/sup\u003e. The isolation list was generated by first identifying the primary interactors of SET8, isolated from the STRING human interactome dataset using Python 3 and the Pandas package\u003csup\u003e50,71,72\u003c/sup\u003e. The interactors of the primary interactors of SET8, or secondary interactors were also included in this list. Proteins which contained the predicted and validated SET8 \u003cem\u003ein vitro\u003c/em\u003e methylation sites were selected from within the primary and secondary interactors of SET8, and the resulting network was mapped with the NetworkX package for Python\u003csup\u003e91\u003c/sup\u003e. The verified sites contained within this subnetwork were then used to generate an isolation list for targeted mass spectrometry analysis. To quantify relative mono-methylation levels of detectable target peptides, Skyline software was used to design an isolation list (\u003cb\u003eSupplementary Table\u0026nbsp;15\u003c/b\u003e) and was used to derive total peak areas for modified target substrates as well as other unmodified peptides along the protein sequence. The total peak area of each modification site of interest was divided by that of all other reliably detected unmodified peptides within a given parental protein. The MS-verified sites for SIRT2 were obtained from the dataset provided in Zhang et. al., 2022\u003csup\u003e56\u003c/sup\u003e.\u003c/p\u003e \u003cdiv id=\"Sec25\" class=\"Section3\"\u003e \u003ch2\u003eApplication to the Cancer Proteome\u003c/h2\u003e \u003cp\u003eTargeted screen mutant repositories were obtained for breast cancer from v96 of the COSMIC database (cancer.sanger.ac.uk)\u003csup\u003e92\u003c/sup\u003e. In Python 3, the data were cleaned to isolate missense mutations, which were then applied to the full human proteome obtained from UniProt\u003csup\u003e71,85\u003c/sup\u003e. Mutations which occur within \u0026plusmn;7 AAs from either (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) a lysine, or (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) a position mutated to a lysine, were accepted within the oncoproteome set. An accompanying dataset for the healthy proteome was generated. The ML-hybrid ensemble model for SET8 methylation prediction was applied to both the healthy and cancerous dataset. The resulting scores were compared, yielding several potential states for mutations; (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) mutations predicted to have a null effect on SET8 methylation (i.e., no change from healthy population), (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) mutations predicted to remove SET8 methylation activity (i.e., loss of SET8 substrate from healthy population), and (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) mutations predicted to induce SET8 methylation (i.e., gain of SET8 from health population). The human interactome from the STRING database was applied to the mutated proteins and the probability scores from the SET8 ML-hybrid ensemble model were scaled to resemble STRING scores and included within the interaction list as well. The network was mapped with Cytoscape (v3.9.1), and molecular subcomplexes were isolated using the Molecular Complex Detection cluster generator from the clusterMaker (v2.0) toolset\u003csup\u003e93\u0026ndash;95\u003c/sup\u003e. The subcomplex which contained SET8 was further investigated for proteins implicated in pathways associated with the cancer type, and resulting pathways were represented in Adobe Illustrator.\u003c/p\u003e \u003c/div\u003e "},{"header":"Declarations","content":"\u003cdiv id=\"Sec26\" class=\"Section3\"\u003e \u003ch2\u003eDATA AVAILABILITY\u003c/h2\u003e \u003cp\u003eThe data conveying all results reported may be readily obtained within the github repository \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/nashirag/ML-Hybrid_Ensemble_Method\u003c/span\u003e\u003cspan address=\"https://github.com/nashirag/ML-Hybrid_Ensemble_Method\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Source data for all figures is accessible within the same repository, including all SPOT peptide array densitometry figures are available. All mass spectrometry data is available through the PeptideAtlas repository (project ID: PASS05848).\u003c/p\u003e \u003c/div\u003e \u003ch2\u003eCODE AVAILABILITY\u003c/h2\u003e\n\u003cp\u003eThe Python code implemented to produce all results may be accessed at https://github.com/nashirag/ML-Hybrid_Ensemble_Method_SET8.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBrandi, J., Noberini, R., Bonaldi, T. \u0026amp; Cecconi, D. Advances in enrichment methods for mass spectrometry-based proteomics analysis of post-translational modifications. \u003cem\u003eJ. Chromatogr. A\u003c/em\u003e\u003cstrong\u003e1678\u003c/strong\u003e, 463352 (2022).\u003c/li\u003e\n\u003cli\u003eDeribe, Y. L., Pawson, T. \u0026amp; Dikic, I. Post-translational modifications in signal integration. \u003cem\u003eNat. Struct. Mol. Biol.\u003c/em\u003e\u003cstrong\u003e17\u003c/strong\u003e, 666\u0026ndash;672 (2010).\u003c/li\u003e\n\u003cli\u003eLiu, J., Qian, C. \u0026amp; Cao, X. Post-Translational Modification Control of Innate Immunity. \u003cem\u003eImmunity\u003c/em\u003e\u003cstrong\u003e45\u003c/strong\u003e, 15\u0026ndash;30 (2016).\u003c/li\u003e\n\u003cli\u003eQian, M. \u003cem\u003eet al.\u003c/em\u003e Targeting post-translational modification of transcription factors as cancer therapy. \u003cem\u003eDrug Discov. Today\u003c/em\u003e\u003cstrong\u003e25\u003c/strong\u003e, 1502\u0026ndash;1512 (2020).\u003c/li\u003e\n\u003cli\u003eRauh, D. \u003cem\u003eet al.\u003c/em\u003e An acetylome peptide microarray reveals specificities and deacetylation substrates for all human sirtuin isoforms. \u003cem\u003eNat. Commun.\u003c/em\u003e\u003cstrong\u003e4\u003c/strong\u003e, 2327 (2013).\u003c/li\u003e\n\u003cli\u003eMerbl, Y. \u0026amp; Kirschner, M. W. Large-scale detection of ubiquitination substrates using cell extracts and protein microarrays. \u003cem\u003eProc. Natl. Acad. Sci.\u003c/em\u003e\u003cstrong\u003e106\u003c/strong\u003e, 2543\u0026ndash;2548 (2009).\u003c/li\u003e\n\u003cli\u003eMoore, K. E. \u0026amp; Gozani, O. An unexpected journey: Lysine methylation across the proteome. \u003cem\u003eBiochim. Biophys. Acta BBA - Gene Regul. Mech.\u003c/em\u003e\u003cstrong\u003e1839\u003c/strong\u003e, 1395\u0026ndash;1403 (2014).\u003c/li\u003e\n\u003cli\u003ePolo, S. \u003cem\u003eet al.\u003c/em\u003e A single motif responsible for ubiquitin recognition and monoubiquitination in endocytic proteins. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e416\u003c/strong\u003e, 451\u0026ndash;455 (2002).\u003c/li\u003e\n\u003cli\u003eMitchell, C. J. \u003cem\u003eet al.\u003c/em\u003e Unbiased identification of substrates of protein tyrosine phosphatase ptp‐3 in C. elegans. \u003cem\u003eMol. Oncol.\u003c/em\u003e\u003cstrong\u003e10\u003c/strong\u003e, 910\u0026ndash;920 (2016).\u003c/li\u003e\n\u003cli\u003eYu-Ying, Y., Markus, G. \u0026amp; Howard, H. C. Identification of lysine acetyltransferase p300 substrates using 4-pentynoyl-coenzyme A and bioorthogonal proteomics. \u003cem\u003eBioorg. Med. Chem. Lett.\u003c/em\u003e\u003cstrong\u003e21\u003c/strong\u003e, 4976\u0026ndash;4979 (2011).\u003c/li\u003e\n\u003cli\u003eBiggar, K. K. \u003cem\u003eet al.\u003c/em\u003e Proteome-wide Prediction of Lysine Methylation Leads to Identification of H2BK43 Methylation and Outlines the Potential Methyllysine Proteome. \u003cem\u003eCell Rep.\u003c/em\u003e\u003cstrong\u003e32\u003c/strong\u003e, 107896 (2020).\u003c/li\u003e\n\u003cli\u003eJamal, S., Ali, W., Nagpal, P., Grover, A. \u0026amp; Grover, S. Predicting phosphorylation sites using machine learning by integrating the sequence, structure, and functional information of proteins. \u003cem\u003eJ. Transl. Med.\u003c/em\u003e\u003cstrong\u003e19\u003c/strong\u003e, 218 (2021).\u003c/li\u003e\n\u003cli\u003eKiemer, L., Bendtsen, J. D. \u0026amp; Blom, N. NetAcet: prediction of N-terminal acetylation sites. \u003cem\u003eBioinformatics\u003c/em\u003e\u003cstrong\u003e21\u003c/strong\u003e, 1269\u0026ndash;1270 (2005).\u003c/li\u003e\n\u003cli\u003eNeely, B. A. \u003cem\u003eet al.\u003c/em\u003e Toward an Integrated Machine Learning Model of a Proteomics Experiment. \u003cem\u003eJ. Proteome Res.\u003c/em\u003e\u003cstrong\u003e22\u003c/strong\u003e, 681\u0026ndash;696 (2023).\u003c/li\u003e\n\u003cli\u003eDeng, W. \u003cem\u003eet al.\u003c/em\u003e GPS-PAIL: prediction of lysine acetyltransferase-specific modification sites from protein sequences. \u003cem\u003eSci. Rep.\u003c/em\u003e\u003cstrong\u003e6\u003c/strong\u003e, 39787 (2016).\u003c/li\u003e\n\u003cli\u003eWu, Z., Lu, M. \u0026amp; Li, T. Prediction of substrate sites for protein phosphatases 1B, SHP-1, and SHP-2 based on sequence features. \u003cem\u003eAmino Acids\u003c/em\u003e\u003cstrong\u003e46\u003c/strong\u003e, 1919\u0026ndash;1928 (2014).\u003c/li\u003e\n\u003cli\u003eWang, X. \u003cem\u003eet al.\u003c/em\u003e UbiBrowser 2.0: a comprehensive resource for proteome-wide known and predicted ubiquitin ligase/deubiquitinase\u0026ndash;substrate interactions in eukaryotic species. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e50\u003c/strong\u003e, D719\u0026ndash;D728 (2022).\u003c/li\u003e\n\u003cli\u003eFerrari, E. \u003cem\u003eet al.\u003c/em\u003e Identification of New Substrates of the Protein-tyrosine Phosphatase PTP1B by Bayesian Integration of Proteome Evidence. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e\u003cstrong\u003e286\u003c/strong\u003e, 4173\u0026ndash;4185 (2011).\u003c/li\u003e\n\u003cli\u003eSmith, K., Rhoads, N. \u0026amp; Chandrasekaran, S. Protocol for CAROM: A machine learning tool to predict post-translational regulation from metabolic signatures. \u003cem\u003eSTAR Protoc.\u003c/em\u003e\u003cstrong\u003e3\u003c/strong\u003e, 101799 (2022).\u003c/li\u003e\n\u003cli\u003eLanouette, S. \u003cem\u003eet al.\u003c/em\u003e Discovery of Substrates for a SET Domain Lysine Methyltransferase Predicted by Multistate Computational Protein Design. \u003cem\u003eStructure\u003c/em\u003e\u003cstrong\u003e23\u003c/strong\u003e, 206\u0026ndash;215 (2015).\u003c/li\u003e\n\u003cli\u003eKudithipudi, S., Dhayalan, A., Kebede, A. F. \u0026amp; Jeltsch, A. The SET8 H4K20 protein lysine methyltransferase has a long recognition sequence covering seven amino acid residues. \u003cem\u003eBiochimie\u003c/em\u003e\u003cstrong\u003e94\u003c/strong\u003e, 2212\u0026ndash;2218 (2012).\u003c/li\u003e\n\u003cli\u003eFang, J. \u003cem\u003eet al.\u003c/em\u003e Purification and Functional Characterization of SET8, a Nucleosomal Histone H4-Lysine 20-Specific Methyltransferase. \u003cem\u003eCurr. Biol.\u003c/em\u003e\u003cstrong\u003e12\u003c/strong\u003e, 1086\u0026ndash;1099 (2002).\u003c/li\u003e\n\u003cli\u003eMilite, C. \u003cem\u003eet al.\u003c/em\u003e The emerging role of lysine methyltransferase SETD8 in human diseases. \u003cem\u003eClin. Epigenetics\u003c/em\u003e\u003cstrong\u003e8\u003c/strong\u003e, 102 (2016).\u003c/li\u003e\n\u003cli\u003eBiggar, K. K., Wang, Z. \u0026amp; Li, S. S.-C. SnapShot: Lysine Methylation beyond Histones. \u003cem\u003eMol. Cell\u003c/em\u003e\u003cstrong\u003e68\u003c/strong\u003e, 1016-1016.e1 (2017).\u003c/li\u003e\n\u003cli\u003eZhang, H. \u003cem\u003eet al.\u003c/em\u003e SET8 prevents excessive DNA methylation by methylation-mediated degradation of UHRF1 and DNMT1. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e47\u003c/strong\u003e, 9053\u0026ndash;9068 (2019).\u003c/li\u003e\n\u003cli\u003eChin, H. G. \u003cem\u003eet al.\u003c/em\u003e The microtubule-associated histone methyltransferase SET8, facilitated by transcription factor LSF, methylates \u0026alpha;-tubulin. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e\u003cstrong\u003e295\u003c/strong\u003e, 4748\u0026ndash;4759 (2020).\u003c/li\u003e\n\u003cli\u003eHornbeck, P. V. \u003cem\u003eet al.\u003c/em\u003e PhosphoSitePlus, 2014: mutations, PTMs and recalibrations. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e43\u003c/strong\u003e, D512\u0026ndash;D520 (2015).\u003c/li\u003e\n\u003cli\u003eYin, Y. \u003cem\u003eet al.\u003c/em\u003e SET8 recognizes the sequence RHRK20VLRDN within the N terminus of histone H4 and mono-methylates lysine 20. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e\u003cstrong\u003e280\u003c/strong\u003e, 30025\u0026ndash;30031 (2005).\u003c/li\u003e\n\u003cli\u003eTopcu, E., Ridgeway, N. H. \u0026amp; Biggar, K. K. PeSA 2.0: A software tool for peptide specificity analysis implementing positive and negative motifs and motif-based peptide scoring. \u003cem\u003eComput. Biol. Chem.\u003c/em\u003e\u003cstrong\u003e101\u003c/strong\u003e, 107753 (2022).\u003c/li\u003e\n\u003cli\u003eBurkov, A. \u003cem\u003eThe hundred-page machine learning book\u003c/em\u003e. (Andriy Burkov, 2019).\u003c/li\u003e\n\u003cli\u003eBrownlee, J. \u003cem\u003eImbalanced Classification with Python: Choose Better Metrics, Balance Skewed Classes, and Apply Cost-Sensitive Learning\u003c/em\u003e. (Machine Learning Mastery, 2021).\u003c/li\u003e\n\u003cli\u003eYang, K. K., Wu, Z. \u0026amp; Arnold, F. H. Machine-learning-guided directed evolution for protein engineering. \u003cem\u003eNat. Methods\u003c/em\u003e\u003cstrong\u003e16\u003c/strong\u003e, 687\u0026ndash;694 (2019).\u003c/li\u003e\n\u003cli\u003eErjavac, I., Kalafatovic, D. \u0026amp; Mau\u0026scaron;a, G. Coupled encoding methods for antimicrobial peptide prediction: How sensitive is a highly accurate model? \u003cem\u003eArtif. Intell. Life Sci.\u003c/em\u003e\u003cstrong\u003e2\u003c/strong\u003e, 100034 (2022).\u003c/li\u003e\n\u003cli\u003eDurant, J. L., Leland, B. A., Henry, D. R. \u0026amp; Nourse, J. G. Reoptimization of MDL Keys for Use in Drug Discovery. \u003cem\u003eJ. Chem. Inf. Comput. Sci.\u003c/em\u003e\u003cstrong\u003e42\u003c/strong\u003e, 1273\u0026ndash;1280 (2002).\u003c/li\u003e\n\u003cli\u003eRuiz-Blanco, Y. B., Paz, W., Green, J. \u0026amp; Marrero-Ponce, Y. ProtDCal: A program to compute general-purpose-numerical descriptors for sequences and 3D-structures of proteins. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e\u003cstrong\u003e16\u003c/strong\u003e, 162 (2015).\u003c/li\u003e\n\u003cli\u003eRomero‐Molina, S., Ruiz‐Blanco, Y. B., Green, J. R. \u0026amp; Sanchez‐Garcia, E. ProtDCal‐Suite: A web server for the numerical codification and functional analysis of proteins. \u003cem\u003eProtein Sci.\u003c/em\u003e pro.3673 (2019) doi:10.1002/pro.3673.\u003c/li\u003e\n\u003cli\u003eSzeghalmy, S. \u0026amp; Fazekas, A. A Comparative Study of the Use of Stratified Cross-Validation and Distribution-Balanced Stratified Cross-Validation in Imbalanced Learning. \u003cem\u003eSensors\u003c/em\u003e\u003cstrong\u003e23\u003c/strong\u003e, 2333 (2023).\u003c/li\u003e\n\u003cli\u003eIzenman, A. J. Linear Discriminant Analysis. in \u003cem\u003eModern Multivariate Statistical Techniques\u003c/em\u003e 237\u0026ndash;280 (Springer New York, 2013). doi:10.1007/978-0-387-78189-1_8.\u003c/li\u003e\n\u003cli\u003eKamalov, F., Leung, H.-H. \u0026amp; Cherukuri, A. K. Keep it simple: random oversampling for imbalanced data. in \u003cem\u003e2023 Advances in Science and Engineering Technology International Conferences (ASET)\u003c/em\u003e 1\u0026ndash;4 (IEEE, 2023). doi:10.1109/ASET56582.2023.10180891.\u003c/li\u003e\n\u003cli\u003eCereto-Massagu\u0026eacute;, A. \u003cem\u003eet al.\u003c/em\u003e Molecular fingerprint similarity search in virtual screening. \u003cem\u003eMethods\u003c/em\u003e\u003cstrong\u003e71\u003c/strong\u003e, 58\u0026ndash;63 (2015).\u003c/li\u003e\n\u003cli\u003eWright, R. E. Logistic regression. \u003cem\u003eRead. Underst. Multivar. Stat.\u003c/em\u003e 217\u0026ndash;244 (1995).\u003c/li\u003e\n\u003cli\u003eChawla, N. V., Bowyer, K. W., Hall, L. O. \u0026amp; Kegelmeyer, W. P. SMOTE: Synthetic Minority Over-sampling Technique. \u003cem\u003eJ. Artif. Intell. Res.\u003c/em\u003e\u003cstrong\u003e16\u003c/strong\u003e, 321\u0026ndash;357 (2002).\u003c/li\u003e\n\u003cli\u003eNguyen, H. M., Cooper, E. W. \u0026amp; Kamei, K. Borderline over-sampling for imbalanced data classification. \u003cem\u003eInt J Knowl Eng Soft Data Paradig.\u003c/em\u003e\u003cstrong\u003e3\u003c/strong\u003e, 4\u0026ndash;21 (2009).\u003c/li\u003e\n\u003cli\u003eBaryshnikova, A. Spatial Analysis of Functional Enrichment (SAFE) in Large Biological Networks. in \u003cem\u003eComputational Cell Biology\u003c/em\u003e (eds. von Stechow, L. \u0026amp; Santos Delgado, A.) vol. 1819 249\u0026ndash;268 (Springer New York, 2018).\u003c/li\u003e\n\u003cli\u003eThe Gene Ontology Consortium \u003cem\u003eet al.\u003c/em\u003e The Gene Ontology resource: enriching a GOld mine. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e49\u003c/strong\u003e, D325\u0026ndash;D334 (2021).\u003c/li\u003e\n\u003cli\u003eLuck, K. \u003cem\u003eet al.\u003c/em\u003e A reference map of the human binary protein interactome. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e580\u003c/strong\u003e, 402\u0026ndash;408 (2020).\u003c/li\u003e\n\u003cli\u003eCouture, J.-F., Collazo, E., Brunzelle, J. S. \u0026amp; Trievel, R. C. Structural and functional analysis of SET8, a histone H4 Lys-20 methyltransferase. \u003cem\u003eGenes Dev.\u003c/em\u003e\u003cstrong\u003e19\u003c/strong\u003e, 1455\u0026ndash;1465 (2005).\u003c/li\u003e\n\u003cli\u003eKaczmarek Michaels, K., Mohd Mostafa, S., Ruiz Capella, J. \u0026amp; Moore, C. L. Regulation of alternative polyadenylation in the yeast Saccharomyces cerevisiae by histone H3K4 and H3K36 methyltransferases. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e48\u003c/strong\u003e, 5407\u0026ndash;5425 (2020).\u003c/li\u003e\n\u003cli\u003eSzklarczyk, D. \u003cem\u003eet al.\u003c/em\u003e STRING v10: protein\u0026ndash;protein interaction networks, integrated over the tree of life. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e43\u003c/strong\u003e, D447\u0026ndash;D452 (2015).\u003c/li\u003e\n\u003cli\u003eLiu, B. \u003cem\u003eet al.\u003c/em\u003e A functional single nucleotide polymorphism of SET8 is prognostic for breast cancer. \u003cem\u003eOncotarget\u003c/em\u003e\u003cstrong\u003e7\u003c/strong\u003e, 34277\u0026ndash;34287 (2016).\u003c/li\u003e\n\u003cli\u003eYang, C., Wang, K., Zhou, Y. \u0026amp; Zhang, S.-L. Histone lysine methyltransferase SET8 is a novel therapeutic target for cancer treatment. \u003cem\u003eDrug Discov. Today\u003c/em\u003e\u003cstrong\u003e26\u003c/strong\u003e, 2423\u0026ndash;2430 (2021).\u003c/li\u003e\n\u003cli\u003eBogliolo, M. \u003cem\u003eet al.\u003c/em\u003e Mutations in ERCC4, Encoding the DNA-Repair Endonuclease XPF, Cause Fanconi Anemia. \u003cem\u003eAm. J. Hum. Genet.\u003c/em\u003e\u003cstrong\u003e92\u003c/strong\u003e, 800\u0026ndash;806 (2013).\u003c/li\u003e\n\u003cli\u003eFaridounnia, M., Folkers, G. \u0026amp; Boelens, R. Function and Interactions of ERCC1-XPF in DNA Damage Response. \u003cem\u003eMolecules\u003c/em\u003e\u003cstrong\u003e23\u003c/strong\u003e, 3205 (2018).\u003c/li\u003e\n\u003cli\u003eXu, L. \u003cem\u003eet al.\u003c/em\u003e Roles for the methyltransferase SETD8 in DNA damage repair. \u003cem\u003eClin. Epigenetics\u003c/em\u003e\u003cstrong\u003e14\u003c/strong\u003e, 34 (2022).\u003c/li\u003e\n\u003cli\u003eLevy, D. \u003cem\u003eet al.\u003c/em\u003e A proteomic approach for the identification of novel lysine methyltransferase substrates. \u003cem\u003eEpigenetics Chromatin\u003c/em\u003e\u003cstrong\u003e4\u003c/strong\u003e, 19 (2011).\u003c/li\u003e\n\u003cli\u003eMeng, L. \u003cem\u003eet al.\u003c/em\u003e Mini-review: Recent advances in post-translational modification site prediction based on deep learning. \u003cem\u003eComput. Struct. Biotechnol. J.\u003c/em\u003e\u003cstrong\u003e20\u003c/strong\u003e, 3522\u0026ndash;3532 (2022).\u003c/li\u003e\n\u003cli\u003eSchwartz, D. Prediction of lysine post-translational modifications using bioinformatic tools. \u003cem\u003eEssays Biochem.\u003c/em\u003e\u003cstrong\u003e52\u003c/strong\u003e, 165\u0026ndash;177 (2012).\u003c/li\u003e\n\u003cli\u003eShilatifard, A. The COMPASS Family of Histone H3K4 Methylases: Mechanisms of Regulation in Development and Disease Pathogenesis. \u003cem\u003eAnnu. Rev. Biochem.\u003c/em\u003e\u003cstrong\u003e81\u003c/strong\u003e, 65\u0026ndash;95 (2012).\u003c/li\u003e\n\u003cli\u003eWeber, L. M. \u003cem\u003eet al.\u003c/em\u003e The histone acetyltransferase KAT6A is recruited to unmethylated CpG islands via a DNA binding winged helix domain. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e51\u003c/strong\u003e, 574\u0026ndash;594 (2023).\u003c/li\u003e\n\u003cli\u003eShinsky, S. A., Monteith, K. E., Viggiano, S. \u0026amp; Cosgrove, M. S. Biochemical Reconstitution and Phylogenetic Comparison of Human SET1 Family Core Complexes Involved in Histone Methylation. \u003cem\u003eJ. Biol. Chem.\u003c/em\u003e\u003cstrong\u003e290\u003c/strong\u003e, 6361\u0026ndash;6375 (2015).\u003c/li\u003e\n\u003cli\u003eRienzo, M. \u003cem\u003eet al.\u003c/em\u003e PRDM12 in Health and Diseases. \u003cem\u003eInt. J. Mol. Sci.\u003c/em\u003e\u003cstrong\u003e22\u003c/strong\u003e, 12030 (2021).\u003c/li\u003e\n\u003cli\u003eHashimoto, K., Wada, K., Matsumoto, K. \u0026amp; Moriya, M. Physical interaction between SLX4 (FANCP) and XPF (FANCQ) proteins and biological consequences of interaction-defective missense mutations. \u003cem\u003eDNA Repair\u003c/em\u003e\u003cstrong\u003e35\u003c/strong\u003e, 48\u0026ndash;54 (2015).\u003c/li\u003e\n\u003cli\u003eBakker, J. L. \u003cem\u003eet al.\u003c/em\u003e Analysis of the Novel Fanconi Anemia Gene \u003cem\u003eSLX4\u003c/em\u003e / \u003cem\u003eFANCP\u003c/em\u003e in Familial Breast Cancer Cases. \u003cem\u003eHum. Mutat.\u003c/em\u003e\u003cstrong\u003e34\u003c/strong\u003e, 70\u0026ndash;73 (2013).\u003c/li\u003e\n\u003cli\u003eChopra, A. \u003cem\u003eet al.\u003c/em\u003e A peptide array pipeline for the development of Spike-ACE2 interaction inhibitors. \u003cem\u003ePeptides\u003c/em\u003e\u003cstrong\u003e158\u003c/strong\u003e, 170898 (2022).\u003c/li\u003e\n\u003cli\u003eHilpert, K., Winkler, D. F. \u0026amp; Hancock, R. E. Cellulose-bound Peptide Arrays: Preparation and Applications. \u003cem\u003eBiotechnol. Genet. Eng. Rev.\u003c/em\u003e\u003cstrong\u003e24\u003c/strong\u003e, 31\u0026ndash;106 (2007).\u003c/li\u003e\n\u003cli\u003eBradford, M. M. A rapid and sensitive method for the quantitation of microgram quantities of protein utilizing the principle of protein-dye binding. \u003cem\u003eAnal. Biochem.\u003c/em\u003e\u003cstrong\u003e72\u003c/strong\u003e, 248\u0026ndash;254 (1976).\u003c/li\u003e\n\u003cli\u003eHornbeck, P. V. \u003cem\u003eet al.\u003c/em\u003e PhosphoSitePlus: a comprehensive resource for investigating the structure and function of experimentally determined post-translational modifications in man and mouse. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e40\u003c/strong\u003e, D261\u0026ndash;D270 (2012).\u003c/li\u003e\n\u003cli\u003eRossum, G. V. \u0026amp; Drake, F. L. \u003cem\u003ePython 3 Reference Manual\u003c/em\u003e. (CreateSpace, 2009).\u003c/li\u003e\n\u003cli\u003eMcKinney, W. Data Structures for Statistical Computing in Python. in 56\u0026ndash;61 (2010). doi:10.25080/Majora-92bf1922-00a.\u003c/li\u003e\n\u003cli\u003eRowe, E. M. \u0026amp; Biggar, K. K. An optimized method using peptide arrays for the identification of in vitro substrates of lysine methyltransferase enzymes. \u003cem\u003eMethodsX\u003c/em\u003e\u003cstrong\u003e5\u003c/strong\u003e, 118\u0026ndash;124 (2018).\u003c/li\u003e\n\u003cli\u003ePedregosa, F. \u003cem\u003eet al.\u003c/em\u003e Scikit-learn: Machine Learning in Python. \u003cem\u003eArXiv12010490 Cs\u003c/em\u003e (2018).\u003c/li\u003e\n\u003cli\u003eLemaitre, G., Nogueira, F. \u0026amp; Aridas, C. K. Imbalanced-learn: A Python Toolbox to Tackle the Curse of Imbalanced Datasets in Machine Learning. (2016) doi:10.48550/ARXIV.1609.06570.\u003c/li\u003e\n\u003cli\u003eHarris, C. R. \u003cem\u003eet al.\u003c/em\u003e Array programming with NumPy. \u003cem\u003eNature\u003c/em\u003e\u003cstrong\u003e585\u003c/strong\u003e, 357\u0026ndash;362 (2020).\u003c/li\u003e\n\u003cli\u003eHunter, J. D. Matplotlib: A 2D Graphics Environment. \u003cem\u003eComput. Sci. Eng.\u003c/em\u003e\u003cstrong\u003e9\u003c/strong\u003e, 90\u0026ndash;95 (2007).\u003c/li\u003e\n\u003cli\u003eWang, H., Yan, L., Huang, H. \u0026amp; Ding, C. From Protein Sequence to Protein Function via Multi-Label Linear Discriminant Analysis. \u003cem\u003eIEEE/ACM Trans. Comput. Biol. Bioinform.\u003c/em\u003e\u003cstrong\u003e14\u003c/strong\u003e, 503\u0026ndash;513 (2017).\u003c/li\u003e\n\u003cli\u003e\u0026Aacute;lvarez, \u0026Oacute;., Fern\u0026aacute;ndez-Mart\u0026iacute;nez, J. L., Corbeanu, A. C., Fern\u0026aacute;ndez-Mu\u0026ntilde;iz, Z. \u0026amp; Kloczkowski, A. Predicting protein tertiary structure and its uncertainty analysis via particle swarm sampling. \u003cem\u003eJ. Mol. Model.\u003c/em\u003e\u003cstrong\u003e25\u003c/strong\u003e, 79 (2019).\u003c/li\u003e\n\u003cli\u003eXu, Y., Ding, Y.-X., Deng, N.-Y. \u0026amp; Liu, L.-M. Prediction of sumoylation sites in proteins using linear discriminant analysis. \u003cem\u003eGene\u003c/em\u003e\u003cstrong\u003e576\u003c/strong\u003e, 99\u0026ndash;104 (2016).\u003c/li\u003e\n\u003cli\u003eBergstra, J. \u0026amp; Bengio, Y. Random Search for Hyper-Parameter Optimization. \u003cem\u003eJ Mach Learn Res\u003c/em\u003e\u003cstrong\u003e13\u003c/strong\u003e, 281\u0026ndash;305 (2012).\u003c/li\u003e\n\u003cli\u003eDietterich, T. G. Ensemble Methods in Machine Learning. in \u003cem\u003eMultiple Classifier Systems\u003c/em\u003e vol. 1857 1\u0026ndash;15 (Springer Berlin Heidelberg, 2000).\u003c/li\u003e\n\u003cli\u003eThe UniProt Consortium \u003cem\u003eet al.\u003c/em\u003e UniProt: the Universal Protein Knowledgebase in 2023. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e51\u003c/strong\u003e, D523\u0026ndash;D531 (2023).\u003c/li\u003e\n\u003cli\u003eKlausen, M. S. \u003cem\u003eet al.\u003c/em\u003e NetSurfP‐2.0: Improved prediction of protein structural features by integrated deep learning. \u003cem\u003eProteins Struct. Funct. Bioinforma.\u003c/em\u003e\u003cstrong\u003e87\u003c/strong\u003e, 520\u0026ndash;527 (2019).\u003c/li\u003e\n\u003cli\u003eRosario, F. J. \u003cem\u003eet al.\u003c/em\u003e Placental Remote Control of Fetal Metabolism: Trophoblast mTOR Signaling Regulates Liver IGFBP-1 Phosphorylation and IGF-1 Bioavailability. \u003cem\u003eInt. J. Mol. Sci.\u003c/em\u003e\u003cstrong\u003e24\u003c/strong\u003e, 7273 (2023).\u003c/li\u003e\n\u003cli\u003eMacLean, B. \u003cem\u003eet al.\u003c/em\u003e Skyline: an open source document editor for creating and analyzing targeted proteomics experiments. \u003cem\u003eBioinformatics\u003c/em\u003e\u003cstrong\u003e26\u003c/strong\u003e, 966\u0026ndash;968 (2010).\u003c/li\u003e\n\u003cli\u003eHagberg, A. A., Schult, D. A. \u0026amp; Swart, P. J. Exploring Network Structure, Dynamics, and Function using NetworkX. in \u003cem\u003eProceedings of the 7th Python in Science Conference\u003c/em\u003e (eds. Varoquaux, G., Vaught, T. \u0026amp; Millman, J.) 11\u0026ndash;15 (2008).\u003c/li\u003e\n\u003cli\u003eTate, J. G. \u003cem\u003eet al.\u003c/em\u003e COSMIC: the Catalogue Of Somatic Mutations In Cancer. \u003cem\u003eNucleic Acids Res.\u003c/em\u003e\u003cstrong\u003e47\u003c/strong\u003e, D941\u0026ndash;D947 (2019).\u003c/li\u003e\n\u003cli\u003eShannon, P. \u003cem\u003eet al.\u003c/em\u003e Cytoscape: A Software Environment for Integrated Models of Biomolecular Interaction Networks. \u003cem\u003eGenome Res.\u003c/em\u003e\u003cstrong\u003e13\u003c/strong\u003e, 2498\u0026ndash;2504 (2003).\u003c/li\u003e\n\u003cli\u003eMorris, J. H. \u003cem\u003eet al.\u003c/em\u003e clusterMaker: a multi-algorithm clustering plugin for Cytoscape. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e\u003cstrong\u003e12\u003c/strong\u003e, 436 (2011).\u003c/li\u003e\n\u003cli\u003eBader, G. D. \u0026amp; Hogue, C. W. An automated method for finding molecular complexes in large protein interaction networks. \u003cem\u003eBMC Bioinformatics\u003c/em\u003e\u003cstrong\u003e4\u003c/strong\u003e, 2 (2003).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"nature-portfolio","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"","title":"Nature Portfolio","twitterHandle":"","acdcEnabled":false,"dfaEnabled":false,"editorialSystem":"ejp","reportingPortfolio":"","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-3771179/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3771179/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe exploration of post-translational modifications (PTMs) within the proteome is pivotal for advancing disease and cancer therapeutics. However, identifying genuine PTM sites amid numerous candidates is challenging. Integrating machine learning (ML) models with high-throughput in vitro peptide synthesis has introduced an ML-hybrid search methodology, enhancing enzyme-substrate selection prediction. In this study we have developed a ML-hybrid search methodology to better predict enzyme-substrate selection. This model achieved a 37.4% experimentally validated precision, unveiling 885 SET8 candidate methylation sites in the human proteome—marking a 19-fold accuracy increase over traditional \u003cem\u003ein vitro\u003c/em\u003e methods. Mass spectrometry analysis confirmed the methylation status of several sites, responding positively to SET8 overexpression in mammalian cells. This approach to substrate discovery has also shed light on the changing SET8-regulated substrate network in breast cancer, revealing a predicted gain (376) and loss (62) of substrates due to missense mutations. By unraveling enzyme selection features, this approach offers transformative potential, revolutionizing enzyme-substrate discovery across diverse PTMs while capturing crucial biochemical substrate properties.\u003c/p\u003e","manuscriptTitle":"Machine learning-based exploration of enzyme-substrate networks: SET8-mediated methyllysine and its changing impact within cancer proteomes","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-02-27 09:25:46","doi":"10.21203/rs.3.rs-3771179/v1","editorialEvents":[],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"communications-chemistry","isNatureJournal":true,"hasQc":false,"allowDirectSubmit":false,"externalIdentity":"commschem","sideBox":"Learn more about [Communications Chemistry](http://www.nature.com/commschem/)","snPcode":"","submissionUrl":"","title":"Communications Chemistry","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"ejp","reportingPortfolio":"Communications Series","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"d4021d8b-5597-4341-a365-7e925c8c7d69","owner":[],"postedDate":"February 27th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":27666906,"name":"Biological sciences/Biochemistry/Enzymes/Transferases"},{"id":27666907,"name":"Biological sciences/Computational biology and bioinformatics/Protein function predictions"},{"id":27666908,"name":"Biological sciences/Computational biology and bioinformatics/Machine learning"}],"tags":[],"updatedAt":"2025-11-08T08:10:39+00:00","versionOfRecord":{"articleIdentity":"rs-3771179","link":"https://doi.org/10.1038/s42004-025-01717-6","journal":{"identity":"communications-chemistry","isVorOnly":false,"title":"Communications Chemistry"},"publishedOn":"2025-11-07 05:00:00","publishedOnDateReadable":"November 7th, 2025"},"versionCreatedAt":"2024-02-27 09:25:46","video":"","vorDoi":"10.1038/s42004-025-01717-6","vorDoiUrl":"https://doi.org/10.1038/s42004-025-01717-6","workflowStages":[]},"version":"v1","identity":"rs-3771179","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3771179","identity":"rs-3771179","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.