{"paper_id":"120fd875-758a-4c8e-b6c4-d124f1e9ec59","body_text":"Leveraging learned representations and multitask learning for lysine\nmethylation site discovery\nFrançois Charih1,2,3, Mullen Boulter², Kyle K. Biggar2,3 and James R. Green1,3\n¹ Department of Systems and Computer Engineering, Carleton University, Ottawa, ON, Canada\n² Institute of Biochemistry, Department of Biology, Carleton University, Ottawa, ON, Canada\n³ NuvoBio Corp., Ottawa, ON, Canada\nCorrespondence to: François Charih (\nfrancois@charih.ca)\nKeywords: Lysine methylation, lysine methylome, deep learning, transformers, multitask learning\nAbstract\nLysine methylation is a dynamic and reversible post-translational modification of proteins carried out\nby lysine methyltransferase enzymes. The role of this modification in epigenetics and gene regulation is\nrelatively well understood, but our understanding of the extent and the role of lysine methylation of non-\nhistone substrates remains fairly limited. Several lysine methyltransferases which methylate non-histone\nsubstrates are overexpressed in a number of cancers and are believed to be key drivers of cancer progression.\nThere is great incentive to identify the lysine methylome, as this is a key step in identifying drug targets.\nWhile numerous computational models have been developed in the last decade to identify novel lysine\nmethylation sites, the accuracy of these model has been modest, leaving much room for improvement. In this\nwork, we leverage the most recent advancements in deep learning and present a transformer-based model\nfor lysine methylation site prediction which achieves state-of-the-art accuracy. In addition, we show that\nother post-translational modifications of lysine are informative and that multitask learning is an effective\nway to integrate this prior knowledge into our lysine methylation site predictor, MethylSight 2.0. Finally,\nwe validate our model by means of mass spectrometry experiments and identify 68 novel lysine methylation\nsites. This work constitutes another contribution towards the completion of a comprehensive map of the\nlysine methylome.\nIntroduction\nLysine methylation extends far beyond the realm of histone proteins and that it may be more prevalent\nthan previously believed [1]. Studies have uncovered the involvement of non-histone lysine methylation in\noncogenic processes chemoresistance and cancer cell proliferation [2], [3], [4], making it a very attractive\ntarget for anti-cancer therapies. Therapies targeting non-histone KMTs are emerging, with some even having\nreached the clinical trial stage with promising signs of efficacy [5], [6]. For instance, Tazemetostat, an\ninhibitor of EZH2, was trialed and received approval for the treatment of blood and solid malignancies [5].\nEZH2 promotes tumorigenesis in glioblastoma and prostate cancer models via STAT3 methylation [7], [8]\nand in diffuse large B-cell and follicular lymphomas via methylation of the PRC2 complex [5]. Considering\nthat many cancers are driven by KMT overexpression, uncovering the human lysine methylome and the\nassociated KMTs/KDMs would have profound implications in drug discovery and facilitate the identification\nof drug targets for therapeutic intervention, but would also broaden our general understanding of the lysine\nmethylation-dependent biological processes.\nIdentification of novel lysine methylation sites is possible through experimental methods such as mass\nspectrometry (MS) [9]. That said, the process is sufficiently resource-intensive that identifying new sites at\nthe proteome scale experimentally is logistically impractical. For that reason, a number of machine learning\nprediction models have been developed over the years to tackle the lysine methylation prediction problem.\n1\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nThe rationale behind these models is that computation can guide our efforts so that time and resources can\nbe invested on validating the most promising potential lysine methylation sites.\nA wide range of machine learning models trained on lysine methylation datasets to predict lysine methy-\nlation sites from sequence only have been published over the last two decades. Most of them have relied\non the use of “traditional” machine learning algorithms such as support vector machines (SVMs) or random\nforests (RFs) and human-crafted numerical features. The first predictor, MeMO [10], was built using an SVM\nclassifier using what is now refered to as “one-hot” encoding for a 15 amino acids lysine-centered window.\nThe training set used in that study consisted of a total of 156 positive lysine methylation sites, which\nrepresents only an infinitesimal fraction of the space of all lysine-centered 15-mers (156/1420 ≈ 10−19% of\nthe space of possible windows). More recent models used different strategies to represent lysine-centered\nwindows. For example, iMethyl-PseAAC [11] used a SVM model in conjunction with a representation the\nauthors termed “pseudo amino acid composition” (PseAAC). This representation combines evolutionary in-\nformation from the position-specific scoring matrix (PSSM), physicochemical properties of individual amino\nacids from the AAIndex [12], and disorder scores to generate a 346-dimensional feature vector. Another well-\ncited method is GPS-MSP (Group-based Prediction System Methyl-group Specific Predictor), an algorithm\npublished in 2017 [13], was trained on 1,521 methyllysines sites and used an alignment-based custom scoring\nfunction to measure window similarity in conjunction with 𝑘-means clustering (unsupervised learning) to\npredict not only methylation sites, but also the methylation state (mono-, di-, or tri-), an ambitious task given\nthe scarcity of data available to train models to this level of granularity.\nAttempts to leverage deep learning methods to address the challenging task of accurately predicting the\nlysine methylome remain scarce as of the time of writing. Recently, Spadaro et al. applied convolutional\nneural networks (CNNs) to representations combining phylogenetic, physicochemical, structural, and binary\nencodings to predict lysine methylation sites [14]. The PTM-Mamba model  [15] makes use of the Mamba\narchitecture, an attention-free state-space model with architectural optimizations designed to enhance com-\nputational efficiency over long sequences. More specifically, PTM-Mamba fuses Mamba-generated sequence\nembeddings with embeddings generated by the ESM-2 protein language model (pLM) to predict a wide array\nof PTMs which, for some reason, do not include lysine methylation.\nBepler and Berger have shown that multitask language models better capture the semantic organization of\nproteins by training a bidirectional long short-term memory (LSTM) to complete three tasks simultaneously:\nmasked language modeling, residue-residue contact prediction, and structural similarity prediction [16].\nHere, we hypothesize that knowledge about other PTMs of lysine residues such as acetylation, ubiquitina-\ntion, sumoylation, and phosphorylation can inform models designed to predict methylation. This hypothesis\narises as a result of the fact that enzymes catalyzing PTMs are known to compete to modify the same lysine\nresidues and that this competition constitutes an additional layer of protein function modulation. Consis-\ntent with this fact, the PhosphoSitePlus  [17], a well-maintained database which curates post-translational\nmodifications of proteins, lists 31,845 lysine residues in human proteins with annotations for two or more\npost-translational modifications (PTMs) (Figure 1). Of the 4,958 validated methylation sites in the database,\n2,375 sites (48%) are subject to at least one other known modification.\nIt is a reasonable supposition that there might be some partial overlap in the physico-chemical environments\nthat promote lysine methylation and other PTMs, such as solvent accessibility, surrounding amino acid\nproperties, and steric (spatial) constraints. As a result, the lysine methylation prediction problem is amenable\nto a multitask learning formulation, wherein lysine methylation prediction is one task among several PTM\nprediction tasks, and that jointly training a single model on several such tasks concomitantly could lead to\nbetter prediction accuracy through knowledge transfer.\nTo our knowledge, this idea has been exploited only once for the uncommon propionylation PTM of lysines\n[18]. In that work, the authors trained a recurrent neural network (RNN) on lysine malonylation sites and\nfine-tuned it using a dataset of lysine propionylation sites to extract features that are then fed into an SVM\nclassifier. That work did exploit transfer learning, but did not use a multitask learning scheme, as the training\n2\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nFigure 1. Co-occurence of common post-translational modification of lysines in the Phospho-\nSitePlus database\nThe overlap of four major PTMs of lysines among human proteins are shown in this Venn diagram. These\nnumbers were computed using the 10/17/24 update of the PhosphoSitePlus database  [17]. Many yet-to-be-\ndiscovered modifications remain to be deposited. Of the four major modifications of lysines, methylation is\nthe one with the fewest annotations.\nwas not performed for both tasks (i.e. propionylation and malonylation prediction) simultaneously. A model\nwas trained for the malonylation prediction task first, and subsequently trained to predict propionylation\nsites, of which there were fewer known instances (431 as opposed to 9,584).\nCurrently, no one has proposed a lysine methylation site prediction model that leverages 1) state-of-the-art\nneural architecture, i.e. the transformer and 2) domain adaptation by means of transfer learning techniques\nsuch as multitask learning. In addition, very few groups have proven with in vitro experiments that the\nestimates of accuracy of their models translate into the lab upon deployment.\nIn this work, we address these opportunities to develop a more accurate and robust lysine methylation\npredictor. Our contributions are summarized below.\nContribution 1 - Improved prediction accuracy: We bootstrap embeddings generated with pLMs\ntrained on millions of protein sequences to train a model, MethylSight 2.0, which produces dramatically\nmore accurate predictions than previous lysine methylation predictors;\nContribution 2 - Use of multitask learning to enhance model accuracy:\n We demonstrate that\ntraining a transformer-based neural network architecture with a multitask learning strategy can lead to\nmore accurate predictions;\nContribution 3 - Experimental validation:  We show, by means of MS  validation experiments\nperformed on sites predicted to be methylated by our model, MethylSight 2.0, that the accuracy our\nmodel translates experimentally and identify 68 novel lysine methylation sites;\nContribution 4 - Bioinformatics analysis of the MethylSight 2.0 predicted lysine methylome:\nWe deploy MethylSight 2.0 at the proteome scale to identify previously unknown methylation sites and\nconduct analyses to identify biological processes wherein lysine methylation sites are overrepresented.\n3\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nTable 1. Composition of the PhosphoSitePlus dataset (human lysines)\nModification Important functions Unique\nproteins\nPositive\nsites\nNegative\nsites (low\nconfidence)\nMethylation Chromatin and gene expression regulation (histones), signaling,\nenzyme (in)activation [19] 2,751 4,966 157,385\nAcetylation Protein stability, regulation of PPIs and protein-DNA interac-\ntions (histones) [20] 7,047 22,547 333,013\nSumoylation Alteration of molecular interactions of substrate through addi-\ntion/hiding interaction surfaces [21] 2,646 8,391 124,274\nUbiquitination Regulation of protein degradation, autophagy, protein traffick-\ning [22] 11,712 96,545 377,547\nMethods\nLysine methylation dataset preparation\nWith the intent of creating a dataset for multitask learning involving multiple PTM types, we retrieved PTM\ndata for lysines occurring in human proteins by mining the PhosphoSitePlus database [17] (10/17/24 update)\nfor methylation, ubiquitination, sumoylation, and acetylation, which are all known to be modifications of\nlysines. The composition of the PhosphoSitePlus dataset is summarized in Table 1.\nGathering positive lysine modification data is relatively straightforward, but identifying the “negative” sites\nto complete the training set needed to train a binary classifier is more arduous. It is difficult to ascertain with\nconfidence that a lysine not known to be modified never\n is. In reality, sites taken to be “negative” may corre-\nspond to yet-to-be discovered methylation sites. Some groups simply take as negative training examples sites\nnot known to be modified [23], [24], which we argue might disproportionately bias the learning algorithm\ntowards making negative predictions. For this reason, it is typical to only use a subset of sites without PTM\nannotations as negative in PTM site prediction challenges, following some heuristics [25]. A typical practice\nis to only label as negatives unlabeled sites that occur within a protein containing a known modified site\nelsewhere [10], [25], [26], [27], [28]. To train and test MethylSight [25], a lab-validated SVM-based model that\nachieved state-of-the-art performance upon publication, we applied two additional criteria to label potential\nlysine methylation as “negative” in addition to the latter. More specifically, sites were considered “negative”\nin the training set if they were not known to be substrates for another PTM (ubiquitination, sumoylation,\nor acetylation) and\n were predicted to be buried (relative solvent accessibility factor < 0.2, as predicted with\nNetSurfP v1.0  [29]). In this work, we applied the same curation method, but used an updated version of\nNetSurfP (v3.0 [30]). This approach allowed us to build a dataset with high-confidence negatives.\nTo address the issue of redundancy in the data, which could cause overrepresentation of certain patterns\nin the dataset and data leakage, i.e. similar patterns in the training and test data, we clustered the windows\nbased on sequence identity at a similarity threshold of 70% with CD-HIT  [31], as done previously [25], and\nselected one representative from each cluster at random, favouring a positive representative (methylation\nsite) if one occurred within a cluster. Finally, 20% of the non-redundant sites were set aside for testing.\nWe used an identical workflow to assemble the training sets for ubiquitination, acetylation, and sumoylation\nsites required for multitask learning, but did not set any data aside for testing, given that we are only\ninterested in methylation site prediction.\nThe composition of the final dataset is presented in Table 2.\n4\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nFigure 2. Preparation of a high-quality lysine modification dataset\nTo create the dataset used as part of this study, we sourced data from the PhosphoSitePlus database (for PTM\nannotations) and the UniProt database (for protein sequences). Only proteins with at least one methylated\nlysine residue were included in the dataset. Exposed residues of unknown status and/or having an anotation\nfor another PTM were discarded, while the remaining lysines not known to be methylated were selected\nto make up the negative training data. The redundancy in the dataset was reduced with CD-HIT, using a\nwindow size of 31 for clustering and a 70% identity threshold. A blind test set was created by setting aside\n20% of this data.\nTable 2. Composition of the high-confidence dataset used to train and test the models\nMethylation Ubiquitination Acetylation Sumoylation\nTraining set (pos/neg) 2,415/15,699 68,539/58,314 15,791/51,498 5,737/15,862\nValidation set (pos/neg) 604/3,925\nTest set (pos/neg) 755/4,906\nPre-trained protein-language model embeddings\npLMs have been shown to generate rich embeddings that capture physicochemical, phylogenetic and\nstructural information that are extremely useful for a variety of downstream tasks, including structure\nprediction [32], [33], property prediction (e.g. viscosity [34], stability [35], etc.), localization prediction [35],\n[36], and peptide binder design [37], [38], [39], to cite a few.\nGiven that these representations were learned on massive collections of protein sequences and performed\nwell on these tasks, we hypothesized that they may also contain useful information for the prediction of\nlysine methylation sites. Moreover, these embeddings capture more context about the potential methylation\nsites than traditional human-engineered representations. They consider a large portion (or all) of the protein,\ndepending on the pLM’s context length (window size) and the protein length.\n5\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nWe leveraged representations learned by three state-of-the-art foundational model pLMs: ProtT5  [40],\nESM-2 [32], and Ankh [36] (Table 3) and fed the human proteome to all three models to generate embeddings\nfor each lysine in the training and test sets – and to later predict the comprehensive human lysine methylome.\nTable 3. Foundational protein language models used to embed potential lysine methylation sites\nModel Architecture Version Embedding\ndimension\nParameters\n(approx.) Training strategy Training data\nProtT5 [40] Encoder-\ndecoder\nProtT5-XL-\nBFD 1,024 3B\n1-gram random\nmasking with\ndemasking\nBFD (pre-training;\n~2.1B sequences)\nand UniRef50 (fine-\ntuning; ~45M\nsequences)\nESM-2 [32] Encoder-only\nESM-2-\nT33-650M-\nUR50D\n1,280 650M\n1-gram random\nmasking with\ndemasking\nUniRef50+90 (~65M\nsequences)\nAnkh [36] Encoder-\ndecoder Ankh Large 1,536 1B\n1-gram random\nmasking with full\nsequence\nreconstruction\nUniRef50 (~45M\nsequences)\nTraining multilayer perceptrons leveraging pLM-generated embeddings\nUsing Optuna  [41], we trained 50 MLPs  (MLPs) on the embedding vectors of lysine sites extracted with\nProtT5, ESM-2 and Ankh, sampling at random the learning rate, number and width of of hidden layers, and\nthe dropout rate used for regularization. We selected as our final models the ones with the highest validation\narea under the precision-recall curve (AUPRC). All models were trained using PyTorch [42] with the Adam\noptimizer [43] with a batch size of 64, using binary cross-entropy as the loss function:\nℒCE = − 1\n𝑁 ∑\n𝑁\n𝑖=1\n𝑦𝑖 log(̂𝑦𝑖) + (1 − 𝑦𝑖) log(1 −̂𝑦𝑖)\nwhere 𝑦𝑖 = 1 if the site 𝑖 is methylated and 𝑦𝑖 = 0 otherwise, while ̂𝑦𝑖 ∈ [0, 1] is the predicted probability\nof that the site 𝑖 is methylated.\nWe selected the model using an early stopping strategy, using the validation loss to monitor for overfitting.\nWe repeated this procedure using the embeddings generated by all three aforementioned pLMs. In addition,\nwe trained MLPs on a “combined” representation resulting from a concatenation of all three embeddings,\nfor a total of four final MLP models.\nTraining a transformer model\nTo determine whether training a transformer-based model could further improve the quality of the predic-\ntions, we implemented a model which leverages this architecture.\nWe used a context size of 31 amino acids, the ProtT5 embeddings as representations for the individual amino\nacids in the sequence (\ni.e. the tokens), and padded with a zero-filled 1,024-D vector, if the lysine site was\ntoo close to the end of the protein chain (Figure 3A). To capture the positional information of the individual\ntokens, we used the cannonical positional embedding strategy described in  [44]. We used 4 heads in each\nattention block. A schematic representation of this architecture is shown in Figure 3B. The Adam optimizer\nwas used, but with a batch size of 128.\nSimilarly to the approach used to train the MLPs, we conducted hyperparameter tuning in a randomized\nfashion and varied the number of encoder transformer blocks, the learning rate, the number of hidden layers\nin the classification module (i.e. the dense layers that follow the transformer layers), and the width of the\n“embedding layer” and trained a total of 50 models.\nWe used the same loss function and early stopping stratedy as for the MLPs to select the final model for each\nrun. The final transformer architecture selected was the one with the highest AUPRC on the validation set.\n6\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nMultitask learning with a transformer model\nTo investigate whether a multitask learning strategy could enhance the quality of the predictions, we\nenriched our training set with sites and their annotation for the three other PTMs of interest: acetylation,\nubiquitination, and sumoylation.\nIn this context, the “tasks” consist in predicting the four different PTMs of interest. We do not know or can’t\nassume with a satisfying level of certainty the true label for each task for all sites. For example, we may\nknow that a site is acetylated, but not know whether it is also ubiquitinated. Consequently, we chose to not\nassociate each instance or site with 4 labels. Instead, each instance in the dataset is a site associated with a\nPTM and a label associated to that site and PTM. Consequently, a given site may appear up to 4 times in the\ntraining set, in the specific case where a label for each PTM is known.\nWe implemented another transformer model where the last transformer block is followed by a flattening\nlayer whose output is sent to one of four classification heads, depending on the task, i.e.\n prediction of\nmethylation, acetylation, etc. (Figure 3C). Each classification head is designed to predict whether a site is\nsubjected to the corresponding PTM . For each instance in the training set, we only probe the probability\noutput by the head corresponding to the PTM (“task”) associated with the instance.\nWe use a custom batch sampling strategy to train the model wherein all instances in a batch are associated\nwith only one of the four PTM. This ensures that the loss over a batch is only used to update the parameters\nof the classification head associated with the PTM (and the upstream parameters), but not the three classi-\nfication heads which are used for the other tasks. In other terms, we use partial parameter sharing, i.e. only\nthe parameters in the transformer layers and upstream are shared across tasks.\nGiven that the methylation sites are vastly outnumbered by other sites, we multiply the loss for methylation\nbatches by a factor 𝛾 in order to produce larger updates for the shared model weights when methylation\nsites are misclassified relative to misclassified instances of other PTMs. We tried 𝛾 ∈ {1, 13.5, 20}, 1:13.5\nbeing approximatively the methylation-to-other PTMs ratio.\nThe loss function for the multitask learning strategy effectively takes the form:\nℒ = 𝛾ℒCE, me + ℒCE, ub + ℒCE, ac + ℒCE, su\nwhere\nℒCE, 𝑡 = {− 1\n𝑁 ∑𝑁\n𝑖=1 𝑦𝑖 log(̂𝑦𝑖) + (1 − 𝑦𝑖) log(1 −̂𝑦𝑖) , if batch is for task 𝑡\n0 , otherwise\nThe rest of the model selection was done as for the transformer model without the multitask learning training\nstrategy described in the previous section. We henceforth refer to this model as MethylSight 2.0.\nEstimation of the expected imbalance\nTo accurately estimate the precision of MethylSight  2.0 upon deployment on the human proteome, an\nestimate of the class imbalance is required. The human proteome in UniProt/Swiss-Prot database (2025_01\nrelease) [45] comprises 654,185 lysines residues in 20,417 unique proteins, of which an unknown fraction can\nbe methylated under specific biological circumstances such as in response to a biological event, in a stage of\ndevelopment, or in specific tissue types.\nBerryhill et al. [46] published a study which provides some useful insight into the ratio of methylated\nto unmethylated lysines observable through mass spectrometry experiments. In their study, they assessed\nthe sequence bias of commercially available pan-methyllysine antibodies and performed global profiling\nof lysine methylation in HEK293T (human embryonic kidney) and U2OS (human osteosarcoma) cells with\nsamples enriched with combinations of less biased anti-Kme1, anti-Kme2, and anti-Kme3 antibodies and\ntheir combinations. They identified a total of 5,089 lysine methylation sites evenly distributed through the\nproteome, of which 4,862 are novel.\n7\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nFigure 3. Transformer model architectures for methylation site prediction without and with\nmultitask learning\n(A) The inputs to our transformer-based models are the ProtT5 embeddings extracted from the full protein,\nwith a context window of 31, centered around the lysine residue of interest. When a lysine residue is too\nclose to an end of the protein sequence, null embeddings (<PAD>) are appended to complete the context\nwindow. (B)\n Architecture of our vanilla transformer-based model for methylation probability prediction.\n(C) Modified transformer architecture designed to enable a multitask learning strategy. More specifically,\nafter the flattening layer, instance representations are sent to one of four PTM-specific classification heads,\ndepending on the task associated with the individual instances.\nUsing the data collected in this study, we made the assumption that the estimated the imbalance ratio of\nmethylated-to-unmethylated lysines detectable via mass spectrometry without and following enrichment with\nthe antibodies currently in use to be roughly 1:36. This corresponds to the ratio of methylated lysines to\nlysines not found to be methylated in the proteins that were pulled down in the samples (i.e. with at least\none epitope for the anti-Kme antibodies used). It is difficult to speculate about what lysines are or are not\nmethylated in proteins that were not pulled down, so we only estimate what one may observe in a global\nprofiling experiment with mass spectrometry. We use this ratio to evaluate the anticipated precision of\nMethylSight-2.0, when coupled with a mass spectrometry experiment.\nThis figure is an approximation derived from samples extracted from two specific cell types, and as such, it\nmay not apply uniformly across all tissue types and across the entire proteome.\n8\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nSelection of predicted methylation sites for in vitro validation\nWe subsequently sought to estimate the actual precision of MethylSight  2.0 upon deployment onto the\nhuman proteome. To achieve this, we selected 100 sites predicted to be methylated by MethylSight 2.0, but\nwhich were not known methylation sites. Using a conservative threshold on predicted methylation proba-\nbility (i.e. PCPr1:36 = 0.75), we sampled 50 sites at random from each of the following two sets:\n1. Set 1: Exposed lysine residues known to be acetylated, ubiquitinated, and/or sumoylated;\n2. Set 2: Exposed lysine residues with no known modification.\nFurthermore, under the hypothesis supported by the phenomena of PTM  competition that these other\nmodifications provide useful information for the identification of novel lysine methylation sites, one would\nexpect to detect more methylation events within sites sampled from Set 1 than within sites sampled from\nSet 2. To allow for this comparison, we ensured that the methylation probabilities were similarly distributed\nin both samples.\nValidation of predicted methylation sites via mass spectrometry\nUsing the Pyteomics package for Python [47], we generated an isolation list tabulating the peptide fragments\nand their mass-to-charge ratios for the +2, +3 and +4 charged states and for all four methylation states\n(null, mono-, di-, tri-), resulting in a total of 1,200 predicted peaks. The isolation list can be found in the\nsupplementary materials.\nparallel reaction monitoring mass spectrometry   (PRM-MS) experiments were conducted at the John L.\nHolmes Mass Spectrometry Facility at the University of Ottawa with a Q Exactive™ Plus Hybrid Quadrupole-\nOrbitrap™ mass spectrometer, using the aforementioned isolation list to guide the scanning. The results\nwere obtained from a single injection of a Thermo-Fisher Pierce™ HeLa protein digest standard. We opted\nto monitor for methylation in this sample because it is guaranteed to have a low missed tryptic cleavage rate\n(<10%) and minimal methionine oxidation and lysine carbamylation (<10%). Furthermore, these standards\nare thoroughly tested for quality, which improves the reproducibility of the results.\nResults and discussion\nMethylSight 2.0: performance and benchmarking\nThe predictive performances on the blind test set of the final models are tabulated in Table 4.\nAll models trained as part of this work perform significantly better on the blind test set than methods\npublished previously (GPS-MSP [13], MethylSight 1.0 [25], Met-Predictor [23]), in spite of the fact that these\nmethods have likely encountered some of the sites in our test set during training, which would confer them\nan unfair advantage. In fact, our worst performing model, a MLP using lysine embeddings produce by the\nESM-2 pLM was associated with a 25% improvement over the state of the art (SOTA) in terms of both AUPRC\nand precision at 0.5 recall (Pr@0.5Re).\nAmong the MLP models we trained on the four representations produced by the three pLMs considered, the\nmodel trained on ESM-2 embeddings performed noticeably worse relative to the other three representations\nwhich produced similar results, though the model trained on the combined embeddings appears to have a\nperformance modestly superior to that of the MLPs trained on ProtT5 and Ankh embeddings. The relative\nperformance of the different representations is consistent with the sizes of the pLMs, ProtT5 being 3 times\nthe size of Ankh, and 4.6 times the size of ESM-2, in terms of learnable parameter counts. This observation\nis consistent with evidence that pLMs performance scales with model size following a power law [48], [49].\nThe use of a transformer architecture trained “from scratch”specifically for the task of predicting lysine\nmethylation prediction did allow for an improvement in performance over the use of lysine embeddings\ngenerated by all three foundational pLMs in MLPs. Our best single-task transformer model slightly outper-\nformed the best MLP (i.e. the MLP trained on concatenated ProtT5-Ankh-ESM-2 embeddings), using AUPRC\nas a metric. However, it produced a more significant improvement in precision at the a 50% recall threshold\nof nearly 10%. This showcases the power of the self-attention mechanism, as attention layer parameters in\n9\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nTable 4. Prediction accuracy of our models and publicly available methods on the blind test set\n(1:6.5 imbalance)\nMethod AUPRC AUROC Pr@0.5Re\nTransformer (w/ multitask learning) 0.672 0.874 0.769\nTransformer (w/o multitask learning) 0.642 0.867 0.734\nMLP (Combined embeddings) 0.624 0.859 0.637\nMLP (ProtT5 embeddings) 0.620 0.854 0.620\nMLP (Ankh embeddings) 0.611 0.848 0.609\nThis work\nMLP (ESM-2 embeddings) 0.517 0.829 0.520\nMethylSight 1.0 [25] 0.267 0.677 0.242\nMet-Predictor [23] 0.254 0.646 0.221Previous SOTA\nGPS-MSP [13] 0.179 0.588 0.175\nour transformer model were learned specifically for the lysine methylation prediction task as opposed to the\nmore general masked language modeling objective, as is was case for the foundational models.\nLooking at the PRCs (PRCs) assessing model performance on our blind test set (Figure 4A), we see that\nimplementing a multitask learning strategy leveraging knowledge about other PTMs  is useful, as our best\nmodel (hyperparameters tabulated in Table 5) outperforms all other models over the entire range of possible\nrecall values (or operating thresholds). However, this advantage is anticipated to dissipate at higher recall\nvalues assuming that the true class imbalance of methylation sites to non-methylation sites in the proteome\nis higher than that in the test set (e.g., 1:100 as opposed to 1:6.5; Figure 4C). This suggests that knowledge of\nother PTMs of lysines can indeed transfer to lysine methylation. This is consistent with our initial hypothesis\nas well as with the establish phenomenon of PTM competition wherein lysine modifying enzyme “compete”\nto modify specific lysine residues [50], [51], [52].\nIdentification and validation of novel lysine methylation sites\nThe PRM-MS experiments on a HeLa cell lysate guided with an isolation list listing tryptic peptides corre-\nsponding to MethylSight 2.0-predicted methylation sites revealed a significant number of hits (Figure 5A). In\nfact, 68 of the 100 sites predicted to be methylated produced transitions consistent with methylated peptides\nwith a fair or better level of confidence. In contrast, for only 6 of the sites could evidence of the unmethylated\npeptide and no evidence of methylation be found. The results were inconclusive for 26 peptides, i.e.\n the\npeptide could not be detected, neither in a unmethylated state, nor in a methylated state. In the worst case\nwhere we consider all inconclusive sites to be negative, MethylSight 2.0 would achieve a precision of 68%.\nThe precision becomes 91.9% if we discard sites for which no transitions could be detected from the analysis.\nIt is likely that the precision we would have observed, if all results had been conclusive, would lie somewhere\nwithin that range. This suggests that an imbalance ratio of 1:36 is a reasonable estimate.\nInterestingly, we found that more methylation sites were found for lysines that were not known to be\notherwise modified (39) than were found for lysines known to be acetylated, ubiquitinated, or sumoylated\n(29). This observation contradicts our initial hypothesis that more methylation sites would be detected\namong proteins which are known to be otherwise modified, because of their “modifiable” character. One\nplausible explanation is that one or more of these other modifications might have in fact competed with\nmethylation, thus reducing its abundance and preventing its detection.\nWe choose here to highlight two sites occuring within proteins of high biological and clinical significance\nwhich produced transitions unequivocally consistent with methylation: the 𝛾 subunit of the Eukaryotic\nelongation factor 1 (eEF1\n𝛾) and DnaJ homolog subfamily B member 11 (DNAJB11).\neEF1𝛾 is one of four subunit of the eEF1 complex, along with the 𝛼, 𝛽 and 𝛿 subunits [53]. Though not be-\nlieved to be a catalytically active member of the complex [54], eEF1𝛾 is believed to act as a structural scafford\nfor the 𝛼 subunit and to facilitate the complex’s function of bringing aminoacyl tRNAs to the ribosome for\n10\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nFigure 4. Precision-recall curves and receiver operating characteristic curves of the models on the\nindependent test set\nThe precision-recall and receiver operating characteristic curves for the best performing models and/or\ntraining strategies were computed using a blind and independent test set of potential lysine methylation\nsites not seen during training. (A) The precision-recall curves are shown for the test set (1:6.5 imbalance). (B)\nThe associated ROC curves are shown. (C)\n Prevalence-corrected precision assuming a true 1:36 imbalance\nbetween methylated and unmethylated sites provides more pessimistic estimate of performance. (D) The\nperformance metrics are show for two different imbalance ratios (1:6.5 and 1:36).\ntranslation [53]. Beyond its role in the elongation factor 1 complex, eEF1𝛾 is predicted to have “moonlighting”\nroles and be involved in several other biological processes including viral ribonucleic acid (RNA) transcrip-\ntion, oxidative stress response, cytoskeleton-membrane linking, and cellular trafficking [55]. MethylSight 2.0\ncorrectly predicted the methylation of K428 in eEF1𝛾  (Figure 5B), which is located within the C-terminus\ndomain of the protein. While the N-terminus end of eEF1 𝛾 is known to interact with eEF1𝛼 and anchor\nit into the complex, the role of the C-terminus end of the protein is not a clearly understood, aside from\nthe fact that it is highly conserved and protease resistant [56]. There is some evidence that interaction with\nthe 𝛽 subunit of the eEF1 complex may occur at the C-terminus domain  [57]. Therefore, it is plausible that\nthe methylation status of eEF1𝛾 -K428 could modulate the interaction between these the 𝛽 and 𝛾 subunits.\nGiven that eEF1𝛾  has been found to interact with actin [58], it is possible that methylation of K428 could\n11\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nTable 5. Hyperparameters used to train the most accurate model (MethylSight 2.0)\nParameter Value\nLearning rate 8 × 10−7\nNumber of epochs 100\nWeight factor of methylation loss (𝛾) 20\nBatch size 128\nNumber of transformer blocks 2\nHeads per self-attention layer 4\nWidth of the first (pre-attention) dense layer 1,600\nWidths of the dense layers (classification heads) 1,797; 1,803; 338; 493\nEmbedding dimension 1,024\nDropout rate 0.15\nmodulate this interaction and influence cytoskeleton dynamics if it indeed occurs at the C-terminus. eEF1𝛾’s\nclinical significance is supported by the observation that it overexpressed in gastric carcinoma  [59], colon\nadenocarcinoma [60], and pancreatic cancer [61], in all likelihood so that cancer cells can satisfy the higher\ntranslation load required to adapt and proliferate. Altogether, our observation that eEF1𝛾  is methylated on\nK438 combined with its involvement in key biological processes and cancer warrants further investigations\ninto the biological significance of the modification.\nClear transitions consistent with the presence of a methyllysine were also recorded for the tryptic peptide\nfragment from DNAJB11 containing K66 (Figure 5C). DNAJB11 is a member of the DNAJ (or HSP40) subclass\nof family of heat shock proteins which all share a J-domain. The role of this highly conserved domain\nis to stimulate the hydrolysis of ATP by chaperones in the HSP70 protein family whose main function is\nstabilize or restore the native protein conformation of potentially misfolded client proteins under cellular\nstress [62]. Proteins in the DNAJ family have been implicated in tumor progression and metastasis  [62].\nDNAJB11 in particular has been overexpressed in pancreatic cancer, where exosomal DNAJB11 regulates\nexpression of EGFR activates the MAPK pathway [63] and in liver cancer, by preventing alpha-1-antitrypsin\ndegradation [64]. In contrast, low DNAJB11 messenger RNA  (mRNA) levels appear to be correlated with\nworse outcomes in thyroid carcinoma [65]. In addition to it role in several cancers, research has shown that\nphosphorylation of T188 in DNAJB11 can reduce the aggregation of 𝛼-synuclein in Parkinson’s disease [66].\nAs such, an association between K66 methylation and Parkinson’s disease is possible, either directly via\nan unknown mechanism, or indirectly, through the modulation of T188 phosphorylation via cross-talk, for\nexample. In all cases, since it is located within the J-domain, it is likely that the methylation of K66 is\nbiologically significant, and this result provides a rationale for the characterization of K66.\nThe predicted human lysine methylome\nUsing conservative settings (i.e. at PCPr1:36 = 0.75), MethylSight  2.0 identified a total of 62,567 lysine\nmethylation sites within 13,791 different proteins in the human proteome (Figure 6A). Based on our perfor-\nmance assessment and an estimated imbalance ratio of 1:36, we anticipate that out of these predicted sites,\n∼47,000 (75%) are actual methylation sites. This figure, alone, is significantly higher than the ∼30,000 sites\npredicted with 63% precision by MethylSight 1.0. Given that the estimated recall of MethylSight 2.0 under\nthese conditions is ~30%, we estimate the size of the lysine methylome at ~155,000.\nStatistically significant enrichment in several biological process and molecular function GO terms were\nidentified (Figure 6B). Of note, we observed overrepresentation in methylated proteins of terms related to\ntranslation, ribosomal biogenesis and structure, and cytoskeleton structure. Enrichment of related terms in\nmethylated proteins was also observed within the MethylSight 1.0-predicted lysine methylome  [25], and is\nalso supported by the literature. For instance, methylation of several lysine residues in the human elongation\nfactor 1A (eEF1A) is known to regulate ribosome biogenesis and actin cytoskeleton dynamics, among\nothers [67]. The methylation of elongation factor eEF2 by the lysine methyltransferase (KMT) FAM86A is\n12\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nFigure 5. MethylSight 2.0-enabled discovery of novel lysine methylation sites with PRM-MS\n(A) High-level overview of the methodology employed to validate 100 methylation sites identified with\nMethylSight 2.0 and distribution of the compiled results. (B) Transitions for the tryptic peptide containing\neEF1𝛾-K428 (EYFSWEGAFQHVGK). The mass-to-charge ratios of the y ions are shown. The transitions of the\nnull, mono-, di-, and tri-methylation states are plotted separately for clarity, because the measured intensity\nvaries in scale. (C) Transitions for the tryptic peptide containing DNAJB11-K66 (NPDDPQAQEK).\nalso known to regulate translation dynamics  [68]. The literature also supports the involvement of lysine\nmethylation in cytoskeleton regulation. The role of lysine methylation in cytoskeleton regulation is well-\nestablished [69]. The α-tubulin cytoskeletal protein is known to be tri-methylated by SETD2 (and acetylated)\non K40, and loss of methylation has been associated with “catastrophic microtubule defects” which impair\nDNA repair mechanisms [70] and cell cycle progression  [71]. Recently, methylation of BCAR3 on K334 by\nSMYD2 in breast cancer was shown to enhance lamellipodia dynamics of breast cancer cells through the\nrecruitment of Formin-like proteins which accelerate actin polymerization and facilitate cell proliferation\nand metastasis in vivo [72].\nInterestingly, our SAFE analysis of methylated proteins mapped onto the HuRI human interactome\n(Figure 6C) also shows subnetworks where overrepresentation of GO terms related to RNA processing and\nregulation and cytoskeleton organization is observed.\nIn Figure 6D, we illustrate the predicted prevalence of lysine methylation events in a subset of the actin\ncytoskeleton pathway (KEGG: hsa04810). Most proteins in this important subset of the pathway contain\nseveral lysine methylation sites.\n13\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nFigure 6. The lysine methylome as predicted by MethylSight 2.0\n(A) Distribution of methylation probability for all 654,185 lysines in the human proteome. The sites predicted\nto be methylated only while operating under conservative settings (PCPr = 0.75) are shown in green. (B)\nTop 10 overrepresented gene ontology (GO) terms for the biological process (purple) and molecular function\n(blue) categories. Overrepresentation is statistically significant ( 𝑝-value < 0.05; Fisher’s exact test with\nBonferroni correction for multiple testing). (C) Visual representation of the spatial analysis of functional\nenrichment (SAFE) analysis results; i.e. functional domains within the interaction network of methylated\nproteins and the most frequent words present in the GO terms associated with proteins in the domains. (D)\nSubset of the actin cytoskeleton regulation pathway (KEGG: hsa04810). Proteins are colored on a white-to-\nred scale, with darker shades indicating a higher degree of methylation.\n14\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nCancer mutations associated with a predicted loss of methylation at a proximal site\nMethylSight 2.0 was used to predict the impact of the 1,000 most frequent missense mutations in the COSMIC\ndatabase [73] on the predicted methylation score of lysines within the mutant protein. The COSMIC database\ncontains a curated list of somatic mutations and their impact in cancer.\nWe found that scores were relatively insensitive to these mutations except in select cases where the mutations\nwere in close proximity with a lysine. Interestingly, we only observed predicted loss of methylation (using\nthe conservative operating threshold). We detected 25 lysines that were no longer predicted to be methylated\nand associated with a decrease in methylation score ≥ 0.02 in presence of a mutation (Figure 7A).\nThe most striking loss of predicted methylation are associated with the RhoAG17V and RhoAG17E mutations.\nIn these mutants, the neighbouring K18 is no longer predicted to be a methylation site. K18, like the mutated\nG17, is located amidst the GDP binding pocket (Figure 7B).\nWhile methylation of RhoA-K18 has never been – to the best of our knowledge – confirmed experimentally,\nit is possible that methylated K18 could modulate RhoA activity. In fact, though it is believed that mutations\nin G17 impair GDP/GTP binding  [74], it is not known exactly how this mutation impairs binding of GDP/\nGTP. Given the proximity of K18 to the GTP/GDP binding site, it is not implausible that alteration of the\nmethylation status could alter RhoA’s ability to bind and release GTP/GDP or coordinate Mg2+.\nTaken together, this provides an interesting avenue for further investigation so as to determine whether 1\nRhoA-K18 is a true methylation site and 2  loss of methylation occurs in these mutants in vitro, and 3  this\nloss of methylation directly impacts GTP/GDP binding.\nPerformance of MethylSight 2.0 on non-human proteins\nWe deployed MethylSight 2.0 on the set of all known non-human lysine methylation sites catalogued in the\nPhosphoSitePlus database (360 sites). Interestingly, MethylSight 2.0 achieved a recall of 0.383 when applied to\nthese sites (at PCPr1:36 = 0.75). This is on the same order as its predicted recall (i.e. 0.299) on human lysines\noperating at the same decision threshold. This suggests that methylation sites in non-human organisms\nshare homology with sites in human proteins.\nFigure 7. Predicted loss of methylation in oncogenic proteins\n(A) Changes in predicted methylation score resulting in predicted loss of methylation associated with the\n1,000 most commonly reported cancer-associated single amino acid mutations in the Catalog of Somatic\nMutations reported in the Cancer (COSMIC) database. Shown are score changes of at least 2%. \n(B) Position\nof K18 (in purple) in the X-ray structure of RhoA (PDB: 1DPF). The GDP co-factor is in beige with the two\nphosphate groups in orange.\n15\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nThe observation that MethylSight 2.0 achieved a better recall than predicted on this set of sites provides\nsome evidence that the imbalance between methylated and non-methylated lysines could actually be lower\nthan the one we estimated (\ni.e. 1:36) at the proteome scale, but further experiments would be required to\nconfirm this.\nThe MethylSight 2.0 server\nTo make MethylSight 2.0 accessible to the broader community, we implemented a web server (Figure 8)\nwhich can be accessed at https://methylsight2.cubic.ca. The server is easy-to-use and allows users to select\nthe operating threshold (precision and recall) that suits them best, depending on the application.\nThe web server processes individual protein sequences. Users interested in batch predictions may run\nMethylSight  2.0 as a standalone software on their own hardware. The Methylsight  2.0 source code,\nmodel weights, and instructions on how to use the software can be found on GitHub: https://github.com/\nGreenCUBIC/MethylSight2.\nConclusion\nOur work demonstrates that deep representations learned by pLMs trained on tens of millions of protein are\nrich in information directly relevant for the task of lysine methylation site prediction and significantly im-\nprove the quality of the predictions. In fact, using these deep representations, we successfully trained model\nthat achieved more than double the AUPRC of previous models trained on human-engineered descriptors\nof protein sequences, such as those generated by ProtDCal  [75] alongside the SVM-based MethylSight 1.0\npredictor. We also showed that leveraging knowledge about other PTMs by means of a multitask learning\nstrategy can further enhance the quality of the predictions. To the best of our knowledge, our model\nMethylSight 2.0 is the first lysine methylation prediction model to leverage pLM-generated representations\nFigure 8. MethylSight 2.0 web server\n(A) The user can either provide the UniProt accession ID or the sequence of the human protein of interest. If\nthe protein is from another organism or is a non-canonical human protein (e.g., isoform or mutant), the user\nmust provide a FASTA-formatted sequence through a text box or uploading the equivalent FASTA file. (B)\nOnce the results are available, a precision-recall curve computed for MethylSight 2.0 on our test set with an\nassumed real imbalance ratio of 1:36 is presented. This provides the user with an visual interpretation of the\noperating threshold and the optimal threshold which can be tuned with a slider. The results are presented\nas a table which can be downloaded in CSV format.\n16\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\nand to employ a multitask learning strategy to extract knowledge from useful data that would otherwise be\nleft unexploited.\nIn an effort to validate the predictions made by MethylSight 2.0 and show that it can successfully guide lysine\nmethylation site discovery in vivo, we performed a validation PRM-MS experiment guided by MethylSight 2.0\npredictions on a high-quality HeLa cell lysate. We uncovered 68 previously unidentified lysine methylation\nsites, some of which among proteins of high biological and/or therapeutic relevance, further showing the\nusefulness of our model as a drug target identification tool.\nApplying MethylSight 2.0 to the human proteome provides insight into the extent of the lysine methylome.\nIn fact, our analyses of the MethylSight  2.0-predicted lysine methylome suggests that the number of\nmethyllysines in the human proteome may be even larger than previously believed (∼50,000  [25]), though\nthis number remains difficult to estimate, given that our validation experiment was limited to 100 sites,\nand it seems unlikely that so few sites would be representative of the entire human proteome. Additional\nvalidation experiments would be required to get a more accurate portrait of the human lysine methylome\nlandscape.\nLysine methylation is a highly dynamic process which competes with several other PTMs, and a given lysine\nmay be methylated to varying degrees under different conditions, e.g., during development, in response to an\nenvironmental trigger, or in specific tissue types. Given that our model was trained in a tissue-blind fashion,\ni.e.\n positive sites in the training set were known to be methylated in at least one tissue type, we anticipate\nthat validation experiments on methylation sites identified with MethylSight 2.0 may need to be performed\nin more than one cell type. There could be significant value in training a cell-specific lysine methylation\npredictor that could predict whether a lysine is methylated in a given tissue type\n, but training such a model\nwould require tissue-specific datasets which are currently not publicly available.\nFurthermore, in this study, we did not distinguish between the three possible methylation states. It is impor-\ntant to acknowledge that different methylation states can be associated with different – sometimes opposite\n– phenotypes. Several other groups have attempted to address this problem [13], [23], [76], but with limited\nsuccess. Further research is needed to design an accurate predictor of mono-, di-, and trimethylated lysines.\nFinally, MethylSight 2.0 does not attempt to associate a KMT with sites predicted to be methylated. This\nchallenge is of prime importance, as therapeutic intervention would normally target the KMT or lysine\ndemethylase (KDM) responsible for the modification. However, it is non-trivial given the scarcity of data\nfor some KMTs, which in certain cases only have a few dozen known substrates or fewer. At the time of\nwriting, we are aware of only one KMT -specific model for SET8  [77] which could be used in conjunction\nwith MethylSight 2.0.\nTaken together, this work constitutes a significant contribution toward the elucidation of the human lysine\nmethylome. In addition, MethylSight  2.0 can be deployed in a targeted fashion to determine whether\nlysines within proteins involved in a pathway of interest are probable methylation site. It therefore affords\nexperimentalists with a tool which can help formulate rational hypotheses, guide experiments, and cut costs\nthrough prioritization of candidates for validation experiments.\nBibliography\n[1] K. K. Biggar and S. S.-C. Li, “Non-Histone Protein Methylation as a Regulator of Cellular Signalling\nand Function, ” Nature Reviews Molecular Cell Biology, vol. 16, no. 1, pp. 5–17, Jan. 2015, doi: 10.1038/\nnrm3915\n.\n[2] S. M. Carlson and O. Gozani, “Nonhistone Lysine Methylation in the Regulation of Cancer Pathways, ”\nCold Spring Harbor Perspectives in Medicine, vol. 6, no. 11, p. a26435, Nov. 2016, doi: 10.1101/\ncshperspect.a026435\n.\n[3] D. Han et al., “Lysine Methylation of Transcription Factors in Cancer, ” Cell Death & Disease, vol. 10, no.\n4, p. 290, Apr. 2019, doi: 10.1038/s41419-019-1524-2.\n17\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n[4] M. Huang et al., “Methylation Modification of Non-Histone Proteins in Breast Cancer: An Emerging\nTargeted Therapeutic Strategy, ” Pharmacological Research, vol. 208, p. 107354, Oct. 2024, doi: 10.1016/\nj.phrs.2024.107354.\n[5] R. Straining and W. Eighmy, “Tazemetostat: EZH2 Inhibitor, ” Journal of the Advanced Practitioner in\nOncology, vol. 13, no. 2, p. 158, Mar. 2022, doi: 10.6004/jadpro.2022.13.2.7.\n[6] A. Feoli, M. Viviano, A. Cipriano, C. Milite, S. Castellano, and G. Sbardella, “Lysine Methyltransferase\nInhibitors: Where We Are Now, ” RSC Chemical Biology, vol. 3, no. 4, pp. 359–406, 2022, doi: 10.1039/\nD1CB00196E.\n[7] K. Xu et al., “EZH2 Oncogenic Activity in Castration-Resistant Prostate Cancer Cells Is Polycomb-\nindependent, ” Science (New York, N.Y.), vol. 338, no. 6113, pp. 1465–1469, Dec. 2012, doi: 10.1126/\nscience.1227604.\n[8] E. Kim et al., “Phosphorylation of EZH2 Activates STAT3 Signaling via STAT3 Methylation and\nPromotes Tumorigenicity of Glioblastoma Stem-like Cells, ” Cancer Cell, vol. 23, no. 6, pp. 839–852, Jun.\n2013, doi: 10.1016/j.ccr.2013.04.008.\n[9] S. Lanouette, V. Mongeon, D. Figeys, and J.-F. Couture, “The Functional Diversity of Protein Lysine\nMethylation, ” Molecular Systems Biology, vol. 10, no. 4, p. 724, Apr. 2014, doi: 10.1002/msb.134974.\n[10] H. Chen, Y. Xue, N. Huang, X. Yao, and Z. Sun, “MeMo: A Web Tool for Prediction of Protein\nMethylation Modifications, ” Nucleic Acids Research, vol. 34, no. suppl_2, pp. W249–W253, Jul. 2006, doi:\n10.1093/nar/gkl233\n.\n[11] W.-R. Qiu, X. Xiao, W.-Z. Lin, and K.-C. Chou, “iMethyl-PseAAC: Identification of Protein Methylation\nSites via a Pseudo Amino Acid Composition Approach, ” BioMed Research International\n, vol. 2014, p.\n947416, May 2014, doi: 10.1155/2014/947416.\n[12] S. Kawashima, P. Pokarowski, M. Pokarowska, A. Kolinski, T. Katayama, and M. Kanehisa, “AAindex:\nAmino Acid Index Database, Progress Report 2008, ” Nucleic Acids Research, vol. 36, no. Database issue,\npp. D202–205, Jan. 2008, doi: \n10.1093/nar/gkm998.\n[13] W. Deng, Y. Wang, L. Ma, Y. Zhang, S. Ullah, and Y. Xue, “Computational Prediction of Methylation\nTypes of Covalently Modified Lysine and Arginine Residues in Proteins, ” Briefings in Bioinformatics,\nvol. 18, no. 4, pp. 647–658, Jul. 2017, doi: \n10.1093/bib/bbw041.\n[14] A. Spadaro, A. Sharma, and I. Dehzangi, “Predicting Lysine Methylation Sites Using a Convolutional\nNeural Network, ” Methods, vol. 226, pp. 127–132, Jun. 2024, doi: 10.1016/j.ymeth.2024.04.007.\n[15] Z. Peng, B. Schussheim, and P. Chatterjee, “PTM-Mamba: A PTM-Aware Protein Language Model with\nBidirectional Gated Mamba Blocks, ” Feb. 2024. doi: 10.1101/2024.02.28.581983.\n[16] T. Bepler and B. Berger, “Learning the Protein Language: Evolution, Structure, and Function, ” Cell\nSystems, vol. 12, no. 6, pp. 654–669, Jun. 2021, doi: 10.1016/j.cels.2021.05.017.\n[17] P. V. Hornbeck, B. Zhang, B. Murray, J. M. Kornhauser, V. Latham, and E. Skrzypek, “PhosphoSitePlus,\n2014: Mutations, PTMs and Recalibrations, ” Nucleic Acids Research, vol. 43, no. Database issue, pp.\nD512–520, Jan. 2015, doi: \n10.1093/nar/gku1267.\n[18] A. Li, Y. Deng, Y. Tan, and M. Chen, “A Transfer Learning-Based Approach for Lysine Propionylation\nPrediction, ” Frontiers in Physiology, vol. 12, Apr. 2021, doi: 10.3389/fphys.2021.658633.\n[19] V. Lukinović, A. G. Casanova, G. S. Roth, F. Chuffart, and N. Reynoird, “Lysine Methyltransferases\nSignaling: Histones Are Just the Tip of the Iceberg, ” Current Protein and Peptide Science, vol. 21, no. 7,\npp. 655–674, Jul. 2020, doi: \n10.2174/1871527319666200102101608.\n18\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n[20] T. Narita, B. T. Weinert, and C. Choudhary, “Functions and Mechanisms of Non-Histone Protein\nAcetylation, ” Nature Reviews Molecular Cell Biology, vol. 20, no. 3, pp. 156–174, Mar. 2019, doi: 10.1038/\ns41580-018-0081-3.\n[21] R. Geiss-Friedlander and F. Melchior, “Concepts in Sumoylation: A Decade On, ” Nature Reviews Mole-\ncular Cell Biology, vol. 8, no. 12, pp. 947–956, Dec. 2007, doi: 10.1038/nrm2293.\n[22] R. B. Damgaard, “The Ubiquitin System: From Cell Signalling to Disease Biology and New Therapeutic\nOpportunities, ” Cell Death & Differentiation, vol. 28, no. 2, pp. 423–426, Feb. 2021, doi: 10.1038/\ns41418-020-00703-w\n.\n[23] W. Zheng, Q. Wuyun, M. Cheng, G. Hu, and Y. Zhang, “Two-Level Protein Methylation Prediction\nUsing Structure Model-Based Features, ” Scientific Reports\n, vol. 10, no. 1, p. 6008, Apr. 2020, doi: 10.1038/\ns41598-020-62883-2.\n[24] P. Shrestha, J. Kandel, H. Tayara, and K. T. Chong, “DL-SPhos: Prediction of Serine Phosphorylation\nSites Using Transformer Language Model, ” Computers in Biology and Medicine, vol. 169, p. 107925, Feb.\n2024, doi: \n10.1016/j.compbiomed.2024.107925.\n[25] K. K. Biggar et al., “Proteome-Wide Prediction of Lysine Methylation Leads to Identification of H2BK43\nMethylation and Outlines the Potential Methyllysine Proteome, ” Cell Reports, vol. 32, no. 2, p. 107896,\nJul. 2020, doi: 10.1016/j.celrep.2020.107896.\n[26] Y. Xue, F. Zhou, M. Zhu, K. Ahmed, G. Chen, and X. Yao, “GPS: A Comprehensive Www Server for\nPhosphorylation Sites Prediction, ” Nucleic Acids Research, vol. 33, no. suppl_2, pp. W184–W187, Jul.\n2005, doi: \n10.1093/nar/gki393.\n[27] S.-P. Shi, J.-D. Qiu, X.-Y. Sun, S.-B. Suo, S.-Y. Huang, and R.-P. Liang, “PMeS: Prediction of Methylation\nSites Based on Enhanced Feature Encoding Scheme, ” PLOS ONE, vol. 7, no. 6, p. e38772, Jun. 2012, doi:\n10.1371/journal.pone.0038772\n.\n[28] Y. Shi, Y. Guo, Y. Hu, and M. Li, “Position-Specific Prediction of Methylation Sites from Sequence\nConservation Based on Information Theory, ” Scientific Reports, vol. 5, no. 1, p. 12403, Dec. 2015, doi:\n10.1038/srep12403\n.\n[29] B. Petersen, T. Petersen, P. Andersen, M. Nielsen, and C. Lundegaard, “A Generic Method for Assign-\nment of Reliability Scores Applied to Solvent Accessibility Predictions, ” BMC Structural Biology\n, vol. 9,\nno. 1, p. 51, 2009, doi: 10.1186/1472-6807-9-51.\n[30] M. H. Høie et al., “NetSurfP-3.0: Accurate and Fast Prediction of Protein Structural Features by Protein\nLanguage Models and Deep Learning, ” Nucleic Acids Research, vol. 50, no. W1, pp. W510–W515, Jul.\n2022, doi: 10.1093/nar/gkac439.\n[31] L. Fu, B. Niu, Z. Zhu, S. Wu, and W. Li, “CD-HIT: Accelerated for Clustering the next-Generation\nSequencing Data, ” Bioinformatics, vol. 28, no. 23, pp. 3150–3152, Dec. 2012, doi: 10.1093/bioinformatics/\nbts565\n.\n[32] Z. Lin et al., “Evolutionary-Scale Prediction of Atomic-Level Protein Structure with a Language Model, ”\nScience, vol. 379, no. 6637, pp. 1123–1130, Mar. 2023, doi: 10.1126/science.ade2574.\n[33] O. Avraham, T. Tsaban, Z. Ben-Aharon, L. Tsaban, and O. Schueler-Furman, “Protein Language Models\nCan Capture Protein Quaternary State, ” BMC bioinformatics, vol. 24, no. 1, p. 433, Nov. 2023, doi:\n10.1186/s12859-023-05549-w\n.\n[34] X. Hao and L. Fan, “ProtT5 and Random Forests-Based Viscosity Prediction Method for Therapeutic\nmAbs, ” European Journal of Pharmaceutical Sciences, vol. 194, p. 106705, Mar. 2024, doi: 10.1016/\nj.ejps.2024.106705\n.\n19\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n[35] R. Schmirler, M. Heinzinger, and B. Rost, “Fine-Tuning Protein Language Models Boosts Predictions\nacross Diverse Tasks, ” Nature Communications, vol. 15, no. 1, p. 7407, Aug. 2024, doi: 10.1038/\ns41467-024-51844-2.\n[36] A. Elnaggar et al., “Ankh: Optimized Protein Language Model Unlocks General-Purpose Modelling, ”\nJan. 2023. doi: 10.1101/2023.01.16.524265.\n[37] G. Brixi et al., “SaLT&PepPr Is an Interface-Predicting Language Model for Designing Peptide-\nGuided Protein Degraders, ” Communications Biology, vol. 6, no. 1, p. 1081, Oct. 2023, doi: 10.1038/\ns42003-023-05464-z.\n[38] S. Bhat et al., “De Novo Generation and Prioritization of Target-Binding Peptide Motifs from Sequence\nAlone, ” Jun. 2023. doi: 10.1101/2023.06.26.546591.\n[39] T. Chen et al., “PepMLM: Target Sequence-Conditioned Generation of Peptide Binders via Masked\nLanguage Modeling, ” 2024.\n[40] A. Elnaggar et al., “ProtTrans: Towards Cracking the Language of Lifes Code Through Self-Supervised\nDeep Learning and High Performance Computing, ” IEEE Transactions on Pattern Analysis and Machine\nIntelligence, p. 1, 2021, doi: 10.1109/TPAMI.2021.3095381.\n[41] T. Akiba, S. Sano, T. Yanase, T. Ohta, and M. Koyama, “Optuna: A Next-generation Hyperparameter\nOptimization Framework, ” in Proceedings of the 25th ACM SIGKDD International Conference on Knowl-\nedge Discovery & Data Mining, in KDD '19. New York, NY, USA: Association for Computing Machinery,\nJul. 2019, pp. 2623–2631. doi: \n10.1145/3292500.3330701.\n[42] A. Paszke et al., “PyTorch: An Imperative Style, High-Performance Deep Learning Library, ” arXiv, 2019,\ndoi: 10.48550/arxiv.1912.01703.\n[43] D. P. Kingma and J. Ba, “Adam: A Method for Stochastic Optimization, ” no. arXiv:1412.6980. arXiv,\n2015. doi: 10.48550/arXiv.1412.6980.\n[44] A. Vaswani et al., “Attention Is All You Need, ” arXiv:1706.03762 [cs], Dec. 2017.\n[45] The UniProt Consortium, “UniProt: The Universal Protein Knowledgebase in 2025, ” Nucleic Acids\nResearch, vol. 53, no. D1, pp. D609–D617, Jan. 2025, doi: 10.1093/nar/gkae1010.\n[46] C. A. Berryhill et al., “Global Lysine Methylome Profiling Using Systematically Characterized Affinity\nReagents, ” Scientific Reports, vol. 13, no. 1, p. 377, Jan. 2023, doi: 10.1038/s41598-022-27175-x.\n[47] L. I. Levitsky, J. A. Klein, M. V. Ivanov, and M. V. Gorshkov, “Pyteomics 4.0: Five Years of Development\nof a Python Proteomics Framework, ” Journal of Proteome Research, vol. 18, no. 2, pp. 709–714, 2018, doi:\n10.1021/acs.jproteome.8b00717\n.\n[48] Q. Fournier, R. M. Vernon, A. Van Der Sloot, B. Schulz, S. Chandar, and C. J. Langmead, “Protein\nLanguage Models: Is Scaling Necessary?. ” Molecular Biology, Sep. 2024. doi: 10.1101/2024.09.23.614603.\n[49] X. Cheng, B. Chen, P. Li, J. Gong, J. Tang, and L. Song, “Training Compute-Optimal Protein Language\nModels. ” Bioinformatics, Jun. 2024. doi: 10.1101/2024.06.06.597716.\n[50] M. Leutert, S. W. Entwisle, and J. Villén, “Decoding Post-Translational Modification Crosstalk With\nProteomics, ” Molecular & Cellular Proteomics : MCP, vol. 20, p. 100129, Jul. 2021, doi: 10.1016/\nj.mcpro.2021.100129\n.\n[51] A. H. Shukri, V. Lukinović, F. Charih, and K. K. Biggar, “Unraveling the Battle for Lysine: A Review of\nthe Competition among Post-Translational Modifications, ” Biochimica et Biophysica Acta (BBA) - Gene\nRegulatory Mechanisms\n, vol. 1866, no. 4, p. 194990, Dec. 2023, doi: 10.1016/j.bbagrm.2023.194990.\n20\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n[52] J. M. Lee, H. M. Hammarén, M. M. Savitski, and S. H. Baek, “Control of Protein Stability by Post-\nTranslational Modifications, ” Nature Communications, vol. 14, no. 1, p. 201, Jan. 2023, doi: 10.1038/\ns41467-023-35795-8.\n[53] A. N. Sasikumar, W. B. Perez, and T. G. Kinzy, “The Many Roles of the Eukaryotic Elongation Factor 1\nComplex, ” WIREs RNA, vol. 3, no. 4, pp. 543–555, 2012, doi: 10.1002/wrna.1118.\n[54] O. Olarewaju, P. A. Ortiz, W. Q. Chowdhury, I. Chatterjee, and T. G. Kinzy, “The Translation Elongation\nFactor eEF1B Plays a Role in the Oxidative Stress Response Pathway, ” RNA biology, vol. 1, no. 2, pp.\n89–94, Jul. 2004, doi: \n10.4161/rna.1.2.1033.\n[55] B. S. Negrutskii, V. F. Shalak, O. V. Novosylna, L. V. Porubleva, D. M. Lozhko, and A. V. El'skaya, “The\neEF1 Family of Mammalian Translation Elongation Factors, ” BBA Advances, vol. 3, p. 100067, Jan. 2023,\ndoi: \n10.1016/j.bbadva.2022.100067.\n[56] S. Vanwetswinkel et al., “Solution Structure of the 162 Residue C-terminal Domain of Human Elonga-\ntion Factor 1B\\gamma*”, Journal of Biological Chemistry, vol. 278, no. 44, pp. 43443–43451, Oct. 2003,\ndoi: 10.1074/jbc.M306031200.\n[57] I. Achilonu, N. Elebo, B. Hlabano, G. R. Owen, M. Papathanasopoulos, and H. W. Dirr, “An Update\non the Biophysical Character of the Human Eukaryotic Elongation Factor 1 Beta: Perspectives from\nInteraction with Elongation Factor 1 Gamma, ” Journal of Molecular Recognition, vol. 31, no. 7, p. e2708,\n2018, doi: \n10.1002/jmr.2708.\n[58] O. A. Olatona, S. R. Choudhury, R. Kresman, and C. A. Heckman, “Candidate Proteins Interacting with\nCytoskeleton in Cells from the Basal Airway Epithelium in Vitro, ” Frontiers in Molecular Biosciences,\nvol. 11, p. 1423503, Jul. 2024, doi: \n10.3389/fmolb.2024.1423503.\n[59] K. Mimori, M. Mori, S. Tanaka, T. Akiyoshi, and K. Sugimachi, “The Overexpression of Elongation\nFactor 1 Gamma mRNA in Gastric Carcinoma, ” Cancer, vol. 75, no. 6Suppl, pp. 1446–1449, Mar. 1995,\ndoi: \n10.1002/1097-0142(19950315)75:6+<1446::aid-cncr2820751509>3.0.co;2-p.\n[60] K. Chi, D. V. Jones, and M. L. Frazier, “Expression of an Elongation Factor 1 Gamma-Related\nSequence in Adenocarcinomas of the Colon, ” Gastroenterology\n, vol. 103, no. 1, pp. 98–102, Jul. 1992,\ndoi: 10.1016/0016-5085(92)91101-9.\n[61] Y. Lew, D. V. Jones, W. M. Mars, D. Evans, D. Byrd, and M. L. Frazier, “Expression of Elongation Factor-1\nGamma-Related Sequence in Human Pancreatic Cancer, ” Pancreas, vol. 7, no. 2, pp. 144–152, 1992,\ndoi: \n10.1097/00006676-199203000-00003.\n[62] H.-Y. Kim and S. Hong, “Multi-Faceted Roles of DNAJB Protein in Cancer Metastasis and Clinical\nImplications, ” International Journal of Molecular Sciences, vol. 23, no. 23, p. 14970, Nov. 2022, doi:\n10.3390/ijms232314970\n.\n[63] P. Liu, F. Zu, H. Chen, X. Yin, and X. Tan, “Exosomal DNAJB11 Promotes the Development of Pancreatic\nCancer by Modulating the EGFR/MAPK Pathway, ” Cellular & Molecular Biology Letters, vol. 27, p. 87,\nOct. 2022, doi: \n10.1186/s11658-022-00390-0.\n[64] J. Pan, D. Cao, and J. Gong, “The Endoplasmic Reticulum Co-Chaperone ERdj3/DNAJB11 Promotes\nHepatocellular Carcinoma Progression through Suppressing AATZ Degradation, ” Future Oncology\n, vol.\n14, no. 29, pp. 3001–3013, Dec. 2018, doi: 10.2217/fon-2018-0401.\n[65] R. Sun, L. Yang, Y. Wang, Y. Zhang, J. Ke, and D. Zhao, “DNAJB11 Predicts a Poor Prognosis\nand Is Associated with Immune Infiltration in Thyroid Carcinoma: A Bioinformatics Analysis, ”\nJournal of International Medical Research, vol. 49, no. 11, p. 03000605211053722, Nov. 2021, doi:\n10.1177/03000605211053722\n.\n21\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n[66] H.-Y. Chen, C.-Y. Liao, H. Li, Y.-C. Ke, C.-H. Lin, and S.-C. Teng, “ATM-mediated Co-Chaperone\nDNAJB11 Phosphorylation Facilitates \\alpha-Synuclein Folding upon DNA Double-Stranded Breaks”,\nNAR Molecular Medicine\n, vol. 1, no. 2, p. ugae7, Apr. 2024, doi: 10.1093/narmme/ugae007.\n[67] J. J. Hamey, B. Wienert, K. G. R. Quinlan, and M. R. Wilkins, “METTL21B Is a Novel Human Lysine\nMethyltransferase of Translation Elongation Factor 1A: Discovery by CRISPR/Cas9 Knockout*, ” Mole-\ncular & Cellular Proteomics\n, vol. 16, no. 12, pp. 2229–2242, Dec. 2017, doi: 10.1074/mcp.M116.066308.\n[68] J. W. Francis et al., “FAM86A Methylation of eEF2 Links mRNA Translation Elongation to Tumorige-\nnesis, ” Molecular Cell, vol. 84, no. 9, pp. 1753–1763, May 2024, doi: 10.1016/j.molcel.2024.02.037.\n[69] C. Michail, F. Rodrigues Lima, M. Viguier, and F. Deshayes, “Structure and Function of the Lysine\nMethyltransferase SETD2 in Cancer: From Histones to Cytoskeleton, ” Neoplasia\n, vol. 59, p. 101090, Jan.\n2025, doi: 10.1016/j.neo.2024.101090.\n[70] I. Y. Park et al., “Dual Chromatin and Cytoskeletal Remodeling by SETD2, ” Cell, vol. 166, no. 4, pp. 950–\n962, Aug. 2016, doi: 10.1016/j.cell.2016.07.005.\n[71] L. X. Li and X. Li, “Epigenetically Mediated Ciliogenesis and Cell Cycle Regulation, and Their Trans-\nlational Potential, ” Cells, vol. 10, no. 7, p. 1662, Jul. 2021, doi: 10.3390/cells10071662.\n[72] A. G. Casanova et al., “Cytoskeleton Remodeling Induced by SMYD2 Methyltransferase Drives Breast\nCancer Metastasis, ” Cell Discovery, vol. 10, no. 1, pp. 1–22, Jan. 2024, doi: 10.1038/s41421-023-00644-x.\n[73] Z. Sondka et al., “COSMIC:~A Curated Database of Somatic Variants and Clinical Data for Cancer, ”\nNucleic Acids Research, vol. 52, no. D1, pp. D1210–D1217, Jan. 2024, doi: 10.1093/nar/gkad986.\n[74] M. Sakata-Yanagimoto et al., “Somatic RHOA Mutation in Angioimmunoblastic T Cell Lymphoma, ”\nNature Genetics, vol. 46, no. 2, pp. 171–175, Feb. 2014, doi: 10.1038/ng.2872.\n[75] Y. B. Ruiz-Blanco, W. Paz, J. Green, and Y. Marrero-Ponce, “ProtDCal: A Program to Compute General-\nPurpose-Numerical Descriptors for Sequences and 3D-structures of Proteins, ” BMC Bioinformatics, vol.\n16, no. 1, p. 162, May 2015, doi: \n10.1186/s12859-015-0586-0.\n[76] Z. Ju, J.-Z. Cao, and H. Gu, “iLM-2L: A Two-Level Predictor for Identifying Protein Lysine Methylation\ns General\nPseAAC, ” Journal of Theoretical Biology, vol. 385, pp. 50–57, Nov. 2015, doi: 10.1016/j.jtbi.2015.07.030.\n[77] N. Ridgeway et al., “A Machine Learning-Enhanced Methodology for Accurate Prediction of Enzyme-\nSubstrate Selection, ” Nature Communications (In Review), 2024.\nAuthor contributions\nFrançois Charih: Conceptualization, Methodology, Software, Investigation, Formal analysis, Data Curation,\nVisualization, Writing - Original Draft, Mullen Boulter: Formal analysis, Kyle K. Biggar: Conceptual-\nization, Resources, Formal analysis, Writing - Review & Editing, Funding acquisition, James R. Green:\nConceptualization, Writing - Review & Editing, Funding acquisition\nAll authors approved of the manuscript.\nFunding\nThis research was funded by the National Science and Engineering Research Council (NSERC) Canada\nDiscovery grant awarded to Kyle K. Biggar (RGPIN-2023-04651) and James R. Green (RGPIN-2021-04184).\nConflicts of interest\nThe authors have no conflicts of interest to disclose.\n22\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint \n\n.CC-BY 4.0 International licenseperpetuity. It is made available under a \npreprint (which was not certified by peer review) is the author/funder, who has granted bioRxiv a license to display the preprint in \nThe copyright holder for thisthis version posted September 1, 2025. ; https://doi.org/10.1101/2025.08.27.672583doi: bioRxiv preprint","source_license":"CC-BY-4.0","license_restricted":false}