CMC-WDTK: CpG methylation change prediction by a weight-sharing dual-branch Transformer-Kolmogorov–Arnold network model

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-14

CMC-WDTK, a deep learning model integrating sequence and SNV information, accurately predicts DNA methylation changes between sequences and identified a motif promoting methylation.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-14 · read from full text

The paper proposes CMC-WDTK, a deep learning framework that predicts allele-associated CpG methylation changes by integrating DNA sequences flanking CpG sites with adjacent single-nucleotide variants, using a weight-sharing dual-branch Transformer encoder plus a Kolmogorov–Arnold network to model high-order nonlinear relationships. The model was trained and evaluated on eight ASM-derived datasets spanning four tissue types and four cancer cell lines, where allele-specific methylation change labels were defined from bisulfite sequencing read counts mapped to reference versus alternate alleles. CMC-WDTK achieved AUC values greater than 0.8 on all datasets and identified a repeated cytosine–guanine sequence motif linked to increased methylation, with claimed strong cross-dataset generalizability; the explicit caveat is that the work is based on ASM sites identified from cancer BS-Seq datasets rather than other contexts. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Differential methylation is a key epigenetic process contributing to cancer development. Most DNA methylation prediction methods rely on DNA sequences from the background reference genome, neglecting individual genetic variation, which limits their ability to capture methylation differences. To address this, we propose CMC-WDTK, a deep learning framework that combines a weight-sharing dual-branch Transformer with a Kolmogorov‒Arnold network (KAN) to integrate sequences flanking CpG sites and adjacent single nucleotide variation (SNV) information to predict methylation changes between DNA sequences. CMC-WDTK captures global and local features of both reference and variant sequences and models high-dimensional relationships, offering accurate predictions of methylation changes. CMC-WDTK accurately predicted DNA methylation changes in eight real datasets (AUC greater than 0.8 for all datasets), with strong generalizability across datasets. Additionally, it identified a repeated cytosine and guanine sequence motif that promotes increased methylation. CMC-WDTK is the first computational tool used to predict methylation changes between sequences, offering significant advancements in understanding and comparing DNA methylation across diverse datasets and biological conditions.
Full text 134,892 characters · extracted from preprint-html · click to expand
CMC-WDTK: CpG methylation change prediction by a weight-sharing dual-branch Transformer-Kolmogorov–Arnold network model | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article CMC-WDTK: CpG methylation change prediction by a weight-sharing dual-branch Transformer-Kolmogorov–Arnold network model Jianmei Zhao, Di Liu, Yiming Wang, Hongfei Li, Guohua Wang This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7586988/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 07 Feb, 2026 Read the published version in BMC Genomics → Version 1 posted 13 You are reading this latest preprint version Abstract Differential methylation is a key epigenetic process contributing to cancer development. Most DNA methylation prediction methods rely on DNA sequences from the background reference genome, neglecting individual genetic variation, which limits their ability to capture methylation differences. To address this, we propose CMC-WDTK, a deep learning framework that combines a weight-sharing dual-branch Transformer with a Kolmogorov‒Arnold network (KAN) to integrate sequences flanking CpG sites and adjacent single nucleotide variation (SNV) information to predict methylation changes between DNA sequences. CMC-WDTK captures global and local features of both reference and variant sequences and models high-dimensional relationships, offering accurate predictions of methylation changes. CMC-WDTK accurately predicted DNA methylation changes in eight real datasets (AUC greater than 0.8 for all datasets), with strong generalizability across datasets. Additionally, it identified a repeated cytosine and guanine sequence motif that promotes increased methylation. CMC-WDTK is the first computational tool used to predict methylation changes between sequences, offering significant advancements in understanding and comparing DNA methylation across diverse datasets and biological conditions. Deep learning NLP CpG methylation changes Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1 Introduction DNA methylation regulates many cellular processes, with a particularly crucial role in gene transcription, chromatin structure, X chromosome inactivation, and genomic imprinting[ 1 ]. It has also emerged as a promising biomarker for the diagnosis of various diseases, particularly cancer[ 2 – 4 ]. Compared with those in normal tissues, the promoters of many tumor suppressor genes, such as the retinoblastoma suppressor gene (RB1)[ 5 ], the cell cycle regulators p16INK4a (CDKN2A) and p15INK4a (CDKN2B) [ 6 – 8 ], the apoptosis signaling pathway regulator DAPK1[ 9 ], and various transcription factors[ 10 – 12 ], are significantly hypermethylated in cancer tissues[ 13 ]. In contrast, high expression of oncogenes has also recently been reported to be associated with DNA demethylation. For example, several novel hypomethylated regions in intergenic areas have been linked to the expression of androgen receptor (AR), the oncogene MYC, and the transcription factor ERG in cancer cells. These regions contain binding motifs for transcription factors, including HOXB13, FOXA1, AR, and ERG, and are colocalized with enhancer signals, such as H3K27ac [ 14 ]. Although abnormal DNA methylation and its downstream molecular functions have been extensively studied, the mechanisms underlying the regulation of aberrant DNA methylation have not been fully elucidated. The development of predictive models for changes in DNA methylation levels using bioinformatics approaches is essential for uncovering these mechanisms and understanding the role of DNA methylation in tumorigenesis. Emerging evidence suggests that genetic variations, especially SNVs, play crucial roles in modulating DNA methylation changes at CpG sites adjacent to these variations [ 15 ] [ 16 – 18 ]. For example, the T allele at rs1001179 in the CAT promoter alters promoter methylation at specific CpG sites, which increases the levels of ETS-1 and GR-β transcription factors (TFs) and leads to increased CAT expression in chronic lymphocytic leukemia (CLL), a change associated with a more aggressive clinical course[ 19 ]. Another study demonstrated that rs7247241 within the PPP1R14A promoter leads to a decrease in the transcriptional activity of PPP1R14A, with the T allele associated with increased CpG methylation and reduced binding of the transcription factors CTCF and PLAGL2 compared with the C allele[ 20 ]. Furthermore, high CpG methylation in the intergenic region, together with the risk allele of the nearby SNP rs11986220, synergistically disrupts CTCF-mediated insulating loop formation, resulting in the loss of insulation. This enables the cis-regulatory element containing rs11986220 to upregulate the oncogenes MYC and PVT1, thereby increasing the risk of cancer[ 21 ]. These studies indicate that changes in the methylation levels of CpG sites are sensitive to variations in nearby SNVs and that their synergistic interaction regulates the transcription of related genes. This synergistic relationship is widespread across the genome, making it crucial to fully consider SNVs when predicting changes in CpG methylation levels for accurate results. Allele-specific methylation (ASM) analysis uses bisulfite sequencing (BS-Seq) data to identify differences in CpG methylation between alleles of adjacent heterozygous SNV sites in homologous genomes, providing an effective strategy to identify CpGs impacted by SNVs[ 22 – 24 ]. CpG sites exhibiting ASM events are ideal loci for modeling DNA methylation alterations. Although traditional experimental methods, such as whole-genome bisulfite sequencing (WGBS), provide accurate detection of genome-wide methylation states and identification of methylation changes at CpG sites, these techniques are costly and time-consuming. In recent years, deep learning based on natural language processing (NLP) has emerged as a powerful computational tool, overcoming the limitations of traditional machine learning approaches by reducing dependence on well-annotated genomes, and consequently has been increasingly applied to model and predict complex biological data[ 25 – 29 ]. Notably, NLP based deep learning techniques have the ability to learn directly and extract high-level features from DNA sequences, enabling accurate prediction of DNA methylation even in the absence of genomic annotation data [ 30 – 34 ]. Transformer-based models, in particular, have demonstrated exceptional performance in DNA methylation prediction owing to their capacity to model sequence context and capture long-range dependencies. For example, iDNA-ABF effectively identifies methylation-associated sequence patterns through a multiscale deep learning model [ 30 ], StableDNAm significantly enhances model stability and generalizability using feature fusion and contrastive learning techniques[ 32 ], and Methyl-GP improves generalization by examining sequence pattern conservation across species with a BERT-based language model[ 31 ]. While these computational approaches have significantly increased the accuracy of DNA methylation prediction and expanded its applications, a major limitation remains: most methods still rely solely on a background reference genome and fail to account for the local impact of DNA sequence variations on CpG site methylation. Inspired by this, we propose CMC-WDTK, an innovative deep learning framework focused on predicting methylation level changes between reference and alternate sequences for CpG sites with ASM events and identifying sequential patterns affecting DNA methylation. CMC-WDTK uses a dual-branch Transformer architecture with a weight-sharing mechanism to simultaneously extract features from reference and variant sequences surrounding CpGs, capturing global context features and local disturbances induced by SNVs. Furthermore, the framework incorporates a Kolmogorov‒Arnold network (KAN)[ 35 ] to further explore high-order nonlinear relationships between features, enhancing the model’s ability to express and generalize methylation change trends. Experimental validation on eight real-world datasets demonstrated that CMC-WDTK offers significant advantages in predicting the direction and magnitude of methylation changes induced by SNPs. This study effectively predicts, for the first time, CpG site methylation changes between DNA sequences by integrating SNVs and DNA sequences. It identifies the repeated C and G sequence patterns that promote increased methylation, offering a novel and efficient computational tool for understanding the mechanisms underlying DNA methylation alterations. 2 Materials and methods 2.1 Dataset For model training and evaluation, datasets were selected from a previous study conducted by our group, in which we identified ASM sites using BS-Seq datasets across 31 cancer types and their corresponding normal tissue samples[ 36 ]. ASM analysis employs Fisher's exact test to evaluate the significance of differences in the counts of methylated and unmethylated reads covering the reference and variant alleles of nearby heterogeneous SNVs in a BS-Seq sample. This determines whether CpG site methylation levels exhibit allele specificity (Fig. 1 A)[ 18 , 36 ]. A total of eight datasets were selected for modeling ASM CpG sites, encompassing four tissue types (esophageal, gastric, intestinal, and pancreatic) and four cancer cell lines (MCF7, HepG2, OCI-LY7, and HeLa) (Suppl. Table S1 ). Tissue datasets are a mixed collection of cancerous and normal samples. The obtained ASM CpG information includes the genomic coordinates of the CpGs and SNVs, the reference and alternate alleles of the SNVs, and the total and methylated read counts for each allele. Suppl. Table S2 provides a summary of the total number of CpG samples in each dataset, along with the corresponding counts of positive and negative samples. As shown in Fig. 1 B, for each CpG, we first extracted a 101 bp DNA sequence centered on the CpG site, consisting of 50 bases upstream and 50 bases downstream from the reference genome (GRCh38), denoted as \(se{q_r}\) . Then, according to the SNV genomic position information, the corresponding base in the reference sequence is substituted with the alternative allele to generate the alternative sequence \(se{q_a}\) . The change in methylation between \(se{q_r}\) and \(se{q_a}\) of the CpG site is defined as: $$~VA{R_i}={\beta _i}(se{q_a}) - {\beta _i}(se{q_r})$$ 1 where \({\beta _i}(se{q_a})\) refers to the proportion of methylated reads out of the total reads mapped to the alternative allele and where \({\beta _i}(se{q_r})\) refers to the proportion of methylated reads out of the total reads mapped to the reference allele. The true class label for CpG \(\:i\) is 1 (positive) if \(VA{R_i}\) is greater than 0; otherwise, it is 0 (negative). 2.2 Methods This study proposes an end-to-end deep learning framework for CpG contextual sequences containing SNVs. It employs a weight-sharing, dual-branch Transformer encoder to extract global and local features concurrently from both the reference sequence and the alternative sequence. A two-layer KAN subsequently maps the high-dimensional vectors to continuous methylation level predictions. WDTK accurately captures methylation fluctuations at neighboring CpG sites jointly induced by SNVs and their surrounding sequences. The organic combination of Transformer and KAN achieves high-precision prediction of methylation gain and loss types. Specifically, WDTK first uses a shared embedding layer to split each of the two allele sequences ( \(\:{seq}_{r}\) / \(\:{seq}_{a}\) ) into k-mer subsequences and project them into the same d-dimensional vector space. Both sequences are then passed in parallel through identical Transformer encoders for deep feature extraction: in each layer, the multi-head self-attention module adaptively captures long-range dependencies and local microvariations within the sequence, and the subsequent feed-forward network—consisting of two linear transformations with ReLU activation—further enriches the nonlinear representation. All sublayers employ residual connections and layer normalization to ensure stable gradients. Finally, the output tensor of shape (B, L, d) is transposed and compressed via adaptive average pooling into a fixed-length (B, d) vector. This preserves the key positional interactions highlighted by multi-head attention while retaining the global context. We subsequently concatenate the two branch pooling vectors into (B, 2d), which are then fed into a two-layer KAN network. First, a linear mapping of \((2d \to 1d)\) with ReLU activation performs dimension compression and feature fusion. Then, a linear layer of \((d \to 1)\) directly outputs continuous methylation level prediction scores. The entire framework is end-to-end, optimized with AdamW, and incorporates appropriate dropout and learning rate scheduling, further enhancing the model's generalization performance and biological interpretability. 2.2.1 Dual-branch Transformer As shown in Fig. 1 C, the aforementioned reference sequence and variant sequence are first segmented into continuous k-mer fragments to extract local contextual information. These k-mers are subsequently treated as discrete "vocabularies" and fed into a learnable embedding layer, which maps them into fixed-length dense vectors. This process captures the potential semantic features of the sequences and provides efficient vectorized representations for subsequent models, denoted as \({X_r}\) and \({X_a}\) . Subsequently, \({X_r}\) and \({X_a}\) are fed in parallel into a shared-weight Transformer encoder, where multi-head self-attention and feed-forward layers extract their respective context-aware feature representations: $${H_i}=Transformer\left( {{X_{\mathbf{i}}}} \right).i \in \left\{ {r,a} \right\}$$ 2 After the sequences are processed through the Transformer encoder, the resulting feature representations \({H_i}\) are passed through an adaptive average pooling layer to reduce their dimensionality. This pooling operation compresses the sequence information into fixed-length vectors of size 1, each encoding the overall methylation pattern of the sequence. The pooling is applied separately to each \({H_i}\) : $${P_i}=AdaptiveAvePooild\left( {{H_i}} \right).i \in \left\{ {r,a} \right\}$$ 3 The pooling layer compresses each sequence’s representation into a single scalar, preserving its most salient information. The pooled scalars from the reference and alternative sequences are then concatenated along the feature dimension to form a joint representation that captures their interrelationship. $${P_{concat}}=cat\left( {{P_i},{\kern 1pt} {\kern 1pt} {\kern 1pt} \dim =1} \right),i \in \left\{ {r,a} \right\}$$ 4 The concatenated representation is then fed into the KAN module for further processing and feature fusion. The KAN module consists of fully connected layers that are capable of transforming the concatenated features into the final output. The final output is a scalar, reflecting the methylation analysis result between the reference and variant sequences. 2.2.2 Kolmogorov–Arnold network As shown in Fig. 1 D, upon receiving the feature vector extracted and fused by the dual-branch Transformer, the module undergoes nonlinear transformation followed by linear mapping through two layers of learnable mapping functions. This process extracts key features from the high-dimensional, context-aware representation, ultimately generating logit outputs that can be used for determining methylation gain or loss. The KAN consists of two layers of learnable mapping functions \({\varphi ^{\left( l \right)}}\) , and its computation process is as follows: First Layer: $$z_{q}^{{\left( 1 \right)}}=\sum\limits_{{p=1}}^{{{n_0}}} {\varphi _{{q,p}}^{{\left( 1 \right)}}} \left( {{p_{concat,p}}} \right),{\kern 1pt} {\kern 1pt} {\kern 1pt} {\kern 1pt} q=1,...,{n_1}$$ 5 Second Layer: $$y=\sum\limits_{{q=1}}^{{{n_1}}} {\varphi _{{1,q}}^{{\left( 2 \right)}}} \left( {z_{q}^{{\left( 1 \right)}}} \right).$$ 6 Overall, the KAN mapping can be expressed as the composition of the two mapping functions: $$KAN\left( x \right)=({\Phi ^{\left( 2 \right)}} \circ {\Phi ^{\left( 1 \right)}})(x),{\kern 1pt} {\kern 1pt} {\kern 1pt} {\kern 1pt} {\kern 1pt} {\kern 1pt} {\Phi ^{\left( l \right)}}{(x)_j}=\sum\limits_{{i=1}}^{{{n_{l - 1}}}} {\varphi _{{j,i}}^{{\left( l \right)}}} ({x_i}),{\kern 1pt} {\kern 1pt} l=1,2.$$ 7 The model’s final output is the result produced by the KAN module. The predicted value represents the type of methylation change (0, 1) between the variant sequence relative to the reference sequence or between DNA sequences from different individuals. A value closer to 1 indicates a greater increase, whereas a value closer to 0 indicates a greater decrease. 2.2.3 Optimization In this study, model training was optimized using standard supervised learning techniques. For the binary classification task, we employed binary cross-entropy loss, which measures the discrepancy between the model’s predictions and the true labels. The formula for binary cross-entropy loss is as follows: $$L= - \frac{1}{N}\sum\limits_{{i=1}}^{N} {\left[ {{y_i}\log ({{\widehat {y}}_i})+(1 - {y_i})\log (1 - {{\widehat {y}}_i})} \right]} .$$ 8 Here, \({y_i}\) denotes the true label, and \({\widehat {y}_i}\) denotes the model’s predicted value. To optimize the model parameters, we employed the AdamW optimizer. AdamW is an adaptive learning-rate method that integrates the concepts of momentum and RMSprop. During training, it effectively adjusts the learning rate for each parameter, making it particularly suitable for handling sparse gradient problems. In the optimization process, the AdamW optimizer uses the following update rule: $${\theta _t}={\theta _{t - 1}} - \eta \cdot \left( {\frac{{{{\widehat {m}}_t}}}{{\sqrt {{{\widehat {v}}_t}} +\varepsilon }}+\lambda {\theta _{t - 1}}} \right).$$ 9 Here, \({\theta _t}\) denotes the parameter at the current timestep, \({\theta _{t - 1}}\) is the parameter at the previous timestep, \(\eta\) is the learning rate, \({\widehat {m}_t}\) and \({\widehat {v}_t}\) are the bias-corrected estimates of the first and second moments of the gradients, \(\varepsilon\) is a small constant to prevent division-by-zero errors, and \(\lambda\) is the weight decay coefficient. After each forward pass, the model calculates the loss function, and this loss is then propagated back through each layer of the model via backpropagation to compute the gradient for each parameter. The optimizer subsequently updates the model's parameters on the basis of the calculated gradients. With each update, previous gradients are first cleared, new gradients are computed through backpropagation, and finally, the parameter values are adjusted according to the optimizer's rules. The gradients are cleared after each update to prevent accumulation, allowing successive updates to guide the model toward its optimum. Through multiple gradient updates, the model is progressively optimized, approaching the optimal solution. To enhance the model's generalization ability and prevent overfitting, we incorporated the Dropout technique during training. Dropout, by randomly dropping a subset of neurons during each training step, reduces the model's overreliance on specific neurons, thereby improving its performance on unseen data. The use of Dropout contributes to enhancing the model's stability and generalizability. 2.3 Performance assessment To assess the performance of the CMC-WDTK, we applied fivefold cross-validation in the different experiments. The following four commonly used metrics were employed to evaluate our method in models for different experimental datasets: area under the receiver operating characteristic curve (AUC), area under the precision-recall curve (AUPR), accuracy (ACC), and F1 score. The computation of these metrics is as follows: $$\begin{gathered} Recall=\frac{{TP}}{{TP+FN}}, \hfill \\ Specificity=\frac{{TN}}{{TP+FP}}, \hfill \\ Precision=\frac{{TP}}{{TP+FP}}, \hfill \\ F1=\frac{{2*Recall*Precision}}{{Recall+Precision}}, \hfill \\ ACC=\frac{{TP+TN}}{{TP+TN+FP+FN}}. \hfill \\ \end{gathered}$$ 10 where \(TP\) , \(TN\) , \(FP\) , and \(FN\) represent the number of true-positive, true-negative, false-positive, and false-negative predicted CpGs, respectively. We used the receiver operating characteristic (ROC) curve and the precision‒recall (PR) curve to evaluate the overall performance of the model[ 37 ]. These curves provide an intuitive visualization of the model’s superiority. 3 Results Allele-specific methylated CpG sites are widely distributed in tissues and cell lines In this study, CpG sites exhibiting ASM events from eight datasets were used for model training and evaluation. The number of CpG sites in each dataset ranged from 14,551 to 353,746, each affected by one or more SNVs (Fig. 2 A, Suppl Table S2 ). In this study, we set the DNA sequence length to 50 bp upstream and downstream of the CpG sites to extract global features. To better explore the local impact of SNVs on CpG site methylation levels, we calculated the distance of associated SNVs relative to CpG sites. As shown in Fig. 2 B, the SNVs are predominantly located within 20 bases upstream and downstream of the CpG sites, with more than 98% of them being SNPs (dbSNP build 156 release). For all the SNVs, we assessed the frequency of 12 different types of nucleotide substitutions and found that C/T and G/A substitutions occurred more frequently than other substitution types did (Fig. 2 C). This finding suggests that CpG methylation is more sensitive to these two substitutions in the surrounding nucleotides. Additionally, we utilized ANNOVAR [ 38 ] to annotate these SNVs to the GRCh38 reference genes, using genomic annotation information for reference genes obtained from the UCSC database[ 39 ]. The results showed that approximately half of the ASM sites across the eight datasets occurred within or near genes, with the remaining sites located in intergenic regions (Fig. 2 D). These findings suggest that these SNV-affected CpG sites may serve as potential functional sites for regulating gene expression. Therefore, focusing on these sites for modeling, rather than all CpG sites across the genome, enhances the identification of DNA sequence patterns influencing methylation. Experiment Design To comprehensively evaluate the model’s performance, two sets of experiments were designed: a tissue group and a cell line group. Within-cohort validation: To assess the model's performance on a specific tissue or cell line, we conducted experiments on eight real-world datasets. Cross-cohort validation: To further test generalizability, we adopted a cross-dataset scheme within each group. Specifically, we trained the model on one tissue type(e.g., gastric) and then evaluated its performance on a separate, independent tissue dataset (e.g., pancreas), simulating the model’s behavior when confronted with novel tissues or unseen cell types in practical applications. By comparing performance across cross-dataset and within-cohort experiments, we obtain an objective measure of the model’s transferability and robustness. All the experiments were conducted using 5-fold cross-validation to assess the model’s performance accurately. The proposed framework was implemented in PyTorch 1.13.1, and both training and evaluation were conducted on a high-performance server equipped with an NVIDIA RTX 4090 GPU. We also employed Bayesian optimization to jointly search and automatically tune multiple key hyperparameters, using the AUC as the objective metric. The optimal hyperparameter set is summarized in Table 1 . Table 1 The optimal hyperparameter set. Parameter Range of Values Optimal Value(cell line) Optimal Value(tissue) Learning rate 1e-4, 1e-3, 1e-2 1e-3 1e-4 Hidden layer dimension [64, 256] 80 64 Number of attention heads 4, 8 8 4 Number of Transformer layers 3, 4, 5, 6, 7, 8 6 3 Training epochs [64, 256] 256 256 Batch size 128, 256, 512 256 128 k-mer 3, 4, 5, 6 4 4 Dropout 0.1, 0.3, 0.5 0.1 0.3 Performance evaluation of CMC-WDTK We first plotted the ROC (Fig. 3 A) and PR curves (Fig. 3 B) for all the models in different tissues and cell lines separately. All the AUC values are greater than 0.8, reaching above 0.9, indicating good overall model performance in different models, whereas the AUPR is approximately 0.6, reflecting a reasonable ability to identify the positive class. For the cell line models, the ACC ranges from above 0.6 to above 0.7, indicating acceptable model performance. However, the F1 score is approximately 0.6, which suggests a reasonable balance between precision and recall (Fig. 3 C). In the tissue models, the ACC ranged from above 0.7 to above 0.8, showing a slight increase compared with that in the cell line model (Fig. 3 D). For datasets with better performance, such as those for gastric and pancreatic tissues, the cancer and normal samples are balanced in the data source. In contrast, in the other models, the proportion of cancer samples is greater, or the dataset consists entirely of cancer samples (Suppl. Table S1 ). This may result from complex dynamic changes in global DNA methylation during cancer development[ 40 – 42 ]. To present a more intuitive view of how models make decisions, we used uniform manifold approximation and projection (UMAP, a widely used visualization tool) to visualize the distribution of the feature representations learned by each model in the feature space (UMAP). The feature representations learned by most of the models, such as those for gastric tissue and the HeLa cell line, generally have clear boundaries and are well clustered (Fig. 3 E). In other words, our model clearly separates the CpG sites with increased methylation and decreased methylation due to variation and every cluster together rather than dispersing them. This finding indicates that CMC - WDTK has the potential to make accurate predictions for different experimental datasets. Cross-dataset validation To assess whether the learned sequence patterns are applicable for methylation prediction across different datasets, we performed cross-model validation to determine whether a model trained on one tissue or cell line dataset can achieve satisfactory results on testing sets from other tissue or cell line datasets. Specifically, we conducted separate cross-model validation for the tissue and cell line datasets, evaluating the models using four metrics: AUC, AUPR, accuracy, and F1 score. The results show that the cross-model validation outcomes are highly consistent across tissue and cell line datasets for all the evaluation metrics (Fig. 4 ). This suggests that models trained on one tissue or cell line dataset can generalize well to other datasets, indicating that the sequence patterns captured by CMC-WDTK significantly contribute to methylation alteration prediction. Furthermore, these results highlight the model's strong generalizability. The importance of SNVs in the model's decision-making process The attention scores of the Transformer model reflect the weight of the model's focus on features at different positions in the sequence during decision-making. To further explore the model's decision-making mechanism and clarify the importance of SNVs in their local effects on CpG site methylation prediction, this study analyzes the attention scores to investigate how the model's focus on features at SNV sites differs between reference and variant sequences. We will then illustrate this using TP73 as an example. By analyzing sequences from the testing set, we identified six CpG sites located within the introns or exons of the tumor suppressor gene TP73 , each associated with a corresponding ASM SNP. These SNPs are located in close proximity to the CpG sites at distances of 5, 1, 40, 0, and 1 bp and include rs751035, rs6671482, rs10910008, rs11580811, rs3765742, and rs61735051 (Fig. 5 A). Among them, rs751035 comes from the gastric tissue dataset, and the others are from the intestinal tissue dataset. According to the Fisher’s test statistics for SNP and CpG results from the CanASM database, methylation significantly decreases at the variant alleles for all six of these CpG sites, meaning that their true labels are all negative. Simultaneously, the corresponding model predictions show that their predicted labels are consistently negative, matching the true labels. The core task of this study is to capture the global contextual DNA sequence around CpG sites and local SNV features to predict the variation trend of CpG site methylation levels. To quantify the importance of SNVs in the model decision-making process, self-attention score analysis was subsequently performed on above 6 SNV sites. By calculating the attention scores, we obtained a matrix representing the attention scores between all tokens for the reference and alternative sequences. Here, each token represents a k-mer extracted from the DNA sequence. First, the model calculates the self-attention scores of the reference and the variant sequence respectively. The specific process is as follows: the two sequences are input into the model separately, and the self-attention matrix output by the last Transformer block is extracted. This matrix is obtained by averaging the results of independent calculations from multiple attention heads, with a dimension of L×L (where L is the number of k-mers), and is used to characterize the pairwise correlation strength between all k-mers in the sequence. In this study, the sequence length was set to 101 bp, and a sliding window with a step size of 1 was used to segment the sequence into k-mers, where k = 4, resulting in L = 98. Since the step size was 1, each SNV is covered by k consecutive k-mers. Subsequently, local attention submatrices for each SNV were extracted. The last k-mer covering the SNV was taken as the center of the extracted the submatrices, with 10 k-mers selected upstream and downstream of this center to form a 21×21 matrix. Finally, the difference between the corresponding submatrices of the variant and the reference sequence was calculated (by subtracting the two matrices) to generate a differential attention score matrix, based on which a heatmap was plotted (Fig. 5 B). The results suggest that the model can effectively focus on local sequence patterns near SNVs, accurately identifying their potential impact on methylation changes, thereby validating the reasonableness and specificity of the model's decisions. CMC-WDTK helps capture global sequential patterns in methylation alteration Finding the underlying sequence patterns surrounding the CpG sites is an effective step in understanding the reasons for their modification and in revealing the biological functions of modifications[ 43 ]. To this end, this study divided the sequences near CpG sites into three groups based on SNV variations and CpG methylation levels: the reference sequence group, the variant group predicted by CMC-WDTK to gain methylation, and the variant group predicted to lose methylation to identify the sequence features that influence DNA methylation levels by comparing these three groups of sequence motifs. The specific implementation process is as follows. First, 41 bp DNA sequences were extracted with CpG sites as the center, encompassing 20 bp of sequence both upstream and downstream of the CpG sites. This 41 bp length was selected because it can effectively capture key information associated with DNA methylation detection[ 44 , 45 ]. Next, the KpLogo tool[ 46 ] was employed to generate motif logo plots for each group of sequences, with the plots constructed based on the occurrence probabilities of the four nucleotide types at each position within the sequences (Fig. 6 ). After comparing the motif logo plots generated for the three groups of sequences in each dataset, it was found that the motifs for the three corresponding groups exhibited consistent sequence features across datasets. The reference sequence group and the increased methylation group both showed highly conserved C and G motifs near the target CpG sites. In contrast, in the decreased methylation group, this conserved C and G motif tended to be replaced by T or A. This finding not only indicates that CpG site methylation across diverse tissues and cell lines may follow a shared sequence logic centered on conserved C and G motifs, but also offers a model-supported approach to unravel the common molecular mechanisms driving methylation alterations at the DNA sequence level. Discussion DNA methylation is a crucial epigenetic modification involved in the regulation of gene expression and various biological processes. To date, several methods have been proposed to predict DNA methylation and identify sequence patterns. While these methods have provided valuable insights into methylation patterns, they typically rely on reference genomes, which do not account for the genetic variations that exist between individuals and diverse populations. In contrast, CMC-WDTK incorporates SNV-induced variations with sequential contexts, which allows for precise predictions of methylation changes driven by local genetic diversity and global sequence features. We present CMC-WDTK, an innovative deep learning framework designed to predict methylation changes at CpG sites by integrating sequence features of CpG sites and adjacent single nucleotide variants. The framework combines a dual-branch Transformer architecture with a Kolmogorov-Arnold Network (KAN), enabling it to capture the global sequence context and local variation features. This approach allows for accurate prediction of methylation changes between reference and variant sequences, with a particular focus on understanding the impact of SNP variations on CpG site methylation levels. Additionally, cross-model validation revealed that models trained on one tissue or cell line dataset generalize well to others, indicating the model’s ability to perform consistently across different biological contexts. Furthermore, attention score analysis revealed that the model’s dual-branch architecture assigns greater differences in attention scores to local sequences near SNVs, emphasizing the critical role of the local variant sequence context in predicting methylation changes. Additionally, our study identified specific sequence motifs associated with changes in CpG methylation. Using KpLogo, we found that sequences associated with increased methylation are characterized by repeating C and G motifs, whereas sequences showing decreased methylation exhibit substitutions of C and G with A and T. These patterns were consistent across different datasets, providing further validation of the model’s ability to capture biologically relevant sequence features. Despite the promising results of CMC-WDTK, several potential limitations exist. First, while the eight datasets used in this study cover different tissue types and cell lines, there may still be an underrepresentation of other genomic backgrounds. Therefore, future research could expand the datasets. Second, while our approach has shown excellent performance in predicting methylation changes between variant and reference sequences, its application has yet to be fully explored in disease versus normal group comparisons. Future work will focus on extending CMC-WDTK to analyze methylation differences between disease and normal groups, further validating its potential for broad biological and clinical applications. Overall, by integrating DNA sequence features and genomic variations, CMC-WDTK is the first computational tool to predict methylation changes between sequences, offering a significant advancement in the field of DNA methylation comparison. The identified sequence pattern of repeating C and G motifs, which is associated with increased methylation, also helps in understanding the mechanisms underlying DNA methylation. The model’s ability to generalize across different datasets and biological conditions further enhances its potential for broad applications in comparing the methylation levels between two DNA sequences. Declarations Ethics approval and consent to participate Not applicable Consent for publication Not applicable Funding This work was supported by the National Natural Science Foundation of China (Grant Nos. 62225109, 62402345, and 62302342). We would like to express our sincere thanks to the funding agencies for their financial support, which made this research possible. Author Contribution JZ: Conceptualization, Methodology, Formal analysis, Data curation, Writing. DL: Methodology, model development. YW: Methodology. HL: Conceptualization, Investigation. GW: Conceptualization, Supervision, Funding acquisition, Resources. All authors reviewed the manuscript. Acknowledgements Not applicable Data Availability The original Whole-Genome Bisulfite Sequencing (WGBS) data supporting the findings of this study have been deposited in the NCBI GEO database under the primary accession codes provided in Supplementary Table S1.The preprocessed input data obtained from whole-genome bisulfite sequencing (WGBS), together with the model code used in this study, are available at: https://github.com/LiuDiBio/CMC-WDTK. References Meng, H., et al., DNA methylation, its mediators and genome integrity. Int J Biol Sci, 2015. 11 (5): p. 604-17. Liu, J., et al., Evaluating the comprehensive diagnosis efficiency of lung cancer, including measurement of SHOX2 and RASSF1A gene methylation. BMC Cancer, 2024. 24 (1): p. 282. Zhang, F. and T. Evans, Stage-specific DNA methylation dynamics in mammalian heart development. Epigenomics, 2025. 17 (5): p. 359-371. Zhang, W.Z., C.Y. Wu, and H. Lai, A Review on the Role of DNA Methylation in Aortic Disease Associated With Marfan Syndrome. Cardiol Res, 2025. 16 (3): p. 169-177. Greger, V., et al., Epigenetic changes may contribute to the formation and spontaneous regression of retinoblastoma. Hum Genet, 1989. 83 (2): p. 155-8. Herman, J.G., et al., Inactivation of the CDKN2/p16/MTS1 gene is frequently associated with aberrant DNA methylation in all common human cancers. Cancer Res, 1995. 55 (20): p. 4525-30. Merlo, A., et al., 5' CpG island methylation is associated with transcriptional silencing of the tumour suppressor p16/CDKN2/MTS1 in human cancers. Nat Med, 1995. 1 (7): p. 686-92. Baur, A.S., et al., Frequent methylation silencing of p15(INK4b) (MTS2) and p16(INK4a) (MTS1) in B-cell and T-cell lymphomas. Blood, 1999. 94 (5): p. 1773-81. Tan, Y., et al., Alternative polyadenylation reprogramming of MORC2 induced by NUDT21 loss promotes KIRC carcinogenesis. JCI Insight, 2023. 8 (18). Hesham, D., et al., Epigenetic silencing of ZIC4 unveils a potential tumor suppressor role in pediatric choroid plexus carcinoma. Sci Rep, 2024. 14 (1): p. 21293. Diakiw, S.M., et al., The granulocyte-associated transcription factor Kruppel-like factor 5 is silenced by hypermethylation in acute myeloid leukemia. Leuk Res, 2012. 36 (1): p. 110-6. Fan, X.Y., et al., Association between RUNX3 promoter methylation and gastric cancer: a meta-analysis. BMC Gastroenterol, 2011. 11 : p. 92. Herman, J.G. and S.B. Baylin, Gene silencing in cancer in association with promoter hypermethylation. N Engl J Med, 2003. 349 (21): p. 2042-54. Stelloo, S., et al., Endogenous androgen receptor proteomic profiling reveals genomic subcomplex involved in prostate tumorigenesis. Oncogene, 2018. 37 (3): p. 313-322. Villicana, S. and J.T. Bell, Genetic impacts on DNA methylation: research findings and future perspectives. Genome Biol, 2021. 22 (1): p. 127. Zeng, Y., et al., DNA methylation modulated genetic variant effect on gene transcriptional regulation. Genome Biol, 2023. 24 (1): p. 285. Villicana, S., et al., Genetic impacts on DNA methylation help elucidate regulatory genomic processes. Genome Biol, 2023. 24 (1): p. 176. Do, C., et al., Allele-specific DNA methylation is increased in cancers and its dense mapping in normal plus neoplastic cells increases the yield of disease-associated regulatory SNPs. Genome Biol, 2020. 21 (1): p. 153. Galasso, M., et al., The rs1001179 SNP and CpG methylation regulate catalase expression in chronic lymphocytic leukemia. Cell Mol Life Sci, 2022. 79 (10): p. 521. Tian, Y., et al., Novel role of prostate cancer risk variant rs7247241 on PPP1R14A isoform transition through allelic TF binding and CpG methylation. Hum Mol Genet, 2022. 31 (10): p. 1610-1621. Ahmed, M., et al., CRISPRi screens reveal a DNA methylation-mediated 3D genome dependent causal mechanism in prostate cancer. Nat Commun, 2021. 12 (1): p. 1781. Onuchic, V., et al., Allele-specific epigenome maps reveal sequence-dependent stochastic switching at regulatory loci. Science, 2018. 361 (6409). Dumont, E.L.P., B. Tycko, and C. Do, CloudASM: an ultra-efficient cloud-based pipeline for mapping allele-specific DNA methylation. Bioinformatics, 2020. 36 (11): p. 3558-3560. Rosenski, J., et al., Atlas of imprinted and allele-specific DNA methylation in the human body. Nat Commun, 2025. 16 (1): p. 2141. Zhang, W., et al., Predicting genome-wide DNA methylation using methylation marks, genomic position, and DNA regulatory elements. Genome Biol, 2015. 16 (1): p. 14. Abbas, Z., et al., XGBoost framework with feature selection for the prediction of RNA N5-methylcytosine sites. Mol Ther, 2023. 31 (8): p. 2543-2551. Gao, Z., et al., EpiGePT: a pretrained transformer-based language model for context-specific human epigenomics. Genome Biol, 2024. 25 (1): p. 310. Li, W., W.H. Wong, and R. Jiang, DeepTACT: predicting 3D chromatin contacts via bootstrapping deep learning. Nucleic Acids Res, 2019. 47 (10): p. e60. Ji, Y., et al., DNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome. Bioinformatics, 2021. 37 (15): p. 2112-2120. Jin, J., et al., iDNA-ABF: multi-scale deep biological language learning model for the interpretable prediction of DNA methylations. Genome Biol, 2022. 23 (1): p. 219. Xie, H., et al., Methyl-GP: accurate generic DNA methylation prediction based on a language model and representation learning. Nucleic Acids Res, 2025. 53 (6). Zhuo, L., et al., StableDNAm: towards a stable and efficient model for predicting DNA methylation based on adaptive feature correction learning. BMC Genomics, 2023. 24 (1): p. 742. Yu, X., et al., iDNA-ITLM: An interpretable and transferable learning model for identifying DNA methylation. PLoS One, 2024. 19 (10): p. e0301791. Yu, X., et al., iDNA-OpenPrompt: OpenPrompt learning model for identifying DNA methylation. Front Genet, 2024. 15 : p. 1377285. Liu, Z., et al., KAN: Kolmogorov-Arnold Networks. ArXiv, 2024. abs/2404.19756 . Zhao, J., et al., CanASM: a comprehensive database for genome-wide allele-specific DNA methylation identification and annotation in cancer. BMC Genomics, 2025. 26 (1): p. 648. Kumar, R. and A. Indrayan, Receiver operating characteristic (ROC) curve for medical researchers. Indian Pediatr, 2011. 48 (4): p. 277-87. Wang, K., M. Li, and H. Hakonarson, ANNOVAR: functional annotation of genetic variants from high-throughput sequencing data. Nucleic Acids Res, 2010. 38 (16): p. e164. Nassar, L.R., et al., The UCSC Genome Browser database: 2023 update. Nucleic Acids Res, 2023. 51 (D1): p. D1188-D1195. Mishra, N.K. and C. Guda, Genome-wide DNA methylation analysis reveals molecular subtypes of pancreatic cancer. Oncotarget, 2017. 8 (17): p. 28990-29012. Chen, K., et al., Individualized dynamic methylation-based analysis of cell-free DNA in postoperative monitoring of lung cancer. BMC Med, 2023. 21 (1): p. 255. Rodriguez-Lloveras, H., et al., DNA Methylation Dynamics and Prognostic Implications in Metastatic Differentiated Thyroid Cancer. Thyroid, 2025. 35 (5): p. 494-507. Smith, Z.D. and A. Meissner, DNA methylation: roles in mammalian development. Nat Rev Genet, 2013. 14 (3): p. 204-20. Zeng, W., A. Gautam, and D.H. Huson, MuLan-Methyl-multiple transformer-based language models for accurate DNA methylation prediction. Gigascience, 2022. 12 . Li, W. and A. Godzik, Cd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences. Bioinformatics, 2006. 22 (13): p. 1658-9. Wu, X. and D.P. Bartel, kpLogo: positional k-mer analysis reveals hidden specificity in biological sequences. Nucleic Acids Res, 2017. 45 (W1): p. W534-W538. Additional Declarations No competing interests reported. Supplementary Files TableS1Datasetsinfo.xlsx TableS2Samplestasitics.xlsx Cite Share Download PDF Status: Published Journal Publication published 07 Feb, 2026 Read the published version in BMC Genomics → Version 1 posted Editorial decision: Revision requested 05 Nov, 2025 Reviews received at journal 30 Oct, 2025 Reviews received at journal 15 Oct, 2025 Reviews received at journal 14 Oct, 2025 Reviewers agreed at journal 14 Oct, 2025 Reviewers agreed at journal 13 Oct, 2025 Reviewers agreed at journal 08 Oct, 2025 Reviewers agreed at journal 07 Oct, 2025 Reviewers invited by journal 06 Oct, 2025 Editor assigned by journal 15 Sep, 2025 Editor invited by journal 12 Sep, 2025 Submission checks completed at journal 11 Sep, 2025 First submitted to journal 11 Sep, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7586988","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":530734690,"identity":"927e037d-3dcc-4beb-ac7f-856da7bc35c2","order_by":0,"name":"Jianmei Zhao","email":"","orcid":"","institution":"Northeast Forestry University","correspondingAuthor":false,"prefix":"","firstName":"Jianmei","middleName":"","lastName":"Zhao","suffix":""},{"id":530734691,"identity":"f3d77442-7a4d-4898-800b-7587a014c027","order_by":1,"name":"Di Liu","email":"","orcid":"","institution":"Northeast Forestry University","correspondingAuthor":false,"prefix":"","firstName":"Di","middleName":"","lastName":"Liu","suffix":""},{"id":530734693,"identity":"81407ed0-2364-4433-85a9-d6298759838b","order_by":2,"name":"Yiming Wang","email":"","orcid":"","institution":"Northeast Forestry University","correspondingAuthor":false,"prefix":"","firstName":"Yiming","middleName":"","lastName":"Wang","suffix":""},{"id":530734694,"identity":"8b365816-6f51-4804-abae-fa7e5b0c6d83","order_by":3,"name":"Hongfei Li","email":"","orcid":"","institution":"University of Electronic Science and Technology of China","correspondingAuthor":false,"prefix":"","firstName":"Hongfei","middleName":"","lastName":"Li","suffix":""},{"id":530734699,"identity":"b21421e5-76e1-4124-b4d1-85e8d154b7d3","order_by":4,"name":"Guohua Wang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA10lEQVRIiWNgGAWjYDACCTBpw8DAfABIsxGvJQ2oOoE0LYdJ0CIf3WP4ueDX+Tx+Nh4Dhg9lhxn4Zzfg12J454yx9My+28WSbTwGjDPOHWaQuHOAgJYZuRukeXtuJ26432PAzNt2mMFAIoGgls2/eXvOJe4/xmPA/JcYLfISudukeX4cSNwA9AszIzFaDCTyv1nzNiQnzjjGVnCw51w6j8QNQrbMSEu+zfPHLrG/jXnjgx9l1nL8MwjZcgBIMLZBOCA2D371IFsaQOQfgupGwSgYBaNgJAMA3btDXFWffN4AAAAASUVORK5CYII=","orcid":"","institution":"Northeast Forestry University","correspondingAuthor":true,"prefix":"","firstName":"Guohua","middleName":"","lastName":"Wang","suffix":""}],"badges":[],"createdAt":"2025-09-11 02:38:21","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7586988/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7586988/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12864-026-12627-9","type":"published","date":"2026-02-07T15:59:08+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":93765551,"identity":"cc175487-3bfd-450f-b9c1-49ca349c684a","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1666158,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/3a38401d0d89e91b802bac94.docx"},{"id":93766468,"identity":"acd8f7a7-1ff8-4e63-b408-db0d0b653bfe","added_by":"auto","created_at":"2025-10-17 10:41:42","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6631,"visible":true,"origin":"","legend":"","description":"","filename":"ee6c09ee2f5741459f01368cee7f3da2.json","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/36fb9966291a2ac9cb755033.json"},{"id":93765153,"identity":"93fec2fd-6a4f-457b-9928-9faf28fd33ef","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":10949,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1Datasetsinfo.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/fe896d45341596439ba51baf.xlsx"},{"id":93765159,"identity":"3322873d-9287-4150-abc6-983f2a17da31","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":9626,"visible":true,"origin":"","legend":"","description":"","filename":"TableS2Samplestasitics.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/5330efbda3f981c868c2647b.xlsx"},{"id":93765562,"identity":"7f29ea53-430f-43c8-935d-86c90f7a8729","added_by":"auto","created_at":"2025-10-17 10:33:52","extension":"xml","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":107125,"visible":true,"origin":"","legend":"","description":"","filename":"ee6c09ee2f5741459f01368cee7f3da21enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/6430923d93b328855dba7519.xml"},{"id":93765163,"identity":"e50d750c-a057-449a-ac06-adfef98d6c52","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"wmf","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":624,"visible":true,"origin":"","legend":"","description":"","filename":"image26.wmf","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/50f3c6b0e5ff6b32276e94d7.wmf"},{"id":93765567,"identity":"e35f568a-f52e-49b7-b749-bbfde4447c2c","added_by":"auto","created_at":"2025-10-17 10:34:12","extension":"wmf","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":624,"visible":true,"origin":"","legend":"","description":"","filename":"image5.wmf","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/4706206f668f7a8efae9df04.wmf"},{"id":93765559,"identity":"f8786c49-113e-40a0-a5eb-a4d3b38a52ce","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":90847,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/94c1ed9f1969b5ef759c266f.png"},{"id":93766471,"identity":"68ce2d93-0df2-43be-84f7-33a31cf3d42a","added_by":"auto","created_at":"2025-10-17 10:41:42","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":61055,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/625de723bd9a02fe553e0385.png"},{"id":93765172,"identity":"fc1f8fb0-937b-418e-a41b-43a277512bda","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":94199,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/9951d6bddb97e7d50db82100.png"},{"id":93765165,"identity":"92916f12-8600-41c9-9413-bbb0ec47cca2","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":64687,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/3b3cfc002c8b2b580c89e8c5.png"},{"id":93765555,"identity":"dd7b70f4-9194-4c69-b678-c751a99d797e","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":44851,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/27cd83498d0b4613897eafda.png"},{"id":93765558,"identity":"31d96ded-5018-4ae0-87d9-4294bb434383","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":57186,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/c22bf8c713c4d4a53f4a25b6.png"},{"id":93765167,"identity":"f255151f-e4d6-472d-b9da-d7ee706435e7","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"png","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":272,"visible":true,"origin":"","legend":"","description":"","filename":"Onlineimage26.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/29346176338315c9c96f6dfe.png"},{"id":93765557,"identity":"585a8274-7d5d-4665-870f-2257d2f6fcae","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"png","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":220,"visible":true,"origin":"","legend":"","description":"","filename":"Onlineimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/9eec77be1270fcbac4377845.png"},{"id":93765169,"identity":"ceaf06ee-fd0e-4f8d-8186-c06aa44f26a2","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"xml","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":106479,"visible":true,"origin":"","legend":"","description":"","filename":"ee6c09ee2f5741459f01368cee7f3da21structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/e78158248ae0be3bcde0808c.xml"},{"id":93765171,"identity":"c97c3762-c1a8-40e2-a9e9-1b7245c0a619","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"html","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":117196,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/4b5abe9de25e1a6f50046391.html"},{"id":93765151,"identity":"8c9722e2-bdb7-4dd6-b1ec-ebefa8b197f6","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":324510,"visible":true,"origin":"","legend":"\u003cp\u003eWorkflow of CMC-WDTK. In this study, CpG sites exhibiting allele-specific methylation at SNVs on heterogeneous homologous chromosomes were selected to predict methylation changes between reference DNA sequences flanking the CpG sites and the corresponding alternative DNA sequences (A). Tissue and cell line datasets were used, with separate models constructed for each dataset. The reference and alternative sequences flanking the CpG sites were used as inputs, from which k-mer sequences with a step size of 1 were extracted (B). The dual-branch Transformer processes reference and alternative sequences simultaneously. It tokenizes them into k-mers, embeds them into dense vectors, and extracts features through shared Transformer encoders to capture global context and local SNV-induced variations (C). The KAN processes the features from the Transformer and predicts methylation changes, with values near 1 indicating methylation gain and values near 0 indicating methylation loss (D).\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/d39a6b73eee1c7f293ff2f72.png"},{"id":93766469,"identity":"38e0419a-afa4-40c0-afef-b6897f266011","added_by":"auto","created_at":"2025-10-17 10:41:42","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":211814,"visible":true,"origin":"","legend":"\u003cp\u003eCpG information. The number of CpG sites in each dataset (A). Distance of associated SNVs relative to CpG sites, showing the genomic coordinate of the CpG site minus the coordinate of the SNV (B). The frequency of 12 different types of related SNVs in each dataset (C). Genomic distribution of all the sites used in each dataset(D).\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/765462131a4b877310dd7549.png"},{"id":93765157,"identity":"a5c936e2-74cd-43cc-9b4d-53f5a04b4f95","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":323955,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance evaluation. AUC performance for the tissue and cell line datasets (A). AUPR performance for the tissue and cell line datasets (B). Performancesof all the metrics (C, D). UMAP visualization for the distribution of the feature representations learned by each model in the feature space (E).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/93a17aa3725d2634a7ec8fc2.png"},{"id":93765154,"identity":"48931dfc-2c10-4143-b0ca-494fa95b3652","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":198378,"visible":true,"origin":"","legend":"\u003cp\u003eIndependent validation. Tissue datasets are independently validated across different tissue datasets, and cell line datasets are independently validated across different cell line datasets.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/a5bf4edb46633cc2d9a87fdc.png"},{"id":93765553,"identity":"e000af42-f49f-464a-ace1-01e25cb5ca01","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":198677,"visible":true,"origin":"","legend":"\u003cp\u003eAttention scores differ at SNV sites between the reference and alternative sequences. Six CpG sites exhibiting allele-specific methylation (with rs6671482, rs10910008, rs11580811, rs3765742, and rs61735051 originating from the intestine dataset and rs751035 from the gastric tissue dataset), located within the introns or exons of the tumor suppressor gene TP73 (A). The differential attention matrix, shown in the upper panel of (B), displays the differences between the variant and the reference sequence in the 21×21 k-mer local attention matrix around the SNV, with the 8th to 11th k-mers covering the SNV site. The vector below represents the result of summing the upper matrix by column, characterizing the overall contribution of each k-mer position to the attention difference.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/29d533fd37b2e27efb4a67c9.png"},{"id":93766470,"identity":"dce84ead-9898-44d6-b319-ed869ba9e9c1","added_by":"auto","created_at":"2025-10-17 10:41:42","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":190109,"visible":true,"origin":"","legend":"\u003cp\u003eMotif identification. Probability-based motif logos for sequences surrounding 20 bp upstream and downstream of the CpG sites were generated using KpLogo. The logos were generated individually from three sets of DNA sequences: the reference sequence group (top), the variant group predicted by CMC-WDTK to gain methylation (middle), and the variant group predicted to lose methylation (middle).\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/f6b3ba8df33cbc4421a19619.png"},{"id":102235032,"identity":"4c85ce91-a326-4c63-a206-4de68e9dedec","added_by":"auto","created_at":"2026-02-09 16:14:53","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2270464,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/d35aa833-aaff-48ee-8694-35b5c966d15f.pdf"},{"id":93765550,"identity":"81dbaa2d-6cb5-4d7b-93a3-d004c6abf956","added_by":"auto","created_at":"2025-10-17 10:33:42","extension":"xlsx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":10949,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1Datasetsinfo.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/ad2e22b10d2b02d2b796ee4b.xlsx"},{"id":93765149,"identity":"ab8defd2-dd86-4254-8009-131cdae24417","added_by":"auto","created_at":"2025-10-17 10:25:42","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":9626,"visible":true,"origin":"","legend":"","description":"","filename":"TableS2Samplestasitics.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-7586988/v1/5366bea92af24316fc0bb44d.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"CMC-WDTK: CpG methylation change prediction by a weight-sharing dual-branch Transformer-Kolmogorov–Arnold network model","fulltext":[{"header":"1 Introduction","content":"\u003cp\u003eDNA methylation regulates many cellular processes, with a particularly crucial role in gene transcription, chromatin structure, X chromosome inactivation, and genomic imprinting[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. It has also emerged as a promising biomarker for the diagnosis of various diseases, particularly cancer[\u003cspan additionalcitationids=\"CR3\" citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Compared with those in normal tissues, the promoters of many tumor suppressor genes, such as the retinoblastoma suppressor gene (RB1)[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e], the cell cycle regulators p16INK4a (CDKN2A) and p15INK4a (CDKN2B) [\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], the apoptosis signaling pathway regulator DAPK1[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], and various transcription factors[\u003cspan additionalcitationids=\"CR11\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e], are significantly hypermethylated in cancer tissues[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e]. In contrast, high expression of oncogenes has also recently been reported to be associated with DNA demethylation. For example, several novel hypomethylated regions in intergenic areas have been linked to the expression of androgen receptor (AR), the oncogene MYC, and the transcription factor ERG in cancer cells. These regions contain binding motifs for transcription factors, including HOXB13, FOXA1, AR, and ERG, and are colocalized with enhancer signals, such as H3K27ac [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. Although abnormal DNA methylation and its downstream molecular functions have been extensively studied, the mechanisms underlying the regulation of aberrant DNA methylation have not been fully elucidated. The development of predictive models for changes in DNA methylation levels using bioinformatics approaches is essential for uncovering these mechanisms and understanding the role of DNA methylation in tumorigenesis.\u003c/p\u003e\u003cp\u003eEmerging evidence suggests that genetic variations, especially SNVs, play crucial roles in modulating DNA methylation changes at CpG sites adjacent to these variations [\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] [\u003cspan additionalcitationids=\"CR17\" citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. For example, the T allele at rs1001179 in the CAT promoter alters promoter methylation at specific CpG sites, which increases the levels of ETS-1 and GR-β transcription factors (TFs) and leads to increased CAT expression in chronic lymphocytic leukemia (CLL), a change associated with a more aggressive clinical course[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. Another study demonstrated that rs7247241 within the PPP1R14A promoter leads to a decrease in the transcriptional activity of PPP1R14A, with the T allele associated with increased CpG methylation and reduced binding of the transcription factors CTCF and PLAGL2 compared with the C allele[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. Furthermore, high CpG methylation in the intergenic region, together with the risk allele of the nearby SNP rs11986220, synergistically disrupts CTCF-mediated insulating loop formation, resulting in the loss of insulation. This enables the cis-regulatory element containing rs11986220 to upregulate the oncogenes MYC and PVT1, thereby increasing the risk of cancer[\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. These studies indicate that changes in the methylation levels of CpG sites are sensitive to variations in nearby SNVs and that their synergistic interaction regulates the transcription of related genes. This synergistic relationship is widespread across the genome, making it crucial to fully consider SNVs when predicting changes in CpG methylation levels for accurate results. Allele-specific methylation (ASM) analysis uses bisulfite sequencing (BS-Seq) data to identify differences in CpG methylation between alleles of adjacent heterozygous SNV sites in homologous genomes, providing an effective strategy to identify CpGs impacted by SNVs[\u003cspan additionalcitationids=\"CR23\" citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. CpG sites exhibiting ASM events are ideal loci for modeling DNA methylation alterations.\u003c/p\u003e\u003cp\u003eAlthough traditional experimental methods, such as whole-genome bisulfite sequencing (WGBS), provide accurate detection of genome-wide methylation states and identification of methylation changes at CpG sites, these techniques are costly and time-consuming. In recent years, deep learning based on natural language processing (NLP) has emerged as a powerful computational tool, overcoming the limitations of traditional machine learning approaches by reducing dependence on well-annotated genomes, and consequently has been increasingly applied to model and predict complex biological data[\u003cspan additionalcitationids=\"CR26 CR27 CR28\" citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e]. Notably, NLP based deep learning techniques have the ability to learn directly and extract high-level features from DNA sequences, enabling accurate prediction of DNA methylation even in the absence of genomic annotation data [\u003cspan additionalcitationids=\"CR31 CR32 CR33\" citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]. Transformer-based models, in particular, have demonstrated exceptional performance in DNA methylation prediction owing to their capacity to model sequence context and capture long-range dependencies. For example, iDNA-ABF effectively identifies methylation-associated sequence patterns through a multiscale deep learning model [\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e], StableDNAm significantly enhances model stability and generalizability using feature fusion and contrastive learning techniques[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e], and Methyl-GP improves generalization by examining sequence pattern conservation across species with a BERT-based language model[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. While these computational approaches have significantly increased the accuracy of DNA methylation prediction and expanded its applications, a major limitation remains: most methods still rely solely on a background reference genome and fail to account for the local impact of DNA sequence variations on CpG site methylation.\u003c/p\u003e\u003cp\u003eInspired by this, we propose CMC-WDTK, an innovative deep learning framework focused on predicting methylation level changes between reference and alternate sequences for CpG sites with ASM events and identifying sequential patterns affecting DNA methylation. CMC-WDTK uses a dual-branch Transformer architecture with a weight-sharing mechanism to simultaneously extract features from reference and variant sequences surrounding CpGs, capturing global context features and local disturbances induced by SNVs. Furthermore, the framework incorporates a Kolmogorov‒Arnold network (KAN)[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e] to further explore high-order nonlinear relationships between features, enhancing the model\u0026rsquo;s ability to express and generalize methylation change trends. Experimental validation on eight real-world datasets demonstrated that CMC-WDTK offers significant advantages in predicting the direction and magnitude of methylation changes induced by SNPs. This study effectively predicts, for the first time, CpG site methylation changes between DNA sequences by integrating SNVs and DNA sequences. It identifies the repeated C and G sequence patterns that promote increased methylation, offering a novel and efficient computational tool for understanding the mechanisms underlying DNA methylation alterations.\u003c/p\u003e"},{"header":"2 Materials and methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1 Dataset\u003c/h2\u003e\u003cp\u003eFor model training and evaluation, datasets were selected from a previous study conducted by our group, in which we identified ASM sites using BS-Seq datasets across 31 cancer types and their corresponding normal tissue samples[\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. ASM analysis employs Fisher's exact test to evaluate the significance of differences in the counts of methylated and unmethylated reads covering the reference and variant alleles of nearby heterogeneous SNVs in a BS-Seq sample. This determines whether CpG site methylation levels exhibit allele specificity (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA)[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. A total of eight datasets were selected for modeling ASM CpG sites, encompassing four tissue types (esophageal, gastric, intestinal, and pancreatic) and four cancer cell lines (MCF7, HepG2, OCI-LY7, and HeLa) (Suppl. Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). Tissue datasets are a mixed collection of cancerous and normal samples. The obtained ASM CpG information includes the genomic coordinates of the CpGs and SNVs, the reference and alternate alleles of the SNVs, and the total and methylated read counts for each allele. Suppl. Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e provides a summary of the total number of CpG samples in each dataset, along with the corresponding counts of positive and negative samples.\u003c/p\u003e\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eB, for each CpG, we first extracted a 101 bp DNA sequence centered on the CpG site, consisting of 50 bases upstream and 50 bases downstream from the reference genome (GRCh38), denoted as \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(se{q_r}\\)\u003c/span\u003e\u003c/span\u003e. Then, according to the SNV genomic position information, the corresponding base in the reference sequence is substituted with the alternative allele to generate the alternative sequence \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(se{q_a}\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e\u003cp\u003eThe change in methylation between \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(se{q_r}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(se{q_a}\\)\u003c/span\u003e\u003c/span\u003e of the CpG site \u003cspan class=\"InlineEquation\"\u003e\u003c/span\u003e is defined as:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$~VA{R_i}={\\beta _i}(se{q_a}) - {\\beta _i}(se{q_r})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\beta _i}(se{q_a})\\)\u003c/span\u003e\u003c/span\u003e refers to the proportion of methylated reads out of the total reads mapped to the alternative allele and where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\beta _i}(se{q_r})\\)\u003c/span\u003e\u003c/span\u003e refers to the proportion of methylated reads out of the total reads mapped to the reference allele. The true class label for CpG \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:i\\)\u003c/span\u003e\u003c/span\u003e is 1 (positive) if \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(VA{R_i}\\)\u003c/span\u003e\u003c/span\u003e is greater than 0; otherwise, it is 0 (negative).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2 Methods\u003c/h2\u003e\u003cp\u003eThis study proposes an end-to-end deep learning framework for CpG contextual sequences containing SNVs. It employs a weight-sharing, dual-branch Transformer encoder to extract global and local features concurrently from both the reference sequence and the alternative sequence. A two-layer KAN subsequently maps the high-dimensional vectors to continuous methylation level predictions. WDTK accurately captures methylation fluctuations at neighboring CpG sites jointly induced by SNVs and their surrounding sequences. The organic combination of Transformer and KAN achieves high-precision prediction of methylation gain and loss types. Specifically, WDTK first uses a shared embedding layer to split each of the two allele sequences (\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{seq}_{r}\\)\u003c/span\u003e\u003c/span\u003e/\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{seq}_{a}\\)\u003c/span\u003e\u003c/span\u003e) into k-mer subsequences and project them into the same d-dimensional vector space. Both sequences are then passed in parallel through identical Transformer encoders for deep feature extraction: in each layer, the multi-head self-attention module adaptively captures long-range dependencies and local microvariations within the sequence, and the subsequent feed-forward network\u0026mdash;consisting of two linear transformations with ReLU activation\u0026mdash;further enriches the nonlinear representation. All sublayers employ residual connections and layer normalization to ensure stable gradients. Finally, the output tensor of shape (B, L, d) is transposed and compressed via adaptive average pooling into a fixed-length (B, d) vector. This preserves the key positional interactions highlighted by multi-head attention while retaining the global context. We subsequently concatenate the two branch pooling vectors into (B, 2d), which are then fed into a two-layer KAN network. First, a linear mapping of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\((2d \\to 1d)\\)\u003c/span\u003e\u003c/span\u003e with ReLU activation performs dimension compression and feature fusion. Then, a linear layer of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\((d \\to 1)\\)\u003c/span\u003e\u003c/span\u003e directly outputs continuous methylation level prediction scores. The entire framework is end-to-end, optimized with AdamW, and incorporates appropriate dropout and learning rate scheduling, further enhancing the model's generalization performance and biological interpretability.\u003c/p\u003e\u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\u003ch2\u003e2.2.1 Dual-branch Transformer\u003c/h2\u003e\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eC, the aforementioned reference sequence and variant sequence are first segmented into continuous k-mer fragments to extract local contextual information. These k-mers are subsequently treated as discrete \"vocabularies\" and fed into a learnable embedding layer, which maps them into fixed-length dense vectors. This process captures the potential semantic features of the sequences and provides efficient vectorized representations for subsequent models, denoted as\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X_r}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X_a}\\)\u003c/span\u003e\u003c/span\u003e. Subsequently, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X_r}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({X_a}\\)\u003c/span\u003e\u003c/span\u003e are fed in parallel into a shared-weight Transformer encoder, where multi-head self-attention and feed-forward layers extract their respective context-aware feature representations:\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$${H_i}=Transformer\\left( {{X_{\\mathbf{i}}}} \\right).i \\in \\left\\{ {r,a} \\right\\}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAfter the sequences are processed through the Transformer encoder, the resulting feature representations \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({H_i}\\)\u003c/span\u003e\u003c/span\u003e are passed through an adaptive average pooling layer to reduce their dimensionality. This pooling operation compresses the sequence information into fixed-length vectors of size 1, each encoding the overall methylation pattern of the sequence. The pooling is applied separately to each \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({H_i}\\)\u003c/span\u003e\u003c/span\u003e:\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$${P_i}=AdaptiveAvePooild\\left( {{H_i}} \\right).i \\in \\left\\{ {r,a} \\right\\}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe pooling layer compresses each sequence\u0026rsquo;s representation into a single scalar, preserving its most salient information. The pooled scalars from the reference and alternative sequences are then concatenated along the feature dimension to form a joint representation that captures their interrelationship.\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$${P_{concat}}=cat\\left( {{P_i},{\\kern 1pt} {\\kern 1pt} {\\kern 1pt} \\dim =1} \\right),i \\in \\left\\{ {r,a} \\right\\}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe concatenated representation is then fed into the KAN module for further processing and feature fusion. The KAN module consists of fully connected layers that are capable of transforming the concatenated features into the final output. The final output is a scalar, reflecting the methylation analysis result between the reference and variant sequences.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section3\"\u003e\u003ch2\u003e2.2.2 Kolmogorov\u0026ndash;Arnold network\u003c/h2\u003e\u003cp\u003eAs shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eD, upon receiving the feature vector extracted and fused by the dual-branch Transformer, the module undergoes nonlinear transformation followed by linear mapping through two layers of learnable mapping functions. This process extracts key features from the high-dimensional, context-aware representation, ultimately generating logit outputs that can be used for determining methylation gain or loss. The KAN consists of two layers of learnable mapping functions \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\varphi ^{\\left( l \\right)}}\\)\u003c/span\u003e\u003c/span\u003e, and its computation process is as follows:\u003c/p\u003e\u003cp\u003eFirst Layer:\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$z_{q}^{{\\left( 1 \\right)}}=\\sum\\limits_{{p=1}}^{{{n_0}}} {\\varphi _{{q,p}}^{{\\left( 1 \\right)}}} \\left( {{p_{concat,p}}} \\right),{\\kern 1pt} {\\kern 1pt} {\\kern 1pt} {\\kern 1pt} q=1,...,{n_1}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eSecond Layer:\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$y=\\sum\\limits_{{q=1}}^{{{n_1}}} {\\varphi _{{1,q}}^{{\\left( 2 \\right)}}} \\left( {z_{q}^{{\\left( 1 \\right)}}} \\right).$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eOverall, the KAN mapping can be expressed as the composition of the two mapping functions:\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$KAN\\left( x \\right)=({\\Phi ^{\\left( 2 \\right)}} \\circ {\\Phi ^{\\left( 1 \\right)}})(x),{\\kern 1pt} {\\kern 1pt} {\\kern 1pt} {\\kern 1pt} {\\kern 1pt} {\\kern 1pt} {\\Phi ^{\\left( l \\right)}}{(x)_j}=\\sum\\limits_{{i=1}}^{{{n_{l - 1}}}} {\\varphi _{{j,i}}^{{\\left( l \\right)}}} ({x_i}),{\\kern 1pt} {\\kern 1pt} l=1,2.$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe model\u0026rsquo;s final output is the result produced by the KAN module. The predicted value \u003cspan class=\"InlineEquation\"\u003e\u003c/span\u003e represents the type of methylation change (0, 1) between the variant sequence relative to the reference sequence or between DNA sequences from different individuals. A value closer to 1 indicates a greater increase, whereas a value closer to 0 indicates a greater decrease.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section3\"\u003e\u003ch2\u003e2.2.3 Optimization\u003c/h2\u003e\u003cp\u003eIn this study, model training was optimized using standard supervised learning techniques. For the binary classification task, we employed binary cross-entropy loss, which measures the discrepancy between the model\u0026rsquo;s predictions and the true labels. The formula for binary cross-entropy loss is as follows:\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$L= - \\frac{1}{N}\\sum\\limits_{{i=1}}^{N} {\\left[ {{y_i}\\log ({{\\widehat {y}}_i})+(1 - {y_i})\\log (1 - {{\\widehat {y}}_i})} \\right]} .$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eHere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({y_i}\\)\u003c/span\u003e\u003c/span\u003e denotes the true label, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\widehat {y}_i}\\)\u003c/span\u003e\u003c/span\u003e denotes the model\u0026rsquo;s predicted value. To optimize the model parameters, we employed the AdamW optimizer. AdamW is an adaptive learning-rate method that integrates the concepts of momentum and RMSprop. During training, it effectively adjusts the learning rate for each parameter, making it particularly suitable for handling sparse gradient problems. In the optimization process, the AdamW optimizer uses the following update rule:\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$${\\theta _t}={\\theta _{t - 1}} - \\eta \\cdot \\left( {\\frac{{{{\\widehat {m}}_t}}}{{\\sqrt {{{\\widehat {v}}_t}} +\\varepsilon }}+\\lambda {\\theta _{t - 1}}} \\right).$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eHere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\theta _t}\\)\u003c/span\u003e\u003c/span\u003e denotes the parameter at the current timestep, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\theta _{t - 1}}\\)\u003c/span\u003e\u003c/span\u003e is the parameter at the previous timestep, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\eta\\)\u003c/span\u003e\u003c/span\u003e is the learning rate, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\widehat {m}_t}\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\({\\widehat {v}_t}\\)\u003c/span\u003e\u003c/span\u003e are the bias-corrected estimates of the first and second moments of the gradients, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\varepsilon\\)\u003c/span\u003e\u003c/span\u003e is a small constant to prevent division-by-zero errors, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\lambda\\)\u003c/span\u003e\u003c/span\u003e is the weight decay coefficient.\u003c/p\u003e\u003cp\u003eAfter each forward pass, the model calculates the loss function, and this loss is then propagated back through each layer of the model via backpropagation to compute the gradient for each parameter. The optimizer subsequently updates the model's parameters on the basis of the calculated gradients. With each update, previous gradients are first cleared, new gradients are computed through backpropagation, and finally, the parameter values are adjusted according to the optimizer's rules.\u003c/p\u003e\u003cp\u003eThe gradients are cleared after each update to prevent accumulation, allowing successive updates to guide the model toward its optimum. Through multiple gradient updates, the model is progressively optimized, approaching the optimal solution. To enhance the model's generalization ability and prevent overfitting, we incorporated the Dropout technique during training. Dropout, by randomly dropping a subset of neurons during each training step, reduces the model's overreliance on specific neurons, thereby improving its performance on unseen data. The use of Dropout contributes to enhancing the model's stability and generalizability.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e2.3 Performance assessment\u003c/h2\u003e\u003cp\u003eTo assess the performance of the CMC-WDTK, we applied fivefold cross-validation in the different experiments. The following four commonly used metrics were employed to evaluate our method in models for different experimental datasets: area under the receiver operating characteristic curve (AUC), area under the precision-recall curve (AUPR), accuracy (ACC), and F1 score. The computation of these metrics is as follows:\u003cdiv id=\"Equ10\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ10\" name=\"EquationSource\"\u003e\n$$\\begin{gathered} Recall=\\frac{{TP}}{{TP+FN}}, \\hfill \\\\ Specificity=\\frac{{TN}}{{TP+FP}}, \\hfill \\\\ Precision=\\frac{{TP}}{{TP+FP}}, \\hfill \\\\ F1=\\frac{{2*Recall*Precision}}{{Recall+Precision}}, \\hfill \\\\ ACC=\\frac{{TP+TN}}{{TP+TN+FP+FN}}. \\hfill \\\\ \\end{gathered}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(TP\\)\u003c/span\u003e\u003c/span\u003e, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(TN\\)\u003c/span\u003e\u003c/span\u003e,\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(FP\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(FN\\)\u003c/span\u003e\u003c/span\u003e represent the number of true-positive, true-negative, false-positive, and false-negative predicted CpGs, respectively. We used the receiver operating characteristic (ROC) curve and the precision‒recall (PR) curve to evaluate the overall performance of the model[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. These curves provide an intuitive visualization of the model\u0026rsquo;s superiority.\u003c/p\u003e\u003c/div\u003e"},{"header":"3 Results","content":"\u003cp\u003e\u003cb\u003eAllele-specific methylated CpG sites are widely distributed in tissues and cell lines\u003c/b\u003e\u003c/p\u003e\u003cp\u003eIn this study, CpG sites exhibiting ASM events from eight datasets were used for model training and evaluation. The number of CpG sites in each dataset ranged from 14,551 to 353,746, each affected by one or more SNVs (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eA, Suppl Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e). In this study, we set the DNA sequence length to 50 bp upstream and downstream of the CpG sites to extract global features. To better explore the local impact of SNVs on CpG site methylation levels, we calculated the distance of associated SNVs relative to CpG sites. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eB, the SNVs are predominantly located within 20 bases upstream and downstream of the CpG sites, with more than 98% of them being SNPs (dbSNP build 156 release). For all the SNVs, we assessed the frequency of 12 different types of nucleotide substitutions and found that C/T and G/A substitutions occurred more frequently than other substitution types did (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eC). This finding suggests that CpG methylation is more sensitive to these two substitutions in the surrounding nucleotides. Additionally, we utilized ANNOVAR [\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] to annotate these SNVs to the GRCh38 reference genes, using genomic annotation information for reference genes obtained from the UCSC database[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The results showed that approximately half of the ASM sites across the eight datasets occurred within or near genes, with the remaining sites located in intergenic regions (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003eD). These findings suggest that these SNV-affected CpG sites may serve as potential functional sites for regulating gene expression. Therefore, focusing on these sites for modeling, rather than all CpG sites across the genome, enhances the identification of DNA sequence patterns influencing methylation.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eExperiment Design\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo comprehensively evaluate the model\u0026rsquo;s performance, two sets of experiments were designed: a tissue group and a cell line group.\u003c/p\u003e\u003cp\u003eWithin-cohort validation: To assess the model's performance on a specific tissue or cell line, we conducted experiments on eight real-world datasets.\u003c/p\u003e\u003cp\u003eCross-cohort validation: To further test generalizability, we adopted a cross-dataset scheme within each group. Specifically, we trained the model on one tissue type(e.g., gastric) and then evaluated its performance on a separate, independent tissue dataset (e.g., pancreas), simulating the model\u0026rsquo;s behavior when confronted with novel tissues or unseen cell types in practical applications. By comparing performance across cross-dataset and within-cohort experiments, we obtain an objective measure of the model\u0026rsquo;s transferability and robustness.\u003c/p\u003e\u003cp\u003eAll the experiments were conducted using 5-fold cross-validation to assess the model\u0026rsquo;s performance accurately. The proposed framework was implemented in PyTorch 1.13.1, and both training and evaluation were conducted on a high-performance server equipped with an NVIDIA RTX 4090 GPU. We also employed Bayesian optimization to jointly search and automatically tune multiple key hyperparameters, using the AUC as the objective metric. The optimal hyperparameter set is summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eThe optimal hyperparameter set.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eParameter\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRange of Values\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eOptimal Value(cell line)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eOptimal Value(tissue)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLearning rate\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e1e-4, 1e-3, 1e-2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e1e-3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1e-4\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eHidden layer dimension\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[64, 256]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e64\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNumber of attention heads\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e4, 8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNumber of Transformer layers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e3, 4, 5, 6, 7, 8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eTraining epochs\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e[64, 256]\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e256\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e256\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eBatch size\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e128, 256, 512\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e256\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e128\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003ek-mer\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e3, 4, 5, 6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e4\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eDropout\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e0.1, 0.3, 0.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.3\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003ePerformance evaluation of CMC-WDTK\u003c/b\u003e\u003c/p\u003e\u003cp\u003eWe first plotted the ROC (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eA) and PR curves (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eB) for all the models in different tissues and cell lines separately. All the AUC values are greater than 0.8, reaching above 0.9, indicating good overall model performance in different models, whereas the AUPR is approximately 0.6, reflecting a reasonable ability to identify the positive class. For the cell line models, the ACC ranges from above 0.6 to above 0.7, indicating acceptable model performance. However, the F1 score is approximately 0.6, which suggests a reasonable balance between precision and recall (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eC). In the tissue models, the ACC ranged from above 0.7 to above 0.8, showing a slight increase compared with that in the cell line model (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eD). For datasets with better performance, such as those for gastric and pancreatic tissues, the cancer and normal samples are balanced in the data source. In contrast, in the other models, the proportion of cancer samples is greater, or the dataset consists entirely of cancer samples (Suppl. Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). This may result from complex dynamic changes in global DNA methylation during cancer development[\u003cspan additionalcitationids=\"CR41\" citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eTo present a more intuitive view of how models make decisions, we used uniform manifold approximation and projection (UMAP, a widely used visualization tool) to visualize the distribution of the feature representations learned by each model in the feature space (UMAP). The feature representations learned by most of the models, such as those for gastric tissue and the HeLa cell line, generally have clear boundaries and are well clustered (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003eE). In other words, our model clearly separates the CpG sites with increased methylation and decreased methylation due to variation and every cluster together rather than dispersing them. This finding indicates that CMC\u003cb\u003e-\u003c/b\u003eWDTK has the potential to make accurate predictions for different experimental datasets.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eCross-dataset validation\u003c/b\u003e\u003c/p\u003e\u003cp\u003eTo assess whether the learned sequence patterns are applicable for methylation prediction across different datasets, we performed cross-model validation to determine whether a model trained on one tissue or cell line dataset can achieve satisfactory results on testing sets from other tissue or cell line datasets. Specifically, we conducted separate cross-model validation for the tissue and cell line datasets, evaluating the models using four metrics: AUC, AUPR, accuracy, and F1 score. The results show that the cross-model validation outcomes are highly consistent across tissue and cell line datasets for all the evaluation metrics (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). This suggests that models trained on one tissue or cell line dataset can generalize well to other datasets, indicating that the sequence patterns captured by CMC-WDTK significantly contribute to methylation alteration prediction. Furthermore, these results highlight the model's strong generalizability.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eThe importance of SNVs in the model's decision-making process\u003c/b\u003e\u003c/p\u003e\u003cp\u003eThe attention scores of the Transformer model reflect the weight of the model's focus on features at different positions in the sequence during decision-making. To further explore the model's decision-making mechanism and clarify the importance of SNVs in their local effects on CpG site methylation prediction, this study analyzes the attention scores to investigate how the model's focus on features at SNV sites differs between reference and variant sequences. We will then illustrate this using TP73 as an example.\u003c/p\u003e\u003cp\u003eBy analyzing sequences from the testing set, we identified six CpG sites located within the introns or exons of the tumor suppressor gene \u003cem\u003eTP73\u003c/em\u003e, each associated with a corresponding ASM SNP. These SNPs are located in close proximity to the CpG sites at distances of 5, 1, 40, 0, and 1 bp and include rs751035, rs6671482, rs10910008, rs11580811, rs3765742, and rs61735051 (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eA). Among them, rs751035 comes from the gastric tissue dataset, and the others are from the intestinal tissue dataset. According to the Fisher\u0026rsquo;s test statistics for SNP and CpG results from the CanASM database, methylation significantly decreases at the variant alleles for all six of these CpG sites, meaning that their true labels are all negative. Simultaneously, the corresponding model predictions show that their predicted labels are consistently negative, matching the true labels.\u003c/p\u003e\u003cp\u003eThe core task of this study is to capture the global contextual DNA sequence around CpG sites and local SNV features to predict the variation trend of CpG site methylation levels. To quantify the importance of SNVs in the model decision-making process, self-attention score analysis was subsequently performed on above 6 SNV sites. By calculating the attention scores, we obtained a matrix representing the attention scores between all tokens for the reference and alternative sequences. Here, each token represents a k-mer extracted from the DNA sequence. First, the model calculates the self-attention scores of the reference and the variant sequence respectively. The specific process is as follows: the two sequences are input into the model separately, and the self-attention matrix output by the last Transformer block is extracted. This matrix is obtained by averaging the results of independent calculations from multiple attention heads, with a dimension of L\u0026times;L (where L is the number of k-mers), and is used to characterize the pairwise correlation strength between all k-mers in the sequence. In this study, the sequence length was set to 101 bp, and a sliding window with a step size of 1 was used to segment the sequence into k-mers, where k\u0026thinsp;=\u0026thinsp;4, resulting in L\u0026thinsp;=\u0026thinsp;98. Since the step size was 1, each SNV is covered by k consecutive k-mers. Subsequently, local attention submatrices for each SNV were extracted. The last k-mer covering the SNV was taken as the center of the extracted the submatrices, with 10 k-mers selected upstream and downstream of this center to form a 21\u0026times;21 matrix. Finally, the difference between the corresponding submatrices of the variant and the reference sequence was calculated (by subtracting the two matrices) to generate a differential attention score matrix, based on which a heatmap was plotted (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003eB). The results suggest that the model can effectively focus on local sequence patterns near SNVs, accurately identifying their potential impact on methylation changes, thereby validating the reasonableness and specificity of the model's decisions.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003cb\u003eCMC-WDTK helps capture global sequential patterns in methylation alteration\u003c/b\u003e\u003c/p\u003e\u003cp\u003eFinding the underlying sequence patterns surrounding the CpG sites is an effective step in understanding the reasons for their modification and in revealing the biological functions of modifications[\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e]. To this end, this study divided the sequences near CpG sites into three groups based on SNV variations and CpG methylation levels: the reference sequence group, the variant group predicted by CMC-WDTK to gain methylation, and the variant group predicted to lose methylation to identify the sequence features that influence DNA methylation levels by comparing these three groups of sequence motifs. The specific implementation process is as follows. First, 41 bp DNA sequences were extracted with CpG sites as the center, encompassing 20 bp of sequence both upstream and downstream of the CpG sites. This 41 bp length was selected because it can effectively capture key information associated with DNA methylation detection[\u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e44\u003c/span\u003e, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e45\u003c/span\u003e]. Next, the KpLogo tool[\u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e46\u003c/span\u003e] was employed to generate motif logo plots for each group of sequences, with the plots constructed based on the occurrence probabilities of the four nucleotide types at each position within the sequences (Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e).\u003c/p\u003e\u003cp\u003eAfter comparing the motif logo plots generated for the three groups of sequences in each dataset, it was found that the motifs for the three corresponding groups exhibited consistent sequence features across datasets. The reference sequence group and the increased methylation group both showed highly conserved C and G motifs near the target CpG sites. In contrast, in the decreased methylation group, this conserved C and G motif tended to be replaced by T or A. This finding not only indicates that CpG site methylation across diverse tissues and cell lines may follow a shared sequence logic centered on conserved C and G motifs, but also offers a model-supported approach to unravel the common molecular mechanisms driving methylation alterations at the DNA sequence level.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eDNA methylation is a crucial epigenetic modification involved in the regulation of gene expression and various biological processes. To date, several methods have been proposed to predict DNA methylation and identify sequence patterns. While these methods have provided valuable insights into methylation patterns, they typically rely on reference genomes, which do not account for the genetic variations that exist between individuals and diverse populations. In contrast, CMC-WDTK incorporates SNV-induced variations with sequential contexts, which allows for precise predictions of methylation changes driven by local genetic diversity and global sequence features. We present CMC-WDTK, an innovative deep learning framework designed to predict methylation changes at CpG sites by integrating sequence features of CpG sites and adjacent single nucleotide variants. The framework combines a dual-branch Transformer architecture with a Kolmogorov-Arnold Network (KAN), enabling it to capture the global sequence context and local variation features. This approach allows for accurate prediction of methylation changes between reference and variant sequences, with a particular focus on understanding the impact of SNP variations on CpG site methylation levels. Additionally, cross-model validation revealed that models trained on one tissue or cell line dataset generalize well to others, indicating the model\u0026rsquo;s ability to perform consistently across different biological contexts.\u003c/p\u003e\u003cp\u003eFurthermore, attention score analysis revealed that the model\u0026rsquo;s dual-branch architecture assigns greater differences in attention scores to local sequences near SNVs, emphasizing the critical role of the local variant sequence context in predicting methylation changes. Additionally, our study identified specific sequence motifs associated with changes in CpG methylation. Using KpLogo, we found that sequences associated with increased methylation are characterized by repeating C and G motifs, whereas sequences showing decreased methylation exhibit substitutions of C and G with A and T. These patterns were consistent across different datasets, providing further validation of the model\u0026rsquo;s ability to capture biologically relevant sequence features.\u003c/p\u003e\u003cp\u003eDespite the promising results of CMC-WDTK, several potential limitations exist. First, while the eight datasets used in this study cover different tissue types and cell lines, there may still be an underrepresentation of other genomic backgrounds. Therefore, future research could expand the datasets. Second, while our approach has shown excellent performance in predicting methylation changes between variant and reference sequences, its application has yet to be fully explored in disease versus normal group comparisons. Future work will focus on extending CMC-WDTK to analyze methylation differences between disease and normal groups, further validating its potential for broad biological and clinical applications.\u003c/p\u003e\u003cp\u003eOverall, by integrating DNA sequence features and genomic variations, CMC-WDTK is the first computational tool to predict methylation changes between sequences, offering a significant advancement in the field of DNA methylation comparison. The identified sequence pattern of repeating C and G motifs, which is associated with increased methylation, also helps in understanding the mechanisms underlying DNA methylation. The model\u0026rsquo;s ability to generalize across different datasets and biological conditions further enhances its potential for broad applications in comparing the methylation levels between two DNA sequences.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003cp\u003eNot applicable\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003cp\u003eNot applicable\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e\u003cp\u003eThis work was supported by the National Natural Science Foundation of China (Grant Nos. 62225109, 62402345, and 62302342). We would like to express our sincere thanks to the funding agencies for their financial support, which made this research possible.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eJZ: Conceptualization, Methodology, Formal analysis, Data curation, Writing. DL: Methodology, model development. YW: Methodology. HL: Conceptualization, Investigation. GW: Conceptualization, Supervision, Funding acquisition, Resources. All authors reviewed the manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e\u003cp\u003eNot applicable\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe original Whole-Genome Bisulfite Sequencing (WGBS) data supporting the findings of this study have been deposited in the NCBI GEO database under the primary accession codes provided in Supplementary Table S1.The preprocessed input data obtained from whole-genome bisulfite sequencing (WGBS), together with the model code used in this study, are available at: https://github.com/LiuDiBio/CMC-WDTK.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eMeng, H., et al., \u003cem\u003eDNA methylation, its mediators and genome integrity.\u003c/em\u003e Int J Biol Sci, 2015. \u003cstrong\u003e11\u003c/strong\u003e(5): p. 604-17.\u003c/li\u003e\n\u003cli\u003eLiu, J., et al., \u003cem\u003eEvaluating the comprehensive diagnosis efficiency of lung cancer, including measurement of SHOX2 and RASSF1A gene methylation.\u003c/em\u003e BMC Cancer, 2024. \u003cstrong\u003e24\u003c/strong\u003e(1): p. 282.\u003c/li\u003e\n\u003cli\u003eZhang, F. and T. Evans, \u003cem\u003eStage-specific DNA methylation dynamics in mammalian heart development.\u003c/em\u003e Epigenomics, 2025. \u003cstrong\u003e17\u003c/strong\u003e(5): p. 359-371.\u003c/li\u003e\n\u003cli\u003eZhang, W.Z., C.Y. Wu, and H. Lai, \u003cem\u003eA Review on the Role of DNA Methylation in Aortic Disease Associated With Marfan Syndrome.\u003c/em\u003e Cardiol Res, 2025. \u003cstrong\u003e16\u003c/strong\u003e(3): p. 169-177.\u003c/li\u003e\n\u003cli\u003eGreger, V., et al., \u003cem\u003eEpigenetic changes may contribute to the formation and spontaneous regression of retinoblastoma.\u003c/em\u003e Hum Genet, 1989. \u003cstrong\u003e83\u003c/strong\u003e(2): p. 155-8.\u003c/li\u003e\n\u003cli\u003eHerman, J.G., et al., \u003cem\u003eInactivation of the CDKN2/p16/MTS1 gene is frequently associated with aberrant DNA methylation in all common human cancers.\u003c/em\u003e Cancer Res, 1995. \u003cstrong\u003e55\u003c/strong\u003e(20): p. 4525-30.\u003c/li\u003e\n\u003cli\u003eMerlo, A., et al., \u003cem\u003e5\u0026apos; CpG island methylation is associated with transcriptional silencing of the tumour suppressor p16/CDKN2/MTS1 in human cancers.\u003c/em\u003e Nat Med, 1995. \u003cstrong\u003e1\u003c/strong\u003e(7): p. 686-92.\u003c/li\u003e\n\u003cli\u003eBaur, A.S., et al., \u003cem\u003eFrequent methylation silencing of p15(INK4b) (MTS2) and p16(INK4a) (MTS1) in B-cell and T-cell lymphomas.\u003c/em\u003e Blood, 1999. \u003cstrong\u003e94\u003c/strong\u003e(5): p. 1773-81.\u003c/li\u003e\n\u003cli\u003eTan, Y., et al., \u003cem\u003eAlternative polyadenylation reprogramming of MORC2 induced by NUDT21 loss promotes KIRC carcinogenesis.\u003c/em\u003e JCI Insight, 2023. \u003cstrong\u003e8\u003c/strong\u003e(18).\u003c/li\u003e\n\u003cli\u003eHesham, D., et al., \u003cem\u003eEpigenetic silencing of ZIC4 unveils a potential tumor suppressor role in pediatric choroid plexus carcinoma.\u003c/em\u003e Sci Rep, 2024. \u003cstrong\u003e14\u003c/strong\u003e(1): p. 21293.\u003c/li\u003e\n\u003cli\u003eDiakiw, S.M., et al., \u003cem\u003eThe granulocyte-associated transcription factor Kruppel-like factor 5 is silenced by hypermethylation in acute myeloid leukemia.\u003c/em\u003e Leuk Res, 2012. \u003cstrong\u003e36\u003c/strong\u003e(1): p. 110-6.\u003c/li\u003e\n\u003cli\u003eFan, X.Y., et al., \u003cem\u003eAssociation between RUNX3 promoter methylation and gastric cancer: a meta-analysis.\u003c/em\u003e BMC Gastroenterol, 2011. \u003cstrong\u003e11\u003c/strong\u003e: p. 92.\u003c/li\u003e\n\u003cli\u003eHerman, J.G. and S.B. Baylin, \u003cem\u003eGene silencing in cancer in association with promoter hypermethylation.\u003c/em\u003e N Engl J Med, 2003. \u003cstrong\u003e349\u003c/strong\u003e(21): p. 2042-54.\u003c/li\u003e\n\u003cli\u003eStelloo, S., et al., \u003cem\u003eEndogenous androgen receptor proteomic profiling reveals genomic subcomplex involved in prostate tumorigenesis.\u003c/em\u003e Oncogene, 2018. \u003cstrong\u003e37\u003c/strong\u003e(3): p. 313-322.\u003c/li\u003e\n\u003cli\u003eVillicana, S. and J.T. Bell, \u003cem\u003eGenetic impacts on DNA methylation: research findings and future perspectives.\u003c/em\u003e Genome Biol, 2021. \u003cstrong\u003e22\u003c/strong\u003e(1): p. 127.\u003c/li\u003e\n\u003cli\u003eZeng, Y., et al., \u003cem\u003eDNA methylation modulated genetic variant effect on gene transcriptional regulation.\u003c/em\u003e Genome Biol, 2023. \u003cstrong\u003e24\u003c/strong\u003e(1): p. 285.\u003c/li\u003e\n\u003cli\u003eVillicana, S., et al., \u003cem\u003eGenetic impacts on DNA methylation help elucidate regulatory genomic processes.\u003c/em\u003e Genome Biol, 2023. \u003cstrong\u003e24\u003c/strong\u003e(1): p. 176.\u003c/li\u003e\n\u003cli\u003eDo, C., et al., \u003cem\u003eAllele-specific DNA methylation is increased in cancers and its dense mapping in normal plus neoplastic cells increases the yield of disease-associated regulatory SNPs.\u003c/em\u003e Genome Biol, 2020. \u003cstrong\u003e21\u003c/strong\u003e(1): p. 153.\u003c/li\u003e\n\u003cli\u003eGalasso, M., et al., \u003cem\u003eThe rs1001179 SNP and CpG methylation regulate catalase expression in chronic lymphocytic leukemia.\u003c/em\u003e Cell Mol Life Sci, 2022. \u003cstrong\u003e79\u003c/strong\u003e(10): p. 521.\u003c/li\u003e\n\u003cli\u003eTian, Y., et al., \u003cem\u003eNovel role of prostate cancer risk variant rs7247241 on PPP1R14A isoform transition through allelic TF binding and CpG methylation.\u003c/em\u003e Hum Mol Genet, 2022. \u003cstrong\u003e31\u003c/strong\u003e(10): p. 1610-1621.\u003c/li\u003e\n\u003cli\u003eAhmed, M., et al., \u003cem\u003eCRISPRi screens reveal a DNA methylation-mediated 3D genome dependent causal mechanism in prostate cancer.\u003c/em\u003e Nat Commun, 2021. \u003cstrong\u003e12\u003c/strong\u003e(1): p. 1781.\u003c/li\u003e\n\u003cli\u003eOnuchic, V., et al., \u003cem\u003eAllele-specific epigenome maps reveal sequence-dependent stochastic switching at regulatory loci.\u003c/em\u003e Science, 2018. \u003cstrong\u003e361\u003c/strong\u003e(6409).\u003c/li\u003e\n\u003cli\u003eDumont, E.L.P., B. Tycko, and C. Do, \u003cem\u003eCloudASM: an ultra-efficient cloud-based pipeline for mapping allele-specific DNA methylation.\u003c/em\u003e Bioinformatics, 2020. \u003cstrong\u003e36\u003c/strong\u003e(11): p. 3558-3560.\u003c/li\u003e\n\u003cli\u003eRosenski, J., et al., \u003cem\u003eAtlas of imprinted and allele-specific DNA methylation in the human body.\u003c/em\u003e Nat Commun, 2025. \u003cstrong\u003e16\u003c/strong\u003e(1): p. 2141.\u003c/li\u003e\n\u003cli\u003eZhang, W., et al., \u003cem\u003ePredicting genome-wide DNA methylation using methylation marks, genomic position, and DNA regulatory elements.\u003c/em\u003e Genome Biol, 2015. \u003cstrong\u003e16\u003c/strong\u003e(1): p. 14.\u003c/li\u003e\n\u003cli\u003eAbbas, Z., et al., \u003cem\u003eXGBoost framework with feature selection for the prediction of RNA N5-methylcytosine sites.\u003c/em\u003e Mol Ther, 2023. \u003cstrong\u003e31\u003c/strong\u003e(8): p. 2543-2551.\u003c/li\u003e\n\u003cli\u003eGao, Z., et al., \u003cem\u003eEpiGePT: a pretrained transformer-based language model for context-specific human epigenomics.\u003c/em\u003e Genome Biol, 2024. \u003cstrong\u003e25\u003c/strong\u003e(1): p. 310.\u003c/li\u003e\n\u003cli\u003eLi, W., W.H. Wong, and R. Jiang, \u003cem\u003eDeepTACT: predicting 3D chromatin contacts via bootstrapping deep learning.\u003c/em\u003e Nucleic Acids Res, 2019. \u003cstrong\u003e47\u003c/strong\u003e(10): p. e60.\u003c/li\u003e\n\u003cli\u003eJi, Y., et al., \u003cem\u003eDNABERT: pre-trained Bidirectional Encoder Representations from Transformers model for DNA-language in genome.\u003c/em\u003e Bioinformatics, 2021. \u003cstrong\u003e37\u003c/strong\u003e(15): p. 2112-2120.\u003c/li\u003e\n\u003cli\u003eJin, J., et al., \u003cem\u003eiDNA-ABF: multi-scale deep biological language learning model for the interpretable prediction of DNA methylations.\u003c/em\u003e Genome Biol, 2022. \u003cstrong\u003e23\u003c/strong\u003e(1): p. 219.\u003c/li\u003e\n\u003cli\u003eXie, H., et al., \u003cem\u003eMethyl-GP: accurate generic DNA methylation prediction based on a language model and representation learning.\u003c/em\u003e Nucleic Acids Res, 2025. \u003cstrong\u003e53\u003c/strong\u003e(6).\u003c/li\u003e\n\u003cli\u003eZhuo, L., et al., \u003cem\u003eStableDNAm: towards a stable and efficient model for predicting DNA methylation based on adaptive feature correction learning.\u003c/em\u003e BMC Genomics, 2023. \u003cstrong\u003e24\u003c/strong\u003e(1): p. 742.\u003c/li\u003e\n\u003cli\u003eYu, X., et al., \u003cem\u003eiDNA-ITLM: An interpretable and transferable learning model for identifying DNA methylation.\u003c/em\u003e PLoS One, 2024. \u003cstrong\u003e19\u003c/strong\u003e(10): p. e0301791.\u003c/li\u003e\n\u003cli\u003eYu, X., et al., \u003cem\u003eiDNA-OpenPrompt: OpenPrompt learning model for identifying DNA methylation.\u003c/em\u003e Front Genet, 2024. \u003cstrong\u003e15\u003c/strong\u003e: p. 1377285.\u003c/li\u003e\n\u003cli\u003eLiu, Z., et al., \u003cem\u003eKAN: Kolmogorov-Arnold Networks.\u003c/em\u003e ArXiv, 2024. \u003cstrong\u003eabs/2404.19756\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eZhao, J., et al., \u003cem\u003eCanASM: a comprehensive database for genome-wide allele-specific DNA methylation identification and annotation in cancer.\u003c/em\u003e BMC Genomics, 2025. \u003cstrong\u003e26\u003c/strong\u003e(1): p. 648.\u003c/li\u003e\n\u003cli\u003eKumar, R. and A. Indrayan, \u003cem\u003eReceiver operating characteristic (ROC) curve for medical researchers.\u003c/em\u003e Indian Pediatr, 2011. \u003cstrong\u003e48\u003c/strong\u003e(4): p. 277-87.\u003c/li\u003e\n\u003cli\u003eWang, K., M. Li, and H. Hakonarson, \u003cem\u003eANNOVAR: functional annotation of genetic variants from high-throughput sequencing data.\u003c/em\u003e Nucleic Acids Res, 2010. \u003cstrong\u003e38\u003c/strong\u003e(16): p. e164.\u003c/li\u003e\n\u003cli\u003eNassar, L.R., et al., \u003cem\u003eThe UCSC Genome Browser database: 2023 update.\u003c/em\u003e Nucleic Acids Res, 2023. \u003cstrong\u003e51\u003c/strong\u003e(D1): p. D1188-D1195.\u003c/li\u003e\n\u003cli\u003eMishra, N.K. and C. Guda, \u003cem\u003eGenome-wide DNA methylation analysis reveals molecular subtypes of pancreatic cancer.\u003c/em\u003e Oncotarget, 2017. \u003cstrong\u003e8\u003c/strong\u003e(17): p. 28990-29012.\u003c/li\u003e\n\u003cli\u003eChen, K., et al., \u003cem\u003eIndividualized dynamic methylation-based analysis of cell-free DNA in postoperative monitoring of lung cancer.\u003c/em\u003e BMC Med, 2023. \u003cstrong\u003e21\u003c/strong\u003e(1): p. 255.\u003c/li\u003e\n\u003cli\u003eRodriguez-Lloveras, H., et al., \u003cem\u003eDNA Methylation Dynamics and Prognostic Implications in Metastatic Differentiated Thyroid Cancer.\u003c/em\u003e Thyroid, 2025. \u003cstrong\u003e35\u003c/strong\u003e(5): p. 494-507.\u003c/li\u003e\n\u003cli\u003eSmith, Z.D. and A. Meissner, \u003cem\u003eDNA methylation: roles in mammalian development.\u003c/em\u003e Nat Rev Genet, 2013. \u003cstrong\u003e14\u003c/strong\u003e(3): p. 204-20.\u003c/li\u003e\n\u003cli\u003eZeng, W., A. Gautam, and D.H. Huson, \u003cem\u003eMuLan-Methyl-multiple transformer-based language models for accurate DNA methylation prediction.\u003c/em\u003e Gigascience, 2022. \u003cstrong\u003e12\u003c/strong\u003e.\u003c/li\u003e\n\u003cli\u003eLi, W. and A. Godzik, \u003cem\u003eCd-hit: a fast program for clustering and comparing large sets of protein or nucleotide sequences.\u003c/em\u003e Bioinformatics, 2006. \u003cstrong\u003e22\u003c/strong\u003e(13): p. 1658-9.\u003c/li\u003e\n\u003cli\u003eWu, X. and D.P. Bartel, \u003cem\u003ekpLogo: positional k-mer analysis reveals hidden specificity in biological sequences.\u003c/em\u003e Nucleic Acids Res, 2017. \u003cstrong\u003e45\u003c/strong\u003e(W1): p. W534-W538.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Deep learning, NLP, CpG methylation changes","lastPublishedDoi":"10.21203/rs.3.rs-7586988/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7586988/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eDifferential methylation is a key epigenetic process contributing to cancer development. Most DNA methylation prediction methods rely on DNA sequences from the background reference genome, neglecting individual genetic variation, which limits their ability to capture methylation differences. To address this, we propose CMC-WDTK, a deep learning framework that combines a weight-sharing dual-branch Transformer with a Kolmogorov‒Arnold network (KAN) to integrate sequences flanking CpG sites and adjacent single nucleotide variation (SNV) information to predict methylation changes between DNA sequences. CMC-WDTK captures global and local features of both reference and variant sequences and models high-dimensional relationships, offering accurate predictions of methylation changes. CMC-WDTK accurately predicted DNA methylation changes in eight real datasets (AUC greater than 0.8 for all datasets), with strong generalizability across datasets. Additionally, it identified a repeated cytosine and guanine sequence motif that promotes increased methylation. CMC-WDTK is the first computational tool used to predict methylation changes between sequences, offering significant advancements in understanding and comparing DNA methylation across diverse datasets and biological conditions.\u003c/p\u003e","manuscriptTitle":"CMC-WDTK: CpG methylation change prediction by a weight-sharing dual-branch Transformer-Kolmogorov–Arnold network model","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-10-17 10:25:37","doi":"10.21203/rs.3.rs-7586988/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-11-05T09:18:43+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-30T04:04:31+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-16T01:39:35+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-10-14T13:53:50+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"300973840835648595773995261156302981170","date":"2025-10-14T13:24:32+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"261777113537077878863156591632313353143","date":"2025-10-13T07:42:57+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"265020616252024822302390785737194001287","date":"2025-10-08T15:03:23+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"241810004475470644961389656994068004259","date":"2025-10-07T07:43:10+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-10-06T14:56:12+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-09-15T14:34:55+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-09-12T09:11:14+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-09-12T01:51:44+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Genomics","date":"2025-09-12T01:48:44+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-genomics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"gics","sideBox":"Learn more about [BMC Genomics](http://bmcgenomics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/gics","title":"BMC Genomics","twitterHandle":"#BMCGenomics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"b8f4ca50-b50f-41db-a27e-c2ba1a153dfc","owner":[],"postedDate":"October 17th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2026-02-09T16:07:51+00:00","versionOfRecord":{"articleIdentity":"rs-7586988","link":"https://doi.org/10.1186/s12864-026-12627-9","journal":{"identity":"bmc-genomics","isVorOnly":false,"title":"BMC Genomics"},"publishedOn":"2026-02-07 15:59:08","publishedOnDateReadable":"February 7th, 2026"},"versionCreatedAt":"2025-10-17 10:25:37","video":"","vorDoi":"10.1186/s12864-026-12627-9","vorDoiUrl":"https://doi.org/10.1186/s12864-026-12627-9","workflowStages":[]},"version":"v1","identity":"rs-7586988","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7586988","identity":"rs-7586988","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00
unpaywall
last seen: 2026-05-24T02:00:01.246996+00:00
License: CC-BY-4.0