Identification of monitoring panels and hub genes for heterogeneous relapse outcomes in acute lymphoblastic leukemia

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Background Relapse is the leading cause of mortality in acute lymphoblastic leukemia (ALL) patients. Minimal residual disease (MRD) detection is effective for bone marrow relapse (BMR) but less so for central nervous system relapse (CR) and combined bone marrow-central nervous system relapse (BMCR). Furthermore, there remains an unmet need for more effective therapeutic targets tailored to distinct relapse outcomes in ALL. Consequently, the identification of minimally invasive, serological biomarkers capable of monitoring diverse relapse outcomes, coupled with the exploration of hub genes underlying these outcomes, holds significant clinical relevance. Methods Clinical (n = 927), whole-exome mutation (n = 618), copy number variation (CNV, n = 250), and transcriptomic (n = 343) data were collected from the TARGET database for ALL patients with four outcomes: no relapse (NR), BMR, CR, and BMCR. The characteristics of each dataset were analysed. Serological panels and monitoring models for each relapse outcome were screened and constructed using optimized XGBoost and SHAP algorithms, with 50 random training iterations and top 5 indicator combinations. Hub genes were identified via the STRING database and Boruta algorithm (50 iterations). Functional impacts of representative hub genes on ALL progression were validated in cellular models. Results Multi-omics analysis showed higher mortality in relapse groups than in NR, with BMCR posing the highest risk. Elevated mutation burden in BMCR was linked to poor survival. Distinct mutational profiles, CNV signatures, and transcriptomic dysregulation patterns were observed across groups, contributing to heterogeneous relapse outcomes. Optimal serological panels and monitoring models were constructed for each group: NR (LRBA/SH2B1A/IL36G/ADH1A/CEP85, AUC = 0.985), BMR (CD22/CRMP1/EPHA2/KLK14, AUC = 0.997), CR (ACVRL1/FAT3/SKAP1, AUC = 0.990), and BMCR (AGRN/CP/EFEMP1/LGALS7/ST6GAL2, AUC = 0.999), all superior to existing models (AUC < 0.903). Hub gene screening yielded 18 candidates for BMR (e.g., VAMP2, EFNB2), 13 for CR (e.g., MT-ATP8, CXCR3), and 3 for BMCR (e.g., TGFB1, PLCG1). Experimental validation demonstrated that knockdown of CD74 (BMR) and TGFB1 (BMCR)-highly expressed hub genes-significantly inhibited ALL cell proliferation and induced apoptosis, whereas knockdown of MT-CO1 (CR)-a low-expressed hub gene-produced the opposite effect. Conclusion This study established serological panels for diverse ALL relapse outcomes and identified hub genes with therapeutic potential, providing a foundation for early relapse surveillance and targeted interventions in ALL.
Full text 123,872 characters · extracted from preprint-html · click to expand
Identification of monitoring panels and hub genes for heterogeneous relapse outcomes in acute lymphoblastic leukemia | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Identification of monitoring panels and hub genes for heterogeneous relapse outcomes in acute lymphoblastic leukemia Xin He, Jianhong Zhang, Guowei He, Jing Shi, Bo Tian, Jing Huang, and 6 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8999580/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 3 You are reading this latest preprint version Abstract Background Relapse is the leading cause of mortality in acute lymphoblastic leukemia (ALL) patients. Minimal residual disease (MRD) detection is effective for bone marrow relapse (BMR) but less so for central nervous system relapse (CR) and combined bone marrow-central nervous system relapse (BMCR). Furthermore, there remains an unmet need for more effective therapeutic targets tailored to distinct relapse outcomes in ALL. Consequently, the identification of minimally invasive, serological biomarkers capable of monitoring diverse relapse outcomes, coupled with the exploration of hub genes underlying these outcomes, holds significant clinical relevance. Methods Clinical (n = 927), whole-exome mutation (n = 618), copy number variation (CNV, n = 250), and transcriptomic (n = 343) data were collected from the TARGET database for ALL patients with four outcomes: no relapse (NR), BMR, CR, and BMCR. The characteristics of each dataset were analysed. Serological panels and monitoring models for each relapse outcome were screened and constructed using optimized XGBoost and SHAP algorithms, with 50 random training iterations and top 5 indicator combinations. Hub genes were identified via the STRING database and Boruta algorithm (50 iterations). Functional impacts of representative hub genes on ALL progression were validated in cellular models. Results Multi-omics analysis showed higher mortality in relapse groups than in NR, with BMCR posing the highest risk. Elevated mutation burden in BMCR was linked to poor survival. Distinct mutational profiles, CNV signatures, and transcriptomic dysregulation patterns were observed across groups, contributing to heterogeneous relapse outcomes. Optimal serological panels and monitoring models were constructed for each group: NR (LRBA/SH2B1A/IL36G/ADH1A/CEP85, AUC = 0.985), BMR (CD22/CRMP1/EPHA2/KLK14, AUC = 0.997), CR (ACVRL1/FAT3/SKAP1, AUC = 0.990), and BMCR (AGRN/CP/EFEMP1/LGALS7/ST6GAL2, AUC = 0.999), all superior to existing models (AUC < 0.903). Hub gene screening yielded 18 candidates for BMR (e.g., VAMP2, EFNB2), 13 for CR (e.g., MT-ATP8, CXCR3), and 3 for BMCR (e.g., TGFB1, PLCG1). Experimental validation demonstrated that knockdown of CD74 (BMR) and TGFB1 (BMCR)-highly expressed hub genes-significantly inhibited ALL cell proliferation and induced apoptosis, whereas knockdown of MT-CO1 (CR)-a low-expressed hub gene-produced the opposite effect. Conclusion This study established serological panels for diverse ALL relapse outcomes and identified hub genes with therapeutic potential, providing a foundation for early relapse surveillance and targeted interventions in ALL. Acute Lymphoblastic Leukemia Relapse Outcome Monitoring Model Serological Panel Hub Gene Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Introduction Relapse remains a pivotal challenge impacting long-term survival and quality of life in acute lymphoblastic leukemia (ALL), a malignancy originating from lymphoid progenitor cells with complex pathogenesis involving dysregulated gene expression networks. Despite therapeutic advances, 15–30% of ALL patients experience relapse, with post-relapse 5-year survival rates plummeting to 36% ( 1 – 2 ). Relapse manifests in diverse anatomical patterns, including isolated bone marrow relapse (BMR, 70–80% of cases), isolated central nervous system relapse (CR, ~ 10–15% of cases), and combined bone marrow-central nervous system relapse (BMCR, ~ 5% of cases) ( 3 ). Notably, adult ALL patients with central nervous system (CNS) relapse involvement exhibit dismal prognoses, with a median survival duration of less than one year ( 4 – 8 ). The heterogeneity of relapse patterns and persistent high relapse rates pose a significant challenge to improving ALL patients survival outcomes. Consequently, monitoring different recurrence outcomes and identifying hub genes from networks holds critical implications for early therapeutic intervention and survival improvement in high-risk ALL populations. Minimal residual disease (MRD) detection, currently the cornerstone of ALL relapse surveillance, demonstrates high sensitivity in monitoring BMR. However, its performance declines in CR and BMCR, with reduced sensitivity and specificity for distinguishing distinct relapse outcomes ( 9 – 10 ). Furthermore, MRD assessment requires invasive bone marrow aspiration, posing procedural risks and patient discomfort while risking false-negative results due to sampling variability in timing and anatomical location ( 11 – 12 ). Serological biomarkers, as a non-invasive alternative, could circumvent these limitations ( 13 – 17 ), yet no validated serological monitoring tool exists for ALL relapse. Consequently, the development of minimally invasive, outcome-specific serological models with high predictive accuracy remains an unmet clinical need. Beyond diagnostic limitations, the pathophysiological mechanisms underlying ALL relapse outcomes remain poorly characterized. Although mutations in TP53, NOTCH1, and epigenetic alterations have been implicated in relapse pathogenesis ( 18 – 23 ), these studies predominantly adopt single-omics approaches or focus on isolated genes, lacking multidimensional integration and systemic network analysis ( 24 ). This fragmented perspective hinders comprehensive understanding of relapse heterogeneity. To address this limitation, multi-omics integration and holistic network modeling are essential to delineate relapse-specific molecular signatures and identify hub genes regulating these distinct outcomes. Such insights would enable precision intervention strategies targeting relapse-driving genes, thereby improving clinical outcomes in high-risk ALL. The Therapeutically Applicable Research to Generate Effective Treatments (TARGET) database integrates multi-dimensional datasets across seven malignancies, including ALL, providing a critical resource for identifying predictive, therapeutic, and prognostic targets in cancer progression ( 25 – 26 ). This study utilized clinical, whole-exome mutation, copy number variation (CNV), and transcriptomic data from ALL patients to systematically investigate the biological foundations of four relapse outcomes: no relapse (NR), BMR, CR, and BMCR. Specifically, we firstly elucidated associations between clinical variables and relapse outcomes, alongside survival prognosis evaluation for each relapse group. Simultaneously, our research characterized molecular alterations across genomic, epigenomic, and transcriptomic layers to delineate relapse-specific patterns. Subsequently, following the selection of the optimal parameters for the XGBoost training model using 5-fold cross-validation and grid search, the prioritisation of features through SHAP interpretability analysis, the evaluation of the model's discrimination and calibration using AUC and HL p-values, the 50 random trainings based on optimised algorithm ( 27 – 31 ), and the further training of the top five indicators in various combinations, the optimal serum panels for monitoring different recurrence outcomes were screened. Ultimately, we constructed protein-protein interaction networks and applied the Boruta algorithm to identify hub genes significantly associated with each relapse group, and the functional validation of candidate hub genes (e.g., CD74, TGFB1, MT-CO1) was performed using in vitro cellular assays to confirm their biological relevance in ALL pathogenesis. Our research employs a multi-omics perspective to analyze clinical issues, constructing serological models for minimally invasive relapse monitoring and seeking mechanistic insights into relapse outcomes. These findings hold significant potential to transform clinical management strategies by enabling early relapse detection and precision targeting of relapse-driving pathways in ALL. Materials and methods 1. Clinical Data Acquisition Data for this study were sourced from the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) database ( https://portal.gdc.cancer.gov/analysis_page?app=CohortBuilder&tab=general ). We retrieved open-access samples with a primary diagnosis of acute lymphoblastic leukaemia (ALL) from the TARGET-ALL-P1, TARGET-ALL-P2, and TARGET-ALL-P3 datasets in this database. The samples encompassed clinical, whole-exome mutation, CNV, and transcriptomics data. Following the processes of data merging and collation, samples that satisfied the subsequent criteria were selected for further analysis. Firstly, the cancer type was primary blood-derived cancer. Secondly, the primary tumor site was bone marrow. Thirdly, the follow-up recurrence outcomes were clearly classified as one of the following four types: NR, BMR, CR, and BMCR. Ultimately, a total of 927 clinical samples met these inclusion criteria, including 618 with whole-exome mutation data, 250 with CNV profiles, and 343 with transcriptomic data. With the exception of transcriptomics data, the sample sizes for the remaining data were consistent with the number of clinical cases. The detailed clinical characteristics are summarized in supplementary materials. 2. Clinical Data Analysis The clinical dataset comprised seven categorical variables, including Gender and Race . We tabulated the frequency of each categorical variable across the four relapse outcomes and quantified significance and strength of association using the `assocstats` function from the R package `vcd`. Pearson's chi-squared test was employed for contingency tables with sample sizes ≥ 40 and all cell counts ≥ 5 to analyse the association significance. While the likelihood ratio chi-squared test was used for tables containing cell counts < 5 to analyse the association significance. The chi-square test p < 0.05 denoted statistical significance. Association strength was measured using the Contingency Coefficient and Cramer's V, with values approaching 1 indicating stronger associations and values near 0 suggesting weaker associations. 3. Genomic Mutation Analysis To investigate mutation profiles across ALL relapse outcomes, whole-exome mutation data were analyzed using the R package `maftools`. After importing data via the `read.maf` function, mutation burden was compared between ALL relapse groups and TCGA pan-cancer cohorts using `tcgaCompare`. Differential mutation burden across relapse groups was assessed with the `tmb` function, and mutation summary plots were generated using `plotmafSummary`. Mutational signature analysis was performed using cosine similarity scoring against 60 COSMIC mutational signatures (v3.2), incorporating the `BSgenome.Hsapiens.UCSC.hg38` reference genome and `NMF` package. Pathway enrichment analysis evaluated the contribution of mutated genes to 24 established oncogenic signaling pathways using the `pathways` function. The mutant data are available in the supplementary materials. 4. CNV Analysis To analyze CNV differences across ALL relapse outcomes, genes with a copy number of 2 were labeled as "neutral", while those with copy numbers other than 2 were classified as "change". Four pairwise comparisons were performed: NR vs. Other (non-NR samples), BMR vs. Other (non-BMR samples), CR vs. Other (non-CR samples), and BMCR vs. Other (non-BMCR samples). For each comparison, 2×2 contingency tables (rows: CNV status; columns: relapse groups) were constructed, and chi-squared tests were applied to identify differentially altered CNV genes. Additionally, genes were categorized as "loss" (copy number 2). The frequency of gain/loss events was calculated as the proportion of affected samples within each relapse group to quantify genomic instability patterns. The CNV data are available in the supplementary materials. 5. Transcriptomic Differential Expression Analysis RNA-seq count data were normalized and analyzed for differential expression using the R package DESeq2. Pairwise comparisons (NR vs. Other, BMR vs. Other, CR vs. Other, BMCR vs. Other) identified differentially expressed genes (DEGs) with thresholds of p 1. The RNA-seq data are available in the supplementary materials. 6. Principal Component Analysis (PCA) To visualize transcriptomic heterogeneity across relapse groups, DESeq2-normalized counts underwent log₂(x + 1) transformation. Dimensionality reduction and visualization were performed using PCA from the R package factoextra, with the `fviz_pca_ind` function generating PCA plot to illustrate sample clustering patterns. 7. Gene Biotype Annotation Gene biotypes for DEGs were obtained using the `annoGene` function from the R package AnnoProbe, enabling functional classification of differentially expressed transcripts (e.g., protein-coding, lncRNA, pseudogene). 8. Enrichment Analysis For differentially altered CNV genes, chromosome enrichment was assessed using the `MSigDB c1.all.v2023.2.Hs.symbols.gmt` file and the `enricher` function from the R package clusterProfiler. Gene Ontology(GO) enrichment was performed via the DAVID database ( https://david.ncifcrf.gov/ ) with parameters counts = 2 and ease = 0.1. Both adopt over-representation analysis (ORA) for enrichment. For DEGs, GO and Reactome pathway enrichments were conducted using MSigDB's `c5.go.bp.v2024.1.Hs.symbols.gmt` and `c2.cp.reactome.v2024.1.Hs.symbols.gmt` files, respectively, with Gene Set Enrichment Analysis (GSEA). This approach enabled assessment of pathway activation or inhibition status and quantification of effect magnitude via normalized enrichment scores (NES): NES > 0 indicates pathway activation, NES < 0 indicates suppression, and absolute NES values positively correlate with effect strength. To characterize genes in optimal monitoring panels, we retrieved functional annotations from the GeneCards database ( https://www.genecards.org/ ) and highlighted representative GO pathways from the 'Pathways' section. To explore regulatory roles of hub genes, GO ORA were conducted through DAVID database with parameters counts = 1 and ease = 1. 9. Plasma Protein Curation and Data Integration Plasma protein-coding genes were compiled from the Human Protein Atlas (HPA, https://www.proteinatlas.org/ ) and the Human Plasma Proteome Project (HPPP, https://www.hupo.org/plasma-proteome-project/ ). The union of these datasets was used to filter plasma protein-coding candidates from DEGs. The plasma protein-coding genes are available in the supplementary materials. 10. Disease Database Mining The MalaCards human disease annotation database ( https://www.malacards.org/ ) was queried to assess associations between hematopoietic malignancies and genes in optimal panels/hub gene sets, utilizing gene-disease association scores. 11. Optimal Panel Selection and Monitoring Model Development Four relapse outcomes (NR vs. Other, BMR vs. Other, CR vs. Other, BMCR vs. Other) were modeled as binary endpoints, with upregulated plasma protein-coding DEGs serving as predictive features. Pre-trained XGBoost models were optimized via 5-fold cross-validation and grid search for hyperparameter tuning. Feature importance was quantified using SHAP values from the R package SHAPforxgboost. Model discrimination (AUC) and calibration (Hosmer-Lemeshow p-value) were evaluated with the tidymodels and PredictABEL packages. To enhance robustness, pre-training was repeated across 50 random seeds and results were retained with criteria of AUC < 0.95 and HL p < 0.05. Genes were ranked in descending order by mean SHAP values calculated under retention criteria, and the top 5 genes from each group were combined to generate 31 candidate panels. Each panel underwent retraining under reserved pre-training seeds with optimal parameters, The optimal panel for each relapse outcome was selected based on maximum mean AUC with mean HL p > 0.05. The training data are available in the supplementary materials. 12. Hub Gene Identification Interaction networks for DEGs in each relapse group were constructed using the STRINGdb package (STRING database version 12.0, Homo sapiens, minimum interaction score threshold of 700). Degree values (number of direct interactions) for each gene node were calculated using the igraph package. For BMR vs. Other, CR vs. Other, and BMCR vs. Other comparisons, DEGs with degree > 10 were selected as input features for the Boruta algorithm. To enhance stability, Boruta training was repeated across 50 random seeds, with features classified as "Confirmed" (definitively important), "Tentative" (potentially important), or "Rejected" (not important) based on variable importance. Genes that were classified as "Confirmed" in all 50 iterations and were unique to one group were retained as group-specific hub genes.The networks and Boruta results are available in the supplementary materials. 13. Quantitative real-time PCR (qPCR) Total RNA was isolated by TRIzol® Reagent (Tiangen, China) and converted to cDNA using a reverse transcription kit (Tiangen, China). Then, the expression of target gene was detected by real-time PCR (SYBR green) with specific primers. The PCR reaction conditions were as follows: 95°C for 3 min, followed by 40 cycles of 95°C for 15s, 60°C for 15s, and 72°C for 20 s, concluding with an extension step of 72°C for 5min. The data were analyzed using 2 −∆∆Ct method. Primers were listed in the supplementary materials. 14. Cell lines and culture conditions The human Ph + ALL cell line SUP-B15 and human embryonic kidney cell line 293FT were purchased from the American Type Culture Collection. SUP-B15 cells were cultured in Iscove's modified Dulbecco's medium (IMDM) (Cytiva, Cat# SH30228.FS) supplemented with 1% penicillin-streptomycin, 20% FBS (Gibco; Thermo Fisher Scientific, Inc.), and 0.05 mM β-mercaptoethanol (Gibco, Cat# 21985-023). 293FT cells were cultured in DMEM (HyClone, Cat# SH30022.01) with 10% FBS. All cells were maintained in a humidified incubator at 37°C in an atmosphere of 5% CO₂. 15. Lentivirus Production and Cell Transduction ShRNA sequences targeting CD74, MT-CO1, and TGFB1, along with a negative control sequence, were respectively cloned into the pLKO.1 lentiviral vector (addgene, #8453). Lentiviral particles were produced by co-transfecting 293FT cells with the lentiviral vector, pMD2.G, and psPAX2 viral packaging plasmids using Lipofectamine 3000 (Invitrogen, USA) according to the manufacturer's protocol. Viral supernatants were harvested at 48 h and 72 h post-transfection, and viral titers were determined before storage at -80°C. For cell infection, SUP-B15 cells were incubated with viral culture medium supplemented with 10 µg/mL polybrene. Stably transduced cells were selected using 2 µg/mL puromycin for 14 days. 16. Cell Proliferation Assay Stable SUP-B15 cell lines were seeded at 2,000 cells/well in 96-well plates (30 replicates per sample). Cell viability was measured daily for 5 days using the Cell Counting Kit-8 (MCE, HY-K0301) according to the manufacturer's instructions, with five wells measured per day. 17. Flow Cytometry Analysis For S-phase analysis, stable cells were pulsed with 100 µM 5-bromo-2'-deoxyuridine (BrdU, Thermo Fisher Scientific) for 1 h at 37°C before harvesting. Then cells were fixed with 70% ethanol at 4°C overnight. Next, permeabilized with 0.3% Triton X-100 for 30 min, and stained with FITC-anti-BrdU (Thermo Fisher Scientific) for 1 h at room temperature. Finally, cells were labeled with 50 ng/mL propidium iodide (PI, Thermo Fisher Scientific) and 10 µg/mL RNase A (Thermo Fisher Scientific) for 30 min at 37°C, and then analysed by flow cytometry. For apoptosis analysis, cells were stained with PI and APC-conjugated annexin V (BestBio, C1062) according to the manufacturer's protocol prior to flow cytometry. Data acquisition was performed on a CytExpert flow cytometer (Beckman Coulter, Inc), with a minimum of 20,000 events analyzed per sample. Results 1. Clinical Characteristics of ALL Relapse Outcomes To investigate clinical features associated with distinct relapse outcomes in ALL, we analyzed clinical data from 927 TARGET-ALL patients categorized into four groups: 708 cases in the NR group, 167 with BMR, 30 with CR, and 22 with BMCR (Fig. 1A). Assessment of the significance and strength of associations between categorical clinical variables and various recurrence outcomes revealed no significant intergroup differences in Race , MLL status , and Down syndrome . However, Gender , CNS status at diagnosis , ETV6/RUNX1 fusion status , TRISOMY 4/10 status , and TGF3 PBX1 status differed significantly across groups (Fig. 1B), suggesting these differential variables may contribute to relapse heterogeneity. Additionally, Vital status showed significant difference among groups and the strongest association with relapse outcome, with BMCR patients exhibiting the highest mortality rate (77.3%, 17/22), followed by BMR (61.1%, 102/167) and CR (43.3%, 13/30) (Fig. 1B). Survival analyses also confirmed that BMCR patients had the poorest overall survival (OS) and event free survival (EFS) outcomes (Fig. 1C). The findings emphasize the significance of recurrence patterns, particularly BMCR, in determining prognosis, further highlighting the urgent need for early detection and intervention in ALL recurrence. 2. Exome-Wide Mutation Analysis in ALL Relapse Outcomes 618 ALL samples with mutation data and varying relapse outcomes from the TARGET database were analyzed to characterize mutation profiles (Fig. 1A). Pan-cancer mutation burden comparison revealed that ALL relapse groups exhibited lower overall mutation rates compared to most solid tumors but higher than acute myeloid leukemia (LAML), suggesting relatively elevated genomic instability in ALL among hematologic malignancies (Fig. 2A). Notably, the BMCR group displayed the highest differential mutation burden (Fig. 2B), aligning with its poorest survival outcomes. These suggest a potential link between elevated genomic alterations and ALL malignant progression. Mutation signature analysis presented high similarity among groups in terms of variant classification, variant type, and SNV class but divergent top 10 mutated genes (Fig. 2C, Fig. S1). Mutational etiology assessment indicated similar exposure profiles among NR, BMR, and BMCR groups, with the exception of the CR patients, implying that mutation triggers do not explain group mutation heterogeneity (Fig. 2D, Fig. S2). Pathway enrichment analysis revealed group-specific oncogenic signaling alterations: MAPK signaling (25.22% mutation frequency) in NR, genome integrity pathways (13.51%) in BMR, NOTCH signaling (35.29%) in CR, and chromatin regulation/NOTCH/ubiquitin-proteasome systems (40% combined frequency) in BMCR (Fig. 2E, Fig. S3), suggesting that distinct oncogenic pathway dysregulation may drive relapse-specific outcomes in ALL. 3. CNV Analysis in ALL Relapse Outcomes 250 ALL samples with CNV data from the TARGET database were analyzed to characterize CNV profiles across relapse outcomes (Fig. 1A). Differential CNV analysis identified 2,414, 924, 3,212, and 43 significantly altered CNV genes in the NR, BMR, CR, and BMCR groups, respectively (Fig. 3A, Fig. S4). The limited number of differential CNVs in BMCR may be influenced by its smaller sample size. Chromosomal enrichment, gene biotype annotation, and GO enrichment analyses were performed to interpret group-specific CNV patterns. Notably, NR and CR groups exhibited similar CNV profiles in chromosomal distribution, gene biotype composition, and functional enrichment, whereas BMR and BMCR groups showed distinct patterns (Fig. 3B-D). Comprehensive analysis of the top 30 differential CNVs revealed predominant CNV loss in NR and CR groups, contrasting with CNV gain in BMR and BMCR groups. However, comparisons to control cohorts indicated increased loss frequency in NR, reduced loss frequency in CR, decreased gain frequency in BMR, and elevated gain frequency in BMCR (Fig. 3E). These findings elucidate CNV characteristics associated with distinct relapse outcomes, providing novel insights into the mechanisms driving ALL relapse heterogeneity. 4. Transcriptomic Analysis of ALL Relapse Outcomes A comprehensive transcriptomic analysis was performed on 343 ALL samples with RNA-sequencing data and relapse outcomes from the TARGET database (Fig. 1A). PCA plots revealed group-specific expression patterns. However, the NR and CR groups exhibited high similarity, as did the BMR and BMCR groups (Fig. 4A). This transcriptional similarity aligns with previous CNV-based group characterizations, suggesting potential regulatory links between genomic alterations and gene expression profiles. Differential expression analysis identified group-specific DEGs (Fig. 4B, Fig. S5), which were subjected to biotype annotation and functional enrichment analysis using GO and Reactome databases. The biotype classification exhibited diversity, yet protein-coding genes were confirmed as the predominant DEG category across groups (Fig. 4C). GO enrichment revealed distinct biological processes. NR group showed suppressed immune response, cell differentiation, and apoptosis pathways. BMR exhibited enhanced cell activation, adhesion and oxidative phosphorylation. CR displayed activated type II interferon production and mitochondrial functions. BMCR demonstrated augmented cell activation and adhesion, leukocyte chemotaxis, and B-cell activation (Fig. 4D). Reactome pathway analysis corroborated these findings, showing predominant pathway suppression in NR versus activation in other groups (Fig. 4E). These transcriptomic disparities provide novel mechanistic insights into the molecular drivers of relapse heterogeneity in ALL. 5. Optimized XGBoost Combined with SHAP Algorithm Identifies Optimal Serum Monitoring Panels for ALL Relapse Outcomes Current ALL relapse monitoring primarily relies on MRD detection, which is mainly applied to BMR surveillance but shows limited efficacy for CR/BMCR monitoring and requires invasive bone marrow sampling. Development of minimally invasive serological panels capable of accurately monitoring diverse relapse outcomes holds significant clinical value. To achieve this, we intersected plasma protein-coding genes from the HPA and HPPP databases with group-specific upregulated genes to identify candidate biomarkers for panel screening (Fig. 5A-B, Fig. S6A). Heatmaps revealed marked expression differences in the top 10 candidate genes across groups, suggesting their potential as relapse-specific markers (Fig. S6B). We optimized XGBoost models using 5-fold cross-validation and grid search, followed by SHAP analysis to rank feature importance. Fifty iterations of model pre-training were performed, with AUC values and Hosmer-Lemeshow (HL) test p-values used to filter suboptimal results and ensure robustness (Fig. S7). Mean SHAP values from retained iterations were calculated for final feature ranking (Fig. 5C). Combining top 5 features from each group generated 31 candidate panels, which were evaluated through iterative pre-training. Optimal panels satisfied with HL p<0.05 included: NR panel (LRBA, SH2B1A, IL36G, ADH1A, CEP85, AUC=0.985); BMR panel (CD22, CRMP1, EPHA2, KLK14, AUC=0.997); CR panel (ACVRL1, FAT3, SKAP1, AUC=0.990); and BMCR panel (AGRN, CP, EFEMP1, LGALS7, ST6GAL2, AUC=0.999) (Fig. 5D). Genomic analysis of panel genes revealed significant expression differences across groups but rare mutations (1% missense in FAT) and limited CNV alterations (statistically significant only in SH2D1A), suggesting expression regulation of panel genes caused by non-genomic factors (Fig. S8A-C). Functional characterization using the HPA protein class identified cancer/disease-associated genes (e.g., LRBA, SH2D1A) and FDA-approved drug targets (e.g., ADH1A, CD22) within optimal panels, supporting therapeutic relevance. Protein half-life predictions via ExPASy ProtParam indicated serum suitability of panel genes (>30 hours). Malacards analysis linked most panel genes to hematologic disorders, while GeneCards GO enrichment highlighted cell adhesion/migration pathways in BMR/CR/BMCR panels and CNS development processes in CR/BMCR panels, aligning with clinical relapse phenotypes (Fig. S8D). Survival analysis revealed poor prognosis associations with elevated levels of panel genes (CD22, LGALS7, CRMP1, AGRN, and KLK14), suggesting dual monitoring and prognostic potential (Fig. S9). 6. Identification and Validation of Hub Genes in ALL Relapse Outcomes To identify hub genes driving ALL relapse progression and potential therapeutic targets, we constructed protein-protein interaction networks using the STRING database and group-specific DEGs from BMR, CR, and BMCR groups. Subnetworks were extracted with nodes exhibiting degree>10, representing 14.44%, 5.62%, and 10.56% of total nodes in BMR/CR/BMCR groups, respectively. These subnetworks governed 39.23%, 15.42%, and 50.41% of total interactions, highlighting their regulatory importance (Fig. S10). Boruta algorithm was applied across 50 iterations to identify group-specific hub genes with confirmed importance scores of 50, yielding 18, 13, and 3 hub genes for BMR, CR, and BMCR groups, respectively (Fig. 6A). Multi-omics analysis revealed downregulated hub genes in CR and upregulated hub genes in BMCR, while BMR showed both up- and downregulated patterns. Notably, only ITPR2, ADCY5, and HLA-DRA exhibited missense mutations in NR samples, and CXCR3 was the sole hub gene with CNV differences across groups, suggesting expression regulation of hub genes caused by non-genomic factors (Fig. S11A, S11C-D). Functional characterization via HPA protein classes classified 28 hub genes (including ADCY5) as cancer/disease-related and 17 (including CD19) as FDA-approved/potential drug targets. Malacards analysis confirmed hematologic disease associations for all hub genes, underscoring their therapeutic relevance (Fig. S11B). DAVID GO enrichment of subnetwork-regulated hub genes revealed pathway convergence on signal transduction, biological function, immune responses, DNA replication, and transcription in BMR/BMCR groups, with additional CNS-related pathway enrichment in CR group (Fig. S12). Representative hub genes CD74 (BMR), MT-CO1 (CR), and TGFB1 (BMCR) were selected for functional validation. CD74 and TGFB1 demonstrate elevated levels of expression in the BMR and BMCR groups, respectively, while MT-CO1 exhibits reduced expression in the CR group (Fig. 6B). Consistent with expression profiles, shRNA-mediated knockdown in SUP-B15 cells (Fig. 6C) demonstrated that CD74/TGFB1 depletion enhanced proliferation and inhibited apoptosis, while MT-CO1 knockdown showed opposite effects (Fig. 6D-F). These results confirm the regulatory roles of identified hub genes in ALL pathogenesis, validating the screening approach. Discussion Relapse remains a critical cause of mortality in ALL, with distinct outcomes including BMR, CR, and BMCR ( 32 ). Current clinical surveillance predominantly relies on MRD monitoring, which is primarily limited to BMR detection and requires invasive bone marrow aspiration ( 11 – 12 ). Consequently, there is an urgent clinical need to develop minimally invasive serological biomarkers capable of monitoring diverse relapse outcomes. Furthermore, mechanistic insights into ALL relapse heterogeneity remain insufficient, necessitating the identification of hub genes driving relapse processes to enable early targeted interventions ( 33 ). Leveraging multi-omics data from the TARGET database, we systematically characterized clinical, genomic, CNV, and transcriptomic features across ALL relapse outcomes. This analysis yielded optimal serological monitoring panels and predictive models for each relapse group, alongside the identification of group-specific hub genes, thereby establishing novel pathways for early relapse detection and intervention in ALL. Utilizing machine learning algorithms and transcriptomic data, we identified serological monitoring panels tailored to ALL relapse outcomes. Predictive models constructed using these panels achieved mean AUC values of 0.985, 0.997, 0.990, and 0.999 for distinguishing NR, BMR, CR, and BMCR groups from their respective controls. These panels outperform existing models such as the Seven-lncRNA-mRNA Signature (AUC = 0.901) ( 34 ) and clinical variable-based models incorporating age, WBC count, and hemoglobin levels (AUC = 0.902) ( 35 ), demonstrating superior discriminatory capacity. The high AUC values underscore the panels' efficacy in identifying distinct relapse outcomes, offering clinicians a more precise monitoring tool. Compared to conventional MRD detection methods, serum-based panels avoid invasive bone marrow sampling, reduce procedural costs, and facilitate broader clinical implementation, highlighting their significant translational potential. Notably, several panel genes (e.g., CD22, CRMP1, KLK14, AGRN, LGALS7) exhibit prognostic relevance in ALL patients. Targeted therapies against CD22 have demonstrated efficacy in refractory/relapsed ALL ( 36 – 37 ), validating the therapeutic relevance of our findings and prompting similar potential for other panel components. Additionally, through network analysis and Boruta algorithm training, we identified hub genes significantly associated with ALL relapse outcomes, including VAMP2 and EFNB2 in BMR, MT-ATP8 and CXCR3 in CR, and HLA-DRA and HLA-F in BMCR. CR-associated hub genes showed enrichment in mitochondrial functional pathways, aligning with reported alterations in mitochondrial DNA levels in cerebrospinal fluid from patients with CNS relapse ( 38 ), implicating mitochondrial dysfunction in ALL CNS relapse pathogenesis. Functional validation experiments confirmed that knockdown of representative hub genes (CD74, MT-CO1, TGFB1) significantly impacts ALL cell proliferation and apoptosis, further emphasizing their critical roles in ALL progression. The identification of these hub genes not only advances mechanistic understanding of ALL relapse but also provides actionable therapeutic targets for early intervention, particularly for mitochondrial dysfunction-related strategies in CNS relapse management. Furthermore, our methodology incorporated algorithmic optimizations, including 50 randomized pre-trainings, 5-fold cross-validation, and grid search parameter tuning ( 39 – 41 ), which substantially enhanced the performance and stability of XGBoost models and Boruta algorithms compared to non-optimized approaches. This study establishes a robust framework for biomarker discovery in heterogeneous diseases, offering clinically translatable solutions for precision oncology ( 34 – 35 ). This study has several limitations. First, the data were sourced from the TARGET database, which contains relatively fewer samples in the CR and BMCR groups compared to the NR and BMR cohorts. To mitigate overfitting risks associated with small sample sizes, we did not partition internal training/validation sets. Furthermore, no external validation datasets containing ALL relapse outcomes were identified in public repositories beyond the TARGET database. To partially address this limitation, we employed 5-fold cross-validation to enhance model robustness. Fusion genes have been reported to play significant roles in ALL pathogenesis and relapse ( 42 – 43 ). However, due to data type limitations in the TARGET database, we were unable to investigate their specific contributions across relapse outcomes. Additionally, this study did not validate the serological panels using clinical samples or establish primary ALL cell lines from relapse-specific cohorts for functional hub gene verification. Nevertheless, functional assessments in ALL cell lines partially compensated for these experimental gaps. Conclusions In conclusion, this study analyzed multi-omics characteristics of ALL relapse outcomes from the TARGET database and developed optimized algorithms to construct serological monitoring panels and identify therapeutically relevant hub genes. While clinical and experimental validations remain limited, multi-dimensional data analysis, literature corroboration, and preliminary functional experiments suggest strong translational potential for the identified panels and hub genes, providing novel theoretical foundations for early relapse monitoring and precision interventions in ALL. Future studies should incorporate multi-center clinical trials to validate panel external validity, elucidate hub gene regulatory mechanisms, and explore targeted therapeutic strategies based on these genes, with the ultimate goal of improving clinical outcomes for ALL patients. Abbreviations Acute lymphoblastic leukemia (ALL) Minimal residual disease (MRD) Bone marrow relapse(BMR) Central nervous system relapse(CR) Bone marrow relapse with concurrent central nervous system relapse/combined relapse(BMCR) No relapse (NR) Central nervous system (CNS) Circulating tumor DNA (ctDNA) Overall survival (OS) Event free survival (EFS) Skin cutaneous melanoma (SKCM) Acute myeloid leukemia (LAML) Single nucleotide polymorphisms (SNPs) Oligonucleotide polymorphisms (ONPs) Copy number variation (CNV) Gene Ontology(GO) Principal Component Analysis (PCA) Gene Set Enrichment Analysis (GSEA) Differentially expressed genes (DEGs) The Human Protein Atlas (HPA) Haemoglobin (HB) Declarations Ethics approval and consent to participate Not applicable Consent for publication Not applicable Availability of data and materials The clinical, whole-exome mutation, CNV and transcriptomic data utilized in our research were obtained from the TARGET database. The data generated during the analysis and training phases , the source code, and the improved algorithms are available at https://github.com/Vera8023/Article.git. Competing interests The authors declare that they have no competing interests. Funding This study was supported by the Joint Key Project of the Chongqing Science and Technology Bureau and the Chongqing Municipal Health Commission (2025ZDXM001), and Key Project of Chongqing Municipal Education Commission (KJZD-K202400103), Chongqing Technology Innovation and Application Development, Chuan-Yu (Sichuan-Chongqing) Scientific and Technological Innovation Cooperation Program (CSTB2024TIAD-CYKJCXX0031), 2024 Hospital-level Cultivation Project of Chongqing University Jiangjin Hospital (2024YCXM010), Research Startup Funding Project of Chongqing University Jiangjin Hospital (2025qdjfxm001, 2025qdjfxm002), Key Research Project for Enhancing Medical Service Capabilities of County-Level Medical Institutions in 2025 (PS202511), Guangdong Medical Science and Technology Research Fund Project (B2021181), and Chongqing Municipal Science and Technology Bureau, Natural Science Fund (Chongqing Science and Technology Development Foundation) Project (CSTB2024NSCQ-KJFZMSX0018), Chongqing Health Commision and Science and Technology Bureau (2026MSXM018), Chongqing Youth Outstanding Medical Talent Project (YXQN2025049)。 Authors' contributions Xin He and Pu Li wrote the paper and conceived the design of the experiment; Xin He,Jianhong Zhang and Guowei He analyzed the data; Jing Shi, Bo Tian, Jianhong Zhang and Qian Lou collated the data; Delu Gan collated the pictures, Guowei He collated the tables; Bin Tang, Jing Huang, Feng Li and Debing Xiang gave guidelines for revising the paper; Pu Li fixed the paper. All authors have read and approved the final submitted manuscript. Acknowledgements Not applicable References Hunger SP and Raetz EA. How I treat relapsed acute lymphoblastic leukemia in the pediatric population. Blood. 2020;136:1803-1812. Sidhu J, Gogoi MP, Krishnan S and Saha V. Relapsed Acute Lymphoblastic Leukemia. Indian J Pediatr. 2024;91:158-167. Rivera GK, Zhou Y, Hancock ML, Gajjar A, Rubnitz J, Ribeiro RC, et al. Bone marrow recurrence after initial intensive treatment for childhood acute lymphoblastic leukemia. Cancer. 2005;103:368-76. Kopmar NE and Cassaday RD. How I prevent and treat central nervous system disease in adults with acute lymphoblastic leukemia. Blood. 2023;141:1379-1388. Kantarjian HM, O'Brien S, Smith TL, Cortes J, Giles FJ, Beran M, et al. Results of treatment with hyper-CVAD, a dose-intensive regimen, in adult acute lymphocytic leukemia. J Clin Oncol. 2000;18:547-61. Reman O, Pigneux A, Huguet F, Vey N, Delannoy A, Fegueux N, et al. Central nervous system involvement in adult acute lymphoblastic leukemia at diagnosis and/or at first relapse: results from the GET-LALA group. Leuk Res. 2008;32:1741-50. Fielding AK, Richards SM, Chopra R, Lazarus HM, Litzow MR, Buck G, et al. Outcome of 609 adults after relapse of acute lymphoblastic leukemia (ALL); an MRC UKALL12/ECOG 2993 study. Blood. 2007;109:944-50. Surapaneni UR, Cortes JE, Thomas D, O'Brien S, Giles FJ, Koller C, et al. Central nervous system relapse in adults with acute lymphoblastic leukemia. Cancer. 2002;94:773-9. van der Velden VH, de Launaij D, de Vries JF, de Haas V, Sonneveld E, Voerman JS, et al. New cellular markers at diagnosis are associated with isolated central nervous system relapse in paediatric B-cell precursor acute lymphoblastic leukaemia. Br J Haematol. 2016;172:769-81. Saygin C, Cannova J, Stock W and Muffly L. Measurable residual disease in acute lymphoblastic leukemia: methods and clinical context in adult patients. Haematologica. 2022;107:2783-2793. Pierce E, Mautner B, Mort J, Blewett A, Morris A, Keng M, et al. MRD in ALL: Optimization and Innovations. Curr Hematol Malig Rep. 2022;17:69-81. Bartram J, Patel B and Fielding AK. Monitoring MRD in ALL: Methodologies, technical aspects and optimal time points for measurement. Semin Hematol. 2020;57:142-148. Pessoa LS, Heringer M and Ferrer VP. ctDNA as a cancer biomarker: A broad overview. Crit Rev Oncol Hematol. 2020;155:103109. Cohen SA, Liu MC and Aleshin A. Practical recommendations for using ctDNA in clinical decision making. Nature. 2023;619:259-268. Kamtchum-Tatuene J and Jickling GC. Blood Biomarkers for Stroke Diagnosis and Management. Neuromolecular Med. 2019;21:344-368. Foster JB, Koptyra MP and Bagley SJ. Recent Developments in Blood Biomarkers in Neuro-oncology. Curr Neurol Neurosci Rep. 2023;23:857-867. Sztolsztener K, Żywno H, Hodun K, Konończuk K, Muszyńska-Rosłan K and Latoch E. Apolipoproteins-New Biomarkers of Overweight and Obesity among Childhood Acute Lymphoblastic Leukemia Survivors? Int J Mol Sci. 2022;23. Saiki R and Ogawa S. Adult Low-Hypodiploid Acute Lymphoblastic Leukemia Evolves from TP53-Mutated Clonal Hematopoiesis. Blood Cancer Discov. 2023;4:102-105. Kim R, Bergugnat H, Larcher L, Duchmann M, Passet M, Gachet S, et al. Adult Low-Hypodiploid Acute Lymphoblastic Leukemia Emerges from Preleukemic TP53-Mutant Clonal Hematopoiesis. Blood Cancer Discov. 2023;4:134-149. Baran N, Lodi A, Dhungana Y, Herbrich S, Collins M, Sweeney S, et al. Inhibition of mitochondrial complex I reverses NOTCH1-driven metabolic reprogramming in T-cell acute lymphoblastic leukemia. Nat Commun. 2022;13:2801. Campagnari A and Belver L. NOTCH1-Induced T-Cell Acute Lymphoblastic Leukemia In Vivo Models. Methods Mol Biol. 2024;2773:9-24. Ruan Y, Xie L and Zou A. Association of CDKN2A/B mutations, PD-1, and PD-L1 with the risk of acute lymphoblastic leukemia in children. J Cancer Res Clin Oncol. 2023;149:10841-10850. Ampatzidou M, Papadhimitriou SI, Paisiou A, Paterakis G, Tzanoudaki M, Papadakis V, et al. The Prognostic Effect of CDKN2A/2B Gene Deletions in Pediatric Acute Lymphoblastic Leukemia (ALL): Independent Prognostic Significance in BFM-Based Protocols. Diagnostics (Basel). 2023;13. Mengxuan S, Fen Z and Runming J. Novel Treatments for Pediatric Relapsed or Refractory Acute B-Cell Lineage Lymphoblastic Leukemia: Precision Medicine Era. Front Pediatr. 2022;10:923419. Brady SW, Roberts KG, Gu Z, Shi L, Pounds S, Pei D, et al. The genomic landscape of pediatric acute lymphoblastic leukemia. Nat Genet. 2022;54:1376-1389. Alexander TB, Gu Z, Iacobucci I, Dickerson K, Choi JK, Xu B, et al. The genetic basis and cell of origin of mixed phenotype acute leukaemia. Nature. 2018;562:373-379. Widman AJ, Shah M, Frydendahl A, Halmos D, Khamnei CC, Øgaard N, et al. Ultrasensitive plasma-based monitoring of tumor burden using machine-learning-guided signal enrichment. Nat Med. 2024;30:1655-1666. Álvez MB, Edfors F, von Feilitzen K, Zwahlen M, Mardinoglu A, Edqvist PH, et al. Next generation pan-cancer blood proteome profiling using proximity extension assay. Nat Commun. 2023;14:4308. Bhat M, Rabindranath M, Chara BS and Simonetto DA. Artificial intelligence, machine learning, and deep learning in liver transplantation. J Hepatol. 2023;78:1216-1233. Silva GFS, Fagundes TP, Teixeira BC and Chiavegatto Filho ADP. Machine Learning for Hypertension Prediction: a Systematic Review. Curr Hypertens Rep. 2022;24:523-533. Ouyang Y, Li X, Zhou W, Hong W, Zheng W, Qi F, et al. Integration of machine learning XGBoost and SHAP models for NBA game outcome prediction and quantitative analysis methodology. PLoS One. 2024;19:e0307478. Lew G, Chen Y, Lu X, Rheingold SR, Whitlock JA, Devidas M, et al. Outcomes after late bone marrow and very early central nervous system relapse of childhood B-acute lymphoblastic leukemia: a report from the Children's Oncology Group phase III study AALL0433. Haematologica. 2021;106:46-55. Verma D, Kapoor S, Kumari S, Sharma D, Singh J, Benjamin M, et al. Decoding the genetic symphony: Profiling protein-coding and long noncoding RNA expression in T-acute lymphoblastic leukemia for clinical insights. PNAS Nexus. 2024;3:pgae011. Qi H, Chi L, Wang X, Jin X, Wang W and Lan J. Identification of a Seven-lncRNA-mRNA Signature for Recurrence and Prognostic Prediction in Relapsed Acute Lymphoblastic Leukemia Based on WGCNA and LASSO Analyses. Anal Cell Pathol (Amst). 2021;2021:6692022. Pan L, Liu G, Lin F, Zhong S, Xia H, Sun X, et al. Machine learning applications for prediction of relapse in childhood acute lymphoblastic leukemia. Sci Rep. 2017;7:7402. Hu Y, Zhou Y, Zhang M, Ge W, Li Y, Yang L, et al. CRISPR/Cas9-Engineered Universal CD19/CD22 Dual-Targeted CAR-T Cell Therapy for Relapsed/Refractory B-cell Acute Lymphoblastic Leukemia. Clin Cancer Res. 2021;27:2764-2772. Kantarjian HM, DeAngelo DJ, Stelljes M, Martinelli G, Liedtke M, Stock W, et al. Inotuzumab Ozogamicin versus Standard Therapy for Acute Lymphoblastic Leukemia. N Engl J Med. 2016;375:740-53. Egan K, Kusao I, Troelstrup D, Agsalda M and Shiramizu B. Mitochondrial DNA in residual leukemia cells in cerebrospinal fluid in children with acute lymphoblastic leukemia. J Clin Med Res. 2010;2:225-9. Liu Q, Yang L, Shi Z, Yu J, Si H, Jin Y, et al. Development and validation of a preliminary clinical support system for measuring the probability of incident 2-year (pre)frailty among community-dwelling older adults: A prospective cohort study. Int J Med Inform. 2023;177:105138. Bhagat SK, Tiyasha T, Awadh SM, Tung TM, Jawad AH and Yaseen ZM. Prediction of sediment heavy metal at the Australian Bays using newly developed hybrid artificial intelligence models. Environ Pollut. 2021;268:115663. Tarwidi D, Pudjaprasetya SR, Adytia D and Apri M. An optimized XGBoost-based machine learning method for predicting wave run-up on a sloping beach. MethodsX. 2023;10:102119. Gale RP. Progress in Transplants for Acute Lymphoblastic Leukemia. Clin Cancer Res. 2022;28:813-815. Beyermann B, Agthe AG, Adams HP, Seeger K, Linderkamp C, Goetze G, et al. Clinical features and outcome of children with first marrow relapse of acute lymphoblastic leukemia expressing BCR-ABL fusion transcripts. BFM Relapse Study Group. Blood. 1996;87:1532-8. Additional Declarations No competing interests reported. Supplementary Files Supplementaryfiguresandfigurelegends.docx Graphicalabstract.png Graphical abstract Cite Share Download PDF Status: Under Review Version 1 posted Reviewers invited by journal 14 Apr, 2026 Submission checks completed at journal 13 Apr, 2026 First submitted to journal 13 Apr, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8999580","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":626317823,"identity":"5530ed40-fd21-46f6-ab05-4c26a83e1f41","order_by":0,"name":"Xin He","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Xin","middleName":"","lastName":"He","suffix":""},{"id":626317828,"identity":"68cef920-f90b-43d8-aa33-7ad082820129","order_by":1,"name":"Jianhong Zhang","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Jianhong","middleName":"","lastName":"Zhang","suffix":""},{"id":626317836,"identity":"d6e835be-a1e7-4b18-936b-ee92d4599f56","order_by":2,"name":"Guowei He","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Guowei","middleName":"","lastName":"He","suffix":""},{"id":626317837,"identity":"e8562508-8af5-4a5b-8935-9ee4736c5f5f","order_by":3,"name":"Jing Shi","email":"","orcid":"","institution":"First Affiliated Hospital of Chongqing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Jing","middleName":"","lastName":"Shi","suffix":""},{"id":626317840,"identity":"f69464d0-ccde-4e5b-8b6b-3cdeaf142d60","order_by":4,"name":"Bo Tian","email":"","orcid":"","institution":"Chongqing Tenth People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Bo","middleName":"","lastName":"Tian","suffix":""},{"id":626317842,"identity":"115f0d66-82b3-4f8e-a3f9-ae6c835ce8a7","order_by":5,"name":"Jing Huang","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Jing","middleName":"","lastName":"Huang","suffix":""},{"id":626317844,"identity":"21d3266d-2afd-4e80-a9c7-f2174a9d4f47","order_by":6,"name":"Feng Li","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Feng","middleName":"","lastName":"Li","suffix":""},{"id":626317848,"identity":"ccfdeedf-1cc3-4cb7-8bb5-ea19fdaa22e8","order_by":7,"name":"Debing Xiang","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Debing","middleName":"","lastName":"Xiang","suffix":""},{"id":626317853,"identity":"7fdbc190-1cc5-4d66-a6d4-f5faec02c796","order_by":8,"name":"Bin Tang","email":"","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":false,"prefix":"","firstName":"Bin","middleName":"","lastName":"Tang","suffix":""},{"id":626317854,"identity":"19ddf13a-22b6-412d-a010-507fecc33f0e","order_by":9,"name":"Qian Lou","email":"","orcid":"","institution":"Chongqing Fifth People's Hospital","correspondingAuthor":false,"prefix":"","firstName":"Qian","middleName":"","lastName":"Lou","suffix":""},{"id":626317865,"identity":"8c2f437b-6fb3-408f-871a-152e41482b9e","order_by":10,"name":"Delu Gan","email":"","orcid":"","institution":"Second Affiliated Hospital of Chongqing Medical University","correspondingAuthor":false,"prefix":"","firstName":"Delu","middleName":"","lastName":"Gan","suffix":""},{"id":626317869,"identity":"bc86a05c-1c60-4271-8ec0-5c8d39f6f9b8","order_by":11,"name":"Pu Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvElEQVRIiWNgGAWjYBACPmYogx9CMeNUCQdsMDWSDURrgTEMDhCthZ33mARDjXXi5uOHt0kwVFgnNrCfPUDAYXxpEgzH0hO3nUkrk2A4k57YwJOXQEALj5kEY8PhxG03QIy2w4kNEjwGxGnZPAPE+EeKlg0SUAYRWviSLRKOpRvPOJNWDGa08eTg18LPf/bgjQ811rL97Yc3QhjsZ/BrYWDgYWBIgESHAZCBFFN4tTDAtIyCUTAKRsEowAYAIVY4W/p4H08AAAAASUVORK5CYII=","orcid":"","institution":"Chongqing University Jiangjin Hospital","correspondingAuthor":true,"prefix":"","firstName":"Pu","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2026-03-01 06:38:13","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8999580/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8999580/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":107617102,"identity":"238a3eac-53f8-4c77-b40e-56440bd06129","added_by":"auto","created_at":"2026-04-23 09:17:19","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":670693,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eClinical Characteristics of ALL Relapse Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Sample distribution of ALL patients across relapse outcomes in the TARGET multidimensional data. (B) Association significance and strength between clinical features and ALL relapse outcomes. (C) OS (top) and EFS (bottom) curves for ALL patients across relapse groups.\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/44a51fc850cd8c070bb507d8.png"},{"id":107617100,"identity":"55188d71-63f0-4afe-9d29-c8cfdcf0b9ca","added_by":"auto","created_at":"2026-04-23 09:17:19","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":817190,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eExome-Wide Mutation Analysis in ALL Relapse Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Comparative mutation burden profiles of ALL relapse groups and pan-cancer cohorts. (B) Mutation burden analysis of ALL samples across relapse outcomes. (C) Mutation type distribution and top 10 mutated genes in ALL relapse groups. (D) Cosine similarity analysis of mutational signatures in each relapse group. (E) Oncogenic pathway alteration profiles (mutation frequency\u0026gt;0) across relapse groups.\u003c/p\u003e\n\u003cp\u003eTMB: Tumor mutation burden; SKCM: Skin Cutaneous Melanoma; LAML: Acute Myeloid Leukemia; ONP: Oligonucleotide Polymorphism\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/b54a3d7fa2f7ffa9c2388c7b.png"},{"id":107617121,"identity":"1e72ff23-6461-4bcc-8ab8-6d907ca94f27","added_by":"auto","created_at":"2026-04-23 09:17:20","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":926587,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCNV Analysis in ALL Relapse Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Bar plot displaying the number of differential CNV genes in ALL relapse outcome groups. (B) Network plot illustrating chromosomal enrichment results of differential CNV genes per group. (C) Heatmap showing biotype classification of differential CNV genes in each group. (D) Bubble plot of top 10 GO enrichment terms for differential CNV genes in each group, ranked by p-value. (E) Heatmap depicting gain/loss frequencies of top 30 differential CNV genes (ranked by p-value) across groups.\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/ef911e1df61b26fe2aa80ed1.png"},{"id":107617101,"identity":"6002665d-9153-40d9-854f-924de420b007","added_by":"auto","created_at":"2026-04-23 09:17:19","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":979242,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eTranscriptomic Characterization of ALL Relapse Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) PCA plot displaying global expression pattern differences among ALL relapse groups at the transcriptomic level. (B) Bar plot showing the number of DEGs in each group. (C) Heatmap illustrating biotype distribution of DEGs across groups. Bubble plots of top 10 pathways from GO (D) and Reactome (E) GSEA analyses, ranked by absolute NES scores.\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/402b40cf2a6c0ea31e3b3b00.png"},{"id":107707181,"identity":"25039ae0-2a40-4131-b006-0386edd81e55","added_by":"auto","created_at":"2026-04-24 09:19:44","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1544002,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eOptimized XGBoost Combined with SHAP Algorithm Identifies Optimal Monitoring Panels for ALL Relapse Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Network plot showing union of plasma protein-coding genes from HPA and HPPP databases. (B) Venn diagram illustrating intersections between plasma protein-coding genes and group-specific upregulated genes. (C) Bubble plot displaying final SHAP values of top 10 candidate genes per group after pre-training. (D) Bar plot showing mean AUC values and HL p-values for candidate panels in each group.\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/ba42ed3b6f60ff76c7c2fa14.png"},{"id":107707305,"identity":"6637fc26-f049-4c90-bb34-d38ec900fc0b","added_by":"auto","created_at":"2026-04-24 09:20:02","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":1543624,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eIdentification and Validation of Hub Genes in ALL Relapse Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e(A) Bar plot showing hub gene screening results using the Boruta algorithm in each group. Group-specific genes with confirmed importance scores of 50 were designated as hub genes. (B) Bar plot displaying differential expression fold changes of representative hub genes (CD74, MT-CO1, TGFB1) across ALL relapse outcomes. (C) QPCR validation of stable CD74, MT-CO1, and TGFB1 knockdown efficiency in SUP-B15 cells. Functional validation through CCK8 proliferation assays (D), BrdU incorporation assays (E), and apoptosis assays (F) post-knockdown. All experiments were performed in triplicate with paired t-test statistics. Data are presented as mean ± SD. *: p\u0026lt;0.05; **: p\u0026lt;0.01; ***: p\u0026lt;0.001.\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/e73aa845563c4bb218d3e4d7.png"},{"id":107709165,"identity":"f27aa319-acc3-4298-92ca-63100fda0552","added_by":"auto","created_at":"2026-04-24 09:34:53","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":7009746,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/1f4fd17f-b1aa-4c8c-a028-9a788b2b9022.pdf"},{"id":107617104,"identity":"e5f48f45-1713-4ef4-812f-451731f66fbf","added_by":"auto","created_at":"2026-04-23 09:17:19","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":9202873,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementaryfiguresandfigurelegends.docx","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/8b027780083de35f07514161.docx"},{"id":107705815,"identity":"ce159611-ce85-40b5-ac55-df9919f8b75f","added_by":"auto","created_at":"2026-04-24 09:15:23","extension":"png","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":312608,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGraphical abstract\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Graphicalabstract.png","url":"https://assets-eu.researchsquare.com/files/rs-8999580/v1/8f13a768f0e3dfbc0dc5eaa1.png"}],"financialInterests":"No competing interests reported.","formattedTitle":"Identification of monitoring panels and hub genes for heterogeneous relapse outcomes in acute lymphoblastic leukemia","fulltext":[{"header":"Introduction","content":"\u003cp\u003eRelapse remains a pivotal challenge impacting long-term survival and quality of life in acute lymphoblastic leukemia (ALL), a malignancy originating from lymphoid progenitor cells with complex pathogenesis involving dysregulated gene expression networks. Despite therapeutic advances, 15\u0026ndash;30% of ALL patients experience relapse, with post-relapse 5-year survival rates plummeting to 36% (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e). Relapse manifests in diverse anatomical patterns, including isolated bone marrow relapse (BMR, 70\u0026ndash;80% of cases), isolated central nervous system relapse (CR, ~\u0026thinsp;10\u0026ndash;15% of cases), and combined bone marrow-central nervous system relapse (BMCR, ~\u0026thinsp;5% of cases) (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e). Notably, adult ALL patients with central nervous system (CNS) relapse involvement exhibit dismal prognoses, with a median survival duration of less than one year (\u003cspan additionalcitationids=\"CR5 CR6 CR7\" citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e). The heterogeneity of relapse patterns and persistent high relapse rates pose a significant challenge to improving ALL patients survival outcomes. Consequently, monitoring different recurrence outcomes and identifying hub genes from networks holds critical implications for early therapeutic intervention and survival improvement in high-risk ALL populations.\u003c/p\u003e \u003cp\u003eMinimal residual disease (MRD) detection, currently the cornerstone of ALL relapse surveillance, demonstrates high sensitivity in monitoring BMR. However, its performance declines in CR and BMCR, with reduced sensitivity and specificity for distinguishing distinct relapse outcomes (\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e). Furthermore, MRD assessment requires invasive bone marrow aspiration, posing procedural risks and patient discomfort while risking false-negative results due to sampling variability in timing and anatomical location (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Serological biomarkers, as a non-invasive alternative, could circumvent these limitations (\u003cspan additionalcitationids=\"CR14 CR15 CR16\" citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e), yet no validated serological monitoring tool exists for ALL relapse. Consequently, the development of minimally invasive, outcome-specific serological models with high predictive accuracy remains an unmet clinical need. Beyond diagnostic limitations, the pathophysiological mechanisms underlying ALL relapse outcomes remain poorly characterized. Although mutations in TP53, NOTCH1, and epigenetic alterations have been implicated in relapse pathogenesis (\u003cspan additionalcitationids=\"CR19 CR20 CR21 CR22\" citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e), these studies predominantly adopt single-omics approaches or focus on isolated genes, lacking multidimensional integration and systemic network analysis (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e). This fragmented perspective hinders comprehensive understanding of relapse heterogeneity. To address this limitation, multi-omics integration and holistic network modeling are essential to delineate relapse-specific molecular signatures and identify hub genes regulating these distinct outcomes. Such insights would enable precision intervention strategies targeting relapse-driving genes, thereby improving clinical outcomes in high-risk ALL.\u003c/p\u003e \u003cp\u003eThe Therapeutically Applicable Research to Generate Effective Treatments (TARGET) database integrates multi-dimensional datasets across seven malignancies, including ALL, providing a critical resource for identifying predictive, therapeutic, and prognostic targets in cancer progression (\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e). This study utilized clinical, whole-exome mutation, copy number variation (CNV), and transcriptomic data from ALL patients to systematically investigate the biological foundations of four relapse outcomes: no relapse (NR), BMR, CR, and BMCR. Specifically, we firstly elucidated associations between clinical variables and relapse outcomes, alongside survival prognosis evaluation for each relapse group. Simultaneously, our research characterized molecular alterations across genomic, epigenomic, and transcriptomic layers to delineate relapse-specific patterns. Subsequently, following the selection of the optimal parameters for the XGBoost training model using 5-fold cross-validation and grid search, the prioritisation of features through SHAP interpretability analysis, the evaluation of the model's discrimination and calibration using AUC and HL p-values, the 50 random trainings based on optimised algorithm (\u003cspan additionalcitationids=\"CR28 CR29 CR30\" citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e), and the further training of the top five indicators in various combinations, the optimal serum panels for monitoring different recurrence outcomes were screened. Ultimately, we constructed protein-protein interaction networks and applied the Boruta algorithm to identify hub genes significantly associated with each relapse group, and the functional validation of candidate hub genes (e.g., CD74, TGFB1, MT-CO1) was performed using in vitro cellular assays to confirm their biological relevance in ALL pathogenesis. Our research employs a multi-omics perspective to analyze clinical issues, constructing serological models for minimally invasive relapse monitoring and seeking mechanistic insights into relapse outcomes. These findings hold significant potential to transform clinical management strategies by enabling early relapse detection and precision targeting of relapse-driving pathways in ALL.\u003c/p\u003e"},{"header":"Materials and methods","content":"\n\u003ch3\u003e1. Clinical Data Acquisition\u003c/h3\u003e\n\u003cp\u003eData for this study were sourced from the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portal.gdc.cancer.gov/analysis_page?app=CohortBuilder\u0026amp;tab=general\u003c/span\u003e\u003cspan address=\"https://portal.gdc.cancer.gov/analysis_page?app=CohortBuilder\u0026amp;tab=general\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). We retrieved open-access samples with a primary diagnosis of acute lymphoblastic leukaemia (ALL) from the TARGET-ALL-P1, TARGET-ALL-P2, and TARGET-ALL-P3 datasets in this database. The samples encompassed clinical, whole-exome mutation, CNV, and transcriptomics data. Following the processes of data merging and collation, samples that satisfied the subsequent criteria were selected for further analysis. Firstly, the cancer type was primary blood-derived cancer. Secondly, the primary tumor site was bone marrow. Thirdly, the follow-up recurrence outcomes were clearly classified as one of the following four types: NR, BMR, CR, and BMCR. Ultimately, a total of 927 clinical samples met these inclusion criteria, including 618 with whole-exome mutation data, 250 with CNV profiles, and 343 with transcriptomic data. With the exception of transcriptomics data, the sample sizes for the remaining data were consistent with the number of clinical cases. The detailed clinical characteristics are summarized in supplementary materials.\u003c/p\u003e\n\u003ch3\u003e2. Clinical Data Analysis\u003c/h3\u003e\n\u003cp\u003eThe clinical dataset comprised seven categorical variables, including \u003cem\u003eGender\u003c/em\u003e and \u003cem\u003eRace\u003c/em\u003e. We tabulated the frequency of each categorical variable across the four relapse outcomes and quantified significance and strength of association using the `assocstats` function from the R package `vcd`. Pearson's chi-squared test was employed for contingency tables with sample sizes\u0026thinsp;\u0026ge;\u0026thinsp;40 and all cell counts\u0026thinsp;\u0026ge;\u0026thinsp;5 to analyse the association significance. While the likelihood ratio chi-squared test was used for tables containing cell counts\u0026thinsp;\u0026lt;\u0026thinsp;5 to analyse the association significance. The chi-square test p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 denoted statistical significance. Association strength was measured using the Contingency Coefficient and Cramer's V, with values approaching 1 indicating stronger associations and values near 0 suggesting weaker associations.\u003c/p\u003e\n\u003ch3\u003e3. Genomic Mutation Analysis\u003c/h3\u003e\n\u003cp\u003eTo investigate mutation profiles across ALL relapse outcomes, whole-exome mutation data were analyzed using the R package `maftools`. After importing data via the `read.maf` function, mutation burden was compared between ALL relapse groups and TCGA pan-cancer cohorts using `tcgaCompare`. Differential mutation burden across relapse groups was assessed with the `tmb` function, and mutation summary plots were generated using `plotmafSummary`. Mutational signature analysis was performed using cosine similarity scoring against 60 COSMIC mutational signatures (v3.2), incorporating the `BSgenome.Hsapiens.UCSC.hg38` reference genome and `NMF` package. Pathway enrichment analysis evaluated the contribution of mutated genes to 24 established oncogenic signaling pathways using the `pathways` function. The mutant data are available in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e4. CNV Analysis\u003c/h3\u003e\n\u003cp\u003eTo analyze CNV differences across ALL relapse outcomes, genes with a copy number of 2 were labeled as \"neutral\", while those with copy numbers other than 2 were classified as \"change\". Four pairwise comparisons were performed: NR vs. Other (non-NR samples), BMR vs. Other (non-BMR samples), CR vs. Other (non-CR samples), and BMCR vs. Other (non-BMCR samples). For each comparison, 2\u0026times;2 contingency tables (rows: CNV status; columns: relapse groups) were constructed, and chi-squared tests were applied to identify differentially altered CNV genes. Additionally, genes were categorized as \"loss\" (copy number\u0026thinsp;\u0026lt;\u0026thinsp;2), \"neutral\" (copy number\u0026thinsp;=\u0026thinsp;2), or \"gain\" (copy number\u0026thinsp;\u0026gt;\u0026thinsp;2). The frequency of gain/loss events was calculated as the proportion of affected samples within each relapse group to quantify genomic instability patterns. The CNV data are available in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e5. Transcriptomic Differential Expression Analysis\u003c/h3\u003e\n\u003cp\u003eRNA-seq count data were normalized and analyzed for differential expression using the R package DESeq2. Pairwise comparisons (NR vs. Other, BMR vs. Other, CR vs. Other, BMCR vs. Other) identified differentially expressed genes (DEGs) with thresholds of p\u0026thinsp;\u0026lt;\u0026thinsp;0.05 and |log₂FC|\u0026gt;1. The RNA-seq data are available in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e6. Principal Component Analysis (PCA)\u003c/h3\u003e\n\u003cp\u003eTo visualize transcriptomic heterogeneity across relapse groups, DESeq2-normalized counts underwent log₂(x\u0026thinsp;+\u0026thinsp;1) transformation. Dimensionality reduction and visualization were performed using PCA from the R package factoextra, with the `fviz_pca_ind` function generating PCA plot to illustrate sample clustering patterns.\u003c/p\u003e\n\u003ch3\u003e7. Gene Biotype Annotation\u003c/h3\u003e\n\u003cp\u003eGene biotypes for DEGs were obtained using the `annoGene` function from the R package AnnoProbe, enabling functional classification of differentially expressed transcripts (e.g., protein-coding, lncRNA, pseudogene).\u003c/p\u003e\n\u003ch3\u003e8. Enrichment Analysis\u003c/h3\u003e\n\u003cp\u003eFor differentially altered CNV genes, chromosome enrichment was assessed using the `MSigDB c1.all.v2023.2.Hs.symbols.gmt` file and the `enricher` function from the R package clusterProfiler. Gene Ontology(GO) enrichment was performed via the DAVID database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://david.ncifcrf.gov/\u003c/span\u003e\u003cspan address=\"https://david.ncifcrf.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) with parameters counts\u0026thinsp;=\u0026thinsp;2 and ease\u0026thinsp;=\u0026thinsp;0.1. Both adopt over-representation analysis (ORA) for enrichment. For DEGs, GO and Reactome pathway enrichments were conducted using MSigDB's `c5.go.bp.v2024.1.Hs.symbols.gmt` and `c2.cp.reactome.v2024.1.Hs.symbols.gmt` files, respectively, with Gene Set Enrichment Analysis (GSEA). This approach enabled assessment of pathway activation or inhibition status and quantification of effect magnitude via normalized enrichment scores (NES): NES\u0026thinsp;\u0026gt;\u0026thinsp;0 indicates pathway activation, NES\u0026thinsp;\u0026lt;\u0026thinsp;0 indicates suppression, and absolute NES values positively correlate with effect strength. To characterize genes in optimal monitoring panels, we retrieved functional annotations from the GeneCards database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.genecards.org/\u003c/span\u003e\u003cspan address=\"https://www.genecards.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and highlighted representative GO pathways from the 'Pathways' section. To explore regulatory roles of hub genes, GO ORA were conducted through DAVID database with parameters counts\u0026thinsp;=\u0026thinsp;1 and ease\u0026thinsp;=\u0026thinsp;1.\u003c/p\u003e\n\u003ch3\u003e9. Plasma Protein Curation and Data Integration\u003c/h3\u003e\n\u003cp\u003ePlasma protein-coding genes were compiled from the Human Protein Atlas (HPA, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.proteinatlas.org/\u003c/span\u003e\u003cspan address=\"https://www.proteinatlas.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and the Human Plasma Proteome Project (HPPP, \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.hupo.org/plasma-proteome-project/\u003c/span\u003e\u003cspan address=\"https://www.hupo.org/plasma-proteome-project/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). The union of these datasets was used to filter plasma protein-coding candidates from DEGs. The plasma protein-coding genes are available in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e10. Disease Database Mining\u003c/h3\u003e\n\u003cp\u003eThe MalaCards human disease annotation database (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.malacards.org/\u003c/span\u003e\u003cspan address=\"https://www.malacards.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was queried to assess associations between hematopoietic malignancies and genes in optimal panels/hub gene sets, utilizing gene-disease association scores.\u003c/p\u003e\n\u003ch3\u003e11. Optimal Panel Selection and Monitoring Model Development\u003c/h3\u003e\n\u003cp\u003eFour relapse outcomes (NR vs. Other, BMR vs. Other, CR vs. Other, BMCR vs. Other) were modeled as binary endpoints, with upregulated plasma protein-coding DEGs serving as predictive features. Pre-trained XGBoost models were optimized via 5-fold cross-validation and grid search for hyperparameter tuning. Feature importance was quantified using SHAP values from the R package SHAPforxgboost. Model discrimination (AUC) and calibration (Hosmer-Lemeshow p-value) were evaluated with the tidymodels and PredictABEL packages. To enhance robustness, pre-training was repeated across 50 random seeds and results were retained with criteria of AUC\u0026thinsp;\u0026lt;\u0026thinsp;0.95 and HL p\u0026thinsp;\u0026lt;\u0026thinsp;0.05. Genes were ranked in descending order by mean SHAP values calculated under retention criteria, and the top 5 genes from each group were combined to generate 31 candidate panels. Each panel underwent retraining under reserved pre-training seeds with optimal parameters, The optimal panel for each relapse outcome was selected based on maximum mean AUC with mean HL p\u0026thinsp;\u0026gt;\u0026thinsp;0.05. The training data are available in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e12. Hub Gene Identification\u003c/h3\u003e\n\u003cp\u003eInteraction networks for DEGs in each relapse group were constructed using the STRINGdb package (STRING database version 12.0, Homo sapiens, minimum interaction score threshold of 700). Degree values (number of direct interactions) for each gene node were calculated using the igraph package. For BMR vs. Other, CR vs. Other, and BMCR vs. Other comparisons, DEGs with degree\u0026thinsp;\u0026gt;\u0026thinsp;10 were selected as input features for the Boruta algorithm. To enhance stability, Boruta training was repeated across 50 random seeds, with features classified as \"Confirmed\" (definitively important), \"Tentative\" (potentially important), or \"Rejected\" (not important) based on variable importance. Genes that were classified as \"Confirmed\" in all 50 iterations and were unique to one group were retained as group-specific hub genes.The networks and Boruta results are available in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e13. Quantitative real-time PCR (qPCR)\u003c/h3\u003e\n\u003cp\u003eTotal RNA was isolated by TRIzol\u0026reg; Reagent (Tiangen, China) and converted to cDNA using a reverse transcription kit (Tiangen, China). Then, the expression of target gene was detected by real-time PCR (SYBR green) with specific primers. The PCR reaction conditions were as follows: 95\u0026deg;C for 3 min, followed by 40 cycles of 95\u0026deg;C for 15s, 60\u0026deg;C for 15s, and 72\u0026deg;C for 20 s, concluding with an extension step of 72\u0026deg;C for 5min. The data were analyzed using 2\u003csup\u003e\u0026minus;∆∆Ct\u003c/sup\u003e method. Primers were listed in the supplementary materials.\u003c/p\u003e\n\u003ch3\u003e14. Cell lines and culture conditions\u003c/h3\u003e\n\u003cp\u003eThe human Ph\u0026thinsp;+\u0026thinsp;ALL cell line SUP-B15 and human embryonic kidney cell line 293FT were purchased from the American Type Culture Collection. SUP-B15 cells were cultured in Iscove's modified Dulbecco's medium (IMDM) (Cytiva, Cat# SH30228.FS) supplemented with 1% penicillin-streptomycin, 20% FBS (Gibco; Thermo Fisher Scientific, Inc.), and 0.05 mM β-mercaptoethanol (Gibco, Cat# 21985-023). 293FT cells were cultured in DMEM (HyClone, Cat# SH30022.01) with 10% FBS. All cells were maintained in a humidified incubator at 37\u0026deg;C in an atmosphere of 5% CO₂.\u003c/p\u003e\n\u003ch3\u003e15. Lentivirus Production and Cell Transduction\u003c/h3\u003e\n\u003cp\u003eShRNA sequences targeting CD74, MT-CO1, and TGFB1, along with a negative control sequence, were respectively cloned into the pLKO.1 lentiviral vector (addgene, #8453). Lentiviral particles were produced by co-transfecting 293FT cells with the lentiviral vector, pMD2.G, and psPAX2 viral packaging plasmids using Lipofectamine 3000 (Invitrogen, USA) according to the manufacturer's protocol. Viral supernatants were harvested at 48 h and 72 h post-transfection, and viral titers were determined before storage at -80\u0026deg;C. For cell infection, SUP-B15 cells were incubated with viral culture medium supplemented with 10 \u0026micro;g/mL polybrene. Stably transduced cells were selected using 2 \u0026micro;g/mL puromycin for 14 days.\u003c/p\u003e\n\u003ch3\u003e16. Cell Proliferation Assay\u003c/h3\u003e\n\u003cp\u003eStable SUP-B15 cell lines were seeded at 2,000 cells/well in 96-well plates (30 replicates per sample). Cell viability was measured daily for 5 days using the Cell Counting Kit-8 (MCE, HY-K0301) according to the manufacturer's instructions, with five wells measured per day.\u003c/p\u003e\n\u003ch3\u003e17. Flow Cytometry Analysis\u003c/h3\u003e\n\u003cp\u003eFor S-phase analysis, stable cells were pulsed with 100 \u0026micro;M 5-bromo-2'-deoxyuridine (BrdU, Thermo Fisher Scientific) for 1 h at 37\u0026deg;C before harvesting. Then cells were fixed with 70% ethanol at 4\u0026deg;C overnight. Next, permeabilized with 0.3% Triton X-100 for 30 min, and stained with FITC-anti-BrdU (Thermo Fisher Scientific) for 1 h at room temperature. Finally, cells were labeled with 50 ng/mL propidium iodide (PI, Thermo Fisher Scientific) and 10 \u0026micro;g/mL RNase A (Thermo Fisher Scientific) for 30 min at 37\u0026deg;C, and then analysed by flow cytometry. For apoptosis analysis, cells were stained with PI and APC-conjugated annexin V (BestBio, C1062) according to the manufacturer's protocol prior to flow cytometry. Data acquisition was performed on a CytExpert flow cytometer (Beckman Coulter, Inc), with a minimum of 20,000 events analyzed per sample.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003e1. Clinical Characteristics of ALL Relapse Outcomes\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo investigate clinical features associated with distinct relapse outcomes in ALL, we analyzed clinical data from 927 TARGET-ALL patients categorized into four groups: 708 cases in the NR group, 167 with BMR, 30 with CR, and 22 with BMCR (Fig. 1A). Assessment of the significance and strength of associations between categorical clinical variables and various recurrence outcomes revealed no significant intergroup differences in \u003cem\u003eRace\u003c/em\u003e, \u003cem\u003eMLL status\u003c/em\u003e, and \u003cem\u003eDown syndrome\u003c/em\u003e. However, \u003cem\u003eGender\u003c/em\u003e, \u003cem\u003eCNS status at diagnosis\u003c/em\u003e, \u003cem\u003eETV6/RUNX1 fusion status\u003c/em\u003e, \u003cem\u003eTRISOMY 4/10 status\u003c/em\u003e, and \u003cem\u003eTGF3 PBX1 status\u003c/em\u003e differed significantly across groups (Fig. 1B), suggesting these differential variables may contribute to relapse heterogeneity. Additionally, \u003cem\u003eVital status\u003c/em\u003e showed significant difference among groups and the strongest association with relapse outcome, with BMCR patients exhibiting the highest mortality rate (77.3%, 17/22), followed by BMR (61.1%, 102/167) and CR (43.3%, 13/30) (Fig. 1B). Survival analyses also confirmed that BMCR patients had the poorest overall survival (OS) and event free survival (EFS) outcomes (Fig. 1C). The findings emphasize the significance of recurrence patterns, particularly BMCR, in determining prognosis, further highlighting the urgent need for early detection and intervention in ALL recurrence.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2. Exome-Wide Mutation Analysis in ALL Relapse Outcomes\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e618 ALL samples with mutation data and varying relapse outcomes from the TARGET database were analyzed to characterize mutation profiles (Fig. 1A). Pan-cancer mutation burden comparison revealed that ALL relapse groups exhibited lower overall mutation rates compared to most solid tumors but higher than acute myeloid leukemia (LAML), suggesting relatively elevated genomic instability in ALL among hematologic malignancies (Fig. 2A). Notably, the BMCR group displayed the highest differential mutation burden (Fig. 2B), aligning with its poorest survival outcomes. These suggest a potential link between elevated genomic alterations and ALL malignant progression. Mutation signature analysis presented high similarity among groups in terms of variant classification, variant type, and SNV class but divergent top 10 mutated genes (Fig. 2C, Fig. S1). Mutational etiology assessment indicated similar exposure profiles among NR, BMR, and BMCR groups, with the exception of the CR patients, implying that mutation triggers do not explain group mutation heterogeneity (Fig. 2D, Fig. S2). Pathway enrichment analysis revealed group-specific oncogenic signaling alterations: MAPK signaling (25.22% mutation frequency) in NR, genome integrity pathways (13.51%) in BMR, NOTCH signaling (35.29%) in CR, and chromatin regulation/NOTCH/ubiquitin-proteasome systems (40% combined frequency) in BMCR (Fig. 2E, Fig. S3), suggesting that distinct oncogenic pathway dysregulation may drive relapse-specific outcomes in ALL.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3. CNV Analysis in ALL Relapse Outcomes\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e250 ALL samples with CNV data from the TARGET database were analyzed to characterize CNV profiles across relapse outcomes (Fig. 1A). Differential CNV analysis identified 2,414, 924, 3,212, and 43 significantly altered CNV genes in the NR, BMR, CR, and BMCR groups, respectively (Fig. 3A, Fig. S4). The limited number of differential CNVs in BMCR may be influenced by its smaller sample size. Chromosomal enrichment, gene biotype annotation, and GO enrichment analyses were performed to interpret group-specific CNV patterns. Notably, NR and CR groups exhibited similar CNV profiles in chromosomal distribution, gene biotype composition, and functional enrichment, whereas BMR and BMCR groups showed distinct patterns (Fig. 3B-D). Comprehensive analysis of the top 30 differential CNVs revealed predominant CNV loss in NR and CR groups, contrasting with CNV gain in BMR and BMCR groups. However, comparisons to control cohorts indicated increased loss frequency in NR, reduced loss frequency in CR, decreased gain frequency in BMR, and elevated gain frequency in BMCR (Fig. 3E). These findings elucidate CNV characteristics associated with distinct relapse outcomes, providing novel insights into the mechanisms driving ALL relapse heterogeneity.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4. Transcriptomic Analysis of ALL Relapse Outcomes\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA comprehensive transcriptomic analysis was performed on 343 ALL samples with RNA-sequencing data and relapse outcomes from the TARGET database (Fig. 1A). PCA plots revealed group-specific expression patterns. However, the NR and CR groups exhibited high similarity, as did the BMR and BMCR groups (Fig. 4A). This transcriptional similarity aligns with previous CNV-based group characterizations, suggesting potential regulatory links between genomic alterations and gene expression profiles. Differential expression analysis identified group-specific DEGs (Fig. 4B, Fig. S5), which were subjected to biotype annotation and functional enrichment analysis using GO and Reactome databases. The biotype classification exhibited diversity, yet protein-coding genes were confirmed as the predominant DEG category across groups (Fig. 4C). GO enrichment revealed distinct biological processes. NR group showed suppressed immune response, cell differentiation, and apoptosis pathways. BMR exhibited enhanced cell activation, adhesion and oxidative phosphorylation. CR displayed activated type II interferon production and mitochondrial functions. BMCR demonstrated augmented cell activation and adhesion, leukocyte chemotaxis, and B-cell activation (Fig. 4D). Reactome pathway analysis corroborated these findings, showing predominant pathway suppression in NR versus activation in other groups (Fig. 4E). These transcriptomic disparities provide novel mechanistic insights into the molecular drivers of relapse heterogeneity in ALL.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e5. Optimized XGBoost Combined with SHAP Algorithm Identifies Optimal Serum Monitoring Panels for ALL Relapse Outcomes\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCurrent ALL relapse monitoring primarily relies on MRD detection, which is mainly applied to BMR surveillance but shows limited efficacy for CR/BMCR monitoring and requires invasive bone marrow sampling. Development of minimally invasive serological panels capable of accurately monitoring diverse relapse outcomes holds significant clinical value. To achieve this, we intersected plasma protein-coding genes from the HPA and HPPP databases with group-specific upregulated genes to identify candidate biomarkers for panel screening (Fig. 5A-B, Fig. S6A). Heatmaps revealed marked expression differences in the top 10 candidate genes across groups, suggesting their potential as relapse-specific markers (Fig. S6B). We optimized XGBoost models using 5-fold cross-validation and grid search, followed by SHAP analysis to rank feature importance. Fifty iterations of model pre-training were performed, with AUC values and Hosmer-Lemeshow (HL) test p-values used to filter suboptimal results and ensure robustness (Fig. S7). Mean SHAP values from retained iterations were calculated for final feature ranking (Fig. 5C). Combining top 5 features from each group generated 31 candidate panels, which were evaluated through iterative pre-training. Optimal panels satisfied with HL p\u0026lt;0.05 included: NR panel (LRBA, SH2B1A, IL36G, ADH1A, CEP85, AUC=0.985); BMR panel (CD22, CRMP1, EPHA2, KLK14, AUC=0.997); CR panel (ACVRL1, FAT3, SKAP1, AUC=0.990); and BMCR panel (AGRN, CP, EFEMP1, LGALS7, ST6GAL2, AUC=0.999) (Fig. 5D).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eGenomic analysis of panel genes revealed significant expression differences across groups but rare mutations (1% missense in FAT) and limited CNV alterations (statistically significant only in SH2D1A), suggesting expression regulation of panel genes caused by non-genomic factors (Fig. S8A-C). Functional characterization using the HPA protein class identified cancer/disease-associated genes (e.g., LRBA, SH2D1A) and FDA-approved drug targets (e.g., ADH1A, CD22) within optimal panels, supporting therapeutic relevance. Protein half-life predictions via ExPASy ProtParam indicated serum suitability of panel genes (\u0026gt;30 hours). Malacards analysis linked most panel genes to hematologic disorders, while GeneCards GO enrichment highlighted cell adhesion/migration pathways in BMR/CR/BMCR panels and CNS development processes in CR/BMCR panels, aligning with clinical relapse phenotypes (Fig. S8D). Survival analysis revealed poor prognosis associations with elevated levels of panel genes (CD22, LGALS7, CRMP1, AGRN, and KLK14), suggesting dual monitoring and prognostic potential (Fig. S9).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e6. Identification and Validation of Hub Genes in ALL Relapse Outcomes\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo identify hub genes driving ALL relapse progression and potential therapeutic targets, we constructed protein-protein interaction networks using the STRING database and group-specific DEGs from BMR, CR, and BMCR groups. Subnetworks were extracted with nodes exhibiting degree\u0026gt;10, representing 14.44%, 5.62%, and 10.56% of total nodes in BMR/CR/BMCR groups, respectively. These subnetworks governed 39.23%, 15.42%, and 50.41% of total interactions, highlighting their regulatory importance (Fig. S10). Boruta algorithm was applied across 50 iterations to identify group-specific hub genes with confirmed importance scores of 50, yielding 18, 13, and 3 hub genes for BMR, CR, and BMCR groups, respectively (Fig. 6A).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMulti-omics analysis revealed downregulated hub genes in CR and upregulated hub genes in BMCR, while BMR showed both up- and downregulated patterns. Notably, only ITPR2, ADCY5, and HLA-DRA exhibited missense mutations in NR samples, and CXCR3 was the sole hub gene with CNV differences across groups, suggesting expression regulation of hub genes caused by non-genomic factors (Fig. S11A, S11C-D). Functional characterization via HPA protein classes classified 28 hub genes (including ADCY5) as cancer/disease-related and 17 (including CD19) as FDA-approved/potential drug targets. Malacards analysis confirmed hematologic disease associations for all hub genes, underscoring their therapeutic relevance (Fig. S11B). DAVID GO enrichment of subnetwork-regulated hub genes revealed pathway convergence on signal transduction, biological function, immune responses, DNA replication, and transcription in BMR/BMCR groups, with additional CNS-related pathway enrichment in CR group (Fig. S12).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eRepresentative hub genes CD74 (BMR), MT-CO1 (CR), and TGFB1 (BMCR) were selected for functional validation. CD74 and TGFB1 demonstrate elevated levels of expression in the BMR and BMCR groups, respectively, while MT-CO1 exhibits reduced expression in the CR group (Fig. 6B). Consistent with expression profiles, shRNA-mediated knockdown in SUP-B15 cells (Fig. 6C) demonstrated that CD74/TGFB1 depletion enhanced proliferation and inhibited apoptosis, while MT-CO1 knockdown showed opposite effects (Fig. 6D-F). These results confirm the regulatory roles of identified hub genes in ALL pathogenesis, validating the screening approach.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eRelapse remains a critical cause of mortality in ALL, with distinct outcomes including BMR, CR, and BMCR (\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e). Current clinical surveillance predominantly relies on MRD monitoring, which is primarily limited to BMR detection and requires invasive bone marrow aspiration (\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e). Consequently, there is an urgent clinical need to develop minimally invasive serological biomarkers capable of monitoring diverse relapse outcomes. Furthermore, mechanistic insights into ALL relapse heterogeneity remain insufficient, necessitating the identification of hub genes driving relapse processes to enable early targeted interventions (\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e). Leveraging multi-omics data from the TARGET database, we systematically characterized clinical, genomic, CNV, and transcriptomic features across ALL relapse outcomes. This analysis yielded optimal serological monitoring panels and predictive models for each relapse group, alongside the identification of group-specific hub genes, thereby establishing novel pathways for early relapse detection and intervention in ALL.\u003c/p\u003e \u003cp\u003eUtilizing machine learning algorithms and transcriptomic data, we identified serological monitoring panels tailored to ALL relapse outcomes. Predictive models constructed using these panels achieved mean AUC values of 0.985, 0.997, 0.990, and 0.999 for distinguishing NR, BMR, CR, and BMCR groups from their respective controls. These panels outperform existing models such as the Seven-lncRNA-mRNA Signature (AUC\u0026thinsp;=\u0026thinsp;0.901) (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e) and clinical variable-based models incorporating age, WBC count, and hemoglobin levels (AUC\u0026thinsp;=\u0026thinsp;0.902) (\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e), demonstrating superior discriminatory capacity. The high AUC values underscore the panels' efficacy in identifying distinct relapse outcomes, offering clinicians a more precise monitoring tool. Compared to conventional MRD detection methods, serum-based panels avoid invasive bone marrow sampling, reduce procedural costs, and facilitate broader clinical implementation, highlighting their significant translational potential. Notably, several panel genes (e.g., CD22, CRMP1, KLK14, AGRN, LGALS7) exhibit prognostic relevance in ALL patients. Targeted therapies against CD22 have demonstrated efficacy in refractory/relapsed ALL (\u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e), validating the therapeutic relevance of our findings and prompting similar potential for other panel components. Additionally, through network analysis and Boruta algorithm training, we identified hub genes significantly associated with ALL relapse outcomes, including VAMP2 and EFNB2 in BMR, MT-ATP8 and CXCR3 in CR, and HLA-DRA and HLA-F in BMCR. CR-associated hub genes showed enrichment in mitochondrial functional pathways, aligning with reported alterations in mitochondrial DNA levels in cerebrospinal fluid from patients with CNS relapse (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e), implicating mitochondrial dysfunction in ALL CNS relapse pathogenesis. Functional validation experiments confirmed that knockdown of representative hub genes (CD74, MT-CO1, TGFB1) significantly impacts ALL cell proliferation and apoptosis, further emphasizing their critical roles in ALL progression. The identification of these hub genes not only advances mechanistic understanding of ALL relapse but also provides actionable therapeutic targets for early intervention, particularly for mitochondrial dysfunction-related strategies in CNS relapse management. Furthermore, our methodology incorporated algorithmic optimizations, including 50 randomized pre-trainings, 5-fold cross-validation, and grid search parameter tuning (\u003cspan additionalcitationids=\"CR40\" citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e), which substantially enhanced the performance and stability of XGBoost models and Boruta algorithms compared to non-optimized approaches. This study establishes a robust framework for biomarker discovery in heterogeneous diseases, offering clinically translatable solutions for precision oncology (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThis study has several limitations. First, the data were sourced from the TARGET database, which contains relatively fewer samples in the CR and BMCR groups compared to the NR and BMR cohorts. To mitigate overfitting risks associated with small sample sizes, we did not partition internal training/validation sets. Furthermore, no external validation datasets containing ALL relapse outcomes were identified in public repositories beyond the TARGET database. To partially address this limitation, we employed 5-fold cross-validation to enhance model robustness. Fusion genes have been reported to play significant roles in ALL pathogenesis and relapse (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e). However, due to data type limitations in the TARGET database, we were unable to investigate their specific contributions across relapse outcomes. Additionally, this study did not validate the serological panels using clinical samples or establish primary ALL cell lines from relapse-specific cohorts for functional hub gene verification. Nevertheless, functional assessments in ALL cell lines partially compensated for these experimental gaps.\u003c/p\u003e"},{"header":"Conclusions","content":"\u003cp\u003eIn conclusion, this study analyzed multi-omics characteristics of ALL relapse outcomes from the TARGET database and developed optimized algorithms to construct serological monitoring panels and identify therapeutically relevant hub genes. While clinical and experimental validations remain limited, multi-dimensional data analysis, literature corroboration, and preliminary functional experiments suggest strong translational potential for the identified panels and hub genes, providing novel theoretical foundations for early relapse monitoring and precision interventions in ALL. Future studies should incorporate multi-center clinical trials to validate panel external validity, elucidate hub gene regulatory mechanisms, and explore targeted therapeutic strategies based on these genes, with the ultimate goal of improving clinical outcomes for ALL patients.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eAcute lymphoblastic leukemia (ALL)\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eMinimal residual disease (MRD)\u003c/p\u003e\n\u003cp\u003eBone marrow relapse(BMR)\u003c/p\u003e\n\u003cp\u003eCentral nervous system relapse(CR)\u003c/p\u003e\n\u003cp\u003eBone marrow relapse with concurrent central nervous system relapse/combined relapse(BMCR)\u003c/p\u003e\n\u003cp\u003eNo relapse (NR)\u003c/p\u003e\n\u003cp\u003eCentral nervous system (CNS)\u003c/p\u003e\n\u003cp\u003eCirculating tumor DNA (ctDNA)\u003c/p\u003e\n\u003cp\u003eOverall survival (OS)\u003c/p\u003e\n\u003cp\u003eEvent free survival (EFS)\u003c/p\u003e\n\u003cp\u003eSkin cutaneous melanoma (SKCM)\u003c/p\u003e\n\u003cp\u003eAcute myeloid leukemia (LAML)\u003c/p\u003e\n\u003cp\u003eSingle nucleotide polymorphisms (SNPs)\u003c/p\u003e\n\u003cp\u003eOligonucleotide polymorphisms (ONPs)\u003c/p\u003e\n\u003cp\u003eCopy number variation (CNV)\u003c/p\u003e\n\u003cp\u003eGene Ontology(GO)\u003c/p\u003e\n\u003cp\u003ePrincipal Component Analysis (PCA)\u003c/p\u003e\n\u003cp\u003eGene Set Enrichment Analysis (GSEA)\u003c/p\u003e\n\u003cp\u003eDifferentially expressed genes (DEGs)\u003c/p\u003e\n\u003cp\u003eThe Human Protein Atlas (HPA)\u003c/p\u003e\n\u003cp\u003eHaemoglobin (HB)\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe clinical, whole-exome mutation, CNV and transcriptomic data utilized in our research were obtained from the TARGET database. The data generated during the analysis and training phases , \u0026nbsp;the source code, \u0026nbsp;and the improved algorithms are available at https://github.com/Vera8023/Article.git.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was supported by the Joint Key Project of the Chongqing Science and Technology Bureau and the Chongqing Municipal Health Commission (2025ZDXM001), and Key Project of Chongqing Municipal Education Commission (KJZD-K202400103), Chongqing Technology Innovation and Application Development, Chuan-Yu (Sichuan-Chongqing) Scientific and Technological Innovation Cooperation Program (CSTB2024TIAD-CYKJCXX0031), 2024 Hospital-level Cultivation Project of Chongqing University Jiangjin Hospital (2024YCXM010), Research Startup Funding Project of Chongqing University Jiangjin Hospital (2025qdjfxm001, 2025qdjfxm002), Key Research Project for Enhancing Medical Service Capabilities of County-Level Medical Institutions in 2025 (PS202511), Guangdong Medical Science and Technology Research Fund Project (B2021181), and Chongqing Municipal Science and Technology Bureau, Natural Science Fund (Chongqing Science and Technology Development Foundation) Project (CSTB2024NSCQ-KJFZMSX0018), Chongqing Health Commision and Science and Technology Bureau (2026MSXM018), Chongqing Youth Outstanding Medical Talent Project (YXQN2025049)。\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eXin He and Pu Li wrote the paper and conceived the design of the experiment; Xin He,Jianhong Zhang and Guowei He analyzed the data; Jing Shi, Bo\u0026nbsp;Tian, Jianhong Zhang and Qian Lou collated the data; Delu Gan collated the pictures, Guowei He collated the tables; Bin Tang, Jing Huang, Feng Li and Debing Xiang gave guidelines for revising the paper; Pu Li fixed the paper. All authors have read and approved the final submitted manuscript.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eHunger SP and Raetz EA. How I treat relapsed acute lymphoblastic leukemia in the pediatric population. Blood. 2020;136:1803-1812.\u003c/li\u003e\n\u003cli\u003eSidhu J, Gogoi MP, Krishnan S and Saha V. Relapsed Acute Lymphoblastic Leukemia. Indian J Pediatr. 2024;91:158-167.\u003c/li\u003e\n\u003cli\u003eRivera GK, Zhou Y, Hancock ML, Gajjar A, Rubnitz J, Ribeiro RC, et al. Bone marrow recurrence after initial intensive treatment for childhood acute lymphoblastic leukemia. Cancer. 2005;103:368-76.\u003c/li\u003e\n\u003cli\u003eKopmar NE and Cassaday RD. How I prevent and treat central nervous system disease in adults with acute lymphoblastic leukemia. Blood. 2023;141:1379-1388.\u003c/li\u003e\n\u003cli\u003eKantarjian HM, O\u0026apos;Brien S, Smith TL, Cortes J, Giles FJ, Beran M, et al. Results of treatment with hyper-CVAD, a dose-intensive regimen, in adult acute lymphocytic leukemia. J Clin Oncol. 2000;18:547-61.\u003c/li\u003e\n\u003cli\u003eReman O, Pigneux A, Huguet F, Vey N, Delannoy A, Fegueux N, et al. Central nervous system involvement in adult acute lymphoblastic leukemia at diagnosis and/or at first relapse: results from the GET-LALA group. Leuk Res. 2008;32:1741-50.\u003c/li\u003e\n\u003cli\u003eFielding AK, Richards SM, Chopra R, Lazarus HM, Litzow MR, Buck G, et al. Outcome of 609 adults after relapse of acute lymphoblastic leukemia (ALL); an MRC UKALL12/ECOG 2993 study. Blood. 2007;109:944-50.\u003c/li\u003e\n\u003cli\u003eSurapaneni UR, Cortes JE, Thomas D, O\u0026apos;Brien S, Giles FJ, Koller C, et al. Central nervous system relapse in adults with acute lymphoblastic leukemia. Cancer. 2002;94:773-9.\u003c/li\u003e\n\u003cli\u003evan der Velden VH, de Launaij D, de Vries JF, de Haas V, Sonneveld E, Voerman JS, et al. New cellular markers at diagnosis are associated with isolated central nervous system relapse in paediatric B-cell precursor acute lymphoblastic leukaemia. Br J Haematol. 2016;172:769-81.\u003c/li\u003e\n\u003cli\u003eSaygin C, Cannova J, Stock W and Muffly L. Measurable residual disease in acute lymphoblastic leukemia: methods and clinical context in adult patients. Haematologica. 2022;107:2783-2793.\u003c/li\u003e\n\u003cli\u003ePierce E, Mautner B, Mort J, Blewett A, Morris A, Keng M, et al. MRD in ALL: Optimization and Innovations. Curr Hematol Malig Rep. 2022;17:69-81.\u003c/li\u003e\n\u003cli\u003eBartram J, Patel B and Fielding AK. Monitoring MRD in ALL: Methodologies, technical aspects and optimal time points for measurement. Semin Hematol. 2020;57:142-148.\u003c/li\u003e\n\u003cli\u003ePessoa LS, Heringer M and Ferrer VP. ctDNA as a cancer biomarker: A broad overview. Crit Rev Oncol Hematol. 2020;155:103109.\u003c/li\u003e\n\u003cli\u003eCohen SA, Liu MC and Aleshin A. Practical recommendations for using ctDNA in clinical decision making. Nature. 2023;619:259-268.\u003c/li\u003e\n\u003cli\u003eKamtchum-Tatuene J and Jickling GC. Blood Biomarkers for Stroke Diagnosis and Management. Neuromolecular Med. 2019;21:344-368.\u003c/li\u003e\n\u003cli\u003eFoster JB, Koptyra MP and Bagley SJ. Recent Developments in Blood Biomarkers in Neuro-oncology. Curr Neurol Neurosci Rep. 2023;23:857-867.\u003c/li\u003e\n\u003cli\u003eSztolsztener K, Żywno H, Hodun K, Konończuk K, Muszyńska-Rosłan K and Latoch E. Apolipoproteins-New Biomarkers of Overweight and Obesity among Childhood Acute Lymphoblastic Leukemia Survivors? Int J Mol Sci. 2022;23.\u003c/li\u003e\n\u003cli\u003eSaiki R and Ogawa S. Adult Low-Hypodiploid Acute Lymphoblastic Leukemia Evolves from TP53-Mutated Clonal Hematopoiesis. Blood Cancer Discov. 2023;4:102-105.\u003c/li\u003e\n\u003cli\u003eKim R, Bergugnat H, Larcher L, Duchmann M, Passet M, Gachet S, et al. Adult Low-Hypodiploid Acute Lymphoblastic Leukemia Emerges from Preleukemic TP53-Mutant Clonal Hematopoiesis. Blood Cancer Discov. 2023;4:134-149.\u003c/li\u003e\n\u003cli\u003eBaran N, Lodi A, Dhungana Y, Herbrich S, Collins M, Sweeney S, et al. Inhibition of mitochondrial complex I reverses NOTCH1-driven metabolic reprogramming in T-cell acute lymphoblastic leukemia. Nat Commun. 2022;13:2801.\u003c/li\u003e\n\u003cli\u003eCampagnari A and Belver L. NOTCH1-Induced T-Cell Acute Lymphoblastic Leukemia In Vivo Models. Methods Mol Biol. 2024;2773:9-24.\u003c/li\u003e\n\u003cli\u003eRuan Y, Xie L and Zou A. Association of CDKN2A/B mutations, PD-1, and PD-L1 with the risk of acute lymphoblastic leukemia in children. J Cancer Res Clin Oncol. 2023;149:10841-10850.\u003c/li\u003e\n\u003cli\u003eAmpatzidou M, Papadhimitriou SI, Paisiou A, Paterakis G, Tzanoudaki M, Papadakis V, et al. The Prognostic Effect of CDKN2A/2B Gene Deletions in Pediatric Acute Lymphoblastic Leukemia (ALL): Independent Prognostic Significance in BFM-Based Protocols. Diagnostics (Basel). 2023;13.\u003c/li\u003e\n\u003cli\u003eMengxuan S, Fen Z and Runming J. Novel Treatments for Pediatric Relapsed or Refractory Acute B-Cell Lineage Lymphoblastic Leukemia: Precision Medicine Era. Front Pediatr. 2022;10:923419.\u003c/li\u003e\n\u003cli\u003eBrady SW, Roberts KG, Gu Z, Shi L, Pounds S, Pei D, et al. The genomic landscape of pediatric acute lymphoblastic leukemia. Nat Genet. 2022;54:1376-1389.\u003c/li\u003e\n\u003cli\u003eAlexander TB, Gu Z, Iacobucci I, Dickerson K, Choi JK, Xu B, et al. The genetic basis and cell of origin of mixed phenotype acute leukaemia. Nature. 2018;562:373-379.\u003c/li\u003e\n\u003cli\u003eWidman AJ, Shah M, Frydendahl A, Halmos D, Khamnei CC, \u0026Oslash;gaard N, et al. Ultrasensitive plasma-based monitoring of tumor burden using machine-learning-guided signal enrichment. Nat Med. 2024;30:1655-1666.\u003c/li\u003e\n\u003cli\u003e\u0026Aacute;lvez MB, Edfors F, von Feilitzen K, Zwahlen M, Mardinoglu A, Edqvist PH, et al. Next generation pan-cancer blood proteome profiling using proximity extension assay. Nat Commun. 2023;14:4308.\u003c/li\u003e\n\u003cli\u003eBhat M, Rabindranath M, Chara BS and Simonetto DA. Artificial intelligence, machine learning, and deep learning in liver transplantation. J Hepatol. 2023;78:1216-1233.\u003c/li\u003e\n\u003cli\u003eSilva GFS, Fagundes TP, Teixeira BC and Chiavegatto Filho ADP. Machine Learning for Hypertension Prediction: a Systematic Review. Curr Hypertens Rep. 2022;24:523-533.\u003c/li\u003e\n\u003cli\u003eOuyang Y, Li X, Zhou W, Hong W, Zheng W, Qi F, et al. Integration of machine learning XGBoost and SHAP models for NBA game outcome prediction and quantitative analysis methodology. PLoS One. 2024;19:e0307478.\u003c/li\u003e\n\u003cli\u003eLew G, Chen Y, Lu X, Rheingold SR, Whitlock JA, Devidas M, et al. Outcomes after late bone marrow and very early central nervous system relapse of childhood B-acute lymphoblastic leukemia: a report from the Children\u0026apos;s Oncology Group phase III study AALL0433. Haematologica. 2021;106:46-55.\u003c/li\u003e\n\u003cli\u003eVerma D, Kapoor S, Kumari S, Sharma D, Singh J, Benjamin M, et al. Decoding the genetic symphony: Profiling protein-coding and long noncoding RNA expression in T-acute lymphoblastic leukemia for clinical insights. PNAS Nexus. 2024;3:pgae011.\u003c/li\u003e\n\u003cli\u003eQi H, Chi L, Wang X, Jin X, Wang W and Lan J. Identification of a Seven-lncRNA-mRNA Signature for Recurrence and Prognostic Prediction in Relapsed Acute Lymphoblastic Leukemia Based on WGCNA and LASSO Analyses. Anal Cell Pathol (Amst). 2021;2021:6692022.\u003c/li\u003e\n\u003cli\u003ePan L, Liu G, Lin F, Zhong S, Xia H, Sun X, et al. Machine learning applications for prediction of relapse in childhood acute lymphoblastic leukemia. Sci Rep. 2017;7:7402.\u003c/li\u003e\n\u003cli\u003eHu Y, Zhou Y, Zhang M, Ge W, Li Y, Yang L, et al. CRISPR/Cas9-Engineered Universal CD19/CD22 Dual-Targeted CAR-T Cell Therapy for Relapsed/Refractory B-cell Acute Lymphoblastic Leukemia. Clin Cancer Res. 2021;27:2764-2772.\u003c/li\u003e\n\u003cli\u003eKantarjian HM, DeAngelo DJ, Stelljes M, Martinelli G, Liedtke M, Stock W, et al. Inotuzumab Ozogamicin versus Standard Therapy for Acute Lymphoblastic Leukemia. N Engl J Med. 2016;375:740-53.\u003c/li\u003e\n\u003cli\u003eEgan K, Kusao I, Troelstrup D, Agsalda M and Shiramizu B. Mitochondrial DNA in residual leukemia cells in cerebrospinal fluid in children with acute lymphoblastic leukemia. J Clin Med Res. 2010;2:225-9.\u003c/li\u003e\n\u003cli\u003eLiu Q, Yang L, Shi Z, Yu J, Si H, Jin Y, et al. Development and validation of a preliminary clinical support system for measuring the probability of incident 2-year (pre)frailty among community-dwelling older adults: A prospective cohort study. Int J Med Inform. 2023;177:105138.\u003c/li\u003e\n\u003cli\u003eBhagat SK, Tiyasha T, Awadh SM, Tung TM, Jawad AH and Yaseen ZM. Prediction of sediment heavy metal at the Australian Bays using newly developed hybrid artificial intelligence models. Environ Pollut. 2021;268:115663.\u003c/li\u003e\n\u003cli\u003eTarwidi D, Pudjaprasetya SR, Adytia D and Apri M. An optimized XGBoost-based machine learning method for predicting wave run-up on a sloping beach. MethodsX. 2023;10:102119.\u003c/li\u003e\n\u003cli\u003eGale RP. Progress in Transplants for Acute Lymphoblastic Leukemia. Clin Cancer Res. 2022;28:813-815.\u003c/li\u003e\n\u003cli\u003eBeyermann B, Agthe AG, Adams HP, Seeger K, Linderkamp C, Goetze G, et al. Clinical features and outcome of children with first marrow relapse of acute lymphoblastic leukemia expressing BCR-ABL fusion transcripts. BFM Relapse Study Group. Blood. 1996;87:1532-8.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"cancer-cell-international","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ccin","sideBox":"Learn more about [Cancer Cell International](http://cancerci.biomedcentral.com/)","snPcode":"12935","submissionUrl":"https://submission.nature.com/new-submission/12935/3","title":"Cancer Cell International","twitterHandle":"@OncoBioMed","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Acute Lymphoblastic Leukemia, Relapse Outcome, Monitoring Model, Serological Panel, Hub Gene","lastPublishedDoi":"10.21203/rs.3.rs-8999580/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8999580/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRelapse is the leading cause of mortality in acute lymphoblastic leukemia (ALL) patients. Minimal residual disease (MRD) detection is effective for bone marrow relapse (BMR) but less so for central nervous system relapse (CR) and combined bone marrow-central nervous system relapse (BMCR). Furthermore, there remains an unmet need for more effective therapeutic targets tailored to distinct relapse outcomes in ALL. Consequently, the identification of minimally invasive, serological biomarkers capable of monitoring diverse relapse outcomes, coupled with the exploration of hub genes underlying these outcomes, holds significant clinical relevance.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eClinical (n = 927), whole-exome mutation (n = 618), copy number variation (CNV, n = 250), and transcriptomic (n = 343) data were collected from the TARGET database for ALL patients with four outcomes: no relapse (NR), BMR, CR, and BMCR. The characteristics of each dataset were analysed. Serological panels and monitoring models for each relapse outcome were screened and constructed using optimized XGBoost and SHAP algorithms, with 50 random training iterations and top 5 indicator combinations. Hub genes were identified via the STRING database and Boruta algorithm (50 iterations). Functional impacts of representative hub genes on ALL progression were validated in cellular models.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMulti-omics analysis showed higher mortality in relapse groups than in NR, with BMCR posing the highest risk. Elevated mutation burden in BMCR was linked to poor survival. Distinct mutational profiles, CNV signatures, and transcriptomic dysregulation patterns were observed across groups, contributing to heterogeneous relapse outcomes. Optimal serological panels and monitoring models were constructed for each group: NR (LRBA/SH2B1A/IL36G/ADH1A/CEP85, AUC = 0.985), BMR (CD22/CRMP1/EPHA2/KLK14, AUC = 0.997), CR (ACVRL1/FAT3/SKAP1, AUC = 0.990), and BMCR (AGRN/CP/EFEMP1/LGALS7/ST6GAL2, AUC = 0.999), all superior to existing models (AUC \u0026lt; 0.903). Hub gene screening yielded 18 candidates for BMR (e.g., VAMP2, EFNB2), 13 for CR (e.g., MT-ATP8, CXCR3), and 3 for BMCR (e.g., TGFB1, PLCG1). Experimental validation demonstrated that knockdown of CD74 (BMR) and TGFB1 (BMCR)-highly expressed hub genes-significantly inhibited ALL cell proliferation and induced apoptosis, whereas knockdown of MT-CO1 (CR)-a low-expressed hub gene-produced the opposite effect.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study established serological panels for diverse ALL relapse outcomes and identified hub genes with therapeutic potential, providing a foundation for early relapse surveillance and targeted interventions in ALL.\u003c/p\u003e","manuscriptTitle":"Identification of monitoring panels and hub genes for heterogeneous relapse outcomes in acute lymphoblastic leukemia","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-23 09:17:10","doi":"10.21203/rs.3.rs-8999580/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewersInvited","content":"","date":"2026-04-15T01:50:02+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-04-13T12:41:48+00:00","index":"","fulltext":""},{"type":"submitted","content":"Cancer Cell International","date":"2026-04-13T07:06:33+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"cancer-cell-international","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"ccin","sideBox":"Learn more about [Cancer Cell International](http://cancerci.biomedcentral.com/)","snPcode":"12935","submissionUrl":"https://submission.nature.com/new-submission/12935/3","title":"Cancer Cell International","twitterHandle":"@OncoBioMed","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"BMC/SO AJ","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e6eaeba3-6f4d-41c2-b4c5-385f91aae4b4","owner":[],"postedDate":"April 23rd, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-23T09:17:10+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-23 09:17:10","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8999580","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8999580","identity":"rs-8999580","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00