Artificial intelligence-driven integration of multi-biofluid omics and clinical phenotype enables stratification of endometrial cancer.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by claude@2026-07, 2026-07-28

An AI platform integrating multi-biofluid omics and clinical data achieved 95.65% sensitivity for EC screening and an AUC of 0.94 for confirmation in external validation.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

Abstract

Endometrial cancer (EC) incidence is rising, yet current diagnostics lack precision and scalability. We develop an artificial intelligence (AI)-based platform integrating multi-biofluid omics and clinical data for EC stratification. Using two independent cohorts from different clinical centers (531 participants for model development, 204 for external validation), we collect 1,179 samples (plasma, cervical/uterine secretions) and the corresponding clinical data (age, ultrasound, etc.). Machine learning identifies EC-specific signatures, and the AI framework fuses omics features with clinical factors to enable multilevel risk stratification. On the external validation cohort, the platform achieves 95.65% sensitivity for minimally invasive EC screening and balanced performance with an area under the curve (AUC) value of 0.94 for EC confirmation. The model also shows potential for high-risk subtype detection. Biological plausibility is supported by identified omics signatures. A web tool is developed to support clinical translation. This platform demonstrates the potential of multi-omics and AI in precision oncology.
Full text 90,207 characters · extracted from pmc-nxml · 8 sections · click to expand

Author

Conceptualization, D.L., X.C., and L.Q.; methodology, D.L. and P.W.; investigation, D.L., P.W., J.Z., Y.Y., W.S., J.Y., D.Z., S.Y., Y.H., X.Z., R.F., P.F., and C.B.; writing – original draft, D.L. and P.W.; writing – review and editing, X.C. and L.Q.; funding acquisition, X.C. and L.Q.; resources, P.W., W.S., X.C., and L.Q.; supervision, X.C. and L.Q.

Results

Between November 2022 and January 2024, we enrolled a cohort of 531 patients associated with endometrial disease, including 273 cancer and 258 benign. The inclusion criteria for participants are detailed in the STAR Methods section. Clinical phenotype data were collected, including patient demographics, medical history, ultrasound imaging, and tumor biomarker test results (HE4 and CA125 levels) ( Data S1 ). Concurrently, multi-source body fluids were collected, including in situ uterine cavity secretions, cervical secretions, and plasma, providing a comprehensive molecular characterization of the patients’ physiological states. Prior to ML, we implemented a stratified cohort design comprising a modeling cohort (M, n = 436) and a test cohort (TS, n = 95), Figure 1 , left. The modeling cohort ( n = 436) was partitioned into three sub-cohorts: M1, M2, and M3. The M1 cohort (148 EC, 166 control) included patients sampled with cervical secretions or plasma, serving as the basis for minimally invasive screening of EC. M2 (219 EC, 217 control) comprised patients sampled with cervical secretions, plasma, or uterine cavity secretions, enabling the development of a comprehensive EC diagnostic approach. M3 (77 EC) comprised patients with molecular subtype information who were sampled with uterine cavity secretions, providing data for high-risk EC subtype detection. Within the modeling cohort, 80% of samples were randomly assigned for model training, while 20% were reserved for internal validation, using a 5-fold cross-validation approach. The model’s generalization capability was further evaluated using the hold-out test cohorts: TS1 (54 cancer, 41 control), which included patients sampled with cervical secretions and plasma, and TS2 (30 cancer, 15 control), which included patients sampled with cervical secretions, plasma, and uterine cavity secretions. Figure 1 Study design of sampling, modeling, and development of the 2M-EC platform for task-specific EC risk stratification Left: Clinical cohort design, multi-biofluid multi-omics profiling, and clinical data collection. Right: Development of an AI-driven clinical decision support tool for EC risk stratification by engineering heterogeneous multi-biofluid multi-omics and clinical data. Created partially in BioRender. Qiao, L. (2026) https://BioRender.com/5g7bdob and https://BioRender.com/8tn28lh . Study design of sampling, modeling, and development of the 2M-EC platform for task-specific EC risk stratification Left: Clinical cohort design, multi-biofluid multi-omics profiling, and clinical data collection. Right: Development of an AI-driven clinical decision support tool for EC risk stratification by engineering heterogeneous multi-biofluid multi-omics and clinical data. Created partially in BioRender. Qiao, L. (2026) https://BioRender.com/5g7bdob and https://BioRender.com/8tn28lh . The metabolic features of uterine cavity secretions and cervical secretions, as well as the metabolic and peptidomic features of plasma, were acquired by MALDI-TOF MS to create multi-omics profiles. Instrumental stability was assessed using a pooled mixture of plasma obtained from six patients with EC as the quality control (QC) sample. The sample was repeatedly analyzed over 1 month to check the stability of the MALDI-TOF MS measurements. The Pearson correlation coefficients between the MALDI-TOF MS spectra of the QC samples all exceeded 0.90, confirming consistent instrument performance throughout the spectral acquisition process ( Figure S1 ). The obtained MS profiles were processed via binning. 29 MS features for EC risk assessment were selected by ML from the MS bins and then converted into feature vectors, which were used to develop ML models for EC risk assessment based on cervical secretions, plasma, or uterine cavity secretions (MS models). Based on clinical expertise and data-driven analysis, we identified key clinical factors, including age, HE4 levels, menopausal status, hormone replacement therapy (HRT) history, and gynecologic ultrasound findings, which were also converted into factor vectors and used to develop the clinical ML model for EC risk assessment (CL model). Then, we deployed the MS and CL models via both feature-level and decision-level fusion strategies ( Figure 1 , right). We first fused the clinical vector and the MS vector (cervical secretions, plasma, or uterine cavity secretions) from the same patient to develop the combined ML model for EC risk assessment (Fusion model). Building upon these fusion models, we implemented a decision-level fusion strategy for multiple task-specific modes: the level-1 mode for EC screening based on cervical secretion and plasma samples, the level-2 mode for EC confirmation based on cervical secretion, plasma, and uterine cavity secretion samples, and the level-3 mode for high-risk subtype detection based on uterine cavity secretion samples. The integration of these multiple task-specific modes formed the integrative 2M-EC platform. A dedicated web-based tool has been developed to implement the 2M-EC platform. Benchmarking against the Risk of Ovarian Malignancy Algorithm (ROMA), which incorporates menopausal status, CA-125, and HE4 levels, along with multi-center external validation, was sufficiently evaluated. Multivariate analysis was performed to identify clinical factors most relevant to EC progression ( Figure 2 ). Continuous numerical data (age, endometrial thickness, HE4 level, and CA125 level) were log-transformed and normalized as Z scores, followed by unsupervised clustering analysis among patients with EC and patients in the control group. The clustering heatmap demonstrated a distinct subset of control group patients on the left side, where plasma HE4 expression was notably reduced, highlighting its role as an EC discriminative biomarker ( Figure 2 A). Compared to HE4, CA125 showed lower significance in distinguishing EC from the control. In menopausal patients, both markers exhibited significant differences between the EC and control groups, while in non-menopausal patients, CA125 showed less pronounced differences ( Figure S2 ). Figure 2 Analysis of clinical data (A) Continuous numerical clinical data, including age, endometrial thickness, HE4 level, and CA125 level, were log-transformed and normalized, followed by clustering analysis among patients with EC and controls. Categorical factors were represented by one-hot encoding in different colors. MS data collection of plasma (P), cervical secretions (C), and uterine secretions (U) were also displayed by one-hot encoding. (B) From top to bottom, each row encodes a patient, with horizontal yellow bars representing age. The colors of “menopause” and “HRT” indicate “yes” and “no,” consistent with Figure 2 A. The pie charts represent the proportion of “yes” or “no” within the entire group of patients, with color labels consistent with Figure 2 A. See also Figure S2 . Analysis of clinical data (A) Continuous numerical clinical data, including age, endometrial thickness, HE4 level, and CA125 level, were log-transformed and normalized, followed by clustering analysis among patients with EC and controls. Categorical factors were represented by one-hot encoding in different colors. MS data collection of plasma (P), cervical secretions (C), and uterine secretions (U) were also displayed by one-hot encoding. (B) From top to bottom, each row encodes a patient, with horizontal yellow bars representing age. The colors of “menopause” and “HRT” indicate “yes” and “no,” consistent with Figure 2 A. The pie charts represent the proportion of “yes” or “no” within the entire group of patients, with color labels consistent with Figure 2 A. See also Figure S2 . Radiomics-based ultrasound examination also plays a critical role in EC early screening and staging, as it assesses endometrial thickness, morphological abnormalities, and tumor invasion depth. 36 Figure 2 A shows a positive correlation between cancer status and uterine cavity occupation (UCO1) and between cancer status and UCO lesions with rich blood flow (UCO2), especially for the distinct control subset clustering on the left side, emphasizing the statistical evidence for the inclusion of ultrasound imaging findings in EC modeling. Clinically, patient baseline factors, such as age, menopausal status, and HRT, are also considered risk factors relevant to EC. As the heatmap represents, many patients with cancer were aligned with menopausal status ( Figure 2 A). A higher EC incidence was observed among menopausal women ( Figure 2 B). Age-stratified analysis found that more EC cases occurred than controls among older individuals ( Figure 2 B). This confirmed that the EC group had a higher proportion of menopausal patients with increasing age, consistent with well-established clinical epidemiological patterns of EC. Besides, HRT experience constitutes an independent risk factor for EC incidence. 37 Younger patients with EC indicated higher HRT adoption rates than those in the control group, showing that younger women receiving hormone therapy have a higher risk of EC ( Figure 2 B). As for patient health state, diabetes and hypertension showed no significant distribution trends in the heatmap and were excluded from modeling. After distilling clinical factors, age, HRT experience, menopausal status, HE4 level, and ultrasound findings (endometrial heterogeneity, endometrial thickness, uterine cavity fluid (UCF), UCO1, and UCO2) were included as key clinical factors for subsequent modeling. The M2 (219 cancer, 217 control) subset was used for model training using 5-fold cross validation, while the TS1 cohort (54 cancer, 41 control) was used for model evaluation. The best performance of the clinical model on TS1 showed an accuracy of 78.57%, sensitivity of 79.59%, and specificity of 77.55% using the eXtreme Gradient Boosting (XGBoost) algorithm ( Figure S3 ). From the 531 patients, we acquired 1,160 multi-omics MS profiles across multiple biofluids ( Figure S4 , Data S2 ), providing a comprehensive molecular characterization of the cohort’s physiological states. Plasma achieved the highest sample coverage, while cervical and uterine secretions exhibited varying degrees of missing profiles, primarily due to clinical sampling practices. Uterine cavity secretions, in particular, are more challenging to collect compared to plasma and cervical secretions. This highlights the need for developing multi-model approaches that can accommodate diverse biological fluid sources, better aligning with various clinical scenarios. With the acquired multi-omics MS profiles, we explored various analyte-algorithm combinations ( Figure 3 A) based on a reported framework by Osipov et al. 38 We began with multi-omics dataset preparation, followed by conversion of the original MALDI-TOF mass spectrum into equal-width interval bins. Unlike conventional peak-picking methods, which require intensive peak alignment calculations, the binning method 29 sums the intensity of all data points within each bin to obtain MS features. This method is better suited for processing large datasets. After spectra processing, the combination of 17 ML algorithms and multi-source analytes (plasma metabolites, plasma peptides, uterine secretion metabolites, cervical secretion metabolites, as well as clinical factors) was systematically compared. Accuracy and F1-score were used for analyte-algorithm combination evaluation ( Figures 3 B and 3C). According to the selected algorithm, the MS feature panel was further optimized by Monte Carlo simulation. Accuracy, sensitivity, and specificity were used for evaluation. Based on grid searching ( Figures 3 B and 3C), we found that decision tree-based ensemble methods outperformed other ML algorithms. XGBoost (XGB), using plasma features and clinical indices, showed the highest metrics (acc = 0.70, F1 = 0.68). Based on these results, we implemented four decision tree-based ensemble algorithms, random forest (RF), XGB, light gradient boosting machine (LightGBM), and categorical boosting (CatBoost) to train subsequent predictive models. Figure 3 Analyte-algorithm combination selection using multi-omics data from multiple biofluid sources coupled with clinical data (A) Pipeline of the analyte-algorithm combination selection procedure, including multi-omics MS data preparation, MS data processing using binning methods, analyte-algorithm selection by grid searching, feature panel optimization by Monte Carlo simulation, and model construction. Benchmarking results (accuracy and F1-score) of analyte-algorithm combination, colored by (B) analyte combinations and (C) ML algorithms. Abbreviations are as follows: PP, plasma peptides; PM, plasma metabolites; UM, uterine metabolites; CM, cervical metabolites; CI, clinical indices. See also Figure S3 . Analyte-algorithm combination selection using multi-omics data from multiple biofluid sources coupled with clinical data (A) Pipeline of the analyte-algorithm combination selection procedure, including multi-omics MS data preparation, MS data processing using binning methods, analyte-algorithm selection by grid searching, feature panel optimization by Monte Carlo simulation, and model construction. Benchmarking results (accuracy and F1-score) of analyte-algorithm combination, colored by (B) analyte combinations and (C) ML algorithms. Abbreviations are as follows: PP, plasma peptides; PM, plasma metabolites; UM, uterine metabolites; CM, cervical metabolites; CI, clinical indices. See also Figure S3 . The 17 ML algorithms employed in this study were carefully selected and organized to reflect a progressive increase both in terms of feature selection strategy and model architecture. The framework begins with robust baseline models, such as RF_Regression and SVM_Model, which provide stable performance benchmarks due to their inherent resistance to overfitting. Building on this foundation, models with embedded feature selection, including L1_Norm_SVM_Model, L1_Norm_RF_Model, and L1_Norm_LR_Model, were incorporated, utilizing L1 regularization to perform intrinsic feature selection during training. Further advancing the approach, wrapper-based models such as RFE_LR_Model and RFE_RG_Model were applied, which iteratively refine feature subsets by eliminating the least important features. Finally, the framework culminates in two-stage meta-ensemble models (LR_Logic, LR_MLP, LR_KNei, LR_SVM_Lin, LR_SVM_RBF, LR_SVM_Sig, LR_SVM_Poly, LR_GBDT, LR_XGB, LR_RF), where feature selection is explicitly decoupled from classification. In these models, a meta-learner, logistic regression, integrates predictions from diverse base classifiers to capture both linear and nonlinear relationships. This progression systematically encapsulates the evolution of feature engineering from foundational stability to advanced comprehensiveness. Moreover, the selected algorithms leverage complementary feature selection mechanisms (weight-based, frequency-based, and ranking-based) to deliver multifaceted insights into feature importance, while collectively spanning core ML paradigms, including tree-based methods, support vector machines, linear models, ensemble techniques, and basic neural network variants. Newer architectures like deep learning were not chosen due to the insufficient training data size. After selecting analytes and algorithms, we first developed models for rapid EC screening using the M1 and TS1 cohorts. Among the 531 patients, 409 were collected with plasma or cervical secretions but without uterine cavity secretions, including 361 (176 cancer and 185 control) sampled with plasma and 192 (101 cancer and 91 control) sampled with cervical secretions ( Figure 4 A). Venn analysis shows that a subset of 144 patients provided both plasma and cervical secretion samples ( Figure 4 B). From this subset, 95 were randomly selected to form the TS1 cohort for model testing. Excepting these 95, 314 patients comprised the M1 cohort for model training. Using the M1 cohort, orthogonal partial least squares-discriminant analysis (OPLS-DA) revealed greater discrimination of plasma metabolites and peptides profiles between patients with EC and controls compared to cervical profiles ( Figures 4 C and 4D). Monte Carlo cross-validation with 1,000 iterations using RF identified 13 cervical MS features and 49 plasma MS features ( Figures 4 E and 4F, Data S3 ), which were prioritized as omics feature vectors for subsequent modeling. Using a 5-fold cross-validation on M1, the cervical MS models based on the four decision tree-based ensemble algorithms (RF, XGB, LightGBM, and CatBoost) showed balanced but moderate performance across all evaluation metrics. In contrast, the plasma MS model based on CatBoost exhibited good predictive capability, achieving 79.31% accuracy, 75.86% specificity, and 82.76% sensitivity ( Figure S5 ). Figure 4 Minimally invasive EC screening (A) Clinical cohort for EC screening. (B) Venn illustration of cervical secretion and plasma MS data and TS1 split ( n = 95) for test. (C) OPLS-DA analysis of cervical MS data distinguishing patients with EC from control patients. (D) OPLS-DA analysis of plasma MS data distinguishing patients with EC from control patients. (E and F) Identification of (E) 13 key cervical MS features and (F) 49 key plasma MS features based on Monte Carlo cross-validation with RF. (G) AUC-ROC curves of the cervical and plasma Fusion models on the TS1 cohort. (H) Integration of the EC-Screen model using “OR” logic and its predictive performance on the TS1 cohort ( n = 95). (I) Benchmarking of the EC-Screen model against ROMA with menopausal stratification. (J–M) Waterfall plots visualizing the contribution of features during EC-Screen decision-making for patient 571, patient 611, patient 722, and patient 527. Abbreviations are as follows: PM, plasma metabolite; PP, plasma peptide; CM, cervical metabolite. See also Figures S1 , S4 , S5 , and S6 . Minimally invasive EC screening (A) Clinical cohort for EC screening. (B) Venn illustration of cervical secretion and plasma MS data and TS1 split ( n = 95) for test. (C) OPLS-DA analysis of cervical MS data distinguishing patients with EC from control patients. (D) OPLS-DA analysis of plasma MS data distinguishing patients with EC from control patients. (E and F) Identification of (E) 13 key cervical MS features and (F) 49 key plasma MS features based on Monte Carlo cross-validation with RF. (G) AUC-ROC curves of the cervical and plasma Fusion models on the TS1 cohort. (H) Integration of the EC-Screen model using “OR” logic and its predictive performance on the TS1 cohort ( n = 95). (I) Benchmarking of the EC-Screen model against ROMA with menopausal stratification. (J–M) Waterfall plots visualizing the contribution of features during EC-Screen decision-making for patient 571, patient 611, patient 722, and patient 527. Abbreviations are as follows: PM, plasma metabolite; PP, plasma peptide; CM, cervical metabolite. See also Figures S1 , S4 , S5 , and S6 . Although individual MS models showed predictive potential, their performance as standalone tools for cancer screening lacks sufficient sensitivity. To address this, we fused the clinical vectors (age, HRT, menopausal status, HE4, and ultrasound findings) with the MS features (13 cervical MS features or 49 plasma MS features) to develop Fusion models (MS plus clinical vectors) using the four ML algorithms (RF, XGB, LightGBM, and CatBoost), Figure S6 . XGB exhibited the best performance when evaluated across sensitivity, accuracy, and specificity using 5-fold cross-validation on M1 and was thereby selected for the development of the subsequent Fusion models. The area under the receiver operating characteristic curve (AUC-ROC) values of the cervical and plasma Fusion models were 0.85 and 0.78 in the TS1 cohort, respectively ( Figure 4 G). The cervical and plasma Fusion models were then combined into the EC-Screen model using an “OR” logic, meaning that a positive result from either model would classify the result as cancer-positive ( Figure 4 H). In the TS1 cohort ( n = 95), the EC-Screen model demonstrated 98.15% sensitivity and 82.11% accuracy. This high sensitivity guarantees the EC-Screen model as a practical tool for EC risk screening using easily accessible clinical and MS profiling data. To benchmark the EC-Screen model, we compared its predictive performance with ROMA using the same test set (54 cancer cases vs. 30 controls), after excluding samples with missing HE4 and CA125 values in TS1. ROMA is an FDA-approved tool for ovarian cancer risk stratification based on HE4, CA125, and menopausal status (details in the STAR Methods section), which has also demonstrated potential in EC risk stratification. 39 , 40 , 41 Compared with ROMA, the EC-Screen model achieved higher sensitivity (98.15% vs. 59.26%), specificity (66.67% vs. 63.33%), and accuracy (86.90% vs. 60.71%) ( Figure 4 I). When stratified by menopausal status, the EC-Screen model performed much better in non-menopausal patients than ROMA, with 96.15% sensitivity, 86.67% accuracy, and 73.68% specificity, whereas all values of ROMA were below 53%. In menopausal patients, the EC-Screen model exhibited better sensitivity and accuracy than ROMA but lower specificity. These results suggested that the EC-Screen model is particularly suited for EC screening in non-menopausal women, potentially addressing the current limitations of EC screening in premenopausal women. Waterfall plots visualized the decision-making process of the EC-Screen model using Shapley additive explanations (SHAP) values ( Figures 4 J–4M). Modeling features were ranked by their contribution, with color-coded arrows indicating the direction and magnitude of each feature’s contribution. Rightward red arrows represent positive contributions, favoring EC outcomes, while leftward blue arrows represent negative contributions, supporting control outcomes. In the plasma Fusion model, EC sample 571 showed consistent decision-making, with HE4 providing the strongest correct contribution ( Figure 4 J). For control sample 611, although the final prediction was correct, the primary HE4 contribution was incorrect, with other molecular features leading to the correct classification ( Figure 4 K). In EC sample 722 ( Figure 4 L), cervical metabolite CM628 and elevated HE4 levels were the main drivers of the correct EC prediction. Conversely, in control sample 527 ( Figure 4 M), cervical metabolites CM628 and CM4160 predominated in the negative classification. Age and HE4 levels supported the control outcome, while endometrial thickness and endometrial heterogeneity contributed weakly to the EC outcome. These results illustrate the importance of correlating clinical information with MS profiles for EC diagnosis. To identify the MS features for EC screening, we selected biofluid samples (sample 577, 677, 912, 939, 975 for cervical secretions and sample 31, 53, 490, 519, 638 for plasma) that showed strong signals for the key MALDI-TOF MS features to perform LC-MS/MS analysis. We employed two proteomic preprocessing strategies to identify plasma proteins: SDS-PAGE-based separation (10–25 kDa) with typical bottom-up proteomics and direct 10 kDa filtration of plasma peptides followed by top-down analysis ( Figure 5 A). The molecular weight range selection was based on the m/z values of the features from MALDI-TOF MS. By matching the MALDI-TOF MS features with proteomic identification results ( Figure 5 B, Data S4 and S5 ), we identified 14 proteins corresponding to the 18 plasma MALDI-TOF MS features, including fibrinogen alpha chain (FGA), serum amyloid P component (APCS), serum amyloid A-4 protein (SAA4), retinoic acid (RA) receptor responder protein 2 (RARRES2), ubiquitin-conjugating enzyme E2 variant 2 (UBE2V2), and leucine-rich alpha-2-glycoprotein (LRG) ( Table S1 ). FGA is a component of the extracellular matrix protein fibrinogen. Our findings are consistent with previous studies reporting elevated plasma fibrinogen levels, especially FGA, in various tumors. 42 Elevated phosphorylated FGA has also been reported in ovarian cancer. 43 Moreover, Chen et al. identified FGA as both a coding and non-coding driver of hepatocellular carcinoma progression and metastasis. 44 We hypothesize that EC cells may synthesize FGA and release it into peripheral blood. Recent clinical studies indicate that high plasma fibrinogen levels are associated with poor disease-free survival in several cancers. 45 Although APCS is not a widely studied biomarker for EC, another related protein, serum amyloid A (SAA), shows potential as an EC biomarker. Omer et al. reported significantly higher serum SAA levels in patients with EC compared with controls. 46 RA has been shown to induce differentiation, inhibit proliferation, and promote apoptosis in endometrial carcinoma. Studies indicate that RA can induce differentiation of endometrial adenocarcinoma cells, 47 , 48 a process involving phosphatidylinositol 3-kinase (PI3K) and epidermal growth factor receptor (EGFR) signaling pathways. 49 , 50 Furthermore, fenofibrate, a peroxisome proliferator-activated receptor α agonist, can inhibit proliferation and induce apoptosis, and RA enhances this effect. 51 Nevertheless, the exact molecular mechanisms of RA-induced apoptosis and growth inhibition in EC cells remain to be explored. Research indicates that the UBE2 family plays an important role in EC prognosis. 52 , 53 Specifically, UBE2C expression serves as an independent predictor of unfavorable overall survival. 54 LRG has also been reported as a useful diagnostic biomarker for endometriosis. 52 Figure 5 Molecular feature identification and biological relevance of the EC screening models (A) Sample preparation for two proteomics strategies: 10–25 kDa molecular weight-based extraction by SDS-PAGE and plasma free peptide purification via membrane filtration. (B–D) (B) Aligning of the MALDI-TOF MS peak m/z with the proteomic identification results for feature identification. Metabolic pathway enrichment analysis of (C) cervical significant metabolites and (D) plasma significant metabolites. Created partially in BioRender. Qiao, L. (2026) https://bioRender.com/huvinwy . See also Table S1 . Molecular feature identification and biological relevance of the EC screening models (A) Sample preparation for two proteomics strategies: 10–25 kDa molecular weight-based extraction by SDS-PAGE and plasma free peptide purification via membrane filtration. (B–D) (B) Aligning of the MALDI-TOF MS peak m/z with the proteomic identification results for feature identification. Metabolic pathway enrichment analysis of (C) cervical significant metabolites and (D) plasma significant metabolites. Created partially in BioRender. Qiao, L. (2026) https://bioRender.com/huvinwy . See also Table S1 . We also analyzed the cervical and plasma metabolite features. By matching the MALDI-TOF MS features with the LC-MS/MS-based metabolomics results ( Data S6 ), 44 cervical metabolites and 114 plasma metabolites were identified ( Data S7 ). Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analysis of the identified cervical and plasma metabolites highlighted functional pathways with potential biological and diagnostic relevance to EC ( Figures 5 C and 5D, Data S8 ). The metabolic pathways enriched in the cervical metabolites were primarily associated with fatty acid oxidation. Lipid metabolism is well established in cancer, as cancer cells require lipids for energy, survival, and growth. 55 Lipid upregulation is also linked to cancer progression and metastasis. 56 In contrast, amino acid metabolism was enriched in the plasma metabolites, differing from the findings in the cervical metabolites. It is hypothesized that cancer cells utilize amino acids for proliferation and to maintain redox homeostasis. 57 Altered amino acid pathways included the metabolism of glycine, serine, threonine, arginine, proline, histidine, and tryptophan. Proline helps protect cells from oxidative stress, which cancer cells can exploit to promote growth. 58 Tryptophan metabolism is also known to be elevated in cancer, contributing to immune suppression. 59 The metabolomic analysis of plasma samples aligned with our previous findings in EC tissue samples, 60 reinforcing the feasibility of non-invasive EC screening using plasma metabolites. The EC-Screen model exhibits high sensitivity for EC screening but is limited by specificity. To improve detection accuracy, we further established a uterine Fusion model by fusing clinical factors with uterine secretions metabolomic features. We collected uterine secretions from 246 patients (147 cancer, 99 control) for MALDI-TOF MS analysis of metabolic profiles ( Figure 6 A). As in situ samples from the endometrial cavity, uterine secretions provide rich molecular information from cancer cells. OPLS-DA analysis of uterine secretion MS data effectively distinguished patients with EC from controls ( Figure S7 A). The 246 patients were divided into a modeling set of 201 patients and a test set of 45 patients (TS2) ( Figure 6 B). The 45 patients of the TS2 cohort were sampled with plasma, cervical secretions, and uterine secretions. The 201 patients of the modeling cohort were combined with the M1 cohort to form the M2 cohort ( n = 436), where plasma, cervical secretions, or uterine secretions were sampled. Through Monte Carlo cross-validation with RF on the M2 cohort, a panel of 23 uterine secretion MS features was identified to differentiate EC from controls ( Figure 6 C and Data S3 ). The 23 MS features were integrated with the clinical factors (age, HRT, menopausal status, HE4, and ultrasound findings) for ML training, establishing the uterine Fusion model. The uterine Fusion model consistently demonstrated predictive performance across all metrics in 5-fold cross-validation on the M2 cohort ( Figure S6 C). On the TS2 cohort ( n = 45), the uterine Fusion model achieved an AUC-ROC of 0.88 ( Figure S7 B). Next, we constructed an integrated EC-Diagnosis model by integrating the cervical, plasma, and uterine Fusion models using majority-voting logic ( Figure 6 D). Validated on the TS2 cohort ( n = 45), which shared all body fluid data, the EC-Diagnosis model achieved 93.3% accuracy, 90.0% sensitivity, and 100% specificity ( Figure 6 E). Compared to the EC-Screen model, which emphasizes sensitivity for minimally invasive EC screening, the EC-Diagnosis model provides enhanced precision by incorporating broader factors for more accurate EC confirmation. Figure 6 EC diagnosis and high-risk subtype detection using uterine secretions (A) Clinical cohort for EC diagnosis. (B) Venn illustration of three-body fluid MS data and TS2 cohort ( n = 45) for testing. (C) Identification of 23 key uterine MS features for EC diagnosis based on Monte Carlo cross-validation with RF. (D) Integration of uterine, cervical, and plasma Fusion models into the EC-Diagnosis model using majority-voting logic. (E) EC-Diagnosis model prediction performance on the TS2 cohort ( n = 45) and benchmarking against the uterine Fusion model. (F) Enriched pathways of identified uterine metabolic features for EC detection. (G) Clinical cohort for EC high-risk subtype detection. (H) OPLS-DA analysis of uterine MS data distinguishing NSMP from other EC subtype patients. (I) OPLS-DA analysis of uterine MS data distinguishing p53abn from other EC subtype patients. (J and K) EC high-risk model prediction performance using 5-fold cross-validation. (L) Enriched pathways of identified p53abn uterine metabolic features. See also Figures S1 and S4–S8 . EC diagnosis and high-risk subtype detection using uterine secretions (A) Clinical cohort for EC diagnosis. (B) Venn illustration of three-body fluid MS data and TS2 cohort ( n = 45) for testing. (C) Identification of 23 key uterine MS features for EC diagnosis based on Monte Carlo cross-validation with RF. (D) Integration of uterine, cervical, and plasma Fusion models into the EC-Diagnosis model using majority-voting logic. (E) EC-Diagnosis model prediction performance on the TS2 cohort ( n = 45) and benchmarking against the uterine Fusion model. (F) Enriched pathways of identified uterine metabolic features for EC detection. (G) Clinical cohort for EC high-risk subtype detection. (H) OPLS-DA analysis of uterine MS data distinguishing NSMP from other EC subtype patients. (I) OPLS-DA analysis of uterine MS data distinguishing p53abn from other EC subtype patients. (J and K) EC high-risk model prediction performance using 5-fold cross-validation. (L) Enriched pathways of identified p53abn uterine metabolic features. See also Figures S1 and S4–S8 . Employing a strategy analogous to that used for identifying key cervical and plasma metabolic features, we identified key uterine metabolic features, which were annotated to 93 metabolites ( Data S7 ). KEGG pathway enrichment analysis of these metabolites ( Figure 6 F and Data S8 ) revealed alterations in galactose, fructose/mannose, and amino acid metabolism (methionine, tyrosine, phenylalanine), reflecting metabolic reprogramming in EC for energy production and proliferation. Our previous metabolomic profiling of uterine cavity secretions from patients with EC also identified changes in key metabolic pathways such as purine, tyrosine, and phenylalanine metabolism. 60 Additionally, research indicates that high glucose promotes glycolysis and cholesterol synthesis while suppressing the autophagy-lysosomal pathway in EC, a process regulated by estrogen-related receptor α. This mechanism underlies the glucose and cholesterol metabolic disturbances in patients with EC and diabetes mellitus. 52 Hormone-related pathways (androgen/estrogen metabolism, steroidogenesis) and PI3K/AKT-related lipids metabolism (phosphatidylinositol phosphate metabolism) further highlight the hormone-driven pathogenesis of EC. Androgen and estrogen metabolism play crucial roles in the development and progression of EC. In premenopausal women, hyperandrogenemia leads to progesterone deficiency; while in postmenopausal women, aromatization promotes increased estrogen synthesis. 61 Androgens play a complex role in EC, regulated by both their metabolism and receptor expression. 62 Postmenopausal adrenal androgens primarily serve as precursors to estrogen synthesis, leading to an increased EC risk. 63 The PI3K/protein kinase B (AKT)/mechanistic target of rapamycin (mTOR) pathway is a central driver in endometrial carcinogenesis and a promising therapeutic target. 64 , 65 Alterations in this pathway are common across histological subtypes, with specific mutations in phosphatidylinositol-4,5-bisphosphate 3-kinase-catalytic subunit alpha (PIK3CA) and loss of phosphatase and tensin homolog (PTEN) serving as predictive biomarkers for targeted therapies. 65 , 66 Moreover, the co-occurrence of PI3K-AKT pathway dysregulation and p53 alterations defines a molecular subgroup with significantly worse prognosis, integrating this pathway into the molecular classification of EC. 66 EC is classified into four molecular subtypes according to the TCGA system: POLE-mutant, MMRd, NSMP, and p53abn. 8 Among these, the NSMP and p53abn subtypes present a high risk for patients with EC due to their aggressive tumor behavior and poor prognosis. 2 Of the 147 patients with EC who had uterine secretions samples, 77 had known molecular subtype information ( Figure 6 G). We established two classifiers for rapid detection of NSMP or p53abn EC subtypes using uterine cavity secretions. MS profiles of the uterine secretions effectively discriminated p53abn and non-p53abn subtypes, as well as NSMP and non-NSMP subtypes by OPLS-DA ( Figures 6 H and 6I). Given the limited number of samples with molecular subtype information, we employed cross-validation to assess the potential of EC subtyping using uterine secretion metabolic profiling by MALDI-TOF MS. Through Monte Carlo simulation by RF, 10 key MS features were identified for p53abn and 52 for NSMP ( Figure S8 , Data S3 ). Based on these features, high-risk subtype detection models for p53abn and NSMP were developed using the XGB algorithm. The p53abn model achieved 77.8% accuracy, and the NSMP model reached 91.7% accuracy in 5-fold cross-validation ( Figures 6 J and 6K). The two uterine metabolic signature-based classifiers enabling rapid detection of NSMP and p53abn EC subtypes together form the high-risk-subtyping model, informing treatment decisions efficiently. Metabolite annotation of the uterine metabolic features specific to the p53abn subtype identified 14 metabolites ( Data S7 ). The KEGG-enriched pathways were primarily involved in cell cycle regulation, DNA repair, and metabolic reprogramming, particularly one-carbon and lipid metabolism ( Figure 6 L and Data S8 ). In EC, p53 is a master regulator. Wild-type p53 regulates both the G1/S and G2/M cell cycle checkpoints, while its mutation defines the aggressive p53abn molecular subtype, which accounts for the majority of EC-related deaths despite representing only about 15% of cases. 67 In the context of DNA repair, p53abn EC often exhibits high genomic instability and homologous recombination deficiency, creating therapeutic opportunities analogous to those in serous ovarian cancer, including potential sensitivity to poly ADP-ribose polymerase (PARP) inhibitors. 67 , 68 Metabolically, p53 can influence critical pathways in one-carbon and lipid metabolism, which support cancer cell migration, metastasis, and survival, adding another dimension to its complex role in EC pathogenesis. 69 , 70 To further demonstrate the clinical potential of the EC risk stratification models, we enrolled an external cohort of 204 participants (80 EC and 124 control) from the Shanghai Tenth People’s Hospital of Tongji University, a different clinical center compared to the original cohort used for model development. The required sample size for external validation was assessed based on the method published by Riley et al., 71 which suggested at least 196 participants, including 79 patients with EC (see details in the STAR Methods section). From the 204 participants, we collected 204 plasma samples and 88 paired cervical/uterine secretion samples (with 23 EC cases) ( Figure 7 A). Heatmap visualization of the clinical data ( Figure 7 B and Data S9 ) of the external validation cohort revealed that among patients with EC, the prevalence of postmenopausal status, EH, UCO1, UCO2, and UCF was higher. Age and HE4 levels were also higher in patients with EC. These clinical features were consistent with the original cohort, further emphasizing the correlation between these clinical indicators and the incidence of EC. Figure 7 External validation cohort for EC risk stratification (A) External validation cohort design ( n = 204). (B) Heatmap visualization of clinical characteristics. (C) Accuracy, sensitivity, and specificity of EC detection by the three biofluid Fusion models. (D) AUC of prediction results for the three biofluid Fusion models. (E–G) (E) EC-Screen and EC-Diagnosis model performance on the external validation cohort. AUC-ROC curves of (F) the EC-Screen model and (G) the EC-Diagnosis model on the external validation cohort. External validation cohort for EC risk stratification (A) External validation cohort design ( n = 204). (B) Heatmap visualization of clinical characteristics. (C) Accuracy, sensitivity, and specificity of EC detection by the three biofluid Fusion models. (D) AUC of prediction results for the three biofluid Fusion models. (E–G) (E) EC-Screen and EC-Diagnosis model performance on the external validation cohort. AUC-ROC curves of (F) the EC-Screen model and (G) the EC-Diagnosis model on the external validation cohort. Following the same standardized workflow established on the original cohort, we performed sample pretreatment, MALDI-TOF data acquisition, and processing. Using the collected external validation dataset, we first examined the fusion model integrating multi-omics data from biofluids and clinical features, including age, HRT, menopausal status, HE4, and ultrasound findings. The pre-trained models and fitted scalers from the original cohort were adopted to generate prediction results on the external validation cohort. The results showed that the prediction accuracy based on plasma, cervical secretion, and uterine secretion samples was 76.47%, 81.82%, and 76.14%, respectively ( Figure 7 C and Data S10 ), with AUC values of 0.81, 0.82, and 0.89, respectively ( Figure 7 D), similar to those obtained in the original cohort, where the AUC values for plasma, cervical secretion, and uterine secretion were 0.78, 0.85, and 0.88, respectively ( Figures 4 G and S7 B). For the EC-Screen model based on paired plasma and cervical secretion cases, the sensitivity was 95.65% and specificity was 61.54% ( Figure 7 E), with an AUC value of 0.83 ( Figure 7 F); while those in the original cohort were 98.15% (sensitivity) and 60.98% (specificity) ( Figure 4 H). The similar sensitivity and specificity confirm that the EC-Screen model exhibits stable performance across multi-center cohorts. For the EC-Diagnosis model using paired plasma, cervical, and uterine secretion cases, the sensitivity was 91.30% and specificity was 83.08% ( Figure 7 E), with an AUC value of 0.94 ( Figure 7 G); while those in the original cohort were 90% (sensitivity) and 100% (specificity) ( Figure 6 E), also demonstrating robust performance of the EC-Diagnosis model across multi-center cohorts. Based on the models, we have developed the 2M-EC platform by integrating clinical risk factors, clinical laboratory test results, and omics-derived molecular signatures from multiple biofluids analyzed by MALDI-TOF MS. The EC-Screen model of the 2M-EC platform, utilizing clinical risk factors, laboratory test results, as well as plasma and cervical secretion samples, prioritizes high sensitivity for initial EC screening, minimizing missed diagnoses in at-risk populations and enhancing the predictability of suspected EC. Building on this, the EC-Diagnosis model, which integrates clinical risk factors, laboratory test results, as well as plasma, cervical, and uterine secretion samples, achieves balanced performance in sensitivity, specificity, and accuracy for cancer confirmation, reducing false positives and overdiagnosis. For high-risk molecular stratification, the EC high-risk subtype model, using uterine secretion samples, detects p53abn and NSMP subtypes to guide personalized therapeutic plans prior to sequencing-based subtyping, shortening the timeline for precision medicine implementation. For clinical translation, we implemented the 2M-EC predictive platform as a Streamlit-based web tool ( https://2e-mc-web.streamlit.app/ ). This platform allows users to input multiple types of information, including age, HRT, menopausal status, HE4 levels, ultrasound findings, and molecular features from multiple biofluids, providing interpretable risk prediction results. Code packages are provided to extract molecular features from raw MALDI-TOF MS spectrum for input into the web tool ( https://github.com/lmsac/2M-EC ). The platform also visualizes feature contributions through interactive dashboards, highlighting individual-specific factors and fostering the development of precision diagnostics and treatment in oncology. The 2M-EC platform offers scalable decision support for both resource-limited and personalized healthcare settings.

Resource

Further information and requests for resources should be directed to and will be fulfilled by the lead contact, Liang Qiao ( [email protected] ). This study did not generate new unique reagents. • The proteomics data generated in this study have been deposited in ProteomeXchange via the iProX 79 partner repository as IPX0011779001 and are publicly available as of the date of publication at https://www.iprox.cn//page/subproject.html?id=IPX0011779001 . • The metabolomics data generated in this study have been deposited to the National Metabolomics Data Repository (Metabolomics Workbench) as PR002537 and are publicly available as of the date of publication at https://doi.org/10.21228/M8Q839 . • All original code has been deposited on GitHub ( https://github.com/lmsac/2M-EC ) and is publicly available at https://doi.org/10.5281/zenodo.19592109 ( https://doi.org/10.5281/zenodo.19592109 ) as of the date of publication. • Any additional information required to reanalyze the data reported in this work is available from the lead contact upon request. The proteomics data generated in this study have been deposited in ProteomeXchange via the iProX 79 partner repository as IPX0011779001 and are publicly available as of the date of publication at https://www.iprox.cn//page/subproject.html?id=IPX0011779001 . The metabolomics data generated in this study have been deposited to the National Metabolomics Data Repository (Metabolomics Workbench) as PR002537 and are publicly available as of the date of publication at https://doi.org/10.21228/M8Q839 . All original code has been deposited on GitHub ( https://github.com/lmsac/2M-EC ) and is publicly available at https://doi.org/10.5281/zenodo.19592109 ( https://doi.org/10.5281/zenodo.19592109 ) as of the date of publication. Any additional information required to reanalyze the data reported in this work is available from the lead contact upon request.

Discussion

During the past years, multi-omics integration has been widely applied in research on disease pathogenesis, progression, diagnosis, subtyping, and prognosis. Recent breakthroughs in artificial intelligence (AI) have revolutionized multi-dimensional data integration. The AI-driven integration of multi-biofluid omics and clinical phenotype aligns with current precision oncology frameworks that seek to unify molecular, tissue or biofluid, and clinical outcome-related information. Recent translational reviews have demonstrated how liquid biopsy platforms and ML algorithms can transform diagnostic and prognostic workflows in gynecologic oncology. 72 , 73 In this work, we propose the 2M-EC framework, which integrates multi-biofluid omics profiling with clinical data for EC risk stratification using ML. From the MS data, we identified EC-associated omics features in uterine secretions, cervical secretions, and plasma. These MS-derived features were combined with clinically relevant factors, curated through expert knowledge and statistical analysis, into unified models. Employing a multi-level fusion strategy, the 2M-EC platform comprises three key components: a minimally invasive EC screening model (using cervical secretions, plasma, and clinical data), an EC diagnosis model (augmented with uterine secretions), and a high-risk EC subtype detection model (based on uterine cavity secretions). Among the three models, the EC screening model carries the most immediate clinical relevance. Current EC diagnosis relies on endometrial biopsy or hysteroscopy, which is typically triggered by a positive screening result. Abnormal uterine bleeding is the most common screening symptom; however, fewer than 1% of women presenting with it are eventually diagnosed with cancer. Existing screening relies heavily on transvaginal ultrasound; while the final cancer-positive rate is <10% among suspected cases screened by transvaginal ultrasound. These limitations highlight a critical gap in EC screening. The EC-Screen model was developed and tested using a cohort from the Obstetrics and Gynecology Hospital of Fudan University and externally validated on a second cohort from the Shanghai Tenth People’s Hospital of Tongji University. On the external validation cohort, the model achieved a sensitivity of 95.65%, an accuracy of 70.45%, and a specificity of 61.54%, demonstrating the promising clinical potential of the 2M-EC framework for EC screening. It is important to note that the multi-omics data in this study is collected by MALDI-TOF MS rather than LC-MS/MS. While LC-MS/MS-based omics analysis can provide in-depth molecular profiling, MALDI-TOF MS offers advantages such as ease of operation, rapid turnaround, high throughput, and low cost, which are key factors for clinical applicability. In this work, the whole assay from sampling to clinical decision support can be finished in 1 h, including ∼45 min of sample pretreatment, ∼10 min of MS analysis, and ∼2 min of model prediction. Actually, MALDI-TOF MS has been licensed for clinical use such as bacterial pathogen identification. Studies show that MALDI-TOF MS has the potential to serve as a diagnostic platform for various purpose, enabling antimicrobial resistance prediction, 29 severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) detection, 22 Alzheimer’s biomarker quantification, 74 heterogeneous risk stratification of metabolic syndrome, 75 as well as diagnostic and prognostic prediction of gastric cancer. 76 . The clinical translation of the developed models must address the generalizability issue. This study has incorporated some strategies to mitigate it. The binning preprocessing aids in standardizing MS signals and mitigating heterogeneity across patients. The matching of analyte-algorithm pairs synergistically optimized the model’s computational efficiency. Besides, the models were assessed both on an internal hold-out cohort and an independent external cohort from another clinical center. Future validations across diverse international cohorts are necessary, accounting for key EC confounding factors such as ethnicity, BMI, and hormonal background that could influence metabolomic signatures. The molecular classification of EC was originated with TCGA, which provided the foundational research by defining four distinct molecular subtypes with profound prognostic significance. 8 To translate this complex genomic data into a clinically applicable tool, the Proactive Molecular Risk Classifier for Endometrial Cancer (ProMisE) was developed in 2015, 77 offering a method using immunohistochemistry and targeted sequencing to approximate the TCGA subgroups. The clinical validation of ProMisE ultimately led to its formal adoption into global clinical practice, culminating in the 2023 FIGO staging revision, 33 , 78 which officially integrated molecular subtypes into the EC staging system, marking a paradigm shift from a purely histologic to a molecular-driven diagnostic and prognostic framework. Our high-risk subtyping model based on uterine cavity secretions demonstrates strong potential for early identification of aggressive tumors. Compared with ProMisE, MALDI-TOF MS greater convenience for clinical application. However, broader clinical translation will require larger sample sizes and long-term accumulation of genotyped cases with complete TCGA-based molecular classification. We are conducting an ongoing study to continuously expand the endometrial disease biobank, with a particular focus on accumulating EC cases annotated with molecular subtyping. Future work will focus on the expansion of the 2M-EC framework into efficient EC subtyping to assist in the treatment of patients with EC. A current technical limitation of the study is the need for preprocessing of MALDI-TOF MS and clinical data before input into the predictive model, underscoring the importance of developing an integrated in vitro diagnostic device that combines MS with electronic clinical record systems and a reporting interface embedded with a computational module for automated data processing and model prediction. Furthermore, although the models were validated using both an internal hold-out cohort and an independent external cohort from multiple clinical centers, future assessments across diverse international cohorts are warranted. Such studies should account for regional differences in socioeconomic development, which often correspond to variations in lifestyle factors, including ethnicity, BMI, and hormonal background, that may influence endometrial cancer risk and metabolomic profiles.

Introduction

Endometrial cancer (EC) is a prevalent gynecological malignancy, with its incidence rising due to socioeconomic development and population aging. According to Global Cancer Statistics 2022 , there were 420,242 new EC cases worldwide in 2022, with over 90% of the cases occurring in women over the age of 50. 1 The current gold standard for diagnosing EC is endometrial biopsy or hysteroscopy-guided biopsy, which is highly invasive. 2 The decision to proceed with tissue biopsy is guided by findings from initial screening of EC. The current initial screening of EC mainly relies on transvaginal ultrasound, 3 which is primarily effective in postmenopausal women but less efficient in younger populations, 4 possibly due to intra-tumor heterogeneity. 5 Abnormal uterine bleeding is the most common presenting symptom for EC initial screening. 6 However, fewer than 1% of women with abnormal bleeding are ultimately diagnosed with EC. 7 Developing accurate EC screening and diagnosis methods remains a critical challenge in EC early management. On top of diagnosis, molecular subtyping of EC is essential for patient prognosis management. In 2013, The Cancer Genome Atlas (TCGA) proposed a classification of EC with four molecular subgroups 8 : POLE-ultramutated (POLEmut), microsatellite-unstable (MMRd), copy-number high tumors with TP53 mutations (p53abn), and a residual group lacking these alterations (NSMP). These subgroups, each with distinct clinical and prognostic features, are now incorporated into clinical guidelines. Current molecular classification of EC relies on sequencing, which is time-consuming and, hence, significantly delays the implementation of precision therapies. Several studies have developed deep learning methods on histopathological images to predict EC molecular subtypes, 9 , 10 which, however, suffer from limitations such as tissue variability and insufficient accuracy. Rapid subtyping is another urgent requirement in EC management. EC is intricately associated with the hormonal regulation of the reproductive system. Molecular analysis of body fluids offers a promising approach to overcoming current diagnostic challenges. In plasma, human epididymis protein 4 (HE4) has emerged as one of the most promising EC biomarkers. 11 Other potential plasma-based liquid biopsy markers include exosomal microRNAs (miRNAs), 12 circulating tumor DNA, 13 , 14 and circulating tumor RNA, 15 all showing encouraging potential for early EC detection. However, although plasma biomarker assays represent a valuable alternative for EC diagnosis, 16 reliance on single-marker strategies often lacks adequate diagnostic accuracy. Body fluids collected in proximity to tumor sites are particularly rich in molecular signatures of malignancy. Cervicovaginal specimens can provide tumor-specific information through the natural shedding of malignant cells into the lower genital tract. 17 Intact cells originating from the upper reproductive tract may also be detected in vaginal fluid and urine, especially in women presenting with abnormal uterine or postmenopausal bleeding. 18 Cervicovaginal methylated DNA markers have been reported as promising indicators for EC detection. 19 , 20 Beyond gene and protein biomarkers, metabolomic profiling can capture dynamic physiological phenotypes and disease-specific biochemical alterations, offering powerful insights into active pathological processes. 21 Studies have shown that metabolomic analysis of cervicovaginal and uterine secretions holds considerable promise for improving EC stratification. Furthermore, integrating multi-omics datasets enables comprehensive mapping of disease mechanisms and patient heterogeneity. 22 When combined with clinical features, these multi-omics profiles allow precise patient classification into distinct risk subgroups, establishing a robust foundation for the advancement of precision medicine in EC. 23 Recent advances in high-throughput analysis technologies have accelerated the discovery of multi-omics biomarkers from body fluids. 24 Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF MS) is a widely used analytical tool for detecting a broad spectrum of molecules, from small compounds like amino acids and lipids to macromolecules such as peptides, nucleic acids, and proteins. 25 , 26 While MALDI-TOF MS is less comprehensive than liquid chromatography coupled tandem mass spectrometry (LC-MS/MS) for omics analysis, it offers significant advantages, such as high throughput, minimal sample preparation, robustness, and ease of operation, making it particularly suitable for clinical applications. 27 Actually, the technique has been successfully employed in various clinical areas, including microbial identification, 28 , 29 screening of infectious disease, 30 detection of alanine abnormality in ocular adnexal lymphoma, 31 and delineation of pituitary tumors. 32 Recent advances in the 2023 International Federation of Gynecology and Obstetrics (FIGO 2023) staging have improved surgical and pathological risk stratification in EC, emphasizing the interplay among histology, stage, and molecular profile. 33 Computational algorithms for molecular classification are now being implemented in real-world settings, bridging traditional pathology and genomics. 34 Furthermore, guideline comparisons across major international societies underscore the urgent need for standardized, data-driven stratification approaches. 35 These developments provide the rationale for designing AI-integrated, multi-omics models that enable personalized stratification through multi-omics integration. In this study, we developed a machine learning (ML)-driven platform for multimodal, multilevel EC risk stratification (2M-EC). Two clinical cohorts from different clinical centers were enrolled, including a cohort of 531 participants (273 EC and 258 control) from the Obstetrics and Gynecology Hospital of Fudan University for model development, and a cohort of 204 participants (80 EC and 124 control) from the Shanghai Tenth People’s Hospital of Tongji University for external model validation. In situ uterine cavity secretions, cervical secretions, and plasma were collected from the participants for MALDI-TOF MS-based multi-omics profiling. Clinical information was fused with the multi-omics data for EC risk assessment. ML was used to identify EC-specific signatures from both the omics and clinical data. The method achieved 95.65% sensitivity for minimally invasive screening of EC using plasma and cervical secretion samples, while demonstrating balanced performance with an area under the curve (AUC) = 0.94 for EC confirmation using plasma, cervical, and uterine secretion samples on the external validation cohort. Additionally, it showed promising potential for detecting high-risk subtypes of EC, including p53abn and NSMP. We also employed LC-MS/MS to identify the metabolic and peptidomic features, and demonstrated the biomedical plausibility of the models. A Streamlit-based web tool has been developed to facilitate clinical translation, providing scalable decision support for diverse healthcare environments.

Coi Statement

The authors declare no competing interests.

Star★Methods

REAGENT or RESOURCE SOURCE IDENTIFIER Biological samples Human plasma samples Obstetrics and Gynecology Hospital of Fudan University, Shanghai Tenth People’s Hospital of Tongji University N/A Human cervical secretion samples Obstetrics and Gynecology Hospital of Fudan University, Shanghai Tenth People’s Hospital of Tongji University N/A Human uterine cavity secretion samples Obstetrics and Gynecology Hospital of Fudan University, Shanghai Tenth People’s Hospital of Tongji University N/A Chemicals, peptides, and recombinant proteins Sinapinic acid Sigma-Aldrich 85429 15 nm gold nanoparticle XFNAN XF J60 Methanol Sigma-Aldrich 34860 Acetonitrile Sigma-Aldrich 34851 TiO 2 nanoparticle Sigma-Aldrich Degussa P25 Phosphate-buffered saline Beyotime Biotechnology C0221A MOPS Beyotime Biotechnology ST1533 SDS-PAGE loading buffer Beyotime Biotechnology P0015L NH 4 HCO 3 Sigma-Aldrich 5.43835 Dithiothreitol Sigma-Aldrich D0632 Iodoacetamide Sigma-Aldrich I1149 Sequencing grade modified trypsin Promega V5111 Trifluoroacetic acid Sigma-Aldrich 302031 Formic acid Sigma-Aldrich 5.43804 Acetic acid Sigma-Aldrich 5.43808 10 kDa ultrafiltration tube Sigma-Aldrich UFC5010 C18 spin column Thermo Fisher Scientific 89878 Reversed-phase column (15 cm × 75 μm) IonOpticks AUR4-15075C18-CSI C8 beads Dr. Maisch r119.b8. Amide column (100 mm × 2.1 mm) Waters Corporation 176001908 Ammonium acetate Sigma-Aldrich 5.43834 Ammonium hydroxide Sigma-Aldrich 5.43830 C18 column (2.6 μm, 2.1 × 100 mm) Phenomenex 00D-4462-AN Critical commercial assays Coomassie Brilliant Blue rapid staining solution Beyotime Biotechnology P0003M Pierce TM Quantitative Colorimetric Peptide Assay Thermo Fisher Scientific 23275 iRT Kit Biognosys N/A Deposited data Proteomics data This paper https://www.iprox.cn//page/subproject.html?id=IPX0011779001 Metabolomics data This paper https://doi.org/10.21228/M8Q839 Software and algorithms R software R Foundation https://www.r-project.org/ ; version 4.2.1; RRID: SCR_001905 Python Python Software Foundation https://www.python.org/ ; version 3.9.0; RRID: SCR_008394 StandardScaler (scikit-learn) scikit-learn developers https://scikit-learn.org/ ; version 1.5.1; RRID: SCR_002577 XGBoost XGBoost developers https://xgboost.readthedocs.io/ ; version 1.7.0; RRID: SCR_023175 MetaboAnalyst Xia Lab https://www.metaboanalyst.ca/ ; version 6.0; RRID: SCR_015539 Adobe Illustrator Adobe Inc. https://www.adobe.com/products/illustrator.html ; version 2024; RRID: SCR_010279 BioRender BioRender https://www.biorender.com/ ; RRID: SCR_018361 ROMA calculation script This paper https://github.com/lmsac/2M-EC/blob/main/ROMA.py https://doi.org/10.5281/zenodo.19592109 ( https://doi.org/10.5281/zenodo.19592109 ) 2M-EC platform code This paper https://github.com/lmsac/2M-EC https://doi.org/10.5281/zenodo.19592109 ( https://doi.org/10.5281/zenodo.19592109 ) ROMA algorithm ROMA web server https://xema.com.ua/en/roma/ Spectronaut Biognosys AG version 19.0 SpectraMine Biognosys AG version 4.5 flexAnalysis software Bruker Corporation v3.4; RRID: SCR_017365 MALDIquant Sebastian Gibb https://codeberg.org/sgibb/MALDIquant ; v1.22.3 MALDIquantForeign Sebastian Gibb https://codeberg.org/sgibb/MALDIquantForeign ; v0.14.1 ProteoWizard MacCoss Lab, University of Washington https://proteowizard.sourceforge.io/ ; version 3.0.6150; RRID: SCR_012056 XCMS Scripps Research https://github.com/sneumann/xcms/issues/194 ; version 1.46.0; RRID: SCR_015538 MetDNA Zhu Lab, Interdisciplinary Research Center on Biology and Chemistry, Chinese Academy of Sciences http://metdna.zhulab.cn/ SHAP Scott Lundberg (PyPI) https://shap.readthedocs.io/ ; v0.46.0 This study enrolled women participants who underwent total hysterectomy with endometrial cancer or benign endometrial diseases at two independent clinical centers. The discovery cohort was recruited from the Obstetrics and Gynecology Hospital of Fudan University, comprising 273 participants with endometrial cancer and 258 participants with benign endometrial diseases. The external validation cohort was recruited from Shanghai Tenth People’s Hospital of Tongji University, comprising 80 participants with endometrial cancer and 124 participants with benign endometrial diseases. Exclusion criteria for all participants included prior malignancy, history of hysterectomy, atypical endometrial hyperplasia, or cervical lesions. Participant age ranged from 22 to 80 years in the discovery cohort and from 21 to 91 years in the external validation cohort. Age was included as a clinical variable in the multimodal machine learning models, without separate statistical testing for between-group differences. All participants were women (gender based on clinical records). As endometrial cancer occurs exclusively in individuals with female reproductive organs, gender is inherently a biological constant in this disease context and was accounted for in the study design. Endometrial uterine secretions were collected using a disposable brush before uterine removal under sterile conditions in the operating room. The brush was advanced 2 cm through the cervical canal and rotated three times in each uterine segment. The brush was then transferred to a 15 mL centrifuge tube containing 7 mL of 80% methanol. Brush heads were detached with sterile forceps and stored together with the samples. All samples were immediately flash-frozen in laparoscopic specimen liquid nitrogen containers, transported on dry ice, and stored at −80°C until metabolites extraction. During preoperative examination, cervical secretions containing cells were collected using HPV brushes before bimanual palpation. These brush heads were immersed directly in 7 mL centrifuge tubes with 80% methanol. Brush heads were stored together with the samples. All samples were immediately flash-frozen in laparoscopic specimen liquid nitrogen containers, transported on dry ice, and stored at −80°C until metabolites extraction. Peripheral blood (8 mL) was collected preoperatively into EDTA tubes (BD Biosciences, USA), centrifuged at 2000 g for 10 min (4°C), and plasma aliquots were stored in liquid nitrogen for subsequent peptidomic/metabolomic analyses. This study was conducted in accordance with the principles of the Declaration of Helsinki. Written informed consent was obtained from all participants prior to enrollment. The study protocol was approved by the Institutional Review Boards of both participating institutions: the Ethics Committee of Obstetrics and Gynecology Hospital of Fudan University (Ethics approval number: 2022-166) and the Ethics Committee of Shanghai Tenth People’s Hospital (Ethics approval number: SHSY-IEC-6.0/25K266/P01). Plasma samples for peptide analysis were pretreated using the PMFpre 1010305 kit (Bioyong, Beijing, China). 5 μL of plasma was transferred into 45 μL of sample processing buffer and vortexed for 30 s. Subsequently, 10 μL of the processed mixture was taken and mixed with 10 μL of sinapinic acid (SA) matrix (Sigma-Aldrich, St. Louis, MO, USA). 1 μL of the final mixture was deposited onto a Bruker MALDI target plate (Bruker, Bremen, Germany) to dry under the ambient condition. For uterine and cervical metabolites, frozen samples were thawed at 4°C and centrifuged at 12,000 g for 5 min 200 μL of the supernatant was transferred to a new centrifuge tube (1.5 mL), lyophilized, and reconstituted in 50 μL of deionized water. 1 μL of the sample was spotted onto a Bruker MALDI target plate and overlaid with 15 nm gold nanoparticle matrix (1 μL, 0.05 mg/mL, XF J60; XFNAN, Nanjing, China). For blood metabolites, 50 μL of blood plasma was mixed with 200 μL of ice-cold methanol:acetonitrile (1:1, v/v). The mixture was vortexed, sonicated for 2 min, and incubated at −20°C for 30 min. After centrifugation at 16,000 g (4°C, 20 min), 100 μL of the supernatant was lyophilized and reconstituted in 20 μL of methanol:acetonitrile:water (1:1:2, v/v/v). 1 μL of the sample was spotted onto a Bruker MALDI target plate and overlaid with Degussa P25 TiO 2 matrix (1 μL, 1 mg/mL, Sigma-Aldrich, St. Louis, MO, USA). Two technical replicates (analyzed twice from one preparation) were prepared for all samples. The one exhibiting the stronger peak densities was used for ML models. MALDI-TOF MS analysis was performed using two instruments. Bruker Microflex LRF MALDI-TOF MS (Bruker, Bremen, Germany) was used for plasma peptides (linear mode: 3–30 kDa, 1000 laser shots, 40 Hz, 52% laser intensity) and blood metabolites (reflection mode: 100–1000 Da, 200 laser shots, 40 Hz, 67% laser intensity). Bruker ultrafleXtreme MALDI-TOF/TOF MS (Bruker, Bremen, Germany) was used for uterine/cervical metabolites (reflection mode: 100–1000 Da, laser frequency 2000 Hz, 70% laser intensity). Detector gains were maintained at 10.0× for blood analyses and 3.0× for uterine/cervical analyses. All spectra were acquired in positive ion mode. Raw MALDI-TOF MS spectra were converted to .txt format using the flexAnalysis software (Bruker, Germany). Spectral preprocessing was performed with R packages MALDIquant and MALDIquantForeign, 80 implementing: 1) square root transformation for variance stabilization, 2) Savitzky-Golay filtering (15-point window) for smoothing, and 3) SNIP algorithm (20 iterations) for baseline correction. After the preprocessing, we conducted mass-to-charge ratio ( m / z ) binning with sample-specific resolutions. Specifically, for the plasma metabolic profiles (100–1000 Da), a bin width of 1 Da was applied, resulting in 900 evenly distributed bins. This choice was made considering the resolution of the Bruker Microflex LRF MALDI-TOF MS system used for data acquisition. For the uterine/cervical metabolic profiles (100–1000 Da), a bin width of 0.1 Da was employed, yielding 9,000 bins. This binning strategy accounted for the high resolution achievable with the Bruker ultrafleXtreme MALDI-TOF/TOF MS instrument used in this subset of samples. Regarding the plasma peptidomic profiles (3–30 kDa), the bin width was 40 Da, resulting in 675 bins to cover the broad mass range. This binning scheme aligned with established practices for processing MALDI-TOF MS data from human plasma peptidome studies, where the number of detected peaks is typically less than 500. 81 , 82 The usage of 675 bins ensured sufficient resolution to capture meaningful peptidome features without over-fragmentation of the MS peaks. All binned data were processed with StandardScaler (scikit-learn v1.5.1) before ML. For bottom-up proteomics analysis, 10 μL of plasma was mixed with 6 μL of phosphate-buffered saline (PBS) and 4 μL of 5X SDS-PAGE loading buffer (Beyotime Biotechnology, Shanghai, China). The diluted samples were denatured by heating at 95°C for 5 min. Precast 12% SDS-PAGE gels were loaded with 5 μL pre-stained protein marker and an equal volume of denatured samples. Electrophoresis was performed using MOPS running buffer (prepared from Tris-MOPS powder according to the manufacturer’s protocol, Beyotime Biotechnology, Shanghai, China) under a constant voltage of 120 V until target protein bands were sufficiently separated (∼1 h, adjusted based on gel length). After electrophoresis, the gel was rinsed three times with deionized water to remove residual buffer. Proteins were visualized by staining with 20 mL BeyoBlue Coomassie Brilliant Blue rapid staining solution (Beyotime Biotechnology, Shanghai, China) for 30 min at room temperature with gentle agitation. Excess dye was removed by repeated washing with deionized water until a clear background was achieved. Target protein bands (10 kDa–25 kDa) were manually excised and diced into 0.5–1 mm 3 gel pieces, which were then transferred to 1.5 mL microcentrifuge tubes. Gel pieces were washed twice with 1 mL water (5 min per wash under agitation) to remove impurities. Destaining was performed by incubating the gel pieces with 1 mL destaining solution (50% acetonitrile (ACN) 50% 25 mM NH 4 HCO 3 aqueous solution) at 37°C with 500 rpm shaking until the gel became transparent. The destaining solution was discarded, and the gel pieces were washed twice with 1 mL 25 mM NH 4 HCO 3 (5 min per wash). To dehydrate the gel, 500 μL ACN was added and incubated for 5 min at room temperature. This step was repeated once, and the gel pieces were left in a fume hood for air-dry. Protein disulfide bonds were reduced by incubating the dried gel pieces with 50 mM dithiothreitol (DTT, dissolved in 25 mM NH 4 HCO 3 ) at 56°C for 1 h with 500 rpm shaking. After centrifugation, the supernatant was removed, and the gel pieces were washed twice with 25 mM NH 4 HCO 3 followed by ACN dehydration. For alkylation, 100 mM iodoacetamide (IAA, prepared in 25 mM NH 4 HCO 3 ) was added to the gel pieces, and the reaction proceeded for 30 min at room temperature in dark. The gel pieces were then washed and dehydrated again. The dehydrated gel pieces were rehydrated with sequencing-grade trypsin (1 μg per 25 μL gel volume, dissolved in 25 mM NH 4 HCO 3 ) and incubated at 37°C overnight (12–16 h) with 500 rpm shaking. After digestion, the supernatant was collected by centrifugation. To maximize peptide recovery, the gel pieces were sequentially extracted with 300 μL of 50% ACN/0.1% trifluoroacetic acid (TFA) (ultrasonicated for 10 min, twice) and 200 μL pure ACN (ultrasonicated for 10 min). All supernatants were pooled and concentrated to dryness using a vacuum centrifugal concentrator. The dried peptides were reconstituted in 0.1% formic acid and desalted using C18 spin columns (Thermo Fisher Scientific, Waltham, MA, USA). Peptide content was quantified using Pierce Quantitative Colorimetric Peptide Assay (Thermo Fisher Scientific, Waltham, MA, USA). For top-down proteomic analysis, 200 μL of plasma was vortexed thoroughly, and placed on ice for 10 min. The sample was then centrifuged at 12,000 g for 20 min at 4°C, and the supernatant was collected. The collected supernatant was filtered using Amicon Ultra 10 kDa ultrafiltration tubes (Merck KGaA, Germany) by centrifugation at 14,000 g for 45 min at room temperature to remove peptides with molecular weights exceeding 10 kDa. The filter membrane was subsequently washed with 100 μL of 32% acetic acid and centrifuged again for 30 min. The resulting filtrate was loaded onto a desalting column, where peptide washing and sample elution were performed sequentially. Peptide content was quantified using Pierce Quantitative Colorimetric Peptide Assay (Thermo Fisher Scientific, Waltham, MA, USA), and the samples were concentrated by vacuum freezing. For both samples, the peptides were reconstituted in phase A (0.1% formic acid in water) containing the iRT standard (Biognosys, Switzerland) to a final concentration of 100 ng/μL. After centrifugation at 15,000 g for 10 min, the supernatant was transferred to injection vials for subsequent LC-MS/MS analysis. For bottom-up proteomics, 2 μL digested peptides were separated using an Aurora Elite reversed-phase column (15 cm × 75 μm, IonOpticks, Australia) packed with 1.7 μm C18 materials, connected to an UltiMate 3000 ultra-high-performance liquid chromatography (UHPLC) system (Thermo Fisher Scientific, Waltham, MA, USA). The column chamber was maintained at 60°C, and peptide separation was performed using a binary gradient at a constant flow rate of 300 nL/min. The mobile phases consisted of phase A (0.1% formic acid in water) and phase B (0.1% formic acid in acetonitrile). The separation gradient was executed as follows: 0–3 min: 2.2–8.0% B, 3–16 min: 8–28% B, 16–24 min: 28–38% B, 24–30 min: 38–90% B. Peptides were analyzed using a trapped ion mobility spectrometry time-of-flight mass spectrometer (TIMS TOF Pro2, Bruker Daltonics, Bremen, Germany). The instrument was operated in parallel accumulation serial fragmentation data independent acquisition (diaPASEF) scan mode. 83 The mass spectrometry data were acquired with the following parameters: m/z 100–1700 mass range, 1500 V ion transfer voltage, 100 ms accumulation time, 20 eV–59 eV collision energy, 0.6–1.6 Vs.⋅cm −2 ion mobility range. For top-down proteomics, 2 μL peptides were separated on a self-packed 15 cm × 75 μm C8 column (1.9 μm, Dr. Maisch, Ammerbuch, Germany). The column chamber was maintained at 60°C, and peptide separation was performed using a binary gradient at a constant flow rate of 300 nL/min. The mobile phases consisted of phase A (0.1% formic acid in water) and phase B (0.1% formic acid in acetonitrile). The separation gradient was executed as follows: 0–3 min: 2.2–8.0% B, 3–16 min: 8–28% B, 16–24 min: 28–38% B, 24–30 min: 38–90% B. Peptides were analyzed using a TIMS TOF Pro2 mass spectrometer (Bruker Daltonics, Bremen, Germany). The instrument was operated in parallel accumulation serial fragmentation data dependent acquisition (DDA-PASEF) scan mode, 83 with both accumulation and ramp time set to 166 ms. The mass scan range was set to m/z 300–2500, and ion mobility scans were performed over a range of 0.6–1.6 Vs.⋅cm −2 . Data acquisition relied on precursor ions being fragmented at ion mobility-dependent collision energy, which increased linearly from 20 to 59 eV. The total acquisition cycle lasted 1.03 s, comprising one full TIMS-MS scan followed by five PASEF MS/MS scans with charge counts ranging from 0 to 5. Low-abundance precursor ions with intensities higher than the 1000-count threshold but below the 10,000-count target were repeatedly selected for fragmentation, while ions exceeding the target intensity were dynamically excluded for 0.4 min. For bottom-up proteomics, protein identification and quantification were performed using Spectronaut (version 19.0, Biognosys AG, Switzerland). The raw MS/MS spectra of peptides were searched against the Homo sapiens reference proteome database (UniProt, downloaded on 14 November 2024). FDR threshold was 1% for peptide-spectrum matches (PSMs), peptides, and protein groups. Other parameters were default. For top-down proteomics, protein identification and quantification were performed using SpectraMine (version 4.5, Biognosys AG, Switzerland) for data-dependent acquisition (DDA) analysis. The raw MS/MS spectra of peptides were searched against the Homo sapiens reference proteome database (UniProt, downloaded on 14 November 2024). Threshold was 1% for peptide-spectrum matches (PSMs), peptides, and protein groups. No enzymatic digestion and fixed modifications were set. Other parameters were default. MALDI-TOF MS peaks were selected based on the featured mass spectra bins, and then matched against the identified serum free peptides within the mass tolerance of 2000 ppm. Only singly charged ions were considered for the MALDI-TOF MS peaks. If multiple peptides can match with a single peak, the one with the smallest mass difference was chosen as the identification result. The MALDI-TOF MS peaks not matched to the serum free peptides were further analyzed using the SDS-PAGE-based bottom-up proteomics identification results referring to a reported strategy. 81 Identified peptides from bottom-up proteomics were mapped to the corresponding protein sequences from UniProt. The first mapped amino acid position in a protein sequence was recorded as the min position; while the last mapped amino acid position in a protein sequence was recorded as the max position. Possible protein fragments were listed considering the mapping results and the SDS-PAGE molecular weight range (10k to 25k Da), including min-to-max, N-terminus-to-max, min-to-C-terminus and full-length proteins. Then, the MALDI-TOF MS peaks were matched against the possible protein fragments within the mass tolerance of 2000 ppm. Only singly charged ions were considered for the MALDI-TOF MS peaks. If multiple protein fragments can match with a single peak, the preference order was full-length protein > N-terminus-to-max or min-to-C-terminus > min-to-max. If further decision was required, the one with the smallest mass difference was chosen as the identification result. Protein matching was implemented using the R package “match_proteins”, while peptide calculation and matching were performed using the R package “Calculating molecular weight & matching molecular weight”, which can be found from the GitHub ( https://github.com/lmsac/2M-EC/tree/main/Identification ). Uterine or cervical secretion samples were thawed at 4°C and homogenized by gentle vortex mixing. After centrifugation at 13,800 g (4°C, 5 min), the supernatant was carefully collected. The supernatant was concentrated using a vacuum centrifugal concentrator until lyophilized powder was obtained. After reconstitution in 50 μL organic solution (methanol:acetonitrile:water = 1:1:2), samples were centrifuged at 16,000 g (4°C, 30 min) and the supernatant was used for LC-MS/MS analysis. For plasma sample, 100 μL plasma was mixed with 400 μL of ice-cold methanol/acetonitrile (1:1, v/v), vortexed for 30 s, and sonicated for 2 min. The mixture was incubated at −20°C for 30 min for protein precipitation, then centrifuged at 16,000 g (4°C, 20 min). 300 μL supernatant was collected, concentrated to dryness by vacuum centrifugation. After reconstitution in 50 μL organic solution (methanol:acetonitrile:water = 1:1:2), samples were centrifuged at 16,000 g (4°C, 30 min) and the supernatant was subjected to LC-MS/MS analysis. The extracted metabolites were analyzed by HPLC-MS/MS using Orbitrap Exploris 480 mass spectrometer coupled to a Vanquish liquid chromatography system (Thermo Fisher Scientific, USA). For HILIC separation, the ACQUITY UPLC BEH Amide column (100 mm × 2.1 mm i.d., 1.7 μm; Waters Corporation, Framingham, MA, USA) was used. 2 μL sample was injected and separated by LC. Phase A was 100% water with 25 mM ammonium acetate and 25 mM ammonium hydroxide. Phase B was 100% acetonitrile. Linear gradient started from 5% A to 60% A over 8 min, followed by a return to 5% A at 9.1 min. The column flow rate was maintained at 500 μL/min with the column temperature of 40°C. For reverse phase separation, the Kinetex C18 column (2.6 μm, 2.1 × 100 mm, Phenomenex, Torrance, CA, USA) was used. 2 μL sample was injected and separated by LC. Phase A was 100% water and 0.01% acetic acid. Phase B was 50% acetonitrile and 50% toluene dicarboxylic acid. Linear gradient started from 99% A to 1% A over 8 min, followed by a return to 99% A at 9.1 min. The column flow rate was maintained at 300 μL/min with the column temperature of 40°C. The MS analysis was performed in positive ion mode and negative ion mode, respectively. Data dependent acquisition (DDA) was used to collect full scan MS and MS/MS. The electrospray voltage was set to 3500 V for positive mode and −2800 V for negative mode. The survey of full scan MS spectra (m/z 70–1200) was acquired in the Orbitrap with 60,000 resolutions. The normalized automatic gain control (AGC) target at 100% and the maximum injection time was 100 ms. Then the precursor ions were selected into collision cell for fragmentation by higher-energy collision dissociation (HCD), the collision energy was 20, 30 and 40. The MS/MS resolution was set at 30,000. The maximum injection time was 60 ms. The isolation window was 1 m/z, and the dynamic exclusion was 4 s. ProteoWizard (version 3.0.6150) 84 was used to convert raw MS data files to mzXML format and MS2 data files to mgf format. All MS files (mzXML format) were processed using R package “XCMS” (version 1.46.0) for peak detection and alignment. Metabolites identification was achieved by MetDNA ( http://metdna.zhulab.cn/ ) with the MS1 peak table and MS2 data files (mgf format). MALDI-TOF MS peaks were selected based on the featured mass spectra bins, and then matched against the identified metabolites within the mass tolerance of 5000 ppm. Only singly charged ions were considered for the MALDI-TOF MS peaks. If multiple metabolites can be matched to a single MS peak, all of them were retained as the metabolic identification result. ML models were optimized through a grid search approach following the protocol described by Osipov et al. 38 This systematic evaluation tested combinations of analytes (plasma metabolites, plasma peptides, uterine metabolites, cervical metabolites, and clinical data) and 17 different ML algorithms to assess model performance. The 17 ML algorithms include: RF_Regression, SVM_Model, L1_Norm_SVM_Model, L1_Norm_RF_Model, L1_Norm_LR_Model, RFE_LR_Model, RFE_RG_Model, LR_Logic, LR_MLP, LR_KNei, LR_SVM_Lin, LR_SVM_RBF, LR_SVM_Sig, LR_SVM_Poly, LR_GBDT, LR_XGB, LR_RF. Random forest was firstly adopted for feature selection from the MS data by considering up to 100 features through Monte Carlo cross validation (1000 times sampling). Feature combination leading to the best performance was selected for MS model development. A 5-fold cross validation (80% training and 20% validation) was adopted for model training using the selected features. To address class imbalance, borderline-SMOTE1 85 oversampling was applied prior to model training. The MS models were developed using four ML algorithms: Random Forest, XGBoost, LightGBM, and CatBoost. Clinical data preprocessing involved: 1) one-hot encoding for categorical variables with mode imputation for missing values, and 2) mean imputation for missing numerical data. The processed clinical variables were fused with binned mass spectral features. Fused data were standardized using StandardScaler (scikit-learn v1.5.1), with the fitted scaler preserved for future analyses. A 5-fold cross validation (80% training and 20% validation) was adopted for model training. To address class imbalance, borderline-SMOTE1 85 oversampling was applied prior to model training. Fusion models were developed using XGBoost. The final model incorporated both clinical and MS features. The trained models were tested on both test set from the Obstetrics and Gynecology Hospital of Fudan University cohort and the external validation cohort from Shanghai Tenth People’s Hospital of Tongji University. Model performance was evaluated using four metrics: accuracy, sensitivity, specificity, and AUC-ROC. For interpretability, SHAP explainers (SHAP v0.46.0) were implemented to calculate feature contributions, with decision pathways visualized through waterfall plots for representative samples. Through decision-level late fusion, we developed three task-driven models: EC-Screen, EC-Diagnosis, and EC-Subtyping. These models were integrated into the 2M-EC platform by deploying locally trained model files, preprocessed data pipelines, and application code to a remote repository (GitHub, https://github.com/lmsac/2M-EC ), followed by Streamlit-based app ( https://2e-mc-web.streamlit.app/ ) deployment with version-controlled dependencies. We also provide processed demonstration cases for platform evaluation in the GitHub repository ( https://github.com/lmsac/2M-EC/blob/main/2M-EC/Demonstration%20case.csv ). The raw mass spectra must first undergo preprocessing and binning to ensure analytical consistency with the established models available in the GitHub repository. ROMA web server ( https://xema.com.ua/en/roma/ ) was used to assess the risk of EC based on the levels of two biomarkers, HE4 and CA125. The ROMA score was calculated using the following formulas: For non-menopausal women: ROMA = 0.0049 × HE4 + 0.018 × CA125. For menopausal women: ROMA = 0.049 × HE4 + 0.022 × CA125. The predictive performance of our EC-Screen model was compared with the ROMA algorithm using the same test set. After excluding samples with missing clinical values in TS1, the comparison was performed on 54 cancer cases and 30 controls. Performance metrics including sensitivity, specificity, and AUC were calculated for both approaches. To facilitate the calculation of ROMA scores for our cohort, we wrote a custom script to automatically collect and process the predictive results from the ROMA online tool, which is also available in the GitHub repository ( https://github.com/lmsac/2M-EC/blob/main/ROMA.py ). The MetaboAnalyst 6.0 interactive platform ( https://www.metaboanalyst.ca/ ), 86 along with R (version 4.2.1) were employed for data processing and KEGG enrichment. Python (version 3.9.0) were employed for data processing and ML. Adobe Illustrator 2024 were used for mapping. BioRender was used for graphical asset development ( https://www.biorender.com ). The sample size for external validation was calculated a priori using the method proposed by Riley et al. 71 The calculation was based on the following assumptions to ensure sufficient statistical power: a well-calibrated model (calibration slope = 1, O/E = 1), an anticipated AUC of 0.88, and an estimated disease prevalence of 0.40. The Stata command pmvalsampsize was employed with parameters targeting precise estimates for calibration and discrimination (command: pmvalsampsize, type(b) prevalence(0.4) cstatistic(0.88) lpbeta(1.1, 1.65) oeciwidth(0.3) csciwidth(0.3) cstatciwidth(0.1)). This calculation indicated that a minimum of 196 participants, including at least 79 EC cases, would be required. It is important to note that these assumed parameters (AUC and prevalence) were used solely for sample size estimation. The actual validation cohort (204 participants with 80 EC) was consecutively enrolled to reflect a realistic clinical population, and the model’s performance was evaluated against the observed outcomes without expectation to match the assumed values. The MetaboAnalyst 6.0 interactive platform ( https://www.metaboanalyst.ca/ ) 86 along with R (version 4.2.1) were employed for statistical analysis. The statistical details of experiments can be found in corresponding figure legends and Results.

Acknowledgments

This work was supported by the 10.13039/501100001809 National Natural Science Foundation of China (NSFC 22525401 ), the Shanghai Municipal Hospital Emerging Frontier Joint Research Project ( SHDC12025141 ), the 10.13039/501100003399 Science and Technology Commission of Shanghai Municipality ( 23JS1400100 ), the 10.13039/501100012166 National Key R&D Program of China ( 2022YFC2704300 ), and the Fudan University Medical and Engineering Integration Project . The computations in this research were performed using the Computing for the Future at Fudan (CFFF) platform of Fudan University.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

SciLite annotations

chemicals 104
peptide palmitoyl amino acid peptide alanine methanol nitrogen methanol nitrogen peptide feruloylacetic acid water gold nanoparticle acetonitrile methanol acetonitrile water peptide water brilliant blue water water disulfide thioredoxin dithiol thioacetamide peptide trifluoroacetic acid peptide formic acid peptide peptide peptide acetic acid peptide peptide formic acid water peptide peptide acetonitrile peptide peptide amino acid amino acid methanol acetonitrile water ammonium acetate ammonium hydroxide acetonitrile acetic acid acetonitrile carboxylic acid peptide peptide peptide peptide leucine glycoprotein +44 more
organisms 10
noordeloos 2009062 human human human human noordeloos 2009062 noordeloos 2009062 noordeloos 2009062 severe acute respiratory syndrome coronavirus 2 severe acute respiratory syndrome coronavirus 2

Source provenance

europepmc
last seen: 2026-08-02T06:10:09.037253+00:00
scilite
last seen: 2026-06-28T09:31:30.222730+00:00
unpaywall
last seen: 2026-06-28T06:30:42.658729+00:00
License: CC-BY-NC-ND-4.0