Results
Serum peptides extracted from peripheral blood samples of ovarian cancer patients and healthy controls were analysed using MALDI-TOF/MS. Profiles in the 1000–10,000 Da range were obtained with high reproducibility across triplicate measurements (Fig. 1 A). Differential peptide peaks, visualized via three-dimensional fitting, were predominantly located between 3000 and 6000 Da (Figure S0). Following fitting and normalization, mass peaks from raw time-of-flight data displayed overlapping distributions (Fig. 1 B). Inter-sample correlation analysis indicated lower between-group than within-group correlation (Fig. 1 C). Volcano plot analysis revealed 17 up-regulated and 26 down-regulated serum peptides in ovarian cancer patients relative to controls (Fig. 1 D, P 1). For further investigation of diagnostic potential, the top 20 peptides with the largest expression differences (Fig. 1 E) and the top 20 most statistically significant peptides (Fig. 1 F) were selected.
Fig. 1 Mass spectrometry data processing and screening of differential peptides. (A) Representative MALDI-TOF/MS peptide profiles (1,000–10,000 Da) and Pearson correlation coefficients of one ovarian cancer sample and one age-matched healthy control. This figure displays a typical result from triplicate measurements per sample, with consistent profiles observed across all biological replicates (intra-sample correlation coefficient > 0.98, as shown in the correlation matrix), verifying the reproducibility of the MALDI-TOF/MS analysis. (B) Top: Peak overlap comparison of mass spectra after normalization; Bottom: Peak superposition of the original time-of-flight mass spectrometry data after fitting. (C) Correlation analysis between samples, in which green represents healthy controls and red represents ovarian cancer samples. (D) Volcano plot of differentially expressed peptides in serum of ovarian cancer and healthy controls, with up-regulated peptides marked in red and down-regulated peptides marked in blue. (E) Top 20 peptides with the most significant differences in expression (log 2 FC). (F) Top 20 differentially expressed peptides with the highest statistical significance (minimum p-value).
Mass spectrometry data processing and screening of differential peptides. (A) Representative MALDI-TOF/MS peptide profiles (1,000–10,000 Da) and Pearson correlation coefficients of one ovarian cancer sample and one age-matched healthy control. This figure displays a typical result from triplicate measurements per sample, with consistent profiles observed across all biological replicates (intra-sample correlation coefficient > 0.98, as shown in the correlation matrix), verifying the reproducibility of the MALDI-TOF/MS analysis. (B) Top: Peak overlap comparison of mass spectra after normalization; Bottom: Peak superposition of the original time-of-flight mass spectrometry data after fitting. (C) Correlation analysis between samples, in which green represents healthy controls and red represents ovarian cancer samples. (D) Volcano plot of differentially expressed peptides in serum of ovarian cancer and healthy controls, with up-regulated peptides marked in red and down-regulated peptides marked in blue. (E) Top 20 peptides with the most significant differences in expression (log 2 FC). (F) Top 20 differentially expressed peptides with the highest statistical significance (minimum p-value).
In this study, principal component analysis (PCA), kernel principal component analysis (KPCA), and t-distributed random neighborhood embedding (t-SNE) were used to reduce the dimension of the aforementioned mass spectrometry data, and the distribution density was visualized by matplotlib to generate a two-dimensional map to support multi-angle comprehensive decision-making (Fig. 2 A). After dividing the dataset into a training set and a test set, eight machine learning models were used to classify the two groups of differential peptide peaks, and the performance of the models was evaluated based on the ROC curve and confusion matrix (Fig. 2 B). Furthermore, the comprehensive analysis of ROC curve was used to show the discrimination ability of each model based on the true positive rate and false positive rate (Fig. 2 C), and multi-dimensional indicators such as accuracy, precision, recall and F1 score were combined to quantitatively evaluate the performance of the models to screen the best model (Fig. 2 D). Decision-curve analysis showed that the net benefit of all models was superior to that of the “all intervention” and “no intervention” reference strategies across thresholds (Fig. 2 E).
Fig. 2 Evaluation of diagnostic efficacy using eight machine learning algorithms. (A) 2D visualization of mass spectrum data after dimensionality reduction by PCA, KPCA and t-SNE. (B) ROC curves and confusion matrices for discriminating categorical differential peptide peaks based on eight machine learning algorithms. The shaded area reflects the fluctuation range of performance indicators in cross validation, reflecting the robustness of the model. The main diagonal of the confusion matrix is the correct classification (TP vs. TN), and the off-diagonal is the misclassification (FP vs. FN). (C) Comprehensive ROC curve and its relationship with true positive rate (TPR) and false positive rate (FPR). (D) Accuracy, precision, recall and F1 score of each machine learning algorithm. (E) DCA of different algorithms, with the shaded area indicating the confidence interval of net benefit compared with “treat all” or “treat nothing” strategy at each clinical threshold probability, reflecting its clinical applicability.
Evaluation of diagnostic efficacy using eight machine learning algorithms. (A) 2D visualization of mass spectrum data after dimensionality reduction by PCA, KPCA and t-SNE. (B) ROC curves and confusion matrices for discriminating categorical differential peptide peaks based on eight machine learning algorithms. The shaded area reflects the fluctuation range of performance indicators in cross validation, reflecting the robustness of the model. The main diagonal of the confusion matrix is the correct classification (TP vs. TN), and the off-diagonal is the misclassification (FP vs. FN). (C) Comprehensive ROC curve and its relationship with true positive rate (TPR) and false positive rate (FPR). (D) Accuracy, precision, recall and F1 score of each machine learning algorithm. (E) DCA of different algorithms, with the shaded area indicating the confidence interval of net benefit compared with “treat all” or “treat nothing” strategy at each clinical threshold probability, reflecting its clinical applicability.
To systematically resolve the decision mechanism of the classifier, an integrated evaluation framework combining the importance of gradient boosting features, SHAP value interpretation and LIME local interpretation was constructed. By performing feature importance cross-validation on four gradient boosting classifiers, random forest, decision tree, LightGBM and XGBoost, four key peptides (m/z = 5354.39, m/z = 4211.41, m/z = 2881.50, m/z = 2662.15) with consistent discrimination ability in multiple algorithms were identified (Fig. 3 A-C). Furthermore, SHAP values were calculated based on the background set of 100 samples to quantify the contribution of each feature to the classification of ovarian cancer at the overall and individual levels, revealing its context-dependent predictive influence (Fig. 3 D). LIME analysis identified directional associations between key features and category predictions through local interpretation of typical samples (Fig. 3 E). The feature consensus strategy in this study, which adopts majority voting across Gini importance, SHAP and LIME, is defined as an exploratory feature stability strategy for screening cross-model stable candidate features from high-dimensional proteomic data, rather than a confirmatory validation analysis of feature importance. The actual biological significance and clinical value of these features need to be further verified by subsequent functional experiments and independent cohorts.
Fig. 3 Feature selection based on different machine learning models. (A , B) Differential peptides with high classification performance were identified and screened by multiple machine learning algorithms. (C) Cross-validation results of feature importance scores obtained by decision tree, random forest, LGBM and XGBoost classifiers. (D , E) The model prediction and interpretation results based on Shapley value and LIME algorithm, respectively.
Feature selection based on different machine learning models. (A , B) Differential peptides with high classification performance were identified and screened by multiple machine learning algorithms. (C) Cross-validation results of feature importance scores obtained by decision tree, random forest, LGBM and XGBoost classifiers. (D , E) The model prediction and interpretation results based on Shapley value and LIME algorithm, respectively.
LC-ESI-MS/MS sequencing followed by UniProt annotation enabled confident protein assignment for eight of the nine discriminatory peptide peaks. We identified the key peptides (excluding m/z = 5354.39), and the encoding genes of the proteins from which these peptides were derived include SERPINA1, Ezrin, C4B, PDE11A, FGA, CCT2, FGB, and SRGN (Fig. 4 A). To evaluate the clinical discrimination ability of the key features in different models, we first plotted the ROC curves for each feature. The results showed clear differences in the area under the curve (AUC) corresponding to each feature (Fig. 4 B). DCA compared net benefits with “no treatment” and “population-wide treatment” strategies and determined thresholds at which each measure was superior to the reference strategy (Fig. 4 C). The net benefit value of each feature was determined as diagnostic by comparing the model prediction curve for each feature with the reference line, including “treatment-none” and “treatment-all”, as shaded areas. The analysis revealed statistically significant differences in the expression levels of the selected peptides (Fig. 4 D). Furthermore, the normalized mass spectrum peak maps of these peptides are presented (Fig. 4 E).
Fig. 4 Identification of differential peptides and results of analysis of clinical indicators for single differential peptides. (A) The corresponding genes of peptide peaks that differed significantly based on the UniProt database. (B) ROC curve was used to evaluate the diagnostic efficacy of individual differentially expressed peptides, in which up-regulated and down-regulated peptides were marked in red and blue, respectively. (C) DCA curve was used to further evaluate the clinical usefulness of each differential peptide and compare it with the reference line. (D) Analysis of the expression levels of individual differential peptides. (E) Overlay of spectral peaks of normalized peptide expression values.
Identification of differential peptides and results of analysis of clinical indicators for single differential peptides. (A) The corresponding genes of peptide peaks that differed significantly based on the UniProt database. (B) ROC curve was used to evaluate the diagnostic efficacy of individual differentially expressed peptides, in which up-regulated and down-regulated peptides were marked in red and blue, respectively. (C) DCA curve was used to further evaluate the clinical usefulness of each differential peptide and compare it with the reference line. (D) Analysis of the expression levels of individual differential peptides. (E) Overlay of spectral peaks of normalized peptide expression values.
To assess the consistency between the selected features and the original unsupervised classification, the datasets were clustered according to the initial labels (Fig. 5 A) and further subjected to bisecting K-means clustering (Fig. 5 B). Results demonstrated high concordance between the unsupervised and supervised clustering outcomes, suggesting that the feature set—comprising SERPINA1, Ezrin, C4B, PDE11A, FGA, CCT2, FGB, and SRGN—is appropriate for classifying ovarian cancer subtypes. Based on this finding, these genes were adopted as discriminant features, and eight machine learning algorithms were applied to plot ROC curves and compute AUC values (Fig. 5 C).
The confusion matrices for all classifiers are presented in Fig. 5 D. The ROC curves across different models were subsequently integrated to illustrate their relationship with the true positive rate (TPR) and false positive rate (FPR) (Fig. 5 E). DCA was employed to quantify the net benefit of the predictive models at varying decision thresholds. Results demonstrated that the net benefits of all models exceeded both the “treat-all” and “treat-none” reference lines, suggesting a relatively low misclassification risk of the machine learning classifiers in diagnosing ovarian cancer (Fig. 5 F). Furthermore, key performance metrics—including accuracy, precision, recall, and F1-score—were computed to provide a comprehensive evaluation of model efficacy (Fig. 5 G).
Fig. 5 Evaluation of Differential Polypeptides’ Diagnostic Efficacy. (A) Original clustering and (B) bisecting K-means clustering of the aforementioned differentially expressed genes. (C) ROC curves and AUC values of the differentially expressed serum peptide under different machine learning algorithms. (D) Confusion matrices under various clustering methods, where diagonal elements represent correctly classified samples and off-diagonal elements indicate misclassifications. (E) ROC curves illustrating their relationship with true positive rate (TPR) and false positive rate (FPR). (F) Decision curve analysis (DCA), with shaded areas denoting confidence intervals for net benefit across clinical decision thresholds. (G) Classification performance metrics—including accuracy, precision, recall, and F1-score—of the differentially expressed genes under multiple machine learning algorithms.
Evaluation of Differential Polypeptides’ Diagnostic Efficacy. (A) Original clustering and (B) bisecting K-means clustering of the aforementioned differentially expressed genes. (C) ROC curves and AUC values of the differentially expressed serum peptide under different machine learning algorithms. (D) Confusion matrices under various clustering methods, where diagonal elements represent correctly classified samples and off-diagonal elements indicate misclassifications. (E) ROC curves illustrating their relationship with true positive rate (TPR) and false positive rate (FPR). (F) Decision curve analysis (DCA), with shaded areas denoting confidence intervals for net benefit across clinical decision thresholds. (G) Classification performance metrics—including accuracy, precision, recall, and F1-score—of the differentially expressed genes under multiple machine learning algorithms.
To clarify the contribution and influence patterns of features in different models, a SHAP-based evaluation was performed on AdaBoost, XGBoost, RandomForest, and LGBM. In Fig. 6 A, AdaBoost, PDE11A and SRGN show significant positive SHAP values, indicating that they play an important role in pushing the model output toward the target category, whereas features such as Erln show a negative contribution. On XGBoost (Fig. 6 B), PDE11A maintained high importance and the SHAP distribution was wide, suggesting the diversity of its influence. Figure 6 C (RandomForest) had relatively balanced feature importance, with significant contributions from CCT2 and PDE11A. In Fig. 6 D (LGBM), PDE11A and SRGN again emerged as key features and showed clear effects at different value levels (high values correlated with positive SHAP values). The dependence plots further showed that numerical changes in most features followed a monotonic trend with model output -for example, elevated PDE11A levels consistently promoted target output in all models. Taken together, these results highlight the cross-model consistency of the characteristic peptide derived from PDE11A-encoded protein (m/z = 4211.41) in category discrimination, and also reveal the specific performance of features in different models, which is critical for elucidating the biological significance of these markers and optimizing the interpretability of models.
Fig. 6 SHAP-based Feature Importance and Association Analysis Across Four Machine Learning Models. (A) AdaBoost, (B) XGBoost, (C) RandomForest and (D) LGBM, which was used to clarify the contribution and influence patterns of features in different models, a SHAP-based evaluation.Left Subplots: SHAP scatter plots illustrating the impact direction and magnitude of each feature on model output. The X-axis represents SHAP values (positive values indicate a promoting effect on the target output, while negative values indicate an inhibitory effect), and the color denotes feature value levels (red for high, blue for low).
SHAP-based Feature Importance and Association Analysis Across Four Machine Learning Models. (A) AdaBoost, (B) XGBoost, (C) RandomForest and (D) LGBM, which was used to clarify the contribution and influence patterns of features in different models, a SHAP-based evaluation.Left Subplots: SHAP scatter plots illustrating the impact direction and magnitude of each feature on model output. The X-axis represents SHAP values (positive values indicate a promoting effect on the target output, while negative values indicate an inhibitory effect), and the color denotes feature value levels (red for high, blue for low).
Ting-Ting Gong et al. 24 performed a comprehensive proteomic analysis of 80 clear cell carcinoma (CCC), 79 endometrioid carcinoma (EC), 80 serous carcinoma (SC) and 30 healthy control samples. They constructed a proteomic profile of epithelial ovarian cancer histological subtypes and reveal the prognostic or diagnostic value of dysregulated proteins and phosphorylation sites in important pathways. After verification, our results were highly consistent with it, and the ROC curves and DAC values were generated, which validate the expression of individual proteins in ovarian cancer (Figure S1). SERPINA1 can promote the proliferation and invasion of cancer cells by regulating the tumor microenvironment and inhibiting cell apoptosis 25 ; The ectopic expression of Ezrin significantly enhanced the cell proliferation ability, invasion ability and degree of epithelial-mesenchymal transition of ovarian cancer cells 26 .
Materials
The present study utilized MALDI-TOF MS as the core technique for serum proteomic profiling. Curve fitting was performed on the raw spectral data to accurately resolve characteristic protein and peptide peaks representing distinct biological states, followed by normalization to minimize technical variations. Subsequently, a machine learning classifier was applied to analyse the normalized peak intensity matrix. Peaks with high feature importance were selected for biological annotation by mapping them to established biological pathways and functions, thereby identifying validated biomarkers that may serve as potential diagnostic or prognostic indicators for ovarian cancer. All data analysis and visualization were conducted in a Python-based analytical framework, with machine learning modeling implemented using XGBoost and visual representations generated using Matplotlib (version 3.7.0) and Seaborn (version 0.13.2).
A prospective case-control study was conducted at Xianyang Hospital of Yan’an University, enrolling 188 patients with histopathologically confirmed ovarian cancer and 208 age-matched healthy controls. Patients with ovarian cancer were consecutively recruited between January 2023 and September 2024 following diagnosis. Exclusion criteria included previous chemotherapy or radiotherapy, as well as comorbidities that could potentially affect serum protein profiles (such as chronic kidney disease). Healthy controls were screened by questionnaire and hematological testing to exclude underlying abnormalities. Venous blood samples(5mL) were collected into non-anticoagulant tubes, allowed to clot at 4 °C for 1 h, and then centrifuged at 3000 × g for 15 min at 4 °C. Within 2 h of collection, 500 µL aliquots of serum were prepared and stored at − 80 °C to ensure protein stability. This study received approval from the Ethics Committee of the Xianyang Hospital of Yan’an University (YDXY-KY-2023-006). All procedures were conducted in compliance with the Declaration of Helsinki and with informed consent. The serum samples collected were clinical laboratory by-products, thus involving no patient personal information or repeated blood draws.
Standardized calibration process: The calibration operation in this study is divided into external calibration and confirmatory internal calibration. Among them, external calibration is conducted once a week and is completed using the fresh extract of Escherichia coli ATCC25922: ①Prepare the fresh extract of Escherichia coli ATCC25922 as the calibration sample; ②Load the calibration sample and collect the mass spectrometry spectrum; ③The software automatically identifies the characteristic peaks; ④The system calculates and optimizes the calibration parameters based on the characteristic peaks; ⑤Save the calibration parameters to complete the calibration; ⑥Select the quality control sample for automatic spectrum acquisition, verify the calibration effect, and only after verification is qualified can sample testing be carried out.
Internal quality control and batch control strategies: ①For each batch of samples, blank controls (pure matrix), positive quality control samples (known ovarian cancer serum samples), and negative quality control samples (healthy control serum samples) are set up for each test. These samples are tested on the same plate and in the same process as the test samples, and the detection deviations within the batch are monitored. ②Cross-batch quality control samples are inserted between adjacent batches to verify the consistency of the tests between batches. The relative deviation of the peak area of the quality control spectrum should be less than 5%. Otherwise, the instrument should be recalibrated and the samples of this batch should be re-tested. ③After external calibration every week, three quality control samples are continuously tested. The retention time and peak intensity variation coefficient (CV) of the characteristic peaks of these samples should be less than 3% to ensure the stability of the instrument’s detection. ④All samples are randomly loaded. The testing personnel perform blind operations on the clinical grouping of the samples to minimize human and batch biases to the greatest extent.
The raw spectral data were preprocessed in Python using NumPy (1.21.5) 18 , SciPy (1.7.3) 19 , and scikit-learn (1.0.2) 20 . Signal smoothing was performed with a Savitzky–Golay filter (window size 21, polynomial order 10), chosen based on literature in MALDI-TOF/MS serum peptidomics and systematic validation of nine parameter combinations. This combination maximized correlation with raw spectra ( r > 0.99) for characteristic peaks (m/z = 2000–8000) and minimized intensity error (RMSE < 0.02). To prevent overfitting and preserve peak morphology, we followed the polynomial order < 1/2 window size criterion and applied a subsequent moving median filter (window size 10). Baseline correction was performed using morphological opening (white_tophat, structuring element size = 11). All spectra were visually reviewed with a standardized checklist to verify peak accuracy. A randomly selected subset (10%, n = 119) was independently reviewed by two researchers, achieving 98.3% consistency. Known matrix cluster regions (m/z = 1500–2000) were excluded a priori based on CHCA matrix characteristics and confirmed clear during visual inspection.
Significant m/z peaks were detected by applying adaptive thresholds (minimum height: 3% of maximum intensity; minimum width: 5 data points) via scipy.signal.find_peaks. Artifactual peaks in matrix cluster regions (m/z 1500–2000) were excluded through visual validation of raw spectra with overlaid markers. Based on the accurate mass calibration of external standards, an iterative greedy algorithm aligned peaks across samples with 2500 ppm mass tolerance, effectively compensating for instrumental drift (± 1500 ppm) and biological variation between samples. The resulting feature matrix (samples×peaks) represented relative peak areas, with missing values imputed as zero or linearly interpolated from adjacent m/z bins.
Normalization strategy for constructing the final data matrix involved two sequential steps: (1) Total Ion Current (TIC) normalization: each peak intensity in a sample was divided by the sample’s total ion current (sum of all peak intensities) and scaled by 1 × 10⁶ to standardize total signal intensity across samples, eliminating technical variation from sample loading differences; (2) median centering: each peak feature (column) was subtracted by its median value across all samples to reduce systematic bias and enhance model convergence. Post-normalization, intra-sample technical replicate correlation coefficients remained > 0.98, confirming effective control of technical variation.
Differential expression analysis of serum peptides for volcano plot construction was performed using a two-sample t-test, with the Benjamini-Hochberg method applied for multiple test correction. The adjusted significance thresholds were set as adjusted P-value 1 for screening differentially expressed peptides.
Employing Principal Component Analysis (PCA), Kernel Principal Component Analysis (KPCA), and t-Distributed Stochastic Neighbor Embedding (t-SNE), we performed dimensionality reduction on the dataset. Data I/O and numerical operations were managed with pandas (1.3.5) 18 , 20 , 21 and NumPy, respectively. Data distributions in the reduced space were visualized as density plots using matplotlib.pyplot (3.7.0) 19 and seaborn’s kdeplot (0.13.2) 21 to enable multi-faceted analysis.
Dataset partitioning takes precedence over all feature selection and modeling operations. The complete data set was divided into a training set (80%) and an independent test set (20%) by stratified sampling to ensure a consistent proportion of ovarian cancer patients to healthy controls in both groups. The independent test set was strictly isolated throughout the whole process and only used for the final model generalization performance evaluation, without participating in any feature selection, model training or hyperparameter optimization steps.
The training set was used for model training and evaluation using 5-fold stratified cross-validation. Eight classification algorithms were compared, including traditional methods and ensemble methods: KNN, SVM, Gaussian naive Bayes, Decision tree, Random forest, XGBoost, AdaBoost and LGBM (3.3.2) 18 . In each fold cross-validation, the training set was further divided into the in-fold training subset (90%) and the in-fold validation subset (10%). All feature selection operations were completed based on the in-fold training subset, without involving the in-fold validation subset or the independent test set. Model performance was quantified by AUC-ROC, accuracy, precision, recall, and F1 score.
To screen biologically meaningful peptides from proteomics data, three complementary computational methods were used to conduct exploratory assessment of feature importance, all done independently within each fold of 5-fold cross validation, based only on the in-fold training subset and without involving the in-fold validation subset or the independent test set: (1) Gini index was calculated by random forest and XGBoost algorithm (1.6.1) 21 to quantitatively evaluate the relative importance of each feature in the classification model. (2) SHAP (Shapley additive explanations) analysis (0.41.0) 21 , 22 was used to extract 100 samples as the background set based on the in-fold training subset, and the marginal contribution of each feature to the model prediction was estimated by KernelExplainer. (3) LIME (Local interpretable model-agnostic explanations, 0.2.0.1) 18 method was used to generate local explanations through LimeTabularExplainer for the typical samples of the training subset in the fold, which were summarized into a global feature ranking to screen the persistently important features in different samples. Feature importance analysis preferentially selected three consensus biomarkers, which were verified through ROC optimization and DCA.
Finally, the top 20 consensus features derived from the three methods were selected for subsequent validation, with any inconsistencies resolved by majority voting.
The key mass spectral peaks contributing to classifier performance were annotated via tandem mass spectrometry, enabling identification of corresponding peptides and proteins. This provided deeper molecular insights and biological context for the proteomic findings.
Subsequently, a relative abundance matrix incorporating the annotated biomarkers was constructed. Unsupervised clustering was performed using bisecting k-means and BIRCH algorithms to partition samples based on biomarker abundance patterns. To visualize the resulting clusters, Uniform Manifold Approximation and Projection (UMAP) was applied, yielding a low-dimensional representation that facilitated comparison between the annotated and original datasets. This comparison helped determine whether biological annotation revealed meaningful new data structures.
Subsequently, individual biomarkers underwent rigorous clinical validation. The optimal diagnostic thresholds were determined by optimizing the ROC curves using the Youden index, which maximizes overall diagnostic accuracy by balancing sensitivity and specificity. Furthermore, DCA was applied to evaluate the clinical utility of these biomarkers. DCA quantified the net benefit of their application in ovarian cancer diagnosis by integrating the trade-offs between true positive and false positive rates, thereby providing a clinically relevant assessment of their practical value.
To assess the impact of biological annotation on diagnostic performance, the machine-learning classifiers were applied to compare the classification efficacy of the original data versus the annotated biomarker matrix. Model performance was evaluated using the AUC-ROC, accuracy, precision, recall, and F1 score to determine whether biological annotation improved the discrimination between ovarian cancer and healthy samples. This analysis validated the utility of annotated biomarkers in enhancing ovarian cancer diagnostics.
Peptide separation was performed using a nanoACQUITY UPLC system (Waters, Milford, MA, USA) with a 60-minute acetonitrile gradient from 5% to 80%. The eluted peptides were then analysed on an Orbitrap Fusion Lumos mass spectrometer (Thermo Fisher Scientific, Waltham, MA, USA). Full-scan MS1 spectra were collected at a resolution of 100,000 across the m/z 200–2000 range. From each cycle, the top 10 precursor ions were selected for HCD fragmentation with 25% collision energy.
The resulting spectra were processed using Proteome Discoverer 2.5 and queried against the UniProt database (accessed December 4, 2024) 23 . Search parameters included a precursor mass tolerance of 20 ppm and a fragment mass tolerance of 1.0 Da. A target-decoy approach was applied to control the false discovery rate (FDR) below 1% for high-confidence peptide identification.
Two unsupervised clustering methods, bisecting K-means and BIRCH, were applied to analyse the peptide data. Bisecting K-means iteratively partitioned the dataset via recursive binary division to minimize intra-cluster variance. In parallel, BIRCH facilitated efficient large-scale clustering by building a clustering feature tree configured with a threshold of 0.5 and a branching factor of 50.
Cluster quality was quantitatively evaluated using the Rand Index and Adjusted Mutual Information, both computed against clinical annotations. The resulting groupings were visualized in a low-dimensional space generated by UMAP, aiding intuitive interpretation of cluster structures and supporting the identification of patterns significant for ovarian cancer diagnosis.
To comprehensively evaluate the biomarkers, data processing and numerical computations were performed using the pandas and numpy libraries. ROC analysis was conducted with the ROC curve and AUC functions from sklearn.metrics, generating false positive rate (FPR), true positive rate (TPR), corresponding thresholds, and UC. The ROC curves visually represent the discriminatory ability between groups, such as cancer versus healthy samples.
For threshold optimization, 100 equally spaced thresholds were generated using np.linspace. At each threshold, sensitivity (SEN), specificity (SPEC), and the Youden index (SEN + SPEC − 1) were calculated based on confusion matrices. The threshold maximizing the Youden index was selected as optimal and marked on the ROC plots to indicate the biomarker’s peak classification performance.
A 5-fold stratified cross-validation framework was applied to evaluate classifier performance. In each iteration, classifiers were trained on the training subset and prediction probabilities were generated for the validation subset.
Model net benefit:
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\:\mathrm{Net\:Benefit}\mathrm{=}\mathrm{TPR}\mathrm{-}\mathrm{FPR}{\times}\frac{{\tau}}{\mathrm{1}\mathrm{-}{\tau}}$$\end{document}
Where :
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\:\mathrm{TPR}\mathrm{=}\frac{\mathrm{TP}}{\mathrm{TP+FN}}\mathrm{\:(}\mathrm{True\:Positive\:Rate}\mathrm{)}$$\end{document}
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\:\mathrm{FPR}\mathrm{=}\frac{\mathrm{FP}}{\mathrm{TN+FP}}\mathrm{\:(False\:Positive\:Rate)}$$\end{document}
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\tau :{\text{ Treatment threshold }}({\mathrm{range}}:{\text{ }}0-{\mathrm{1}})$$\end{document}
Treat - all strategy:
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\:{\mathrm{Net\:Benefit}}_{\mathrm{all}}\mathrm{=}{\mathrm{TPR}}_{\mathrm{all}}\mathrm{-}{\mathrm{FPR}}_{\mathrm{all}}{\times}\frac{{\tau}}{\mathrm{1}\mathrm{-}{\tau}}$$\end{document}
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\:{\mathrm{TPR}}_{\mathrm{all}}\mathrm{=}\frac{\mathrm{TP+FN}}{\mathrm{Total\:Samples}}$$\end{document}
\documentclass[12pt]{minimal}
\usepackage{amsmath}
\usepackage{wasysym}
\usepackage{amsfonts}
\usepackage{amssymb}
\usepackage{amsbsy}
\usepackage{mathrsfs}
\usepackage{upgreek}
\setlength{\oddsidemargin}{-69pt}
\begin{document}$$\:{\mathrm{FPR}}_{\mathrm{all}}\mathrm{=}\frac{\mathrm{TN+FP}}{\mathrm{Total\:Samples}}$$\end{document}
Net benefit curves with 95% confidence intervals were generated using cross - validated results. Simpson’s method quantified the area where the model outperformed “treat - none” (net benefit>0). Single - feature thresholds were optimized based on Youden index and clinical relevance.
Conclusion
Machine learning models hold considerable promise for early disease diagnosis, yet their clinical translation faces challenges such as data heterogeneity, limited interpretability, and overfitting. Future efforts should focus on integrating multi-omics data, developing interpretable artificial intelligence frameworks, and conducting rigorous external validation across diverse prospective cohorts to improve model generalizability and clinical utility.
In this study, we established a streamlined workflow based on mass spectrometry-based proteomics to identify and validate serum biomarkers for early diagnosis of ovarian cancer. By combining multiple machine learning algorithms, we processed high-dimensional proteomic data and selected a panel of novel candidate biomarkers. This integrated machine learning model provides a new technical approach for proteomics research on ovarian cancer diagnosis.
We successfully identified a set of candidate biomarkers with potential diagnostic value that are closely associated with the pathogenesis of ovarian cancer. These findings provide novel candidate targets for ovarian cancer diagnostic research and offer new insights into the molecular mechanisms underlying ovarian carcinogenesis, and their value in early diagnosis warrants further validation. Consequently, this biomarker panel may serve as a key resource for developing novel therapeutic strategies and enhancing diagnostic precision in ovarian cancer patients.
Discussion
Our study provided a panel of three key serum peptide biomarkers (m/z = 4211.41, 2881.50, 2662.15) for ovarian cancer diagnosis. This discovery was achieved through a robust multi-model machine learning analysis of MALDI-TOF/MS data, which initially revealed 43 differentially expressed peptides. The final biomarker signature was consistently selected by multiple classifiers based on feature importance and validated via cross-validation, the selected biomarker combination demonstrated excellent diagnostic discrimination ability. Furthermore, we quantified feature contributions using both Shapley values and LIME. By aggregating these interpretability techniques with conventional feature importance scores, we implemented a multi-faceted optimization strategy for the model, leveraging their complementary strengths. We identified PDE11A-derived polypeptides as a potential diagnostic biomarker for ovarian cancer. At present, there is a lack of specific research on the PDE11A gene in ovarian cancer 27 . However, building on the established role of PDE11A in cell proliferation 28 , 29 and cAMP/cGMP signaling (closely related to cell proliferation, differentiation and apoptosis) 30 , we hypothesize that it may facilitate ovarian cancer cell survival and malignant transformation by modulating second messenger levels—a mechanistic insight that warrants further investigation. From a clinical perspective, a serum-based PDE11A assay, being minimally invasive and technically straightforward, could overcome the limitation of current imaging techniques (ultrasound/CT/MRI) in detecting microscopic lesions. Thus, PDE11A represents a candidate biomarker warranting further validation in independent prospective cohorts for ovarian cancer diagnosis, which is crucial for a disease often diagnosed at advanced stages due to its asymptomatic onset.
Ovarian cancer is characterized by asymptomatic onset and insidious progression, resulting in approximately 70% of patients being diagnosed at an advanced stage (III/IV), with a 5-year survival rate of less than 30% 31 . In sharp contrast, the 5-year survival rate of patients diagnosed at an early stage (I/II) dramatically exceeds 90% 32 , highlighting that early diagnosis is the most crucial breakthrough for improving prognosis. However, the diagnostic methods currently used in clinical practice and research have significant limitations. Serum Biomarkers, such as CA125 and Human Epididymis Protein 4 (HE4), offer the advantages of being minimally invasive, convenient, and cost-effective 33 . Nevertheless, their utility is hampered by poor specificity and insufficient sensitivity. CA125 can be elevated in benign conditions like endometriosis and pelvic inflammatory disease 34 , and approximately 50% of early-stage cases show normal CA125 levels. Even combined panels fail to detect all early-stage malignancies. Imaging Techniques, primarily transvaginal ultrasound (TVUS) and contrast-enhanced pelvic CT/MRI, provide direct visualization of ovarian morphology and structural abnormalities 35 . Yet, their effectiveness is limited by a low detection rate for small peritoneal metastases or early peritoneal metastases, an inability to reliably differentiate benign from malignant early-stage tumors, and operator-dependent variability that compromises reproducibility 36 . Magnetic resonance spectroscopy (MRS) can noninvasive-detect tissue metabolite profiles (such as choline, lipid, creatine, etc.), reflect the malignant characteristics of tumors through metabolic phenotype, and has shown unique advantages in the differentiation of benign and malignant ovarian lesions at early stage. However, its clinical application is still limited by high detection cost and low accessibility in primary medical institutions 37 , 38 . Genomic/Transcriptomic analyses, including circulating tumor DNA (ctDNA) and microRNA profiling, can reveal tumor-specific molecular signatures and hold promise for early warning 17 . However, their widespread clinical application is constrained by high costs, low concentrations of ctDNA in early-stage patients, and a lack of standardized analytical protocols 39 . Thus, further clarification is needed to define the minimum tumor size detectable through ctDNA analysis. Moreover, large prospective studies are needed to determine the clinical utility of ctDNA detection for early diagnosis of ovarian cancer and its impact on patient outcomes.
The integration of serum proteomics with antibody-independent diagnostic platforms presents a transformative approach for early cancer diagnosis 40 , effectively overcoming the limitations of conventional methods. This methodology enables the high-throughput profiling of the serum proteome to identify subtle alterations in protein expression during the initial stages of carcinogenesis, capturing signals that traditional single-marker assays (e.g., CA125, CEA) often miss 21 . Its antibody-free nature not only circumvents issues of antibody specificity and cross-reactivity but also facilitates the discovery of novel and previously undruggable protein/peptide biomarkers. From a clinical perspective, the assay requires only microliter volumes of serum, is non-invasive, and boasts a high-throughput format that allows for rapid processing and cost-effective large-scale screening. Furthermore, machine learning models integrating multi-dimensional proteomic features demonstrate superior sensitivity and specificity compared to single markers, minimizing false positives and negatives. The platform’s inherent adaptability across cancer types, such as lung, breast, and ovarian cancer, underscores its broad potential to provide a novel technical approach for ovarian cancer diagnosis, and its application value in early diagnosis and screening warrants further validation in large-sample, multi-center cohorts 16 .
MALDI-TOF/MS presents a compelling analytical platform for serum peptidomics and proteomics, characterized by its high sensitivity, high throughput, and minimal sample consumption 41 . These attributes make it ideally suited for detecting subtle protein alterations in ovarian cancer. The utility of this technique is evidenced by its established applications across diverse disease areas, including oncology. For instance, a study by Fudan University employing colloidal nanoparticles with MALDI-TOF/MS quantified free amino acids in serum from 60 serum samples. A machine learning-based classification model was constructed to distinguish patients with triple-negative breast cancer from the control group. The sensitivity, specificity and accuracy of this model were 95%, 100% and 97% respectively 10 . Researchers proposed that MALDI-TOF/MS technology can be applied to analyse lipids in urine samples and Bligh & Dyer method associated with the usage of HCCA matrix was found to be the most effective for lipidomics using MALDI-TOF/MS. The results showed that MALDI-TOF/MS analysis could distinguish healthy individuals and prostate cancer patients. Preliminary statistical models indicated that by using pre-selected mass spectrometry signals, the classification accuracy could reach 83.3% to 100.0% 6 .
Limitations
This study has the following limitations: (1) The single-center origin and limited sample size of the cohort, coupled with the absence of benign and other malignant gynecological disease controls, may constrain the model’s specificity and generalizability. Given the small sample size and high-dimensional proteomic data characteristics of this study, the near-perfect diagnostic performance obtained in the current cohort may be optimistic, and this excellent performance needs to be further verified and confirmed in multi-center, large-sample independent cohorts to avoid over-optimistic estimation of model generalization ability. Meanwhile, this study did not clearly report the distribution of cases by FIGO stage, did not provide the model performance data for stage stratification, and the control group only included age-matched healthy individuals, without including clinical confounding populations such as benign gynecological diseases. The current evidence is insufficient to support the application of the model for early screening. (2) The proteomic workflow lacks standardization, with variability in instrumentation, experimental protocols, and data processing potentially affecting cross-laboratory reproducibility. (3) The biological mechanisms of identified biomarkers (e.g., PDE11A) remain unvalidated, and insufficient integration with clinicopathological features limits the model’s interpretability and clinical applicability. (4) The dependence on specialized mass spectrometry and complex bioinformatics hinders translation into a rapid, cost-effective screening tool suitable for primary care settings. In response to the above issues, subsequent research will focus on the following directions: (1) Conduct multi-center prospective validation with expanded cohorts to assess technical variability and establish standardized analytical protocols. (2) Elucidate biomarker mechanisms through integrated multi-omics and functional studies, and improve model interpretability by incorporating key clinical variables. (3) Develop rapid, point-of-care detection assays based on identified biomarkers and evaluate their feasibility in pilot screening studies to enable clinical translation.
Introduction
Ovarian cancer remains a major global health challenge, ranking as the eighth most common malignancy and the leading cause of gynecological cancer-related mortality 1 . Due to its asymptomatic in early stages, the majority of cases are diagnosed at advanced phases (III/IV), which contributes to a five-year survival rate below 30% 2 . These statistics underscore the urgent need for improved early detection methods to enhance survival and reduce disease burden.
The current diagnosis of ovarian cancer predominantly relies on the assessment of imaging modalities, notably ultrasonography, MRI, CT, serum biomarkers, and physical examination. However, these methods are constrained by limited specificity and sensitivity, especially in early-diagnosis disease 3 . For example, though ultrasonography is widely used, it is operator-dependent, increasing risks of misdiagnosis. CA125 (Carbohydrate antigen 125) is not only inadequate for early detection but also elevated in benign gynecological conditions 4 . Thus, more reliable, non-invasive diagnostic tools are critically needed.
Proteomics-driven biomarker discovery represents a promising strategy for the early diagnosis of gynecologic malignancies 5 – 7 . However, traditional mass spectrometry encounters specific limitations in serum analyses: it often fails to detect low-abundance, tumor-related peptides, potentially missing differential peptides present at low concentrations; its typical resolution (< 50,000) is associated with a high false positive rate (FPR) in peptide identification, challenging the accuracy of sequence matching; and the lack of standardized methodologies globally leads to poor reproducibility across studies, impeding clinical translation 8 . Approaches that detect tumor-specific protein expression patterns may therefore overcome these constraints.
Matrix-assisted laser desorption/ionization time-of-flight mass spectrometry (MALDI-TOF/MS) is a widely used analytical technique in clinical microbiology and proteomics 9 . It can quickly determine the mass-to-charge ratio (m/z) of large molecules including proteins and peptides. The resulting spectral “fingerprints” can be aligned against a reference database to enable the identification of biomolecules, a property that makes them important for diagnostic applications in diseases such as breast cancer, and diabetes 10 , 11 . And it is especially suitable for high-throughput screening of serum peptide biomarkers for ovarian cancer in this study. However, MALDI-TOF/MS has obvious limitations in serum peptidomics analysis: its detection of low-abundance peptides is easily interfered by high-abundance serum proteins (e.g., albumin and immunoglobulin), and there is a peak overlap resolution issue for peptides with similar mass-to-charge ratios, and its performance is largely dependent on the quality of the peptide reference database and the standardization of sample processing. Advances in bioinformatics and machine learning are being increasingly used to address these limitations, thereby improving their accuracy and utility in peptide and protein studies 9 .
Artificial intelligence (AI) and machine learning (ML) have shown great potential in the diagnosis of ovarian cancer, especially when combined with proteomics data 12 , 13 . These tools excel at analyzing high-dimensional proteomics datasets and can identify low-abundance specific biomarker features that cannot be detected through traditional univariate analysis, thereby improving diagnostic accuracy 5 , 13 – 15 . The method that combines proteomics as the core with AI/ML analysis can systematically explore new biomarkers and therapeutic targets, addressing the limitations of single-analyte analysis. Moreover, with the support of AI/ML, the discovery of biomarkers based on proteomics has become a transformative approach for early cancer diagnosis 8 , 16 . Traditional proteomics based on mass spectrometry often faces challenges such as low-abundance peptide detection and the complexity of high-dimensional data, but AI/ML algorithms effectively solve these problems through feature extraction, dimensionality reduction, and pattern recognition. This synergy has been verified in multiple studies, where ML-enhanced proteomics analysis identified powerful biomarkers with high sensitivity and specificity for ovarian cancer and other malignant tumors. However, issues such as the quality of proteomics data (such as reproducibility, noise) and model interpretability still exist, and further research is needed to fully exploit its translational potential 17 . Our research addresses these gaps by combining serum proteomics based on MALDI-TOF/MS with an integrated ML pipeline, focusing on identifying consensus biomarkers for ovarian cancer early diagnosis.
Our study addresses the clinical challenges in the early diagnosis of ovarian cancer and proposes the following core hypotheses: MALDI-TOF/MS combined with multi-model machine learning strategy can screen differential peptide combinations from serum proteome with cross-model consistency, so as to construct a high-precision and non-invasive diagnostic model for ovarian cancer and overcome the limitation of single biomarker. This study aims to: (1) Screen differentially expressed peptides based on MALDI-TOF/MS data for biomarker discovery; (2) Evaluate and optimize the diagnostic performance of multiple machine learning models to determine the optimal model; (3) Identify core biomarker peptides that are stable across multiple algorithms; (4) To verify its clinical practicability and provide a feasible technical scheme for non-invasive and early diagnosis of ovarian cancer. This framework shows good potential for application in ovarian cancer detection, providing a potential clinical solution to overcome the limitations of single biomarkers.
Supplementary Material
Below is the link to the electronic supplementary material.
Supplementary Material 1
Supplementary Material 1
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.