Comprehensive serum glycopeptide spectra analysis to identify early-stage epithelial ovarian cancer.

OA: gold CC-BY-NC-ND-4.0
AI-generated summary by qwen3.7-flash, 2026-08-20

A convolutional neural network analyzing serum glycopeptide spectra with tumor markers differentiated early-stage epithelial ovarian cancer from controls and endometriomas, outperforming conventional biomarkers like CA125 and HE4.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by qwen3.7-flash, 2026-08-20 · read from full text

This study developed a comprehensive serum glycopeptide spectra analysis method to identify early-stage epithelial ovarian cancer by integrating machine learning with conventional tumor markers. The researchers analyzed 1712 enriched glycopeptides from over 1700 participants, including patients with various gynecologic conditions such as endometrioma, uterine myoma, and ovarian cysts, alongside healthy controls and those with other malignancies. The results demonstrated that specific glycosylation patterns could distinguish between benign conditions like endometriomas and clear cell carcinoma of the ovary, although the overall separation of early-stage cancer from non-cancer groups remained limited without advanced computational modeling. Relevance to endometriosis: Endometrioma is included in the control group as a benign differential diagnosis for ovarian masses, but the paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Epithelial ovarian cancer (EOC) is widely recognized as the most lethal gynecological malignancy; however, its early-stage detection remains a considerable clinical challenge. To address this, we have introduced a new method, named Comprehensive Serum Glycopeptide Spectral Analysis (CSGSA), which detects early-stage cancer by combining glycan alterations in serum glycoproteins with tumor markers. We detected 1712 glycopeptides using liquid chromatography-mass spectrometry from the sera obtained from 564 patients with EOC and 1149 controls across 13 institutions. Furthermore, we used a convolutional neural network to analyze the expression patterns of the glycopeptides and tumor markers. Using this approach, we successfully differentiated early-stage EOC (Stage I) from non-EOC, with an area under the curve (AUC) of 0.924 in receiver operating characteristic (ROC) analysis. This method markedly outperforms conventional tumor markers, including cancer antigen 125 (CA125, 0.842) and human epididymis protein 4 (HE4, 0.717). Notably, our method exhibited remarkable efficacy in differentiating early-stage ovarian clear cell carcinoma from endometrioma, achieving a ROC-AUC of 0.808, outperforming CA125 (0.538) and HE4 (0.557). Our study presents a promising breakthrough in the early detection of EOC through the innovative CSGSA method. The integration of glycan alterations with cancer-related tumor markers has demonstrated exceptional diagnostic potential.
Full text 52,525 characters · extracted from pmc-nxml · 6 sections · click to expand

Results

The expression of 1712 EGPs (Supplementally Information) was analyzed in 1713 individuals using several statistical methods, including Principal Component Analysis (PCA), Uniform Manifold Approximation and Projection (UMAP), heatmap, volcano plot, and Orthogonal Partial Least Squares Discriminant Analysis (OPLSDA). The participants comprised 564 with EOC and 1149 without, including healthy women ( n  = 943) and those with gynecologic conditions other than ovarian cancer ( n  = 206). The latter group included 83 with uterine myoma (LE), 72 with endometrioma (EM), and 51 with ovarian cysts (OCY), as detailed in Table 1 . PCA revealed a slight shift in the distribution of the EOC group compared with the non-EOC groups (Fig.  1 A), suggesting that EOC serum glycosylation was altered by cancer development. OPLSDA, aiming to discriminate between the EOC and non-EOC groups, slightly improved the separation of these groups compared with using PCA; however, the extent of this improvement remained limited and did not attain a practical level for screening (Fig.  1 B). The heatmap comprising 1712 EGPs and 1713 samples indicated that some serum glycosylation was overexpressed in the EOC group compared with the non-EOC group (Fig.  1 C). Furthermore, the volcano plot analysis comparing advanced-stage EOC (EOC-A) versus non-EOCs and early-stage EOC (EOC-E) vs. non-EOCs revealed that glycan alterations increased as EOC progressed (Fig. 1 D and E). Additionally, we utilized UMAP, a dimension reduction technique often employed to handle complex datasets, to visualize the glycopeptide profiles of EOC and non-EOC patients. The analysis revealed EOC-specific glycosylation changes, indicating disease-related alterations in glycoprotein structure. These findings highlight their potential utility as effective biomarkers for early detection (Fig.  1 F). Table 1 Demographic characteristics of the patients. Condition Age Number Classification Stage Modeling Test Epithelial ovarian cancer (EOC) 55.8 (± 12.2) 564 Serous (SE, 187) Stage I (276) Stage II (64) Stage III (192) Stage IV (32) ✓ ✓ Endometrioid (EN, 143) Clear cell (CCC, 180) Mucinous (MU, 39) Unclassified (UC, 15) Benign control (non-EOC) 54.2 (± 11.9) 1149 Endometrioma (EM, 72) ✓ ✓ Healthy women (HE, 943) Uterine myoma (LE, 83) Ovarian cyst (OCY, 51) Borderline ovarian tumor (BOT) 50.1 (± 13.9) 61 Serous (16) ✓ Endometrioid (4) Mucinous (39) Other class (2) Other Ovarian cancers 52.2 (± 14.9) 50 Mixed epithelial and Mesenchymal tumor (MEMT, 6) ✓ Sex cord–stromal tumor (SCST, 5) Germ cell tumor (GCT, 12) Metastatic cancer (ME, 7) Fallopian tube cancer (FTC, 9) Peritoneal cancer (PC, 11) Other cancers 52.6 (± 10.8) 19 Breast cancer (BC, 3) Stage II (2) Stage III (17) ✓ Gastric cancer (GC, 4) Colorectal cancer (CRC, 4) Head and neck cancer (HNC, 2) Hepatocellular carcinoma (HCC, 2) Lung cancer (LC, 2) Pancreas cancer (PC, 2) Total 54.5 (± 12.2) 1843 The numbers in the parentheses indicate the standard deviation of the age or number of participants. Checkmarks indicate that the samples were used for modeling or testing purposes. Figure 1 Comparison of EGP expressions between the EOC and non-EOC groups. The analyses included 1713 participants: 564 with EOC and 1149 without, comprising 943 healthy women (HE), 83 patients with uterine myoma (LE), 72 with endometrioma (EM), and 51 with ovarian cysts (OCY). ( A ) Score plot of the principal component analysis (PCA) of the EOC and non-EOC groups plotted based on the first and second components: Red: EOC, Blue: HE, Yellow: OCY, Pink: EM, and Green: LE. ( B ) Score plot of the OPLSDA model plotted using t1 and t01: Red: EOC, Blue: HE, Yellow: OCY, Pink: EM, and Green: LE. ( C ) Heatmap comprising 1712 EGPs and 1713 individuals: The EGPs were rearranged using cluster analysis. Red: relatively increased, Green: relatively decreased, and Black: not changed. ( D ) Volcano plot comparing the EOC-A (advanced-stage) and non-EOC groups, and ( E ) Volcano plot comparing the EOC-E (early-stage )and non-EOC groups: The vertical axis represents the − log10 ( p -value) from the Student’s t-test, and the horizontal axis represents the log2 (mean fold change). The cut-off criteria were set at a − log10 ( p -value) greater than 10 and a log2 (mean fold change) greater than 1 or less than − 1. ( F ) UMAP Analysis: The plots depict a UMAP analysis based on 1712 EGPs from five distinct groups. Each data point corresponds to an individual patient or healthy woman, and the contours illustrate the density of the distribution across all cases. Demographic characteristics of the patients. Stage I (276) Stage II (64) Stage III (192) Stage IV (32) Germ cell tumor (GCT, 12) Metastatic cancer (ME, 7) Fallopian tube cancer (FTC, 9) Peritoneal cancer (PC, 11) Stage II (2) Stage III (17) The numbers in the parentheses indicate the standard deviation of the age or number of participants. Checkmarks indicate that the samples were used for modeling or testing purposes. Comparison of EGP expressions between the EOC and non-EOC groups. The analyses included 1713 participants: 564 with EOC and 1149 without, comprising 943 healthy women (HE), 83 patients with uterine myoma (LE), 72 with endometrioma (EM), and 51 with ovarian cysts (OCY). ( A ) Score plot of the principal component analysis (PCA) of the EOC and non-EOC groups plotted based on the first and second components: Red: EOC, Blue: HE, Yellow: OCY, Pink: EM, and Green: LE. ( B ) Score plot of the OPLSDA model plotted using t1 and t01: Red: EOC, Blue: HE, Yellow: OCY, Pink: EM, and Green: LE. ( C ) Heatmap comprising 1712 EGPs and 1713 individuals: The EGPs were rearranged using cluster analysis. Red: relatively increased, Green: relatively decreased, and Black: not changed. ( D ) Volcano plot comparing the EOC-A (advanced-stage) and non-EOC groups, and ( E ) Volcano plot comparing the EOC-E (early-stage )and non-EOC groups: The vertical axis represents the − log10 ( p -value) from the Student’s t-test, and the horizontal axis represents the log2 (mean fold change). The cut-off criteria were set at a − log10 ( p -value) greater than 10 and a log2 (mean fold change) greater than 1 or less than − 1. ( F ) UMAP Analysis: The plots depict a UMAP analysis based on 1712 EGPs from five distinct groups. Each data point corresponds to an individual patient or healthy woman, and the contours illustrate the density of the distribution across all cases. The 2D barcodes that represent serum sugar chain expression profiles were generated using 1712 EGPs to allow the CNN to learn the features discriminating the EOC and non-EOC groups. Figure  2 A displays the representative patterns of the 2D barcodes for the five groups. We established the CNN model using 70% of the randomly selected samples (training set) and evaluated its performance using the remaining 30% of the samples (test set). This process was repeated 10 times, and the sum of the test results was assessed using receiver operating characteristic (ROC) curve analysis (Fig.  2 B). The area under the ROC curve (AUC) of the CNN model reached 0.924 (95% confidence interval [CI] 0.913–0.934) when patients with EOC-E (Stage I) were compared with those with non-EOC; this significantly surpassed the AUCs of the existing tumor markers, CA125 (0.842, 95% CI 0.812–0.871) and HE4 (0.717, 95% CI 0.674–0.759; Fig.  2 C). The ROC-AUC was 0.982 (95% CI 0.978–0.987) when patients with EOC-A were compared with those with non-EOC; this value also exceeded that of CA125 (0.960, 95% CI 0.946–0.974) and HE4 (0.911, 95% CI 0.884–0.937). We further converted the CNN-predicted values, which range from 0 to 1, into CSGSA scores using the following equation and set two cutoffs (3 and 6) to classify the examinees into three (high, middle, and low) risk groups. \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${\text{CSGSA}}\;{\text{score}} = - \log_{{{1}0}} \left( {{1}{-}{\text{CNN - predicted}}\;{\text{value}}} \right)$$\end{document} CSGSA score = - log 10 1 - CNN - predicted value Figure 2 Establishing and evaluating CSGSA CNN model for identifying early-stage EOC. ( A ) Representative 2D barcode images illustrating EGP expression and tumor markers (CA125 and HE49) within each disease category. ( B ) CNN model establishment and assessment methodology. The model was developed using a training set (70%) and subsequently evaluated on a separate test set (30%). This process was iterated 10 times, and the ROC-AUC was calculated using the sum of 10 trials. ( C ) ROC-AUCs of CA125, HE4, and CSGSA between EOC-E and non-EOC , and between EOC-A and non-EOC; the values in the charts represent AUCs. ( D ) Histograms based on the CSGSA scores for each disease group; cutoffs were set at 3 and 6 and patients were divided into three groups; high, middle, and low-risk. Sensitivity, specificity, and PPV are shown in Table 2 . ( E ) Correlation between CSGSA scores and tumor markers (CA125 and HE4): CA125 and HE4 values were logarithmically converted. ( F ) Gradient-weighted Class Activation Mapping (Grad-CAM) analysis: The portions that CNN used to discriminate EOCs are shown in red and yellow. ( G ) Histograms illustrating CSGSA scores for clear cell carcinoma (CCC), endometrioma (EN), mucinous (MU), and serous (SE) in the EOC-A and EOC-E groups: Scores are classified as high, middle, or low using cutoff values of 3 and 6. The number of patients is indicated in parentheses. ( H ) Correlations between CSGSA scores and the tumor markers CA125 and HE4 for CCC, EN, MU, and SE cases in the EOC-A and EOC-E groups. CA125 and HE4 values were logarithmically converted. Establishing and evaluating CSGSA CNN model for identifying early-stage EOC. ( A ) Representative 2D barcode images illustrating EGP expression and tumor markers (CA125 and HE49) within each disease category. ( B ) CNN model establishment and assessment methodology. The model was developed using a training set (70%) and subsequently evaluated on a separate test set (30%). This process was iterated 10 times, and the ROC-AUC was calculated using the sum of 10 trials. ( C ) ROC-AUCs of CA125, HE4, and CSGSA between EOC-E and non-EOC , and between EOC-A and non-EOC; the values in the charts represent AUCs. ( D ) Histograms based on the CSGSA scores for each disease group; cutoffs were set at 3 and 6 and patients were divided into three groups; high, middle, and low-risk. Sensitivity, specificity, and PPV are shown in Table 2 . ( E ) Correlation between CSGSA scores and tumor markers (CA125 and HE4): CA125 and HE4 values were logarithmically converted. ( F ) Gradient-weighted Class Activation Mapping (Grad-CAM) analysis: The portions that CNN used to discriminate EOCs are shown in red and yellow. ( G ) Histograms illustrating CSGSA scores for clear cell carcinoma (CCC), endometrioma (EN), mucinous (MU), and serous (SE) in the EOC-A and EOC-E groups: Scores are classified as high, middle, or low using cutoff values of 3 and 6. The number of patients is indicated in parentheses. ( H ) Correlations between CSGSA scores and the tumor markers CA125 and HE4 for CCC, EN, MU, and SE cases in the EOC-A and EOC-E groups. CA125 and HE4 values were logarithmically converted. Figure  2 D shows the histograms arranged based on the CSGSA scores for each disease category. When the cutoff was set at “3,” the sensitivity of EOC-A and EOC-E was 86.3% and 56.2%, respectively. The specificities of EM, HE, uterine myoma (LE), and ovarian cyst (OCY) were 79.6%, 99.5%, 91.9%, and 94.0%, respectively (Table 2 a). To estimate the positive predictive value (PPV), the number of cases was corrected for the prevalence of each disease. The prevalence rates of EOC, EM, LE, and OCY in over 40-year-old patients (possible examinees who will undergo the screening test) were assumed to be 30, 1000, 2000, and 1000 among 100,000 women, respectively 3 , 40 – 43 , and the ratio of EOC-E and EOC-A was assumed to be 1:1. Following the correction, the PPV was found to be 2.23% when the cutoff was set at “3” (Table 2 b). The reason behind the low PPV, despite the high sensitivity and specificity, can be attributed to an increase in the number of false positives owing to the leverage effect. This effect is particularly observed when the prevalence rate of the target disease is low. When the cutoff was increased to “6,” the total EOC sensitivity decreased to 52.7%; however, the specificities were remarkably improved to 93.1% (EM), 100.0% (HE), 95.3% (LE), and 98.2% (OCY); the PPV reached 8.03% (Table 2 c and 2d). When we assigned patients to three classes using the two cutoffs (3 and 6), 0.20% of the examinees were classified into the high-risk class (> 6), with a PPV of 8.03%; 0.76% were classified into the middle-risk class (3–6), with a PPV of 0.734%; and over 99% were classified into the low-risk class (< 3), with a PPV of 0.00872% (Table 2 e). At the cutoff value of 6, 52.5% of the patients with total EOC were identified, and at the cutoff value of 3, 56.2% of patients with EOC-E (71.2% of total EOC) were identified. Although there was a slight positive correlation between the CSGSA scores and CA125 or HE4, it is evident that CSGSA accurately identified the patients with EOC, even when their CA125 and HE4 levels were low (Fig.  2 E). In addition, Gradient-weighted Class Activation Mapping (Grad-CAM) analysis was performed to identify the key areas in images in which the CNN recognizes EOC. The CNN recognized the peripheral area of the 2D barcode (Fig.  2 F), which was not surprising because EGP expression was allocated in the order of PCA loadings; the EGPs with high variation were positioned at the peripheral area. Table 2 Evaluation of the CSGSA screening model (CNN). CSGSA (screening model) Sum Sens. or Spec Positive Negative (a) Cutoff = 3 (low), test set EOC Advanced 736 117 853 86.3% 71.4% Early 470 367 837 56.2% Non-EOC EM 44 172 216 79.6% 97.4% HE 15 2794 2809 99.5% LE 21 237 258 91.9% OCY 10 157 167 94.0% Sum 1296 3844 5140 (b) Cutoff = 3 (low), corrected based on prevalence EOC Advanced 13 2 15 86.3% 71.2% Early 8 7 15 56.2% Non-EOC EM 204 796 1000 79.6% 99.1% HE 512 95,458 95,970 99.5% LE 163 1837 2000 91.9% OCY 60 940 1000 94.0% Sum 960 99,040 100,000 PPV or NPV 2.23% 99.99% ( c ) Cutoff = 6 (high), test set EOC Advanced 617 236 853 72.3% 52.7% Early 274 563 837 32.7% Non-EOC EM 15 201 216 93.1% 99.1% HE 0 2809 2809 100.0% LE 12 246 258 95.3% OCY 3 164 167 98.2% Sum 921 4219 5140 (d) Cutoff = 6 (high), corrected based on prevalence EOC Advanced 11 4 15 72.3% 52.5% Early 5 10 15 32.7% Non-EOC EM 69 931 1000 93.1% 99.8% HE 0 95,970 95,970 100.0% LE 93 1907 2000 95.3% OCY 18 982 1000 98.2% Sum 196 99,804 100,000 PPV or NPV 8.03% 99.99% Cutoff Predicted population PPV Sensitivity (EOC all) Sensitivity (EOC early) (e) CSGSA capability High risk  > 6 0.196% 8.03% 52.5% 32.7% Middle risk 3–6 0.764% 0.734% 18.7% 23.4% Low risk  < 3 99.0% 0.00872% 28.8% 43.9% CSGSA: comprehensive serum glycopeptide spectra analysis; CNN: convolutional neural network; PPV: positive predictive value; NPV: negative predictive value; EM: endometrioma; HE: healthy women; LE: uterine myoma; OCY: ovarian cyst. Significant values are in bold. Evaluation of the CSGSA screening model (CNN). CSGSA: comprehensive serum glycopeptide spectra analysis; CNN: convolutional neural network; PPV: positive predictive value; NPV: negative predictive value; EM: endometrioma; HE: healthy women; LE: uterine myoma; OCY: ovarian cyst. Significant values are in bold. When EOC was divided into four major histological types, CCC, endometrioid (EN), mucinous (MU), and serous (SE), the analysis revealed that the sensitivities (with a cutoff of “3”) of EOC-A remained relatively high regardless of the histological type (CCC: 78.2%, EN: 86.6%, MU: 79.5%, and SE: 89.4%). However, in the case of EOC-E, the sensitivities varied depending on the specific histological types (CCC: 50.0%, EN: 65.3%, MU: 39.7%, and SE: 73.7%; Fig.  2 G). This variability in sensitivity may be attributed to the relatively lower levels of CA125 and HE4 in patients, particularly those with MU and CCC, compared with those with EN and SE (Fig.  2 H). We evaluated whether CSGSA could distinguish EOC from other cancers, assuming that multiple patients with various diseases would undergo this screening test. Borderline ovarian tumors (BOT), often referred to as low malignant potential tumors, constitute 10–20% of all EOC cases and tend to occur in younger individuals. While noninvasive, BOT has the potential to progress into cancer. To investigate how CSGSA classifies BOT, we recruited 61 patients with BOT, including those with SE ( n  = 16), MU ( n  = 39), EN ( n  = 4), and other histological types ( n  = 2). Based on the CSGSA scores, 15 of the 61 patients with BOT (24.6%) were considered high risk, 16 (26.2%) were middle risk, and the remaining 30 (49.2%) were low risk (Fig.  3 A). Although CSGSA did not provide a clear distinction between EOC and BOT, it may indicate the potential for BOT malignancy because some BOT cases can progress into EOC. Figure 3 CSGSA selectivity against other cancers. ( A ) Histogram of the CSGSA scores for borderline ovarian tumors (BOT). ( B ) Histograms of the CSGSA scores for ovarian cancers, except for EOC. ( C ) CSGSA scores for other general cancers. ( D ) UMAP analysis of BOT, other ovarian cancer (OVC), and other cancers. The number of patients is indicated in parentheses. ( E ) Correlation between CSGSA scores and the tumor markers CA125 and HE4 for other ovarian cancers, other general cancers, and BOT. CSGSA selectivity against other cancers. ( A ) Histogram of the CSGSA scores for borderline ovarian tumors (BOT). ( B ) Histograms of the CSGSA scores for ovarian cancers, except for EOC. ( C ) CSGSA scores for other general cancers. ( D ) UMAP analysis of BOT, other ovarian cancer (OVC), and other cancers. The number of patients is indicated in parentheses. ( E ) Correlation between CSGSA scores and the tumor markers CA125 and HE4 for other ovarian cancers, other general cancers, and BOT. To evaluate the ability of CSGSA to differentiate EOC from other types of ovarian cancer, we enrolled 50 patients diagnosed with ovarian cancers other than EOC. The ovarian cancer group comprised patients with mixed epithelial and mesenchymal tumors (MEMT, n  = 6), sex cord–stromal tumors (SCST, n  = 5), germ cell tumors (GST, n  = 12), fallopian tube cancer (FTC, n  = 9), peritoneal cancer (PC, n  = 11), and metastatic cancer (ME, n  = 7), which had metastasized to the ovary from primary tumors in other organs such as the stomach and large intestine. Our findings revealed that a high percentage of patients with MEMT (5 out of 6, 83.3%), PC (10 out of 11, 90.9%), and FTC (6 out of 9, 66.7%) were classified as middle-risk or higher (Fig.  3 B). These outcomes were consistent with expectations because MEMT encompasses both epithelial and mesenchymal components, PC exhibits histological similarities to serous ovarian carcinoma, and the majority of FTC cases are classified as epithelial cancers. Furthermore, it was reasonable that all patients with SCST (5 out of 5) were classified as low-risk because SCST has distinct histological characteristics compared with EOC. The patients with GCT (8 of 12, 66.7%), and ME (4 of 7, 57.1%) were also classified as middle-risk or higher (Fig.  3 B). Next, we conducted a study with 19 patients diagnosed with different types of cancer, including breast cancer (BC, n  = 3), gastric cancer (GC, n  = 4), colorectal cancer (CRC, n  = 4), head and neck cancer (HNC, n  = 2), liver cancer (HCC, n  = 2), lung cancer (LC, n  = 2), and pancreatic cancer (PC, n  = 2). We observed that 6 of 19 (31.6%) patients, particularly those with CRC, LC, GC, and PC, were classified as middle-risk or higher (Fig.  3 C). Considering that all these cancers were in advanced stages (BC, GC, CRC, NHC, HCC, LC: stage III, PC: stage II), these findings suggest that CSGSA has a certain level of selectivity in distinguishing EOC from other types of cancer. However, further improvements are required to enhance the discriminatory capacity, particularly when the prevalence rates of these cancers are much higher than those of EOC. When comparing the UMAP distributions of the patients with those of EOC, as illustrated in Fig.  1 C, no significant differences were observed in these patterns (Fig.  3 D). Furthermore, when we examined the correlation between CSGSA scores and CA125 or HE4 in these patients, we observed a strong correlation between CSGSA scores and CA125 in ovarian cancers such as GCT, ME, MEMT, and SCST; however, limited correlations were observed in the other cases (Fig.  3 E). We selected several EGPs located in the upper-right and upper-left positions of the volcano plot and compared their retention times, mass spectra, and MS/MS patterns with those of the glycopeptides obtained from purified human serum proteins. We identified four glycopeptides as follows: alpha-1-acid glycoprotein (AGP) with a fully sialylated tetraantennary glycan containing one fucose attached to the asparagine at position 103, haptoglobin with a fully sialylated triantennary glycan containing three fucoses attached to the asparagine at position 241, complement C9 with a fully sialylated biantennary glycan attached to the asparagine at position 415, and transferrin with a biantennary glycan containing one sialic acid attached to the asparagine at position 432 (Fig.  4 ). However, the detailed sugar chain structure, including sequence and linkage manner, remains unknown. The diagnostic performance of each glycopeptide, as well as the combined model of CA125, HE4, and the four glycopeptides, was evaluated using ROC curves and AUC values. The combination model was established through logistic regression. As illustrated in Supplementary Fig. S1 , the performance of individual markers was slightly lower than that of CA125 alone. Although the combined model of CA125, HE4, and the four glycopeptides significantly improved performance, it did not exceed that of the CSGSA. These four glycopeptides were located in the periphery of the 2D barcodes, suggesting that they play a vital role in discriminating EOC, when considered in conjunction with the results of Grad-CAM analysis (Fig.  2 F). Figure 4 Identification of the glycopeptides contributing to EOC discrimination. Volcano plot displaying 1712 EGPs when comparing all EOC and non-EOC samples (center); the horizontal axis represents log 2 (mean fold ratio) and the vertical axis represents log 10 (Student’s t -test p -value). The identified glycopeptides are marked with red circles. Extracted ion chromatogram (EIC), single mass spectrum (MS), and MS/MS spectrum (MSMS) of the glycopeptides obtained from purified standard serum proteins (in black) and EOC serum (in red) are presented. Proposed proteins, peptide sequences, and sugar chain structures are illustrated in each corner. The positions of the identified glycopeptides on the 2D barcode, along with bar graphs illustrating their relative abundance in EOC and non-EOC groups, are presented. Identification of the glycopeptides contributing to EOC discrimination. Volcano plot displaying 1712 EGPs when comparing all EOC and non-EOC samples (center); the horizontal axis represents log 2 (mean fold ratio) and the vertical axis represents log 10 (Student’s t -test p -value). The identified glycopeptides are marked with red circles. Extracted ion chromatogram (EIC), single mass spectrum (MS), and MS/MS spectrum (MSMS) of the glycopeptides obtained from purified standard serum proteins (in black) and EOC serum (in red) are presented. Proposed proteins, peptide sequences, and sugar chain structures are illustrated in each corner. The positions of the identified glycopeptides on the 2D barcode, along with bar graphs illustrating their relative abundance in EOC and non-EOC groups, are presented. Because CCC can potentially arise from EM, distinguishing these two conditions is another important issue. We established a CSGSA CCC discrimination model using the OPLSDA approach instead of a CNN because the number of specimens (CCC: 180 and EM: 76) was too small to establish a robust CNN model. OPLSDA enables the identification of discriminating factors between two classes by iteratively reducing non-discriminable dimensions, ultimately uncovering an underlying factor that separates the two groups in a 1712-dimensional space. This single dimension can then be used to effectively discriminate between CCC and EM. First, we divided the samples into three groups and established an OPLSDA model using two groups as a training set. Then, we evaluated its performance using the remaining group as a test set (Fig.  5 A). This process was repeated three times by changing the combinations. Finally, the test results were summed to calculate the ROC-AUC, sensitivity, and specificity. The discrimination patterns of the OPLADA model in the three trials were similar, and the distributions in the training sets were reproduced in the test sets (Fig.  5 B). The ROC-AUC calculated using the OPLSDA t1 score (horizontal axis value) of the test sets reached 0.808 (95% CI 0.746–0.870), when CCC-E (early stage) was compared with EM, which significantly exceeded that of CA125 (0.538, 95% CI 0.454–0.622) and HE4 (0.557, 95% CI 0.476–0.638). When comparing CCC-A (advanced stage) with EM, the ROC-AUC reached 0.927 (95% CI 0.883–0.971), whereas that of CA125 and HE4 was 0.670 (95% CI 0.579–0.761) and 0.718 (95% CI 0.625–0.811), respectively (Fig.  5 C). The sensitivities were 73.4% and 43.1% when CCC-A and CCC-E were compared with EM, respectively, and the specificity of EM was 95.8% when the cutoff was set at zero. When these numbers were corrected with the prevalence rates, the PPV and sensitivity were 11.2% and 52.2%, respectively. Figure 5 Evaluating the CSGSA OPLSDA model for distinguishing CCC from EM. ( A ) Method to establish and assess the OPLSDA model: The samples were randomly divided into three groups; the OPLSDA model was established using two groups as a training set. The remaining group was used as a test set. ( B ) OPLSDA score plots of the training and test sets for three trials. Red: CCC and Green: EM. ( C ) Box plot and ROC-AUC values between CCC-E and EM, and between CCC-A and EM for CA125, HE4, and CSGSA; CA125 and HE4 values were logarithmically converted. The p -values of Student’s t -test for between CCC-E and EM and between CCC-A and EM were calculated. ( D ) Proposed scheme of CSGSA: 1712 EGPs, obtained through proteolysis, were subjected to liquid chromatography–mass spectrometry analysis. The resulting data, along with CA125 and HE4 values, were used to generate 2D barcodes. A pretrained CNN model was used to interpret the 2D barcode pattern to classify whether it corresponds to EOC. The OPLSDA model was used to differentiate between CCC and EM based on the expression patterns of these 1712 EGPs. Evaluating the CSGSA OPLSDA model for distinguishing CCC from EM. ( A ) Method to establish and assess the OPLSDA model: The samples were randomly divided into three groups; the OPLSDA model was established using two groups as a training set. The remaining group was used as a test set. ( B ) OPLSDA score plots of the training and test sets for three trials. Red: CCC and Green: EM. ( C ) Box plot and ROC-AUC values between CCC-E and EM, and between CCC-A and EM for CA125, HE4, and CSGSA; CA125 and HE4 values were logarithmically converted. The p -values of Student’s t -test for between CCC-E and EM and between CCC-A and EM were calculated. ( D ) Proposed scheme of CSGSA: 1712 EGPs, obtained through proteolysis, were subjected to liquid chromatography–mass spectrometry analysis. The resulting data, along with CA125 and HE4 values, were used to generate 2D barcodes. A pretrained CNN model was used to interpret the 2D barcode pattern to classify whether it corresponds to EOC. The OPLSDA model was used to differentiate between CCC and EM based on the expression patterns of these 1712 EGPs. The proposed scheme for EOC screening using CSGSA is illustrated in Fig.  5 D. In total, 1712 EGPs derived from protein proteolysis were analyzed using liquid chromatography–quadrupole time-of-flight mass spectrometry (MS), and their intensities were transformed into 2D barcodes incorporating the CA125 and HE4 values. These barcodes were then input into a pretrained CNN model, which discriminates EOC from non-EOC. Additionally, the OPLSDA model was used to differentiate CCC from EM based on the expression patterns of the 1712 EGPs. To assess the practicality of CSGSA as a screening test, we conducted several evaluations, including (i) reproducibility, (ii) inter-machine error, (iii) inter-operator error, (iv) effects of circadian rhythm and diet, (v) short-term stability in blood, and (vi) long-term storage stability in serum. For reproducibility assessment, we performed analysis on both HE and EOC samples, repeating the process 10 times for each sample. The results showed that the variations in the CSGSA scores were consistently within a range of ± 1 level (Supplementary Fig. S2 A). To evaluate inter-machine and inter-operator error, two operators analyzed 10 HE samples and 10 EOC samples (comprising 5 EOC-E and 5 EOC-A) twice using different mass spectrometers. The results indicate that the level of the errors remained within an acceptable range (Supplementary Fig. S2 B and S2 C). In terms of circadian and dietary influences, there was not a significant change in the levels of glycopeptide expression between samples collected at 8 AM and 2 PM, as well as between 8 AM and 9 PM, for all four participants (R 2  > 0.991, Supplementary Fig. S3 A). The CSGSA scores remained stable, except for one spike observed in Woman A at 9 PM (Supplementary Fig. S3 B). Additionally, the scattering of the OPLSDA spots was limited to each individual (Supplementary Fig. S3 C). In the analysis of short-term stability of blood at room temperature (22–25 °C), glycopeptide expression did not show any noticeable changes within 6 h, 24 h (R 2  > 0.937, Supplementary Fig. S4 A, S4 B and S4 C), and remained stable even within a period of 3 days (Supplementary Fig. S4 D). In the analysis for long-term storage stability in serum, the glycopeptides were stable both within 3 days and 1 month at − 20 °C and 4 °C (R 2  > 0.946, Supplementary Fig. S5 and S6 ). We did not observe noticeable differences in CSGSA outcomes among different facilities (Supplementary Fig. S7 ), demonstrating that the CSGSA test is robust enough to yield consistent results regardless of the variations in blood collection methods or tubes used across different locations.

Material

This was a retrospective, observational, case–control study. All patients and healthy volunteers who visited 13 facilities during a designated period were eligible for inclusion. Random sampling and blinding were not implemented. Under a framework assuming an alpha error of 0.05 and a beta error of 0.2, considering that the average expression levels of glycopeptides would differ by no more than half of the standard deviations between the EOC and non-EOC groups, the minimum sample size was estimated to be approximately 100. Detailed estimations are described in the Supplementary Information. Furthermore, given that we prepare the training and test sets and the test set constitutes 30% of all, the ideal sample size of EOC cases was estimated to be three times larger than the minimum size, which is more than 300. In this study, 564 serum samples were collected from patients with EOC at the time of ovarian mass detection and before any treatment or surgery. The sera were obtained from Tokai University Hospital (Kanagawa, Japan), Kanagawa Cancer Center (Kanagawa, Japan), Kyorin University Hospital (Tokyo, Japan), National Cancer Center Japan (Tokyo, Japan), St. Marianna University Hospital (Kanagawa, Japan), Tsukuba University Hospital (Ibaraki, Japan), Tohoku Medical Megabank Organization (Miyagi, Japan), Fujita Health University Hospital (Aichi, Japan), Yokohama City University Hospital (Kanagawa, Japan), LSI Medience Corporation (Tokyo, Japan), SOIKEN (Osaka, Japan), KAC (Kyoto, Japan), and Sanfco (Tokyo, Japan), with informed consent from each patient or volunteer. The study-specific exclusion criteria included individuals with (i) a combination of several types of cancers; (ii) a history of hormonal drug administration owing to a malignant tumor, autoimmune disease, or thyroid abnormality; (iii) abnormal blood test results (white blood cell count > 9600/μL; platelet count > 48 × 10 4 /μL; lactate dehydrogenase level > 263 units/L; hemoglobin level  3.0 mg/dL); (iv) renal dysfunction; and (v)liver dysfunction. We followed the International Federation of Gynecology and Obstetrics guidelines 60 when diagnosing EOC and for stage classification. Blood samples were collected via venous puncture before surgery or treatment. The sera were separated from blood cells through centrifugation and stored at − 80 °C until analysis. All samples were only analyzed once because we had previously validated their accuracy and reproducibility. The Institutional Review Board (IRB) of Tokai University approved the use of patient clinical information and serum/tumor samples (IRB registration number: 09R-082; September 17, 2009). All research methods and procedures were conducted in strict accordance with the ethical principles outlined in the Declaration of Helsinki, as well as in full compliance with relevant institutional, national, and international guidelines. CA125 and HE4 tests were performed using the chemiluminescent immunoassay (CLIA) method in LSI Medience corporation, the clinical testing laboratory in Japan. The preparation and analysis were performed as described previously 33 , 36 , 39 . Briefly, EGPs were analyzed by liquid chromatography quadrupole time-of-flight mass spectrometer (HP1200 + 6520, Agilent Technologies, Palo Alto, CA, USA) equipped with a C18 column (Inertsil ODS-4, 2 μm, 100 Å, 100 mm × 1.5 mm ID, GL Science, Tokyo, Japan). The EGPs were eluted at a flow rate of 0.15 mL/min at 40 °C with the following gradient program: 0–7 min, 15–30% mobile phase B; 7–12 min, 30–50% mobile phase B; and 2 min hold at 100% mobile phase B, where mobile phase A was 0.1% formic acid in water and mobile phase B was 0.1% formic acid in 9.9% water and 90% acetonitrile. The mass spectrometer was set to operate at a capillary voltage of 4000 V in negative ion mode, specifically to enhance the detection sensitivity of glycopeptides containing sialic acids. All samples were analyzed only once, as their accuracy has already been confirmed, as shown in the Supplementary Information. A detailed description of the data processing has been reported previously 33 , 36 . Briefly, liquid chromatography–MS raw data were converted to CSV-type data using Mass Hunter Export (Agilent Technologies); the peak positions (retention times and m/z) and peak areas were obtained using R software (R 3.2.2, R Foundation). Marker Analysis, a proprietary software provided by LSI Medience Corporation, aligned all the peak areas, reduced noise, and corrected errors 33 . The tolerance of m/z and retention time was set at 0.06 Da and 0.3 min, respectively, during alignment and peak assignment. The relative expression of 1712 EGPs was obtained by calculating the ratio of expression between each sample and the quality control. The peak list of 1712 EGPs are shown in Supplementary Information. EGP expressions were allocated to 41 × 42 cells of an Excel sheet based on the order of PCA loadings. The method used for generating 2D barcodes and training a CNN model was based on our previous study 39 . MATLAB and its deep learning toolbox (R2022b, Mathworks, Natick, MA, USA) were used for training and testing the CNN model to identify EOCs. AlexNet, a CNN algorithm, was introduced as a Mathworks add-in program. We removed the last three layers of the pretrained AlexNet and added a new fully connected layer with neurons for the number of classes, followed by the softmax and classification layers. A mini-batch size of 20, maximum epochs of 30, squared gradient decay factor of 0.9995, and adaptive moment estimation (Adam) were used as a learning method. The sample was randomly divided into two sets: training (70%) and testing (30%). After training the CNN with the training dataset, the accuracy of the CNN model was evaluated using the test set. This process was repeated 10 times. The results obtained from each trial were summed, and the ROC-AOC was calculated for the sum of 10 trials. The CNN score obtained from zero (low risk) to one (high risk) was further converted into the CSGSA score using the following equation: \documentclass[12pt]{minimal} \usepackage{amsmath} \usepackage{wasysym} \usepackage{amsfonts} \usepackage{amssymb} \usepackage{amsbsy} \usepackage{mathrsfs} \usepackage{upgreek} \setlength{\oddsidemargin}{-69pt} \begin{document}$${\text{CSGSA}}\;{\text{score}} = - \log_{10} \left( {{1}{-}{\text{CNN}}\;{\text{score}}} \right)$$\end{document} CSGSA score = - log 10 1 - CNN score If the CNN score was over 0.99999999, a CSGSA score of 8 was assigned automatically. We compared the retention times, single mass spectra, and MS/MS patterns of target glycopeptides with those of the glycopeptides obtained from purified human serum proteins. The purified human serum proteins, including alpha-1-acid glycoprotein, complement C8, complement C9, complement factor H, fibrinogen, haptoglobin, alpha-2-macroglobulin, and transferrin, were purchased from Sigma-Aldrich (St. Louis, MO, US). These proteins were digested according to the method described above and subsequently analyzed along with the glycopeptides obtained from patients with EOC. When the retention times, single mass spectra, and MS/MS patterns of the glycopeptides obtained from EOC serum matched those obtained from the purified protein standards, the molecular weights of multiple glycopeptides were calculated, considering all the various possible glycan combinations. The structure was confirmed when the theoretical molecular weight of the glycopeptide matched the observed molecular weight in the range of 0.03 Da. SIMCA software (version 13.0.3; Umetrics; Umeå, Sweden) was used to analyze the OPLSDA model for differentiating CCC from EM. First, we randomly divided the samples into three groups and established an OPLSDA model using two groups as a training set. Then, we evaluated its performance using the remaining group as the test set. This process was repeated three times by changing the combinations. Finally, we calculated the ROC-AUC, sensitivity, and specificity using the test results. OPLSDA showed two-dimensional differentiation using the t1 and t01 values. We defined the t1 values as the CSGSA scores using box–whisker plots and ROC analysis. To evaluate circadian and dietary influences, we obtained sera from four healthy female volunteers (women A, B, C, and D) at 8 AM, 2 PM, and 9 PM. They fasted from 9 PM the previous day to 8 AM the next day, followed by meals at 9 AM, 12 PM, and 7 PM. To evaluate the short-term stability of blood, blood samples collected from five healthy female volunteers were allowed to rest in a blood tube at room temperature for 1, 6, and 24 h; the serum was separated from the blood clot through centrifugation for 10 min at 2000 ×  g . To evaluate the long-term storage stability of serum, we stored the five serum samples at 4 °C and − 20 °C in a refrigerator/freezer for 3 days or 1 month. All samples were analyzed using mass spectrometry, and EGP levels were compared. The samples were evaluated using the CNN and OPLSDA models. To assess reproducibility, we performed analysis on both HE and EOC samples, repeating the process 10 times for each sample. Additionally, to evaluate inter-machine and inter-operator errors, two operators analyzed 10 HE samples and 10 EOC samples (comprising 5 EOC-E and 5 EOC-A) two times using different mass spectrometers. All glycopeptide expression levels (intensities obtained via mass spectrometry) were normalized to QC values to adjust for inter-day variability. Additionally, CA125 and HE4 values were log-transformed before conducting the Student's t-test. Although not all parameters conform to a normal distribution, for simplicity, we assumed normal distribution for all parameters. Furthermore, all parameters were converted to Z scores prior to performing PCA and OPLSDA analyses. The CNN model was developed using MATLAB and its deep learning toolbox (R2022b, Mathworks, Natick, MA, USA). SIMCA software (version 13.0.3; Umetrics) was used to compare CCC and EM and establish the OPLSDA model. Missing data values, which were mostly below the detection limit, were substituted with zero. The Statistical Package for Social Sciences (SPSS, version 27.0, Chicago, IL, USA) and proprietary statistical software 33 were used to perform all other statistical analyses.

Discussion

In this study, we aimed to develop a novel screening method to enable the early detection of EOC, which integrated the cancer-specific tumor markers and glycan alterations. We classified patients into three risk classes using two cutoffs, we found that 0.20% of individuals belonged to the high-risk class with a PPV of 8.03%; 0.76% belonged to the middle-risk class with a PPV of 0.734%; and over 99% belonged to the low-risk class with a PPV of 0.00872%. To the best of our knowledge, few commercially available blood cancer screening methods report their PPVs, and the reported PPVs are all below or around 1% 44 , highlighting the significant challenge of providing screening tests with high PPVs. We identified four glycopeptides that contributed to the differentiation of EOC. High levels of sialylation and fucosylation in AGP, induced by cancer development, have been reported in several types of cancer, including liver cancer 33 , melanoma 45 , and lung cancer 46 . Similarly, the fucosylation of haptoglobin is increased in patients with various cancers, including liver 47 , ovarian 48 , gastric 49 , 50 , colorectal 50 , pancreatic 50 , 51 , and prostate 50 cancers. The observed increase in the glycopeptide derived from complement C9 and the decrease in the glycopeptide derived from transferrin could be attributed to changes in protein expression rather than glycan alterations. This is supported by an increase in complement C9 protein levels in colorectal cancer 52 , 53 and a decrease in serum transferrin levels during ovarian cancer development 54 . The CSGSA OPLSDA model provided a clearer distinction between CCC and EM than screening for the available tumor markers, CA125 and HE4. The ROC-AUC reached 0.808 when comparing CCC-E with EM, which was a significant improvement over that of the current tumor markers, CA125 (0.538) and HE4 (0.557). Although the PPV and sensitivity were 11.2% and 52.2%, respectively, this would assist doctors in distinguishing between malignant and benign tumors more accurately than using the current tumor markers. This study has several limitations. (i) The retrospective approach is limited to verifying the capability of CSGSA. A prospective approach is warranted to verify whether it can truly identify EOC in asymptomatic women. (ii) Discriminating EOC from other cancers, including CRC, LC, GC, or PC, is vital to increase selectivity because the prevalence rates of these cancers are one digit higher than that of EOC. (iii) The number of specimens used to develop the CNN model was insufficient to obtain a fully optimized model because the training performance reached almost 100%, indicating the occurrence of partial overfitting. The application of omics approaches in cancer screening has been controversial 55 ; however, the integration of omics with artificial intelligence can advance this field of research. (iv) We have only identified four glycopeptides and have not yet identified many glycopeptides that contribute to EOC discrimination. Glycopeptides pose a significant identification challenge compared to peptides. While proteomics primarily focuses on identifying proteins, glycopeptide analysis requires not only identifying the protein but also determining the specific site of sugar attachment and the structure of the glycan itself. Additionally, while proteomics benefits from comprehensive MS/MS libraries and databases, glycopeptides lack such resources due to the complexity of their MS/MS patterns. Although we have successfully identified several glycopeptide structures in past studies 33 , 34 , 36 , these efforts were time-consuming and labor-intensive. Given these challenges, it was not feasible to identify all 1712 glycopeptides at this time. However, identification of these glycopeptides will help us establish assays that do not require high-cost mass spectrometry. (v) BRCA gene mutations 1/2 are known risk factors for EOC pathogenesis; therefore, it is necessary to investigate the relationship between these mutations and CSGSA determination. (vi) Achieving more precise differentiation between CCC-E and EM is essential. The incorporation of tissue factor pathway inhibitor 2 (TFPI2), which is widely used as a marker to discriminate between CCC and EM 56 , or Wisteria floribunda agglutinin-reactive ceruloplasmin 57 could enhance the discriminatory performance of the CSGSA OPLSDA model. Our main focus was to understand whether early-stage identification contributes to a decrease in EOC mortality. Menon et al. stressed the significance of reducing EOC mortality as an outcome of screening trials in the UKCTOCS study 58 . The screening test involving the use of CA125 not only decreased the incidence of stage IV cases but also reduced deaths in stage IV by almost 25% compared with the non-screening group. However, the overall EOC mortality rate did not show any improvement. We believe CSGSA has the potential to contribute to a decrease in EOC mortality. In response, we initiated a new prospective study in 2023, supported by the Japan Agency for Medical Research and Development 59 .

Conclusions

The CSGSA CNN model effectively distinguished early-stage EOC (Stage I) from non-EOC controls, achieving a ROC–AUC of 0.924, which is higher than the ROC-AUCs of existing tumor markers CA125 (0.842) and HE4 (0.717). By setting two cutoffs to classify examinees into three categories, we found that 0.20% of the examinees were classified into the high-risk class with a PPV of 8.03%, 0.76% into the middle-risk class with a PPV of 0.734%, and over 99% into the low-risk class with a PPV of 0.00872%. Additionally, the CSGSA OPLSDA model demonstrated the ability to distinguish patients with early-stage CCC (Stage I) from patients with EM, with an ROC–AUC of 0.808, outperforming CA125 (0.538) and HE4 (0.557). This approach is expected to significantly enhance ovarian cancer diagnosis by enabling early detection in asymptomatic women and by improving the accuracy of differentiating between CCC and EM.

Introduction

Ovarian cancer (OC) is recognized as a significant challenge in gynecological oncology due to its often asymptomatic progression, which complicates early detection efforts 1 , 2 .The fact that over half of OC cases are detected at an advanced stage underscores the critical need for improved early detection methods 3 . Imaging modalities such as positron emission tomography (PET) and computed tomography (CT) are potential avenues as screening tests 4 , 5 . However, their widespread use has been hampered by their economic cost and physical burden on patients 6 . Noninvasive blood tests using tumor markers such as cancer antigen 125 (CA125) or human epididymis protein 4 (HE4) 7 , 8 are preferred alternatives; however, they exhibit limited effectiveness in detecting early-stage OC 9 , 10 . To accomplish this objective, several researchers have discovered various biomarkers to identify OC; including nectin-4 11 , CCL20 12 , Adam17 13 , sialic acids 14 , anti-PDLIM1 autoantibody 15 , and ciRS-7 16 . Additionally, the combination assays targeting multiple biomarkers has been proposed as a potential solution to overcome the limitations of using a single marker 17 – 22 . Particularly, machine learning and deep learning have been increasingly utilized in the field of ovarian cancer diagnostics 21 , 23 . However, these technologies have not been adopted in clinical practice due to their lack of proven effectiveness and insufficient validation across a broad range of cases. Epithelial ovarian cancer (EOC) is the most common type of ovarian cancer, accounting for about 90% of cases. It includes several subtypes, such as serous (SE), clear cell carcinoma (CCC), mucinous (MU), and endometrioid (EN), each with different prognoses and treatment responses. The accurately differentiating between benign conditions such as endometrioma (EM) and endometriosis-associated ovarian cancer is urgently warranted 24 . Distinguishing CCC of the ovary 25 , a histologic type of EOC that is more common in younger women 26 , from EM is particularly concerning. The prognosis of CCC is discouraging primarily owing to its noted resistance to platinum-based chemotherapy treatments 27 , 28 . Despite numerous proposals to differentiate CCC from EM, solid empirical evidence supporting their efficacy remains scarce 29 , 30 . Aberrant glycosylation in serum proteins emerge with cancer development 31 , 32 ; nevertheless, the precise characterization of these glycan alterations has remained elusive owing to several technical issues. We have overcome the obstacles using a proteomic approach that analyzes glycopeptides instead of glycans, which offers the advantage of detecting both glycan alterations and changes in serum protein expression induced by cancer development 33 – 35 . However, separating glycopeptides from non-glycosylated peptides presents challenges in conventional glycoproteomics. Typically, lectins are used for the separation, but no single lectin can selectively bind to all glycan structures, and processing large samples with lectins is labor-intensive. We prioritized the practical application of this test and aimed to develop a simple, robust, and high-throughput method by addressing two key issues. Firstly, we enriched glycopeptides based on their molecular weight differences. The molecular size of glycopeptides, when digested with trypsin, are significantly larger than non-glycosylated peptides and can be concentrated using ultrafiltration membranes 33 . Although this method permits some contamination by non-glycosylated peptides, it greatly improved throughput compared to lectin enrichment. Secondly, we only identified key glycopeptides that were significant to the discriminative model, rather than attempting to identify every glycopeptide detected by mass spectrometry. This selective approach is due to the fact that glycopeptide identification is more challenging than that of peptides in proteomics, owing to the complexity of their structures 34 , 36 . These decisions substantially enhanced our analysis efficiency, enabling us to test over 1000 cases. The comprehensive serum glycopeptide spectra analysis (CSGSA) analyzes 1712 enriched glycopeptides (EGPs) using either orthogonal partial least squares discriminant analysis (OPLSDA) 37 , 38 or convolutional neural network (CNN) 39 . It integrates conventional tumor markers with glycopeptides derived from serum glycoproteins, which reflect an individual's physical state. This method aims to challenge conventional notions of EOC diagnosis and overcome the limitations of single biomarkers. We aimed to validate the diagnostic ability of CSGSA for the early detection of EOC and to differentiate between CCC malignancies and benign EM using a comprehensive dataset comprising more than 500 EOC specimens and 1000 benign control specimens obtained from various medical institutions.

Supplementary Material

Supplementary Information 1. Supplementary Information 2. Supplementary Information 3. Supplementary Information 1. Supplementary Information 2. Supplementary Information 3.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-08-30T09:23:35.175841+00:00
unpaywall
last seen: 2026-05-21T05:10:58.409756+00:00
License: CC-BY-NC-ND-4.0