Constructing a prognostic survival model for patients with stage II-IV epithelial ovarian cancer: a study based on the SEER database and external validation in China

Other OA: gold CC-BY-NC-ND-4.0
⚙ AI-generated summary by qwen3.7-flash, 2026-09-18 ⓘ

An externally validated nomogram predicting 1-, 3-, and 5-year overall survival in stage II-IV epithelial ovarian cancer was developed using SEER data, demonstrating moderate discriminative performance and clinical utility for personalized prognosis.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

⚙ AI-generated deep summary by qwen3.7-flash, 2026-09-18 · read from full text ⓘ

This study developed and validated a prognostic survival model for patients with stage II–IV epithelial ovarian cancer using data from the SEER database and an external Chinese cohort. The researchers analyzed 17 clinicopathological variables, employing multiple imputation by chained equations to handle significant missing data in key factors such as CA125 levels and tumor size. Key findings identified age, histological subtype, stage, and treatment modalities as independent predictors of overall survival, with the resulting model demonstrating superior predictive accuracy compared to traditional staging alone. Relevance to endometriosis: endometrioid carcinoma, a subtype linked to endometriosis, is listed among the histotypes analyzed for prognostic differences.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

This study utilized the Surveillance, Epidemiology, and End Results (SEER) database to identify clinicopathological prognostic factors and develop a prognostic nomogram for predicting 1-, 3-, and 5-year overall survival (OS) in patients with stage II-IV epithelial ovarian cancer (EOC), aiming to enhance survival prediction accuracy and inform personalized therapeutic strategies. Clinical data of patients diagnosed between 2004 and 2020 were extracted and randomly divided into training, tuning, and internal validation cohorts at a 7:2:1 ratio. Univariate and multivariate Cox regression analyses identified age, tumor differentiation grade, AJCC stage, tumor size, number of positive lymph nodes, number of lymph nodes examined, surgery, chemotherapy, sequence of systemic therapy and surgery, and time from diagnosis to treatment as independent prognostic factors for OS (all P < 0.05). A prognostic nomogram was subsequently developed and externally validated using an independent cohort of 115 EOC patients from Chengde Medical University (2014-2020). The model exhibited acceptable discriminative performance, with concordance indexes (C-index) of 0.693, 0.683, and 0.700 for the training, internal validation, and external validation cohorts, respectively. Calibration curves, receiver operating characteristic (ROC) curves, and area under the curve (AUC) values all exceeded 0.7, indicating reliable prediction of 1-, 3-, and 5-year survival rates. Kaplan-Meier curves revealed significant survival differences across risk groups, and decision curve analysis (DCA) confirmed the clinical utility of the nomogram. In conclusion, this externally validated nomogram can assist in predicting OS in patients with stage II-IV EOC, demonstrating moderate yet clinically useful performance with a C-index of 0.700. It addresses the need for personalized prognostic assessment and aids in tailoring therapeutic strategies, thereby potentially improving patient outcomes.
Full text 88,943 characters · extracted from pmc-nxml · 6 sections · click to expand

Results

Of the total, 28,972 patients met the eligibility criteria and were included in the study. These patients were stratified and randomly sampled using R Studio and were divided into the training set (n = 20,281), the tuning set (n = 5,794), and the internal validation set (n = 2,897) (Fig.  1 A). As shown in Table 1 , the three sets all exhibited a high degree of balance in all clinical pathological features. The standardized mean difference (SMD) analysis revealed that the SMD values of all variables ranged from 0.001 to 0.050 (all < 0.1), indicating that the data segmentation process effectively maintained the balance of feature distributions between groups. Table 1 Clinicopathological data of patients with EOC. Stratified by cohort Level Training Tuning Validation SMD n 20,281 5794 2897 Age (%)  < 40 830 (4.09) 239 (4.12) 114 (3.94) 0.019 40–69 13,741 (67.75) 3866 (66.72) 1969 (67.97)  ≥ 70 5710 (28.15) 1689 (29.15) 814 (28.10) Race (%) White 16,791 (82.79) 4819 (83.17) 2367 (81.71) 0.035 Black 1402 (6.91) 421 (7.27) 229 (7.90) Others 2088 (10.30) 554 (9.56) 301 (10.39) Marital (%) Single 3649 (17.99) 1011 (17.45) 558 (19.26) 0.043 Married 10,879 (53.64) 3099 (53.49) 1525 (52.64) Widowed 2682 (13.22) 811 (14.00) 409 (14.12) Other 3071 (15.14) 873 (15.07) 405 (13.98) Histologic (%) Serous 16,909 (83.37) 4828 (83.33) 2374 (81.95) 0.040 Mucinous 577 (2.85) 183 (3.16) 79 (2.73) Endometrioid 1209 (5.96) 342 (5.90) 194 (6.70) Clear cell 1030 (5.08) 283 (4.88) 160 (5.52) Other epithelial 556 (2.74) 158 (2.73) 90 (3.11) AJCC (%) Ⅳ 6603 (32.56) 1920 (33.14) 964 (33.28) 0.019 Ⅲ 11,392 (56.17) 3184 (54.95) 1595 (55.06) Ⅱ 2286 (11.27) 690 (11.91) 338 (11.67) Grade (%) Highly 765 (3.77) 231 (3.99) 115 (3.97) 0.019 Moderate 1783 (8.79) 519 (8.96) 266 (9.18) Poorly/undifferentiated 12,163 (59.97) 3502 (60.44) 1750 (60.41) Blank 5570 (27.46) 1542 (26.61) 766 (26.44) CA125 (%) Negative/Borderline 641 (3.16) 178 (3.07) 88 (3.04) 0.005 Positive 11,024 (54.36) 3143 (54.25) 1574 (54.33) Unknown/Blank 8616 (42.48) 2473 (42.68) 1235 (42.63) Laterality (%) Unilateral 8474 (41.78) 2408 (41.56) 1208 (41.70) 0.021 Bilateral 9945 (49.04) 2902 (50.09) 1440 (49.71) Paired 1862 (9.18) 484 (8.35) 249 (8.60) Tumor Size (%) Blank 6169 (30.42) 1743 (30.08) 852 (29.41) 0.020  ≤ 10 8656 (42.68) 2502 (43.18) 1240 (42.80)  > 10 5456 (26.90) 1549 (26.73) 805 (27.79) Number of examined LN (%) 0 10,381 (51.19) 2931 (50.59) 1489 (51.40) 0.047 1–9 5078 (25.04) 1483 (25.60) 695 (23.99)  ≥ 10 4222 (20.82) 1180 (20.37) 637 (21.99) Unknown 600 (2.96) 200 (3.45) 76 (2.62) Number of positive LN (%) Negative 4853 (23.93) 1407 (24.28) 689 (23.78) 0.010  ≥ 1 4853 (23.93) 1398 (24.13) 692 (23.89) Unknown 10,575 (52.14) 2989 (51.59) 1516 (52.33) Residual lesion (%) No/Blank 12,884 (63.53) 3691 (63.70) 1818 (62.75) 0.041 R0 4180 (20.61) 1185 (20.45) 596 (20.57) R1: ≤ 1 cm 1477 (7.28) 425 (7.34) 231 (7.97) R2: > 1 cm 905 (4.46) 246 (4.25) 108 (3.73) Macroscopic Residual size Unknown 835 (4.12) 247 (4.26) 144 (4.97) Chemotherapy (%) Yes 17,398 (85.78) 4948 (85.40) 2506 (86.50) 0.021 No 2883 (14.22) 846 (14.60) 391 (13.50) Surgery (%) Yes 18,137 (89.43) 5193 (89.63) 2607 (89.99) 0.012 No 2144 (10.57) 601 (10.37) 290 (10.01) LN surgery scope (%) No 10,554 (52.04) 2981 (51.45) 1500 (51.78) 0.028 1–3 2244 (11.06) 709 (12.24) 336 (11.60)  ≥ 4 6907 (34.06) 1931 (33.33) 974 (33.62) Sentinel LN biopsy 449 (2.21) 139 (2.40) 67 (2.31) Blank 127 (0.63) 34 (0.59) 20 (0.69) Systemic treatment Surgery sequence (%) No 6547 (32.28) 1920 (33.14) 915 (31.58) 0.034 Systemic treatment before surgery 1480 (7.30) 415 (7.16) 227 (7.84) Systemic treatment intraoperative/after surgery 9592 (47.30) 2723 (47.00) 1372 (47.36) Systemic treatment both before and after surgery 2508 (12.37) 702 (12.12) 358 (12.36) Surgery both before and after systemic treatment 154 (0.76) 34 (0.59) 25 (0.86) Status (%) Survival 8114 (40.01) 2303 (39.75) 1137 (39.25) 0.010 death 12,167 (59.99) 3491 (60.25) 1760 (60.75) Months from diagnosis to treatment (%)   3 months 182 (0.90) 53 (0.91) 25 (0.86) Blank 581 (2.86) 158 (2.73) 78 (2.69) SMD = Standardized Mean Difference; Balance Threshold: SMD < 0.1; All continuous variables are expressed as mean (standard deviation); The SMD calculation uses Cohen’s d method: (mean₁—mean₂) / pooled SD. Clinicopathological data of patients with EOC. SMD = Standardized Mean Difference; Balance Threshold: SMD < 0.1; All continuous variables are expressed as mean (standard deviation); The SMD calculation uses Cohen’s d method: (mean₁—mean₂) / pooled SD. Substantial missingness was observed in key prognostic variables (CA125: 42.5%; tumor size: 30.4%; positive lymph nodes: 52.1%), resulting in only 22.4% complete cases (4,544/20,281). Systematic analysis (Table 2 ) revealed distributional biases: complete cases exhibited elevated tumor burden (+ 7% size, + 6% positive LNs) but lower CA125 (–2%) versus the full cohort. Validation of the Missing at Random (MAR) mechanism was confirmed through: 1. Non-significant associations between missing indicators and survival in Cox models (CA125: HR = 1.02[0.95–1.09], P  = 0.572; tumor size: HR = 0.97[0.91–1.04], P  = 0.372; positive LNs: HR = 1.04[0.98–1.11], P  = 0.218); 2. Stability of effect estimates post-imputation (76.5% variables with < 10% HR change between CCA and MI); 3. Robustness across sensitivity scenarios (pattern mixture models, inverse probability weighting, and CCA). Table 2 Sensitivity analysis: comparison of results between complete cases and multiple imputation. Variable Complete case analysis (CCA) Multiple imputation analysis (MI) HR difference (%) Missing-survival association HR 95%CI HR 95%CI HR (95% CI) p -value Age 1.452 1.32–1.60 1.498 1.44–1.55 3.1 1.01 (0.97–1.05) 0.681 Race 1.013 0.94–1.09 0.966 0.94–1.00 -4.6 0.98 (0.94–1.02) 0.412 Marital 1.039 0.99–1.09 1.028 1.01–1.05 -1.1 1.03 (0.99–1.07) 0.198 AJCC 0.601 0.55–0.65 0.707 0.68–0.73 17.7 0.95 (0.91–1.00) 0.063 CA125 1.364 1.11–1.68 1.459 1.25–1.70 7.0 1.02 (0.95–1.09) 0.572 Chemotherapy 0.925 0.72–1.18 1.159 1.09–1.23 25.3 1.05 (0.98–1.12) 0.175 Grade 1.201 1.11–1.30 1.113 1.08–1.15 -7.3 0.99 (0.95–1.03) 0.584 Histologic 1.254 1.21–1.31 1.123 1.10–1.15 -10.4 1.04 (1.00–1.08) 0.057 Laterality 1.215 1.11–1.33 1.075 1.04–1.11 -11.5 0.97 (0.93–1.01) 0.165 LN surgery scope 0.863 0.79–0.94 0.878 0.85–0.91 1.8 1.02 (0.98–1.06) 0.387 Residual lesion 1.134 1.09–1.18 1.041 1.02–1.06 -8.2 0.96 (0.93–1.00) 0.043* Surgery 3.559 2.39–5.31 1.745 1.62–1.88 -51.0 1.08 (1.01–1.15) 0.022* Systemic treatment Surgery seq 0.879 0.80–0.97 0.891 0.87–0.91 1.4 1.01 (0.97–1.05) 0.721 Tumor Size 0.773 0.70–0.85 0.883 0.84–0.93 14.2 0.97 (0.91–1.04) 0.372 Number of examined LN 0.886 0.81–0.97 0.889 0.86–0.92 0.4 1.03 (0.99–1.07) 0.155 Number of positive LN 1.365 1.24–1.51 1.357 1.27–1.45 -0.6 1.04 (0.98–1.11) 0.218 Months from diagnosis to treatment 1.044 0.95–1.15 1.137 1.10–1.17 8.9 0.99 (0.94–1.04) 0.638 LN, Lymph Node: Missing mechanism: Cox regression of missing indicators on survival; MAR supported if p  > 0.05 (bolded for key biomarkers); Clinical interpretation: All HR estimates between 0.91–1.15 indicate < 15% risk change; * p  < 0.05 but HR ≈1.0 indicates marginal clinical significance. Sensitivity analysis: comparison of results between complete cases and multiple imputation. 1.01 (0.97–1.05) 0.98 (0.94–1.02) 1.03 (0.99–1.07) 0.95 (0.91–1.00) 1.02 (0.95–1.09) 1.05 (0.98–1.12) 0.99 (0.95–1.03) 1.04 (1.00–1.08) 0.97 (0.93–1.01) 1.02 (0.98–1.06) 0.96 (0.93–1.00) 1.08 (1.01–1.15) 1.01 (0.97–1.05) 0.97 (0.91–1.04) 1.03 (0.99–1.07) 1.04 (0.98–1.11) 0.99 (0.94–1.04) LN, Lymph Node: Missing mechanism: Cox regression of missing indicators on survival; MAR supported if p  > 0.05 (bolded for key biomarkers); Clinical interpretation: All HR estimates between 0.91–1.15 indicate < 15% risk change; * p  < 0.05 but HR ≈1.0 indicates marginal clinical significance. To visually compare the effect estimates across three missing data handling methods (complete case analysis [CCA], multiple imputation [MICE], and mode/median imputation), we constructed a forest plot (Fig.  2 ). For the majority of covariates – including age, grade, CA125, positive lymph node count, residual lesion, LN surgery scope, and treatment sequence – hazard ratios (HRs) and 95% confidence intervals (CIs) were highly consistent across the three methods, with overlapping CIs and no qualitative differences. For surgery, CCA produced a markedly inflated HR (≈3.5) with a wide CI, whereas MICE and mode/median imputation yielded HRs around 1.8 with narrower CIs (the largest HR change between CCA and MICE was observed for surgery: CCA: HR = 3.56 → MICE: HR = 1.75). For chemotherapy, CCA indicated a protective but non‑significant effect (HR  1). Multiple imputation (MICE, 20 datasets) significantly enhanced model discrimination compared to CCA: the global C-index increased from 0.686 to 0.698, and the 1‑year AUC improved by 0.058 to 0.780 (Table 3 ). The mode/median imputation model yielded nearly identical performance (C‑index = 0.696; 1‑year AUC = 0.780), further supporting the robustness of the MICE results. Although Brier scores rose with MICE (e.g., + 0.040 at 1‑year), this reflected the intentional retention of high‑risk cases that were excluded in CCA (Table 3 ). Collectively, the sensitivity analyses (Table 2 , Fig.  2 , Table 3 ) confirm that while missing data affect complete case analysis, the use of multiple imputation (or even a simple imputation) yields consistent and reliable prognostic estimates. Table 3 Performance metrics of Cox models under different missing data handling methods. Method C-index AUC (1-year) AUC (3-year) AUC (5-year) Brier Score (1-year) Brier Score (3-year) Brier Score (5-year) Complete case analysis (CCA) 0.686 0.722 0.718 0.715 0.065 0.163 0.211 Multiple imputation (MICE) 0.698 0.780 0.737 0.739 0.105 0.204 0.226 Mode/Median Imputation 0.696 0.780 0.735 0.734 0.105 0.205 0.227 Missing data were handled by three methods: complete case analysis (CCA), multiple imputation by chained equations (MICE), and mode/median imputation. Details of MICE parameters are provided in section “ Statistical analysis ”. All metrics are rounded to three decimal places. Performance metrics of Cox models under different missing data handling methods. Missing data were handled by three methods: complete case analysis (CCA), multiple imputation by chained equations (MICE), and mode/median imputation. Details of MICE parameters are provided in section “ Statistical analysis ”. All metrics are rounded to three decimal places. Univariable Cox regression analysis of the MICE-imputed training set demonstrated statistically significant associations ( P  < 0.05) between all candidate variables and survival outcomes in stage II-IV ovarian cancer (Table 4 ). To mitigate overfitting, least absolute shrinkage and selection operator (LASSO) regression with tenfold cross‑validation was applied to the tuning set, yielding values of λ.min = 0.00321 and λ.1se = 0.05234. As shown in Fig.  3 A, the coefficient paths exhibited progressive shrinkage as the penalty parameter λ increased. Figure  3 B displays the cross‑validation error curve, which reached its minimum at λ.min, while λ.1se (the largest λ within one standard error of the minimum) provided a more parsimonious model. Using the λ.1se criterion, 10 non‑redundant prognostic factors were selected (compared to 17 under λ.min): patient age, tumor differentiation grade, AJCC stage, tumor size, number of positive lymph nodes, total lymph nodes examined, surgical approach, chemotherapy regimen, sequence of systemic therapy relative to surgery, and time from diagnosis to treatment initiation. These covariates were subsequently incorporated into a multivariable Cox proportional hazards model using the MICE-processed training set. Rubin’s rules were applied to pool estimates across 20 imputed datasets. The final model confirmed all 10 factors as statistically independent prognostic determinants (P < 0.05; Table 4 ). Results were visualized via a forest plot generated with the R “forestplot” package (Fig.  4 ), displaying hazard ratios with 95% confidence intervals. Table 4 Univariate and multivariate Cox regression analyses of patients in the training set after MICE. Variable Single factor analysis Multi-factor analysis HR 95%CI P HR 95%CI P Age (years)  < 40 1.00 Reference 40–69 1.76 (1.58–1.97)  < 0.001*** 1.53 (1.36–1.72)  < 0.001***  ≥ 70 3.16 (2.82–3.54)  < 0.001*** 2.28 (2.02–2.56)  < 0.001*** Race White 1.00 Black 1.24 (1.15–1.32)  < 0.001*** Others 0.82 (0.77–0.88)  < 0.001*** Marital Single 1.00 Married 0.94 (0.90–0.99) 0.019* Widowed 1.67 (1.57–1.77)  < 0.001*** Others 1.09 (1.03–1.16)  < 0.005** Histologic Serous 1.00 Mucinous 1.40 (1.26–1.55)  < 0.001*** Endometrioid 0.48 (0.44–0.52)  < 0.001*** Clear cell 0.99 (0.91–1.08) 0.789 Other epithelial 1.48 (1.33–1.65)  < 0.001*** AJCC Ⅳ 1.00 Reference Ⅲ 0.62 (0.60–0.65)  < 0.001*** 0.83 (0.80–0.87)  < 0.001*** Ⅱ 0.22 (0.20–0.24)  < 0.001*** 0.37 (0.34–0.41)  < 0.001*** Grade Highly 1.00 Reference Moderate 2.07 (1.79–2.40)  < 0.001*** 1.85 (1.59–2.14)  < 0.001*** Poorly/undifferentiated 2.62 (2.29–3.00)  < 0.001*** 2.06 (1.80–2.36)  < 0.001*** Blank 3.66 (3.20–4.20)  < 0.001*** 2.00 (1.74–2.30)  < 0.001*** CA125 Negative/Borderline 1.00 Positive 1.67 (1.53–1.82)  < 0.001*** Laterality Unilateral 1.00 Bilateral 1.20 (1.15–1.25)  < 0.001*** Paired 2.29 (2.15–2.43)   10 0.81 (0.78–0.84)  < 0.001*** 0.92 (0.88–0.96)  < 0.001*** Number of LN examined 0 1.00 Reference 1–9 0.58 (0.56–0.61)  < 0.001*** 0.72 (0.68–0.75)  < 0.001***  ≥ 10 0.44 (0.42–0.46)  < 0.001*** 0.56 (0.53–0.60)  < 0.001*** Unknown 0.82 (0.74–0.91)  < 0.001*** 0.77 (0.70–0.86)  < 0.001*** Number of positive LN Negative 1.00 Reference  ≥ 1 1.47 (1.42–1.53)  < 0.001*** 1.32 (1.25–1.39)  < 0.001*** Residual lesion (cm) No/Blank 1.00 R0 0.50 (0.47–0.52)   1 1.09 (1.00–1.18) 0.048* Macroscopic Residual size Unknown 1.23 (1.13–1.34)  < 0.001*** Chemotherapy No 1.00 Reference Yes 1.35 (1.28–1.41)  < 0.001*** 1.08 (1.02–1.15) P < 0.05* Surgery Yes 1.00 Reference No 3.94 (3.75–4.15)  < 0.001*** 1.85 (1.72–2.00)  < 0.001*** LN surgery No 1.00 1–3 0.65 (0.61–0.69)  < 0.001***  ≥ 4 0.47 (0.45–0.49)  < 0.001*** Sentinel LN biopsy 0.64 (0.56–0.72)  < 0.001*** Blank 0.99 (0.80–1.22) 0.911 Systemic treatment/Surgery sequence No 1.00 Reference Systemic treatment before surgery 0.93 (0.87–1.00) 0.045* 0.94 (0.87–1.02) 0.138 Systemic treatment intraoperative/after surgery 0.57 (0.54–0.59)  < 0.001** 0.76 (0.72–0.80)  < 0.001*** Systemic treatment both before and after surgery 0.75 (0.70–0.80)  < 0.001*** 0.79 (0.73–0.85)  < 0.001*** Surgery both before and after systemic treatment 0.56 (0.43–0.73)  < 0.001*** 0.62 (0.48–0.80)  < 0.001*** Months from diagnosis to treatment  < 1 month 1.00 Reference 1–3 months 1.35 (1.30–1.40)   3 months 1.52 (1.27–1.81)  < 0.001*** 1.02 (0.85–1.22) 0.836 Blank 5.65 (5.16–6.18)  < 0.001*** 1.99 (1.78–2.22)  < 0.001*** * P  < 0.05, ** p < 0.01, *** p < 0.001, Statistically significant difference, the significance increased successively. Univariate and multivariate Cox regression analyses of patients in the training set after MICE. * P  < 0.05, ** p < 0.01, *** p  10 cm) was associated with improved survival (HR = 0.92, 95% CI: 0.88–0.96, P  < 0.001), a finding contrary to the clinical expectation that larger tumors generally portend poorer prognosis. To further investigate this unexpected finding, we performed a subgroup analysis stratified by AJCC stage. The stage‑specific hazard ratios for tumor size > 10 cm (versus ≤ 10 cm) were: stage II: HR = 0.993 (95% CI 0.841–1.174, P  = 0.94); stage III: HR = 0.840 (95% CI 0.795–0.888, P  < 0.001); stage IV: HR = 0.938 (95% CI 0.866–1.015, P  = 0.11). A significant protective effect was observed only in stage III ( P  < 0.001), whereas no significant association was found in stage II or stage IV. The hazard ratio for stage III (0.840) was also lower than those for stage II (0.993) and stage IV (0.938). These results suggest that the overall protective effect of larger tumor size (HR = 0.92) was largely driven by stage III patients. Based on the 10 independent prognostic factors identified through LASSO regression analysis, a nomogram was developed using the “rms” and “regplot” packages in R Studio. The scoring system of the nomogram was based on the final multivariable Cox proportional hazards regression model to provide visual predictions of individual survival risk. The scoring of the nomogram was derived as follows. First, using the regression coefficient with the largest absolute value (β_max) from the model, the full clinical range of the corresponding variable was mapped to a predefined interval of 0–100 points to define a scaling factor (Factor = 100 / β_max). Next, the point value for each variable (or variable level) was calculated by multiplying its regression coefficient (β) by the scaling factor (i.e., points = β × Factor). Each variable in the nomogram was assigned a specific score corresponding to its contribution to prognosis. The total score, calculated by summing the individual scores for a given patient, was used to estimate 1-year, 3-year, and 5-year cumulative death probabilities (i.e., the risk of dying within that time). The corresponding overall survival (OS) probabilities are obtained as 1 – death probability. For instance, a 65-year-old patient diagnosed with EOC and exhibiting the following characteristics (Fig.  5 ) was assigned scores as follows: age—65 (40 points); tumor differentiation grade—moderate (37 points); FIGO stage—Stage IV (77.5 points); tumor size—greater than 10 cm (45 points); number of lymph nodes examined—more than 10 (22 points); number of positive lymph nodes— ≥ 1 (64 points); receipt of chemotherapy (47 points) and surgery (43.5 point); surgery both before and after systemic treatment (23 points); and time from diagnosis to treatment— < 1 month (44 points). The total score for this patient was 443. According to the nomogram, this total score corresponds to cumulative death probabilities of 7.96% at 1 year, 26.9% at 3 years, and 42.5% at 5 years. Thus, the estimated OS probabilities are 92.04%, 73.1%, and 57.5%, respectively. Fig. 5 Nomogram of 1-year, 3-year, and 5-year overall survival in women with Stage II–IV metastatic EO0C. Notes: Values shown on the axes are cumulative death probabilities. Overall survival probability = 1 – death probability. The example patient (total points = 443, highlighted in the figure) has death probabilities of 7.96% (1-year), 26.9% (3-year), and 42.5% (5-year), corresponding to survival probabilities of 92.04%, 73.1%, and 57.5%, respectively. Nomogram of 1-year, 3-year, and 5-year overall survival in women with Stage II–IV metastatic EO0C. Notes: Values shown on the axes are cumulative death probabilities. Overall survival probability = 1 – death probability. The example patient (total points = 443, highlighted in the figure) has death probabilities of 7.96% (1-year), 26.9% (3-year), and 42.5% (5-year), corresponding to survival probabilities of 92.04%, 73.1%, and 57.5%, respectively. To validate and evaluate the predictive performance of this nomogram, its accuracy must be assessed in both the imputed training set and the independent internal validation set of Stage II-IV epithelial ovarian cancer patients, where identical variable selection results derived from the training set were applied to the validation set. After deleting 7 variables, the recalculated C-index for the training set was 0.693, representing a slight decrease of 0.005 (from 0.698 to 0.693). This difference is not clinically significant and has improved the practicality of the model. On the independent internal validation set, the C-index of this prognostic model was 0.683 (95% CI 0.669—0.697), indicating that the model has moderate predictive ability. Compared with the performance of the training set (C-index = 0.693), the model demonstrates moderate generalization ability. The “timeROC” package was subsequently used to generate ROC curves and calculate the area under the curve (AUC) values for 1-year, 3-year, and 5-year OS. For the internal validation set, the AUC values were 0.755 (1-year), 0.712 (3-year), and 0.732 (5-year). These consistently high AUC values across internal validation set indicated good discriminatory power of the nomogram (Fig.  6 ). Fig. 6 ROC curves for the internal and external validation cohorts at 1, 3, and 5 years ( A – C ) ROC curves for the external validation cohort at 1-, 3-, and 5-year survival, respectively.( D – F ) ROC curves for the internal validation cohort at 1-, 3-, and 5-year survival, respectively. ROC curves for the internal and external validation cohorts at 1, 3, and 5 years ( A – C ) ROC curves for the external validation cohort at 1-, 3-, and 5-year survival, respectively.( D – F ) ROC curves for the internal validation cohort at 1-, 3-, and 5-year survival, respectively. Calibration curves for 1-year, 3-year, and 5-year OS using the internal validation set were plotted with the “rms” package (Fig.  7 ). These curves showed strong agreement between predicted and observed survival probabilities, as they closely aligned with the ideal reference line (slope = 1). The mean absolute error (MAE) values were 0.051 (1‑year), 0.093 (3‑year), and 0.079 (5‑year). In addition, the Brier scores (Table 3 ) were low: 0.105 (1‑year), 0.204 (3‑year), and 0.226 (5‑year). Together, these metrics confirm good calibration of the model. Collectively, these results confirmed that the nomogram-based survival prediction model is both reliable and clinically valuable for assessing prognosis in patients with Stage II–IV EOC. Fig. 7 Calibration curves for the internal and external validation cohorts at 1, 3, and 5 years. ( A – C ) Internal validation cohort at 1, 3, and 5 years, respectively. ( D – F ) External validation cohort at 1, 3, and 5 years, respectively. MAE = Mean Absolute Error; The error bars represent the 95% confidence interval; The dotted line is the ideal calibration line. Calibration curves for the internal and external validation cohorts at 1, 3, and 5 years. ( A – C ) Internal validation cohort at 1, 3, and 5 years, respectively. ( D – F ) External validation cohort at 1, 3, and 5 years, respectively. MAE = Mean Absolute Error; The error bars represent the 95% confidence interval; The dotted line is the ideal calibration line. Patients were stratified into low‑, intermediate‑, and high‑risk groups based on tertiles of the nomogram‑derived risk scores (low risk: ≤ 33rd percentile; intermediate risk: 34th–66th percentile; high risk: ≥ 67th percentile). To compare this tertile‑based approach with statistically optimal cutoffs, we performed an X‑tile analysis. The X‑tile algorithm identified cutoffs at 0.686 and 3.195 (log‑rank χ 2  = 649.3, p  < 0.001), compared to the tertile cutoffs of 0.900 and 1.776 (χ 2  = 438.8, p  < 0.001). Although the X‑tile cutoffs yielded a higher χ 2 statistic, we retained the tertile‑based stratification. This approach generates three groups of equal size, which is more straightforward for clinical application, and still achieved a highly significant separation in survival (log‑rank p  < 0.0001). Kaplan‑Meier curves based on the X‑tile cutoffs are shown in Supplementary Figure S1 . Kaplan–Meier analysis using tertile-based risk stratification confirmed significantly divergent survival outcomes (log-rank P  < 0.0001; Fig.  8 ) among the nomogram-stratified groups. The high-risk group demonstrated a median survival of 25 months with  50% 5-year survival, collectively validating the model’s discriminative capacity. Fig. 8 Kaplan–Meier curves of risk stratification for the internal and external validation cohorts. ( A ) Internal validation cohort (Log-rank P < 2e-16). ( B ) External validation cohort (Log-rank P = 1.58e-05). Kaplan–Meier curves of risk stratification for the internal and external validation cohorts. ( A ) Internal validation cohort (Log-rank P < 2e-16). ( B ) External validation cohort (Log-rank P = 1.58e-05). Table 7 presents the distribution of key clinicopathological variables across the three risk groups. Compared to the low-risk group, patients in the high-risk group were older and had more advanced AJCC stage (where a lower numerical value denotes a more advanced stage according to our coding scheme), higher tumor grade, smaller tumor size, a higher number of positive lymph nodes, and a lower number of examined lymph nodes. They were also more frequently treated with chemotherapy and surgery (all p  < 0.001). Both the systemic treatment sequence and the time from diagnosis to treatment initiation differed significantly among the groups ( p  < 0.001). Decision curve analysis (DCA; Fig.  9 ) evaluated the clinical utility of the nomogram across 1‑, 3‑, and 5‑year time horizons by comparing its net benefit with two extreme strategies: treating all patients (yellow curve) and treating none (gray curve). For 1‑year predictions, the nomogram curve (blue) was essentially indistinguishable from “treat all” below a 15% threshold, but both became inferior to “treat none” above 15%, indicating negligible clinical benefit for short‑term risk stratification. For 3‑year predictions, the nomogram showed a useful decision window between approximately 25% and 45% thresholds, where its net benefit exceeded both extreme strategies. For 5‑year predictions, the model demonstrated the broadest clinical utility, with a positive net benefit across a 25%–50% threshold range, reflecting its greatest value for long‑term prognosis. Overall, the effective decision range widened progressively from 1‑year (limited) to 5‑year (widest), aligning with the expectation that long‑term survival is more stable and thus more predictable. The example patient in Fig.  5 (total score = 443, high‑risk) had a predicted 5‑year mortality risk of 42.5%, which lies within the DCA‑derived optimal decision threshold range (25%–50%) where the nomogram provides positive net benefit for guiding 5‑year management decisions. The model’s role is risk quantification; actual treatment selection should integrate tumor biology, patient preferences, and clinical guidelines. Fig. 9 Decision curve analysis for the internal and external validation cohorts at 1, 3, and 5 years. ( A – C ) DCA for the internal validation cohort at 1, 3, and 5 years, respectively. ( D – F ) DCA for the external validation cohort at 1, 3, and 5 years, respectively. Across all timepoints in both cohorts, the nomogram (blue line) demonstrated a superior net benefit across a wide range of threshold probabilities when compared to the “Treat None” strategy (gray line) and matched or exceeded the net benefit of the “Treat All” approach (yellow line) within the clinically relevant range. Decision curve analysis for the internal and external validation cohorts at 1, 3, and 5 years. ( A – C ) DCA for the internal validation cohort at 1, 3, and 5 years, respectively. ( D – F ) DCA for the external validation cohort at 1, 3, and 5 years, respectively. Across all timepoints in both cohorts, the nomogram (blue line) demonstrated a superior net benefit across a wide range of threshold probabilities when compared to the “Treat None” strategy (gray line) and matched or exceeded the net benefit of the “Treat All” approach (yellow line) within the clinically relevant range. The nomogram was rigorously evaluated using an independent cohort from Chengde Medical University (n = 115), which presented a distinct clinical profile (Table 5 ) compared to the SEER-derived training cohort (n = 20,281). The external cohort was notably younger (91.3% vs. 67.8% in age category 2, p  < 0.001), exclusively Asian (100% vs. 10.3%, p  < 0.001), and exhibited a different distribution within the study’s stage II-IV spectrum, with a significantly higher proportion of stage II disease (75.7% vs. 56.2%, p  < 0.001) and grade 3 tumors (86.1% vs. 60.0%, p  < 0.001). This distinct profile, characterized by a preponderance of locally advanced (stage II) and high-grade disease, provided a stringent test for the model’s generalizability across different case mixes within the advanced-stage population. Treatment strategies, including the scope of lymph node surgery and the sequencing of systemic therapy, also differed significantly (both p  < 0.001). Critically, despite these pronounced variations, key treatment modalities such as chemotherapy and surgery were similarly utilized, providing a consistent therapeutic context for model validation. The deliberate exclusion of race, marital status, and histologic type—variables with significant inter-cohort disparity but no independent prognostic value in the training set—from the final multivariate model was a crucial step that effectively mitigated potential bias and underscored the model’s reliance on biologically and clinically relevant predictors. Table 5 Baseline characteristics comparison for external validation. level External validation Training P n 115 20,281 Age (%)  < 40 1 (0.9) 830 (4.1)   69 9 (7.8) 5710 (28.2) Marital (%) Single 0 (0.0) 3649 (18.0)  < 0.001 Married 104 (90.4) 10,879 (53.6) Widowed 9 (7.8) 2682 (13.2) Other 2 (1.7) 3071 (15.1) Histologic (%) Serous 107 (93.0) 16,909 (83.4) 0.038 Mucinous 0 (0.0) 577 (2.8) Endometrioid 1 (0.9) 1209 (6.0) Clear cell 4 (3.5) 1030 (5.1) Other epithelial 3 (2.6) 556 (2.7) AJCC (%) Ⅳ 12 (10.4) 6603 (32.6)  < 0.001 Ⅲ 87 (75.7) 11,392 (56.2) Ⅱ 16 (13.9) 2286 (11.3) Grade (%) Highly 4 (3.5) 765 (3.8)  < 0.001 Moderate 4 (3.5) 1783 (8.8) Poorly/ undifferentiated 99 (86.1) 12,163 (60.0) Blank 8 (7.0) 5570 (27.5) CA125 (%) Negative/Borderline 5 (4.3) 641 (5.5)  < 0.001 Positive 108 (93.9) 11,024 (94.5) Unknown/Blank 2 (1.7) 0 (0.0) Laterality (%) Unilateral 50 (43.5) 8474 (41.8) 0.003 Bilateral 65 (56.5) 9945 (49.0) Paired 0 (0.0) 1862 (9.2) Tumor Size (%) Blank 5 (4.3) 0 (0.0)   10 41 (35.7) 5456 (38.7) Number of examined LN (%) 0 36 (31.3) 10,381 (51.2)  < 0.001 1–9 19 (16.5) 5078 (25.0)  ≥ 10 53 (46.1) 4222 (20.8) Unknown 7 (6.1) 600 (3.0) Number of positive LN (%) Negative 45 (39.1) 4853 (50.0)  < 0.001  ≥ 1 28 (24.3) 4853 (50.0) Unknown 42 (36.5) 0 (0.0) Residual lesion (%) No/Blank 11 (9.6) 12,884 (63.5)   1 cm 17 (14.8) 905 (4.5) Macroscopic Residual size Unknown 12 (10.4) 835 (4.1) Chemotherapy (%) Yes 107 (93.0) 17,398 (85.8) 0.036 No 8 (7.0) 2883 (14.2) Surgery (%) Yes 106 (92.2) 18,137 (89.4) 0.422 No 9 (7.8) 2144 (10.6) LN surgery scope (%) No 35 (30.4) 10,554 (52.0)  < 0.001 1–3 7 (6.1) 2244 (11.1)  ≥ 4 60 (52.2) 6907 (34.1) Sentinel LN biopsy 6 (5.2) 449 (2.2) Blank 7 (6.1) 127 (0.6) Systemic treatment Surgery sequence (%) No 80 (69.6) 6547 (32.3)  < 0.001 Systemic treatment before surgery 0 (0.0) 1480 (7.3) Systemic treatment intraoperative/after surgery 20 (17.4) 9592 (47.3) Systemic treatment both before and after surgery 15 (13.0) 2508 (12.4) Surgery both before and after systemic treatment 0 (0.0) 154 (0.8) Months from diagnosis to treatment (%)  < 1 month 97 (84.3) 12,613 (62.2)   3 months 8 (7.0) 182 (0.9) Blank 0 (0.0) 581 (2.9) Status (%) Survival 22 (19.1) 8114 (40.0)  < 0.001 Death 93 (80.9) 12,167 (60.0) Data are n (%). External validation cohort from Chengde Medical University (n = 115) vs SEER training cohort (n = 20,281). Statistical significance was assessed using chi-square tests (categorical). Baseline characteristics comparison for external validation. Poorly/ undifferentiated Data are n (%). External validation cohort from Chengde Medical University (n = 115) vs SEER training cohort (n = 20,281). Statistical significance was assessed using chi-square tests (categorical). Despite these differences, the nomogram demonstrated robust generalizability.The model’s discriminative ability remained robust in external validation, with a C-index of 0.700 (95% CI: 0.672–0.728), statistically comparable to the internal validation performance of 0.683 (Z = 1.064, P  = 0.287). Notably, time-dependent AUC analysis (Fig.  6 ) revealed a consistent enhancement in the external cohort across all time horizons—1-year AUC increased from 0.755 to 0.817 (+ 8.2%), 3-year from 0.712 to 0.760 (+ 6.7%), and 5-year from 0.732 to 0.777 (+ 6.1%). This counter-intuitive improvement in discrimination (Fig.  10 ), despite major cohort differences, suggests the model captures fundamental prognostic relationships that transcend specific population distributions. Fig. 10 Comprehensive performance comparison between internal and external validation cohorts. The figure integrates multiple evaluation metrics to assess the nomogram’s performance. It compares the time-dependent area under the curve (AUC) and calibration accuracy (mean absolute error, MAE) at 1, 3, and 5 years between cohorts, demonstrating improved discrimination and calibration in the external validation set. Additionally, it presents a forest plot summarizing key performance indicators and compares the median survival times across different risk strata, highlighting consistent stratification in both cohorts. Comprehensive performance comparison between internal and external validation cohorts. The figure integrates multiple evaluation metrics to assess the nomogram’s performance. It compares the time-dependent area under the curve (AUC) and calibration accuracy (mean absolute error, MAE) at 1, 3, and 5 years between cohorts, demonstrating improved discrimination and calibration in the external validation set. Additionally, it presents a forest plot summarizing key performance indicators and compares the median survival times across different risk strata, highlighting consistent stratification in both cohorts. To further investigate the factors that may have contributed to the improved performance in the external cohort, we performed subgroup analyses stratified by AJCC stage. In the external validation cohort, the C-index for stage II, III, and IV were 0.734 (n = 16), 0.656 (n = 87), and 0.667 (n = 12), respectively. In the internal validation cohort, the corresponding C-index values were 0.701 (n = 338), 0.642 (n = 1595), and 0.640 (n = 964). Despite the small subgroup sample sizes in the external cohort, the model maintained acceptable discriminative ability (C-index > 0.65) across all stages. The point estimates suggest that the model performs consistently across disease stages, and the overall improvement in discrimination (C-index 0.700 vs. 0.683) is likely attributable to the standardized treatment protocols and rigorous follow-up procedures in the single-center setting rather than a stage-specific effect. Calibration accuracy saw substantial improvement in the external cohort (Fig.  7 ). Mean absolute error (MAE) decreased by 33.3% at 1 year (0.034 vs. 0.051), 59.1% at 3 years (0.038 vs. 0.093), and 5.1% at 5 years (0.075 vs. 0.079). The superior calibration (Fig.  10 ), particularly for short-term predictions (1-year and 3-year MAE < 0.05), is clinically crucial and may be attributed to more standardized treatment protocols and rigorous follow-up procedures in the single-institution validation setting. Risk stratification consistency was confirmed by Kaplan–Meier analysis (Fig.  8 ), which yielded highly significant separation in both cohorts (internal log-rank P  < 2e-16; external P  = 1.58e-05). The identical median survival of 25 months in the high-risk category (Fig.  10 ) across cohorts underscores the model’s precision in identifying the most vulnerable patients. The observed variations in intermediate- and low-risk groups are expected and likely reflect inherent differences in patient mix and management between population-based registries and specialized clinical cohorts. Decision curve analysis for the external validation cohort (Fig.  9 , panels d-f) showed that the nomogram provided no meaningful net benefit for 1-year predictions beyond a 15% threshold. For 3-year predictions, the model exhibited a positive net benefit within a threshold range of approximately 30% to 45%, and for 5-year predictions, the useful range extended to 30%–50%. These results indicate that the clinical utility of the nomogram increases with longer prediction horizons, with the effective decision range widening progressively from 1-year (narrow) to 5-year (widest). In summary, the external validation not only confirmed the model’s transportability but, surprisingly, revealed enhanced performance in several key metrics (Fig.  10 ). This comprehensive assessment, demonstrating resilience to demographic and clinical variations, moderate discrimination, acceptable calibration, and meaningful risk stratification, suggests that the nomogram’s robustness and readiness for clinical application in diverse healthcare environments. Following internal validation cohort, the nomogram was benchmarked against the current clinical standard (AJCC 8th edition staging). Comparative analysis (Table 6 ) demonstrated significantly superior prognostic discrimination by the nomogram. It achieved a markedly higher concordance index (C-index: 0.683, 95% CI 0.669–0.697) compared to AJCC staging (C-index: 0.603, 95% CI 0.590–0.616; p  < 0.001 by Z-test). Time-dependent AUC analysis further confirmed this enhanced discrimination at 1 year (0.755 vs 0.617), 3 years (0.712 vs 0.623), and 5 years (0.732 vs 0.648), with all inter-system differences being statistically significant (DeLong test p  < 0.001 for each time point). This represents a 13.3% relative improvement in discriminative accuracy based on the C-index comparison. (Table 7 ). Table 6 Prognostic performance comparison between nomogram and AJCC staging system. Metric Nomogram AJCC Improvement p -value C-index (95% CI) 0.683 (0.669–0.697) 0.603 (0.590–0.616) Δ + 0.080 (+ 13.3%)  < 0.001 1-year AUC 0.755 0.617 Δ + 0.138 (+ 22.4%)  < 0.001 3-year AUC 0.712 0.623 Δ + 0.089 (+ 14.3%)  < 0.001 5-year AUC 0.732 0.648 Δ + 0.084 (+ 13.0%)  < 0.001 Median survival (months) (Low-risk group) 89 46  + 93.5%  < 0.0001 CI, Confidence interval: AUC, Area under the ROC curve. Improvement calculations: Absolute improvement (Δ) = Nomogram value–AJCC value; Relative improvement (%) = [(Nomogram—AJCC)/AJCC] × 100%. Statistical tests: C-index: Z-test; AUC: Delong test; Survival: Log-rank test. Table 7 Clinical characteristics of patients by risk group (tertile-based). Variable Low risk (n = 966) Intermediate risk (n = 969) High risk (n = 962) P -value* Age 2.00 ± 0.42 2.18 ± 0.45 2.55 ± 0.50  < 0.001 AJCC 2.17 ± 0.58 1.78 ± 0.55 1.40 ± 0.52  < 0.001 Grade 2.82 ± 0.74 3.08 ± 0.66 3.39 ± 0.61  < 0.001 Tumor size 2.28 ± 0.69 2.00 ± 0.75 1.67 ± 0.71  < 0.001 Number of examined LN 2.59 ± 0.62 1.58 ± 0.78 1.10 ± 0.46  < 0.001 Number of Positive LN 1.45 ± 0.56 2.46 ± 0.71 2.94 ± 0.24  < 0.001 Chemotherapy 1.08 ± 0.28 1.11 ± 0.31 1.21 ± 0.41  < 0.001 Surgery 1.00 ± 0.00 1.01 ± 0.08 1.30 ± 0.46  < 0.001 Systemic treatment Surgery sequence 2.72 ± 0.87 2.58 ± 1.09 1.99 ± 1.13  < 0.001 Months from diagnosis to treatment 1.21 ± 0.42 1.33 ± 0.49 1.73 ± 0.84  < 0.001 Data are presented as mean ± standard deviation. * P ‑values were calculated using the Kruskal‑Wallis test for continuous variables. All comparisons showed statistically significant differences (p < 0.001). The sample sizes (n) for each risk group are indicated in the column headers. Prognostic performance comparison between nomogram and AJCC staging system. 0.683 (0.669–0.697) 0.603 (0.590–0.616) Median survival (months) (Low-risk group) CI, Confidence interval: AUC, Area under the ROC curve. Improvement calculations: Absolute improvement (Δ) = Nomogram value–AJCC value; Relative improvement (%) = [(Nomogram—AJCC)/AJCC] × 100%. Statistical tests: C-index: Z-test; AUC: Delong test; Survival: Log-rank test. Clinical characteristics of patients by risk group (tertile-based). Data are presented as mean ± standard deviation. * P ‑values were calculated using the Kruskal‑Wallis test for continuous variables. All comparisons showed statistically significant differences (p < 0.001). The sample sizes (n) for each risk group are indicated in the column headers. We also compared our nomogram with two previously published prognostic models for epithelial ovarian cancer (EOC). Cheng et al. developed a 12-variable nomogram for elderly EOC patients (age ≥ 60 years) using the SEER database, reporting a validation C-index of 0.751 for overall survival 14 . Lee et al. constructed a six-variable nomogram for platinum-resistant ovarian cancer, with validation C-statistics of 0.59 (PENELOPE cohort) and 0.67 (AURELIA bevacizumab-chemotherapy cohort) 15 . Our 10-variable nomogram achieved an internal validation C-index of 0.683 (95% CI: 0.669–0.697) and an external validation C-index of 0.700. In terms of predictor count and discriminative ability, our model uses a moderate number of variables (10) and yielded a C-index (0.683–0.700) that is higher than that of the Lee model (0.59–0.67) but lower than that of the Cheng model (0.751). The Cheng model was developed on a specific elderly population (n = 5,588), and the Lee model was derived from platinum-resistant clinical trial cohorts. In contrast, our model was developed and validated on a broader stage II–IV EOC population (age 18–85 + years; n = 20,281 for training, n = 115 for external validation). Furthermore, our model provides time-dependent AUCs at 1, 3, and 5 years (0.755, 0.712, 0.732) and corresponding Brier scores (0.105, 0.204, 0.226), metrics not reported in the Cheng or Lee studies. These additional metrics allow for a more comprehensive evaluation of predictive performance across multiple time points..

Materials

The study cohort was selected from the SEER Program ( www.seer.cancer.gov ) using the SEER*Stat Database (Incidence—SEER Research Data, 17 Registries, November 2022 Submission [2000–2020], version 8.4.3). We identified 93,133 patients who were pathologically diagnosed with ovarian cancer between 2004 and 2020 based on the International Classification of Diseases for Oncology, Third Edition (ICD-O-3), with the primary site coded as C56.9 (ovary). The study was conducted in accordance with the Declaration of Helsinki and approved by the Ethics Committee of the Affiliated Hospital of Chengde Medical University (No. CYFYLL2022424). Due to the retrospective nature of this study, requirement for informed consent was waived by the Ethics Committee of the Affiliated Hospital of Chengde Medical University. Inclusion Criteria: Patients with a confirmed diagnosis of EOC based on surgical and pathological findings; diagnosis at AJCC Stage II–IV; and complete clinical records, including information on age, histological subtype, race, marital status, disease stage, and treatment regimen. Exclusion Criteria: Patients diagnosed with other primary malignancies; missing data on months of survival; death from non-EOC causes; or absence of tumor-specific survival data (Fig.  1 ). An external validation cohort of 128 patients who met the same inclusion criteria was drawn from Chengde Medical University for the period 2014–2020. The pathological diagnostic criteria for this external cohort were based on the 8th edition of the AJCC staging system; all cases were confirmed by two independent pathologists. Follow-up consisted of outpatient visits and telephone interviews: every 3 months for the first 2 years, every 6 months for the next 3 years, and annually thereafter. Follow‑up was administratively censored on December 31, 2025. The median follow‑up for censored patients was 62.0 months (range: 4–85 months). The number of lost-to-follow-up patients was 13 (10.2%), with reasons including relocation (n = 8) and inability to contact (n = 5). The remaining 115 patients were included in the analysis (Fig.  1 B). Fig. 1 Flow diagrams of patient selection. Note: ( A ) SEER database screening process; ( B ) External validation cohort enrollment and follow-up. Flow diagrams of patient selection. Note: ( A ) SEER database screening process; ( B ) External validation cohort enrollment and follow-up. Seventeen clinicopathological variables were selected as candidate prognostic factors: age (< 40/40–69/ ≥ 70 years; clinical rationale: fertility preservation and geriatric oncology thresholds), race, marital status, histological subtype, tumor differentiation grade, AJCC stage, preoperative CA125 levels (categorized), tumor laterality, tumor size (< 10/ ≥ 10 cm; clinical rationale: surgical complexity benchmark), number of positive lymph nodes, number of lymph nodes examined, surgical intervention, residual tumor size following cytoreductive surgery, extent of lymph node dissection, administration of chemotherapy, sequence of systemic therapy and surgery, and the interval between diagnosis and initiation of treatment. All statistical analyses were performed using R software (version 4.3.2) in the RStudio environment. The patient cohort was randomly split into training (70%), tuning (20%), and internal validation (10%) sets via stratified sampling by survival status (Status). Specifically, using the createDataPartition function from the caret package (version 7.01) with a fixed seed (123), we allocated 70% of the data to the training set, stratified by Status. From the remaining 30%, we allocated two‑thirds (20% of total) to the tuning set and one‑third (10%) to the internal validation set, again stratified by Status. Intergroup feature balance was assessed using standardized mean differences (SMDs); all SMDs were below 0.05 (well below the 0.1 balance threshold), indicating that the baseline feature distributions across the subsets were highly balanced. Given substantial missingness in key prognostic variables (CA125: 42.5%; tumor size: 30.4%; positive lymph node count: 52.1%), we validated the Missing at Random (MAR) mechanism by demonstrating non-significant associations between missing indicators (binary variables: 1 = missing, 0 = observed) and survival outcomes in Cox regression models (all P  > 0.05; Table 2 ). The 17 candidate prognostic factors were all categorical variables. We then performed multiple imputation by chained equations (MICE) using the R package ‘mice’ (version 3.18.0), with 20 imputed datasets (m = 20) and 20 iterations (maxit = 20), under a random seed of 123. As no continuous variables were included, default imputation methods were applied: logistic regression for binary variables and multinomial regression for polytomous variables. A three‑tiered sensitivity analysis framework was implemented. First, a complete case analysis (CCA) excluded any cases with missing values in the 17 prognostic factors. Second, MICE was applied to all cases to generate 20 imputed datasets. Third, an extended missing indicator analysis was performed: for variables with significant missing‑survival associations (surgery and residual lesion), missing indicators were added as covariates; for the three key variables with high missing rates (CA125, tumor size, and positive lymph node count), median imputation was used, and both the imputed values and their missing indicators were entered into multivariate Cox models. In addition, to directly compare MICE with a simpler approach, mode/median imputation (mode for categorical variables, median for continuous variables) was performed on the training set. To visually compare the effect estimates across the three methods, a forest plot was constructed (Fig.  2 ). Hazard ratios (HRs) and 95% confidence intervals (CIs) for the 17 covariates were extracted from the CCA, MICE, and mode/median imputation models and plotted using the R package ggplot2. In the forest plot, blue squares (CCA), orange circles (MICE), and green triangles (mode/median imputation) represent the estimates, with horizontal lines indicating 95% CIs and a vertical dashed line at HR = 1.0. Model performance was evaluated using Harrell’s C‑index, time‑dependent AUC at 1, 3, and 5 years, and Brier scores (Table 3 ). Model training on the imputed training set avoids bias from missing values. Fig. 2 Forest plot comparing hazard ratios among three imputation methods (CCA, MICE, and mode/median). Forest plot comparing hazard ratios (HRs) and 95% confidence intervals (CIs) among complete case analysis (CCA), multiple imputation (MI), and mode/median imputation for 17 covariates. Each covariate is represented by three symbols: blue square for CCA, orange circle for MI, and green triangle for mode/median imputation. Horizontal lines represent 95% CIs. The vertical dashed line indicates the null effect (HR = 1.0). Abbreviations: CCA, complete case analysis; MI, multiple imputation; HR, hazard ratio; CI, confidence interval. Forest plot comparing hazard ratios among three imputation methods (CCA, MICE, and mode/median). Forest plot comparing hazard ratios (HRs) and 95% confidence intervals (CIs) among complete case analysis (CCA), multiple imputation (MI), and mode/median imputation for 17 covariates. Each covariate is represented by three symbols: blue square for CCA, orange circle for MI, and green triangle for mode/median imputation. Horizontal lines represent 95% CIs. The vertical dashed line indicates the null effect (HR = 1.0). Abbreviations: CCA, complete case analysis; MI, multiple imputation; HR, hazard ratio; CI, confidence interval. The multiply imputed training datasets were systematically employed to develop the prognostic model through a multi-stage analytical framework. First, univariate Cox proportional hazards regression analyses were conducted to assess preliminary associations between each candidate prognostic factor and overall survival (OS), with statistical significance established at α = 0.05 (Table 4 ). No adjustment for multiple comparisons (e.g., Bonferroni correction) was applied at this initial screening stage. To mitigate overfitting, we subsequently performed Least Absolute Shrinkage and Selection Operator (LASSO) regression with tenfold cross-validation on a dedicated tuning set (Fig.  3 ). The optimal penalty parameter (λ) was determined by cross‑validation. We selected λ using the 1‑standard‑error rule (λ.1se)—the largest λ value within one standard error of the minimum cross‑validated partial likelihood deviance—rather than the value yielding the minimum error (λ.min). This choice promotes a more parsimonious model with fewer non‑zero coefficients, prioritizing simplicity and generalizability over marginal improvements in fit, which is crucial for clinical applicability. The coefficient trajectory and cross‑validation error curves (Fig.  3 ) were examined to guide this selection. Variables that were both statistically significant in univariate analysis ( p  < 0.05) and selected by the LASSO algorithm were retained for the final model. These consensus predictors were then incorporated into a multivariate Cox proportional hazards model fitted across all multiply imputed datasets, with Rubin’s rules applied to pool hazard ratio (HR) estimates and corresponding 95% confidence intervals (CIs). Model diagnostics included assessment of proportional hazards assumptions and examination of variance inflation factors to ensure statistical validity. The final independent prognostic factors for Stage II-IV ovarian cancer were visually summarized through a forest plot (Fig.  4 ), with clinical relevance evaluated against established biological mechanisms. Fig. 3 LASSO regression analysis. Notes: ( A ) LASSO coefficient distribution map; LASSO coefficient distribution of all variables. ( B ) Variables determined through LASSO analysis. Fig. 4 Multivariate regression model of women with EOC in the training set. *p < 0.05, **p < 0.01, ***p < 0.001, Statistically significant difference, the significance increased successively. LASSO regression analysis. Notes: ( A ) LASSO coefficient distribution map; LASSO coefficient distribution of all variables. ( B ) Variables determined through LASSO analysis. Multivariate regression model of women with EOC in the training set. *p < 0.05, **p < 0.01, ***p < 0.001, Statistically significant difference, the significance increased successively. A prognostic nomogram for stage II-IV ovarian cancer patients was developed using independent prognostic factors selected via univariate/multivariate analyses and LASSO regression on the imputed training set and tuning set. Each variable in the nomogram was assigned a score based on the clinical characteristics of each patient. The total score was calculated by summing the individual scores, enabling the estimation of 1-year, 3-year, and 5-year OS probabilities. To evaluate the predictive performance of the nomogram, several statistical methods were applied: the concordance index (C-index) was used to assess the model’s discriminatory ability; calibration curves were used to evaluate the agreement between predicted and observed survival probabilities; receiver operating characteristic (ROC) curves were used to measure the model’s discriminative performance; Kaplan–Meier (KM) survival curves were used to compare survival outcomes between high-risk and low-risk groups; and decision curve analysis (DCA) was used to determine the clinical utility and net benefit of the nomogram in supporting treatment decision-making. The model underwent rigorous external validation using an independent cohort from Chengde Medical University (n = 115). Baseline characteristics between the SEER-derived training cohort and this external validation cohort were compared using chi-square tests. Continuous variables, including age, CA125, and tumor size, were categorized prior to analysis and treated as categorical variables. Statistical significance was set at p < 0.05. In this external cohort, all performance metrics were reassessed to verify generalizability. Furthermore, the nomogram was benchmarked against the AJCC 8th edition staging using the internal validation cohort (n = 2897) by quantitatively comparing their C-index values (statistical significance assessed via Z-test) and time-dependent AUC values at 1, 3, and 5 years (statistical significance assessed via Delong test). All analyses were performed in R version 4.3.2.

Conclusion

Through a retrospective SEER analysis, this study identified key prognostic factors for Stage II–IV EOC and incorporated them into a predictive nomogram for overall survival. Crucially, external validation on a distinct, single-institution cohort confirmed the model’s discriminative accuracy (C-index: 0.700) and demonstrated good calibration, indicating that the nomogram captures fundamental prognostic relationships beyond database‑specific patterns. Although the C‑index indicates good (rather than excellent) performance, this externally validated tool can support clinical decision‑making and improve prognostic counseling for patients with advanced‑stage EOC.

Discussion

Ovarian cancer remains the most lethal gynecologic malignancy. Owing to the low early detection rates and resistance to treatment, most patients are diagnosed at advanced stages (Stage II–IV), often presenting with metastatic progression by the time symptoms appear 16 . For this study, we employed the staging system of the American Joint Committee on Cancer (AJCC), 8th edition, which incorporates the widely adopted International Federation of Gynecology and Obstetrics (FIGO) staging criteria for ovarian cancer within its TNM framework 5 , 17 , 18 . Primary ovarian cancer typically spreads through direct invasion of adjacent organs, such as the uterus and surrounding connective tissue, followed by transcoelomic, lymphatic, and hematogenous dissemination. Transcoelomic spread is the most common pathway, due to the absence of a physical barrier between the tumor and peritoneal fluid, with the omentum being the most frequent site of metastasis 19 . In advanced stages, peritoneal carcinomatosis and ascites are commonly observed, while the involvement of the liver and lungs is less frequent 20 . Global efforts to improve ovarian cancer prevention and management have intensified with advancements in medical technology 21 . For patients with EOC in the early stage or medically eligible advanced stage, the standard approach typically involves primary debulking surgery (PDS) followed by adjuvant chemotherapy (ACT), depending on individual clinical circumstances. Studies have shown that neoadjuvant chemotherapy (NACT) is associated with lower surgical morbidity than ACT but yields comparable outcomes in terms of progression-free survival and OS. In cases involving pleural effusion, hypoalbuminemia, or poor performance status, NACT may be administered first, followed by interval debulking surgery (IDS) and ACT 22 . According to the American Cancer Society, the overall 5-year survival rate for EOC is less than 30% 23 , 24 . However, early-stage detection is associated with markedly improved outcomes, as Stage I patients have a 5-year survival rate exceeding 90% 25 . Given the heterogeneity in treatment response and survival across disease stages, our study concentrated on Stage II–IV EOC patients and evaluated multiple independent prognostic factors. By developing and validating a predictive survival model, we aim to provide a clinically meaningful tool to support individualized treatment strategies, improve risk stratification, and enhance prognostic assessment in advanced-stage EOC. These findings underscore the value of personalized prognostic evaluation in facilitating more informed clinical decision-making for Stage II–IV EOC patients. Our study identifies age as an independent risk factor for EOC. According to data from the ACS 26 , ovarian cancer may occur at any age but is more frequently diagnosed after the age of 40 years, with the majority of cases occurring in women aged ≥ 65 years. These statistics suggest that the risk of developing ovarian cancer increases with age. In line with this, our nomogram-based survival predictions demonstrate that older age is associated with poorer survival outcomes among EOC patients. This association may be attributed to the age-related decline in physiological function, which reduces tolerance for aggressive treatments 27 , as well as the cumulative damage to ovarian epithelium resulting from prolonged ovulatory cycles, thereby elevating the risk of malignant transformation 28 . The histological types of EOC are classified based on distinct genomic and epigenomic profiles. Although high-grade serous carcinoma (accounting for ~ 70% of new cases) represents the most common aggressive subtype, and histologic types such as endometrioid carcinoma (ENDOC, 15%) and clear cell carcinoma (CCOC, 12%) carry established prognostic value 29 , histological subtype was excluded during variable selection in our study. Several factors account for this. First, our cohort is limited to stage II–IV disease, in which high-grade serous carcinoma (HGSC) predominates; the relatively small numbers of non-HGSC cases (e.g., ENDOC, CCOC) limits the statistical power to detect their independent effect on survival after adjustment for grade, stage, and treatment. Second, tumor differentiation grade—a strong correlate of biological aggressiveness—captures much of the prognostic information that might otherwise be attributed to histologic subtype, resulting in redundancy when both variables are considered together. Third, to maintain model parsimony and improve generalizability across diverse clinical settings (where routine histotyping may vary), we prioritized variables with consistently strong prognostic value across all stage II-IV patients, such as differentiation grade, age, and treatment parameters. This approach aligns with several published SEER-based nomograms for EOC, which also omitted detailed histological subtyping 30 , 31 . Conversely, tumor differentiation grade was retained as a critical factor influencing survival outcomes in Stage II–IV epithelial ovarian cancer. This aligns with Ramalinga P.’s findings 32 that both histological subtype and tumor differentiation grade strongly correlate with metastasis risk, disease progression, chemotherapy response, and overall prognosis, while our study reinforces the pivotal role of differentiation grade in prognostic stratification. One study 33 found that only 13% of patients with serous ovarian carcinoma were diagnosed at an early stage. Among this group, the 10-year survival rate was approximately 55% for patients in early stage and only 15% for those in advanced stage of the disease. Notably, more than 65% of long-term survivors initially presented with advanced-stage disease, underscoring the aggressive nature of serous carcinoma. Our findings are consistent with those of the study by Lauren C. et al. 34 , which showed that, regardless of stage or tumor grade, ENDOC and serous carcinoma were associated with significantly higher survival rates than CCOC, mucinous carcinoma, and other rare subtypes. ENDOC, which is often considered to originate from endometriosis, has been shown to have the most favorable survival prognosis 35 . Additionally, a study by Ling Tang et al. 36 demonstrated that ENDOC is typically less aggressive and more frequently diagnosed at an earlier stage. This earlier detection may be linked to elevated estrogen levels, chronic inflammation, increased iron levels, and detectable tumor markers and imaging findings—many of which are identifiable through routine clinical testing. Our nomogram-based survival model reflects these clinical patterns: higher tumor differentiation grades are associated with improved 1-year, 3-year, and 5-year OS rates, while lower differentiation grades correspond to poorer outcomes. These results reinforce the importance of tumor grading in effective risk stratification. Although CA125 remains a cornerstone biomarker for epithelial ovarian cancer (EOC) screening—with our study confirming elevated preoperative levels as an independent prognostic factor aligning with Morales-Vásquez et al. 37 —its exclusion from our final model stemmed from three evidence-based considerations. First, substantial data gaps (42.5% missing values among 20,281 cases) coupled with stage-distribution skewness (88.8% Stage III–IV vs. 11.2% Stage II) introduced survival prediction bias, reflecting real-world limitations in early-stage detection where CA125 holds maximal diagnostic utility 38 . Second, non-oncological elevations during menstruation or in benign conditions (endometriosis, cardiovascular diseases etc.) inherently compromise its specificity, generating false-positive signals that undermine reliability as a standalone diagnostic tool 39 . Third, in advanced-stage-dominated cohorts like ours, CA125’s prognostic discriminatory power becomes markedly attenuated due to universally elevated levels, diminishing its stratification capacity. Regarding residual tumor burden, we acknowledge its strong prognostic value but omitted it due to confounding effects and its inherent association with treatment sequence, which could obscure the independent contribution of other variables. Consequently, while acknowledging CA125’s established role in EOC management, our modeling prioritizes variables with robust completeness and stage-agnostic prognostic validity to optimize clinical generalizability. This pragmatic approach inevitably excludes some well-established prognostic factors, which may limit direct comparability with models that include CA125 or detailed residual lesion data. Nevertheless, it enhances the model’s applicability across routine clinical settings where such information is often incomplete. The clinical utility of our model is supported by decision curve analysis, demonstrating positive net benefit within relevant threshold ranges. Looking forward, integration of serial CA125 kinetics (e.g., the KELIM score during neoadjuvant chemotherapy) and validated molecular markers (e.g., BRCA1/2 mutation status, homologous recombination deficiency) could further improve predictive accuracy, and such additions would be feasible in prospective or resource-rich settings. Our study suggests that survival prognosis in EOC patients is also influenced by tumor size, the number of positive lymph nodes, and the number of lymph nodes surgically examined. An increased number of lymph nodes removed and assessed during surgery reduces the likelihood of missed diagnoses and increases the probability of positive node detections, which subsequently affects both treatment decisions and survival outcomes. According to the AJCC staging system 5 , when malignancy is confined to the ovary or fallopian tube, a lower number of positive lymph nodes is typically associated with an earlier disease stage and more favorable prognosis. However, most EOC cases are diagnosed at an advanced stage, often characterized by infiltration, metastasis, and widespread dissemination. These factors complicate accurate estimation of the primary tumor size and contribute to missing or incomplete data, which, in turn, correlates with poorer prognosis—findings that are consistent with our results. In the overall cohort, larger tumor size (> 10 cm) was unexpectedly associated with improved survival (HR = 0.92). Subgroup analysis stratified by AJCC stage showed that this protective effect was statistically significant only in stage III disease (HR = 0.84), with no significant association observed in stage II or IV. This stage-specific pattern indicates that the overall hazard ratio (HR = 0.92) was largely driven by stage III patients. This paradoxical finding warrants further discussion of three potential explanations. First, measurement bias within the SEER database may contribute. Larger tumors are generally measured and documented more accurately, whereas smaller tumors may be underreported or associated with incomplete staging, potentially biasing the hazard ratio to suggest a spurious protective effect. Second, confounding by resectability is another key factor.. In stage III ovarian cancer, bulky but localized tumor masses are often more amenable to complete (R0) resection than diffuse small lesions, which are frequently associated with peritoneal carcinomatosis and suboptimal cytoreduction. Thus, the apparent survival benefit may reflect surgical resectability rather than tumor size per se. Third, treatment heterogeneity—specifically the more frequent use of neoadjuvant chemotherapy (NACT) in stage III bulky disease 40 —may also contribute. Patients with large tumors are more likely to receive NACT followed by interval debulking surgery, which can downstage the disease and improve survival outcomes compared with primary debulking of smaller but more disseminated lesions 41 . These mechanisms help explain why the apparent protective effect of larger tumor size was observed only in stage III patients. Interestingly, some studies have reported that EOC tumors detected at an earlier stage tend to have a larger diameter than those identified later 42 . This may be due to a subset of larger tumors causing noticeable symptoms such as pelvic pain, urinary tract infections, or compression of adjacent organs, which facilitates earlier clinical detection 43 —a mechanism that could also contribute to the paradoxical survival benefit observed in stage III patients. Despite this heterogeneity, tumor size was retained in the final nomogram because it remained a significant independent predictor in the overall (all-stage) multivariable model; excluding it would reduce predictive accuracy for the intended stage II–IV population. We fully acknowledge that residual tumor burden (R0/R1/R2 classification) post-cytoreduction is a critical determinant of first-line EOC treatment strategy 42 , 43 . Its exclusion from the final nomogram therefore constitutes a major limitation of our study. This variable was excluded from the final model for three principled reasons: The inclusion of “macroscopic residual (unspecified size)” and “missing data (Blank)” categories introduced confounding effects in survival analysis; The efficacy equivalence debate between NACT-IDS and PDS 44 reveals that residual tumor is not an independent prognostic factor—as demonstrated by Vergote et al. 45 where PDS outperformed NACT-IDS in patients with ≤ 5 cm pelvic metastases but the reverse held for Stage IV disease, proving residual status is a consequence of treatment decisions rather than a  causal  biomarker; Since residual grading inherently depends on surgical completeness, and our model focuses on preoperative prognostic stratification to guide initial therapy selection (PDS vs. NACT pathways), prioritizing primary tumor characteristics (e.g., differentiation grade, tumor size) offers greater clinical immediacy. Crucially, in advanced-stage EOC, complete cytoreduction (R0) is often not achievable, and residual disease is frequently present regardless of treatment choices. Therefore, the exclusion of residual lesion variables does not compromise the model’s validity for its intended preoperative risk assessment purpose. Nonetheless, future studies incorporating detailed, standardized residual tumor data are strongly encouraged. The primary treatment modalities for EOC include surgery and adjuvant chemotherapy, followed by maintenance therapy, targeted therapy, and immunotherapy 46 . In our study, the treatment regimen included surgery, chemotherapy, and systemic therapy, with particular emphasis on residual tumor size following cytoreductive surgery and the role of lymph node dissection. According to our nomogram, when other variables were controlled for, patients with Stage II–IV EOC who underwent surgery and chemotherapy had a significantly lower risk of poor prognosis than those who did not receive these interventions. In particular, surgical treatment following diagnosis was associated with a marked improvement in OS. With respect to systemic therapy, current SEER database guidelines define it as encompassing chemotherapy, hormone therapy, surgery, radiation, endocrinotherapy, biologic response modifiers, and immunotherapy 47 . Our results indicate that intraoperative or postoperative systemic therapy, as well as systemic therapy administered both before and after surgery, serves as an independent prognostic factor in Stage II–IV EOC ( p  < 0.001). Notably, patients who received systemic therapy following cytoreductive surgery experienced particularly favorable outcomes 48 . For many years, the standard first-line systemic treatment for EOC has been intravenous chemotherapy with carboplatin and paclitaxel, administered every three weeks for a total of 6 to 8 weeks 49 . However, D. K. Armstrong et al. 50 noted that certain elderly patients (aged ≥ 70 years) and those with comorbidities may not tolerate standard chemotherapy regimens. Severe toxicity and premature discontinuation of adjuvant chemotherapy can negatively impact survival. In such cases, dose reductions of paclitaxel and carboplatin should be considered based on an individualized assessment of treatment tolerance. Furthermore, low-grade serous carcinomas and low-grade endometrioid carcinomas may benefit from hormone therapy as a maintenance strategy, particularly with aromatase inhibitors 48 . Ultimately, treatment plans should be tailored to the histopathological subtype, disease stage, and patients’ specific clinical presentation 51 . While acknowledging numerous SEER-based prognostic models for ovarian cancer 52 , our study advances the field through three key innovations. First, benchmarking against the AJCC 8th edition staging system showed that our nomogram provides a statistically significant improvement in discriminative ability, with a relative C-index gain of over 13% ( p  < 0.001). This gain, validated internally and externally, underscores the added value of incorporating a broader set of clinical and treatment variables beyond stage alone. Second, our 10-variable nomogram achieves greater parsimony than existing models (e.g., Cheng et al.’s 12-variable model and Lee et al.’s 6-variable model) while maintaining moderate yet clinically meaningful discrimination (C-index 0.683–0.700). Compared with Cheng’s model, which was developed on a restricted elderly population (age ≥ 60 years) and reported a validation C-index of 0.751, our model applies to a broader stage II-IV population across all ages. Compared with Lee’s model (validation C-statistics 0.59–0.67), our model achieves higher or comparable discrimination and is applicable to general EOC patients, not limited to platinum-resistant disease. The balance of parsimony and performance is enabled through strategic integration of dynamic clinical predictors—treatment sequence, time from diagnosis to treatment initiation, and lymph node counts—which have been underrepresented in some SEER models despite their proven prognostic value 53 , 54 . Third, we objectively frame performance below the ‘excellent’ threshold (C-index > 0.8), with CA125 exclusion representing a deliberate feasibility trade-off; future integration of molecular marker (e.g., BRCA/HRD status) could yield C-index gains 55 . The clinical importance of the dynamic predictors identified in our model deserves further discussion. A shorter interval (< 1 month) from diagnosis to treatment initiation was independently associated with better survival. This principle is reflected in multiple clinical guidelines, which broadly recommend initiating therapy as soon as possible after diagnosis to avoid disease progression 56 , 57 . A SEER-based study supports this approach, reporting a 5-year overall survival probability of 61.4% in patients who received immediate treatment (within one month), compared with 36.4% and 34.8% in those with intermediate or longer delays 58 . At the same time, a more nuanced picture has emerged in the literature. For example, a 2026 cohort study noted that very short diagnostic intervals can sometimes be associated with poorer survival, an effect partly explained by the “wait time paradox”—the observation that patients with especially aggressive disease are often diagnosed more quickly but have inherently worse outcomes 59 . These seemingly contradictory findings underscore the heterogeneity of the patient population. Nevertheless, when clinical judgment is applied appropriately, minimizing diagnostic-to-treatment delays remains a clinically actionable factor to improve prognosis in advanced-stage EOC. In our dataset, the months‑from‑diagnosis‑to‑treatment variable contained a “Blank” category (2.9%) representing undocumented intervals. It is widely recognized that such missing entries cannot be simply treated as valid prognostic categories. Therefore, we treated these entries as missing data and performed a sensitivity analysis using multiple imputation (MICE). A prior SEER ovarian cancer study reported that missing data rates for different variables in the SEER dataset range from 0 to 36.11%, and the results between the complete dataset and the MI datasets were similar 60 . In our sensitivity analysis, excluding these 2.9% of missing records via multiple imputation resulted in no change in the C‑index (0.693 in both the original and the imputed datasets) and no meaningful alteration of any hazard ratio. Given the consistency between the imputed and original results, we elected to proceed with the original dataset (retaining the “Blank” category as originally recorded) for constructing the final nomogram, rather than using the imputed datasets. This decision was based on the principle that imputation should not alter the substantive conclusions; when sensitivity analysis confirms robustness, retaining the original data enhances transparency and reflects real-world data collection. This approach aligns with the practice of a prior SEER-based nomogram study, which also confirmed that no statistical difference existed between the complete-case and imputed data before deciding which dataset to use for further analyses 60 . In real-world clinical data, a “Blank” entry for time from diagnosis to treatment often does not occur at random; it may reflect rapid clinical deterioration, incomplete workup, or patient frailty—all of which are intrinsically linked to poorer prognosis. Therefore, treating “Blank” as a separate category can capture this prognostic signal rather than ignoring it. Moreover, evidence from a large SEER ovarian cancer cohort has shown that missing information can correlate with worse outcomes, and complete-case deletion may introduce selection bias 61 . In that study, over a third of cases had unknown grade and 10–20% had unknown stage, with missing information itself being significantly associated with the worst survival 61 . Deleting the 2.9% of cases with undocumented intervals would discard 581 patients without statistical gain but would reduce transparency and risk removing a patient subgroup with potentially distinct prognostic implications. Therefore, we retained the original data for transparency while maintaining scientific rigour through a well-established sensitivity analysis framework. The sequence of systemic therapy and surgery represents another dynamic predictor with clear clinical relevance. Current guidelines endorse two first-line strategies for advanced EOC: primary debulking surgery (PDS) followed by adjuvant chemotherapy, or neoadjuvant chemotherapy (NACT) followed by interval debulking surgery 56 , 57 . Our model confirmed that either sequence, when appropriately selected, significantly improves survival compared with no systemic therapy. This aligns with the individualized, multidisciplinary approach recommended by current guidelines. Moreover, the category “Surgery both before and after systemic treatment” accounted for only 0.76% of the training cohort. Despite its small size, this category was retained because it represents a clinically distinct treatment pattern—systemic therapy interrupted by surgery, often reflecting complex or refractory disease (e.g., salvage surgery or interval debulking after progression). The hazard ratio for this category remained stable and significant in multivariate analysis (HR = 0.62, 95% CI 0.48–0.80, p  < 0.001), and a sensitivity analysis excluding these patients did not alter the model’s C‑index or any other hazard ratio. Thus, the original classification was kept for completeness and clinical transparency, with the small sample size noted as a limitation. The relative prognostic value of clinical nomograms versus molecular biomarkers in ovarian cancer remains debated. Some studies report that molecular markers (e.g., BRCA1/2 mutation status, homologous recombination deficiency) offer superior prognostic discrimination and can guide targeted therapies, including PARP inhibitors 62 , 63 . In contrast, other work indicates that well-constructed clinical models based on routine variables can achieve performance comparable to molecular signatures, especially in resource-limited settings 64 . Clinical nomograms integrate conventional clinicopathological factors and provide the advantages of accessibility, standardization, and low cost; however, they may overlook the molecular mechanisms underlying tumor heterogeneity. Conversely, molecular markers reflect tumor biology and can offer higher predictive accuracy, but they face challenges such as insufficient assay standardization, high cost, and marked subgroup specificity. Differences in variable dimensions, data sources, and statistical models between the two approaches can lead to inconsistent prognostic stratification for the same patient cohort. Molecular markers may compensate for nomograms’ limitations in predicting chemotherapy resistance and long-term survival, whereas nomograms provide a rapid, cost-effective baseline risk assessment to support clinical decision-making. Our nomogram, which uses only clinical variables, provides a cost-effective and pragmatic alternative, integrating molecular markers in the future may further enhance its predictive accuracy 65 . To further assess the model’s clinical applicability, we evaluated its discriminative ability in clinically relevant subgroups of the external cohort (n = 115). The model achieved a C-index of 0.857 in elderly patients (≥ 70 years, n = 9), 0.667 in stage IV patients (n = 12), 0.682 in serous carcinoma (n = 107), and 0.778 in non-serous subtypes (n = 8). Due to the very small sample sizes in the elderly and non-serous subgroups (each < 10), these estimates are imprecise and should be interpreted with caution. The model’s performance in the larger subgroups (< 70 years, stage II-III, serous) was consistent (C-index ~ 0.66–0.68), supporting its generalizability across most advanced EOC patients. Larger multi-center studies are needed to validate these subgroup findings. This study has several limitations that warrant careful consideration alongside its strengths. The population-based retrospective design may introduce selection bias, while the 7:2:1 data partitioning primarily evaluates intrinsic consistency rather than robustness across diverse healthcare systems. The exclusive reliance on SEER database presents additional constraints, including binary chemotherapy records lacking details on specific regimens and ambiguous measurement standards for variables like tumor size, which may systematically bias prognostic estimates. Moreover, the SEER database does not contain information on molecular markers (e.g., BRCA1/2 mutation status, homologous recombination deficiency, microsatellite instability), subsequent treatment details such as maintenance therapy (e.g., PARP inhibitors) and targeted therapy, nor does it include detailed cytoreduction status (R0/R1/R2) — all of which are increasingly recognized as key factors influencing EOC prognosis. These omissions represent important limitations of our study. Despite these limitations, the successful external validation substantially mitigates many concerns. The model demonstrated robust performance in the Chengde Medical University cohort (n = 115), achieving a C-index of 0.700 and excellent time-dependent AUC values (0.817, 0.760, and 0.777 for 1-, 3-, and 5-year survival, respectively). Notably, the model’s performance improved in this distinct clinical setting, with exceptional calibration (MAE: 0.034, 0.038, and 0.075 for 1-, 3-, and 5-year predictions) and consistent risk stratification observed in Kaplan–Meier analysis. However, the higher AUC in the external cohort (e.g., 1-year AUC 0.817 vs. 0.755 internally) is largely attributable to the different stage distribution: stage IV patients constituted only 10.4% of the external cohort compared to 32.6% in the SEER internal validation set, making the external population easier to discriminate. The better calibration (lower MAE) may reflect more standardized treatment protocols, more complete lymph node dissection, and more rigorous follow-up procedures in the single-center external validation cohort. This enhanced performance in external validation, coupled with sustained clinical utility in decision curve analysis, suggests that the nomogram captures fundamental biological relationships rather than database-specific artifacts. However, the external validation cohort was derived from a single institution with a modest sample size (n = 115) and a distinct clinical profile (predominantly Asian, younger age, higher proportion of stage II disease), which limits the generalizability of our findings to broader, more heterogeneous populations. Furthermore, the C-index of our nomogram (0.683–0.700) indicates only moderate predictive accuracy and does not reach the “excellent” threshold (C-index > 0.8). Therefore, while this single-institution validation represents a valuable step confirming the model’s transportability, these findings require confirmation through larger multi-center studies. To improve predictive accuracy to an excellent level, future research should focus on multi-center international validation incorporating granular clinical data (e.g., detailed surgical records, CA125 kinetics) and validated molecular biomarkers (e.g., BRCA/HRD status, MSI). The model’s performance in resource-limited regions also remains uncertain.

Introduction

Ovarian cancer is the leading cause of mortality among gynecologic malignancies globally 1 . Among its subtypes that originate from different ovarian tissues, including epithelial, stromal, and germ cells, epithelial ovarian cancer (EOC) accounts for approximately 90% of all cases 2 . According to World Health Organization’s histopathological classification of EOC, updated in 2020, the major subtypes are high-grade serous carcinoma, low-grade serous carcinoma, mucinous carcinoma, endometrioid carcinoma, clear cell carcinoma, mixed carcinoma, and several rare variants such as Brenner tumors, mesonephric-like carcinoma, and dedifferentiated carcinoma 3 . These histotypes exhibit markedly different prognostic profiles: high‑grade serous carcinoma carries the worst prognosis, whereas endometrioid and mucinous carcinomas are associated with more favorable outcomes. Unlike other gynecologic malignancies, ovarian cancer typically presents with nonspecific early symptoms, resulting in delayed diagnosis. At the time of symptom onset, most patients have already progressed to metastatic disease (Stage II–IV), which results in a low survival rate 4 . Ovarian cancer staging mainly depends on intraoperative exploration to determine the extent of disease spread. According to the American Joint Committee on Cancer (AJCC) staging system 5 , Stage II–IV ovarian cancer is defined by tumor dissemination beyond the ovaries or fallopian tubes into pelvic or extrapelvic regions. Thus, disease staging is a critical determinant of both treatment strategies and prognosis 6 . Since the early 1980s, the mortality rate of ovarian cancer has been substantially high 7 . However, its incidence has decreased in high-income countries, including the United States, France, Germany, and the United Kingdom, suggesting that specific environmental or genetic risk factors are closely associated with both mortality and incidence rates 8 . These epidemiological patterns demonstrate the importance of identifying independent prognostic factors for EOC to improve survival outcomes and implement targeted diagnostic and therapeutic approaches that enhance patients’ outcomes and quality of life. The Surveillance, Epidemiology, and End Results (SEER) database, maintained by the National Cancer Institute, has provided population-based cancer statistics since 1973. Currently covering approximately 48% of the US population, the SEER Program systematically compiles demographic data, primary tumor sites, histological features, staging information, initial treatment modalities, and survival outcomes. The database is annually updated to include data on cancer incidence, mortality, staging, prevalence, and lifetime risk 9 . According to the most recent SEER data, the 5-year relative survival rate for ovarian cancer between 2014 and 2020 was 50.90% 10 . Although most clinical prognostic assessments for EOC continue to rely on the AJCC staging system, its predictive power remains limited. Recently, complementary approaches including molecular signature development, biomarker identification, and mechanistic immune studies have been explored to improve EOC prognostication 11 – 13 . By using the extensive clinical, treatment, and survival data available in the SEER database, we integrated multiple prognostic variables with AJCC Stage II–IV classifications to construct a survival prediction model. This model was validated and assessed for its accuracy in estimating overall survival (OS) among patients with Stage II–IV EOC. A more individualized prognostic assessment can support clinicians in tailoring cancer management strategies and enhancing diagnostic and therapeutic precision for patients with advanced EOC.

Supplementary Material

Supplementary Information. Supplementary Information.

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: pmc-nxml ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

MeSH descriptors

Carcinoma, Ovarian Epithelial Carcinoma, Ovarian Epithelial Carcinoma, Ovarian Epithelial Ovarian Neoplasms Ovarian Neoplasms Ovarian Neoplasms Adult Aged China China Female Humans Middle Aged Neoplasm Staging Nomograms Prognosis Proportional Hazards Models ROC Curve SEER Program

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2026) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

SciLite annotations

organisms 1
mus sp.
chemicals 9
platinum platinum estrogen iron carboplatin paclitaxel paclitaxel carboplatin platinum

Source provenance

europepmc
last seen: 2026-10-04T09:26:46.659050+00:00
pubmed
last seen: 2026-10-08T21:55:40.749282+00:00
scilite
last seen: 2026-09-13T09:58:29.948030+00:00
unpaywall
last seen: 2026-09-14T06:35:54.356137+00:00
License: CC-BY-NC-ND-4.0