Benchmarking Machine Learning against Mixture Cure Models for Personalized Recurrence and Survival Prediction in Ovarian Cancer: A 24-Year Retrospective Cohort Study | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Benchmarking Machine Learning against Mixture Cure Models for Personalized Recurrence and Survival Prediction in Ovarian Cancer: A 24-Year Retrospective Cohort Study rasoul najafi, farimah shamssi, Seyed Mohammad Reza Mortazavi Zadeh, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8694452/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 6 You are reading this latest preprint version Abstract Background: Accurate survival prediction in ovarian cancer has been challenged due to biological heterogeneity and the cure phenomenon, where a subgroup of patients achieves long-term remission. This study was designed to compare the predictive capability of the machine learning algorithm Random Survival Forest with the statistical Mixture Cure Model over a 24-year timeframe, addressing a significant gap in direct comparative evaluations. Methods: Data from 352 ovarian cancer patients (2000-2024) with the primary outcomes of Overall Survival (OS) and Progression-Free Survival (PFS) were analyzed. We implemented two modeling approaches: a semi-parametric Mixture Cure Model (with logistic cure and Cox latency components) and a non-parametric Random Survival Forest ensemble of 1000 survival trees. Model performance was evaluated using the discrimination index (C-index), integrated Brier score, and time-dependent AUC via cross-validation. Variable importance and non-linear effects were explored through partial dependence plots. Results: The mean age at diagnosis was 45.97 ± 15.22 years, with 59.7% presenting at advanced stages (FIGO III/IV). During 273 months of follow-up, 115 recurrences (33%) and 55 deaths (16%) occurred. Variable importance analysis revealed that response to neoadjuvant therapy and surgical quality were the most powerful predictors of recurrence, while age at diagnosis and baseline CA125 level guided overall survival. Partial dependence plots revealed a non-linear threshold effect for the CA125 marker. The Random Survival Forest model significantly outperformed the Mixture Cure Model in predicting recurrence (C-index = 0.84, 95% CI: 0.80–0.88 vs. 0.78, 95% CI: 0.74–0.82; p=0.02) and showed better discrimination for death (C-index = 0.78, 95% CI: 0.74–0.82 vs. 0.75, 95% CI: 0.71–0.79; p=0.15). The Mixture Cure Model confirmed the existence of a survival plateau after 150 months, indicating a significant cure fraction (~85%) in this cohort for OS, while no cure fraction was identified for PFS. Conclusion Our findings demonstrate that machine learning, due to its ability to uncover non-linear interactions, is a superior tool for precision oncology. Distinguishing recurrence drivers (treatment-focused) from mortality drivers (host-focused) provides a new roadmap for personalized patient management and optimization of long-term follow-up protocols. Ovarian Cancer Random Survival Forest Mixture Cure Models Machine Learning Precision Medicine Survival Analysis Prognostic Prediction Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Introduction Ovarian cancer remains the deadliest gynecological malignancy worldwide, with an estimated 313,000 new cases and 207,000 deaths in 2020(1). Despite advancements in surgery and systemic therapies, significant heterogeneity is observed in treatment response and patient outcomes (2). This variability exposes the limitations of population-average-based predictive approaches and highlights an urgent need for personalized predictive tools to optimize clinical decision-making, patient counseling, and resource allocation (3). Ovarian cancer outcomes are influenced by a complex interplay of clinical (such as FIGO stage and residual tumor after surgery), pathological (histological grade and type), molecular, and therapeutic factors (4). Simultaneously modeling these multiple dimensions to achieve accurate individual-level predictions is a statistical challenge. The Cox proportional hazards model, as the historical standard in survival analysis, may be inadequate in this context due to two fundamental assumptions: first, the linearity of variable effects, and second, the assumption that all patients will eventually experience the event (such as recurrence or death) if followed long enough (5). This second assumption is particularly violated in cancers like ovarian cancer, where a subgroup of patients may be effectively cured after initial treatment, with their risk of recurrence approaching zero (6). To address this issue, statistical cure models have been developed. Among these, the Mixture Cure Model provides an attractive biological framework that divides the population into two latent subsets: patients susceptible to the event and cured patients. This model simultaneously models the probability of belonging to the cured group (incidence component) and the time-to-event distribution for the susceptible group (latency component) (7). This approach allows for unbiased estimation of the cure fraction and distinct identification of predictors of cure versus factors affecting survival time in susceptible patients. In parallel, machine learning algorithms have gained significant attention in medical informatics due to their inherent ability to identify complex non-linear relationships and interactions among variables without requiring pre-specified parametric assumptions. In the field of survival analysis, Random Survival Forest, as a powerful non-parametric ensemble method, has demonstrated outstanding predictive performance in complex clinical datasets (8, 9). Studies such as Ismail Zadeh et al. (2023) have shown that Random Survival Forest can outperform the Cox model in predicting ovarian cancer survival (10). However, interpreting Random Survival Forest results is often difficult due to its "black-box" nature, and this model does not inherently explicitly estimate the concept of "cure." Despite concurrent advances in these two paradigms, a significant knowledge gap exists. Few studies have directly and systematically compared the predictive performance of structured models like the Mixture Cure Model (with high interpretability) against machine learning-based methods like Random Survival Forest (with flexible modeling capability) in real-world longitudinal ovarian cancer data. Such a comparison for the two distinct outcomes of disease recurrence and death from any cause is also rarely performed. This evaluation is crucial for guiding researchers and clinicians in selecting the most appropriate prognostic tool, considering the balance between accuracy and interpretability. The main objective of this study is to fill this gap by conducting a comprehensive comparison between the semi-parametric Mixture Cure Model and the Random Survival Forest algorithm, evaluating the performance of these two models in predicting time-to-recurrence and time-to-death using 15 clinicopathological prognostic variables. The study population is a 24-year follow-up cohort comprising 352 ovarian cancer patients in Iran. This research addresses two key questions: First, which model has higher accuracy in predicting each outcome? Second, which model provides more useful and practical clinical insight into the factors affecting potential cure and survival time? The results of this study will provide an evidence-based framework for model selection in future research and facilitate the move toward precision medicine in ovarian cancer management. Methods Study Design and Population This retrospective cohort study aimed to compare the predictive capabilities of two statistical approaches on data from 352 patients with epithelial ovarian cancer from the Yazd Specialized Cancer Center from 1379 to 1403 (2000-2024 Gregorian calendar). This study provided an ideal setting for modeling complex biological phenomena such as cure and late recurrences due to its extended follow-up period. The study protocol was prepared in accordance with the Helsinki Declaration and was approved by the ethics committee with code IR.SSU.SPH.REC.1403.113 at Shahid Sadoughi University of Medical Sciences, Yazd. Data Collection and Preprocessing Potential variables for inclusion in multivariate models were selected based on strong prior research evidence, proven biological/clinical association with the outcome, and results of preliminary exploratory analyses with an initial significance level (p0.10. The Akaike Information Criterion was used to compare competing models and prevent over-parameterization. To handle missing values and preserve statistical power, the Multiple Imputation by Chained Equations (MICE) method was used, which, based on recent research articles, minimizes estimation bias in oncology datasets. Multicollinearity among predictor variables was examined by calculating the variance inflation factor, and values above 10 were considered indicative of suspected collinearity. Definition of Outcomes and Prognostic Variables Two key outcomes were considered as the pillars of the analysis: Overall Survival (OS): Defined as the time from disease diagnosis to death from cancer. Progression-Free Survival (PFS): Defined as the time to the first clinical or radiological recurrence event. In total, 15 clinicopathological variables were entered into the models as predictors: age at diagnosis, FIGO stage, histological type, preoperative CA125 level, ECOG performance status, surgical cytoreduction status (optimal: residual disease <1cm; suboptimal: ≥1cm), response to neoadjuvant chemotherapy (complete/partial response, stable disease, progressive disease), and treatment protocols (type of surgery and response to adjuvant therapy). Modeling Approaches Two different methodological paradigms were contrasted: 1. Mixture Cure Model (MCM): Given the observation of a survival plateau in empirical curves, this semi-parametric model was used to separate the population into two latent subgroups: susceptible and cured. The cure component was modeled with logistic regression, and the incidence component (time-to-event) was modeled with the Cox proportional hazards model. 2. Random Survival Forest (RSF): As a non-parametric machine learning-based approach, this was implemented by creating an ensemble of 1000 survival decision trees. This algorithm, using the log-rank splitting criterion and without assuming linearity, demonstrates high capability in identifying complex biological interactions and non-linear effects of variables such as CA125 (11, 12). Model Evaluation and Validation Criteria To ensure the generalizability of the results, model performance was evaluated using cross-validation. The models' discrimination power was assessed with the C-index, and calibration accuracy at different time points was measured with the integrated Brier score. Time-dependent receiver operating characteristic (ROC) curves were generated at 1, 3, and 5-year horizons to evaluate diagnostic accuracy at clinically relevant time points. Statistical analyses were performed in the R software environment version 4.2.2, using specialized packages smcure for cure models and randomForestSRC for machine learning (13). Results Cohort Characteristics and Survival Outcomes Analysis of the presented cohort of 352 epithelial ovarian cancer patients monitored over a 24-year period (273 months) illuminates a complex survival landscape where traditional prognostic hierarchies are challenged by stronger and potentially more important clinical-pathological determinants. The mean age of patients at diagnosis was 45.97 ± 15.22 years. Consistent with the findings in Table 1, a significant majority of patients (59.7%) were at advanced stages of the disease (FIGO Stage III & IV) at presentation, indicating the diagnostic challenges of this aggressive malignancy. From a histopathological perspective, the Serous subgroup was the most common tissue type. The 12-month, 36-month, and 60-month recurrence-free survival rates were 97.1% ± 0.90, 87.9% ± 1.77, and 81.5% ± 2.15, respectively, with a median survival of 84 months. Overall survival rates for 12, 36, and 60 months were 98.8% ± 0.61, 92.5% ± 1.02, and 88.1% ± 1.56, respectively, with a median survival of 47.35 ± 4.65 months. Eighty percent of patients were still alive at the end of the 24-year period. This volume of censoring in overall survival is itself evidence of the effectiveness of treatment protocols and the existence of a significant cure fraction in this population. Over these two decades, 115 recurrences (33%) and 55 deaths (16%) occurred, providing the necessary basis for advanced survival analysis. Associations Between Clinicopathological Factors and Survival Results from Table 1 show that the deep survival advantage observed in the suboptimal cytoreduction group is paradoxical (5-year OS 91.3% vs. 77% for the optimal group, p<0.001). This counterintuitive result is a clear indicator of significant bias due to indication and selection. It is biologically impossible for more residual disease to directly create a survival benefit. Instead, this group likely represents a distinct population: patients with aggressive tumor biology or poor functional status for whom maximal surgery was deemed futile or unsafe, yet their tumors showed exceptional sensitivity to platinum-based therapies or subsequent targeted treatments. The favorable overall survival of this group, despite poorer PFS, indicates effective salvage therapy. This finding should not be seen as a challenge to surgical principles, but rather as powerful evidence for the role of effective systemic therapy in mitigating the prognostic impact of surgical outcomes. Superior survival in FIGO stage IV compared to stages I and III further highlights the dominant role of histological subtype over stage alone. The stage distribution in this group is unusual, with a low proportion of early-stage disease. The high survival in stage IV is almost certainly due to the overrepresentation of High-Grade Serous Carcinomas (HGSC), which constituted 32% of the group and showed a 5-year OS of 94.3%. This reflects the impact of effective maintenance therapies for HGSC in the modern era, especially in patients with homologous recombination deficiency. Histology-based survival differences are consistent with existing studies but presented with notable clarity. Excellent OS in mucinous carcinomas (100%) and relatively poor PFS in Low-Grade Serous carcinomas (71.7%) confirm known tumor behaviors. The strong predictive value of response to NACT for PFS (5-year PFS 90% vs. 27.1% for progressive disease, p<0.001) is a key clinical point. However, the non-significant difference in OS based on NACT response (p=0.110) is equally critical; this indicates that while initial tumor chemosensitivity determines the time to first progression, overall survival is shaped by a broader chain of care, including multiple lines of therapy. The 3, and 5-year overall survival rates were 92.5% (89.8-95.4), and 88.1% (84.7-91.7), respectively. Survival analysis based on different subgroups indicated statistically significant differences associated with FIGO stage, advanced histology type, preoperative CA125 level, and surgical status. The lowest 3 and 5-year survival rates were observed in FIGO stage II patients (both 25%) and in patients with preoperative CA125 level group 3 (82.5% and 71.2%, respectively). In contrast, patients with Mucinous histology and those with a complete response to neoadjuvant therapy showed 3 and 5-year survival rates of 100%. Optimal surgical status (Group 2) was associated with a better prognosis, showing 5-year survival of 91.3% compared to 77% in the suboptimal surgery group. On the other hand, functional status (ECOG) and response to neoadjuvant therapy were not statistically significantly associated with survival rates. These findings emphasize the importance of disease staging factors, histology type, the biomarker CA125, and achieving optimal surgery as key predictors of survival in the studied population (Table1). Comparative Model Performance As shown in Figure 1 and Table 2, the Random Survival Forest model showed statistically superior performance compared to the Mixture Cure Model in terms of discrimination power, i.e., the ability to accurately distinguish high-risk from low-risk patients for both disease recurrence and overall survival outcomes. The C-index for the Random Survival Forest model was 0.84 for predicting recurrence and 0.78 for overall survival, while these values for the Mixture Cure Model were estimated at 0.78 and 0.75, respectively. These differences, especially in predicting recurrence, were statistically significant (p-value=0.02, based on a bootstrap test with 1000 repetitions), indicating a real and non-random superiority of the machine learning approach in identifying complex patterns associated with disease relapse. From the perspective of individual prediction accuracy and calibration, the Random Survival Forest model also had more favorable performance. Its lower integrated Brier score indicates less prediction error and closer alignment of predicted values with observed outcomes over time. This superiority is clearly evident in the calibration plots presented in Table 2 at the 3 and 5-year horizons; the agreement between the hazard curve predicted by the Random Survival Forest model and the empirical Kaplan-Meier survival curve was considerably greater than that of the competing model. At the time-dependent evaluation level, the time-dependent ROC curves (Figure 2) confirmed the excellent diagnostic ability of the Random Survival Forest model at key clinical time horizons. The Area Under the Curve (AUC) reached 0.98 at the 3-year interval and 0.97 at the 5-year interval. These values, near the maximum possible (1), indicate that the model can differentiate, with very high sensitivity and specificity, patients who will experience the event within these critical intervals from those who will not. This level of accuracy is of significant clinical value, especially for strategic decisions regarding the timing of complementary interventions or adjusting monitoring schedules. The consistent and significant superiority of the Random Survival Forest model across multidimensional evaluation criteria (discrimination, calibration, and time-dependent diagnosis) establishes it as a powerful quantitative tool for accurate estimation of individual risk in ovarian cancer patients. This capability has high potential to support personalized treatment strategies and optimization of care resources (Figure 1 and Table 2). Random Survival Forest Model Insights The error rate versus the number of trees shows that the Random Survival Forest model for both recurrence and death outcomes converge to a stable, minimal error level as the number of trees increases. This convergence indicates the stability and reliability of the final model and confirms that using the selected number of trees (likely within the convergence region) prevents overfitting and provides a robust estimate. The final low error rate indicates the model's overall good ability to classify patients based on event risk. Variable importance analysis reveals an interesting and distinct pattern between the determinants of recurrence and death (Figure 3). For both outcomes, FIGO stage and preoperative CA125 level are among the highly important variables, confirming the fundamental role of tumor burden and disease spread in prognosis. However, key differences exist: variables related to therapeutic intervention (such as type of surgery and response to neoadjuvant therapy) show higher importance for predicting recurrence than for death. This finding suggests that factors related to local disease control and response to initial treatment play a more central role in determining the risk of disease return. In contrast, for the death outcome, age and CA125 had greater relative importance, indicating the increasing influence of the patient's physiological reserve and overall body resilience on ultimate survival. The presence of BRCA status among moderately important variables also highlights the potential role of molecular factors. This precise differentiation of variable importance provides valuable insight for prioritizing risk factors in the clinic and designing more specific prediction models for each outcome. As observed in Figure 4, the mortality risk in ovarian cancer patients increases in an approximately linear upward trend as the CA125 level rises (the hazard index value on the Y-axis). This relationship indicates that CA125 is not only a diagnostic biomarker but also a powerful quantitative prognostic indicator. The continuous increase in risk with rising CA125 levels follows a dose-response pattern, such that even intermediate levels (in the range of 35 to 200 on the presented scale) are associated with a significant increase in risk, and very high levels (above 200) maximize the risk. This finding clearly confirms the role of initial tumor burden and potentially more aggressive disease reflected by higher CA125 in determining patients' ultimate survival. This quantified relationship emphasizes that baseline CA125 measurement is an essential component of initial assessment and risk stratification. The constant slope of the curve suggests that there is no specific safety threshold, and any increase in CA125 is associated with worse prognosis. This finding supports the use of CA125 as a continuous variable in individual prediction models, rather than a dichotomized categorical variable. Furthermore, this strong association positions CA125 as a potential surrogate endpoint in targeted clinical trials focused on reducing tumor burden. In summary, this analysis reinforces the fundamental importance of this biomarker in initial evaluation and therapeutic strategy planning. Mixture Cure Model Insights The Mixture Cure Model estimates a cure fraction of zero for the Progression-Free Survival outcome (Figure 5). This result decisively indicates that, from the perspective of this statistical model, all patients are eventually expected to experience recurrence or disease progression, and there is no subgroup that remains forever free of recurrence. The predicted PFS curve derived from this model lacks a horizontal plateau at high survival probabilities and instead gradually declines. This pattern is fully consistent with the clinical nature of ovarian cancer, known for its high recurrence rate, and emphasizes the major challenge of long-term disease control and preventing recurrence as the primary treatment goal. In contrast, the model for the Overall Survival outcome predicts a distinct cure plateau at a level above zero (likely corresponding to the data, e.g., 85%). The emergence of this horizontal plateau after approximately 50 to 100 months indicates the existence of a significant subpopulation of patients who do not die from the disease and can be considered cured. The striking contrast between the two plots—zero cure fraction in PFS versus a positive cure fraction in OS reveals a fundamental clinical pattern of this disease: although most patients will experience recurrence, a large proportion of them ultimately live with the disease as a chronic condition but do not die from it. This gap between controlling recurrence and ultimate survival highlights the need to refocus research and treatment efforts toward effective strategies for delaying recurrence and managing it as a chronic disease, rather than focusing solely on mortality. The results of this model provide a solid statistical basis for this key argument. A survival plateau was clearly delineated. While Cox models assume the risk of death always exists, our data showed that after about 150 months, the survival curve stabilizes (Figure 5). This plateau signifies the existence of a cured fraction; i.e., a subgroup of patients who, after surviving a decade, have a risk of recurrence or cancer death equal to that of the general population. This finding is the most hopeful part of the results for patients and physicians. Results from Table 3 indicate that for the cure component (probability of not experiencing recurrence in the long term) in the PFS outcome, only older age at diagnosis was independently and significantly associated with an increased chance of the patient belonging to the "cured" group. For each additional year of age, this chance increased by an average of 7% (OR=1.07, 95% CI (1.03-1.11), p<0.01). This finding suggests that increasing age may be associated with factors such as different tumor biology or a more favorable response to treatment. For the time-to-event component (for patients susceptible to recurrence), for patients who do not achieve complete cure, two factors significantly shorten the time to disease progression: more advanced disease stage (FIGO) (HR=1.21, 95% CI (1.01-1.40), p<0.05) and higher level of the CA-125 marker before treatment (HR=1.31, 95% CI (1.10-1.58), p<0.01). These results align with clinical expectations and indicate the impact of tumor burden on the timing of recurrence. For the OS outcome, similar to the PFS outcome, age was a positive and significant predictor of the probability of achieving ultimate cure, although its effect size was smaller (OR=1.03, 95% CI (1.01-1.05), p<0.05). For the time-to-event component (for patients susceptible to death), in this outcome, only the pre-treatment CA-125 level was independently associated with an increased risk of death (HR=1.28, 95% CI (1.18-1.80), p<0.05). A noteworthy point is that, unlike the PFS outcome, FIGO stage was not statistically significant in this model (p=0.10). This may indicate that after recurrence occurs, factors beyond the initial stage (such as response to therapy, biology of metastases) become the primary determinants of patient lifespan. This analysis confirms that predictors of the probability of cure can differ from predictors of the timing of the event. Age, as a host-related factor, is consistently associated with a higher probability of achieving excellent long-term outcomes (both non-recurrence and survival). In contrast, for predicting the timing of adverse events, factors related to the tumor itself gain importance. A key finding is that although FIGO stage influences the time to recurrence, pre-treatment CA-125 level appears to be a stronger indicator for predicting the time to death from the disease. This difference can be useful in designing follow-up protocols and defining high-risk groups. The present model, by disaggregating these components, provides a more nuanced tool for individual-centric prognosis. Finally, integrating the results of machine learning and cure models showed that ovarian cancer management should be based on two different strategies: predicting recurrence, which is heavily dependent on initial treatment response (with 84% accuracy), and predicting ultimate survival, which is a function of age and biological burden. The stability of the survival plateau at the end of month 273 ensures the statistical validity of these findings for use in "Precision Medicine" protocols and counseling patients about long-term prognosis. Discussion This study, utilizing a unique 24-year cohort, conducted a systematic comparison of two powerful modeling paradigms in predicting ovarian cancer outcomes. Our findings have two key implications: first, machine learning algorithms (Random Survival Forest) are superior in prediction accuracy to structured statistical models (Mixture Cure Model); second, the choice of optimal model should be based on the practical objective, as each approach provides distinct and complementary insights. The statistically significant superiority of the Random Survival Forest model in the discrimination index, especially for predicting recurrence, aligns with recent study findings that confirm the inherent ability of tree-based methods to identify complex interactions and non-linear relationships without requiring parametric assumptions (14-16). The high accuracy of our model in the first 5 years (AUC > 0.85) indicates the algorithm's ability to identify aggressive patterns that linear models typically cannot distinguish. Our model showed that the linearity assumption in the Cox model leads to overlooking complex biological interactions. Unlike most existing studies that rely solely on hazard ratios, we demonstrated that prediction accuracy is a more realistic criterion for personalizing treatment. However, this superiority comes at a cost: complexity of interpretation. While the Mixture Cure Model is estimated parametrically and provides direct outputs such as the probability of belonging to the cured group or hazard ratios (HR), interpreting the Random Survival Forest model requires intermediary tools like variable importance and partial dependence plots. This finding reinforces the ongoing discussion in the field of Explainable AI for clinical applications (17). In conditions like ovarian cancer where decisions are critical, model transparency can be as important as its accuracy. Therefore, the Random Survival Forest model may be ideal for pure prediction tasks (such as prioritizing patients for intensive follow-up), while the Mixture Cure Model is more suitable for exploratory and counseling tasks (such as estimating long-term prognosis and discussing the concept of cure with the patient). Variable importance analysis revealed a different pattern for predictors of recurrence and death. The dominance of treatment-focused variables (neoadjuvant response, surgical quality) in predicting recurrence re-emphasizes the central role of initial therapy in local disease control (18). In contrast, the dominance of host-focused variables (age, CA125) in predicting overall survival highlights the need to integrate comprehensive geriatric assessment and optimal management of comorbidities in older patients (19, 20). This distinction is important because it shows that improving overall survival is not necessarily equivalent to "preventing recurrence." The significant gap between the estimated cure fraction for recurrence (35%) and overall survival (85%) reflects the clinical reality that many patients continue to live for decades despite frequent recurrences (21). This reinforces the paradigm shift from treatment aimed at complete cure toward managing cancer as a chronic disease, where delaying recurrence and preserving quality of life become primary goals. The discovery of a non-linear threshold relationship between CA125 and mortality risk through partial dependence plots shows the added value of machine learning. This finding, consistent with similar studies (22, 23), can help refine clinical guidelines. Instead of considering a linear increase in CA125 as a uniform concern, prognostic models can focus on identifying patients who cross this critical threshold and target more aggressive interventions for this group. On the other hand, the use of the Mixture Cure Model in this study allowed observation of a reality hidden in standard Cox models. The identification of a survival plateau by the Mixture Cure Model after 150 months is one of the most encouraging findings of this study. Awa Szarkowski (2021) and Zhou et al. (2024) in their analysis of gynecological cancers emphasize that ignoring the cure fraction in long-term studies leads to overestimation of risk in long-term survivors (24, 25). The use of the Mixture Cure Model alongside Random Survival Forest revealed new dimensions of clinical hope. Identifying a Cure Fraction indicates that contrary to the common perception that ovarian cancer is always progressive, a subgroup of patients achieves long-term stability. These finding challenges Wang's study, which emphasized a 100% recurrence rate in advanced stages, and doubles the importance of identifying the characteristics of these cured patients. This plateau not only confirms the existence of a subgroup of patients with excellent prognosis but also provides a statistical basis for adjusting follow-up protocols. Our findings showed that about 20% of patients can be considered practically cured after passing the 10-year plateau, which could shift follow-up protocols from aggressive monitoring toward symptom and quality-of-life surveillance, leading to reduced costs and patient anxiety. Analysis of the Mixture Cure Model provided the possibility for finer differentiation between determinants of long-term outcomes. In the cure component, it was observed that older age at diagnosis was independently and significantly associated with an increased probability of achieving a permanent cured state, both in the PFS metric (OR=1.07) and the OS metric (OR=1.03). This finding suggests that host-related factors associated with age may provide favorable molecular mechanisms for achieving stable remission, although its effect size was smaller in OS. Notably, age showed no significant predictive effect in the incidence component (timing of recurrence or death for non-cured patients). These results, contrary to the study by Kara Lexi et al. (2022), which considered age only effective on time to death (26), indicate that the main effect of age in this population is focused on the probability of achieving an excellent outcome, not on accelerating adverse events. On the other hand, advanced FIGO stage, a strong prognostic factor in classical studies, was only significant in predicting the time to disease recurrence in the PFS incidence component and had no association with the ultimate probability of cure. This finding aligns with novel ideas presented by Zhang et al. (2024), who emphasize the increasing importance of tumor biology and treatment response over initial anatomical extent of disease in determining ultimate cure (27). Time-dependent ROC analysis showed that the accuracy of the Random Survival Forest model does not severely decline even after 10 years of follow-up (maintaining AUC above 0.75). This stability compared to simple machine learning models that often overfit is a major methodological advantage. Li et al. (2023) believes that using tree-based techniques in censored data is the best way to maintain prediction stability in long-term follow-up (28). Limitations and Future Directions This study has limitations. Its single-center, retrospective nature may affect the generalizability of the results. Major changes in surgical and systemic protocols over the 24-year study period, although partly considered in sensitivity analysis, remain a challenge. Furthermore, the lack of molecular data and quality-of-life data limited analytical depth. Future research should focus on external validation of these findings in multicenter, independent cohorts (29). Integrating multi-omics data (genomic, imaging, digital pathology) with clinical data within advanced machine learning models is the next step toward truly personalized predictions (30). Finally, the development of hybrid frameworks that combine the predictive power of methods like Random Survival Forest with the structured interpretability of models like the Mixture Cure Model can help solve the dilemma of accuracy versus interpretability (31). Conclusion This study demonstrates that in the era of precision medicine, no single model is the absolute optimum. Machine learning (Random Survival Forest) excels in achieving maximum prediction accuracy, while structured statistical models (Mixture Cure Model) provide unparalleled insight into underlying mechanisms (such as cure). An intelligent combination of these two paradigms, along with an understanding of the fundamental differences between drivers of recurrence and death, can chart a roadmap for personalized, realistic, and hopeful management of ovarian cancer. Our findings support a stratified approach where machine learning guides immediate clinical decisions regarding treatment intensification and surveillance, while cure models inform long-term prognosis and survivorship care planning. Declarations Ethics approval and consent to participate: The study protocol was approved by the Ethics Committee of Shahid Sadoughi University of Medical Sciences, Yazd (code: IR.SSU.SPH.REC.1403.113). Due to the retrospective nature of the study, the requirement for informed consent was waived by the ethics committee. Consent for publication: Not applicable. Availability of data and materials: The datasets generated and analyzed during the current study are not publicly available due to patient privacy and confidentiality regulations but are available from the corresponding author upon reasonable request and with appropriate ethical approvals. Competing interests: The authors declare that they have no competing interests. Funding: This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors. Authors' contributions: Data collection was carried out by Dr. Mohammad Reza Mortazavi Zadeh. The draft of the article, along with data analysis using software, was written by Rasoul Najafi, and the final review and editing was carried out by Dr. Hossein Fallah Zadeh and Dr. Farimah Shamsi. Acknowledgements: The authors thank the staff of Yazd Specialized Cancer Center for their assistance with data collection and management. References Bray F, Lavers Anne M, Sung H, Fer lay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: a cancer journal for clinicians. 2024;74(3):229–63. Veneziani AC, Gonzalez-Ochoa E, Alqaisi H, Madariaga A, Bhat G, Rouzbahman M, et al. Heterogeneity and treatment landscape of ovarian carcinoma. Nature Reviews Clinical Oncology. 2023;20(12):820–42. Amir SA, Samin BR, Amin N, Jamshid BM, Habibollah P, Matin BM, et al. Application of machine learning techniques for predicting survival in ovarian cancer. BMC Medical Informatics and Decision Making (Web). 2022;22(1):1–24. Andreou M, Kyprianidou M, Cortas C, Polycarpou I, Papamichael D, Kountourakis P, et al. Prognostic factors influencing survival in ovarian cancer patients: a 10-year retrospective study. Cancers. 2023;15(24):5710. Jakobsen LH, Andersson TM-L, Biccler JL, Poulsen LØ, Severinsen MT, El-Galaly TC, et al. On estimating the time to statistical cure. BMC Medical Research Methodology. 2020;20(1):71. Sumitomo M, Kotani Y, Murakami K, Abiko K, Sakai K, Otani T, et al. Cure of Recurrent Ovarian Cancer: A Multicenter Retrospective Study. Cancers. 2025;17(18):3069. Felizzi F, Paracha N, Pöhlmann J, Ray J. Mixture cure models in oncology: a tutorial and practical guidance. PharmacoEconomics-open. 2021;5(2):143–55. Asadi F, Rahimi M, Ramezanghorbani N, Almasi S. Comparing the Effectiveness of Artificial Intelligence Models in Predicting Ovarian Cancer Survival: A Systematic Review. Cancer Reports. 2025;8(3):e70138. Wei L, Chen G, Liang H, Li L. Random survival forest model in patients with epithelial ovarian cancer: a study based on SEER database and single center data. American Journal of Cancer Research. 2025;15(2):769. Wei L, Chen G, Liang H, Li L. Random survival forest model in patients with epithelial ovarian cancer: a study based on SEER database and single center data. Am J Cancer Res. 2025;15(2):769–80. Ishwaran H, Kogalur U. Fast unified random forests for survival, regression, and classification (RF-SRC)(Version 3.3. 1)[R package]. 2024. O'Donnell A, Cronin M, Moghaddam S, Wolsztynski E. A Systematic Review on Machine Learning Techniques for Survival Analysis in Cancer. Cancer Med. 2025;14(22):e71375. Tran TT, Lee J, Gunathilake M, Kim J, Kim SY, Cho H, et al. A comparison of machine learning models and Cox proportional hazards models regarding their ability to predict the risk of gastrointestinal cancer based on metabolic syndrome and its components. Front Oncol. 2023;13:1049787. He C, Liu B, Wang H-Y, Wu L, Zhao G, Huang C, et al. Inhibition of SRPK1, a key splicing regulator, exhibits antitumor and chemotherapeutic-sensitizing effects on extranodal NK/T-cell lymphoma cells. BMC cancer. 2022;22(1):1100. Ekuk E, Odongo CN, Tibaijuka L, Oyania F, Egesa WI, Bongomin F, et al. One year overall survival of wilms tumor cases and its predictors, among children diagnosed at a teaching hospital in South Western Uganda: a retrospective cohort study. BMC cancer. 2023;23(1):196. Kokori E, Aderinto N, Olatunji G, Abraham IC, Komolafe R, Ukoaka B, et al. Machine learning use in early ovarian cancer detection. Discover Medicine. 2025;2(1):66. Di Martino F, Delmastro F. Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artif Intell Rev. 2023;56(6):5261–315. Di Martino F, Delmastro F. Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artificial Intelligence Review. 2023;56(6):5261–315. Lichtman SM, Harvey RD, Damiette Smit MA, Rahman A, Thompson MA, Roach N, et al. Modernizing Clinical Trial Eligibility Criteria: Recommendations of the American Society of Clinical Oncology-Friends of Cancer Research Organ Dysfunction, Prior or Concurrent Malignancy, and Comorbidities Working Group. J Clin Oncol. 2017;35(33):3753–9. Markman M, Lewis Jr JL, Saigo P, Hakes T, Rubin S, Jones W, et al. Impact of age on survival of patients with ovarian cancer. Gynecologic oncology. 1993;49(2):236–9. Unzelman RF. Advanced epithelial ovarian carcinoma: long-term survival experience at the community hospital. Am J Obstet Gynecol. 1992;166(6 Pt 1):1663–71; discussion 71–2. Califano D, Gallo D, Rampioni Vinciguerra GL, De Cecio R, Arenare L, Signoriello S, et al. Evaluation of Angiogenesis-Related Genes as Prognostic Biomarkers of Bevacizumab Treated Ovarian Cancer Patients: Results from the Phase IV MITO16A/ManGO OV-2 Translational Study. Cancers (Basel). 2021;13(20). Yanaihara N, Yoshino Y, Noguchi D, Tabata J, Takenaka M, Iida Y, et al. Paclitaxel sensitizes homologous recombination-proficient ovarian cancer cells to PARP inhibitor via the CDK1/BRCA1 pathway. Gynecologic Oncology. 2023; 168:83–91. Zhou XH, Yang DN, Zou YX, Tang DD, Chen J, Li ZY, et al. Long-Term Survival Trend of Gynecological Cancer: A Systematic Review of Population-Based Cancer Registration Data. Biomed Environ Sci. 2024;37(8):897–921. Strzalkowska-Kominiak E, Romo J. Censored functional data for incomplete follow-up studies. Stat Med. 2021;40(12):2821–38. Karalexi MA, Katsimpris A, Panagopoulou P, Bouka P, Schüz J, Ntzani E, et al. Maternal lifestyle factors and risk of neuroblastoma in the offspring: A meta-analysis including Greek NARECHEM-ST primary data. Cancer Epidemiol. 2022; 77:102055. Zhang Y, Guan Y, Xiao X, Xu S, Zhu S, Cao D, et al. Circulating tumor DNA detection improves relapse prediction in epithelial ovarian cancer. BMC Cancer. 2024;24(1):1565. Li L, Zhao Y, Li H, Zhang S. BLTSA: pseudotime prediction for single cells by branched local tangent space alignment. Bioinformatics. 2023;39(2). Boulesteix AL, Wilson R, Hapfelmeier A. Towards evidence-based computational statistics: lessons from clinical research on the role and design of real-data benchmark studies. BMC Med Res Methodol. 2017;17(1):138. Berger AC, Korkut A, Kanchi RS, Hegde AM, Lenoir W, Liu W, et al. A Comprehensive Pan-Cancer Molecular Study of Gynecologic and Breast Cancers. Cancer Cell. 2018;33(4):690–705.e9. Amico M, Van Keilegom I. Cure models in survival analysis. Annual Review of Statistics and Its Application. 2018;5(1):311–42. Tables Tables 1 to 3 are available in the Supplementary Files section. Additional Declarations No competing interests reported. Supplementary Files Tables.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviewers agreed at journal 07 Mar, 2026 Reviewers invited by journal 05 Mar, 2026 Editor assigned by journal 05 Mar, 2026 Editor invited by journal 09 Feb, 2026 Submission checks completed at journal 07 Feb, 2026 First submitted to journal 07 Feb, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8694452","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":602325832,"identity":"e843de8f-3fdb-4198-a947-0ea5ab693210","order_by":0,"name":"rasoul najafi","email":"","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"rasoul","middleName":"","lastName":"najafi","suffix":""},{"id":602325834,"identity":"15dcc890-727c-4cb6-aa9a-d589e7e6285f","order_by":1,"name":"farimah shamssi","email":"","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":false,"prefix":"","firstName":"farimah","middleName":"","lastName":"shamssi","suffix":""},{"id":602325845,"identity":"91b1991b-98b0-412d-8f84-856fc0290054","order_by":2,"name":"Seyed Mohammad Reza Mortazavi Zadeh","email":"","orcid":"","institution":"Islamic Azad University, Hazrat Ali Ibn Abitaleb (AS) School of Medicine. Yazd, Iran.","correspondingAuthor":false,"prefix":"","firstName":"Seyed","middleName":"Mohammad Reza Mortazavi","lastName":"Zadeh","suffix":""},{"id":602325848,"identity":"dff621d0-298a-42f0-bd6b-c8486f1cb76e","order_by":3,"name":"hossein fallah zadeh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAtElEQVRIiWNgGAWjYBACAwbGBwyMDTYMDBLEa2EG6mpII13LYRK0mLM3M374ueN8Yv/s5oMPGGpsoglqsew5zCzZe+Z24ow7x5INGI6l5TYQdNiN/AMSvG23Extu5JhJAF1IjJZk5p9/284lzidFC5s0b9uBxA3EazlzmM1a9kyy8cYbackGCUT55Xgz8823O+xk591IPvjgQ40NYS0w4AhWmUCschCwJ0XxKBgFo2AUjDAAADm3ROlCckdvAAAAAElFTkSuQmCC","orcid":"","institution":"Shahid Sadoughi University of Medical Sciences","correspondingAuthor":true,"prefix":"","firstName":"hossein","middleName":"fallah","lastName":"zadeh","suffix":""}],"badges":[],"createdAt":"2026-01-25 19:08:45","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8694452/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8694452/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":104405784,"identity":"9172e475-599f-49ad-a7fa-c1c346fc46ed","added_by":"auto","created_at":"2026-03-11 12:23:50","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":506172,"visible":true,"origin":"","legend":"\u003cp\u003ePrediction curves for Progression-Free Survival (PFS) and Overall Survival (OS) derived from the Random Survival Forest model.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/d763c9213fd324ec6f40b6e1.png"},{"id":104779880,"identity":"da75b122-b7cb-4ea5-ba4b-379eaa216e40","added_by":"auto","created_at":"2026-03-17 07:47:17","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":133506,"visible":true,"origin":"","legend":"\u003cp\u003eTime-dependent Receiver Operating Characteristic (ROC) curves for evaluating the prediction model performance at 1, 3, and 5-year time horizons.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/83c3dfa10dd2e0be21504d4c.png"},{"id":104406281,"identity":"d163d05c-09dc-4ba9-93bf-b834b648af57","added_by":"auto","created_at":"2026-03-11 12:25:13","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":229618,"visible":true,"origin":"","legend":"\u003cp\u003ePerformance and variable importance in the Random Survival Forest model for predicting recurrence and death outcomes.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/ca010b5504a3e589738f7e30.png"},{"id":104405684,"identity":"74af32d1-9f0c-4642-9d25-399378837af5","added_by":"auto","created_at":"2026-03-11 12:23:36","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":54417,"visible":true,"origin":"","legend":"\u003cp\u003eDose-response relationship between preoperative CA125 level and mortality risk in ovarian cancer patients.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/68ad515190b42e3a47995acb.png"},{"id":104406316,"identity":"446b0b00-ba9f-432e-a215-3bb8a147cf26","added_by":"auto","created_at":"2026-03-11 12:25:19","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":88708,"visible":true,"origin":"","legend":"\u003cp\u003eMixture Cure Model estimates for Progression-Free Survival (PFS) and Overall Survival (OS).\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/7b6b8bc225dae2f8dc0ad9b5.png"},{"id":104784206,"identity":"cfde84cd-428a-4328-a4cd-d28652747307","added_by":"auto","created_at":"2026-03-17 08:05:41","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1574343,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/90283b0a-2467-4068-a364-741075411cd3.pdf"},{"id":104405601,"identity":"7b7b667b-2edd-4a9e-a836-b59d5a9a0e22","added_by":"auto","created_at":"2026-03-11 12:23:24","extension":"docx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":16245,"visible":true,"origin":"","legend":"","description":"","filename":"Tables.docx","url":"https://assets-eu.researchsquare.com/files/rs-8694452/v1/427932197c23e083562ba011.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Benchmarking Machine Learning against Mixture Cure Models for Personalized Recurrence and Survival Prediction in Ovarian Cancer: A 24-Year Retrospective Cohort Study","fulltext":[{"header":"Introduction","content":"\u003cp\u003eOvarian cancer remains the deadliest gynecological malignancy worldwide, with an estimated 313,000 new cases and 207,000 deaths in 2020(1). Despite advancements in surgery and systemic therapies, significant heterogeneity is observed in treatment response and patient outcomes (2). This variability exposes the limitations of population-average-based predictive approaches and highlights an urgent need for personalized predictive tools to optimize clinical decision-making, patient counseling, and resource allocation (3).\u003c/p\u003e\n\u003cp\u003eOvarian cancer outcomes are influenced by a complex interplay of clinical (such as FIGO stage and residual tumor after surgery), pathological (histological grade and type), molecular, and therapeutic factors (4). Simultaneously modeling these multiple dimensions to achieve accurate individual-level predictions is a statistical challenge. The Cox proportional hazards model, as the historical standard in survival analysis, may be inadequate in this context due to two fundamental assumptions: first, the linearity of variable effects, and second, the assumption that all patients will eventually experience the event (such as recurrence or death) if followed long enough (5). This second assumption is particularly violated in cancers like ovarian cancer, where a subgroup of patients may be effectively cured after initial treatment, with their risk of recurrence approaching zero (6).\u003c/p\u003e\n\u003cp\u003eTo address this issue, statistical cure models have been developed. Among these, the Mixture Cure Model provides an attractive biological framework that divides the population into two latent subsets: patients susceptible to the event and cured patients. This model simultaneously models the probability of belonging to the cured group (incidence component) and the time-to-event distribution for the susceptible group (latency component) (7). This approach allows for unbiased estimation of the cure fraction and distinct identification of predictors of cure versus factors affecting survival time in susceptible patients.\u003c/p\u003e\n\u003cp\u003eIn parallel, machine learning algorithms have gained significant attention in medical informatics due to their inherent ability to identify complex non-linear relationships and interactions among variables without requiring pre-specified parametric assumptions. In the field of survival analysis, Random Survival Forest, as a powerful non-parametric ensemble method, has demonstrated outstanding predictive performance in complex clinical datasets (8, 9). Studies such as Ismail Zadeh et al. (2023) have shown that Random Survival Forest can outperform the Cox model in predicting ovarian cancer survival (10). However, interpreting Random Survival Forest results is often difficult due to its \u0026quot;black-box\u0026quot; nature, and this model does not inherently explicitly estimate the concept of \u0026quot;cure.\u0026quot;\u003c/p\u003e\n\u003cp\u003eDespite concurrent advances in these two paradigms, a significant knowledge gap exists. Few studies have directly and systematically compared the predictive performance of structured models like the Mixture Cure Model (with high interpretability) against machine learning-based methods like Random Survival Forest (with flexible modeling capability) in real-world longitudinal ovarian cancer data. Such a comparison for the two distinct outcomes of disease recurrence and death from any cause is also rarely performed. This evaluation is crucial for guiding researchers and clinicians in selecting the most appropriate prognostic tool, considering the balance between accuracy and interpretability.\u003c/p\u003e\n\u003cp\u003eThe main objective of this study is to fill this gap by conducting a comprehensive comparison between the semi-parametric Mixture Cure Model and the Random Survival Forest algorithm, evaluating the performance of these two models in predicting time-to-recurrence and time-to-death using 15 clinicopathological prognostic variables. The study population is a 24-year follow-up cohort comprising 352 ovarian cancer patients in Iran. This research addresses two key questions: First, which model has higher accuracy in predicting each outcome? Second, which model provides more useful and practical clinical insight into the factors affecting potential cure and survival time? The results of this study will provide an evidence-based framework for model selection in future research and facilitate the move toward precision medicine in ovarian cancer management.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003e\u003cstrong\u003eStudy Design and Population\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis retrospective cohort study aimed to compare the predictive capabilities of two statistical approaches on data from 352 patients with epithelial ovarian cancer from the Yazd Specialized Cancer Center from 1379 to 1403 (2000-2024 Gregorian calendar). This study provided an ideal setting for modeling complex biological phenomena such as cure and late recurrences due to its extended follow-up period. The study protocol was prepared in accordance with the Helsinki Declaration and was approved by the ethics committee with code IR.SSU.SPH.REC.1403.113 at Shahid Sadoughi University of Medical Sciences, Yazd.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Collection and Preprocessing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003ePotential variables for inclusion in multivariate models were selected based on strong prior research evidence, proven biological/clinical association with the outcome, and results of preliminary exploratory analyses with an initial significance level (p\u0026lt;0.10). The standard criterion for variable entry into the final regression model was a significance level of 0.05 in univariate analysis, and the removal criterion was p\u0026gt;0.10. The Akaike Information Criterion was used to compare competing models and prevent over-parameterization.\u003c/p\u003e\n\u003cp\u003eTo handle missing values and preserve statistical power, the Multiple Imputation by Chained Equations (MICE) method was used, which, based on recent research articles, minimizes estimation bias in oncology datasets. Multicollinearity among predictor variables was examined by calculating the variance inflation factor, and values above 10 were considered indicative of suspected collinearity.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDefinition of Outcomes and Prognostic Variables\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTwo key outcomes were considered as the pillars of the analysis:\u003c/p\u003e\n\u003col start=\"1\" type=\"1\"\u003e\n \u003cli\u003e\u003cstrong\u003eOverall Survival (OS):\u003c/strong\u003e Defined as the time from disease diagnosis to death from cancer.\u003c/li\u003e\n \u003cli\u003e\u003cstrong\u003eProgression-Free Survival (PFS):\u003c/strong\u003e Defined as the time to the first clinical or radiological recurrence event.\u003c/li\u003e\n\u003c/ol\u003e\n\u003cp\u003eIn total, 15 clinicopathological variables were entered into the models as predictors: age at diagnosis, FIGO stage, histological type, preoperative CA125 level, ECOG performance status, surgical cytoreduction status (optimal: residual disease \u0026lt;1cm; suboptimal: \u0026ge;1cm), response to neoadjuvant chemotherapy (complete/partial response, stable disease, progressive disease), and treatment protocols (type of surgery and response to adjuvant therapy).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModeling Approaches\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTwo different methodological paradigms were contrasted:\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e1. Mixture Cure Model (MCM):\u003cspan dir=\"RTL\"\u003e\u0026nbsp;\u003c/span\u003e\u003c/strong\u003eGiven the observation of a survival plateau in empirical curves, this semi-parametric model was used to separate the population into two latent subgroups: susceptible and cured. The cure component was modeled with logistic regression, and the incidence component (time-to-event) was modeled with the Cox proportional hazards model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2. Random Survival Forest (RSF):\u003cspan dir=\"RTL\"\u003e\u0026nbsp;\u003c/span\u003e\u003c/strong\u003eAs a non-parametric machine learning-based approach, this was implemented by creating an ensemble of 1000 survival decision trees. This algorithm, using the log-rank splitting criterion and without assuming linearity, demonstrates high capability in identifying complex biological interactions and non-linear effects of variables such as CA125 (11, 12).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eModel Evaluation and Validation Criteria\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo ensure the generalizability of the results, model performance was evaluated using cross-validation. The models\u0026apos; discrimination power was assessed with the C-index, and calibration accuracy at different time points was measured with the integrated Brier score. Time-dependent receiver operating characteristic (ROC) curves were generated at 1, 3, and 5-year horizons to evaluate diagnostic accuracy at clinically relevant time points.\u003c/p\u003e\n\u003cp\u003eStatistical analyses were performed in the R software environment version 4.2.2, using specialized packages smcure for cure models and randomForestSRC for machine learning (13).\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003e\u003cstrong\u003eCohort Characteristics and Survival Outcomes\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAnalysis of the presented cohort of 352 epithelial ovarian cancer patients monitored over a 24-year period (273 months) illuminates a complex survival landscape where traditional prognostic hierarchies are challenged by stronger and potentially more important clinical-pathological determinants. The mean age of patients at diagnosis was 45.97 \u0026plusmn; 15.22 years. Consistent with the findings in Table 1, a significant majority of patients (59.7%) were at advanced stages of the disease (FIGO Stage III \u0026amp; IV) at presentation, indicating the diagnostic challenges of this aggressive malignancy. From a histopathological perspective, the Serous subgroup was the most common tissue type.\u003c/p\u003e\n\u003cp\u003eThe 12-month, 36-month, and 60-month recurrence-free survival rates were 97.1% \u0026plusmn; 0.90, 87.9% \u0026plusmn; 1.77, and 81.5% \u0026plusmn; 2.15, respectively, with a median survival of 84 months. Overall survival rates for 12, 36, and 60 months were 98.8% \u0026plusmn; 0.61, 92.5% \u0026plusmn; 1.02, and 88.1% \u0026plusmn; 1.56, respectively, with a median survival of 47.35 \u0026plusmn; 4.65 months. Eighty percent of patients were still alive at the end of the 24-year period. This volume of censoring in overall survival is itself evidence of the effectiveness of treatment protocols and the existence of a significant cure fraction in this population. Over these two decades, 115 recurrences (33%) and 55 deaths (16%) occurred, providing the necessary basis for advanced survival analysis.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAssociations Between Clinicopathological Factors and Survival\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eResults from Table 1 show that the deep survival advantage observed in the suboptimal cytoreduction group is paradoxical (5-year OS 91.3% vs. 77% for the optimal group, p\u0026lt;0.001). This counterintuitive result is a clear indicator of significant bias due to indication and selection. It is biologically impossible for more residual disease to directly create a survival benefit. Instead, this group likely represents a distinct population: patients with aggressive tumor biology or poor functional status for whom maximal surgery was deemed futile or unsafe, yet their tumors showed exceptional sensitivity to platinum-based therapies or subsequent targeted treatments. The favorable overall survival of this group, despite poorer PFS, indicates effective salvage therapy. This finding should not be seen as a challenge to surgical principles, but rather as powerful evidence for the role of effective systemic therapy in mitigating the prognostic impact of surgical outcomes.\u003c/p\u003e\n\u003cp\u003eSuperior survival in FIGO stage IV compared to stages I and III further highlights the dominant role of histological subtype over stage alone. The stage distribution in this group is unusual, with a low proportion of early-stage disease. The high survival in stage IV is almost certainly due to the overrepresentation of High-Grade Serous Carcinomas (HGSC), which constituted 32% of the group and showed a 5-year OS of 94.3%. This reflects the impact of effective maintenance therapies for HGSC in the modern era, especially in patients with homologous recombination deficiency.\u003c/p\u003e\n\u003cp\u003eHistology-based survival differences are consistent with existing studies but presented with notable clarity. Excellent OS in mucinous carcinomas (100%) and relatively poor PFS in Low-Grade Serous carcinomas (71.7%) confirm known tumor behaviors. The strong predictive value of response to NACT for PFS (5-year PFS 90% vs. 27.1% for progressive disease, p\u0026lt;0.001) is a key clinical point. However, the non-significant difference in OS based on NACT response (p=0.110) is equally critical; this indicates that while initial tumor chemosensitivity determines the time to first progression, overall survival is shaped by a broader chain of care, including multiple lines of therapy.\u003c/p\u003e\n\u003cp\u003eThe 3, and 5-year overall survival rates were 92.5% (89.8-95.4), and 88.1% (84.7-91.7), respectively. Survival analysis based on different subgroups indicated statistically significant differences associated with FIGO stage, advanced histology type, preoperative CA125 level, and surgical status. The lowest 3 and 5-year survival rates were observed in FIGO stage II patients (both 25%) and in patients with preoperative CA125 level group 3 (82.5% and 71.2%, respectively). In contrast, patients with Mucinous histology and those with a complete response to neoadjuvant therapy showed 3 and 5-year survival rates of 100%. Optimal surgical status (Group 2) was associated with a better prognosis, showing 5-year survival of 91.3% compared to 77% in the suboptimal surgery group. On the other hand, functional status (ECOG) and response to neoadjuvant therapy were not statistically significantly associated with survival rates. These findings emphasize the importance of disease staging factors, histology type, the biomarker CA125, and achieving optimal surgery as key predictors of survival in the studied population (Table1).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eComparative Model Performance\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAs shown in Figure 1 and Table 2, the Random Survival Forest model showed statistically superior performance compared to the Mixture Cure Model in terms of discrimination power, i.e., the ability to accurately distinguish high-risk from low-risk patients for both disease recurrence and overall survival outcomes. The C-index for the Random Survival Forest model was 0.84 for predicting recurrence and 0.78 for overall survival, while these values for the Mixture Cure Model were estimated at 0.78 and 0.75, respectively. These differences, especially in predicting recurrence, were statistically significant (p-value=0.02, based on a bootstrap test with 1000 repetitions), indicating a real and non-random superiority of the machine learning approach in identifying complex patterns associated with disease relapse.\u003c/p\u003e\n\u003cp\u003eFrom the perspective of individual prediction accuracy and calibration, the Random Survival Forest model also had more favorable performance. Its lower integrated Brier score indicates less prediction error and closer alignment of predicted values with observed outcomes over time. This superiority is clearly evident in the calibration plots presented in Table 2 at the 3 and 5-year horizons; the agreement between the hazard curve predicted by the Random Survival Forest model and the empirical Kaplan-Meier survival curve was considerably greater than that of the competing model.\u003c/p\u003e\n\u003cp\u003eAt the time-dependent evaluation level, the time-dependent ROC curves (Figure 2) confirmed the excellent diagnostic ability of the Random Survival Forest model at key clinical time horizons. The Area Under the Curve (AUC) reached 0.98 at the 3-year interval and 0.97 at the 5-year interval. These values, near the maximum possible (1), indicate that the model can differentiate, with very high sensitivity and specificity, patients who will experience the event within these critical intervals from those who will not. This level of accuracy is of significant clinical value, especially for strategic decisions regarding the timing of complementary interventions or adjusting monitoring schedules.\u003c/p\u003e\n\u003cp\u003eThe consistent and significant superiority of the Random Survival Forest model across multidimensional evaluation criteria (discrimination, calibration, and time-dependent diagnosis) establishes it as a powerful quantitative tool for accurate estimation of individual risk in ovarian cancer patients. This capability has high potential to support personalized treatment strategies and optimization of care resources (Figure 1 and Table 2).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRandom Survival Forest Model Insights\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe error rate versus the number of trees shows that the Random Survival Forest model for both recurrence and death outcomes converge to a stable, minimal error level as the number of trees increases. This convergence indicates the stability and reliability of the final model and confirms that using the selected number of trees (likely within the convergence region) prevents overfitting and provides a robust estimate. The final low error rate indicates the model\u0026apos;s overall good ability to classify patients based on event risk.\u003c/p\u003e\n\u003cp\u003eVariable importance analysis reveals an interesting and distinct pattern between the determinants of recurrence and death (Figure 3). For both outcomes, FIGO stage and preoperative CA125 level are among the highly important variables, confirming the fundamental role of tumor burden and disease spread in prognosis. However, key differences exist: variables related to therapeutic intervention (such as type of surgery and response to neoadjuvant therapy) show higher importance for predicting recurrence than for death. This finding suggests that factors related to local disease control and response to initial treatment play a more central role in determining the risk of disease return. In contrast, for the death outcome, age and CA125 had greater relative importance, indicating the increasing influence of the patient\u0026apos;s physiological reserve and overall body resilience on ultimate survival. The presence of BRCA status among moderately important variables also highlights the potential role of molecular factors. This precise differentiation of variable importance provides valuable insight for prioritizing risk factors in the clinic and designing more specific prediction models for each outcome.\u003c/p\u003e\n\u003cp\u003eAs observed in Figure 4, the mortality risk in ovarian cancer patients increases in an approximately linear upward trend as the CA125 level rises (the hazard index value on the Y-axis). This relationship indicates that CA125 is not only a diagnostic biomarker but also a powerful quantitative prognostic indicator. The continuous increase in risk with rising CA125 levels follows a dose-response pattern, such that even intermediate levels (in the range of 35 to 200 on the presented scale) are associated with a significant increase in risk, and very high levels (above 200) maximize the risk. This finding clearly confirms the role of initial tumor burden and potentially more aggressive disease reflected by higher CA125 in determining patients\u0026apos; ultimate survival. This quantified relationship emphasizes that baseline CA125 measurement is an essential component of initial assessment and risk stratification. The constant slope of the curve suggests that there is no specific safety threshold, and any increase in CA125 is associated with worse prognosis. This finding supports the use of CA125 as a continuous variable in individual prediction models, rather than a dichotomized categorical variable. Furthermore, this strong association positions CA125 as a potential surrogate endpoint in targeted clinical trials focused on reducing tumor burden. In summary, this analysis reinforces the fundamental importance of this biomarker in initial evaluation and therapeutic strategy planning.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMixture Cure Model Insights\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Mixture Cure Model estimates a cure fraction of zero for the Progression-Free Survival outcome (Figure 5). This result decisively indicates that, from the perspective of this statistical model, all patients are eventually expected to experience recurrence or disease progression, and there is no subgroup that remains forever free of recurrence. The predicted PFS curve derived from this model lacks a horizontal plateau at high survival probabilities and instead gradually declines. This pattern is fully consistent with the clinical nature of ovarian cancer, known for its high recurrence rate, and emphasizes the major challenge of long-term disease control and preventing recurrence as the primary treatment goal.\u003c/p\u003e\n\u003cp\u003eIn contrast, the model for the Overall Survival outcome predicts a distinct cure plateau at a level above zero (likely corresponding to the data, e.g., 85%). The emergence of this horizontal plateau after approximately 50 to 100 months indicates the existence of a significant subpopulation of patients who do not die from the disease and can be considered cured. The striking contrast between the two plots\u0026mdash;zero cure fraction in PFS versus a positive cure fraction in OS\u003cspan dir=\"RTL\"\u003e\u0026nbsp;\u003c/span\u003ereveals a fundamental clinical pattern of this disease: although most patients will experience recurrence, a large proportion of them ultimately live with the disease as a chronic condition but do not die from it. This gap between controlling recurrence and ultimate survival highlights the need to refocus research and treatment efforts toward effective strategies for delaying recurrence and managing it as a chronic disease, rather than focusing solely on mortality. The results of this model provide a solid statistical basis for this key argument.\u003c/p\u003e\n\u003cp\u003eA survival plateau was clearly delineated. While Cox models assume the risk of death always exists, our data showed that after about 150 months, the survival curve stabilizes (Figure 5). This plateau signifies the existence of a cured fraction; i.e., a subgroup of patients who, after surviving a decade, have a risk of recurrence or cancer death equal to that of the general population. This finding is the most hopeful part of the results for patients and physicians.\u003c/p\u003e\n\u003cp\u003eResults from Table 3 indicate that for the cure component (probability of not experiencing recurrence in the long term) in the PFS outcome, only older age at diagnosis was independently and significantly associated with an increased chance of the patient belonging to the \u0026quot;cured\u0026quot; group. For each additional year of age, this chance increased by an average of 7% (OR=1.07, 95% CI (1.03-1.11), p\u0026lt;0.01). This finding suggests that increasing age may be associated with factors such as different tumor biology or a more favorable response to treatment. For the time-to-event component (for patients susceptible to recurrence), for patients who do not achieve complete cure, two factors significantly shorten the time to disease progression: more advanced disease stage (FIGO) (HR=1.21, 95% CI (1.01-1.40), p\u0026lt;0.05) and higher level of the CA-125 marker before treatment (HR=1.31, 95% CI (1.10-1.58), p\u0026lt;0.01). These results align with clinical expectations and indicate the impact of tumor burden on the timing of recurrence.\u003c/p\u003e\n\u003cp\u003eFor the OS outcome, similar to the PFS outcome, age was a positive and significant predictor of the probability of achieving ultimate cure, although its effect size was smaller (OR=1.03, 95% CI (1.01-1.05), p\u0026lt;0.05). For the time-to-event component (for patients susceptible to death), in this outcome, only the pre-treatment CA-125 level was independently associated with an increased risk of death (HR=1.28, 95% CI (1.18-1.80), p\u0026lt;0.05). A noteworthy point is that, unlike the PFS outcome, FIGO stage was not statistically significant in this model (p=0.10). This may indicate that after recurrence occurs, factors beyond the initial stage (such as response to therapy, biology of metastases) become the primary determinants of patient lifespan. This analysis confirms that predictors of the probability of cure can differ from predictors of the timing of the event.\u003c/p\u003e\n\u003cp\u003eAge, as a host-related factor, is consistently associated with a higher probability of achieving excellent long-term outcomes (both non-recurrence and survival). In contrast, for predicting the timing of adverse events, factors related to the tumor itself gain importance. A key finding is that although FIGO stage influences the time to recurrence, pre-treatment CA-125 level appears to be a stronger indicator for predicting the time to death from the disease. This difference can be useful in designing follow-up protocols and defining high-risk groups. The present model, by disaggregating these components, provides a more nuanced tool for individual-centric prognosis.\u003c/p\u003e\n\u003cp\u003eFinally, integrating the results of machine learning and cure models showed that ovarian cancer management should be based on two different strategies: predicting recurrence, which is heavily dependent on initial treatment response (with 84% accuracy), and predicting ultimate survival, which is a function of age and biological burden. The stability of the survival plateau at the end of month 273 ensures the statistical validity of these findings for use in \u0026quot;Precision Medicine\u0026quot; protocols and counseling patients about long-term prognosis.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThis study, utilizing a unique 24-year cohort, conducted a systematic comparison of two powerful modeling paradigms in predicting ovarian cancer outcomes. Our findings have two key implications: first, machine learning algorithms (Random Survival Forest) are superior in prediction accuracy to structured statistical models (Mixture Cure Model); second, the choice of optimal model should be based on the practical objective, as each approach provides distinct and complementary insights.\u003c/p\u003e\n\u003cp\u003eThe statistically significant superiority of the Random Survival Forest model in the discrimination index, especially for predicting recurrence, aligns with recent study findings that confirm the inherent ability of tree-based methods to identify complex interactions and non-linear relationships without requiring parametric assumptions (14-16). The high accuracy of our model in the first 5 years (AUC \u0026gt; 0.85) indicates the algorithm\u0026apos;s ability to identify aggressive patterns that linear models typically cannot distinguish. Our model showed that the linearity assumption in the Cox model leads to overlooking complex biological interactions. Unlike most existing studies that rely solely on hazard ratios, we demonstrated that prediction accuracy is a more realistic criterion for personalizing treatment. However, this superiority comes at a cost: complexity of interpretation. While the Mixture Cure Model is estimated parametrically and provides direct outputs such as the probability of belonging to the cured group or hazard ratios (HR), interpreting the Random Survival Forest model requires intermediary tools like variable importance and partial dependence plots. This finding reinforces the ongoing discussion in the field of Explainable AI for clinical applications (17). In conditions like ovarian cancer where decisions are critical, model transparency can be as important as its accuracy. Therefore, the Random Survival Forest model may be ideal for pure prediction tasks (such as prioritizing patients for intensive follow-up), while the Mixture Cure Model is more suitable for exploratory and counseling tasks (such as estimating long-term prognosis and discussing the concept of cure with the patient).\u003c/p\u003e\n\u003cp\u003eVariable importance analysis revealed a different pattern for predictors of recurrence and death. The dominance of treatment-focused variables (neoadjuvant response, surgical quality) in predicting recurrence re-emphasizes the central role of initial therapy in local disease control (18). In contrast, the dominance of host-focused variables (age, CA125) in predicting overall survival highlights the need to integrate comprehensive geriatric assessment and optimal management of comorbidities in older patients (19, 20). This distinction is important because it shows that improving overall survival is not necessarily equivalent to \u0026quot;preventing recurrence.\u0026quot; The significant gap between the estimated cure fraction for recurrence (35%) and overall survival (85%) reflects the clinical reality that many patients continue to live for decades despite frequent recurrences (21). This reinforces the paradigm shift from treatment aimed at complete cure toward managing cancer as a chronic disease, where delaying recurrence and preserving quality of life become primary goals.\u003c/p\u003e\n\u003cp\u003eThe discovery of a non-linear threshold relationship between CA125 and mortality risk through partial dependence plots shows the added value of machine learning. This finding, consistent with similar studies (22, 23), can help refine clinical guidelines. Instead of considering a linear increase in CA125 as a uniform concern, prognostic models can focus on identifying patients who cross this critical threshold and target more aggressive interventions for this group. On the other hand, the use of the Mixture Cure Model in this study allowed observation of a reality hidden in standard Cox models. The identification of a survival plateau by the Mixture Cure Model after 150 months is one of the most encouraging findings of this study. Awa Szarkowski (2021) and Zhou et al. (2024) in their analysis of gynecological cancers emphasize that ignoring the cure fraction in long-term studies leads to overestimation of risk in long-term survivors (24, 25).\u003c/p\u003e\n\u003cp\u003eThe use of the Mixture Cure Model alongside Random Survival Forest revealed new dimensions of clinical hope. Identifying a Cure Fraction indicates that contrary to the common perception that ovarian cancer is always progressive, a subgroup of patients achieves long-term stability. These finding challenges Wang\u0026apos;s study, which emphasized a 100% recurrence rate in advanced stages, and doubles the importance of identifying the characteristics of these cured patients. This plateau not only confirms the existence of a subgroup of patients with excellent prognosis but also provides a statistical basis for adjusting follow-up protocols. Our findings showed that about 20% of patients can be considered practically cured after passing the 10-year plateau, which could shift follow-up protocols from aggressive monitoring toward symptom and quality-of-life surveillance, leading to reduced costs and patient anxiety.\u003c/p\u003e\n\u003cp\u003eAnalysis of the Mixture Cure Model provided the possibility for finer differentiation between determinants of long-term outcomes. In the cure component, it was observed that older age at diagnosis was independently and significantly associated with an increased probability of achieving a permanent cured state, both in the PFS metric (OR=1.07) and the OS metric (OR=1.03). This finding suggests that host-related factors associated with age may provide favorable molecular mechanisms for achieving stable remission, although its effect size was smaller in OS. Notably, age showed no significant predictive effect in the incidence component (timing of recurrence or death for non-cured patients). These results, contrary to the study by Kara Lexi et al. (2022), which considered age only effective on time to death (26), indicate that the main effect of age in this population is focused on the probability of achieving an excellent outcome, not on accelerating adverse events.\u003c/p\u003e\n\u003cp\u003eOn the other hand, advanced FIGO stage, a strong prognostic factor in classical studies, was only significant in predicting the time to disease recurrence in the PFS incidence component and had no association with the ultimate probability of cure. This finding aligns with novel ideas presented by Zhang et al. (2024), who emphasize the increasing importance of tumor biology and treatment response over initial anatomical extent of disease in determining ultimate cure (27). Time-dependent ROC analysis showed that the accuracy of the Random Survival Forest model does not severely decline even after 10 years of follow-up (maintaining AUC above 0.75). This stability compared to simple machine learning models that often overfit is a major methodological advantage. Li et al. (2023) believes that using tree-based techniques in censored data is the best way to maintain prediction stability in long-term follow-up (28).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLimitations and Future Directions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study has limitations. Its single-center, retrospective nature may affect the generalizability of the results. Major changes in surgical and systemic protocols over the 24-year study period, although partly considered in sensitivity analysis, remain a challenge. Furthermore, the lack of molecular data and quality-of-life data limited analytical depth. Future research should focus on external validation of these findings in multicenter, independent cohorts (29). Integrating multi-omics data (genomic, imaging, digital pathology) with clinical data within advanced machine learning models is the next step toward truly personalized predictions (30). Finally, the development of hybrid frameworks that combine the predictive power of methods like Random Survival Forest with the structured interpretability of models like the Mixture Cure Model can help solve the dilemma of accuracy versus interpretability (31).\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis study demonstrates that in the era of precision medicine, no single model is the absolute optimum. Machine learning (Random Survival Forest) excels in achieving maximum prediction accuracy, while structured statistical models (Mixture Cure Model) provide unparalleled insight into underlying mechanisms (such as cure). An intelligent combination of these two paradigms, along with an understanding of the fundamental differences between drivers of recurrence and death, can chart a roadmap for personalized, realistic, and hopeful management of ovarian cancer. Our findings support a stratified approach where machine learning guides immediate clinical decisions regarding treatment intensification and surveillance, while cure models inform long-term prognosis and survivorship care planning.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate:\u003c/strong\u003e The study protocol was approved by the Ethics Committee of Shahid Sadoughi University of Medical Sciences, Yazd (code: IR.SSU.SPH.REC.1403.113). Due to the retrospective nature of the study, the requirement for informed consent was waived by the ethics committee.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication:\u003c/strong\u003e Not applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials:\u003c/strong\u003e The datasets generated and analyzed during the current study are not publicly available due to patient privacy and confidentiality regulations but are available from the corresponding author upon reasonable request and with appropriate ethical approvals.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests:\u003c/strong\u003e The authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding:\u003c/strong\u003e This research received no specific grant from any funding agency in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; contributions:\u003c/strong\u003e Data collection was carried out by Dr. Mohammad Reza Mortazavi Zadeh. The draft of the article, along with data analysis using software, was written by Rasoul Najafi, and the final review and editing was carried out by Dr. Hossein Fallah\u003cspan dir=\"RTL\"\u003e\u0026nbsp;\u003c/span\u003eZadeh and Dr. Farimah Shamsi.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements:\u003c/strong\u003e The authors thank the staff of Yazd Specialized Cancer Center for their assistance with data collection and management.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eBray F, Lavers Anne M, Sung H, Fer lay J, Siegel RL, Soerjomataram I, et al. Global cancer statistics 2022: GLOBOCAN estimates of incidence and mortality worldwide for 36 cancers in 185 countries. CA: a cancer journal for clinicians. 2024;74(3):229\u0026ndash;63.\u003c/li\u003e\n\u003cli\u003eVeneziani AC, Gonzalez-Ochoa E, Alqaisi H, Madariaga A, Bhat G, Rouzbahman M, et al. Heterogeneity and treatment landscape of ovarian carcinoma. Nature Reviews Clinical Oncology. 2023;20(12):820\u0026ndash;42.\u003c/li\u003e\n\u003cli\u003eAmir SA, Samin BR, Amin N, Jamshid BM, Habibollah P, Matin BM, et al. Application of machine learning techniques for predicting survival in ovarian cancer. BMC Medical Informatics and Decision Making (Web). 2022;22(1):1\u0026ndash;24.\u003c/li\u003e\n\u003cli\u003eAndreou M, Kyprianidou M, Cortas C, Polycarpou I, Papamichael D, Kountourakis P, et al. Prognostic factors influencing survival in ovarian cancer patients: a 10-year retrospective study. Cancers. 2023;15(24):5710.\u003c/li\u003e\n\u003cli\u003eJakobsen LH, Andersson TM-L, Biccler JL, Poulsen L\u0026Oslash;, Severinsen MT, El-Galaly TC, et al. On estimating the time to statistical cure. BMC Medical Research Methodology. 2020;20(1):71.\u003c/li\u003e\n\u003cli\u003eSumitomo M, Kotani Y, Murakami K, Abiko K, Sakai K, Otani T, et al. Cure of Recurrent Ovarian Cancer: A Multicenter Retrospective Study. Cancers. 2025;17(18):3069.\u003c/li\u003e\n\u003cli\u003eFelizzi F, Paracha N, P\u0026ouml;hlmann J, Ray J. Mixture cure models in oncology: a tutorial and practical guidance. PharmacoEconomics-open. 2021;5(2):143\u0026ndash;55.\u003c/li\u003e\n\u003cli\u003eAsadi F, Rahimi M, Ramezanghorbani N, Almasi S. Comparing the Effectiveness of Artificial Intelligence Models in Predicting Ovarian Cancer Survival: A Systematic Review. Cancer Reports. 2025;8(3):e70138.\u003c/li\u003e\n\u003cli\u003eWei L, Chen G, Liang H, Li L. Random survival forest model in patients with epithelial ovarian cancer: a study based on SEER database and single center data. American Journal of Cancer Research. 2025;15(2):769.\u003c/li\u003e\n\u003cli\u003eWei L, Chen G, Liang H, Li L. Random survival forest model in patients with epithelial ovarian cancer: a study based on SEER database and single center data. Am J Cancer Res. 2025;15(2):769\u0026ndash;80.\u003c/li\u003e\n\u003cli\u003eIshwaran H, Kogalur U. Fast unified random forests for survival, regression, and classification (RF-SRC)(Version 3.3. 1)[R package]. 2024.\u003c/li\u003e\n\u003cli\u003eO\u0026apos;Donnell A, Cronin M, Moghaddam S, Wolsztynski E. A Systematic Review on Machine Learning Techniques for Survival Analysis in Cancer. Cancer Med. 2025;14(22):e71375.\u003c/li\u003e\n\u003cli\u003eTran TT, Lee J, Gunathilake M, Kim J, Kim SY, Cho H, et al. A comparison of machine learning models and Cox proportional hazards models regarding their ability to predict the risk of gastrointestinal cancer based on metabolic syndrome and its components. Front Oncol. 2023;13:1049787.\u003c/li\u003e\n\u003cli\u003eHe C, Liu B, Wang H-Y, Wu L, Zhao G, Huang C, et al. Inhibition of SRPK1, a key splicing regulator, exhibits antitumor and chemotherapeutic-sensitizing effects on extranodal NK/T-cell lymphoma cells. BMC cancer. 2022;22(1):1100.\u003c/li\u003e\n\u003cli\u003eEkuk E, Odongo CN, Tibaijuka L, Oyania F, Egesa WI, Bongomin F, et al. One year overall survival of wilms tumor cases and its predictors, among children diagnosed at a teaching hospital in South Western Uganda: a retrospective cohort study. BMC cancer. 2023;23(1):196.\u003c/li\u003e\n\u003cli\u003eKokori E, Aderinto N, Olatunji G, Abraham IC, Komolafe R, Ukoaka B, et al. Machine learning use in early ovarian cancer detection. Discover Medicine. 2025;2(1):66.\u003c/li\u003e\n\u003cli\u003eDi Martino F, Delmastro F. Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artif Intell Rev. 2023;56(6):5261\u0026ndash;315.\u003c/li\u003e\n\u003cli\u003eDi Martino F, Delmastro F. Explainable AI for clinical and remote health applications: a survey on tabular and time series data. Artificial Intelligence Review. 2023;56(6):5261\u0026ndash;315.\u003c/li\u003e\n\u003cli\u003eLichtman SM, Harvey RD, Damiette Smit MA, Rahman A, Thompson MA, Roach N, et al. Modernizing Clinical Trial Eligibility Criteria: Recommendations of the American Society of Clinical Oncology-Friends of Cancer Research Organ Dysfunction, Prior or Concurrent Malignancy, and Comorbidities Working Group. J Clin Oncol. 2017;35(33):3753\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eMarkman M, Lewis Jr JL, Saigo P, Hakes T, Rubin S, Jones W, et al. Impact of age on survival of patients with ovarian cancer. Gynecologic oncology. 1993;49(2):236\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eUnzelman RF. Advanced epithelial ovarian carcinoma: long-term survival experience at the community hospital. Am J Obstet Gynecol. 1992;166(6 Pt 1):1663\u0026ndash;71; discussion 71\u0026ndash;2.\u003c/li\u003e\n\u003cli\u003eCalifano D, Gallo D, Rampioni Vinciguerra GL, De Cecio R, Arenare L, Signoriello S, et al. Evaluation of Angiogenesis-Related Genes as Prognostic Biomarkers of Bevacizumab Treated Ovarian Cancer Patients: Results from the Phase IV MITO16A/ManGO OV-2 Translational Study. Cancers (Basel). 2021;13(20).\u003c/li\u003e\n\u003cli\u003eYanaihara N, Yoshino Y, Noguchi D, Tabata J, Takenaka M, Iida Y, et al. Paclitaxel sensitizes homologous recombination-proficient ovarian cancer cells to PARP inhibitor via the CDK1/BRCA1 pathway. Gynecologic Oncology. 2023; 168:83\u0026ndash;91.\u003c/li\u003e\n\u003cli\u003eZhou XH, Yang DN, Zou YX, Tang DD, Chen J, Li ZY, et al. Long-Term Survival Trend of Gynecological Cancer: A Systematic Review of Population-Based Cancer Registration Data. Biomed Environ Sci. 2024;37(8):897\u0026ndash;921.\u003c/li\u003e\n\u003cli\u003eStrzalkowska-Kominiak E, Romo J. Censored functional data for incomplete follow-up studies. Stat Med. 2021;40(12):2821\u0026ndash;38.\u003c/li\u003e\n\u003cli\u003eKaralexi MA, Katsimpris A, Panagopoulou P, Bouka P, Sch\u0026uuml;z J, Ntzani E, et al. Maternal lifestyle factors and risk of neuroblastoma in the offspring: A meta-analysis including Greek NARECHEM-ST primary data. Cancer Epidemiol. 2022; 77:102055.\u003c/li\u003e\n\u003cli\u003eZhang Y, Guan Y, Xiao X, Xu S, Zhu S, Cao D, et al. Circulating tumor DNA detection improves relapse prediction in epithelial ovarian cancer. BMC Cancer. 2024;24(1):1565.\u003c/li\u003e\n\u003cli\u003eLi L, Zhao Y, Li H, Zhang S. BLTSA: pseudotime prediction for single cells by branched local tangent space alignment. Bioinformatics. 2023;39(2).\u003c/li\u003e\n\u003cli\u003eBoulesteix AL, Wilson R, Hapfelmeier A. Towards evidence-based computational statistics: lessons from clinical research on the role and design of real-data benchmark studies. BMC Med Res Methodol. 2017;17(1):138.\u003c/li\u003e\n\u003cli\u003eBerger AC, Korkut A, Kanchi RS, Hegde AM, Lenoir W, Liu W, et al. A Comprehensive Pan-Cancer Molecular Study of Gynecologic and Breast Cancers. Cancer Cell. 2018;33(4):690\u0026ndash;705.e9.\u003c/li\u003e\n\u003cli\u003eAmico M, Van Keilegom I. Cure models in survival analysis. Annual Review of Statistics and Its Application. 2018;5(1):311\u0026ndash;42.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Tables","content":"\u003cp\u003eTables 1 to 3 are available in the Supplementary Files section.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Ovarian Cancer, Random Survival Forest, Mixture Cure Models, Machine Learning, Precision Medicine, Survival Analysis, Prognostic Prediction","lastPublishedDoi":"10.21203/rs.3.rs-8694452/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8694452/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground:\u003c/strong\u003e Accurate survival prediction in ovarian cancer has been challenged due to biological heterogeneity and the cure phenomenon, where a subgroup of patients achieves long-term remission. This study was designed to compare the predictive capability of the machine learning algorithm Random Survival Forest with the statistical Mixture Cure Model over a 24-year timeframe, addressing a significant gap in direct comparative evaluations.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMethods:\u003c/strong\u003e Data from 352 ovarian cancer patients (2000-2024) with the primary outcomes of Overall Survival (OS) and Progression-Free Survival (PFS) were analyzed. We implemented two modeling approaches: a semi-parametric Mixture Cure Model (with logistic cure and Cox latency components) and a non-parametric Random Survival Forest ensemble of 1000 survival trees. Model performance was evaluated using the discrimination index (C-index), integrated Brier score, and time-dependent AUC via cross-validation. Variable importance and non-linear effects were explored through partial dependence plots.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eResults:\u003c/strong\u003e The mean age at diagnosis was 45.97 ± 15.22 years, with 59.7% presenting at advanced stages (FIGO III/IV). During 273 months of follow-up, 115 recurrences (33%) and 55 deaths (16%) occurred. Variable importance analysis revealed that response to neoadjuvant therapy and surgical quality were the most powerful predictors of recurrence, while age at diagnosis and baseline CA125 level guided overall survival. Partial dependence plots revealed a non-linear threshold effect for the CA125 marker. The Random Survival Forest model significantly outperformed the Mixture Cure Model in predicting recurrence (C-index = 0.84, 95% CI: 0.80–0.88 vs. 0.78, 95% CI: 0.74–0.82; p=0.02) and showed better discrimination for death (C-index = 0.78, 95% CI: 0.74–0.82 vs. 0.75, 95% CI: 0.71–0.79; p=0.15). The Mixture Cure Model confirmed the existence of a survival plateau after 150 months, indicating a significant cure fraction (~85%) in this cohort for OS, while no cure fraction was identified for PFS.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConclusion\u003c/strong\u003e Our findings demonstrate that machine learning, due to its ability to uncover non-linear interactions, is a superior tool for precision oncology. Distinguishing recurrence drivers (treatment-focused) from mortality drivers (host-focused) provides a new roadmap for personalized patient management and optimization of long-term follow-up protocols.\u003c/p\u003e","manuscriptTitle":"Benchmarking Machine Learning against Mixture Cure Models for Personalized Recurrence and Survival Prediction in Ovarian Cancer: A 24-Year Retrospective Cohort Study","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-03-11 06:11:54","doi":"10.21203/rs.3.rs-8694452/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"69122269266589888744342528912858005897","date":"2026-03-07T12:24:06+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-03-05T12:18:14+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-05T06:04:13+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2026-02-09T09:24:58+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-02-07T08:01:47+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Cancer","date":"2026-02-07T07:54:03+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bmc-cancer","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bcan","sideBox":"Learn more about [BMC Cancer](http://bmccancer.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/bcan/default.aspx","title":"BMC Cancer","twitterHandle":"BMC_series","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e76cf0cb-7a9a-4d34-b319-460e844cbe17","owner":[],"postedDate":"March 11th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-03-11T06:11:54+00:00","versionOfRecord":[],"versionCreatedAt":"2026-03-11 06:11:54","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8694452","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8694452","identity":"rs-8694452","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.