Intro
Despite therapeutic advances, ovarian cancer remains a leading cause of death in gynecologic oncology ( 1 ). High-Grade Serous Ovarian Carcinoma (HGSOC) is the most aggressive histologic subtype ( 2 ). A significant proportion of patients are diagnosed at advanced stages, correlating with unfavorable survival results. The Cancer Genome Atlas (TCGA) has indicated that HGSOC is characterized by almost universal Tumor Protein p53 (TP53) mutations. These molecular traits explain why different malignancies respond differently to certain therapies ( 3 ). These molecular characteristics lead to early treatment resistance and recurrent relapses, which continue to pose significant therapeutic hurdles.
Current management is constrained by delayed diagnosis and variable therapy response. Histopathological assessment of hematoxylin-eosin (H&E) staining and immunohistochemistry (IHC) constitutes the diagnostic gold standard; nonetheless, it is time-intensive and prone to inter-observer variability ( 4 ). Radiologic modalities, including computed tomography (CT), magnetic resonance imaging (MRI), and ultrasound, are crucial for detection and staging; nevertheless, conventional interpretation may overlook nuanced imaging correlates of tumor biology ( 5 ). Commonly utilized biomarkers exhibit restricted sensitivity and specificity for early-stage illness and for forecasting therapy response ( 6 ). These limitations underscore a significant obstacle in the therapy of ovarian cancer: the growing number and complexity of clinical, imaging, and genetic data are challenging to handle with traditional analytical methods.
In this instance, AI, particularly machine learning (ML) and deep learning (DL), can model complex patterns in high-dimensional clinical, imaging, and molecular data, and enable automated complicated connections between large groups of patients ( 7 ). Research indicates that image-based analysis and digital pathology can attain performance levels equivalent to those of human experts. In ovarian cancer, first trials integrating imaging results with fundamental clinical data have already shown significant advantages ( 8 ). However, limited transparency, domain shift, and bias remain key barriers still make AI challenging in regular clinical practice ( 9 ).
Accordingly, this review synthesizes progress across six principal domains (1) Early Detection and Treatment Response Prediction with imaging and pathology; (2) Prediction of HRD from Routine Pathology; (3) Clinical Data Integration using longitudinal records; (4) Multimodal Data Fusion strategies; (5) Spatial Omics and AI for mapping tumor ecosystems; and (6) Model Interpretability and Clinical Translation considerations. In general, ovarian cancer AI is moving away from performance-based models and toward clinically useful systems. This review shows how AI research can be in line with real-world decision-making ( Figure 1 ).
AI-enabled multimodal framework for precision management of ovarian cancer. An integrated framework using AI to facilitate precision management of ovarian cancer. This framework combines multiple data modalities, including radiology (ultrasound, CT, MRI), digital pathology (H&E staining, immunohistochemistry, whole-slide images), molecular profiling (e.g., CA125), clinical data, EHR data, and spatial-temporal multi-omics. AI models process these diverse data sources for early detection, risk stratification, treatment response prediction, immune evasion, and personalized therapeutic decision-making. By integrating imaging, clinical, and molecular data, the framework enhances ovarian cancer diagnosis, prognosis, and treatment selection, leading to more accurate and personalized care.
Model
As new tools are developed for ovarian cancer, questions have shifted toward how their results are obtained and whether they can be used safely in real-world care. However, making these tools understandable, reliable across different patient groups, and practical for routine use remains challenging ( 150 ).
The application of AI models in ovarian cancer has rendered interpretability a crucial necessity for clinical reliability and secure implementation. Nonetheless, the lack of transparency particularly in medical settings, drives efforts to make their decision-making processes easier to understand and trust ( 151 ). Recent research had shown that interpretability alone may not be enough for clinical usage, even though explainable AI (XAI) provides technological ways to show and attribute model predictions. The concept of causal interpretability has been established to distinguish system-level explainability from the human capacity to understand, contextualize, and reason causally about model outputs. From this perspective, explainability refers to a trait of the AI system, whereas causal understanding signifies the caliber of explanations acknowledged and utilized by physicians. This distinction is especially relevant in medicine, as AI-generated explanations should improve clinical rather than merely disclose internal model attributes ( 152 ).
Bias can arise at any stage of medical AI development, from data collection and labeling to model training, validation, and clinical use. If these biases aren’t dealt with, they could affect clinical judgments and make current healthcare disparities worse, especially in high-risk clinical situations. Bias can occur because to skewed or inadequate datasets, incorrect labeling, an undue emphasis on aggregate performance metrics, or reduced accuracy when models are applied to patient cohorts that differ from those utilized during development. To reduce these dangers, individuals need to use a range of representative data, look at how well each subgroup does, make model outputs clearer, report any apparent biases clearly, and do thorough testing before applying them in clinical settings ( 153 ).
Although XAI have improved model transparency, the reliability of current explanation methods remains a concern. Research indicates that little alterations in input data might result in significantly divergent explanations. Such inconsistency can make clinicians less sure of themselves and make it harder to make clinical decisions. These discoveries underscore the necessity to assess explanatory methodologies in the context of realistic data variability and to prioritize those that produce consistent outcomes prior to clinical application ( 154 ).
These findings underscore the necessity to limit bias and ensure clear, trustworthy and equitable use of AI in practice.
The goal of current AI solutions for ovarian cancer is to help clinicians make better judgments, not to take their place. The multimodal, uncertainty-aware AI framework for assessing ovarian cancer risk demonstrates that human-AI collaboration yields superior outcomes compared to individual ( 155 ).
Recent developments in AI have prompted a shift from technology-centric approaches to human-centered AI (HCAI), which emphasizes the design of AI systems to better meet human needs. HCAI focuses on human control, responsibility, and empowerment when using AI-powered technology, rather than just how well the algorithms work. Researchers and practitioners may guarantee that AI systems augment human skills, promote well-being, and provide a positive impact on society by integrating human-centered principles in the design, development, and evaluation of AI. This way, AI systems don’t take over or get rid of human tasks ( 156 ). From a human-centered AI point of view, these kinds of systems can be thought of in terms of their goal (augmentation, automation, and autonomy), values (ethics, safety, and performance), and attributes (oversight, comprehensibility, and integrity). This approach supports doctors in making better decisions, but doctors remain responsible for the final call ( 157 ).
These approaches together make ovarian cancer AI look more like a system that gather doctors work together rather than a replacement for human.
Differences between research datasets and real-world clinical data sometimes lead to worse results, which makes it impossible to use these methods in real life ( 158 ).
Differences in how WSIs are collected, how tissue samples are stained, and where the data come from can lead to shifts in data characteristics, making it difficult for pathology analysis methods to perform reliably across different hospitals and datasets ( 159 ). The WSI-P2P architecture was presented to tackle this difficulty as a domain-robust whole-slide analysis method that reduces performance decline across diverse datasets. The consistent performance in both intra- and inter-domain evaluations underscores the promise of scalable multiple-instance learning frameworks to enhance robustness against domain shifts in practical histopathology applications ( 160 ). Limited reliability of medical imaging tools across clinical settings arises from both overfitting and insufficient constraints, underscoring the need for rigorous evaluation in real-world settings ( 161 ).
The transition of domains through strong modeling and rigorous evaluation is vital for reaching dependable and clinically useful AI.
The acquisition of adequately vast and diversified datasets is frequently obstructed by privacy and data ownership limitations. FL has emerged as a solution by allowing AI models to be trained collaboratively on data from multiple institutions without the data ever leaving its source site ( 162 ). In FL, a common model is initiated and sent to each institution, where it is locally trained on that site’s data; only the model weight updates (and not the raw data) are then sent back and aggregated to form an improved global model while keeping patient data private ( 163 ). In ovarian cancer, FL is increasingly used to combine WSIs from multiple hospitals, addressing the limited availability of rare subtypes at individual sites. HFed-MIL is a heterogeneous federated multiple-instance learning framework for ovarian cancer pathology that enables cross-institutional model training under privacy constraints, improves robustness to site-specific heterogeneity, and maintains model interpretability ( 164 ).
Researchers are addressing these by developing adaptive federated optimizers and strategies like FedProx, FedAvg improvements, as well as ensuring each site’s contribution is weighted properly. Privacy is further safeguarded in FL by techniques such as differential privacy (adding noise to updates so individual patient contribution is obfuscated) and secure multi-party computation, as highlighted by the systematic review ( 165 ) These ensure that even the model updates don’t inadvertently leak patient info.
We foresee FL becoming standard in multi-institutional studies, and possibly a requirement for regulatory approval that a model has been trained/validated on diverse patient data.
Despite rapid methodological advances, most AI models developed for ovarian and pancreatic cancer have not yet been translated into clinical practice, largely due to limited and non-public imaging datasets, low disease prevalence, and the absence of Food and Drug Administration (FDA)-qualified imaging biomarkers ( 166 ). These challenges are intensified by concerns about data quality, bias, and transparency, highlighting the need for more interpretable and reliable models for rare cancers.
Translating laboratory findings into clinically applicable tools requires rigorous clinical validation, clear definition of their capabilities and limitations, and adherence to regulatory standards ( 166 , 167 ).
In summary, translating a promising AI model into an approved clinical tool for ovarian cancer requires an emphasis on explainability (to foster clinician trust and guarantee patient safety), robustness (to manage real-world variability), collaborative design (to enhance clinician expertise rather than circumvent it), and stringent clinical validation ( Figure 6 ).
Model interpretability, generalizability, and clinical translation of AI. This schematic illustrates key components required to translate AI models for ovarian cancer into clinical practice. Model interpretability and trust are enhanced through explainable AI and causability analyses, together with systematic assessment of bias, robustness, and human-AI interaction. Generalizability and domain adaption is addressed by evaluating domain shift between training and real-world data, complemented by stress testing and multi-cohort validation. Path to clinical deployment is supported by privacy-preserving deployment strategies, such as FL, and by rigorous regulatory validation and approval processes. Collectively, these elements define a pathway toward reliable, safe, and clinically actionable AI applications.
Clinical
Clinical data integration underpins prognosis and risk stratification by capturing patient-level risk factors in ovarian cancer. In addition to images data, the longitudinal clinical record of an ovarian cancer patient harbors valuable information for prognosis and treatment personalization ( Figure 3 ). Factors such as a patient’s age, performance status, prior treatment history, comorbidities, and dynamic tumor marker levels (e.g. serial Cancer Antigen 125 (CA125) measurements) all influence outcomes. For example, large prospective cohort studies such as the SOCFCP integrate clinical, genetic, lifestyle, and biospecimen data to enable biomarker discovery and predictive modeling for ovarian cancer diagnosis, treatment, and prognosis ( 63 ). Moreover, AI models are increasingly incorporating these real-world clinical data to improve predictive performance in ovarian cancer, moving beyond static snapshot predictions to a more holistic, longitudinal view of the patient ( 64 ).
Clinical data integration for prognosis and risk stratification in ovarian cancer. Integration of clinical data for prognosis and risk stratification in ovarian cancer. AI models combine clinical information, including demographics, performance status, comorbidities, tumor markers (e.g., CA125, HE4), EHRs (including clinical notes, laboratory results, symptoms, and prior treatments), and dynamic inflammatory and immune-related markers are jointly incorporated into a unified model. By leveraging serial measurements and large-scale EHR data, the framework enables individualized prediction of survival risk, treatment resistance, and recurrence risk, supporting more precise and personalized clinical decision-making.
Overall, routine markers remain central to risk modelling, and their value increases when combined with complementary clinical or imaging information.
Tumor markers are widely used as structured input features for ovarian cancer prediction because they provide quantitative signals related to tumor burden and malignancy. A large retrospective cohort compared multiple statistical and ML models for estimating malignancy risk across pathological categories. Using clinical and ultrasound features, with or without CA125, models showed consistently strong discrimination between benign and malignant tumors (Area Under the Receiver Operating Characteristic Curve (AUROC) ~0.89-0.92) ( 65 ). Although overall diagnostic accuracy was similar across approaches, tree-based ensemble models such as random forest (RF) and XGBoost showed better multiclass calibration and slightly higher clinical net benefit at commonly used risk thresholds and its importance has been consistently confirmed across algorithms such as RF, Recurrent Neural Networks (RNNs), and gradient boosting methods. RF, RNN, and Extreme Gradient Boosting (XGBoost), also achieved high accuracy when incorporating CA-125 ( 66 ). Models that combined CA125, HE4, and routine hematological features also achieved diagnostic accuracies above 85%, with CA125 and HE4 consistently ranking among the most informative variables ( 67 ). Gradient boosting models further supported the diagnostic contribution of HE4 and CA72–4 for distinguishing ovarian cancer from benign tumors. The longitudinal multivariable modeling approaches, including Bayesian change-point detection and RNN, integrated serial CA125 and HE4 measurements for early ovarian cancer detection ( 68 ). The significance of ML in ovarian cancer management lies in improving early detection accuracy, optimizing treatment planning, significantly surpassing traditional statistical methods, and promoting personalized treatment, ultimately improving patient prognosis ( 69 ).
Inflammatory markers reflect the association between systemic inflammation and cancer progression. For blood markers, AI-based models integrating circulating tumor cells and other blood-derived biomarkers achieved good performance in predicting platinum sensitivity and disease stage in ovarian cancer, with AUCs of approximately 0.80 ( 70 ). In Endometrioid Ovarian Carcinoma (EOC), blood-based AI models have been developed for preoperative diagnostic and prognostic prediction by integrating routine peripheral blood biomarkers. Using supervised ML approaches, particularly ensemble tree-based methods such as RF and gradient boosting, these models achieved high accuracy in distinguishing EOC from benign ovarian tumors (AUC up to 0.97), and showed moderate performance in predicting clinical stage, histologic subtype (including high-grade serous and mucinous tumors), residual tumor burden, and survival risk ( 71 ). A 2025 systematic review and meta-analysis evaluated the diagnostic performance of AI-derived blood biomarkers for ovarian cancer and demonstrated overall favorable accuracy. In 40 investigations, AI-based models had a pooled sensitivity of around 85% and a pooled specificity of about 91%. Their performance was particularly strong in analyses based on serum markers with external validation ( 72 ). These findings suggest that blood-based biomarkers could function as noninvasive tools for ovarian cancer diagnosis. Furthermore, inflammatory and hormonal indicators such as C-Reactive Protein (CRP), LDH, the Neutrophil-to-Lymphocyte Ratio (NLR), the Platelet-to-Lymphocyte Ratio (PLR), and the Lymphocyte-to-Monocyte Ratio (LMR), in addition to immune cell ratios, have been incorporated into ML models. This combination enhances the capacity to predict ovarian cancer by revealing inflammation and immunological dysregulation associated with malignancies ( 73 , 74 ).
Epigenetic markers, including transformer-based analysis of cfDNA methylation, have shown considerable diagnostic precision. By identifying disease-specific methylation indicators and transforming them into Droplet Digital Polymerase Chain Reaction (ddPCR)-based assays, these approaches highlight a feasible approach for non-invasive and therapeutically relevant early screening ( 75 ).
When ultrasound results were paired with CA125 and HE4, the diagnosis was more accurate than when imaging was utilized alone. The integrated model exhibited significant accuracy and specificity, indicating that traditional biomarkers like CA125 and HE4 possess increased diagnostic utility when combined with imaging features ( 76 ). A multicenter retrospective study has created a combined prediction method that uses multi-criteria decision making-based classification fusion to merge information from different test findings to make diagnoses more accurate. The AI model reached an AUC of 0.949 in internal validation and steady performance in two external validation cohorts (AUCs ≈0.88). Using 51 routine laboratory markers and age, the model outperformed CA125 and HE4, particularly in early-stage ovarian cancer. Its performance remained robust even without key tumor markers, suggesting a low-cost and widely applicable approach for risk prediction ( 77 ). Table 2 provides a structured overview of representative AI-based diagnostic studies leveraging diverse biomarkers in ovarian cancer.
Representative biomarker cases of AI based ovarian cancer diagnosis.
Beyond imaging and molecular biomarkers, EHRs provide rich longitudinal real-world clinical data that can be leveraged by multivariate AI models for ovarian cancer prediction. EHRs capture patient demographics, laboratory results, comorbidities, treatment history, and healthcare utilization, enabling data-driven models to characterize disease risk and clinical outcomes in a more holistic manner ( 82 ). In ovarian cancer, EHR-based AI approaches can further incorporate prior treatment strategies, treatment-free intervals, performance status, and symptom information documented in clinical notes, thereby supporting individualized risk stratification and outcome prediction across the disease. Building on these advantages, EHR-driven AI models have been explored in several clinically relevant situations ( 83 ).
EHRs are being utilized more and more to develop predictive models that find people who are at a higher risk of ovarian cancer by combining a lot of clinical data that is collected on a regular basis. Unlike traditional models that focus on single markers, EHR-based ML frameworks can accommodate demographics, comorbidities, symptoms, laboratory results, and other longitudinal health information, potentially enhancing early risk stratification ( 84 , 85 ).
A recent study utilized statistical analysis and various ML models on routine clinical data from 349 individuals to facilitate early identification of ovarian cancer. The method used traditional statistical screening along with classifiers like RF, SVM, and gradient boosting models to find important serum tumor markers (like CA-125, CA19-9, Carcinoembryonic Antigen (CEA), and HE4) as well as inflammatory and biochemical blood indices. It was able to correctly tell the difference between malignant and benign ovarian tumors 91% of the time. These findings highlight the prospective efficacy of blood-based markers as supplementary tools for the early detection of ovarian cancer ( 86 ). Combining molecular tumor indicators with standard clinical and laboratory data is important because it makes the model much stronger and allows for a more full, multi-dimensional way to figure out how likely someone is to get ovarian cancer ( 87 ).
A DL-based survival model to forecast the risk of Venous Thromboembolism (VTE) in patients with EOC utilizing longitudinal EHR data. The model attained significantly superior predictive performance (AUROC 0.95-0.98) compared to traditional classification methods by integrating dynamic clinical information and considering competing risks. The system facilitated time-dependent, personalized risk assessment, demonstrating how EHR-driven DL models can aid in tailored complication risk forecasting and guide risk-based clinical treatments in ovarian cancer ( 84 ). Recent studies have shown how ML may use publicly available structured clinical datasets to combine demographic characteristics, tumor markers, and imaging-related aspects to predict ovarian cancer. For instance, a structured clinical dataset of 349 individuals was utilized to create a predictive model that distinguishes benign ovarian tumors from ovarian cancer using routinely gathered demographic and laboratory characteristics. Feature selection indicated HE4 and carcinoembryonic antigen (CEA) as the most informative indicators, with CEA offering supplementary value in patients with low HE4 levels. The resulting two-marker model exhibited enhanced discriminatory performance relative to the risk of ROMA, illustrating that succinct multivariable models can attain precise ovarian cancer risk categorization ( 78 ).
AI has been used to acquire patient-centered information that is vital for making clinical choices but is usually missing from structured datasets. It can also forecast risks and model outcomes. In ovarian cancer, patients’ values, preferences, and goals of care (GOC), encompassing attitudes towards treatment intensity and end-of-life care, are often recorded in free-text clinical notes instead of standardized areas. To fill this gap, recent research has used natural language processing (NLP) methods to get GOC-related information from unstructured EHR data ( 88 ). These methods blend organized clinical data with unstructured clinical notes to better understand what patients want, what they value, and their treatment goals. Text from clinical records is mapped to standardized medical concepts, such as Unified Medical Language System (UMLS) concept unique IDs, to ensure consistent interpretation across different records. To improve accuracy and reduce false positives, the results are reviewed manually and refined through repeated adjustments ( 89 ).
Together, these studies show that information drawn from EHR can be used to better identify high-risk patient groups, and can help guide preventive care and the management of complications in routine clinical practice.
Population-scale EHR studies suggest that longitudinal clinical history can capture weak pre-diagnostic signals at scale, supporting broader and potentially earlier risk stratification. AI models trained on longitudinal EHR data can identify how different clinical factors change and interact over time in ways that simpler risk models often fail to capture ( 90 ).
Conventional screening criteria rely on limited factors (e.g., age or selected risk indicators) and therefore miss many cases. Longitudinal EHR models aim to detect subtle pre-diagnostic patterns that are not captured by traditional risk variables. In the All of Us Research Program (>865, 000 participants), EHR-based models identified 3–6 times more true cancer cases across multiple cancer types, including ovarian cancer, than conventional indicators such as family history or known genetic variants ( 91 ).
Furthermore, neural network models trained on frequently gathered personal health data were created to categorize ovarian cancer risk in high-risk groups. The model exhibited moderate discriminative performance (AUC = 0.71) utilizing extensive population datasets and successfully categorized individuals into high- and low-risk groups, illustrating the potential of non-invasive, EHR-derived AI methodologies to enhance targeted ovarian cancer screening strategies ( 92 ). Large-scale EHR-based prediction models have shown that they can help find people who have a high chance of getting more than one type of cancer, like ovarian cancer. In order to do better than age and family background, these models use real-world data that is collected over time. These results show a lot more risk and help make it easier to use EHRs to predict risk in personalized, scalable cancer screening methods ( 93 ).
More recently, lightweight approaches that use only longitudinal medical history events have been explored for population-scale deployment. Can-SAVE, evaluated across 2.5 million adults from five regions, showed consistent gains in both retrospective (1.9M patients; 4-10× higher detection; AP 0.228 vs. 0.193) and prospective (426K patients; +91% detection; +36% coverage) studies, highlighting its effectiveness at scale. This study shows that using EHR data can support broader, cost-effective cancer detection across large populations, even within the practical constraints of real-world healthcare systems ( 94 ). This process goes beyond simple pattern recognition, capturing the intricate organization and temporal structure of patient health trajectories that may be missed by traditional methods.
AI has been increasingly applied to capture patient values and preferences from diverse data sources, offering potential to enhance patient-centered care, while also facing challenges such as data quality, real-world validation, and ethical concerns ( 95 ). To ensure that AI models used in clinical settings accurately reflect patient experiences, it is essential to incorporate well-defined data dictionaries for PROs ( 96 ). A data dictionary provides a standardized framework for defining and coding variables, ensuring that patient-reported data is consistently interpreted across studies and clinical settings. This standardized approach is applied to other key outcomes, such as pain, fatigue, and emotional distress.
Pain score: A measure of patient-reported pain intensity, categorized as an ordinal variable using a 0–10 numeric scale, based on the PRO-CTCAE ( 97 ). Nausea: A treatment-related symptom, classified as either binary (Yes/No) or ordinal, with severity graded using a numeric scale, according to PRO-CTCAE ( 98 ). Fatigue/Emotional distress: A measure of the severity of fatigue/depression, recorded on a 1–5 Likert scale, as part of the PROMIS system ( 99 ). QoL-physical: Represents physical functioning status, quantified as a continuous variable using a normalized T-score, also part of PROMIS ( 100 ).
With the increasing use of EHRs, clinical information extraction (IE) plays a crucial role in automating data extraction and supporting clinical research ( 101 ). To facilitate this process, advanced systems like the Unified Medical Language System (UMLS) have been developed, which integrate over 2 million names for around 900, 000 concepts from more than 60 biomedical vocabularies and links concepts to external resources like GenBank ( 102 ).
Patient involvement in research has been shown to enhance the relevance and quality of healthcare studies. The Guidance for Reporting Involvement of Patients and the Public 2 (GRIPP2) checklist was developed to improve the transparency and consistency of reporting patient and public involvement in research. It includes both long-form and short-form versions, with 34 and 5 key items, respectively, covering aims, methods, outcomes, and the impact of involvement ( 103 ). Involving patients and stakeholders in research proposal reviews, as seen in the Patient-Centered Outcomes Research Institute (PCORI), showed that while initial scores varied from scientists, in-person discussions improved agreement among reviewers ( 104 ). The potential of AI is in enhancing personalized diagnostics for periodontal diseases, which often require early detection and individualized treatment approaches. Unlike traditional methods that use a “one size fits all” approach, AI can analyze diverse data sets, such as clinical records, imaging, and molecular information, to tailor treatments to each patient’s unique condition ( 105 ).
Data dictionaries not only ensure consistent data capture but also enable cross-study comparisons and data aggregation from diverse sources. Collaboration with patients and individuals with lived experience in developing these criteria ensures that the scales reflect real-world experiences, making the measures more relevant and accurate for clinical applications, and bridging the gap between clinical data and lived experiences in AI models.
Discussion
Although predictive performance in ovarian cancer research has improved substantially, accuracy alone is insufficient to ensure clinical relevance. Many approaches have limited practical impact without explicit linkage to clinical decision-making.
Currently, the AI applications most proximate to clinical implementation are those that conform to existing diagnostic procedures, especially image-based triage and digital pathology. Standardized inputs, clear clinical queries, and dependable reference standards are helpful for tasks like classifying adnexal masses and analyzing whole-slide histopathology. In these situations, AI is mostly used to help people make decisions. It makes things more consistent, efficient, and easier to get expert-level assessments, especially when resources are limited. These systems work best when they add to, rather than replace, the judgment of the clinician.
A high level of prediction accuracy is not enough to make therapeutic use acceptable. For AI outputs to have clinical significance, they must facilitate practical decisions, such referral prioritizing, treatment planning, or follow-up tactics, and exhibit benefits on clinically pertinent outcomes. These include cutting down on unneeded operations, making risk classification better, or making it possible to intervene sooner. To show that anything is better than present standards of care, it needs to be tested in the future and in the actual world.
Early-fusion multimodal models and pan-omics frameworks are still mostly in the testing stage. These methods are based on biology, but they have problems since the sample sizes are small, the modality overlap is not complete, and it is hard to combine the data. Currently, they are more appropriate for hypothesis formulation and biological understanding than for standard clinical application. The main problems with translation between modalities are not model complexity or performance, but external validation, interoperability, interpretability, and integration into clinical workflows.
Successful clinical deployment of AI in ovarian cancer will require more than technical advances. Ethical considerations, bias mitigation, data governance, and patient privacy must be addressed alongside regulatory and implementation challenges. Approaches such as FL, explainable AI, and human-centered design provide practical pathways to address these issues. Progress will be contingent upon the alignment of AI development with well-defined cases and the assessment of systems across varied patient demographics and healthcare environments.
Taken together, the greatest impact in ovarian cancer care is likely to arise from careful integration of AI into existing clinical workflows, where its strengths are leveraged and its limitations are clearly understood. To highlight key aspects of the current research, Box 1 summarizes the criteria for evaluating clinical data in ovarian cancer AI models, Box 2 outlines the limitations of integrating diverse datasets, and Box 3 presents potential future directions for AI in ovarian cancer diagnosis and treatment.
Currently, AI applications in these rare ovarian cancer subtypes are in their infancy. Preliminary studies, mostly radiomics-based, have demonstrated feasibility in subtype classification and outcome prediction. Additionally, some studies have started to explore specific subtypes, such as clear cell carcinoma, mucinous carcinoma, and undifferentiated carcinoma.
Ovarian clear cell carcinoma (OCCC):
OCCC is an uncommon subtype of epithelial ovarian cancer, often arising in an endometriosis setting and difficult to distinguish preoperatively. Emerging AI work has focused on risk stratification and imaging-based classification. Guo et al. (2023) applied ML algorithms to predict distant metastasis in OCCC. The RF model outperformed the others, achieving an accuracy of 0.792, sensitivity of 0.904, specificity of 0.759, and an AUC of 0.834. The Brier Score for the RF model’s calibration curve was 0.06256 ( 168 ). Similarly, Takeyama et al. (2024) applied MRI radiomic analysis to distinguish clear cell vs. endometrioid carcinomas, reporting ROC AUCs of 0.895, 0.910, and 0.956 for their MRI, radiomic, and combined models respectively ( 169 ). However, these studies are limited by small, single-institution cohorts and a lack of independent external validation. No AI models have explicitly incorporated molecular markers, for guiding OCCC prognosis or therapy.
Mucinous ovarian carcinoma (MOC):
AI research in mucinous ovarian carcinoma remains sparse. Available studies have explored imaging-based subtype classification. Yang et al. (2023) developed ultrasound-based radiomics models to predict the five major histological subtypes of epithelial ovarian cancer. For low-grade serous, endometrioid, and clear cell carcinomas, the models had AUCs below 0.80. For mucinous carcinoma, AUCs ranged from 0.80 to 0.89 ( 170 ). Separately, Hallberg et al. (2025) performed a large multi-omic genomic analysis of MOC, identifying recurrent alterations (e.g. SMARCA4, NF1) and showing that ovarian mucinous tumors have distinct mutation/methylation patterns from gastrointestinal mucinous tumors ( 171 ). Critically, distinguishing primary ovarian mucinous carcinoma from metastatic gastrointestinal mucinous tumors remains a significant challenge.
EOC:
EOCs are molecularly similar to uterine endometrial carcinomas, sharing similar alteration burdens and methylation profiles. EOCs accounts for ~10% of epithelial ovarian cancers and displays broad morphologic diversity that complicates diagnosis and grading. Takeyama et al. radiomics study treated endometrioid tumors as the comparator group and achieved AUC ≈ 0.956 for the combined model ( 169 ). In Yang et al.’s ultrasound radiomics analysis, the endometrioid carcinoma model yielded AUC below 0.80 ( 170 ). It is worth noting that no AI model has yet been reported for detecting microsatellite instability (MSI) or mismatch repair status in ovarian endometrioid carcinoma.
In summary, AI applications in these rare ovarian cancer subtypes are in their infancy. Future progress will likely depend on large-scale, multi-modal collaborations (imaging + pathology + genomics) to build robust AI models for these understudied subtypes.
Key criteria for AI-worthy EHRs:
• Structured data: Machine-readable variables (e.g., labs, medications, vital signs).
• Standardized coding: Use of controlled vocabularies (ICD, SNOMED-CT, LOINC).
• Data completeness: Minimal missingness, especially for key clinical variables.
• Longitudinal records: Time-stamped data enabling dynamic modeling.
• Interoperability: Compatibility across systems (e.g., HL7 FHIR).
• Validated outcomes: Reliable ground truth labels (e.g., confirmed diagnoses, survival).
Major limitations of EHR data:
• Heterogeneity: Variability across institutions in data structure and coding.
• Missing data: Incomplete laboratory, symptom, and follow-up information.
• Bias: Unequal data quality across demographic groups.
• Unstructured data burden: Dependence on NLP for clinical notes.
• Privacy constraints: Restricted data sharing and limited external validation.
Practical recommendations:
• Harmonize coding systems and map local terminologies.
• Implement data quality control pipelines.
• Perform external validation across institutions.
• Conduct subgroup and fairness analyses.
• Ensure transparent reporting of preprocessing and data provenance.
The increasing use of AI in healthcare highlights the need for “AI-ready” EHR data. Frameworks such as CODE-EHR emphasize standardized coding, transparent phenotyping, and reproducible data definitions ( 172 ). The use of controlled vocabularies and interoperability standards such as HL7 FHIR further facilitates data integration across institutions. In addition, structured longitudinal data with time-stamped variables and validated outcomes are essential for dynamic modeling and robust prediction ( 173 , 174 ). However, EHR data are often incomplete and heterogeneous, with missing values and inconsistent coding that can introduce bias and reduce generalizability; Unequal data representation across populations may further affect fairness in AI models. Therefore, the effectiveness of EHR-driven AI depends not only on algorithmic performance but critically on data quality, standardization, and fairness.
AI in ovarian cancer research shows great potential, but its clinical application is hindered by challenges like lack of transparency, interpretability, and integration of multi-modal data.
Challenges:
• Lack of transparency and interpretability of AI models.
• Heavy reliance on single-modal data, limiting the model's ability to handle tumor heterogeneity.
• Limited clinical validation and standardization across centers and technologies.
Future directions:
• Integrating multi-omics data with clinical and imaging information for improved prognostic modeling.
• Incorporating spatial biology and tumor-immune ecosystem analysis to enhance cancer diagnosis and therapy.
• Developing AI models with improved interpretability to support clinical decision-making and biological insights.
Multimodal
Multimodal AI represents a transition from precise prediction to clinically relevant decision support. Integrating diverse data types such as medical imaging, pathology results, molecular profiles, and clinical data helps achieve a more comprehensive understanding of ovarian cancer ( 106 ). Combining modalities often results in stronger predictors than relying on a single data source alone, and therefore, the development of effective multimodal AI models depends on how and when data is integrated ( 107 ).
Data fusion in multimodal AI can occur at different stages, which are generally categorized as early fusion, intermediate fusion, and late fusion. These terms refer to when different data streams are combined in the model architecture ( 108 ).
In late fusion (sometimes called decision-level fusion), each modality is processed independently (often by its own specialized model) and the outputs or high-level features from those models are then combined (fused) to make a final prediction. Late fusion models have shown performance gains in ovarian cancer compared to uni-modal models ( 109 ).
For example, Zhao et al. integrated macroscopic imaging features derived from MRI radiomics, microscopic pathology features extracted from biopsy WSIs (pathomics), and immunohistochemical biomarkers (biopsy-adapted immunoscore) into a Cox-based nomogram to predict distant metastasis in locally advanced rectal cancer. The resulting multimodal model achieved strong discriminative performance (Concordance Index (C-index) ≈0.85 in the test cohort; 5-year AUC up to 0.95 in the training cohort) and consistently outperformed models based on any single modality, illustrating the effectiveness of decision-level integration of heterogeneous clinical data ( 110 ). Similarly, the radiologic and pathologic feature layers function as omics-like data modalities. Their AI-driven integration with clinicogenomic predictors at the decision level mirrors multi-omics fusion strategies increasingly used for prognostic modeling, underscoring the role of multimodal AI systems in precision oncology decision making of HGSOC ( 111 ). Recent work introduced a foundation model-driven multimodal framework (FoMu) that integrates clinical variables, MRI data, and H&E whole-slide pathology images for prognostic prediction in HGSOC. The FoMu method used radiologic and pathologic foundation models that had already been trained to extract features and mix modality-specific and cross-modal representations. This led to accurate predictions of overall and progression-free survival across multiple center cohorts, demonstrating how AI-driven multimodal integration could enhance individual risk classification and improve therapeutic decision-making in ovarian cancer ( 112 ).
Intermediate fusion strikes a compromise: each modality is initially processed on its own (like late fusion) to extract intermediate features, but then those features are fused in a joint model layer that continues learning. For example, one might use a CNN to convert a pathology image into a 128-dimensional feature vector, use a separate model to convert genomic data into another feature vector, and then concatenate those vectors and feed them into a fully-connected neural network or transformer that produces the final prediction. The fusion happens in the middle of the pipeline, hence “intermediate.” This approach allows some pre-training or dimensionality reduction per modality (which is helpful if the dataset is small) while still allowing the model to learn interactions between modalities at the feature level ( 113 ).
Many advanced multimodal architectures fall into this category. For example, most image-based prognostic models use either histology or radiology alone, in part due to the gigapixel scale of WSIs and the large spatial mismatch between pathology and radiologic images. Hierarchical Multimodal Co-Attention Transformer (HMCAT) addresses these challenges through a weakly supervised, interpretable multimodal transformer that hierarchically encodes histologic features and integrates them with radiologic representations via co-attention, enabling feature-level fusion and improved survival prediction across multiple cancer datasets, in which modality-specific representations from histology and radiology are fused at the feature level via co-attention before final prognostic prediction ( 114 ). Autoencoder-based approaches are often categorized as intermediate fusion methods, as heterogeneous omics layers are mapped into a shared latent representation that serves as an intermediate feature space for downstream analysis. Using this strategy, Variational Autoencoder (VAE) and Maximum Mean Discrepancy Variational Autoencoder (MMD-VAE) based frameworks have been applied to ovarian cancer, where integrated mono-, di-, and tri-omics data were embedded into a common latent space for molecular subtype classification and survival analysis, achieving high subtype classification accuracy (up to ~95%) and robust prognostic stratification ( 115 ).
In early fusion, data from different modalities are combined at the input or very early in the model, creating a single joint feature space that the model learns over. For example, one could concatenate a patient’s genomic feature vector with her radiomic features before feeding them into a neural network, or even attach image pixels and genomic data as parallel channels into a unified model.
In practice, however, early fusion requires substantially more training data to tune the large number of parameters such models entail. With many modalities, the dimensionality of the joint input explodes, and the model must learn to calibrate disparate data types (image intensities vs. gene counts, for instance) which can be challenging. As a result, early fusion remains largely exploratory in ovarian cancer due to limited availability of fully matched multimodal cohorts ( 116 ).
HEALNet represents an early-fusion multimodal architecture, in which heterogeneous raw inputs from whole-slide histopathology and multi-omic data are jointly integrated at the input level to learn a shared latent representation for survival analysis ( 117 ). Another study proposed an early-fusion DL framework based on a denoising autoencoder, in which multi-omics data from TCGA were jointly embedded at the input level into a shared low-dimensional latent space. These fused representations were subsequently used for k-means clustering to define molecular subtypes, followed by construction of a lightweight L1-penalized logistic regression classifier. Using this approach, the authors identified 34 subtype-associated biomarkers and 19 enriched Kyoto Encyclopedia of Genes and Genomes (KEGG) pathways. Robustness was demonstrated through independent validation across three Gene Expression Omnibus (GEO) datasets, where the derived subtypes showed consistent predictive performance. Literature curation further confirmed that 19 of the 34 biomarkers (56%) and 8 of the 19 KEGG pathways (42.1%) had previously been implicated in ovarian cancer, supporting the biological relevance of the early-fusion multi-omics subtyping strategy ( 118 ). Another early-fusion CNN-Transformer model achieved higher performance than single CNN or Vision Transformer (ViT) models for multiclass ovarian tumor classification on transvaginal ultrasound, with an AUC of 0.990, accuracy of 92.1%, sensitivity of 92.4%, and specificity of 98.9%, supporting its suitability for clinical decision support ( 119 ).
To sum up, multimodal fusion is a key part of AI in ovarian cancer as we progress toward precision treatment. AI models can use the complementary information from radiologic, pathologic, genomic, and clinical data to capture both the “macro” tumor traits that can be seen on scans and the “micro” molecular factors that dictate behavior. Thoughtful choice of fusion strategy (early vs. late) and model design can mitigate data limitations and highlight informative cross-modal patterns. To provide a systematic methodological overview, we summarize the major paradigms of multimodal fusion strategies used in AI-driven ovarian cancer research ( Table 3 ). From a translational perspective, early fusion remains largely exploratory due to data scarcity, intermediate fusion is most suitable for biological discovery and subtype analysis, whereas late fusion currently represents the most clinically deployable paradigm in ovarian cancer ( Figure 4 ).
Overview of temporal integration and representative implementations.
Temporal multimodal integration strategies for ovarian cancer AI models. Temporal multimodal integration strategies for AI models in ovarian cancer, including: (A) Late Fusion: Processes each modality independently with separate models, and combines their outputs at the decision level to produce the final prediction. This strategy allows independent optimization of each modality-specific model before aggregation. (B) Intermediate Fusion: Each modality is processed independently to extract intermediate features, which are then combined at a later stage. This approach enables feature-level fusion, allowing the model to learn interactions between modalities after pre-processing. (C) Early Fusion: Combines data from different modalities (e.g., clinical, imaging, genomic) at the input stage, creating a joint feature space that the model learns from. This approach requires large datasets due to the high dimensionality of the combined inputs.These strategies enhance the model’s ability to handle multimodal data and improve diagnostic, prognostic, and therapeutic decision-making in ovarian cancer.
Spatial omics extends multimodal AI from integrating “different data types” to integrating “different spatial contexts”, enabling AI to model not only what genes are expressed, but where and with whom. While bulk genomics provides important averages of tumor molecular traits, spatial omics technologies add a new dimension by preserving the spatial context of gene/protein expression within tissue sections ( 121 ). In ovarian cancer, AI-spatially resolved data are illuminating how cancer cells and immune cells are organized in the tumor microenvironment ( Figure 5 ).
AI-enabled spatial analysis reveals tumor-immune ecosystems. Spatial omics data generated from multiplexed imaging and transcriptomic platforms (e.g., CosMx SMI and GeoMx) are analyzed using AI-assisted computer method. These analyses include unsupervised clustering, inference of copy number variation, robust cell-type deconvolution, and spatial clustering to delineate tumor subclones and immune cell populations. Integration of spatially resolved gene expression with tissue architecture enables the identification of distinct tumor subclones (Tumor-A and Tumor-B), immune niches (tumor-hot and tumor-cold regions), and ligand-receptor interactions within the tumor microenvironment, providing a comprehensive view of tumor-immune ecosystem organization.1. Data Acquisition: High-resolution tissue images and spatial transcriptomics data are captured to map gene expression across tissue sections. 2. Spatial Mapping: AI models map the locations of tumor and immune cells (e.g., T-cells, macrophages) within the tissue, identifying areas of immune activation or suppression. 3. Cell Type Deconvolution: AI models like Tangram and Cell2location deconvolve the data to estimate the abundance and distribution of different cell types across the tissue. 4. Gene Expression and Tumor Microenvironment Mapping: Gene expression data are integrated with tissue morphology, revealing insights into tumor progression, immune evasion, and the tumor microenvironment. 5. Tumor-Immune Interaction Mapping: AI analysis uncovers interactions between tumor and immune cells, providing insights into immune activation, suppression, and potential therapeutic targets.
AI plays a crucial role in analyzing complex spatial datasets and integrating them to map tumor-immune ecosystems and subclonal structures with greater detail. Developing AI frameworks that are both interpretable and capable of integrating data-driven and mechanistic approaches will be key to unlocking the biological insights and clinical potential of large-scale spatial tissue data ( 122 ).
Spatial omics data can be integrated with other modalities, such as single-cell multi-omics, proteomics, imaging, and clinical features, through various strategies to construct complex AI/ML models.
Pair mapping and deconvolution: By using single-cell RNA-seq data as a reference, it is mapped to spatial data. Representative methods include Tangram (DL) for aligning scRNA and spatial transcriptomics) ( 123 ), Cell2location (Bayesian model integrating scRNA cell types with spatial expression) ( 124 ), and RCTD. Neural networks or Bayesian frameworks are then used to map (deconvolve) single-cell types or expression distributions to spatial locations, outputting cell type abundance or single-cell level expression at each spatial point.
Graph Neural Network (GNN) fusion: SpatialGlue is an AI-driven GNN model that integraes spatial omics data with a dual-attention mechanism to accurately decode spatial domains and cellular properties across multiple omics modalities.
Imaging-omics integrated learning: Using high-resolution tissue images to extract morphological features, which are then combined with spatial transcriptomics data as inputs to a model. For example, DL model StLearn/SpaCell utilizes convolutional networks to extract spatial features from tissue section images and performs smoothing or clustering alongside gene expression data ( 125 , 126 ).
The outputs of spatial biology include cell-cell interactions, spatial gene expression, and tumor microenvironment maps, which help reveal molecular heterogeneity, cell interaction networks, and functional states within tissues and tumors ( 127 ), further enhancing our understanding of complex biological systems.
HGSOC are known to be spatially heterogeneous, often containing multiple distinct clones of tumor cells intermixed with stroma and immune infiltrates. In 2024, a work has demonstrated that AI-guided spatial transcriptomic analysis can help interpret prognostic regions detected by DL in HGSOC. By integrating image-based AI models trained on H&E slides with spatially resolved transcriptomics, Laury et al. showed that AI-identified tumor regions, exhibit distinct transcriptional programs compared with background tumor areas, provides the proof of concept that AI-informed spatial transcriptomics can uncover biologically and clinically relevant intratumoral heterogeneity beyond standard histopathologic assessment ( 128 ).
Then A landmark study by Denisenko et al. in 2024 applied spatial transcriptomics to HGSOC and revealed extensive intratumoural heterogeneity driven by genetically distinct tumor subclones in ovarian cancer. Using 10x Genomics Visium, this work demonstrated that multiple subclones with different copy number alteration profiles coexist within individual tumor sections and occupy distinct spatial niches. These subclones exhibited differential ligand-receptor expression patterns and preferential associations with specific stromal and immune cell populations. High-resolution validation with CosMx single-molecule imaging further confirmed subclone-specific interactions with immune cells, fibroblasts, and endothelial cells, as well as the presence of distinct autocrine signaling loops. Using an AI-enabled computational framework integrating unsupervised expression clustering, copy number variation inference (inferCNV), probabilistic cell-type deconvolution (RCTD), graph-based spatial neighborhood analysis (Squidpy), and ligand-receptor network inference (connectomeDB/NATMI), AI clustering of spatial gene expression was able to delineate these micro-domains, and suggest why certain clones might resist therapy (e.g. a clone shielded by immunosuppressive stroma) ( 129 ).
Moreover, spatially resolved transcriptomic and proteomic analyses have revealed that the progression from serous borderline tumors to low-grade serous carcinoma occurs through an epithelial micropapillary intermediate state and is supported by coordinated tumor-stroma interactions. Integrated spatial profiling identified actionable targets, with combined Cyclin-Dependent Kinase 4/6 (CDK4/6) and FOLR1 inhibition demonstrating therapeutic efficacy in vivo ( 130 ).
This advanced AI approach captures both the spatial organization and the molecular dynamics within the tumor, which are critical for improving therapeutic strategies and predicting treatment resistance. Together, these findings highlight the value of AI-enabled spatial multi-omics for understanding therapy resistance and identifying actionable therapeutic targets.
Complementary to spatial RNA, spatial proteomics (e.g. imaging mass cytometry, Cytometry by Time-of-Flight (CyTOF), multiplex immunofluorescence) can map dozens of proteins at single-cell resolution on tissue sections ( 131 , 132 ).
A 2025 study by Makhmut et al. applied deep spatial proteomics to Serous Tubal Intraepithelial Carcinomas (STICs) and matched invasive HGSOCs, quantifying over 10, 000 proteins in situ . The analysis revealed striking proteomic similarity between STICs and invasive tumors, indicating that many malignant programs are established at the precursor stage. Despite this similarity, two distinct spatial subtypes were identified, characterized by divergent tumor-immune microenvironments, including immune-enriched versus fibrotic, matrix-remodeling niches, and differences in cell-of-origin markers. AI-based clustering of spatial neighborhoods further uncovered early microenvironmental remodeling and highlighted subtype-specific vulnerabilities, such as aberrant activation of cholesterol biosynthesis pathways. Functional validation showed that inhibition of DHCR7 suppressed ovarian cancer cell growth and synergized with carboplatin, illustrating how AI-enabled spatial proteomics can uncover targetable pathways emerging from early tumor-stroma interactions ( 133 ).
Another spatial proteomic study (Yeh et al., 2024) established a large-scale single-cell spatial transcriptomic atlas of high-grade serous tubo-ovarian cancer by profiling ~2.6 million cells across 130 tumors, enabling quantitative mapping of tumor-immune spatial organization. Using AI-enabled computational analysis, including DL-based cell segmentation (Mesmer), unsupervised embedding/clustering (UMAP) to derive immune infiltration-associated cell states, and machine-learning classifiers as SVM to predict T/NK infiltration from malignant-cell programs, the authors identified a malignant transcriptional program that robustly marks the presence of TILs (MTIL) that robustly tracks spatial immune infiltration and is associated with clinical outcome. Integrating spatial cell-state mapping with perturbational transcriptomics (Perturb-seq), they further pinpointed genetic regulators (e.g., PTPN1 and ACTR8) that drive immune-evasive malignant states; genetic knockout or pharmacologic inhibition (PTPN1/PTPN2 inhibitor) increased susceptibility of ovarian cancer cells to T and NK cell cytotoxicity. Collectively, this work illustrates how AI-assisted spatial transcriptomics can connect spatially patterned tumor cell states to immune microenvironment exclusion and nominate actionable regulators of immune evasion ( 134 ).
In summary, integrating spatial omics with AI provides valuable insights into ovarian cancer heterogeneity, highlighting how cancer and immune cells interact and coexist. These approaches are uncovering new substructures (subclones, niches) in ovarian tumors that were invisible to bulk analyses, and they offer a path toward spatially-informed precision therapy , treating not just an “average” tumor, but the specific malignant ecosystems present in each patient.
Spatial AI marks a conceptual shift: from predicting outcomes to explaining why tumors behave differently within the same patient in ovarian cancer. Table 4 summarizes representative multimodal AI data fusion strategies that have been applied to ovarian cancer.
Representative AI-based multimodal data fusion strategies in ovarian cancer.
Beyond ovarian cancer-specific studies, multimodal AI frameworks developed in other cancer types transferable to ovarian cancer research ( Table 5 ). These frameworks illustrate what is technically possible for ovarian cancer.
AI-based multimodal strategies with potential translational value for ovarian cancer.
“Modalities” lists the input data each model uses. “OV Data Used?” indicates if ovarian cancer cases were included in training. The Transfer Rating (High/Med/Low) reflects how well the model is expected to work for OV, and the justification cites our sources and domain reasoning. Model type reflects the primary architectural or learning paradigm; many models are hybrid.
Conclusions
In summary, AI has demonstrated practical utility in ovarian cancer research, particularly in imaging-based triage, digital pathology, and clinical risk stratification. Integrating multimodal data has further improved precision. Collectively, these advances support the role of AI as a strong tool in ovarian cancer detection and management. Future progress will depend not on increasingly complex models, but on rigorous validation and clinically grounded design.
Image Based
Image-based AI techniques mainly facilitate early diagnosis, risk stratification, and clinical triage in ovarian cancer.
The two main categories of contemporary approaches are radiomics and DL models. Radiomics uses hand-crafted quantitative imaging features like shape and texture, whereas DL models automatically learn discriminative patterns straight from raw images ( 10 ). Ultrasound is still the best way for evaluating adnexal masses, and accumulating evidence shows that AI can make it much easier to tell the difference between benign and malignant ovarian lesions. A meta-analysis of histopathology-validated studies reported pooled sensitivity ~81% and specificity ~92% across >15, 000 ultrasound images, supporting clinical utility while underscoring the need for prospective validation ( 11 ).
These findings enhance the diagnostic efficacy of AI and emphasize the need for further clinical validation. AI performance involves a trade-off between sensitivity and specificity. High sensitivity is crucial for early detection but can increase false positives. While a high AUC suggests good performance, it doesn’t guarantee the model is ready for primary screening. Optimizing for sensitivity alone could lead to unnecessary treatments, underscoring the importance of balancing sensitivity and specificity in clinical use.
Residual Network with Feature Pyramid Network (ResNet-FPN) based Convolutional Neural Networks (CNNs) have been used for CT-based tumor classification by capturing multi-scale features relevant to lesion size, shape, and texture ( 12 ). Three-dimensional (3D) DL architectures go beyond slice-based analysis by using volumetric Positron Emission Tomography (PET) data. This makes it possible to estimate tumor load and disease stage more accurately. Models based on ResNet that are tailored for 3D input have shown better performance in diagnosing and staging, and attention-based visualization supports their clinical relevance ( 13 ). Subgroup studies demonstrated that AI models utilized in MRI attained diagnostic performance on par with ultrasound- and CT-based methodologies, highlighting the efficacy of MRI-based AI in ovarian cancer diagnosis ( 14 ). AI-enhanced ADNEX model (ADNEX-AI) was made for a multi-task DL system that automates feature extraction within the International Ovarian Tumor Analysis (IOTA) Assessment of Different NEoplasias in the adneXa (IOTA-ADNEX) ultrasonography framework that is approved by the guidelines. In a large, multicenter study, ADNEX-AI was able to tell the difference between benign and malignant ovarian masses (AUC ≈ 0.93) by automatically separating important tumor parts (lesion, locules, solid tissue, and papillary projections) and measuring their attributes ( 15 ). Importantly, the AI model’s performance was similar to that of expert-derived Assessment of Different NEoplasias in the adneXa (ADNEX) ratings, but it had better calibration and less variability between centers. These results suggest that such approaches could support ovarian mass triage in settings with limited specialist ultrasound expertise. In a multicenter study of 17, 119 images from 3, 652 patients across 20 centers in eight countries, performance matched or exceeded expert assessment and could reduce specialist referrals by up to 63% ( 16 ). These DL architectures, such as ResNet-FPN and 3D CNNs, automatically learn hierarchical spatial features, ranging from low-level texture and edge patterns to high-level structural cues, through stacked convolutional filters and multi-scale feature integration. This enables them to detect subtle lesion characteristics, such as fine texture differences and shape irregularities, that are difficult for handcrafted methods and may be overlooked by human observers ( 17 , 18 ).
Integrating imaging with clinical variables generally improves risk stratification compared with clinical data alone ( 19 ).
Pathologists diagnose ovarian endometrioid carcinoma by examining Hematoxylin and eosin (H&E) slides, where it typically appears as dense glandular clusters with back-to-back arranged glands, showing brick-like or sieve/maze-like patterns, consisting of tall columnar epithelial cells with abundant cytoplasm and oval-shaped nuclei in a pseudostratified arrangement ( 20 ). H&E staining is the gold standard for diagnosing ovarian cancer, and is the primary imaging input used in digital pathology and AI/ML models ( 21 ). These pathological features, including glandular structure, nuclear morphology, necrosis, and invasion patterns, are key diagnostic criteria for ovarian endometrioid carcinoma. Emphasizing the features pathologists focus on in H&E staining and how AI models use them can enhance clinician trust in AI predictions.
Existing studies have used DL to segment tissue/cells in whole-slide images, automatically quantifying the shape, texture, and number of cells and nuclei. These features, such as nuclear size, shape, staining intensity, glandular distribution, and density, can be used to train classifiers to distinguish tumor types or predict molecular characteristics ( 22 , 23 ). AI-driven digital pathology enables fine-grained tissue-level analysis, bridging morphological patterns with molecular and clinical outcomes. In ovarian cancer, H&E whole-slide image (WSI) provide rich morphological information for classification and prognostic modeling ( 24 ).
Using AI-based analysis of digital pathology slides improved diagnosis and prognosis assessment in gynecologic cancers ( 25 ). New imaging analysis methods have helped improve ovarian cancer detection and care, although several challenges still limit their routine use in clinical practice ( 26 ). Recent research has methodically evaluated the efficacy of five pre-trained CNNs (VGG16, VGG19, ResNet50, MobileNet, and DenseNet121) for the classification of ovarian cancer utilizing histopathology images, in comparison to traditional ML classifiers. DenseNet variants often performing best in histopathology-based classification ( 27 ).
Homologous Recombination Repair (HRR) is a high-fidelity repair pathway for double-strand DNA breaks. Dysfunction in HRR forces tumor cells to rely on error-prone repair mechanisms, leading to genomic instability and the accumulation of hallmark features, such as Loss of Heterozygosity (LOH), Telomeric Allelic Imbalance (TAI), and Large-Scale Transitions (LST) within the genome ( 28 , 29 ). Homologous recombination deficiency (HRD) is a significant clinical indicator in ovarian cancer. It most often arises from mutations in BRCA1 or BRCA2, or from alterations in other DNA repair genes, and is associated with increased sensitivity to platinum-based chemotherapy and Poly (ADP-ribose) Polymerase (PARP) inhibitor treatment ( 30 ). HRD, initially characterized by BRCA1/2 mutations, underlies the clinical efficacy of PARP inhibitors in ovarian cancer and serves as a fundamental diagnostic for broadening precision medicine beyond BRCA-mutant illness in ovarian cancer ( 31 ). Approximately 25% of ovarian cancer patients harbor germline BRCA mutations, the majority of which result in HRD, while an additional 5-7% exhibit somatic HRD ( 32 ). Data from TCGA indicate that nearly half of HGSOC display defects in homologous recombination repair ( 33 ). The HRD status is closely related to the efficacy of PARP inhibitors. HRD-positive patients receiving PARP inhibitors achieve Longer Progression-free Survival (PFS) and survival benefits.
DeepHRD and other DL models may directly determine homologous recombination failure from standard H&E-stained pathology slides and have demonstrated strong performance in separate cancer groups. AI-predicted HRD is significantly linked to better results in patients treated with platinum, underscoring histology-based AI as a scalable alternative for HRD evaluation in precision oncology ( 34 ). Patient-level DL models like HRDPath show that HRD can be reliably predicted from routine histopathology in HGSOC, making it a digital biomarker that can be used to guide therapy ( 35 ).
Yang et al. have developed graph-based DL models using H&E-stained WSIs to predict prognosis and therapeutic response in ovarian cancer. The Ovarian Cancer Digital Pathology Index (OCDPI), trained on TCGA data and validated in independent multicenter cohorts, demonstrated robust prognostic value for overall survival and recurrence risk, independent of established clinicopathological factors ( 36 ).
Overall, imaging- and pathology-based AI supports early detection, subtype characterization, and prognostic risk stratification, although standardization and external validation remain essential for clinical deployment ( Figure 2 ).
AI workflows in radiology and digital pathology for ovarian cancer. AI workflows in radiology and digital pathology used for ovarian cancer detection and prognosis. In radiology, AI techniques like radiomics and DL models analyze imaging data (ultrasound, CT, MRI) to detect, classify, and stage tumors. Tasks include lesion segmentation, feature extraction (such as shape, texture, and size), and lesion subtyping, with models such as ResNet ADNEX applied to classify ovarian masses. In digital pathology, AI models process WSI from H&E-stained tissue sections, segmenting and analyzing tissue morphology (e.g., glandular structure, nuclear morphology, necrosis). These AI workflows improve diagnostic accuracy, support risk stratification, and assist in clinical decision-making regarding therapy.
Table 1 provides an overview of the major data modalities in ovarian cancer and the AI approaches and tasks for which they have been employed.
Representative image-based AI studies for ovarian tumor detection.
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.