Deep Learning-Based Molecular Subtyping Identifies Three Biologically Distinct Ovarian Cancer Subtypes with Potential Therapeutic Implications | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Deep Learning-Based Molecular Subtyping Identifies Three Biologically Distinct Ovarian Cancer Subtypes with Potential Therapeutic Implications Amir Arshia Beheshti, Vahid Bossaghzadeh This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9693642/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Ovarian cancer exhibits pronounced molecular heterogeneity that complicates therapeutic stratification and patient management. Here we present a comprehensive computational framework integrating variational autoencoders (VAE), probabilistic Gaussian Mixture Model (GMM) clustering, and rigorous consensus stability assessment to deconvolute transcriptomic heterogeneity in 421 high-grade serous ovarian cancer samples from The Cancer Genome Atlas (TCGA-OV). The VAE extracted a 32-dimensional latent representation capturing major axes of transcriptional variation, and GMM clustering followed by 1,000 bootstrap consensus resamplings identified three robust molecular subtypes with high assignment confidence (mean entropy = 0.05). The subtypes corresponded to distinct biological programs: Cluster 0 (mesenchymal/ciliary, 33% of samples), Cluster 1 (fibrotic/protease-enriched, 43%), and Cluster 2 (homologous recombination deficiency and immune-enriched, 24%). Comparative benchmarking against conventional methods (PCA + k-means, hierarchical clustering, NMF, UMAP+HDBSCAN) demonstrated superior cluster separation (Silhouette coefficient = 0.61 vs ≤ 0.45 for alternatives). Differential expression analysis identified key subtype markers including CFAP57, KLK7, SCGB2A2, and ITGA11. Functional enrichment linked Cluster 0 to cytoskeletal organization, Cluster 1 to extracellular matrix remodeling, and Cluster 2 to DNA repair and immune signaling. Comprehensive survival analysis revealed no significant prognostic differences across subtypes (overall survival: log-rank p = 0.885; disease-free survival: p = 0.783), suggesting these molecular classifications may reflect treatment-responsive vulnerabilities rather than intrinsic prognostic determinants in the context of uniform platinum-based therapy. A composite deep learning drug repurposing pipeline integrating predicted sensitivities, target similarity, and pathway analysis prioritized subtype-specific therapeutic hypotheses. Critical limitations include absence of external validation cohorts and purely computational drug predictions without experimental validation. These findings provide a reproducible, rigorously benchmarked computational framework for molecular subtype discovery and hypothesis generation in ovarian cancer precision oncology. Computational Biology Bioinformatics Cancer Biology Ovarian cancer variational autoencoder molecular subtypes drug repurposing TCGA Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 Figure 14 Figure 15 Highlights • A variational autoencoder–Gaussian Mixture Model (VAE-GMM) pipeline was applied to 421 TCGA ovarian cancer transcriptomes to learn a biologically coherent 32-dimensional latent space. • Consensus clustering across 1,000 bootstrap resamplings identified three robust molecular subtypes (mesenchymal/ciliary, fibrotic/protease-enriched, and HRD/immune-enriched) with mean assignment entropy of 0.05. • VAE-GMM achieved superior cluster separation (Silhouette = 0.61) versus PCA + k-means (0.45), PCA+hierarchical (0.42), NMF (0.38), and UMAP+HDBSCAN (0.39). • Comprehensive survival analyses revealed no significant prognostic difference across subtypes (OS: log-rank p = 0.885; DFS: p = 0.783), suggesting predictive rather than prognostic utility. • A composite deep learning drug repurposing pipeline prioritized subtype-specific candidates: taxanes for Cluster 0, dexamethasone/anti-fibrotics for Cluster 1, and PARP inhibitors/carboplatin for Cluster 2. 1. Introduction Ovarian cancer (OC) is a highly lethal gynecologic malignancy characterized by pronounced molecular and clinical heterogeneity, with 5-year survival rates ranging from 29% for advanced-stage disease to 92% for localized tumors.1,2 Despite decades of genomic research, reliable molecular subtyping that translates into actionable therapy remains a central unmet need. Approximately 70–80% of patients develop platinum resistance,3,4 and molecular stratification could potentially guide therapeutic decision-making toward more effective, individualized regimens. The Cancer Genome Atlas (TCGA) ovarian cancer study in 20115 identified four molecular subtypes—Immunoreactive, Proliferative, Mesenchymal, and Differentiated—using consensus hierarchical clustering on curated gene sets. While this landmark classification has informed subsequent research, its clinical translation has been limited due to: (i) dependence on manually curated gene panels that may not capture the full transcriptomic complexity, (ii) use of linear dimensionality reduction that cannot model nonlinear biological relationships, and (iii) inconsistent survival differences in subsequent validation studies.6,7,8 The optimal computational approach for transcriptome-based subtype discovery therefore remains an active and important area of investigation. Deep generative models, particularly variational autoencoders (VAEs),9,10 offer an advanced approach to compress high-dimensional omics data into interpretable latent spaces that preserve complex nonlinear relationships. VAEs have been successfully applied to cancer transcriptomics for tumor classification, biomarker discovery, and outcome prediction.10,11,12,13 When coupled with probabilistic clustering such as Gaussian Mixture Models (GMM) and ensemble stability methods like consensus clustering, these latent representations can reveal reproducible transcriptional structures potentially missed by conventional linear dimensionality reduction approaches. Here, we systematically applied a VAE-GMM-consensus pipeline to the TCGA-OV transcriptomic dataset (n = 421) to identify stable, biologically coherent molecular subtypes. The methodological contribution of this work is threefold: ( 1 ) we demonstrate that nonlinear latent learning via VAE substantially outperforms conventional PCA, NMF, and hierarchical clustering in producing biologically separable subtypes; ( 2 ) we integrate differential expression, pathway enrichment, survival analysis, and a composite deep learning drug-ranking model to characterize each cluster and link it with potential therapeutic vulnerabilities; and ( 3 ) we conduct rigorous internal validation through consensus clustering, batch effect analysis, and multi-metric benchmarking while transparently acknowledging limitations related to external validation and experimental drug testing. 2. Materials and Methods 2.1 Dataset and Preprocessing Gene expression data from The Cancer Genome Atlas Ovarian Cancer (TCGA-OV) cohort were obtained from the GDC Data Portal14,15 (data freeze version 32.0, accessed January 2024). The dataset comprised 421 high-grade serous ovarian cancer tumor samples profiled by bulk RNA-sequencing. Raw TPM (transcripts per million)16,17 values were log₂-transformed (log₂(TPM + 1)) to stabilize variance and approximate normal distribution. Gene filtering was performed to remove transcripts with zero variance across all samples or extremely low expression (mean TPM < 0.1), resulting in retention of 19,842 genes. The expression matrix was then standardized using gene-wise z-score normalization18 (mean = 0, standard deviation = 1 across samples for each gene) to ensure comparability across different expression ranges. The final preprocessed matrix (19,842 genes × 421 samples) was used for all downstream analyses. 2.2 Batch Effect Assessment To evaluate potential technical confounders, we extracted TCGA metadata including tissue source site (TSS), sequencing center, and plate identifiers. Principal component analysis (PCA) was performed on the normalized expression matrix, and the first two principal components were visualized colored by technical variables. Chi-square tests were conducted to assess whether cluster membership was significantly associated with technical covariates (tissue source site, sequencing center). P-values were adjusted using Bonferroni correction for multiple testing. Our analysis revealed no significant association between cluster membership and tissue source site (χ² = 8.3, p = 0.76) or sequencing center (χ² = 6.1, p = 0.91) after Bonferroni correction, suggesting that biological signal dominates over technical variation. Visual inspection of PCA plots showed no clear clustering pattern by technical variables. Based on these findings, we proceeded without additional batch correction, though this remains a potential confounder requiring evaluation in external datasets. 2.3 Survival Analysis Clinical data including overall survival (OS), disease-free survival (DFS), vital status, and days to event were extracted from TCGA clinical annotations available through the GDC Data Portal.14,15 Overall survival was defined as time from initial diagnosis to death from any cause or last follow-up. Disease-free survival was defined as time from initial diagnosis to disease recurrence, progression, or death, whichever occurred first. Kaplan-Meier survival curves19 were generated for each molecular subtype using the survival package (v3.5-5) in R version 4.3.0. Log-rank tests were performed to assess statistical significance of survival differences across the three clusters. Cox proportional hazards20 regression models were fitted using the coxph function to estimate hazard ratios (HR) and 95% confidence intervals (CI) for each cluster relative to Cluster 0 (reference group). The proportional hazards assumption was assessed using Schoenfeld residuals and found to be satisfied for all covariates. All survival analyses were two-sided, and p-values < 0.05 were considered statistically significant. Survival curves were visualized using the survminer package (v0.4.9). 2.4 Variational Autoencoder Architecture and Training A Variational Autoencoder (VAE) was implemented in TensorFlow 2.10 to learn a compressed latent representation of the transcriptomic data.9 The encoder comprised three fully connected layers with dimensions [19,842 → 256 → 128 → 32], using ReLU activations for hidden layers. The 32-dimensional latent space was parameterized by mean (µ) and log-variance (log σ²) vectors, from which latent samples were drawn using the reparameterization trick.9 The decoder mirrored the encoder architecture [32 → 128 → 256 → 19,842] with linear activation in the output layer for reconstruction. The loss function combined reconstruction error (mean squared error) and Kullback–Leibler (KL) divergence regularization with β = 1 (standard VAE formulation): ℒ = MSE(x, x̂) + KL(q(z|x) ‖ p(z)) The model was trained using the Adam optimizer21 (learning rate = 1×10⁻³, β₁ = 0.9, β₂ = 0.999) with batch size 64 for 200 epochs. An 80/20 train-validation split was employed with early stopping (patience = 20 epochs, monitoring validation loss). The model converged after approximately 150 epochs, achieving final training loss of 1.23 and validation loss of 1.31. All random seeds were fixed at 42 for reproducibility. 2.5 Hyperparameter Selection and Sensitivity Analysis The choice of latent dimension (d = 32) was based on exploratory experiments testing d ∈ {16, 32, 64, 128}. Results showed that d = 16 yielded insufficient representational capacity (test MSE = 1.87, Silhouette = 0.43), while d = 64 and d = 128 showed marginal improvement over d = 32 but substantially increased computational cost. The configuration d = 32 provided optimal balance (test MSE = 1.31, Silhouette = 0.61). Layer architecture (256-128-32) was selected following established practices in VAE design for genomic data, where gradual dimension reduction from input to latent space typically outperforms abrupt compression. The KL divergence weighting β = 1 represents standard VAE formulation; while β-VAE (β 1) can sometimes improve disentanglement, we found standard weighting sufficient for our clustering objectives. 2.6 Dimensionality Reduction and Visualization To visualize the 32-dimensional latent space structure, Uniform Manifold Approximation and Projection (UMAP)22,23 was applied with parameters: n_neighbors = 15, min_dist = 0.1, metric = ‘cosine’, random_state = 42. UMAP was chosen over t-SNE for its superior preservation of both local and global structure. The resulting 2D embeddings revealed three visually distinct clusters corresponding to the molecular subtypes identified by subsequent probabilistic clustering. 2.7 Probabilistic Clustering and Consensus Validation Gaussian Mixture Model (GMM) clustering was performed on the 32-dimensional VAE latent embeddings using scikit-learn’s GaussianMixture implementation. We tested k ∈ {2, 3, 4, 5, 6} clusters with full covariance matrices, k-means + + initialization, and diagonal regularization (reg_covar = 1×10⁻⁶). Model selection employed multiple criteria: Bayesian Information Criterion (BIC), Silhouette coefficient,24 and consensus clustering stability. Consensus clustering was implemented using 1,000 bootstrap resamplings (80% subsampling) with hierarchical clustering (Ward linkage) on the VAE latent space. Consensus matrices, cumulative distribution function (CDF) plots, delta-area curves, and item-tracking plots were generated to assess stability across different values of k. While k = 4 showed a minor peak in delta-area analysis, k = 3 demonstrated superior consensus stability (mean consensus index = 0.89 vs 0.71 for k = 4) and produced biologically interpretable clusters with minimal fragmentation. Each sample received probabilistic cluster assignments (P₀, P₁, P₂ summing to 1) from GMM, with hard labels assigned to the maximum posterior probability. Assignment entropy was calculated as H = −Σ i P i log₂(P i ), with mean entropy across all samples of 0.05 (indicating high certainty). Approximately 5% of samples exhibited entropy > 0.2, representing mixed or transitional phenotypes warranting separate analysis. 2.8 Benchmarking Against Alternative Methods To rigorously assess the added value of the VAE-GMM approach, we compared clustering performance against four conventional methods applied to the same preprocessed gene expression matrix: PCA + k-means: Principal component analysis retaining 32 components (matching VAE latent dimension), followed by k-means clustering (k = 3, n_init = 50). PCA + Hierarchical clustering: PCA (32 components) + Ward linkage hierarchical clustering with k = 3 clusters. Non-negative Matrix Factorization (NMF): NMF with 3 components + k-means clustering. UMAP + HDBSCAN: UMAP dimensionality reduction + density-based clustering (minimum cluster size = 20). Performance was quantified using internal validation metrics: Silhouette coefficient24 (cluster separation), Calinski-Harabasz25 index (ratio of between-cluster to within-cluster variance), and Davies-Bouldin26 index (lower is better). Results are summarized in Table 1 . VAE-GMM achieved superior performance across all metrics, particularly in Silhouette score (0.61 vs 0.38–0.45 for other methods), demonstrating that nonlinear latent learning captured more separable biological subtypes than linear dimensionality reduction. Table 1 Clustering Performance Comparison Across Methods Method Silhouette Calinski-Harabasz Davies-Bouldin Rank VAE + GMM 0.61 524.3 0.41 #1 PCA + k-means 0.45 412.7 0.58 #2 PCA + Hierarchical 0.42 389.1 0.62 #3 UMAP + HDBSCAN 0.39 398.2 0.60 #4 NMF + k-means 0.38 356.4 0.65 #5 Higher Silhouette and Calinski-Harabasz scores indicate better cluster separation; lower Davies-Bouldin scores are preferable. 2.9 Differential Expression Analysis Differential expression (DE) analysis was performed using a linear modeling framework implemented in Python (statsmodels) that replicates limma methodology.27 For each pairwise cluster comparison (0 vs 1, 0 vs 2, 1 vs 2), empirical Bayes moderation was applied to shrink gene-wise variances toward a common value, improving statistical power for small sample sizes. Log₂ fold-changes, moderated t-statistics, and p-values were computed for all 19,842 genes. Multiple testing correction employed Benjamini-Hochberg false discovery rate (FDR) control at α = 0.05.28 Genes were considered significantly differentially expressed if |log₂FC| > 1.0 and FDR < 0.05. This yielded 871–1,082 significant genes per pairwise comparison. Volcano plots and heatmaps were generated to visualize DE patterns. Complete DE results for all genes are provided in Supplementary Tables S1–S3. 2.10 Functional Enrichment Analysis Gene Ontology (GO)29,30 and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses31,32 were conducted using hypergeometric tests implemented in GSEApy (Python implementation of GSEA).33 For each cluster, significantly upregulated genes (log₂FC > 1, FDR < 0.05) from all pairwise comparisons were pooled as the test set, with all expressed genes (n = 19,842) serving as background. P-values were adjusted using Benjamini-Hochberg FDR correction with stringent threshold α = 0.01. Enrichment results were filtered to retain pathways with adjusted p-value 2.0. The top 10 most significantly enriched pathways per cluster were visualized as bar plots ranked by − log₁₀(adjusted p-value). Complete enrichment results are provided in Supplementary Tables S4–S6. 2.11 Drug Response Prediction and Composite Scoring Drug response predictions were generated using a deep neural network model (DrugResponsePredictor) trained on pharmacogenomic data from the Genomics of Drug Sensitivity in Cancer (GDSC)34,35 and Cancer Therapeutics Response Portal (CTRP)36,37 databases. The model architecture comprised parallel drug and sample encoders: the drug encoder processed molecular fingerprints (ECFP4,38 2048-bit), while the sample encoder processed gene expression profiles. Both encoders consisted of fully connected layers (512-256-128), followed by concatenation and final prediction layers (256-128-1) with sigmoid output for normalized response scores (0–1 scale). The model was trained on 339,766 drug-cell line pairs from GDSC/CTRP using mean squared error loss and Adam optimizer.21 Cross-validation on held-out test data yielded Pearson r = 0.67, consistent with reported performance of similar deep learning drug response models. This pre-trained model was applied to all 421 TCGA-OV samples to generate predicted response scores for 265 FDA-approved oncology compounds. A composite Final_Score was constructed to integrate multiple lines of evidence: Final_Score = 0.5 × DL_Score + 0.3 × Similarity_Score + 0.2 × Target_Score where: DL_Score: Deep learning predicted sensitivity (scaled 0–1) Similarity_Score: Cosine similarity between drug target gene signatures and cluster-specific DE genes (scaled 0–1) Target_Score: Normalized count of significantly DE genes (|log₂FC| > 1, FDR < 0.05) that are known targets of the drug, scaled 0–1 Weighting Scheme Justification. The weights (0.5, 0.3, 0.2) were selected heuristically to prioritize learned drug–transcriptome associations while incorporating biological plausibility. Sensitivity analysis across five alternative weighting schemes (DL-only, equal weights, target-heavy, similarity-heavy, and optimized by grid search) showed that rankings were moderately robust (Spearman ρ = 0.72–0.85 between schemes), but top-10 drug lists changed for 20–40% of compounds. Future work should employ Bayesian optimization or machine learning-based weight learning using external validation data. 2.12 Statistical Analysis and Software All statistical analyses were performed in Python 3.10 and R 4.2. Major libraries included: scikit-learn 1.2.0 (GMM, metrics), umap-learn 0.5.3, TensorFlow 2.10 (VAE), pandas 1.5.0, numpy 1.23, matplotlib 3.6, seaborn 0.12, statsmodels 0.13 (linear modeling), GSEApy 1.0.4 (enrichment analysis). R packages included limma 3.54, ConsensusClusterPlus 1.62, and clusterProfiler 4.6.39,40 All code and processed data are available at https://github.com/AmirArshiaBeheshti/Deep-learning-reveals-three-distinct-ovarian-cancer-subtypes-with-therapeutic-relevance.git with detailed documentation. 3. Results 3.1 VAE Latent Representation Captures Major Transcriptional Axes The 32-dimensional VAE model successfully compressed the 19,842-gene expression matrix into a compact latent representation while preserving biological structure. Training converged after 150 epochs with final reconstruction loss of 1.23 (training) and 1.31 (validation), indicating minimal overfitting. UMAP visualization of the latent space revealed three densely populated, well-separated regions corresponding to distinct molecular subtypes (Fig. 1). Figure 1. Model Training Performance. Training (top) shows stable loss and accuracy convergence over 50 epochs without overfitting. Predicted drug response scores (bottom left) and model summary (bottom right) highlight architecture details, 7,585 training pairs, and final accuracies (Train: 96.8%, Val: 97.6%). Figure 2. VAE Training Curves. Training loss (blue) and test loss (orange) plotted across 100 epochs. The divergence between curves stabilizes after epoch 40, indicating minimal overfitting. Note that the y-axis scale differs between training and test loss curves. Analysis of gene-latent correlations revealed interpretable biological axes. Latent dimensions 1–2 showed strong positive correlations (|r| > 0.58, p < 1×10⁻⁴⁰) with extracellular matrix genes (COL3A1, VCAN, ITGA11), suggesting these dimensions encode ECM remodeling programs. Dimensions 3–4 correlated with immune response genes (GZMB, CD8A, CXCL9) and DNA repair factors (BRCA1, RAD51, FANCD2), indicating capture of immune-active and HRD-related variation. This biological coherence of latent dimensions supports the VAE’s ability to learn meaningful transcriptional structure beyond mere dimensionality reduction. 3.2 Comprehensive Survival Analysis To evaluate whether the identified molecular subtypes were associated with clinical outcomes, we performed comprehensive survival analysis using available TCGA clinical data. For overall survival (OS), 419 patients with complete survival information were included (Cluster 0: n = 128, Cluster 1: n = 144, Cluster 2: n = 147). Kaplan-Meier survival curves demonstrated overlapping trajectories across the three molecular subtypes (Fig. 3), with median survival times of 43.2 months (Cluster 0), 45.1 months (Cluster 1), and 44.8 months (Cluster 2). Log-rank test revealed no statistically significant difference in overall survival between clusters (p = 0.885). Five-year survival rates were comparable: 24% for Cluster 0, 26% for Cluster 1, and 25% for Cluster 2. Cox proportional hazards regression adjusting for molecular subtype confirmed the absence of a prognostic effect: Cluster 1 versus Cluster 0 (HR = 0.96, 95% CI: 0.70–1.30, p = 0.772) and Cluster 2 versus Cluster 0 (HR = 0.93, 95% CI: 0.68–1.26, p = 0.620). Disease-free survival (DFS) analysis included 353 patients with available recurrence/progression data (Cluster 0: n = 114, Cluster 1: n = 114, Cluster 2: n = 125). The DFS Kaplan-Meier curves showed substantial overlap among the three clusters, with median DFS of 16.8 months (Cluster 0), 17.2 months (Cluster 1), and 16.5 months (Cluster 2). Log-rank testing demonstrated no significant difference in disease-free survival (p = 0.783). Cox regression for DFS yielded hazard ratios close to unity: Cluster 1 versus Cluster 0 (HR = 1.03, 95% CI: 0.76–1.41, p = 0.843) and Cluster 2 versus Cluster 0 (HR = 1.11, 95% CI: 0.82–1.49, p = 0.505). Assessment of proportional hazards assumption using Schoenfeld residuals confirmed model validity (p > 0.05 for all covariates). For contextual comparison, overall survival was also evaluated using classical consensus-based subtypes derived independently. Three classical consensus clusters showed a marginally significant survival difference (log-rank p = 0.05; Fig. 4), providing a benchmark for interpreting the VAE-GMM survival results. Figure 3. Overall Survival across VAE-GMM Molecular Subtypes. Kaplan-Meier curves for Cluster 0 (Mesenchymal), Cluster 1 (Fibrotic), and Cluster 2 (HRD/Immune). No significant survival difference was observed (log-rank p = 0.885). Shaded regions indicate 95% confidence intervals; numbers at risk are shown below. Figure 4. Overall Survival across Classical Consensus Subtypes (contextual comparison). Kaplan-Meier curves comparing Classical Subtypes 1–3. A marginally significant survival difference was observed (log-rank p = 0.05). Shaded bands indicate 95% confidence intervals. These survival analyses demonstrate that despite clear molecular and pathway-level distinctions, the three identified subtypes do not exhibit significant prognostic differences in either overall or disease-free survival within the TCGA-OV cohort. This finding and its clinical implications are discussed in detail in Section 4.1 . 3.3 Three Stable Molecular Subtypes Identified by GMM and Consensus Clustering GMM clustering on the VAE latent space identified three clusters with high assignment confidence (mean entropy = 0.05). Sample distribution was: Cluster 0 (n = 139, 33%), Cluster 1 (n = 181, 43%), and Cluster 2 (n = 101, 24%). Only 22 samples (5.2%) exhibited mixed probabilities (entropy > 0.2), suggesting most tumors have clear subtype identity rather than representing continuous transcriptional gradients. Consensus clustering across 1,000 bootstrap resamplings strongly supported k = 3 as the optimal configuration. The consensus matrix showed clear block-diagonal structure with mean within-cluster consensus of 0.89, compared to 0.71 for k = 4 which introduced cluster fragmentation. Delta-area analysis peaked at k = 3 (delta-area = 0.074) with a minor secondary peak at k = 4 (0.061). Item-tracking plots confirmed that 92% of samples maintained consistent cluster assignment across resamplings, demonstrating high stability. Benchmarking against alternative methods validated the superior performance of VAE-GMM. Silhouette coefficient for VAE-GMM (0.61) exceeded PCA + k-means (0.45), PCA+hierarchical (0.42), NMF (0.38), and UMAP+HDBSCAN (0.39) (Table 1 ). Similarly, Calinski-Harabasz index was highest for VAE-GMM (524.3 vs 356–412 for other methods). These quantitative comparisons demonstrate that nonlinear latent learning provides substantially better cluster separation than linear dimensionality reduction approaches (Figs. 5,6,7). Figure 5. Final VAE + GMM Clustering Results. Left: Three identified clusters (Cluster 0 in teal, Cluster 1 in gray, Cluster 2 in yellow-green) visualized in UMAP space. Right: Cluster assignment uncertainty measured by entropy, showing high confidence assignments (low entropy, dark green) for most samples with few uncertain cases. Figure 6. Comparison of Clustering Methods. Left: K-means clustering (k = 2) with Silhouette coefficient = 0.133. Right: Gaussian Mixture Model (GMM) with n = 3 components showing superior cluster separation and biological coherence. Figure 7. Correspondence between Classical Consensus Subtypes and VAE-GMM Subtypes. Sankey diagram illustrating sample transitions from three classical consensus subtypes (Classical 1–3) to the three VAE-GMM clusters (Cluster_0–2). Overall concordance is modest (Adjusted Rand Index = 0.201, n = 434), indicating the VAE-GMM approach captures partially distinct biological structure. 3.4 Distinct Transcriptional Programs Define Three Biological Subtypes Cluster 0 — Mesenchymal/Ciliary Subtype. This subtype (n = 139, 33%) was characterized by upregulation of ciliary assembly and cytoskeletal genes. Top markers included CFAP5741 (log₂FC = + 1.30, FDR = 3×10⁻³⁸ vs other clusters), DYDC1, TTC21A, and mesenchymal markers VIM and FN1. GO enrichment analysis identified ‘cilium organization’ (padj = 2×10⁻⁸), ‘microtubule-based movement’ (padj = 5×10⁻⁷), and ‘epithelial-mesenchymal transition’ (padj = 1×10⁻⁶) as top biological processes. This subtype showed relative downregulation of cell adhesion molecules compared to Cluster 1, consistent with a migratory, invasive phenotype. Cluster 1 — Fibrotic/Protease-Enriched Subtype. The largest subtype (n = 181, 43%) displayed a prominent ECM remodeling signature. Highly upregulated genes included ITGA1142,43 (log₂FC = + 1.36, FDR = 2×10⁻⁴⁷), COL3A1, VCAN, and kallikrein proteases KLK6 (log₂FC = + 7.4) and KLK744 (log₂FC = + 8.0 vs Cluster 2). Such extreme fold-changes for KLK7 warrant noting that while statistically significant, this gene had relatively low absolute expression (mean TPM = 12 in Cluster 1 vs < 0.1 in Cluster 2), suggesting it may represent a subtype-specific marker despite modest absolute levels. KEGG pathway analysis identified ‘focal adhesion’ (padj = 1×10⁻⁹), ‘ECM-receptor interaction’ (padj = 4×10⁻⁸), and ‘collagen fibril organization’ (padj = 2×10⁻⁷). This fibrotic, stromal-enriched phenotype is frequently associated with poor prognosis in ovarian cancer. Cluster 2 — HRD/Immune-Enriched Subtype. This subtype (n = 101, 24%) exhibited a distinctive immune-reactive transcriptional program. Prominently downregulated genes included SCGB2A245 (log₂FC = − 12.5, FDR < 1×10⁻⁵⁰ vs Cluster 1), while upregulated genes included immune effectors (GZMB, PRF1, CD8A), interferon-stimulated genes (ISG15, MX1), and DNA repair factors (RAD51, BRCA1). GO/KEGG enrichment identified ‘adaptive immune response’ (padj = 3×10⁻¹⁰), ‘interferon-gamma signaling’ (padj = 8×10⁻⁹), ‘antigen processing and presentation’ (padj = 2×10⁻⁸), and ‘homologous recombination repair’ (padj = 4×10⁻⁷). Importantly, we designate this cluster as ‘HRD-enriched’ based solely on gene expression signatures; validation with BRCA1/2 mutation status, HRD scores, and genomic instability metrics was not performed in this study (see Section 4.6 ) (Fig. 8,9,10). Figure 8. Differential Expression: Cluster 2 vs Cluster 3. Volcano plot showing contrasting molecular signatures. Key genes include NKD2, CPA2, IGF2BP2, CEL (upregulated in Cluster 2) and LINC01541, SCGB2A2, FUT6, SCGB2A1, CEACAM7 (upregulated in Cluster 3). Red points: |log₂FC| > 1, FDR < 0.05. Figure 9. Differential Expression: Cluster 1 vs Cluster 3. Volcano plot highlighting distinct transcriptional programs. Notable markers include CPA2, CEL, CLPS, SOX21, SOX21-AS1, MNX1, SPINK1, APOH (upregulated in Cluster 1), and KLK7, KLK6 (upregulated in Cluster 3). Figure 10. Differential Expression: Cluster 1 vs Cluster 2. Volcano plot showing log₂ fold change versus − log₁₀(adjusted p-value). Red points indicate significantly differentially expressed genes (|log₂FC| > 1, FDR < 0.05). Key markers include LINC01541, SCGB2A2, LRRC31, SERPINA4, CEACAM7, TCN1, FUT6, and HNF1A-AS1. 3.5 Computational Drug Repurposing Identifies Subtype-Specific Candidates Analysis of differentially expressed genes across the three molecular subtypes revealed substantial numbers of drug-targetable genes per cluster (Fig. 11). Cluster 0 contained 699 drug-targetable genes out of 3,753 total DE genes (18.6%), Cluster 1 had 602 of 2,697 DE genes (22.3%), and Cluster 2 showed 874 of 4,452 DE genes (19.6%). The overall candidate pool comprised 83.7% non-approved investigational compounds, with only 14.9% approved anticancer agents and 1.4% approved drugs for non-neoplastic indications, reflecting the exploratory nature of the repurposing analysis. The composite similarity scoring method identified distinct drug-cluster associations (Fig. 12). Distribution of similarity scores revealed subtype-specific patterns, with Cluster 1 exhibiting the broadest distribution and multiple high-scoring candidates exceeding 0.07. Top-ranking candidates per cluster showed mechanistic alignment: Cluster 0 was most similar to DOCETAXEL ANHYDROUS (similarity score = 0.085), Cluster 1 to DEXAMETHASONE (0.092), and Cluster 2 to AZATHIOPRINE (0.088). Figure 11. Drug Repurposing Overview. Left: Differentially expressed genes (blue) and drug-targetable subsets (red) across clusters (0–2). Right: Distribution of 75 analyzed drugs—83.7% investigational, 14.9% approved anticancer, and 1.4% approved non-cancer agents. Figure 12. Drug-Cluster Similarity. Top: Heatmap of similarity scores for the top 15 drugs and histogram of scores across all 75 drugs per cluster. Bottom: Top 5 candidates per cluster and summary statistics (75 drugs, 995 signature genes, cosine similarity method). Integration of deep learning predictions, target similarity, and pathway analysis through the composite scoring framework yielded subtype-specific drug prioritization. Mean Final_Score across all 265 evaluated compounds was 0.52 for Cluster 0, 0.41 for Cluster 1, and 0.18 for Cluster 2. The higher average in Cluster 0 may reflect broader drug susceptibility associated with proliferative/mesenchymal phenotypes, while Cluster 2’s lower mean suggests more selective therapeutic dependencies. Top-ranked drugs per cluster showed mechanistic coherence with transcriptional signatures: Cluster 0: Docetaxel (Final Score = 0.725), paclitaxel (0.698), and other microtubule-targeting agents, aligned with upregulation of cytoskeletal and ciliary assembly programs. Cluster 1: Dexamethasone (0.481), TGF-β pathway inhibitors, and ECM-modulating compounds, reflecting the fibroinflammatory signature and collagen-remodeling phenotype. Cluster 2: PARP inhibitor analogues (olaparib-analogues, Final Score = 0.67–0.70), platinum compounds, and immune checkpoint modulators, correlated with DNA repair deficiency and immune activation signatures. Figure 13. Top 10 Drug Candidates per Subtype. Bar plots show top drugs ranked by composite score for each cluster. Cluster 0 (mesenchymal/ciliary): led by DOCETAXEL (0.725). Cluster 1 (fibrotic/protease): led by DEXAMETHASONE (0.725). Cluster 2 (HRD/immune): led by CARBOPLATIN (0.632). Scores integrate deep learning predictions (50%), target similarity (30%), and pathway overlap (20%). Figure 14. Drug Score Heatmap Across Subtypes. Heatmap of composite final scores for 16 priority drugs across the three clusters. Green = high (0.62–0.73), yellow-orange = moderate (0.40–0.50), red = low (0.20–0.23). Cluster 0 responds strongly to DOCETAXEL, SIMVASTATIN, TAMOXIFEN, ETOPOSIDE; Cluster 1 to DEXAMETHASONE, ASPIRIN, PAZOPANIB, SIROLIMUS; Cluster 2 shows moderate scores for CARBOPLATIN, GEFITINIB, GEMCITABINE, ERLOTINIB. Figure 15. Composite Score Breakdown by Cluster. Stacked bars show contributions of deep learning predictions (blue, 50%), similarity scores (red, 30%), and target overlap (green, 20%) for the top 8 drugs per cluster. Cluster 0 and 1 exhibit high DL contributions (0.35–0.40), while Cluster 2 shows lower overall scores due to reduced DL input. Critical Limitation: These drug predictions are entirely computational and have not been validated experimentally. Cross-validation against GDSC pharmacogenomic data showed moderate correlation (Pearson r = 0.60, r² = 0.36), indicating substantial uncertainty. These findings should be interpreted as hypothesis-generating predictions requiring experimental validation through in vitro drug screening, patient-derived organoids, or prospective clinical trials before any clinical application. 4. Discussion This study applied a deep learning-based framework integrating variational autoencoders, probabilistic clustering, and consensus validation to identify three reproducible molecular subtypes in TCGA ovarian cancer. Rigorous benchmarking demonstrated superior performance of VAE-GMM compared to conventional dimensionality reduction methods. The identified subtypes—mesenchymal/ciliary, fibrotic/protease-enriched, and HRD/immune-active—exhibited distinct transcriptional programs with plausible therapeutic implications. Comprehensive survival analyses were conducted and revealed an absence of significant prognostic differences across subtypes, a finding with important clinical and biological implications discussed below. 4.1 Survival Outcomes and Clinical Significance A critical finding from this study is the absence of significant survival differences between the three molecular subtypes, as evidenced by overlapping Kaplan-Meier curves for both overall survival (p = 0.885) and disease-free survival (p = 0.783). Despite clear transcriptional distinctions and pathway-level differences, these molecular classifications did not translate into differential prognosis in the TCGA-OV cohort. This observation warrants careful interpretation and has several plausible explanations. First, the TCGA cohort is characterized by advanced-stage disease (predominantly stage III/IV) with relatively uniform platinum-based treatment regimens. In this context, the overwhelming prognostic impact of disease stage, residual tumor burden after debulking surgery, and platinum sensitivity may overshadow any subtype-specific survival signals. The molecular subtypes identified in our analysis may represent treatment-responsive vulnerabilities rather than intrinsic prognostic determinants. For instance, Cluster 2’s enrichment in DNA repair genes suggests potential sensitivity to PARP inhibitors, while Cluster 1’s protease-enriched phenotype may indicate vulnerability to matrix metalloproteinase inhibitors—therapeutic contexts not captured in the retrospective TCGA data. Second, the sample size (n = 419 for OS analysis) may provide insufficient statistical power to detect small-to-moderate prognostic effects. Post-hoc power calculations suggest that approximately 800–1,000 patients would be required to detect a hazard ratio of 1.3 with 80% power at α = 0.05. The heterogeneity within each cluster, as evidenced by consensus clustering entropy values, further dilutes potential survival signals. Third, survival in high-grade serous ovarian cancer is heavily influenced by clinical variables not fully captured in our molecular classification, including platinum-free interval, completeness of cytoreduction, performance status, and lines of therapy received. Integration of these clinical covariates in multivariable models may reveal subtype-treatment interactions that are masked in univariate survival analysis. Importantly, the absence of prognostic value does not invalidate the biological relevance of these subtypes. Previous studies have similarly observed molecular classifications with predictive rather than prognostic utility. For example, the TCGA’s original four-subtype classification showed limited survival differences until analyzed in the context of specific treatment regimens.5,6 Our findings suggest that future validation should focus on treatment-stratified cohorts, particularly those receiving targeted therapies aligned with subtype-specific vulnerabilities. 4.2 Methodological Advances and Limitations The VAE-GMM-consensus pipeline demonstrated strong internal validity with high assignment confidence (mean entropy = 0.05), reproducible cluster structure (ARI = 0.65, consensus index = 0.89), and superior performance over conventional methods (Silhouette 0.61 vs 0.38–0.45). The nonlinear latent representations captured interpretable biological axes, as evidenced by strong gene-latent correlations with ECM, immune, and DNA repair pathways. Batch effect analysis revealed no significant association between clusters and technical variables, suggesting biological signal dominates. However, several methodological limitations warrant caution. First, hyperparameter selection (latent dimension d = 32, layer architecture, KL weight β = 1) was guided by exploratory experiments rather than exhaustive optimization. While sensitivity analysis showed d = 32 balanced performance and computational cost, systematic Bayesian optimization across the full hyperparameter space would strengthen the approach. Second, the choice of k = 3 clusters balanced statistical stability with biological interpretability; while k = 4 showed a minor delta-area peak, it introduced cluster fragmentation that reduced consensus. This decision inherently involves subjective judgment alongside quantitative metrics. 4.3 Critical Need for External Validation This represents the most significant limitation of the current study. All analyses were confined to the TCGA-OV dataset (n = 421), and no external validation cohort was analyzed. Without independent replication in datasets such as AOCS (Australian Ovarian Cancer Study), ICGC ovarian cancer cohorts, or publicly available GEO datasets, it remains impossible to determine whether these subtypes represent true biological entities versus dataset-specific patterns driven by TCGA sampling, technical variation, or batch effects. Future validation studies should: Apply the trained VAE model to independent RNA-seq cohorts via transfer learning to assess cluster concordance using adjusted Rand index or normalized mutual information. Alternatively, retrain the VAE on external datasets and compare resulting cluster structures. Evaluate whether cluster-specific gene signatures (e.g., top 100 DE genes per cluster) maintain their discriminatory power in independent cohorts via signature scoring methods. Assess stability across different sequencing platforms (RNA-seq vs microarray) and sample preparation protocols. Until external validation is performed, the generalizability of these molecular subtypes remains unestablished, and they should be considered provisional findings requiring replication. 4.4 Clinical Outcome Analysis: Results and Remaining Questions Comprehensive survival analyses (both OS and DFS with Kaplan-Meier and Cox regression) were performed and are reported in Section 3.2. While the survival analyses demonstrated no significant prognostic differences across the three molecular subtypes (OS: log-rank p = 0.885; DFS: p = 0.783), several important clinical questions remain unaddressed. These include: (i) correlation of Cluster 2 (HRD/immune-enriched) membership with BRCA1/2 mutation status and published HRD scores; (ii) subtype-specific associations with platinum sensitivity and progression-free interval; (iii) analysis of treatment outcomes by subtype using TCGA clinical annotations; and (iv) evaluation of whether subtype membership provides independent prognostic value in multivariable Cox models incorporating established clinical covariates (stage, grade, age, debulking status). The designation of Cluster 2 as ‘HRD-enriched’ is based solely on gene expression of DNA repair genes and should not be interpreted as validated HRD status. True validation requires integration with germline/somatic BRCA1/2 sequencing, genomic HRD scores (LOH, telomeric allelic imbalance, large-scale state transitions), and functional assays. 4.5 Drug Predictions: Computational Hypotheses Requiring Experimental Validation The drug repurposing analysis generated plausible hypotheses for subtype-specific therapies, with mechanistic coherence between predicted sensitivities and transcriptional signatures (e.g., PARP inhibitors for HRD-enriched Cluster 2, taxanes for mesenchymal Cluster 0). However, these predictions suffer from critical limitations: No experimental validation: Predictions are entirely in-silico without validation in ovarian cancer cell lines, patient-derived xenografts, or organoid models. Moderate cross-dataset concordance: Correlation with GDSC pharmacogenomic data was r = 0.60 (r² = 0.36), indicating 64% of variance remains unexplained. Heuristic composite scoring: Weights (0.5 DL + 0.3 Similarity + 0.2 Targets) were selected without systematic optimization. Sensitivity analysis showed 20–40% of top-10 drugs changed under alternative weighting schemes. No clinical validation: Predictions were not correlated with actual patient treatment outcomes. Future validation should include in vitro screening of top predicted drugs in ovarian cancer cell line panels representative of each subtype (e.g., OVCAR3, SKOV3, IGROV1, KURAMOCHI), validation in patient-derived organoids or xenograft models, and optimization of composite scoring weights via machine learning. These drug predictions are explicitly framed as hypothesis-generating and exploratory, not clinically actionable. 4.6 Incomplete Integration of Multi-Omics Data This analysis was limited to RNA-seq data, despite TCGA-OV providing rich multi-omics annotations including mutations, copy number alterations, methylation, and proteomic data. Critically, the HRD designation for Cluster 2 rests entirely on gene expression without validation using BRCA1/2 mutation status, genomic HRD scores (LOH, LST, TAI), copy number patterns (CCNE1 amplification, BRCA1 deletions), immune deconvolution (CIBERSORT, xCell, ESTIMATE), or somatic mutation burden. Integrating these omics layers would substantially strengthen biological interpretations and enable mechanistic hypotheses. 4.7 Comparison with Prior Literature The TCGA 2011 ovarian cancer study identified four subtypes (Immunoreactive, Proliferative, Mesenchymal, Differentiated) using consensus hierarchical clustering on 1,500 curated genes.5 Our VAE-derived subtypes show conceptual overlap: Cluster 2 appears analogous to the Immunoreactive subtype (immune activation, interferon signaling), while Clusters 0 and 1 may represent subdivisions of the Mesenchymal/Proliferative categories. The modest correspondence with classical consensus subtypes (ARI = 0.201) suggests the VAE captures partially distinct biological structure, reflecting the nonlinear nature of transcriptomic variation that linear methods may fail to resolve. The added value of our approach lies in: (1) unsupervised, data-driven learning without manual gene curation; (2) demonstrated superiority over conventional methods through comprehensive benchmarking; (3) probabilistic cluster assignments enabling identification of mixed/transitional phenotypes; and (4) integration with a multi-component drug repurposing framework. Whether these subtypes provide clinical value beyond existing classifications requires validation with survival and treatment response data in independent cohorts. 4.8 Study Strengths Despite the limitations detailed above, this study demonstrates several notable strengths: Rigorous internal validation: Multiple metrics (entropy, consensus stability, silhouette, ARI) consistently supported the three-cluster structure. Comprehensive benchmarking: Quantitative comparison against four alternative methods demonstrated VAE-GMM superiority. Biological coherence: Subtypes exhibited distinct, interpretable pathway enrichments aligned with known ovarian cancer biology. Formal survival analysis: Comprehensive Kaplan-Meier and Cox regression analyses were performed, yielding clinically interpretable results even where no prognostic significance was detected. 13. Batch effect assessment: Analysis demonstrated no significant confounding by technical variables. 14. Reproducible open-science framework: Detailed methodology with open-source implementation enabling replication. Integrated drug repurposing: Provided testable therapeutic hypotheses for each molecular subtype with clearly stated confidence levels. 4.9 Future Directions To advance this work toward clinical translation, the following studies are recommended: External validation: Replicate subtypes in ≥ 2 independent ovarian cancer cohorts (AOCS, ICGC, GEO datasets) with assessment of cluster concordance and stability. Treatment-stratified survival analysis: Kaplan-Meier and multivariate Cox regression in cohorts receiving targeted therapies aligned with subtype-specific vulnerabilities (e.g., PARP inhibitor trials for Cluster 2). 18. Multi-omics integration: Incorporate BRCA/HRD status, copy number, methylation, immune infiltration quantification, and mutation burden. 19. Experimental drug validation: In vitro screening in subtype-representative cell lines, PDX models, and patient-derived organoids. 20. Clinical trial correlation: Retrospective analysis in clinical trial cohorts to validate subtype-treatment associations. Methodological refinements: Systematic hyperparameter optimization, drug scoring weight learning via machine learning, and integration with single-cell RNA-seq for cellular deconvolution. 5. Conclusion This study developed a rigorous computational framework integrating variational autoencoders, consensus clustering, and multi-metric validation to identify three molecularly distinct ovarian cancer subtypes in TCGA-OV data. Comprehensive benchmarking demonstrated that VAE-based nonlinear dimensionality reduction outperformed conventional linear methods (PCA, NMF) in capturing biologically coherent subtypes. The identified clusters—mesenchymal/ciliary, fibrotic/protease-enriched, and HRD/immune-active—exhibited distinct pathway enrichments aligned with known ovarian cancer biology and generated plausible therapeutic hypotheses through integrated drug repurposing analysis. Formal survival analysis revealed no significant prognostic differences across subtypes, suggesting these molecular classifications reflect predictive rather than prognostic biological characteristics in the context of uniform platinum-based therapy. Critical limitations constrain immediate translational impact: the absence of external validation in independent cohorts means these subtypes remain provisional findings requiring replication; drug predictions are entirely computational without experimental validation and should be interpreted as hypothesis-generating only; and the HRD designation for Cluster 2 requires confirmation with BRCA mutation data and genomic HRD scores. With appropriate validation studies—external cohort replication, treatment-stratified survival analysis, multi-omics integration, and experimental drug screening—this framework could contribute meaningfully to precision oncology in ovarian cancer. In its current form, it provides a reproducible, rigorously benchmarked computational approach for molecular subtype discovery and serves as a foundation for hypothesis-driven translational research toward individualized ovarian cancer therapy. Declarations Conflicts of Interest The authors declare no competing financial interests or personal relationships that could have influenced the work reported in this paper. Funding This research received no specific grant from funding agencies in the public, commercial, or not-for-profit sectors. Acknowledgements We thank The Cancer Genome Atlas Research Network and contributing institutions for generating and sharing the TCGA-OV dataset. We acknowledge the GDSC and CTRP consortia for providing pharmacogenomic training data. Data and Code Availability TCGA-OV RNA-seq data are publicly available from the NIH Genomic Data Commons (GDC) Data Portal ( https://portal.gdc.cancer.gov/ ). All analysis code (Python) is available at https://github.com/AmirArshiaBeheshti/Deep-learning-reveals-three-distinct-ovarian-cancer-subtypes-with-therapeutic-relevance.git with detailed documentation and reproducibility instructions. GDSC and CTRP pharmacogenomic data are available from https://www.cancerrxgene.org/ and https://portals.broadinstitute.org/ctrp/ respectively. References Cancer Genome Atlas Research Network (2011) Integrated genomic analyses of ovarian carcinoma. Nature 474(7353):609–615. 10.1038/nature10166 Lheureux S, Gourley C, Vergote I, Oza AM (2019) Epithelial ovarian cancer. Lancet 393(10177):1240–1253. 10.1016/S0140-6736(18)32552-2 Davis A, Tinker AV, Friedlander M (2014) Platinum resistant’ ovarian cancer: what is it, who to treat and how to measure benefit? Gynecol Oncol 133(2):624–631. 10.1016/j.ygyno.2014.02.038 Bowtell DD, Böhm S, Ahmed AA et al (2015) Rethinking ovarian cancer II: reducing mortality from high-grade serous ovarian cancer. Nat Rev Cancer 15(11):668–679. 10.1038/nrc4019 Tothill RW, Tinker AV, George J et al (2008) Novel molecular subtypes of serous and endometrioid ovarian cancer linked to clinical outcome. Clin Cancer Res 14(16):5198–5208. 10.1158/1078-0432.CCR-08-0196 Verhaak RGW, Tamayo P, Yang JY et al (2013) Prognostically relevant gene signatures of high-grade serous ovarian carcinoma. J Clin Invest 123(1):517–525. 10.1172/JCI65833 Konecny GE, Wang C, Hamidi H et al (2014) Prognostic and therapeutic relevance of molecular subtypes in high-grade serous ovarian cancer. J Natl Cancer Inst 106(10):dju249. 10.1093/jnci/dju249 Chen GM, Kannan L, Geistlinger L et al (2018) Consensus on Molecular Subtypes of High-Grade Serous Ovarian Carcinoma. Clin Cancer Res 24(20):5037–5047. 10.1158/1078-0432.CCR-18-0784 Kingma DP, Welling M (2014) Auto-Encoding Variational Bayes. Proceedings of ICLR. arXiv:1312.6114 Way GP, Greene CS (2018) Extracting a biologically relevant latent space from cancer transcriptomes with variational autoencoders. Pac Symp Biocomput 23:80–91. 10.1142/9789813235533_0008 Simidjievski N, Bodnar C, Tariq I et al (2019) Variational Autoencoders for Cancer Data Integration: Design Principles and Computational Practice. Front Genet 10:1205. 10.3389/fgene.2019.01205 Simidjievski N, Shams Z, Jamnik M (2021) Integrated multi-omics analysis of ovarian cancer using variational autoencoders. Sci Rep 11:6265. 10.1038/s41598-021-85285-4 Rampášek L, Hidru D, Smirnov P et al (2019) Dr.VAE: improving drug response prediction via modeling of drug perturbation effects. Bioinformatics 35(19):3743–3751. 10.1093/bioinformatics/btz158 Zhang Z, Hernandez K, Savage J et al (2021) Nat Commun 12:1226. 10.1038/s41467-021-21254-9 . Uniform genomic data analysis in the NCI Genomic Data Commons Heath AP, Ferretti V, Agrawal S et al (2017) The NCI Genomic Data Commons as an engine for precision medicine. Blood 130(4):453–459. 10.1182/blood-2017-03-735654 Li B, Dewey CN (2011) RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics 12:323. 10.1186/1471-2105-12-323 Wagner GP, Kin K, Lynch VJ (2012) Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples. Theory Biosci 131(4):281–285. 10.1007/s12064-012-0162-3 Cheadle C, Vawter MP, Freed WJ, Becker KG (2003) Analysis of microarray data using Z score transformation. J Mol Diagn 5(2):73–81. 10.1016/S1525-1578(10)60455-2 Kaplan EL, Meier P (1958) Nonparametric estimation from incomplete observations. J Am Stat Assoc 53(282):457–481. 10.1080/01621459.1958.10501452 Cox DR (1972) Regression models and life-tables. J R Stat Soc Ser B 34(2):187–220 Kingma DP, Ba J, Adam (2015) A Method for Stochastic Optimization. Proceedings of ICLR arXiv:1412.6980 McInnes L, Healy J, Saul N, Großberger L (2018) UMAP: Uniform Manifold Approximation and Projection. J Open Source Softw 3(29):861. 10.21105/joss.00861 Becht E, McInnes L, Healy J et al (2019) Dimensionality reduction for visualizing single-cell data using UMAP. Nat Biotechnol 37:38–44. 10.1038/nbt.4314 Rousseeuw PJ (1987) Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J Comput Appl Math 20:53–65. 10.1016/0377-0427(87)90125-7 Caliński T, Harabasz J (1974) A dendrite method for cluster analysis. Commun Stat 3(1):1–27 Davies DL, Bouldin DW (1979) A cluster separation measure. IEEE Trans Pattern Anal Mach Intell PAMI–1(2):224–227. 10.1109/TPAMI.1979.4766909 Ritchie ME, Phipson B, Wu D et al (2015) limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 43(7):e47. 10.1093/nar/gkv007 Benjamini Y, Hochberg Y (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B 57(1):289–300 Ashburner M, Ball CA, Blake JA et al (2000) Gene ontology: tool for the unification of biology. Nat Genet 25(1):25–29. 10.1038/75556 The Gene Ontology Consortium (2023) The Gene Ontology knowledgebase in 2023. Genetics 224(1):iyad031. 10.1093/genetics/iyad031 Kanehisa M, Goto S (2000) KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res 28(1):27–30. 10.1093/nar/28.1.27 Kanehisa M, Furumichi M, Sato Y et al (2023) KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res 51(D1):D587–D592. 10.1093/nar/gkac963 Fang Z, Liu X, Peltz G (2023) GSEApy: a comprehensive package for performing gene set enrichment analysis in Python. Bioinformatics 39(1):btac757. 10.1093/bioinformatics/btac757 Yang W, Soares J, Greninger P et al (2013) Genomics of Drug Sensitivity in Cancer (GDSC): a resource for therapeutic biomarker discovery in cancer cells. Nucleic Acids Res 41(Database issue):D955–D961. 10.1093/nar/gks1111 Garnett MJ, Edelman EJ, Heidorn SJ et al (2012) Systematic identification of genomic markers of drug sensitivity in cancer cells. Nature 483(7391):570–575. 10.1038/nature11005 Seashore-Ludlow B, Rees MG, Cheah JH et al (2015) Harnessing Connectivity in a Large-Scale Small-Molecule Sensitivity Dataset. Cancer Discov 5(11):1210–1223. 10.1158/2159-8290.CD-15-0235 Basu A, Bodycombe NE, Cheah JH et al (2013) An interactive resource to identify cancer genetic and lineage dependencies targeted by small molecules. Cell 154(5):1151–1161. 10.1016/j.cell.2013.08.003 Rogers D, Hahn M, Extended-Connectivity Fingerprints (2010) J Chem Inf Model 50(5):742–754. 10.1021/ci100050t Yu G, Wang LG, Han Y, He QY (2012) clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS 16(5):284–287. 10.1089/omi.2011.0118 Wu T, Hu E, Xu S et al (2021) clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innov 2(3):100141. 10.1016/j.xinn.2021.100141 Hassounah NB, Nagle R, Saboda K et al (2019) Primary Cilium in Cancer Hallmarks. Int J Mol Sci 20(6):1336. 10.3390/ijms20061336 Winkler J, Abisoye-Ogunniyan A, Metcalf KJ, Werb Z (2020) Concepts of extracellular matrix remodelling in tumour progression and metastasis. Nat Commun 11:5120. 10.1038/s41467-020-18794-x Huang J, Zhang L, Wan D et al (2023) Extracellular matrix remodeling in tumor progression and immune escape: from mechanisms to treatments. Mol Cancer 22:48. 10.1186/s12943-023-01744-8 Tamir A, Jag U, Sarojini S et al (2014) Kallikrein family proteases KLK6 and KLK7 are potential early detection and diagnostic biomarkers for serous and papillary serous ovarian cancer subtypes. J Ovarian Res 7:109. 10.1186/s13048-014-0109-z Zafrakas M, Petschke B, Donner A et al (2006) Expression analysis of mammaglobin A (SCGB2A2) and lipophilin B (SCGB1D2) in more than 300 human tumors and matching normal tissues reveals their co-expression in gynecologic malignancies. BMC Cancer 6:88. 10.1186/1471-2407-6-88 Prakash R, Zhang Y, Feng W, Jasin M (2015) Homologous recombination and human health: the roles of BRCA1, BRCA2, and associated proteins. Cold Spring Harb Perspect Biol 7(4):a016600. 10.1101/cshperspect.a016600 Dangaj D, Bruand M, Grimm AJ et al (2019) Cooperation between Constitutive and Inducible Chemokines Enables T Cell Engraftment and Immune Attack in Solid Tumors. Cancer Cell 35(6):885–900. 10.1016/j.ccell.2019.05.004 Chow MT, Ozga AJ, Servis RL et al (2020) Macrophage-Derived CXCL9 and CXCL10 Are Required for Antitumor Immune Responses Following Immune Checkpoint Blockade. Clin Cancer Res 26(2):487–504. 10.1158/1078-0432.CCR-19-1868 Moore K, Colombo N, Scambia G et al (2018) Maintenance Olaparib in Patients with Newly Diagnosed Advanced Ovarian Cancer. N Engl J Med 379(26):2495–2505. 10.1056/NEJMoa1810858 Ray-Coquard I, Pautier P, Pignata S et al (2019) Olaparib plus Bevacizumab as First-Line Maintenance in Ovarian Cancer. N Engl J Med 381(25):2416–2428. 10.1056/NEJMoa1911361 Cortez A, Zhou J, Zhang Y, Lim C (2023) PARP Inhibitors in Ovarian Cancer: A Review. Target Oncol 18(4):471–503. 10.1007/s11523-023-00970-w Ozols RF (2018) Carboplatin/Paclitaxel Induction in Ovarian Cancer: The Finer Points. Oncologist 23(9):1003–1008 Chan JK, Brady MF, Penson RT et al (2016) Weekly vs. Every-3-Week Paclitaxel and Carboplatin for Ovarian Cancer. N Engl J Med 374(8):738–748. 10.1056/NEJMoa1505067 Clamp AR, James EC, McNeish IA et al (2019) Weekly dose-dense chemotherapy in first-line epithelial ovarian cancer treatment (ICON8): primary progression free survival results from a GCIG phase 3 randomised controlled trial. Lancet 394(10214):2084–2095. 10.1016/S0140-6736(19)32259-7 Monti S, Tamayo P, Mesirov J, Golub T (2003) Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Mach Learn 52(1–2):91–118. 10.1023/A:1023949509487 Wilkerson MD, Hayes DN (2010) ConsensusClusterPlus: a class discovery tool with confidence assessments and item tracking. Bioinformatics 26(12):1572–1573. 10.1093/bioinformatics/btq170 Pedregosa F, Varoquaux G, Gramfort A et al (2011) Scikit-learn: Machine Learning in Python. J Mach Learn Res 12:2825–2830 Therneau TM, Grambsch PM (2000) Modeling Survival Data: Extending the Cox Model. Springer, New York Scrucca L, Fop M, Murphy TB, Raftery AE (2016) mclust 5: Clustering, Classification and Density Estimation Using Gaussian Finite Mixture Models. R J 8(1):289–317 National Cancer Institute. SEER Cancer Stat Facts (2024) Ovarian Cancer. Available at: https://seer.cancer.gov/statfacts/html/ovary.html . Accessed January Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9693642","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":639116096,"identity":"29821887-9865-40d6-8b64-c0722e3ea060","order_by":0,"name":"Amir Arshia Beheshti","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA6UlEQVRIiWNgGAWjYBACAwYeBoYEBgsexvbmA0C+hAyxWiR4mHuOJYC08BCnBaiSgX2GjwGIRViLOfvZgx8e1EjI8M7g+fzqRo0FDwP74aMb8Gmx7MlLlkg4JsEjObt3m3UOkMHAk5Z2A6/DDuQYSCSwSfAYzjm7zTgHyAB6xwy/lvNvjH8k/JPgsb+R88w45x8xWm7kmEkktknwMM7IYX6c20aEFssZ79IsEvuAWnqOmTHnAhlshPxizp97+OaPbzb2wKh8/DnnW50cP/vhY3i1IAM2CTBJrHIQYP5AiupRMApGwSgYOQAAvGJGORLbIwEAAAAASUVORK5CYII=","orcid":"https://orcid.org/0009-0009-1546-9286","institution":"Students Research Committee, School of Medicine, Ardabil University of Medical Sciences, Ardabil, Iran","correspondingAuthor":true,"prefix":"","firstName":"Amir","middleName":"Arshia","lastName":"Beheshti","suffix":""},{"id":639118707,"identity":"6d109f66-973d-47d3-a31b-2a32509fb78b","order_by":1,"name":"Vahid Bossaghzadeh","email":"","orcid":"","institution":"Department of Obstetrics and Gynaecology, Royal College of Surgeons in Ireland, Dublin, Ireland","correspondingAuthor":false,"prefix":"","firstName":"Vahid","middleName":"","lastName":"Bossaghzadeh","suffix":""}],"badges":[],"createdAt":"2026-05-12 14:51:23","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-9693642/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9693642/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":109275560,"identity":"f3449516-efac-40e9-80f0-879cfb0a8e07","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":122008,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eModel Training Performance. Training (top) shows stable loss and accuracy convergence over 50 epochs without overfitting. Predicted drug response scores (bottom left) and model summary (bottom right) highlight architecture details, 7,585 training pairs, and final accuracies (Train: 96.8%, Val: 97.6%).\u003c/em\u003e\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/0b871a6c8f20484dff734a27.png"},{"id":109296066,"identity":"29987138-4857-422c-b8f0-3fec03b21c79","added_by":"auto","created_at":"2026-05-15 08:45:10","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":114932,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eVAE Training Curves. Training loss (blue) and test loss (orange) plotted across 100 epochs. The divergence between curves stabilizes after epoch 40, indicating minimal overfitting. Note that the y-axis scale differs between training and test loss curves.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/86b58b362d6924f723bec6eb.png"},{"id":109275562,"identity":"10708b51-de24-43a9-85d5-5a1449fc4a13","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":50284,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eOverall Survival across VAE-GMM Molecular Subtypes. Kaplan-Meier curves for Cluster 0 (Mesenchymal), Cluster 1 (Fibrotic), and Cluster 2 (HRD/Immune). No significant survival difference was observed (log-rank p = 0.885). Shaded regions indicate 95% confidence intervals; numbers at risk are shown below.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/695ea2503e74e424f57ded06.png"},{"id":109296935,"identity":"14f0a2da-e191-4112-8842-03601c827ca4","added_by":"auto","created_at":"2026-05-15 08:52:15","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":48942,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eOverall Survival across Classical Consensus Subtypes (contextual comparison). Kaplan-Meier curves comparing Classical Subtypes 1–3. A marginally significant survival difference was observed (log-rank p = 0.05). Shaded bands indicate 95% confidence intervals.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/ca0478007a07fb9b08917261.png"},{"id":109296563,"identity":"7edbc747-e7db-466d-a2fa-840a4bfc6005","added_by":"auto","created_at":"2026-05-15 08:48:14","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":464294,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eFinal VAE+GMM Clustering Results. Left: Three identified clusters (Cluster 0 in teal, Cluster 1 in gray, Cluster 2 in yellow-green) visualized in UMAP space. Right: Cluster assignment uncertainty measured by entropy, showing high confidence assignments (low entropy, dark green) for most samples with few uncertain cases.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/d4f363585a9bc8dd2000dd90.png"},{"id":109275574,"identity":"9954f73f-d9f8-411b-be51-f181c0a4c598","added_by":"auto","created_at":"2026-05-14 14:59:02","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":450764,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eComparison of Clustering Methods. Left: K-means clustering (k = 2) with Silhouette coefficient = 0.133. Right: Gaussian Mixture Model (GMM) with n = 3 components showing superior cluster separation and biological coherence.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/dca648c419f5cc8c24570edc.png"},{"id":109275564,"identity":"771c127b-e2fb-44fe-be91-bba49258cde7","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":94818,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eCorrespondence between Classical Consensus Subtypes and VAE-GMM Subtypes. Sankey diagram illustrating sample transitions from three classical consensus subtypes (Classical 1–3) to the three VAE-GMM clusters (Cluster_0–2). Overall concordance is modest (Adjusted Rand Index = 0.201, n = 434), indicating the VAE-GMM approach captures partially distinct biological structure.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/29ad54a91ad95ccbf1d6b7f6.png"},{"id":109275566,"identity":"64f08e8a-3bf9-4231-944f-c071ef2e0fa4","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":239063,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDifferential Expression: Cluster 2 vs Cluster 3. Volcano plot showing contrasting molecular signatures. Key genes include NKD2, CPA2, IGF2BP2, CEL (upregulated in Cluster 2) and LINC01541, SCGB2A2, FUT6, SCGB2A1, CEACAM7 (upregulated in Cluster 3). Red points: |log₂FC| \u0026gt; 1, FDR \u0026lt; 0.05.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/e6eade32ac23c7a8676b89e1.png"},{"id":109275568,"identity":"5eb2367f-6ad4-4f7f-863a-a204bd0fe8e9","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":295166,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDifferential Expression: Cluster 1 vs Cluster 3. Volcano plot highlighting distinct transcriptional programs. Notable markers include CPA2, CEL, CLPS, SOX21, SOX21-AS1, MNX1, SPINK1, APOH (upregulated in Cluster 1), and KLK7, KLK6 (upregulated in Cluster 3).\u003c/em\u003e\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/0a47d27ef0daac56d4ee6b9a.png"},{"id":109275567,"identity":"599d6c33-93a3-4b72-a604-885619277a86","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":288124,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDifferential Expression: Cluster 1 vs Cluster 2. Volcano plot showing log₂ fold change versus −log₁₀(adjusted p-value). Red points indicate significantly differentially expressed genes (|log₂FC| \u0026gt; 1, FDR \u0026lt; 0.05). Key markers include LINC01541, SCGB2A2, LRRC31, SERPINA4, CEACAM7, TCN1, FUT6, and HNF1A-AS1.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/06cb54061cd997105cc2ec69.png"},{"id":109275570,"identity":"9b4c5ac4-72ea-4edb-9d6e-0124faaf1813","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":52897,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDrug Repurposing Overview. Left: Differentially expressed genes (blue) and drug-targetable subsets (red) across clusters (0–2). Right: Distribution of 75 analyzed drugs—83.7% investigational, 14.9% approved anticancer, and 1.4% approved non-cancer agents.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/e0685a6f5194eebfb4e327b8.png"},{"id":109296521,"identity":"42da5db0-2c58-467a-bbbb-5dc0b42ac579","added_by":"auto","created_at":"2026-05-15 08:47:46","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":131263,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDrug-Cluster Similarity. Top: Heatmap of similarity scores for the top 15 drugs and histogram of scores across all 75 drugs per cluster. Bottom: Top 5 candidates per cluster and summary statistics (75 drugs, 995 signature genes, cosine similarity method).\u003c/em\u003e\u003c/p\u003e","description":"","filename":"12.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/ec9a906a10f1ba7c98eb6931.png"},{"id":109275569,"identity":"a4dfb27b-8bff-428d-976a-0a57756e8b39","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":88564,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eTop 10 Drug Candidates per Subtype. Bar plots show top drugs ranked by composite score for each cluster. Cluster 0 (mesenchymal/ciliary): led by DOCETAXEL (0.725). Cluster 1 (fibrotic/protease): led by DEXAMETHASONE (0.725). Cluster 2 (HRD/immune): led by CARBOPLATIN (0.632). Scores integrate deep learning predictions (50%), target similarity (30%), and pathway overlap (20%).\u003c/em\u003e\u003c/p\u003e","description":"","filename":"13.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/00f72d92102ecafd44505b63.png"},{"id":109296581,"identity":"fa5cbbeb-b4ff-4fad-b7c1-99b5d85393ca","added_by":"auto","created_at":"2026-05-15 08:48:17","extension":"png","order_by":14,"title":"Figure 14","display":"","copyAsset":false,"role":"figure","size":115106,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eDrug Score Heatmap Across Subtypes. Heatmap of composite final scores for 16 priority drugs across the three clusters. Green = high (0.62–0.73), yellow-orange = moderate (0.40–0.50), red = low (0.20–0.23). Cluster 0 responds strongly to DOCETAXEL, SIMVASTATIN, TAMOXIFEN, ETOPOSIDE; Cluster 1 to DEXAMETHASONE, ASPIRIN, PAZOPANIB, SIROLIMUS; Cluster 2 shows moderate scores for CARBOPLATIN, GEFITINIB, GEMCITABINE, ERLOTINIB.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"14.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/2dcf657479a68101cbadc5ba.png"},{"id":109275572,"identity":"427ebc3e-5c7c-4913-9915-2c8a53e572ad","added_by":"auto","created_at":"2026-05-14 14:59:01","extension":"png","order_by":15,"title":"Figure 15","display":"","copyAsset":false,"role":"figure","size":166927,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003eComposite Score Breakdown by Cluster. Stacked bars show contributions of deep learning predictions (blue, 50%), similarity scores (red, 30%), and target overlap (green, 20%) for the top 8 drugs per cluster. Cluster 0 and 1 exhibit high DL contributions (0.35–0.40), while Cluster 2 shows lower overall scores due to reduced DL input.\u003c/em\u003e\u003c/p\u003e","description":"","filename":"15.png","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/ff85540dbfc545ec04fbee17.png"},{"id":109297573,"identity":"e065fcc4-2aa7-4cd7-814e-22c6bc0b685b","added_by":"auto","created_at":"2026-05-15 09:00:01","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2784664,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9693642/v1/639aa0a2-8a04-470f-8444-bc3747b122db.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eDeep Learning-Based Molecular Subtyping Identifies Three Biologically Distinct Ovarian Cancer Subtypes with Potential Therapeutic Implications\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"Highlights","content":"\u003cp\u003e\u0026bull; A variational autoencoder\u0026ndash;Gaussian Mixture Model (VAE-GMM) pipeline was applied to 421 TCGA ovarian cancer transcriptomes to learn a biologically coherent 32-dimensional latent space.\u003c/p\u003e\u003cp\u003e\u0026bull; Consensus clustering across 1,000 bootstrap resamplings identified three robust molecular subtypes (mesenchymal/ciliary, fibrotic/protease-enriched, and HRD/immune-enriched) with mean assignment entropy of 0.05.\u003c/p\u003e\u003cp\u003e\u0026bull; VAE-GMM achieved superior cluster separation (Silhouette\u0026thinsp;=\u0026thinsp;0.61) versus PCA\u0026thinsp;+\u0026thinsp;k-means (0.45), PCA+hierarchical (0.42), NMF (0.38), and UMAP+HDBSCAN (0.39).\u003c/p\u003e\u003cp\u003e\u0026bull; Comprehensive survival analyses revealed no significant prognostic difference across subtypes (OS: log-rank p\u0026thinsp;=\u0026thinsp;0.885; DFS: p\u0026thinsp;=\u0026thinsp;0.783), suggesting predictive rather than prognostic utility.\u003c/p\u003e\u003cp\u003e\u0026bull; A composite deep learning drug repurposing pipeline prioritized subtype-specific candidates: taxanes for Cluster 0, dexamethasone/anti-fibrotics for Cluster 1, and PARP inhibitors/carboplatin for Cluster 2.\u003c/p\u003e"},{"header":"1. Introduction","content":"\u003cp\u003eOvarian cancer (OC) is a highly lethal gynecologic malignancy characterized by pronounced molecular and clinical heterogeneity, with 5-year survival rates ranging from 29% for advanced-stage disease to 92% for localized tumors.1,2 Despite decades of genomic research, reliable molecular subtyping that translates into actionable therapy remains a central unmet need. Approximately 70\u0026ndash;80% of patients develop platinum resistance,3,4 and molecular stratification could potentially guide therapeutic decision-making toward more effective, individualized regimens.\u003c/p\u003e \u003cp\u003eThe Cancer Genome Atlas (TCGA) ovarian cancer study in 20115 identified four molecular subtypes\u0026mdash;Immunoreactive, Proliferative, Mesenchymal, and Differentiated\u0026mdash;using consensus hierarchical clustering on curated gene sets. While this landmark classification has informed subsequent research, its clinical translation has been limited due to: (i) dependence on manually curated gene panels that may not capture the full transcriptomic complexity, (ii) use of linear dimensionality reduction that cannot model nonlinear biological relationships, and (iii) inconsistent survival differences in subsequent validation studies.6,7,8 The optimal computational approach for transcriptome-based subtype discovery therefore remains an active and important area of investigation.\u003c/p\u003e \u003cp\u003eDeep generative models, particularly variational autoencoders (VAEs),9,10 offer an advanced approach to compress high-dimensional omics data into interpretable latent spaces that preserve complex nonlinear relationships. VAEs have been successfully applied to cancer transcriptomics for tumor classification, biomarker discovery, and outcome prediction.10,11,12,13 When coupled with probabilistic clustering such as Gaussian Mixture Models (GMM) and ensemble stability methods like consensus clustering, these latent representations can reveal reproducible transcriptional structures potentially missed by conventional linear dimensionality reduction approaches.\u003c/p\u003e \u003cp\u003eHere, we systematically applied a VAE-GMM-consensus pipeline to the TCGA-OV transcriptomic dataset (n\u0026thinsp;=\u0026thinsp;421) to identify stable, biologically coherent molecular subtypes. The methodological contribution of this work is threefold: (\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e) we demonstrate that nonlinear latent learning via VAE substantially outperforms conventional PCA, NMF, and hierarchical clustering in producing biologically separable subtypes; (\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e) we integrate differential expression, pathway enrichment, survival analysis, and a composite deep learning drug-ranking model to characterize each cluster and link it with potential therapeutic vulnerabilities; and (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e) we conduct rigorous internal validation through consensus clustering, batch effect analysis, and multi-metric benchmarking while transparently acknowledging limitations related to external validation and experimental drug testing.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Dataset and Preprocessing\u003c/h2\u003e \u003cp\u003eGene expression data from The Cancer Genome Atlas Ovarian Cancer (TCGA-OV) cohort were obtained from the GDC Data Portal14,15 (data freeze version 32.0, accessed January 2024). The dataset comprised 421 high-grade serous ovarian cancer tumor samples profiled by bulk RNA-sequencing. Raw TPM (transcripts per million)16,17 values were log₂-transformed (log₂(TPM\u0026thinsp;+\u0026thinsp;1)) to stabilize variance and approximate normal distribution.\u003c/p\u003e \u003cp\u003eGene filtering was performed to remove transcripts with zero variance across all samples or extremely low expression (mean TPM\u0026thinsp;\u0026lt;\u0026thinsp;0.1), resulting in retention of 19,842 genes. The expression matrix was then standardized using gene-wise z-score normalization18 (mean\u0026thinsp;=\u0026thinsp;0, standard deviation\u0026thinsp;=\u0026thinsp;1 across samples for each gene) to ensure comparability across different expression ranges. The final preprocessed matrix (19,842 genes \u0026times; 421 samples) was used for all downstream analyses.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Batch Effect Assessment\u003c/h2\u003e \u003cp\u003eTo evaluate potential technical confounders, we extracted TCGA metadata including tissue source site (TSS), sequencing center, and plate identifiers. Principal component analysis (PCA) was performed on the normalized expression matrix, and the first two principal components were visualized colored by technical variables. Chi-square tests were conducted to assess whether cluster membership was significantly associated with technical covariates (tissue source site, sequencing center). P-values were adjusted using Bonferroni correction for multiple testing.\u003c/p\u003e \u003cp\u003eOur analysis revealed no significant association between cluster membership and tissue source site (χ\u0026sup2; = 8.3, p\u0026thinsp;=\u0026thinsp;0.76) or sequencing center (χ\u0026sup2; = 6.1, p\u0026thinsp;=\u0026thinsp;0.91) after Bonferroni correction, suggesting that biological signal dominates over technical variation. Visual inspection of PCA plots showed no clear clustering pattern by technical variables. Based on these findings, we proceeded without additional batch correction, though this remains a potential confounder requiring evaluation in external datasets.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Survival Analysis\u003c/h2\u003e \u003cp\u003eClinical data including overall survival (OS), disease-free survival (DFS), vital status, and days to event were extracted from TCGA clinical annotations available through the GDC Data Portal.14,15 Overall survival was defined as time from initial diagnosis to death from any cause or last follow-up. Disease-free survival was defined as time from initial diagnosis to disease recurrence, progression, or death, whichever occurred first.\u003c/p\u003e \u003cp\u003eKaplan-Meier survival curves19 were generated for each molecular subtype using the survival package (v3.5-5) in R version 4.3.0. Log-rank tests were performed to assess statistical significance of survival differences across the three clusters. Cox proportional hazards20 regression models were fitted using the coxph function to estimate hazard ratios (HR) and 95% confidence intervals (CI) for each cluster relative to Cluster 0 (reference group). The proportional hazards assumption was assessed using Schoenfeld residuals and found to be satisfied for all covariates. All survival analyses were two-sided, and p-values\u0026thinsp;\u0026lt;\u0026thinsp;0.05 were considered statistically significant. Survival curves were visualized using the survminer package (v0.4.9).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Variational Autoencoder Architecture and Training\u003c/h2\u003e \u003cp\u003eA Variational Autoencoder (VAE) was implemented in TensorFlow 2.10 to learn a compressed latent representation of the transcriptomic data.9 The encoder comprised three fully connected layers with dimensions [19,842 \u0026rarr; 256 \u0026rarr; 128 \u0026rarr; 32], using ReLU activations for hidden layers. The 32-dimensional latent space was parameterized by mean (\u0026micro;) and log-variance (log σ\u0026sup2;) vectors, from which latent samples were drawn using the reparameterization trick.9 The decoder mirrored the encoder architecture [32 \u0026rarr; 128 \u0026rarr; 256 \u0026rarr; 19,842] with linear activation in the output layer for reconstruction.\u003c/p\u003e \u003cp\u003eThe loss function combined reconstruction error (mean squared error) and Kullback\u0026ndash;Leibler (KL) divergence regularization with β\u0026thinsp;=\u0026thinsp;1 (standard VAE formulation):\u003c/p\u003e \u003cp\u003e \u003cb\u003eℒ = MSE(x, x̂)\u0026thinsp;+\u0026thinsp;KL(q(z|x) ‖ p(z))\u003c/b\u003e\u003c/p\u003e \u003cp\u003eThe model was trained using the Adam optimizer21 (learning rate\u0026thinsp;=\u0026thinsp;1\u0026times;10⁻\u0026sup3;, β₁ = 0.9, β₂ = 0.999) with batch size 64 for 200 epochs. An 80/20 train-validation split was employed with early stopping (patience\u0026thinsp;=\u0026thinsp;20 epochs, monitoring validation loss). The model converged after approximately 150 epochs, achieving final training loss of 1.23 and validation loss of 1.31. All random seeds were fixed at 42 for reproducibility.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Hyperparameter Selection and Sensitivity Analysis\u003c/h2\u003e \u003cp\u003eThe choice of latent dimension (d\u0026thinsp;=\u0026thinsp;32) was based on exploratory experiments testing d \u0026isin; {16, 32, 64, 128}. Results showed that d\u0026thinsp;=\u0026thinsp;16 yielded insufficient representational capacity (test MSE\u0026thinsp;=\u0026thinsp;1.87, Silhouette\u0026thinsp;=\u0026thinsp;0.43), while d\u0026thinsp;=\u0026thinsp;64 and d\u0026thinsp;=\u0026thinsp;128 showed marginal improvement over d\u0026thinsp;=\u0026thinsp;32 but substantially increased computational cost. The configuration d\u0026thinsp;=\u0026thinsp;32 provided optimal balance (test MSE\u0026thinsp;=\u0026thinsp;1.31, Silhouette\u0026thinsp;=\u0026thinsp;0.61).\u003c/p\u003e \u003cp\u003eLayer architecture (256-128-32) was selected following established practices in VAE design for genomic data, where gradual dimension reduction from input to latent space typically outperforms abrupt compression. The KL divergence weighting β\u0026thinsp;=\u0026thinsp;1 represents standard VAE formulation; while β-VAE (β\u0026thinsp;\u0026lt;\u0026thinsp;1 or β\u0026thinsp;\u0026gt;\u0026thinsp;1) can sometimes improve disentanglement, we found standard weighting sufficient for our clustering objectives.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Dimensionality Reduction and Visualization\u003c/h2\u003e \u003cp\u003eTo visualize the 32-dimensional latent space structure, Uniform Manifold Approximation and Projection (UMAP)22,23 was applied with parameters: n_neighbors\u0026thinsp;=\u0026thinsp;15, min_dist\u0026thinsp;=\u0026thinsp;0.1, metric = \u0026lsquo;cosine\u0026rsquo;, random_state\u0026thinsp;=\u0026thinsp;42. UMAP was chosen over t-SNE for its superior preservation of both local and global structure. The resulting 2D embeddings revealed three visually distinct clusters corresponding to the molecular subtypes identified by subsequent probabilistic clustering.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e2.7 Probabilistic Clustering and Consensus Validation\u003c/h2\u003e \u003cp\u003eGaussian Mixture Model (GMM) clustering was performed on the 32-dimensional VAE latent embeddings using scikit-learn\u0026rsquo;s GaussianMixture implementation. We tested k \u0026isin; {2, 3, 4, 5, 6} clusters with full covariance matrices, k-means\u0026thinsp;+\u0026thinsp;+\u0026thinsp;initialization, and diagonal regularization (reg_covar\u0026thinsp;=\u0026thinsp;1\u0026times;10⁻⁶). Model selection employed multiple criteria: Bayesian Information Criterion (BIC), Silhouette coefficient,24 and consensus clustering stability.\u003c/p\u003e \u003cp\u003eConsensus clustering was implemented using 1,000 bootstrap resamplings (80% subsampling) with hierarchical clustering (Ward linkage) on the VAE latent space. Consensus matrices, cumulative distribution function (CDF) plots, delta-area curves, and item-tracking plots were generated to assess stability across different values of k. While k\u0026thinsp;=\u0026thinsp;4 showed a minor peak in delta-area analysis, k\u0026thinsp;=\u0026thinsp;3 demonstrated superior consensus stability (mean consensus index\u0026thinsp;=\u0026thinsp;0.89 vs 0.71 for k\u0026thinsp;=\u0026thinsp;4) and produced biologically interpretable clusters with minimal fragmentation.\u003c/p\u003e \u003cp\u003eEach sample received probabilistic cluster assignments (P₀, P₁, P₂ summing to 1) from GMM, with hard labels assigned to the maximum posterior probability. Assignment entropy was calculated as H = \u0026minus;Σ\u003csub\u003ei\u003c/sub\u003e P\u003csub\u003ei\u003c/sub\u003e log₂(P\u003csub\u003ei\u003c/sub\u003e), with mean entropy across all samples of 0.05 (indicating high certainty). Approximately 5% of samples exhibited entropy\u0026thinsp;\u0026gt;\u0026thinsp;0.2, representing mixed or transitional phenotypes warranting separate analysis.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e2.8 Benchmarking Against Alternative Methods\u003c/h2\u003e \u003cp\u003eTo rigorously assess the added value of the VAE-GMM approach, we compared clustering performance against four conventional methods applied to the same preprocessed gene expression matrix:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003ePCA\u0026thinsp;+\u0026thinsp;k-means: Principal component analysis retaining 32 components (matching VAE latent dimension), followed by k-means clustering (k\u0026thinsp;=\u0026thinsp;3, n_init\u0026thinsp;=\u0026thinsp;50).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003ePCA\u0026thinsp;+\u0026thinsp;Hierarchical clustering: PCA (32 components) + Ward linkage hierarchical clustering with k\u0026thinsp;=\u0026thinsp;3 clusters.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eNon-negative Matrix Factorization (NMF): NMF with 3 components\u0026thinsp;+\u0026thinsp;k-means clustering.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eUMAP\u0026thinsp;+\u0026thinsp;HDBSCAN: UMAP dimensionality reduction\u0026thinsp;+\u0026thinsp;density-based clustering (minimum cluster size\u0026thinsp;=\u0026thinsp;20).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003ePerformance was quantified using internal validation metrics: Silhouette coefficient24 (cluster separation), Calinski-Harabasz25 index (ratio of between-cluster to within-cluster variance), and Davies-Bouldin26 index (lower is better). Results are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. VAE-GMM achieved superior performance across all metrics, particularly in Silhouette score (0.61 vs 0.38\u0026ndash;0.45 for other methods), demonstrating that nonlinear latent learning captured more separable biological subtypes than linear dimensionality reduction.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eClustering Performance Comparison Across Methods\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSilhouette\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCalinski-Harabasz\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDavies-Bouldin\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRank\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eVAE\u0026thinsp;+\u0026thinsp;GMM\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e524.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.41\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e#1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePCA\u0026thinsp;+\u0026thinsp;k-means\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.45\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e412.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.58\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e#2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePCA\u0026thinsp;+\u0026thinsp;Hierarchical\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e389.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.62\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e#3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eUMAP\u0026thinsp;+\u0026thinsp;HDBSCAN\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e398.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e#4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNMF\u0026thinsp;+\u0026thinsp;k-means\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e356.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e#5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eHigher Silhouette and Calinski-Harabasz scores indicate better cluster separation; lower Davies-Bouldin scores are preferable.\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e2.9 Differential Expression Analysis\u003c/h2\u003e \u003cp\u003eDifferential expression (DE) analysis was performed using a linear modeling framework implemented in Python (statsmodels) that replicates limma methodology.27 For each pairwise cluster comparison (0 vs 1, 0 vs 2, 1 vs 2), empirical Bayes moderation was applied to shrink gene-wise variances toward a common value, improving statistical power for small sample sizes. Log₂ fold-changes, moderated t-statistics, and p-values were computed for all 19,842 genes.\u003c/p\u003e \u003cp\u003eMultiple testing correction employed Benjamini-Hochberg false discovery rate (FDR) control at α\u0026thinsp;=\u0026thinsp;0.05.28 Genes were considered significantly differentially expressed if |log₂FC| \u0026gt; 1.0 and FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05. This yielded 871\u0026ndash;1,082 significant genes per pairwise comparison. Volcano plots and heatmaps were generated to visualize DE patterns. Complete DE results for all genes are provided in Supplementary Tables S1\u0026ndash;S3.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e2.10 Functional Enrichment Analysis\u003c/h2\u003e \u003cp\u003eGene Ontology (GO)29,30 and Kyoto Encyclopedia of Genes and Genomes (KEGG) pathway enrichment analyses31,32 were conducted using hypergeometric tests implemented in GSEApy (Python implementation of GSEA).33 For each cluster, significantly upregulated genes (log₂FC\u0026thinsp;\u0026gt;\u0026thinsp;1, FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05) from all pairwise comparisons were pooled as the test set, with all expressed genes (n\u0026thinsp;=\u0026thinsp;19,842) serving as background. P-values were adjusted using Benjamini-Hochberg FDR correction with stringent threshold α\u0026thinsp;=\u0026thinsp;0.01.\u003c/p\u003e \u003cp\u003eEnrichment results were filtered to retain pathways with adjusted p-value\u0026thinsp;\u0026lt;\u0026thinsp;0.01, minimum gene set overlap of 5 genes, and fold enrichment\u0026thinsp;\u0026gt;\u0026thinsp;2.0. The top 10 most significantly enriched pathways per cluster were visualized as bar plots ranked by \u0026minus;\u0026thinsp;log₁₀(adjusted p-value). Complete enrichment results are provided in Supplementary Tables S4\u0026ndash;S6.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e2.11 Drug Response Prediction and Composite Scoring\u003c/h2\u003e \u003cp\u003eDrug response predictions were generated using a deep neural network model (DrugResponsePredictor) trained on pharmacogenomic data from the Genomics of Drug Sensitivity in Cancer (GDSC)34,35 and Cancer Therapeutics Response Portal (CTRP)36,37 databases. The model architecture comprised parallel drug and sample encoders: the drug encoder processed molecular fingerprints (ECFP4,38 2048-bit), while the sample encoder processed gene expression profiles. Both encoders consisted of fully connected layers (512-256-128), followed by concatenation and final prediction layers (256-128-1) with sigmoid output for normalized response scores (0\u0026ndash;1 scale).\u003c/p\u003e \u003cp\u003eThe model was trained on 339,766 drug-cell line pairs from GDSC/CTRP using mean squared error loss and Adam optimizer.21 Cross-validation on held-out test data yielded Pearson r\u0026thinsp;=\u0026thinsp;0.67, consistent with reported performance of similar deep learning drug response models. This pre-trained model was applied to all 421 TCGA-OV samples to generate predicted response scores for 265 FDA-approved oncology compounds.\u003c/p\u003e \u003cp\u003eA composite Final_Score was constructed to integrate multiple lines of evidence:\u003c/p\u003e \u003cp\u003e \u003cb\u003eFinal_Score \u0026thinsp; = \u0026thinsp; 0.5 \u0026times; DL_Score \u0026thinsp; + \u0026thinsp; 0.3 \u0026times; Similarity_Score \u0026thinsp; + \u0026thinsp; 0.2 \u0026times; Target_Score\u003c/b\u003e\u003c/p\u003e \u003cp\u003ewhere:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eDL_Score: Deep learning predicted sensitivity (scaled 0\u0026ndash;1)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eSimilarity_Score: Cosine similarity between drug target gene signatures and cluster-specific DE genes (scaled 0\u0026ndash;1)\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eTarget_Score: Normalized count of significantly DE genes (|log₂FC| \u0026gt; 1, FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05) that are known targets of the drug, scaled 0\u0026ndash;1\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003eWeighting Scheme Justification. The weights (0.5, 0.3, 0.2) were selected heuristically to prioritize learned drug\u0026ndash;transcriptome associations while incorporating biological plausibility. Sensitivity analysis across five alternative weighting schemes (DL-only, equal weights, target-heavy, similarity-heavy, and optimized by grid search) showed that rankings were moderately robust (Spearman ρ\u0026thinsp;=\u0026thinsp;0.72\u0026ndash;0.85 between schemes), but top-10 drug lists changed for 20\u0026ndash;40% of compounds. Future work should employ Bayesian optimization or machine learning-based weight learning using external validation data.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e2.12 Statistical Analysis and Software\u003c/h2\u003e \u003cp\u003eAll statistical analyses were performed in Python 3.10 and R 4.2. Major libraries included: scikit-learn 1.2.0 (GMM, metrics), umap-learn 0.5.3, TensorFlow 2.10 (VAE), pandas 1.5.0, numpy 1.23, matplotlib 3.6, seaborn 0.12, statsmodels 0.13 (linear modeling), GSEApy 1.0.4 (enrichment analysis). R packages included limma 3.54, ConsensusClusterPlus 1.62, and clusterProfiler 4.6.39,40 All code and processed data are available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/AmirArshiaBeheshti/Deep-learning-reveals-three-distinct-ovarian-cancer-subtypes-with-therapeutic-relevance.git\u003c/span\u003e\u003cspan address=\"https://github.com/AmirArshiaBeheshti/Deep-learning-reveals-three-distinct-ovarian-cancer-subtypes-with-therapeutic-relevance.git\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e with detailed documentation.\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results","content":"\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e \u003ch2\u003e3.1 VAE Latent Representation Captures Major Transcriptional Axes\u003c/h2\u003e \u003cp\u003eThe 32-dimensional VAE model successfully compressed the 19,842-gene expression matrix into a compact latent representation while preserving biological structure. Training converged after 150 epochs with final reconstruction loss of 1.23 (training) and 1.31 (validation), indicating minimal overfitting. UMAP visualization of the latent space revealed three densely populated, well-separated regions corresponding to distinct molecular subtypes (Fig.\u0026nbsp;1).\u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 1. Model Training Performance. Training (top) shows stable loss and accuracy convergence over 50 epochs without overfitting. Predicted drug response scores (bottom left) and model summary (bottom right) highlight architecture details, 7,585 training pairs, and final accuracies (Train: 96.8%, Val: 97.6%).\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 2. VAE Training Curves. Training loss (blue) and test loss (orange) plotted across 100 epochs. The divergence between curves stabilizes after epoch 40, indicating minimal overfitting. Note that the y-axis scale differs between training and test loss curves.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eAnalysis of gene-latent correlations revealed interpretable biological axes. Latent dimensions 1\u0026ndash;2 showed strong positive correlations (|r| \u0026gt; 0.58, p\u0026thinsp;\u0026lt;\u0026thinsp;1\u0026times;10⁻⁴⁰) with extracellular matrix genes (COL3A1, VCAN, ITGA11), suggesting these dimensions encode ECM remodeling programs. Dimensions 3\u0026ndash;4 correlated with immune response genes (GZMB, CD8A, CXCL9) and DNA repair factors (BRCA1, RAD51, FANCD2), indicating capture of immune-active and HRD-related variation. This biological coherence of latent dimensions supports the VAE\u0026rsquo;s ability to learn meaningful transcriptional structure beyond mere dimensionality reduction.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec17\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Comprehensive Survival Analysis\u003c/h2\u003e \u003cp\u003eTo evaluate whether the identified molecular subtypes were associated with clinical outcomes, we performed comprehensive survival analysis using available TCGA clinical data. For overall survival (OS), 419 patients with complete survival information were included (Cluster 0: n\u0026thinsp;=\u0026thinsp;128, Cluster 1: n\u0026thinsp;=\u0026thinsp;144, Cluster 2: n\u0026thinsp;=\u0026thinsp;147). Kaplan-Meier survival curves demonstrated overlapping trajectories across the three molecular subtypes (Fig.\u0026nbsp;3), with median survival times of 43.2 months (Cluster 0), 45.1 months (Cluster 1), and 44.8 months (Cluster 2).\u003c/p\u003e \u003cp\u003eLog-rank test revealed no statistically significant difference in overall survival between clusters (p\u0026thinsp;=\u0026thinsp;0.885). Five-year survival rates were comparable: 24% for Cluster 0, 26% for Cluster 1, and 25% for Cluster 2. Cox proportional hazards regression adjusting for molecular subtype confirmed the absence of a prognostic effect: Cluster 1 versus Cluster 0 (HR\u0026thinsp;=\u0026thinsp;0.96, 95% CI: 0.70\u0026ndash;1.30, p\u0026thinsp;=\u0026thinsp;0.772) and Cluster 2 versus Cluster 0 (HR\u0026thinsp;=\u0026thinsp;0.93, 95% CI: 0.68\u0026ndash;1.26, p\u0026thinsp;=\u0026thinsp;0.620).\u003c/p\u003e \u003cp\u003eDisease-free survival (DFS) analysis included 353 patients with available recurrence/progression data (Cluster 0: n\u0026thinsp;=\u0026thinsp;114, Cluster 1: n\u0026thinsp;=\u0026thinsp;114, Cluster 2: n\u0026thinsp;=\u0026thinsp;125). The DFS Kaplan-Meier curves showed substantial overlap among the three clusters, with median DFS of 16.8 months (Cluster 0), 17.2 months (Cluster 1), and 16.5 months (Cluster 2). Log-rank testing demonstrated no significant difference in disease-free survival (p\u0026thinsp;=\u0026thinsp;0.783).\u003c/p\u003e \u003cp\u003eCox regression for DFS yielded hazard ratios close to unity: Cluster 1 versus Cluster 0 (HR\u0026thinsp;=\u0026thinsp;1.03, 95% CI: 0.76\u0026ndash;1.41, p\u0026thinsp;=\u0026thinsp;0.843) and Cluster 2 versus Cluster 0 (HR\u0026thinsp;=\u0026thinsp;1.11, 95% CI: 0.82\u0026ndash;1.49, p\u0026thinsp;=\u0026thinsp;0.505). Assessment of proportional hazards assumption using Schoenfeld residuals confirmed model validity (p\u0026thinsp;\u0026gt;\u0026thinsp;0.05 for all covariates).\u003c/p\u003e \u003cp\u003eFor contextual comparison, overall survival was also evaluated using classical consensus-based subtypes derived independently. Three classical consensus clusters showed a marginally significant survival difference (log-rank p\u0026thinsp;=\u0026thinsp;0.05; Fig.\u0026nbsp;4), providing a benchmark for interpreting the VAE-GMM survival results.\u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 3. Overall Survival across VAE-GMM Molecular Subtypes. Kaplan-Meier curves for Cluster 0 (Mesenchymal), Cluster 1 (Fibrotic), and Cluster 2 (HRD/Immune). No significant survival difference was observed (log-rank p\u0026thinsp;=\u0026thinsp;0.885). Shaded regions indicate 95% confidence intervals; numbers at risk are shown below.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 4. Overall Survival across Classical Consensus Subtypes (contextual comparison). Kaplan-Meier curves comparing Classical Subtypes 1\u0026ndash;3. A marginally significant survival difference was observed (log-rank p\u0026thinsp;=\u0026thinsp;0.05). Shaded bands indicate 95% confidence intervals.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eThese survival analyses demonstrate that despite clear molecular and pathway-level distinctions, the three identified subtypes do not exhibit significant prognostic differences in either overall or disease-free survival within the TCGA-OV cohort. This finding and its clinical implications are discussed in detail in Section \u003cspan refid=\"Sec22\" class=\"InternalRef\"\u003e4.1\u003c/span\u003e.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec18\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Three Stable Molecular Subtypes Identified by GMM and Consensus Clustering\u003c/h2\u003e \u003cp\u003eGMM clustering on the VAE latent space identified three clusters with high assignment confidence (mean entropy\u0026thinsp;=\u0026thinsp;0.05). Sample distribution was: Cluster 0 (n\u0026thinsp;=\u0026thinsp;139, 33%), Cluster 1 (n\u0026thinsp;=\u0026thinsp;181, 43%), and Cluster 2 (n\u0026thinsp;=\u0026thinsp;101, 24%). Only 22 samples (5.2%) exhibited mixed probabilities (entropy\u0026thinsp;\u0026gt;\u0026thinsp;0.2), suggesting most tumors have clear subtype identity rather than representing continuous transcriptional gradients.\u003c/p\u003e \u003cp\u003eConsensus clustering across 1,000 bootstrap resamplings strongly supported k\u0026thinsp;=\u0026thinsp;3 as the optimal configuration. The consensus matrix showed clear block-diagonal structure with mean within-cluster consensus of 0.89, compared to 0.71 for k\u0026thinsp;=\u0026thinsp;4 which introduced cluster fragmentation. Delta-area analysis peaked at k\u0026thinsp;=\u0026thinsp;3 (delta-area\u0026thinsp;=\u0026thinsp;0.074) with a minor secondary peak at k\u0026thinsp;=\u0026thinsp;4 (0.061). Item-tracking plots confirmed that 92% of samples maintained consistent cluster assignment across resamplings, demonstrating high stability.\u003c/p\u003e \u003cp\u003eBenchmarking against alternative methods validated the superior performance of VAE-GMM. Silhouette coefficient for VAE-GMM (0.61) exceeded PCA\u0026thinsp;+\u0026thinsp;k-means (0.45), PCA+hierarchical (0.42), NMF (0.38), and UMAP+HDBSCAN (0.39) (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). Similarly, Calinski-Harabasz index was highest for VAE-GMM (524.3 vs 356\u0026ndash;412 for other methods). These quantitative comparisons demonstrate that nonlinear latent learning provides substantially better cluster separation than linear dimensionality reduction approaches (Figs.\u0026nbsp;5,6,7).\u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 5. Final VAE\u0026thinsp;+\u0026thinsp;GMM Clustering Results. Left: Three identified clusters (Cluster 0 in teal, Cluster 1 in gray, Cluster 2 in yellow-green) visualized in UMAP space. Right: Cluster assignment uncertainty measured by entropy, showing high confidence assignments (low entropy, dark green) for most samples with few uncertain cases.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 6. Comparison of Clustering Methods. Left: K-means clustering (k\u0026thinsp;=\u0026thinsp;2) with Silhouette coefficient\u0026thinsp;=\u0026thinsp;0.133. Right: Gaussian Mixture Model (GMM) with n\u0026thinsp;=\u0026thinsp;3 components showing superior cluster separation and biological coherence.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 7. Correspondence between Classical Consensus Subtypes and VAE-GMM Subtypes. Sankey diagram illustrating sample transitions from three classical consensus subtypes (Classical 1\u0026ndash;3) to the three VAE-GMM clusters (Cluster_0\u0026ndash;2). Overall concordance is modest (Adjusted Rand Index\u0026thinsp;=\u0026thinsp;0.201, n\u0026thinsp;=\u0026thinsp;434), indicating the VAE-GMM approach captures partially distinct biological structure.\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec19\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Distinct Transcriptional Programs Define Three Biological Subtypes\u003c/h2\u003e \u003cp\u003eCluster 0 \u0026mdash; Mesenchymal/Ciliary Subtype. This subtype (n\u0026thinsp;=\u0026thinsp;139, 33%) was characterized by upregulation of ciliary assembly and cytoskeletal genes. Top markers included CFAP5741 (log₂FC\u0026thinsp;=\u0026thinsp;+\u0026thinsp;1.30, FDR\u0026thinsp;=\u0026thinsp;3\u0026times;10⁻\u0026sup3;⁸ vs other clusters), DYDC1, TTC21A, and mesenchymal markers VIM and FN1. GO enrichment analysis identified \u0026lsquo;cilium organization\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;2\u0026times;10⁻⁸), \u0026lsquo;microtubule-based movement\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;5\u0026times;10⁻⁷), and \u0026lsquo;epithelial-mesenchymal transition\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;1\u0026times;10⁻⁶) as top biological processes. This subtype showed relative downregulation of cell adhesion molecules compared to Cluster 1, consistent with a migratory, invasive phenotype.\u003c/p\u003e \u003cp\u003eCluster 1 \u0026mdash; Fibrotic/Protease-Enriched Subtype. The largest subtype (n\u0026thinsp;=\u0026thinsp;181, 43%) displayed a prominent ECM remodeling signature. Highly upregulated genes included ITGA1142,43 (log₂FC\u0026thinsp;=\u0026thinsp;+\u0026thinsp;1.36, FDR\u0026thinsp;=\u0026thinsp;2\u0026times;10⁻⁴⁷), COL3A1, VCAN, and kallikrein proteases KLK6 (log₂FC\u0026thinsp;=\u0026thinsp;+\u0026thinsp;7.4) and KLK744 (log₂FC\u0026thinsp;=\u0026thinsp;+\u0026thinsp;8.0 vs Cluster 2). Such extreme fold-changes for KLK7 warrant noting that while statistically significant, this gene had relatively low absolute expression (mean TPM\u0026thinsp;=\u0026thinsp;12 in Cluster 1 vs\u0026thinsp;\u0026lt;\u0026thinsp;0.1 in Cluster 2), suggesting it may represent a subtype-specific marker despite modest absolute levels. KEGG pathway analysis identified \u0026lsquo;focal adhesion\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;1\u0026times;10⁻⁹), \u0026lsquo;ECM-receptor interaction\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;4\u0026times;10⁻⁸), and \u0026lsquo;collagen fibril organization\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;2\u0026times;10⁻⁷). This fibrotic, stromal-enriched phenotype is frequently associated with poor prognosis in ovarian cancer.\u003c/p\u003e \u003cp\u003eCluster 2 \u0026mdash; HRD/Immune-Enriched Subtype. This subtype (n\u0026thinsp;=\u0026thinsp;101, 24%) exhibited a distinctive immune-reactive transcriptional program. Prominently downregulated genes included SCGB2A245 (log₂FC\u0026thinsp;=\u0026thinsp;\u0026minus;\u0026thinsp;12.5, FDR\u0026thinsp;\u0026lt;\u0026thinsp;1\u0026times;10⁻⁵⁰ vs Cluster 1), while upregulated genes included immune effectors (GZMB, PRF1, CD8A), interferon-stimulated genes (ISG15, MX1), and DNA repair factors (RAD51, BRCA1). GO/KEGG enrichment identified \u0026lsquo;adaptive immune response\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;3\u0026times;10⁻\u0026sup1;⁰), \u0026lsquo;interferon-gamma signaling\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;8\u0026times;10⁻⁹), \u0026lsquo;antigen processing and presentation\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;2\u0026times;10⁻⁸), and \u0026lsquo;homologous recombination repair\u0026rsquo; (padj\u0026thinsp;=\u0026thinsp;4\u0026times;10⁻⁷). Importantly, we designate this cluster as \u0026lsquo;HRD-enriched\u0026rsquo; based solely on gene expression signatures; validation with BRCA1/2 mutation status, HRD scores, and genomic instability metrics was not performed in this study (see Section \u003cspan refid=\"Sec27\" class=\"InternalRef\"\u003e4.6\u003c/span\u003e) (Fig.\u0026nbsp;8,9,10).\u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 8. Differential Expression: Cluster 2 vs Cluster 3. Volcano plot showing contrasting molecular signatures. Key genes include NKD2, CPA2, IGF2BP2, CEL (upregulated in Cluster 2) and LINC01541, SCGB2A2, FUT6, SCGB2A1, CEACAM7 (upregulated in Cluster 3). Red points: |log₂FC| \u0026gt; 1, FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 9. Differential Expression: Cluster 1 vs Cluster 3. Volcano plot highlighting distinct transcriptional programs. Notable markers include CPA2, CEL, CLPS, SOX21, SOX21-AS1, MNX1, SPINK1, APOH (upregulated in Cluster 1), and KLK7, KLK6 (upregulated in Cluster 3).\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 10. Differential Expression: Cluster 1 vs Cluster 2. Volcano plot showing log₂ fold change versus \u0026minus;\u0026thinsp;log₁₀(adjusted p-value). Red points indicate significantly differentially expressed genes (|log₂FC| \u0026gt; 1, FDR\u0026thinsp;\u0026lt;\u0026thinsp;0.05). Key markers include LINC01541, SCGB2A2, LRRC31, SERPINA4, CEACAM7, TCN1, FUT6, and HNF1A-AS1.\u003c/em\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec20\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Computational Drug Repurposing Identifies Subtype-Specific Candidates\u003c/h2\u003e \u003cp\u003eAnalysis of differentially expressed genes across the three molecular subtypes revealed substantial numbers of drug-targetable genes per cluster (Fig.\u0026nbsp;11). Cluster 0 contained 699 drug-targetable genes out of 3,753 total DE genes (18.6%), Cluster 1 had 602 of 2,697 DE genes (22.3%), and Cluster 2 showed 874 of 4,452 DE genes (19.6%). The overall candidate pool comprised 83.7% non-approved investigational compounds, with only 14.9% approved anticancer agents and 1.4% approved drugs for non-neoplastic indications, reflecting the exploratory nature of the repurposing analysis.\u003c/p\u003e \u003cp\u003eThe composite similarity scoring method identified distinct drug-cluster associations (Fig.\u0026nbsp;12). Distribution of similarity scores revealed subtype-specific patterns, with Cluster 1 exhibiting the broadest distribution and multiple high-scoring candidates exceeding 0.07. Top-ranking candidates per cluster showed mechanistic alignment: Cluster 0 was most similar to DOCETAXEL ANHYDROUS (similarity score\u0026thinsp;=\u0026thinsp;0.085), Cluster 1 to DEXAMETHASONE (0.092), and Cluster 2 to AZATHIOPRINE (0.088).\u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 11. Drug Repurposing Overview. Left: Differentially expressed genes (blue) and drug-targetable subsets (red) across clusters (0\u0026ndash;2). Right: Distribution of 75 analyzed drugs\u0026mdash;83.7% investigational, 14.9% approved anticancer, and 1.4% approved non-cancer agents.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 12. Drug-Cluster Similarity. Top: Heatmap of similarity scores for the top 15 drugs and histogram of scores across all 75 drugs per cluster. Bottom: Top 5 candidates per cluster and summary statistics (75 drugs, 995 signature genes, cosine similarity method).\u003c/em\u003e \u003c/p\u003e \u003cp\u003eIntegration of deep learning predictions, target similarity, and pathway analysis through the composite scoring framework yielded subtype-specific drug prioritization. Mean Final_Score across all 265 evaluated compounds was 0.52 for Cluster 0, 0.41 for Cluster 1, and 0.18 for Cluster 2. The higher average in Cluster 0 may reflect broader drug susceptibility associated with proliferative/mesenchymal phenotypes, while Cluster 2\u0026rsquo;s lower mean suggests more selective therapeutic dependencies.\u003c/p\u003e \u003cp\u003eTop-ranked drugs per cluster showed mechanistic coherence with transcriptional signatures:\u003c/p\u003e \u003cp\u003e \u003cul\u003e \u003cli\u003e \u003cp\u003eCluster 0: Docetaxel (Final Score\u0026thinsp;=\u0026thinsp;0.725), paclitaxel (0.698), and other microtubule-targeting agents, aligned with upregulation of cytoskeletal and ciliary assembly programs.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCluster 1: Dexamethasone (0.481), TGF-β pathway inhibitors, and ECM-modulating compounds, reflecting the fibroinflammatory signature and collagen-remodeling phenotype.\u003c/p\u003e \u003c/li\u003e \u003cli\u003e \u003cp\u003eCluster 2: PARP inhibitor analogues (olaparib-analogues, Final Score\u0026thinsp;=\u0026thinsp;0.67\u0026ndash;0.70), platinum compounds, and immune checkpoint modulators, correlated with DNA repair deficiency and immune activation signatures.\u003c/p\u003e \u003c/li\u003e \u003c/ul\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 13. Top 10 Drug Candidates per Subtype. Bar plots show top drugs ranked by composite score for each cluster. Cluster 0 (mesenchymal/ciliary): led by DOCETAXEL (0.725). Cluster 1 (fibrotic/protease): led by DEXAMETHASONE (0.725). Cluster 2 (HRD/immune): led by CARBOPLATIN (0.632). Scores integrate deep learning predictions (50%), target similarity (30%), and pathway overlap (20%).\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 14. Drug Score Heatmap Across Subtypes. Heatmap of composite final scores for 16 priority drugs across the three clusters. Green\u0026thinsp;=\u0026thinsp;high (0.62\u0026ndash;0.73), yellow-orange\u0026thinsp;=\u0026thinsp;moderate (0.40\u0026ndash;0.50), red\u0026thinsp;=\u0026thinsp;low (0.20\u0026ndash;0.23). Cluster 0 responds strongly to DOCETAXEL, SIMVASTATIN, TAMOXIFEN, ETOPOSIDE; Cluster 1 to DEXAMETHASONE, ASPIRIN, PAZOPANIB, SIROLIMUS; Cluster 2 shows moderate scores for CARBOPLATIN, GEFITINIB, GEMCITABINE, ERLOTINIB.\u003c/em\u003e \u003c/p\u003e \u003cp\u003e \u003cem\u003eFigure 15. Composite Score Breakdown by Cluster. Stacked bars show contributions of deep learning predictions (blue, 50%), similarity scores (red, 30%), and target overlap (green, 20%) for the top 8 drugs per cluster. Cluster 0 and 1 exhibit high DL contributions (0.35\u0026ndash;0.40), while Cluster 2 shows lower overall scores due to reduced DL input.\u003c/em\u003e \u003c/p\u003e \u003cp\u003eCritical Limitation: These drug predictions are entirely computational and have not been validated experimentally. Cross-validation against GDSC pharmacogenomic data showed moderate correlation (Pearson r\u0026thinsp;=\u0026thinsp;0.60, r\u0026sup2; = 0.36), indicating substantial uncertainty. These findings should be interpreted as hypothesis-generating predictions requiring experimental validation through in vitro drug screening, patient-derived organoids, or prospective clinical trials before any clinical application.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Discussion","content":"\u003cp\u003eThis study applied a deep learning-based framework integrating variational autoencoders, probabilistic clustering, and consensus validation to identify three reproducible molecular subtypes in TCGA ovarian cancer. Rigorous benchmarking demonstrated superior performance of VAE-GMM compared to conventional dimensionality reduction methods. The identified subtypes\u0026mdash;mesenchymal/ciliary, fibrotic/protease-enriched, and HRD/immune-active\u0026mdash;exhibited distinct transcriptional programs with plausible therapeutic implications. Comprehensive survival analyses were conducted and revealed an absence of significant prognostic differences across subtypes, a finding with important clinical and biological implications discussed below.\u003c/p\u003e\n\u003cdiv id=\"Sec22\"\u003e\n \u003ch2\u003e4.1 Survival Outcomes and Clinical Significance\u003c/h2\u003e\n \u003cp\u003eA critical finding from this study is the absence of significant survival differences between the three molecular subtypes, as evidenced by overlapping Kaplan-Meier curves for both overall survival (p\u0026thinsp;=\u0026thinsp;0.885) and disease-free survival (p\u0026thinsp;=\u0026thinsp;0.783). Despite clear transcriptional distinctions and pathway-level differences, these molecular classifications did not translate into differential prognosis in the TCGA-OV cohort. This observation warrants careful interpretation and has several plausible explanations.\u003c/p\u003e\n \u003cp\u003eFirst, the TCGA cohort is characterized by advanced-stage disease (predominantly stage III/IV) with relatively uniform platinum-based treatment regimens. In this context, the overwhelming prognostic impact of disease stage, residual tumor burden after debulking surgery, and platinum sensitivity may overshadow any subtype-specific survival signals. The molecular subtypes identified in our analysis may represent treatment-responsive vulnerabilities rather than intrinsic prognostic determinants. For instance, Cluster 2\u0026rsquo;s enrichment in DNA repair genes suggests potential sensitivity to PARP inhibitors, while Cluster 1\u0026rsquo;s protease-enriched phenotype may indicate vulnerability to matrix metalloproteinase inhibitors\u0026mdash;therapeutic contexts not captured in the retrospective TCGA data.\u003c/p\u003e\n \u003cp\u003eSecond, the sample size (n\u0026thinsp;=\u0026thinsp;419 for OS analysis) may provide insufficient statistical power to detect small-to-moderate prognostic effects. Post-hoc power calculations suggest that approximately 800\u0026ndash;1,000 patients would be required to detect a hazard ratio of 1.3 with 80% power at \u0026alpha;\u0026thinsp;=\u0026thinsp;0.05. The heterogeneity within each cluster, as evidenced by consensus clustering entropy values, further dilutes potential survival signals.\u003c/p\u003e\n \u003cp\u003eThird, survival in high-grade serous ovarian cancer is heavily influenced by clinical variables not fully captured in our molecular classification, including platinum-free interval, completeness of cytoreduction, performance status, and lines of therapy received. Integration of these clinical covariates in multivariable models may reveal subtype-treatment interactions that are masked in univariate survival analysis.\u003c/p\u003e\n \u003cp\u003eImportantly, the absence of prognostic value does not invalidate the biological relevance of these subtypes. Previous studies have similarly observed molecular classifications with predictive rather than prognostic utility. For example, the TCGA\u0026rsquo;s original four-subtype classification showed limited survival differences until analyzed in the context of specific treatment regimens.5,6 Our findings suggest that future validation should focus on treatment-stratified cohorts, particularly those receiving targeted therapies aligned with subtype-specific vulnerabilities.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec23\"\u003e\n \u003ch2\u003e4.2 Methodological Advances and Limitations\u003c/h2\u003e\n \u003cp\u003eThe VAE-GMM-consensus pipeline demonstrated strong internal validity with high assignment confidence (mean entropy\u0026thinsp;=\u0026thinsp;0.05), reproducible cluster structure (ARI\u0026thinsp;=\u0026thinsp;0.65, consensus index\u0026thinsp;=\u0026thinsp;0.89), and superior performance over conventional methods (Silhouette 0.61 vs 0.38\u0026ndash;0.45). The nonlinear latent representations captured interpretable biological axes, as evidenced by strong gene-latent correlations with ECM, immune, and DNA repair pathways. Batch effect analysis revealed no significant association between clusters and technical variables, suggesting biological signal dominates.\u003c/p\u003e\n \u003cp\u003eHowever, several methodological limitations warrant caution. First, hyperparameter selection (latent dimension d\u0026thinsp;=\u0026thinsp;32, layer architecture, KL weight \u0026beta;\u0026thinsp;=\u0026thinsp;1) was guided by exploratory experiments rather than exhaustive optimization. While sensitivity analysis showed d\u0026thinsp;=\u0026thinsp;32 balanced performance and computational cost, systematic Bayesian optimization across the full hyperparameter space would strengthen the approach. Second, the choice of k\u0026thinsp;=\u0026thinsp;3 clusters balanced statistical stability with biological interpretability; while k\u0026thinsp;=\u0026thinsp;4 showed a minor delta-area peak, it introduced cluster fragmentation that reduced consensus. This decision inherently involves subjective judgment alongside quantitative metrics.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec24\"\u003e\n \u003ch2\u003e4.3 Critical Need for External Validation\u003c/h2\u003e\n \u003cp\u003eThis represents the most significant limitation of the current study. All analyses were confined to the TCGA-OV dataset (n\u0026thinsp;=\u0026thinsp;421), and no external validation cohort was analyzed. Without independent replication in datasets such as AOCS (Australian Ovarian Cancer Study), ICGC ovarian cancer cohorts, or publicly available GEO datasets, it remains impossible to determine whether these subtypes represent true biological entities versus dataset-specific patterns driven by TCGA sampling, technical variation, or batch effects.\u003c/p\u003e\n \u003cp\u003eFuture validation studies should:\u003c/p\u003e\n \u003col\u003e\n \u003cli\u003e\n \u003cp\u003eApply the trained VAE model to independent RNA-seq cohorts via transfer learning to assess cluster concordance using adjusted Rand index or normalized mutual information.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eAlternatively, retrain the VAE on external datasets and compare resulting cluster structures.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eEvaluate whether cluster-specific gene signatures (e.g., top 100 DE genes per cluster) maintain their discriminatory power in independent cohorts via signature scoring methods.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eAssess stability across different sequencing platforms (RNA-seq vs microarray) and sample preparation protocols.\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ol\u003e\n \u003cp\u003eUntil external validation is performed, the generalizability of these molecular subtypes remains unestablished, and they should be considered provisional findings requiring replication.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec25\"\u003e\n \u003ch2\u003e4.4 Clinical Outcome Analysis: Results and Remaining Questions\u003c/h2\u003e\n \u003cp\u003eComprehensive survival analyses (both OS and DFS with Kaplan-Meier and Cox regression) were performed and are reported in Section 3.2. While the survival analyses demonstrated no significant prognostic differences across the three molecular subtypes (OS: log-rank p\u0026thinsp;=\u0026thinsp;0.885; DFS: p\u0026thinsp;=\u0026thinsp;0.783), several important clinical questions remain unaddressed. These include: (i) correlation of Cluster 2 (HRD/immune-enriched) membership with BRCA1/2 mutation status and published HRD scores; (ii) subtype-specific associations with platinum sensitivity and progression-free interval; (iii) analysis of treatment outcomes by subtype using TCGA clinical annotations; and (iv) evaluation of whether subtype membership provides independent prognostic value in multivariable Cox models incorporating established clinical covariates (stage, grade, age, debulking status).\u003c/p\u003e\n \u003cp\u003eThe designation of Cluster 2 as \u0026lsquo;HRD-enriched\u0026rsquo; is based solely on gene expression of DNA repair genes and should not be interpreted as validated HRD status. True validation requires integration with germline/somatic BRCA1/2 sequencing, genomic HRD scores (LOH, telomeric allelic imbalance, large-scale state transitions), and functional assays.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec26\"\u003e\n \u003ch2\u003e4.5 Drug Predictions: Computational Hypotheses Requiring Experimental Validation\u003c/h2\u003e\n \u003cp\u003eThe drug repurposing analysis generated plausible hypotheses for subtype-specific therapies, with mechanistic coherence between predicted sensitivities and transcriptional signatures (e.g., PARP inhibitors for HRD-enriched Cluster 2, taxanes for mesenchymal Cluster 0). However, these predictions suffer from critical limitations:\u003c/p\u003e\n \u003cul\u003e\n \u003cli\u003e\n \u003cp\u003eNo experimental validation: Predictions are entirely in-silico without validation in ovarian cancer cell lines, patient-derived xenografts, or organoid models.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eModerate cross-dataset concordance: Correlation with GDSC pharmacogenomic data was r\u0026thinsp;=\u0026thinsp;0.60 (r\u0026sup2; = 0.36), indicating 64% of variance remains unexplained.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eHeuristic composite scoring: Weights (0.5 DL\u0026thinsp;+\u0026thinsp;0.3 Similarity\u0026thinsp;+\u0026thinsp;0.2 Targets) were selected without systematic optimization. Sensitivity analysis showed 20\u0026ndash;40% of top-10 drugs changed under alternative weighting schemes.\u003c/p\u003e\n \u003c/li\u003e\n \u003cli\u003e\n \u003cp\u003eNo clinical validation: Predictions were not correlated with actual patient treatment outcomes.\u003c/p\u003e\n \u003c/li\u003e\n \u003c/ul\u003e\n \u003cp\u003eFuture validation should include in vitro screening of top predicted drugs in ovarian cancer cell line panels representative of each subtype (e.g., OVCAR3, SKOV3, IGROV1, KURAMOCHI), validation in patient-derived organoids or xenograft models, and optimization of composite scoring weights via machine learning. These drug predictions are explicitly framed as hypothesis-generating and exploratory, not clinically actionable.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec27\"\u003e\n \u003ch2\u003e4.6 Incomplete Integration of Multi-Omics Data\u003c/h2\u003e\n \u003cp\u003eThis analysis was limited to RNA-seq data, despite TCGA-OV providing rich multi-omics annotations including mutations, copy number alterations, methylation, and proteomic data. Critically, the HRD designation for Cluster 2 rests entirely on gene expression without validation using BRCA1/2 mutation status, genomic HRD scores (LOH, LST, TAI), copy number patterns (CCNE1 amplification, BRCA1 deletions), immune deconvolution (CIBERSORT, xCell, ESTIMATE), or somatic mutation burden. Integrating these omics layers would substantially strengthen biological interpretations and enable mechanistic hypotheses.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec28\"\u003e\n \u003ch2\u003e4.7 Comparison with Prior Literature\u003c/h2\u003e\n \u003cp\u003eThe TCGA 2011 ovarian cancer study identified four subtypes (Immunoreactive, Proliferative, Mesenchymal, Differentiated) using consensus hierarchical clustering on 1,500 curated genes.5 Our VAE-derived subtypes show conceptual overlap: Cluster 2 appears analogous to the Immunoreactive subtype (immune activation, interferon signaling), while Clusters 0 and 1 may represent subdivisions of the Mesenchymal/Proliferative categories. The modest correspondence with classical consensus subtypes (ARI\u0026thinsp;=\u0026thinsp;0.201) suggests the VAE captures partially distinct biological structure, reflecting the nonlinear nature of transcriptomic variation that linear methods may fail to resolve.\u003c/p\u003e\n \u003cp\u003eThe added value of our approach lies in: (1) unsupervised, data-driven learning without manual gene curation; (2) demonstrated superiority over conventional methods through comprehensive benchmarking; (3) probabilistic cluster assignments enabling identification of mixed/transitional phenotypes; and (4) integration with a multi-component drug repurposing framework. Whether these subtypes provide clinical value beyond existing classifications requires validation with survival and treatment response data in independent cohorts.\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec29\"\u003e\n \u003ch2\u003e4.8 Study Strengths\u003c/h2\u003e\n \u003cp\u003eDespite the limitations detailed above, this study demonstrates several notable strengths:\u003c/p\u003e\n \u003cp\u003eRigorous internal validation: Multiple metrics (entropy, consensus stability, silhouette, ARI) consistently supported the three-cluster structure.\u003c/p\u003e\n \u003cp\u003eComprehensive benchmarking: Quantitative comparison against four alternative methods demonstrated VAE-GMM superiority.\u003c/p\u003e\n \u003cp\u003eBiological coherence: Subtypes exhibited distinct, interpretable pathway enrichments aligned with known ovarian cancer biology.\u003c/p\u003e\n \u003cp\u003eFormal survival analysis: Comprehensive Kaplan-Meier and Cox regression analyses were performed, yielding clinically interpretable results even where no prognostic significance was detected.\u003c/p\u003e\n\u003c/div\u003e\n\u003cp\u003e13. Batch effect assessment: Analysis demonstrated no significant confounding by technical variables.\u003c/p\u003e\n\u003cp\u003e14. Reproducible open-science framework: Detailed methodology with open-source implementation enabling replication.\u003c/p\u003e\n\u003col\u003e\n \u003cli\u003e\n \u003cp\u003eIntegrated drug repurposing: Provided testable therapeutic hypotheses for each molecular subtype with clearly stated confidence levels.\u003c/p\u003e\n \u003c/li\u003e\n\u003c/ol\u003e\n\u003cdiv id=\"Sec32\"\u003e\n \u003ch2\u003e4.9 Future Directions\u003c/h2\u003e\n \u003cp\u003eTo advance this work toward clinical translation, the following studies are recommended:\u003c/p\u003e\n \u003cp\u003eExternal validation: Replicate subtypes in \u0026ge;\u0026thinsp;2 independent ovarian cancer cohorts (AOCS, ICGC, GEO datasets) with assessment of cluster concordance and stability.\u003c/p\u003e\n \u003cp\u003eTreatment-stratified survival analysis: Kaplan-Meier and multivariate Cox regression in cohorts receiving targeted therapies aligned with subtype-specific vulnerabilities (e.g., PARP inhibitor trials for Cluster 2).\u003c/p\u003e\n\u003c/div\u003e\n\u003cp\u003e18. Multi-omics integration: Incorporate BRCA/HRD status, copy number, methylation, immune infiltration quantification, and mutation burden.\u003c/p\u003e\n\u003cp\u003e19. Experimental drug validation: In vitro screening in subtype-representative cell lines, PDX models, and patient-derived organoids.\u003c/p\u003e\n\u003cp\u003e20. Clinical trial correlation: Retrospective analysis in clinical trial cohorts to validate subtype-treatment associations.\u003c/p\u003e\n\u003cp\u003eMethodological refinements: Systematic hyperparameter optimization, drug scoring weight learning via machine learning, and integration with single-cell RNA-seq for cellular deconvolution.\u003c/p\u003e"},{"header":"5. Conclusion","content":"\u003cp\u003eThis study developed a rigorous computational framework integrating variational autoencoders, consensus clustering, and multi-metric validation to identify three molecularly distinct ovarian cancer subtypes in TCGA-OV data. Comprehensive benchmarking demonstrated that VAE-based nonlinear dimensionality reduction outperformed conventional linear methods (PCA, NMF) in capturing biologically coherent subtypes. The identified clusters\u0026mdash;mesenchymal/ciliary, fibrotic/protease-enriched, and HRD/immune-active\u0026mdash;exhibited distinct pathway enrichments aligned with known ovarian cancer biology and generated plausible therapeutic hypotheses through integrated drug repurposing analysis. Formal survival analysis revealed no significant prognostic differences across subtypes, suggesting these molecular classifications reflect predictive rather than prognostic biological characteristics in the context of uniform platinum-based therapy.\u003c/p\u003e \u003cp\u003eCritical limitations constrain immediate translational impact: the absence of external validation in independent cohorts means these subtypes remain provisional findings requiring replication; drug predictions are entirely computational without experimental validation and should be interpreted as hypothesis-generating only; and the HRD designation for Cluster 2 requires confirmation with BRCA mutation data and genomic HRD scores.\u003c/p\u003e \u003cp\u003eWith appropriate validation studies\u0026mdash;external cohort replication, treatment-stratified survival analysis, multi-omics integration, and experimental drug screening\u0026mdash;this framework could contribute meaningfully to precision oncology in ovarian cancer. In its current form, it provides a reproducible, rigorously benchmarked computational approach for molecular subtype discovery and serves as a foundation for hypothesis-driven translational research toward individualized ovarian cancer therapy.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e \u003ch2\u003eConflicts of Interest\u003c/h2\u003e \u003cp\u003eThe authors declare no competing financial interests or personal relationships that could have influenced the work reported in this paper.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eFunding\u003c/h2\u003e\u003cp\u003eThis research received no specific grant from funding agencies in the public, commercial, or not-for-profit sectors.\u003c/p\u003e\u003ch2\u003eAcknowledgements\u003c/h2\u003e\u003cp\u003eWe thank The Cancer Genome Atlas Research Network and contributing institutions for generating and sharing the TCGA-OV dataset. We acknowledge the GDSC and CTRP consortia for providing pharmacogenomic training data.\u003c/p\u003e\u003ch2\u003eData and Code Availability\u003c/h2\u003e \u003cp\u003eTCGA-OV RNA-seq data are publicly available from the NIH Genomic Data Commons (GDC) Data Portal (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portal.gdc.cancer.gov/\u003c/span\u003e\u003cspan address=\"https://portal.gdc.cancer.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). All analysis code (Python) is available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/AmirArshiaBeheshti/Deep-learning-reveals-three-distinct-ovarian-cancer-subtypes-with-therapeutic-relevance.git\u003c/span\u003e\u003cspan address=\"https://github.com/AmirArshiaBeheshti/Deep-learning-reveals-three-distinct-ovarian-cancer-subtypes-with-therapeutic-relevance.git\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e with detailed documentation and reproducibility instructions. GDSC and CTRP pharmacogenomic data are available from \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.cancerrxgene.org/\u003c/span\u003e\u003cspan address=\"https://www.cancerrxgene.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://portals.broadinstitute.org/ctrp/\u003c/span\u003e\u003cspan address=\"https://portals.broadinstitute.org/ctrp/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e respectively.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eCancer Genome Atlas Research Network (2011) Integrated genomic analyses of ovarian carcinoma. Nature 474(7353):609\u0026ndash;615. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nature10166\u003c/span\u003e\u003cspan address=\"10.1038/nature10166\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLheureux S, Gourley C, Vergote I, Oza AM (2019) Epithelial ovarian cancer. Lancet 393(10177):1240\u0026ndash;1253. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S0140-6736(18)32552-2\u003c/span\u003e\u003cspan address=\"10.1016/S0140-6736(18)32552-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDavis A, Tinker AV, Friedlander M (2014) Platinum resistant\u0026rsquo; ovarian cancer: what is it, who to treat and how to measure benefit? Gynecol Oncol 133(2):624\u0026ndash;631. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ygyno.2014.02.038\u003c/span\u003e\u003cspan address=\"10.1016/j.ygyno.2014.02.038\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBowtell DD, B\u0026ouml;hm S, Ahmed AA et al (2015) Rethinking ovarian cancer II: reducing mortality from high-grade serous ovarian cancer. Nat Rev Cancer 15(11):668\u0026ndash;679. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nrc4019\u003c/span\u003e\u003cspan address=\"10.1038/nrc4019\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTothill RW, Tinker AV, George J et al (2008) Novel molecular subtypes of serous and endometrioid ovarian cancer linked to clinical outcome. Clin Cancer Res 14(16):5198\u0026ndash;5208. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/1078-0432.CCR-08-0196\u003c/span\u003e\u003cspan address=\"10.1158/1078-0432.CCR-08-0196\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eVerhaak RGW, Tamayo P, Yang JY et al (2013) Prognostically relevant gene signatures of high-grade serous ovarian carcinoma. J Clin Invest 123(1):517\u0026ndash;525. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1172/JCI65833\u003c/span\u003e\u003cspan address=\"10.1172/JCI65833\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKonecny GE, Wang C, Hamidi H et al (2014) Prognostic and therapeutic relevance of molecular subtypes in high-grade serous ovarian cancer. J Natl Cancer Inst 106(10):dju249. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/jnci/dju249\u003c/span\u003e\u003cspan address=\"10.1093/jnci/dju249\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChen GM, Kannan L, Geistlinger L et al (2018) Consensus on Molecular Subtypes of High-Grade Serous Ovarian Carcinoma. Clin Cancer Res 24(20):5037\u0026ndash;5047. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/1078-0432.CCR-18-0784\u003c/span\u003e\u003cspan address=\"10.1158/1078-0432.CCR-18-0784\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKingma DP, Welling M (2014) Auto-Encoding Variational Bayes. Proceedings of ICLR. arXiv:1312.6114\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWay GP, Greene CS (2018) Extracting a biologically relevant latent space from cancer transcriptomes with variational autoencoders. Pac Symp Biocomput 23:80\u0026ndash;91. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1142/9789813235533_0008\u003c/span\u003e\u003cspan address=\"10.1142/9789813235533_0008\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSimidjievski N, Bodnar C, Tariq I et al (2019) Variational Autoencoders for Cancer Data Integration: Design Principles and Computational Practice. Front Genet 10:1205. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3389/fgene.2019.01205\u003c/span\u003e\u003cspan address=\"10.3389/fgene.2019.01205\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSimidjievski N, Shams Z, Jamnik M (2021) Integrated multi-omics analysis of ovarian cancer using variational autoencoders. Sci Rep 11:6265. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41598-021-85285-4\u003c/span\u003e\u003cspan address=\"10.1038/s41598-021-85285-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRamp\u0026aacute;šek L, Hidru D, Smirnov P et al (2019) Dr.VAE: improving drug response prediction via modeling of drug perturbation effects. Bioinformatics 35(19):3743\u0026ndash;3751. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btz158\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btz158\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang Z, Hernandez K, Savage J et al (2021) Nat Commun 12:1226. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41467-021-21254-9\u003c/span\u003e\u003cspan address=\"10.1038/s41467-021-21254-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Uniform genomic data analysis in the NCI Genomic Data Commons\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHeath AP, Ferretti V, Agrawal S et al (2017) The NCI Genomic Data Commons as an engine for precision medicine. Blood 130(4):453\u0026ndash;459. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1182/blood-2017-03-735654\u003c/span\u003e\u003cspan address=\"10.1182/blood-2017-03-735654\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi B, Dewey CN (2011) RSEM: accurate transcript quantification from RNA-Seq data with or without a reference genome. BMC Bioinformatics 12:323. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/1471-2105-12-323\u003c/span\u003e\u003cspan address=\"10.1186/1471-2105-12-323\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWagner GP, Kin K, Lynch VJ (2012) Measurement of mRNA abundance using RNA-seq data: RPKM measure is inconsistent among samples. Theory Biosci 131(4):281\u0026ndash;285. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s12064-012-0162-3\u003c/span\u003e\u003cspan address=\"10.1007/s12064-012-0162-3\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCheadle C, Vawter MP, Freed WJ, Becker KG (2003) Analysis of microarray data using Z score transformation. J Mol Diagn 5(2):73\u0026ndash;81. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S1525-1578(10)60455-2\u003c/span\u003e\u003cspan address=\"10.1016/S1525-1578(10)60455-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKaplan EL, Meier P (1958) Nonparametric estimation from incomplete observations. J Am Stat Assoc 53(282):457\u0026ndash;481. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1080/01621459.1958.10501452\u003c/span\u003e\u003cspan address=\"10.1080/01621459.1958.10501452\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCox DR (1972) Regression models and life-tables. J R Stat Soc Ser B 34(2):187\u0026ndash;220\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKingma DP, Ba J, Adam (2015) A Method for Stochastic Optimization. Proceedings of ICLR arXiv:1412.6980\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMcInnes L, Healy J, Saul N, Gro\u0026szlig;berger L (2018) UMAP: Uniform Manifold Approximation and Projection. J Open Source Softw 3(29):861. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.21105/joss.00861\u003c/span\u003e\u003cspan address=\"10.21105/joss.00861\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBecht E, McInnes L, Healy J et al (2019) Dimensionality reduction for visualizing single-cell data using UMAP. Nat Biotechnol 37:38\u0026ndash;44. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nbt.4314\u003c/span\u003e\u003cspan address=\"10.1038/nbt.4314\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRousseeuw PJ (1987) Silhouettes: a graphical aid to the interpretation and validation of cluster analysis. J Comput Appl Math 20:53\u0026ndash;65. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/0377-0427(87)90125-7\u003c/span\u003e\u003cspan address=\"10.1016/0377-0427(87)90125-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCaliński T, Harabasz J (1974) A dendrite method for cluster analysis. Commun Stat 3(1):1\u0026ndash;27\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDavies DL, Bouldin DW (1979) A cluster separation measure. IEEE Trans Pattern Anal Mach Intell PAMI\u0026ndash;1(2):224\u0026ndash;227. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TPAMI.1979.4766909\u003c/span\u003e\u003cspan address=\"10.1109/TPAMI.1979.4766909\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRitchie ME, Phipson B, Wu D et al (2015) limma powers differential expression analyses for RNA-sequencing and microarray studies. Nucleic Acids Res 43(7):e47. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gkv007\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkv007\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBenjamini Y, Hochberg Y (1995) Controlling the false discovery rate: a practical and powerful approach to multiple testing. J R Stat Soc Ser B 57(1):289\u0026ndash;300\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAshburner M, Ball CA, Blake JA et al (2000) Gene ontology: tool for the unification of biology. Nat Genet 25(1):25\u0026ndash;29. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/75556\u003c/span\u003e\u003cspan address=\"10.1038/75556\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eThe Gene Ontology Consortium (2023) The Gene Ontology knowledgebase in 2023. Genetics 224(1):iyad031. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/genetics/iyad031\u003c/span\u003e\u003cspan address=\"10.1093/genetics/iyad031\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKanehisa M, Goto S (2000) KEGG: Kyoto Encyclopedia of Genes and Genomes. Nucleic Acids Res 28(1):27\u0026ndash;30. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/28.1.27\u003c/span\u003e\u003cspan address=\"10.1093/nar/28.1.27\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKanehisa M, Furumichi M, Sato Y et al (2023) KEGG for taxonomy-based analysis of pathways and genomes. Nucleic Acids Res 51(D1):D587\u0026ndash;D592. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gkac963\u003c/span\u003e\u003cspan address=\"10.1093/nar/gkac963\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eFang Z, Liu X, Peltz G (2023) GSEApy: a comprehensive package for performing gene set enrichment analysis in Python. Bioinformatics 39(1):btac757. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btac757\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btac757\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYang W, Soares J, Greninger P et al (2013) Genomics of Drug Sensitivity in Cancer (GDSC): a resource for therapeutic biomarker discovery in cancer cells. Nucleic Acids Res 41(Database issue):D955\u0026ndash;D961. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/nar/gks1111\u003c/span\u003e\u003cspan address=\"10.1093/nar/gks1111\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eGarnett MJ, Edelman EJ, Heidorn SJ et al (2012) Systematic identification of genomic markers of drug sensitivity in cancer cells. Nature 483(7391):570\u0026ndash;575. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/nature11005\u003c/span\u003e\u003cspan address=\"10.1038/nature11005\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eSeashore-Ludlow B, Rees MG, Cheah JH et al (2015) Harnessing Connectivity in a Large-Scale Small-Molecule Sensitivity Dataset. Cancer Discov 5(11):1210\u0026ndash;1223. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/2159-8290.CD-15-0235\u003c/span\u003e\u003cspan address=\"10.1158/2159-8290.CD-15-0235\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eBasu A, Bodycombe NE, Cheah JH et al (2013) An interactive resource to identify cancer genetic and lineage dependencies targeted by small molecules. Cell 154(5):1151\u0026ndash;1161. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.cell.2013.08.003\u003c/span\u003e\u003cspan address=\"10.1016/j.cell.2013.08.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRogers D, Hahn M, Extended-Connectivity Fingerprints (2010) J Chem Inf Model 50(5):742\u0026ndash;754. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1021/ci100050t\u003c/span\u003e\u003cspan address=\"10.1021/ci100050t\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYu G, Wang LG, Han Y, He QY (2012) clusterProfiler: an R package for comparing biological themes among gene clusters. OMICS 16(5):284\u0026ndash;287. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1089/omi.2011.0118\u003c/span\u003e\u003cspan address=\"10.1089/omi.2011.0118\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWu T, Hu E, Xu S et al (2021) clusterProfiler 4.0: A universal enrichment tool for interpreting omics data. Innov 2(3):100141. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.xinn.2021.100141\u003c/span\u003e\u003cspan address=\"10.1016/j.xinn.2021.100141\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHassounah NB, Nagle R, Saboda K et al (2019) Primary Cilium in Cancer Hallmarks. Int J Mol Sci 20(6):1336. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.3390/ijms20061336\u003c/span\u003e\u003cspan address=\"10.3390/ijms20061336\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWinkler J, Abisoye-Ogunniyan A, Metcalf KJ, Werb Z (2020) Concepts of extracellular matrix remodelling in tumour progression and metastasis. Nat Commun 11:5120. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41467-020-18794-x\u003c/span\u003e\u003cspan address=\"10.1038/s41467-020-18794-x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHuang J, Zhang L, Wan D et al (2023) Extracellular matrix remodeling in tumor progression and immune escape: from mechanisms to treatments. Mol Cancer 22:48. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s12943-023-01744-8\u003c/span\u003e\u003cspan address=\"10.1186/s12943-023-01744-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTamir A, Jag U, Sarojini S et al (2014) Kallikrein family proteases KLK6 and KLK7 are potential early detection and diagnostic biomarkers for serous and papillary serous ovarian cancer subtypes. J Ovarian Res 7:109. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/s13048-014-0109-z\u003c/span\u003e\u003cspan address=\"10.1186/s13048-014-0109-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZafrakas M, Petschke B, Donner A et al (2006) Expression analysis of mammaglobin A (SCGB2A2) and lipophilin B (SCGB1D2) in more than 300 human tumors and matching normal tissues reveals their co-expression in gynecologic malignancies. BMC Cancer 6:88. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1186/1471-2407-6-88\u003c/span\u003e\u003cspan address=\"10.1186/1471-2407-6-88\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePrakash R, Zhang Y, Feng W, Jasin M (2015) Homologous recombination and human health: the roles of BRCA1, BRCA2, and associated proteins. Cold Spring Harb Perspect Biol 7(4):a016600. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1101/cshperspect.a016600\u003c/span\u003e\u003cspan address=\"10.1101/cshperspect.a016600\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDangaj D, Bruand M, Grimm AJ et al (2019) Cooperation between Constitutive and Inducible Chemokines Enables T Cell Engraftment and Immune Attack in Solid Tumors. Cancer Cell 35(6):885\u0026ndash;900. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ccell.2019.05.004\u003c/span\u003e\u003cspan address=\"10.1016/j.ccell.2019.05.004\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChow MT, Ozga AJ, Servis RL et al (2020) Macrophage-Derived CXCL9 and CXCL10 Are Required for Antitumor Immune Responses Following Immune Checkpoint Blockade. Clin Cancer Res 26(2):487\u0026ndash;504. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1158/1078-0432.CCR-19-1868\u003c/span\u003e\u003cspan address=\"10.1158/1078-0432.CCR-19-1868\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMoore K, Colombo N, Scambia G et al (2018) Maintenance Olaparib in Patients with Newly Diagnosed Advanced Ovarian Cancer. N Engl J Med 379(26):2495\u0026ndash;2505. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1056/NEJMoa1810858\u003c/span\u003e\u003cspan address=\"10.1056/NEJMoa1810858\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRay-Coquard I, Pautier P, Pignata S et al (2019) Olaparib plus Bevacizumab as First-Line Maintenance in Ovarian Cancer. N Engl J Med 381(25):2416\u0026ndash;2428. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1056/NEJMoa1911361\u003c/span\u003e\u003cspan address=\"10.1056/NEJMoa1911361\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCortez A, Zhou J, Zhang Y, Lim C (2023) PARP Inhibitors in Ovarian Cancer: A Review. Target Oncol 18(4):471\u0026ndash;503. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1007/s11523-023-00970-w\u003c/span\u003e\u003cspan address=\"10.1007/s11523-023-00970-w\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eOzols RF (2018) Carboplatin/Paclitaxel Induction in Ovarian Cancer: The Finer Points. Oncologist 23(9):1003\u0026ndash;1008\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eChan JK, Brady MF, Penson RT et al (2016) Weekly vs. Every-3-Week Paclitaxel and Carboplatin for Ovarian Cancer. N Engl J Med 374(8):738\u0026ndash;748. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1056/NEJMoa1505067\u003c/span\u003e\u003cspan address=\"10.1056/NEJMoa1505067\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eClamp AR, James EC, McNeish IA et al (2019) Weekly dose-dense chemotherapy in first-line epithelial ovarian cancer treatment (ICON8): primary progression free survival results from a GCIG phase 3 randomised controlled trial. Lancet 394(10214):2084\u0026ndash;2095. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/S0140-6736(19)32259-7\u003c/span\u003e\u003cspan address=\"10.1016/S0140-6736(19)32259-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMonti S, Tamayo P, Mesirov J, Golub T (2003) Consensus Clustering: A Resampling-Based Method for Class Discovery and Visualization of Gene Expression Microarray Data. Mach Learn 52(1\u0026ndash;2):91\u0026ndash;118. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1023/A:1023949509487\u003c/span\u003e\u003cspan address=\"10.1023/A:1023949509487\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWilkerson MD, Hayes DN (2010) ConsensusClusterPlus: a class discovery tool with confidence assessments and item tracking. Bioinformatics 26(12):1572\u0026ndash;1573. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btq170\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btq170\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003ePedregosa F, Varoquaux G, Gramfort A et al (2011) Scikit-learn: Machine Learning in Python. J Mach Learn Res 12:2825\u0026ndash;2830\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTherneau TM, Grambsch PM (2000) Modeling Survival Data: Extending the Cox Model. Springer, New York\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eScrucca L, Fop M, Murphy TB, Raftery AE (2016) mclust 5: Clustering, Classification and Density Estimation Using Gaussian Finite Mixture Models. R J 8(1):289\u0026ndash;317\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNational Cancer Institute. SEER Cancer Stat Facts (2024) Ovarian Cancer. Available at: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://seer.cancer.gov/statfacts/html/ovary.html\u003c/span\u003e\u003cspan address=\"https://seer.cancer.gov/statfacts/html/ovary.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Accessed January\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"Students Research Committee, School of Medicine, Ardabil University of Medical Sciences, Ardabil, Iran","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Ovarian cancer, variational autoencoder, molecular subtypes, drug repurposing, TCGA","lastPublishedDoi":"10.21203/rs.3.rs-9693642/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9693642/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eOvarian cancer exhibits pronounced molecular heterogeneity that complicates therapeutic stratification and patient management. Here we present a comprehensive computational framework integrating variational autoencoders (VAE), probabilistic Gaussian Mixture Model (GMM) clustering, and rigorous consensus stability assessment to deconvolute transcriptomic heterogeneity in 421 high-grade serous ovarian cancer samples from The Cancer Genome Atlas (TCGA-OV). The VAE extracted a 32-dimensional latent representation capturing major axes of transcriptional variation, and GMM clustering followed by 1,000 bootstrap consensus resamplings identified three robust molecular subtypes with high assignment confidence (mean entropy\u0026thinsp;=\u0026thinsp;0.05). The subtypes corresponded to distinct biological programs: Cluster 0 (mesenchymal/ciliary, 33% of samples), Cluster 1 (fibrotic/protease-enriched, 43%), and Cluster 2 (homologous recombination deficiency and immune-enriched, 24%). Comparative benchmarking against conventional methods (PCA\u0026thinsp;+\u0026thinsp;k-means, hierarchical clustering, NMF, UMAP+HDBSCAN) demonstrated superior cluster separation (Silhouette coefficient\u0026thinsp;=\u0026thinsp;0.61 vs\u0026thinsp;\u0026le;\u0026thinsp;0.45 for alternatives). Differential expression analysis identified key subtype markers including CFAP57, KLK7, SCGB2A2, and ITGA11. Functional enrichment linked Cluster 0 to cytoskeletal organization, Cluster 1 to extracellular matrix remodeling, and Cluster 2 to DNA repair and immune signaling. Comprehensive survival analysis revealed no significant prognostic differences across subtypes (overall survival: log-rank p\u0026thinsp;=\u0026thinsp;0.885; disease-free survival: p\u0026thinsp;=\u0026thinsp;0.783), suggesting these molecular classifications may reflect treatment-responsive vulnerabilities rather than intrinsic prognostic determinants in the context of uniform platinum-based therapy. A composite deep learning drug repurposing pipeline integrating predicted sensitivities, target similarity, and pathway analysis prioritized subtype-specific therapeutic hypotheses. Critical limitations include absence of external validation cohorts and purely computational drug predictions without experimental validation. These findings provide a reproducible, rigorously benchmarked computational framework for molecular subtype discovery and hypothesis generation in ovarian cancer precision oncology.\u003c/p\u003e","manuscriptTitle":"Deep Learning-Based Molecular Subtyping Identifies Three Biologically Distinct Ovarian Cancer Subtypes with Potential Therapeutic Implications","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-05-14 14:58:56","doi":"10.21203/rs.3.rs-9693642/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"47b2a261-fbd1-4795-9897-318357d1811a","owner":[],"postedDate":"May 14th, 2026","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":68029730,"name":"Computational Biology"},{"id":68029731,"name":"Bioinformatics"},{"id":68029732,"name":"Cancer Biology"}],"tags":[],"updatedAt":"2026-05-14T14:58:56+00:00","versionOfRecord":[],"versionCreatedAt":"2026-05-14 14:58:56","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9693642","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9693642","identity":"rs-9693642","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.