Identifying Ovarian Cancer with Machine Learning DNA Methylation Pattern Analysis | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Identifying Ovarian Cancer with Machine Learning DNA Methylation Pattern Analysis Jesus Gonzalez Bosquet, Vincent M. Wagner, Douglas Russo, Henry D. Reyes, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6148442/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 01 Jul, 2025 Read the published version in Scientific Reports → Version 1 posted 11 You are reading this latest preprint version Abstract The majority of patients with epithelial ovarian cancer (EOC) continue to be diagnosed at an advanced stage despite great advances in this disease treatment. To impact overall survival, we need better methods of EOC early diagnosis. We performed a case control study to predict high-grade serous cancer (HGSC) using artificial intelligence (AI) methodology and methylated DNA from surgical specimens. Initial prediction models with MethylNet were accurate but complex (AUC = 100%). We optimized these models by selecting the most informative probes with univariate ANOVA analyses first, and then multivariate lasso regression modelling. This step-wise approach resulted in 9 methylated probes predicting HGSC with an AUC of 100%. These models were validated with different analytics and with an independent DNA-methylation experiment with excellent performances. Biological sciences/Cancer/Gynaecological cancer/Ovarian cancer Biological sciences/Cancer Biological sciences/Cancer/Cancer genomics Figures Figure 1 Figure 2 Introduction Despite notable advances in treatment, patients with epithelial ovarian cancer (EOC) continue to have a high case-fatality rate. This is due in large part to most patients presenting with advanced, symptomatic disease. 1 Conventional methods of diagnosis (imaging, CA125) lack the desired sensitivity and/or specificity to be effective and/or accurate for early diagnosis. Therefore, there is a critical need to develop new strategies with this goal. One of the most exciting novel methods for early diagnosis is detection of cancer genomic material in blood with cell-free DNA (cfDNA) analysis, or ‘liquid biopsy’. 2 To be successful in this endeavor though, adequate markers, analytics and procedures have to be developed. DNA-methylation is an effective screening tool for colon cancer 3 and also has demonstrated potential for ovarian cancer detection. 4 Additionally, artificial intelligence (AI) could be applied to improve the performance of predicting diagnostic tools. 5 The ultimate goal of our study is to develop a method that could be used in early detection of ovarian cancer with cfDNA methodology. Our primary aim was to build and validate a prediction model of high-grade serous ovarian cancer (HGSC - the most common EOC) using deep machine learning (AI) and DNA methylation data. Methods This was a pilot case-control study, including samples from patients with HGSC (N = 99) and normal fallopian tube samples as controls (N = 12)(Fig. 1 A a ). Tissue samples and clinical outcome data were obtained from the Gynecologic Oncology Bank (IRB, ID#200209010) of the University of Iowa (UI) after informed consent from all subjects. Methods were carried out in accordance with UI IRB guidelines and regulations. The experimental protocol was approved by the IRB of the UI, IRB ID# 201804817 (first approved on 05/09/2018) entitled ‘Prediction Models in Ovarian Cancer.’ Clinical and pathological data were collected from the electronic medical record. An Illumina Infinium MethylationEPIC BeadChip Array® was used to determine more than 850,000 DNA methylation features. Methods of patient selection, DNA isolation, bisulfite conversion, and hybridization to the chip are described in detailed elsewhere. 6 This Infinium array has more than 850K methylation datapoints (probes). To create a valid and effective model for a clinical setting extracted the most informative probes. The initial variable reduction was performed with an open-source tool, MethylNet , 7 which has been previously tested in TCGA (the Cancer genome Atlas) datasets successfully. Then, resulting models were optimized for clinical use using univariate and multivariate lasso regression modelling with k-fold cross-validation ( caret and glmnet R packages). Performance of models were measured with the area under the curve (AUC) and their 95% confidence interval (CI). The datasets analyzed during the current study are available in the Genome Expression Omnibus (GEO) repository, database GSE133556 ( https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE133556 ). The resulting models were validated in an independent database also available at the GEO repository, database GSE65820 8 and includes also fallopian tube and HGSC samples and all the most informative probes resulting from the analysis. We used classical learning statistical methods like pROC (R package) and also independent machine learning (ML) analytics, like TensorFlow , for validation of these models. Results The initial variable selection with MethylNet produced a model with 23,397 informative probes and a performance of 100% measured by the AUC ( Figure 1A b ). A model with such a number of variables is impractical. Therefore, we proceeded with multiple ANOVA univariate analysis to select those more informative (11,167 probes at a p-value<0.05) to be included in the multivariate analysis with lasso regression ( Figure 1B c ). The resulting lasso model comprised 9 informative probes ( Figure 1B e ) and had an AUC of 100% ( Figure 1B d ). Validation of prediction models of HGSC created in the UI set and applied to GSE6582 dataset had very good performances, with AUC of 98% (95% CI: 95-100%) for the model with 11,167 probes after the ANOVA analysis, and with an AUC of 84% (CI: 76-93%) for the model with 9 probes after the lasso regression ( Figure 2A ). Validation of these models in a ML platform also had excellent performances as detailed in Figure 2B . Discussion With our step-wise AI methodology, we created and validated a prediction model for HGSC using methylated DNA probes that is accurate and robust across independent datasets. Strengths of our study include the use of a single institution biobank that is well annotated clinically. 6 An homogeneous phenotype (HGSC) in a single population 9 is advantageous when measuring accuracy of prediction models. However, it could detract from the generalizability of the model. Thus, we validated the model in an independent database, from a different geographical location (Australia) but with similar patient ancestry (western Europeans). It is important to note the need for further model validation. For this model to be effective as early diagnosis or screening, we need to optimize its performance in blood, with a diverse population of EOC patients from all stages and appropriate controls. That is our future goal. Declarations Author Contribution Conceptualization, J.G.B. and V.W.; methodology, J.G.B., W.V. and D.R.; validation, J.G.B.; formal analysis, J.G.B. and D.R.; investigation, H.R..; resources, J.G.B., M.J.G. and D.P.B.; data curation, H.R. and J.G.B.; writing—original draft preparation, J.G.B. and D.R.; writing—review and editing, J.G.B., V.W., H.R., A.N., M.J.G. and D.P.B.; visualization, J.G.B.; supervision, J.G.B. and D.R.; project administration, J.G.B. and M.J.G.; funding acquisition, J.G.B. and H.R.; M.J.G. is responsible for assembling and maintaining the tumor bank utilized for this study. All authors have read and agreed to the published version of the manuscript. Data Availability The datasets analyzed during the current study are available in the Genome Expression Omnibus (GEO) repository, database GSE133556 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE133556).The resulting models were validated in an independent database also available at the GEO repository, database GSE658208. Precis We created and validated a prediction model for ovarian cancer using methylated DNA probes and AI methodology that is accurate and robust across independent datasets. References Henley, S. J. et al. Annual report to the nation on the status of cancer, part I: National cancer statistics. Cancer 126 , 2225-2249, doi:10.1002/cncr.32802 (2020). Minato, T. et al. Liquid biopsy with droplet digital PCR targeted to specific mutations in plasma cell-free tumor DNA can detect ovarian cancer recurrence earlier than CA125. Gynecol Oncol Rep 38 , 100847, doi:10.1016/j.gore.2021.100847 (2021). Imperiale, T. F. et al. Multitarget stool DNA testing for colorectal-cancer screening. The New England journal of medicine 370 , 1287-1297, doi:10.1056/NEJMoa1311194 (2014). Marinelli, L. M. et al. Methylated DNA markers for plasma detection of ovarian cancer: Discovery, validation, and clinical feasibility. Gynecologic oncology 165 , 568-576, doi:10.1016/j.ygyno.2022.03.018 (2022). Mendes, J., Domingues, J., Aidos, H., Garcia, N. & Matela, N. AI in Breast Cancer Imaging: A Survey of Different Applications. J Imaging 8 , doi:10.3390/jimaging8090228 (2022). Reyes, H. D. et al. Differential DNA methylation in high-grade serous ovarian cancer (HGSOC) is associated with tumor behavior. Scientific reports 9 , 17996, doi:10.1038/s41598-019-54401-w (2019). Levy, J. J. et al. MethylNet: an automated and modular deep learning approach for DNA methylation analysis. BMC bioinformatics 21 , 108, doi:10.1186/s12859-020-3443-8 (2020). Patch, A. M. et al. Whole-genome characterization of chemoresistant ovarian cancer. Nature 521 , 489-494, doi:10.1038/nature14410 (2015). Miller, M. D. et al. Population Substructure Has Implications in Validating Next-Generation Cancer Genomics Studies with TCGA. Int J Mol Sci 20 , doi:10.3390/ijms20051192 (2019). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Published Journal Publication published 01 Jul, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 21 Apr, 2025 Reviews received at journal 05 Apr, 2025 Reviews received at journal 05 Apr, 2025 Reviewers agreed at journal 27 Mar, 2025 Reviewers agreed at journal 27 Mar, 2025 Reviewers agreed at journal 25 Mar, 2025 Reviewers invited by journal 25 Mar, 2025 Editor assigned by journal 25 Mar, 2025 Editor invited by journal 14 Mar, 2025 Submission checks completed at journal 13 Mar, 2025 First submitted to journal 03 Mar, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6148442","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":435015838,"identity":"69ff74ba-300b-42ec-9b59-bee76e1a63fc","order_by":0,"name":"Jesus Gonzalez Bosquet","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABOklEQVRIie2QMWvCQBTHXwjY5YprgkW/wkkgUjL0qyQIdbmhUBAHK5nMkg9wQ2m+Qlw633EQl2DXlnawFOziELcMtjSnLalGO3e43/L+PO7HvfcAFIp/iOb/ygyAfecBoG2o/alAqbD0uLJDqfDxT6+q6EGwMHO4aXbqYs6y9UuzFZD22+pOnNWN7hyyvqgMFqZ2A8HUOvcTzClaWDhdWpjfC2TSS6zRWVWhpNbQPhMv5j4WyBBebBDbkAp+drF+Oq4q0bscrFDESSbWWHgRJZ2c30qll+kfBxQKtoFg6MUJwgJc4fmPxAbuS4VgXTughMRyELBiBXTFQyZ3WVwbadJDZrQsOrPevtIOpq9POYya+GE6meebi3Un2WDoXNSR7PSdiuJvys7vbhnZ/vuC1raMjigKhUKhKPgCcd9+2o3qLMUAAAAASUVORK5CYII=","orcid":"","institution":"University of Iowa","correspondingAuthor":true,"prefix":"","firstName":"Jesus","middleName":"Gonzalez","lastName":"Bosquet","suffix":""},{"id":435015841,"identity":"fd0334f4-248b-46c8-87b5-49793cf0e6fe","order_by":1,"name":"Vincent M. Wagner","email":"","orcid":"","institution":"University of Iowa","correspondingAuthor":false,"prefix":"","firstName":"Vincent","middleName":"M.","lastName":"Wagner","suffix":""},{"id":435015842,"identity":"293a3596-2388-4f3a-8281-0a2934eafb1c","order_by":2,"name":"Douglas Russo","email":"","orcid":"","institution":"University of Chicago","correspondingAuthor":false,"prefix":"","firstName":"Douglas","middleName":"","lastName":"Russo","suffix":""},{"id":435015843,"identity":"4b7382af-31cf-4ea9-9eed-ebc28ea6887a","order_by":3,"name":"Henry D. Reyes","email":"","orcid":"","institution":"GPPC Network","correspondingAuthor":false,"prefix":"","firstName":"Henry","middleName":"D.","lastName":"Reyes","suffix":""},{"id":435015844,"identity":"5a127418-0558-429c-9b9d-bff6c9e0a174","order_by":4,"name":"Andreea M. Newtson","email":"","orcid":"","institution":"Endeavor Health","correspondingAuthor":false,"prefix":"","firstName":"Andreea","middleName":"M.","lastName":"Newtson","suffix":""},{"id":435015845,"identity":"30410aa4-053b-4d93-b07e-08dcbeacb3ee","order_by":5,"name":"David P. Bender","email":"","orcid":"","institution":"University of Iowa","correspondingAuthor":false,"prefix":"","firstName":"David","middleName":"P.","lastName":"Bender","suffix":""},{"id":435015846,"identity":"97019ee3-43d6-4558-8af9-990981bcfd7a","order_by":6,"name":"Michael J. Goodheart","email":"","orcid":"","institution":"University of Iowa","correspondingAuthor":false,"prefix":"","firstName":"Michael","middleName":"J.","lastName":"Goodheart","suffix":""}],"badges":[],"createdAt":"2025-03-03 17:53:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6148442/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6148442/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-05460-9","type":"published","date":"2025-07-01T15:57:34+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":80050500,"identity":"eba377b5-8fe5-4cbb-ae9f-734b11a5dbd5","added_by":"auto","created_at":"2025-04-07 10:20:43","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1351473,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cu\u003e\u003cstrong\u003eDNA-methylation model to predict HGSC.\u003c/strong\u003e\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA.\u003c/strong\u003e\u003cem\u003e\u003cstrong\u003e MethylNet\u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003e hyperparameter selection and prediction for HGSC: a) \u003c/strong\u003ePreprocessing with normalization, selection of autosomal chromosomal markers, elimination of outliers. Training with 80% of samples, validating with 12.5%. 400 repetitions with most informative 300,000 DNA-methylation features. Representation of all samples with Uniform Manifold Approximation and Projection (UMAP), which is a dimension reduction technique that can be used for visualization of high-dimensional data by giving each datapoint a location in a three-dimensional map (in our sample: x, y and z). \u003cstrong\u003eb)\u003c/strong\u003eAfter \u003cem\u003eMethylNet\u003c/em\u003e prediction analysis, 23,397 DNA-methylated features were selected for the best model to predict HGSC with a performance of 100% measured by the AUC.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB. Variable reduction with \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003elasso \u003c/strong\u003e\u003c/em\u003e\u003cstrong\u003eregression analysis: c) \u003c/strong\u003eHeatmap of the 11,167 significant DNA methylated probes (out of 23,397, p\u0026lt;0.05) after univariate analysis with ANOVA and cross-validation (\u003cem\u003ecaret\u003c/em\u003e R package). \u003cstrong\u003ed) \u003c/strong\u003eAfter \u003cem\u003elasso \u003c/em\u003emultivariate regression analysis (\u003cem\u003eglmnet\u003c/em\u003eR package), there were 9 probes that predicted HGSC (upper axis) with a performance of 100%, measured by AUC (y axis)\u003cem\u003e.\u003c/em\u003e \u003cstrong\u003ec) \u003c/strong\u003eTable with the predictive probes of HGSC in the multivariate analysis: OR: odds Ratio of HGSC: \u0026lt;1, less risk of cancer, \u0026gt;1 more risk of cancer; also, chromosome (chr) position on the chromosome (position); direction of the transcription (strand); UCSC reference gene, where they are located in the gene (group), and what type of DNA methylation the probe is measuring.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6148442/v1/253ac01182eea08a96d55b9e.jpeg"},{"id":80050496,"identity":"66ba6749-e2ab-4c42-bb03-1ed21e10413c","added_by":"auto","created_at":"2025-04-07 10:20:43","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":704273,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cu\u003e\u003cstrong\u003eValidation of prediction models of HGSC with DNA-methylation\u003c/strong\u003e\u003c/u\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eA. Validation of prediction models in GSE65820 dataset: \u003c/strong\u003e(\u003cem\u003eleft panel\u003c/em\u003e) Validation of the model that included 11,167 significant DNA methylated probes in the independent DNA methylation dataset (GSE65820) had an excellent performance of 98% (95% CI: 95%, 100%), measured by the AUC.\u003c/p\u003e\n\u003cp\u003e(\u003cem\u003eright panel\u003c/em\u003e) The validation of the simplified \u003cem\u003elasso \u003c/em\u003emultivariate analysis including only the nine most informative probes, had a very good performance with an AUC of 84% (95% CI: 76%, 93%). Both validations were performed with the \u003cem\u003epROC\u003c/em\u003e R package.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eB. Validation of prediction models with DNA-methylation using a different analytical platform: machine learning (ML) with \u003c/strong\u003e\u003cem\u003e\u003cstrong\u003eTensorFlow\u003c/strong\u003e\u003c/em\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ea)\u003c/strong\u003e Validation of the model using resulting probes from \u003cem\u003eMethylNet\u003c/em\u003e prediction model of HGSC (N=23,397) performed in the UI dataset with ML analytics: the upper panel shows the confusion matrix representing the observed versus the predicted values; the lower panel represents the ROC graphic including models accounting for unbalanced samples: Train R: results of unbalanced (or re-sampling) training model; Test R: results of re-sampling testing model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eb)\u003c/strong\u003e Model using probes selected in the ANOVA univariate analysis with cross-validation (N=11,167) performed in the UI dataset with ML analytics: the upper panel shows the confusion matrix; the lower panel represents the ROC graphic including models accounting for unbalanced samples: Train R: results of unbalanced training model; Test R: results of re-sampling testing model.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ec)\u003c/strong\u003e Model using all 9 probes found to be more informative in the lasso multivariate analysis performed in the UI dataset with ML analytics: the upper panel shows the confusion matrix; the lower panel represents the ROC graphic including models accounting for unbalanced samples.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003ec)\u003c/strong\u003e Validation of the UI model using the most informative 9 probes in the lasso multivariate analysis performed in the GSE65820 dataset with ML analytics: the upper panel shows the confusion matrix; lower panel represents the ROC graphic including: 1) models accounting for weights of the outcome: Train W: results of weighted training model; Test W: results of weighted testing model; 2) models accounting for unbalanced samples.\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6148442/v1/e0418bc44f3a173a833308f0.jpeg"},{"id":86178996,"identity":"3e3e532e-16fa-48ff-9619-492dfcba544d","added_by":"auto","created_at":"2025-07-07 16:14:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":2566947,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6148442/v1/84b866b4-d481-4f69-8fdb-f2c671eddab0.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Identifying Ovarian Cancer with Machine Learning DNA Methylation Pattern Analysis","fulltext":[{"header":"Introduction","content":"\u003cp\u003eDespite notable advances in treatment, patients with epithelial ovarian cancer (EOC) continue to have a high case-fatality rate. This is due in large part to most patients presenting with advanced, symptomatic disease.\u003csup\u003e\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u003c/sup\u003e Conventional methods of diagnosis (imaging, CA125) lack the desired sensitivity and/or specificity to be effective and/or accurate for early diagnosis. Therefore, there is a critical need to develop new strategies with this goal.\u003c/p\u003e \u003cp\u003eOne of the most exciting novel methods for early diagnosis is detection of cancer genomic material in blood with cell-free DNA (cfDNA) analysis, or \u0026lsquo;liquid biopsy\u0026rsquo;.\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e To be successful in this endeavor though, adequate markers, analytics and procedures have to be developed. DNA-methylation is an effective screening tool for colon cancer\u003csup\u003e\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e and also has demonstrated potential for ovarian cancer detection.\u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e Additionally, artificial intelligence (AI) could be applied to improve the performance of predicting diagnostic tools.\u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThe ultimate goal of our study is to develop a method that could be used in early detection of ovarian cancer with cfDNA methodology. Our primary aim was to build and validate a prediction model of high-grade serous ovarian cancer (HGSC - the most common EOC) using deep machine learning (AI) and DNA methylation data.\u003c/p\u003e"},{"header":"Methods","content":"\u003cp\u003eThis was a pilot case-control study, including samples from patients with HGSC (N\u0026thinsp;=\u0026thinsp;99) and normal fallopian tube samples as controls (N\u0026thinsp;=\u0026thinsp;12)(Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003eA \u003cb\u003ea\u003c/b\u003e). Tissue samples and clinical outcome data were obtained from the Gynecologic Oncology Bank (IRB, ID#200209010) of the University of Iowa (UI) after informed consent from all subjects. Methods were carried out in accordance with UI IRB guidelines and regulations. The experimental protocol was approved by the IRB of the UI, IRB ID# 201804817 (first approved on 05/09/2018) entitled \u0026lsquo;Prediction Models in Ovarian Cancer.\u0026rsquo; Clinical and pathological data were collected from the electronic medical record.\u003c/p\u003e \u003cp\u003eAn Illumina Infinium MethylationEPIC BeadChip Array\u0026reg; was used to determine more than 850,000 DNA methylation features. Methods of patient selection, DNA isolation, bisulfite conversion, and hybridization to the chip are described in detailed elsewhere.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e\u003c/p\u003e \u003cp\u003eThis Infinium array has more than 850K methylation datapoints (probes). To create a valid and effective model for a clinical setting extracted the most informative probes. The initial variable reduction was performed with an open-source tool, \u003cem\u003eMethylNet\u003c/em\u003e,\u003csup\u003e\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e\u003c/sup\u003e which has been previously tested in TCGA (the Cancer genome Atlas) datasets successfully. Then, resulting models were optimized for clinical use using univariate and multivariate lasso regression modelling with k-fold cross-validation (\u003cem\u003ecaret\u003c/em\u003e and \u003cem\u003eglmnet\u003c/em\u003e R packages). Performance of models were measured with the area under the curve (AUC) and their 95% confidence interval (CI).\u003c/p\u003e \u003cp\u003eThe datasets analyzed during the current study are available in the Genome Expression Omnibus (GEO) repository, database GSE133556 (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE133556\u003c/span\u003e\u003cspan address=\"https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE133556\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe resulting models were validated in an independent database also available at the GEO repository, database GSE65820\u003csup\u003e8\u003c/sup\u003e and includes also fallopian tube and HGSC samples and all the most informative probes resulting from the analysis. We used classical learning statistical methods like \u003cem\u003epROC\u003c/em\u003e (R package) and also independent machine learning (ML) analytics, like \u003cem\u003eTensorFlow\u003c/em\u003e, for validation of these models.\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eThe initial variable selection with \u003cem\u003eMethylNet\u003c/em\u003e produced a model with 23,397 informative probes and a performance of 100% measured by the AUC (\u003cstrong\u003eFigure 1A b\u003c/strong\u003e). A model with such a number of variables is impractical. Therefore, we proceeded with multiple ANOVA univariate analysis to select those more informative (11,167 probes at a p-value\u0026lt;0.05) to be included in the multivariate analysis with lasso regression (\u003cstrong\u003eFigure 1B c\u003c/strong\u003e). The resulting lasso model comprised 9 informative probes (\u003cstrong\u003eFigure 1B e\u003c/strong\u003e) and had an AUC of 100% (\u003cstrong\u003eFigure 1B d\u003c/strong\u003e).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eValidation of prediction models of HGSC created in the UI set and applied to GSE6582 dataset had very good performances, with AUC of 98% (95% CI: 95-100%) for the model with 11,167 probes after the ANOVA analysis, and with an AUC of 84% (CI: 76-93%) for the model with 9 probes after the lasso regression (\u003cstrong\u003eFigure 2A\u003c/strong\u003e). Validation of these models in a ML platform also had excellent performances as detailed in \u003cstrong\u003eFigure 2B\u003c/strong\u003e.\u0026nbsp;\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eWith our step-wise AI methodology, we created and validated a prediction model for HGSC using methylated DNA probes that is accurate and robust across independent datasets.\u003c/p\u003e \u003cp\u003eStrengths of our study include the use of a single institution biobank that is well annotated clinically.\u003csup\u003e\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u003c/sup\u003e An homogeneous phenotype (HGSC) in a single population\u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e is advantageous when measuring accuracy of prediction models. However, it could detract from the generalizability of the model. Thus, we validated the model in an independent database, from a different geographical location (Australia) but with similar patient ancestry (western Europeans).\u003c/p\u003e \u003cp\u003eIt is important to note the need for further model validation. For this model to be effective as early diagnosis or screening, we need to optimize its performance in blood, with a diverse population of EOC patients from all stages and appropriate controls. That is our future goal.\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eConceptualization, J.G.B. and V.W.; methodology, J.G.B., W.V. and D.R.; validation, J.G.B.; formal analysis, J.G.B. and D.R.; investigation, H.R..; resources, J.G.B., M.J.G. and D.P.B.; data curation, H.R. and J.G.B.; writing\u0026mdash;original draft preparation, J.G.B. and D.R.; writing\u0026mdash;review and editing, J.G.B., V.W., H.R., A.N., M.J.G. and D.P.B.; visualization, J.G.B.; supervision, J.G.B. and D.R.; project administration, J.G.B. and M.J.G.; funding acquisition, J.G.B. and H.R.; M.J.G. is responsible for assembling and maintaining the tumor bank utilized for this study. All authors have read and agreed to the published version of the manuscript.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe datasets analyzed during the current study are available in the Genome Expression Omnibus (GEO) repository, database GSE133556 (https://www.ncbi.nlm.nih.gov/geo/query/acc.cgi?acc=GSE133556).The resulting models were validated in an independent database also available at the GEO repository, database GSE658208.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003e\u003cu\u003ePrecis\u003c/u\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eWe created and validated a prediction model for ovarian cancer using methylated DNA probes and AI methodology that is accurate and robust across independent datasets.\u0026nbsp;\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eHenley, S. J.\u003cem\u003e et al.\u003c/em\u003e Annual report to the nation on the status of cancer, part I: National cancer statistics. \u003cem\u003eCancer\u003c/em\u003e \u003cstrong\u003e126\u003c/strong\u003e, 2225-2249, doi:10.1002/cncr.32802 (2020).\u003c/li\u003e\n\u003cli\u003eMinato, T.\u003cem\u003e et al.\u003c/em\u003e Liquid biopsy with droplet digital PCR targeted to specific mutations in plasma cell-free tumor DNA can detect ovarian cancer recurrence earlier than CA125. \u003cem\u003eGynecol Oncol Rep\u003c/em\u003e \u003cstrong\u003e38\u003c/strong\u003e, 100847, doi:10.1016/j.gore.2021.100847 (2021).\u003c/li\u003e\n\u003cli\u003eImperiale, T. F.\u003cem\u003e et al.\u003c/em\u003e Multitarget stool DNA testing for colorectal-cancer screening. \u003cem\u003eThe New England journal of medicine\u003c/em\u003e \u003cstrong\u003e370\u003c/strong\u003e, 1287-1297, doi:10.1056/NEJMoa1311194 (2014).\u003c/li\u003e\n\u003cli\u003eMarinelli, L. M.\u003cem\u003e et al.\u003c/em\u003e Methylated DNA markers for plasma detection of ovarian cancer: Discovery, validation, and clinical feasibility. \u003cem\u003eGynecologic oncology\u003c/em\u003e \u003cstrong\u003e165\u003c/strong\u003e, 568-576, doi:10.1016/j.ygyno.2022.03.018 (2022).\u003c/li\u003e\n\u003cli\u003eMendes, J., Domingues, J., Aidos, H., Garcia, N. \u0026amp; Matela, N. AI in Breast Cancer Imaging: A Survey of Different Applications. \u003cem\u003eJ Imaging\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, doi:10.3390/jimaging8090228 (2022).\u003c/li\u003e\n\u003cli\u003eReyes, H. D.\u003cem\u003e et al.\u003c/em\u003e Differential DNA methylation in high-grade serous ovarian cancer (HGSOC) is associated with tumor behavior. \u003cem\u003eScientific reports\u003c/em\u003e \u003cstrong\u003e9\u003c/strong\u003e, 17996, doi:10.1038/s41598-019-54401-w (2019).\u003c/li\u003e\n\u003cli\u003eLevy, J. J.\u003cem\u003e et al.\u003c/em\u003e MethylNet: an automated and modular deep learning approach for DNA methylation analysis. \u003cem\u003eBMC bioinformatics\u003c/em\u003e \u003cstrong\u003e21\u003c/strong\u003e, 108, doi:10.1186/s12859-020-3443-8 (2020).\u003c/li\u003e\n\u003cli\u003ePatch, A. M.\u003cem\u003e et al.\u003c/em\u003e Whole-genome characterization of chemoresistant ovarian cancer. \u003cem\u003eNature\u003c/em\u003e \u003cstrong\u003e521\u003c/strong\u003e, 489-494, doi:10.1038/nature14410 (2015).\u003c/li\u003e\n\u003cli\u003eMiller, M. D.\u003cem\u003e et al.\u003c/em\u003e Population Substructure Has Implications in Validating Next-Generation Cancer Genomics Studies with TCGA. \u003cem\u003eInt J Mol Sci\u003c/em\u003e \u003cstrong\u003e20\u003c/strong\u003e, doi:10.3390/ijms20051192 (2019).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"","lastPublishedDoi":"10.21203/rs.3.rs-6148442/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6148442/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe majority of patients with epithelial ovarian cancer (EOC) continue to be diagnosed at an advanced stage despite great advances in this disease treatment. To impact overall survival, we need better methods of EOC early diagnosis. We performed a case control study to predict high-grade serous cancer (HGSC) using artificial intelligence (AI) methodology and methylated DNA from surgical specimens. Initial prediction models with \u003cem\u003eMethylNet\u003c/em\u003e were accurate but complex (AUC\u0026thinsp;=\u0026thinsp;100%). We optimized these models by selecting the most informative probes with univariate ANOVA analyses first, and then multivariate lasso regression modelling. This step-wise approach resulted in 9 methylated probes predicting HGSC with an AUC of 100%. These models were validated with different analytics and with an independent DNA-methylation experiment with excellent performances.\u003c/p\u003e","manuscriptTitle":"Identifying Ovarian Cancer with Machine Learning DNA Methylation Pattern Analysis","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-04-07 10:20:38","doi":"10.21203/rs.3.rs-6148442/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-04-21T09:05:31+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-06T02:14:28+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-04-05T14:36:39+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"153808454481352311488324794696198936850","date":"2025-03-27T20:05:59+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"330192010856398054869536092524276904652","date":"2025-03-27T09:37:00+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"187687307512706120264146334367458472802","date":"2025-03-25T14:33:12+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-03-25T13:48:00+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-03-25T06:16:01+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-03-14T17:13:34+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-03-13T13:02:38+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-03-03T17:43:38+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"e7345d9b-0144-4f2d-9e2a-86d393c18f36","owner":[],"postedDate":"April 7th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":46323463,"name":"Biological sciences/Cancer/Gynaecological cancer/Ovarian cancer"},{"id":46323464,"name":"Biological sciences/Cancer"},{"id":46323465,"name":"Biological sciences/Cancer/Cancer genomics"}],"tags":[],"updatedAt":"2025-07-07T16:02:20+00:00","versionOfRecord":{"articleIdentity":"rs-6148442","link":"https://doi.org/10.1038/s41598-025-05460-9","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-07-01 15:57:34","publishedOnDateReadable":"July 1st, 2025"},"versionCreatedAt":"2025-04-07 10:20:38","video":"","vorDoi":"10.1038/s41598-025-05460-9","vorDoiUrl":"https://doi.org/10.1038/s41598-025-05460-9","workflowStages":[]},"version":"v1","identity":"rs-6148442","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6148442","identity":"rs-6148442","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.