Classifying Breast Cancer Molecular Subtypes using Deep Clustering Approach

preprint OA: closed CC-BY-4.0
📄 Open PDF Full text JSON View at publisher

Abstract

Abstract Background: Cancer is a complex disease with a high rate of mortality. The characteristics of tumor masses are very heterogeneous; thus, the appropriate classification of tumors is a critical point in the correct treatment. A high level of heterogeneity has also been observed in breast cancer. Therefore, detecting the molecular subtypes of this disease is a worthwhile issue for medicine that could be facilitated using bioinformatics. Method: Numerous methods have already classified breast cancer based on gene expression data; however, they are not reliable due to the dynamic nature of these data. In contrast, gene mutation data are relatively stable and may lead to better classification. The aim of this study is to introduce a novel method for detecting the molecular subtypes of breast cancer. In this study, somatic mutation profiles of tumors are used; nonetheless, the somatic mutation profiles are very sparse. To address this issue, we made use of the network propagation method on gene interaction network and made the mutation profiles dense. Afterward, we used deep embedded clustering (DEC) method to classify breast tumors into four subtypes. In the next step, gene signatures of each subtype obtained by Fisher exact test and Benjamini-Hochberg procedure. Results: Clinical and molecular analyses are executed, besides enrichment of results in numerous databases have shown that the proposed method, using mutation profiles can efficiently detect the molecular subtypes of breast cancer. Finally, a supervised classifier is proposed based on discovered subtypes to predict the molecular subtype of a new patient.
Full text 15,756 characters · extracted from preprint-html · click to expand
Classifying Breast Cancer Molecular Subtypes using Deep Clustering Approach | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research article Classifying Breast Cancer Molecular Subtypes using Deep Clustering Approach Narjes Rohani, Changiz Eslahchi This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.2.19530/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 25 Nov, 2020 Read the published version in Frontiers in Genetics → Version 1 posted You are reading this latest preprint version Abstract Background: Cancer is a complex disease with a high rate of mortality. The characteristics of tumor masses are very heterogeneous; thus, the appropriate classification of tumors is a critical point in the correct treatment. A high level of heterogeneity has also been observed in breast cancer. Therefore, detecting the molecular subtypes of this disease is a worthwhile issue for medicine that could be facilitated using bioinformatics. Method: Numerous methods have already classified breast cancer based on gene expression data; however, they are not reliable due to the dynamic nature of these data. In contrast, gene mutation data are relatively stable and may lead to better classification. The aim of this study is to introduce a novel method for detecting the molecular subtypes of breast cancer. In this study, somatic mutation profiles of tumors are used; nonetheless, the somatic mutation profiles are very sparse. To address this issue, we made use of the network propagation method on gene interaction network and made the mutation profiles dense. Afterward, we used deep embedded clustering (DEC) method to classify breast tumors into four subtypes. In the next step, gene signatures of each subtype obtained by Fisher exact test and Benjamini-Hochberg procedure. Results: Clinical and molecular analyses are executed, besides enrichment of results in numerous databases have shown that the proposed method, using mutation profiles can efficiently detect the molecular subtypes of breast cancer. Finally, a supervised classifier is proposed based on discovered subtypes to predict the molecular subtype of a new patient. Oncology Machine learning Cancer Molecular subtypes Breast cancer Tumor classi cation Cancer heterogeneity Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Full Text Supplementary Files Additional.pdf Cite Share Download PDF Status: Published Journal Publication published 25 Nov, 2020 Read the published version in Frontiers in Genetics → Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-10143","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":267114,"identity":"11fb32d0-db25-4023-afdd-8b1a2780c990","order_by":1,"name":"Narjes Rohani","email":"","orcid":"","institution":"Shahid Beheshti University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Narjes","middleName":"","lastName":"Rohani","suffix":""},{"id":267115,"identity":"9a210840-63a9-4aef-8d9e-72a999f23d26","order_by":2,"name":"Changiz Eslahchi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7ElEQVRIie3RsWrDMBCA4ROGdJHxeuCSvMKFjIE8h0eJDt6ahC4pGeJicMashr5ER482GrzoAQzpUD+Cl2BvdVzaJSBn7KB/uEHwwR0CsNn+YV56nRsQgOzti0U/r7mJYHWdNJCY7iIU/BGY4C8xRn7E6pZWz957nLx2mQLvmDO1M5HH3FlwenrBzyI5u1oBagGFNhFcX3wgR36gTM4sUQD9dYVpQULx0LV0GMi268lslARiApzUQMDtCY0RrITjcyplWsnYd3XI51pGRuKlgjXtbi9PaVg3XbacTkulGhO5iQPc9Ts2m81mM/UNE9JQDMPVHq4AAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-8913-3904","institution":"Shahid Beheshti University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Changiz","middleName":"","lastName":"Eslahchi","suffix":""}],"badges":[],"createdAt":"2019-12-22 13:22:35","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.2.19530/v1","doiUrl":"https://doi.org/10.21203/rs.2.19530/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.3389/fgene.2020.553587","type":"published","date":"2020-11-25T19:06:03+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":296712,"identity":"f7a17121-906c-4415-bc77-d2c5f7b0bcd4","added_by":"auto","created_at":"2019-12-23 20:00:07","extension":"jpg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1108894,"visible":true,"origin":"","legend":"The schema of the proposed method","description":"","filename":"1.jpg","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/1.jpg"},{"id":296713,"identity":"7a895d0d-d1cd-4398-bcde-83ec88597352","added_by":"auto","created_at":"2019-12-23 20:00:08","extension":"jpg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":16555,"visible":true,"origin":"","legend":"Bipartite graph between method to be evaluated and PAM50","description":"","filename":"2.jpg","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/2.jpg"},{"id":296714,"identity":"3364fab9-58aa-4904-9e01-367dc86e0369","added_by":"auto","created_at":"2019-12-23 20:00:08","extension":"jpg","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":3702397,"visible":true,"origin":"","legend":"Subnetworks of all subtypes","description":"","filename":"3.jpg","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/3.jpg"},{"id":296715,"identity":"3903e23a-fbdd-422d-85f7-a2fc1b9509e6","added_by":"auto","created_at":"2019-12-23 20:00:08","extension":"jpg","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":96656,"visible":true,"origin":"","legend":"Cox Hazard Regression Survival Diagram","description":"","filename":"4.jpg","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/4.jpg"},{"id":296716,"identity":"be79258a-32b5-4abd-baf4-9c3a76e5adf1","added_by":"auto","created_at":"2019-12-23 20:00:09","extension":"jpg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":28051,"visible":true,"origin":"","legend":"Visualization of each Classi\fcation","description":"","filename":"5.jpg","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/5.jpg"},{"id":296717,"identity":"7765c164-1554-4a38-b8eb-9607678a4ac5","added_by":"auto","created_at":"2019-12-23 20:00:09","extension":"jpg","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":36836,"visible":true,"origin":"","legend":"Area under curve plot of random forest","description":"","filename":"6.jpg","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/6.jpg"},{"id":779285,"identity":"a623291a-4228-4910-952e-c961950777b3","added_by":"auto","created_at":"2020-03-30 22:10:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":5902938,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-10143/v1/manuscript.pdf"},{"id":449699,"identity":"d29a1783-bb6b-4f3a-aed8-38d8656b7aa2","added_by":"auto","created_at":"2020-02-05 14:16:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":23600845,"visible":true,"origin":"","legend":"","description":"","filename":"bmcarticle.pdf","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/bmc_article.pdf"},{"id":296718,"identity":"8b8a190d-61c4-4277-a45c-341f45813faf","added_by":"auto","created_at":"2019-12-23 20:00:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":23681918,"visible":true,"origin":"","legend":"","description":"","filename":"bmcarticle.pdf","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/Manuscript.pdf"},{"id":13483512,"identity":"fa3ee8b8-d22b-4557-8658-3dc61a93e98a","added_by":"auto","created_at":"2021-09-16 21:54:00","extension":"pdf","order_by":1,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":24026945,"visible":true,"origin":"","legend":"","description":"","filename":"bmcarticle.pdf","url":"https://assets-eu.researchsquare.com/files/rs-10143/v1_covered.pdf"},{"id":296711,"identity":"2d37f5c3-21aa-4213-bc0c-64924761c9c9","added_by":"auto","created_at":"2019-12-23 20:00:07","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":1648706,"visible":true,"origin":"","legend":"","description":"","filename":"Additional.pdf","url":"https://assets-eu.researchsquare.com/files/0c077e85-96a7-4811-b2d5-ece0121741d4/v1/Additional.pdf"}],"financialInterests":"","formattedTitle":"Classifying Breast Cancer Molecular Subtypes using Deep Clustering Approach","fulltext":[{"header":"Full Text","content":"\u003cp\u003eThis preprint is available for \u003ca href='/article/rs-10143/latest.pdf' target='_blank'\u003edownload as a PDF\u003c/a\u003e.\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":false,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":true,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Machine learning, Cancer, Molecular subtypes, Breast cancer, Tumor classi\fcation, Cancer heterogeneity","lastPublishedDoi":"10.21203/rs.2.19530/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.2.19530/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eBackground: Cancer is a complex disease with a high rate of mortality. The characteristics of tumor masses are very heterogeneous; thus, the appropriate classification of tumors is a critical point in the correct treatment. A high level of heterogeneity has also been observed in breast cancer. Therefore, detecting the molecular subtypes of this disease is a worthwhile issue for medicine that could be facilitated using bioinformatics. \u003c/p\u003e\u003cp\u003eMethod: Numerous methods have already classified breast cancer based on gene expression data; however, they are not reliable due to the dynamic nature of these data. In contrast, gene mutation data are relatively stable and may lead to better classification. The aim of this study is to introduce a novel method for detecting the molecular subtypes of breast cancer. In this study, somatic mutation profiles of tumors are used; nonetheless, the somatic mutation profiles are very sparse. To address this issue, we made use of the network propagation method on gene interaction network and made the mutation profiles dense. Afterward, we used deep embedded clustering (DEC) method to classify breast tumors into four subtypes. In the next step, gene signatures of each subtype obtained by Fisher exact test and Benjamini-Hochberg procedure. \u003c/p\u003e\u003cp\u003eResults: Clinical and molecular analyses are executed, besides enrichment of results in numerous databases have shown that the proposed method, using mutation profiles can efficiently detect the molecular subtypes of breast cancer. Finally, a supervised classifier is proposed based on discovered subtypes to predict the molecular subtype of a new patient.\u003c/p\u003e","manuscriptTitle":"Classifying Breast Cancer Molecular Subtypes using Deep Clustering Approach","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2019-12-23 20:00:06","doi":"10.21203/rs.2.19530/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d2dd3dd2-bee9-477f-8007-1bb9d3d3fdfb","owner":[],"postedDate":"December 23rd, 2019","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":45020,"name":"Oncology"}],"tags":[],"updatedAt":"2021-07-22T19:06:03+00:00","versionOfRecord":{"articleIdentity":"rs-10143","link":"https://doi.org/10.3389/fgene.2020.553587","journal":{"identity":"frontiers-in-genetics","isVorOnly":true,"title":"Frontiers in Genetics"},"publishedOn":"2020-11-25 19:06:03","publishedOnDateReadable":"November 25th, 2020"},"versionCreatedAt":"2019-12-23 20:00:06","video":"","vorDoi":"10.3389/fgene.2020.553587","vorDoiUrl":"https://doi.org/10.3389/fgene.2020.553587","workflowStages":[]},"version":"v1","identity":"rs-10143","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"identity":"rs-10143","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00
unpaywall
last seen: 2026-05-29T02:00:03.542394+00:00
License: CC-BY-4.0