EARN: an ensemble machine learning algorithm to predict driver genes in metastatic breast cancer

preprint OA: closed
Full text JSON View at publisher
⚙ AI-generated deep summary by qwen3.7-flash, 2026-09-28 · read from full text ⓘ

The authors developed EARN, an ensemble machine learning algorithm combining artificial neural networks, random forests, and support vector machines to identify driver genes in metastatic breast cancer. Using somatic mutation data from 450 metastatic samples and 983 primary tumor samples, the study validated the model's high accuracy compared to individual classifiers through gene set enrichment analysis. The research proposed a targeted gene panel including HDAC3, KRAS, and others to assist precision oncologists in designing compact sequencing panels for prognosis and diagnosis. Relevance to endometriosis: listed as one indication for GnRH antagonists, though the paper's main focus is uterine fibroids.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract BackgroundToday, there are a lot of markers on the prognosis and diagnosis of complex diseases such as primary breast cancer. However, our understanding of the drivers that influence cancer aggression is limited.MethodsIn this work, we study somatic mutation data consists of 450 metastatic breast tumor samples from cBio Cancer Genomics Portal. We use four software tools to extract features from this data. Then, an ensemble classifier (EC) learning algorithm called EARN (Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine) is proposed to evaluate plausible driver genes for metastatic breast cancer (MBCA). ResultsThis study is an attempt to focus on the findings in several aspects of MBCA prognosis and diagnosis. First, drivers and passengers predicted by SVM, ANN, RF, and EARN are introduced. Second, biological inferences of predictions based on gene set enrichment analysis are discussed. Third, statistical validation and comparison of all learning methods based on evaluation metrics are done. Finally, the pathway enrichment analysis (PEA) using ReactomeFIVIz tool (FDR<0.03) for the top 100 genes predicted by EARN leads us to propose a new gene set panel for MBCA, including HDAC3, ABAT, GRIN1, PLCB1, and KPNA2 as well as NCOR1, TBL1XR1, SIRT4, KRAS, CACNA1E, PRKCG, GPS2, SIN3A, ACTB, KDM6B, and PRMT1. Furthermore, we compare results for MBCA to other outputs regarding 983 primary tumor samples of breast invasive carcinoma (BRCA) obtained from the Cancer Genome Atlas (TCGA). The comparison between outputs shows that ROC-AUC reached 99.24% using EARN for MBCA and 99.79% for BRCA. This statistical result is better than three individual classifiers in each case.ConclusionsThis research using an integrative approach assists precision oncologists to design compact targeted panels that eliminate the need for whole-genome/exome sequencing.
Full text 254,155 characters · extracted from preprint-html · click to expand
EARN: an ensemble machine learning algorithm to predict driver genes in metastatic breast cancer | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Help Center Sign In Submit a Preprint Cite Share Download PDF Research article EARN: an ensemble machine learning algorithm to predict driver genes in metastatic breast cancer Leila Mirsadeghi, Reza Haji Hosseini, Ali Mohammad Banaei-Moghaddam, and 1 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-113748/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Background Today, there are a lot of markers on the prognosis and diagnosis of complex diseases such as primary breast cancer. However, our understanding of the drivers that influence cancer aggression is limited. Methods In this work, we study somatic mutation data consists of 450 metastatic breast tumor samples from cBio Cancer Genomics Portal. We use four software tools to extract features from this data. Then, an ensemble classifier (EC) learning algorithm called EARN (Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine) is proposed to evaluate plausible driver genes for metastatic breast cancer (MBCA). Results This study is an attempt to focus on the findings in several aspects of MBCA prognosis and diagnosis. First, drivers and passengers predicted by SVM, ANN, RF, and EARN are introduced. Second, biological inferences of predictions based on gene set enrichment analysis are discussed. Third, statistical validation and comparison of all learning methods based on evaluation metrics are done. Finally, the pathway enrichment analysis (PEA) using ReactomeFIVIz tool ( FDR <0.03) for the top 100 genes predicted by EARN leads us to propose a new gene set panel for MBCA, including HDAC3, ABAT, GRIN1, PLCB1, and KPNA2 as well as NCOR1, TBL1XR1, SIRT4, KRAS, CACNA1E, PRKCG, GPS2, SIN3A, ACTB, KDM6B, and PRMT1. Furthermore, we compare results for MBCA to other outputs regarding 983 primary tumor samples of breast invasive carcinoma (BRCA) obtained from the Cancer Genome Atlas (TCGA). The comparison between outputs shows that ROC-AUC reached 99.24% using EARN for MBCA and 99.79% for BRCA. This statistical result is better than three individual classifiers in each case. Conclusions This research using an integrative approach assists precision oncologists to design compact targeted panels that eliminate the need for whole-genome/exome sequencing. Molecular Genetics Molecular Biology Epigenetics & Genomics Metastasis breast tumor Mutation data Ensemble classifier Plausible drivers Targeted clinical panel sequencing Figures Figure 1 Figure 1 Figure 1 Figure 1 Figure 1 Figure 2 Figure 2 Figure 2 Figure 2 Figure 2 Figure 3 Figure 3 Figure 3 Figure 3 Figure 3 Figure 4 Figure 4 Figure 4 Figure 4 Figure 4 Figure 5 Figure 5 Figure 5 Figure 5 Figure 5 Figure 6 Figure 6 Figure 6 Figure 6 Figure 6 Figure 7 Figure 7 Figure 7 Figure 7 Figure 7 1. Background The mutations induce small changes to the genes. If they cause damage and remain untreated, it drives multifactorial anomalies which are called complex diseases. Cancers are one kind of these complex diseases which are induced by defective driver genes. Among cancers, primary breast cancer as a complex disease is the most commonly diagnosed carcinoma in women worldwide and will be fatal if it progresses towards the secondary-stage. It is the most common and the second most common cause of cancer death in women in developing regions and developed regions, respectively (1). Over the last 10 years, the incidence of breast cancer has increased almost 10 times (2). The concern about this growing trend has prompted oncologists to seek early detection. Nowadays, the molecular technique of next-generation sequencing (NGS), including whole-exome sequencing could generate a large amount of data related to mutated genes called mutation data (3). The analysis of this massive data requires the use of robust computational approaches to exploit the information effectively. Precision oncology by focusing on targeted clinical panel sequencing can be helpful, i.e., a breast cancer-specific NGS panel, including 79 genes has been validated for use in primary and metastatic breast cancer (4). In this way, the advent of bioinformatics tools in parallel with the development of molecular techniques could lead to discovering biomarkers that are efficient in cancer diagnosis and prognosis (5). The machine learning algorithms as one of the computational approaches can be trained with data from countless patients whereas it is too difficult for human physicians and biologists to gain such experience in an entire career or their researches. These models equip experts to make better decisions (6). Some of them are the ensemble classifier (EC) machine learning methods that combine two or several models to optimize the performance of the base components in order to improve data analysis. In previous studies, it has been mentioned that committee approaches can outperform even powerful individual models in many cases (7). Also, investigations show using ensemble models (a.k.a fusion systems) is widely increasing in many fields of inquiry, including the detection of cancers and their subtypes and especially in the area of breast cancer detection. In 1996, a breast cancer dataset, including 699 samples, were analyzed by bagging nearest neighbor classifiers as a fusion system (8). Since then, many ensemble classification methods have been applied to breast cancer prognosis (9). In this regard, we reviewed 42 ensemble methods related to 18 cancers (10). Among these, 22 approaches have been reported for analyzing breast cancer data in the literature (Table 1). Table 1 22 ensemble learning methods concerned with the detection of breast cancer Method name Publication year 1 Bayesian networks-based model integration (11, 12) 2006 and 2019 2 RSS-SCS method (13) 2016 3 Collective approach (correlation, color palette, color proportion, and SVM) (14) 2016 4 Kernel-based Data Fusion Method for Gene Prioritization (15) 2015 5 DECORATE method* (16) 2015 6 HyDRA method* (17) 2015 7 GenEnsemble method* (NBS-IB3-SVM-C4.5 DT) (18) 2014 8 NB (Naïve Bayes) combiner method (19) 2014 9 Evolutionary Ensemble Model (20) 2014 10 smoothed t-statistic SVM (stSVM) (21) 2013 11 SVM Classifiers Fusion (three SVM) (22) 2013 12 COMBINER (Core Module Biomarker Identification)* (23) 2012 13 Ensembles of BioHEL Rule Set (24) 2012 14 Stacking IB3-NBS-RF-SVM method (25) 2012 15 REIS-based ensemble method (26) 2011 16 MRS method (27) 2010 17 Boosting-TWSVM method (28) 2009 18 Bagging and boosting-based TWSVM (29) 2009 19 Feature Subsets Method (30) 2008 20 BNCE method (31) 2007 21 Bayesian Network Classifier (32) 2006 22 enSVM (200 SVM) (33) 2006 * Some proposed methods to discover genomic markers related to breast cancer. In some of these studies, ECs have been used for introducing driver genes associated with breast cancer and the evaluation of genomic biomarkers regarding this cancer. In this work, we propose the EC learning approach called EARN (Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine). It is used to find candidate drivers in primary breast invasive carcinoma (BRCA) and metastatic breast cancer (MBCA) samples from mutation data available in the Cancer Genome Atlas (TCGA) (https://portal.gdc.cancer.gov) and cBioPortal (http://cbioportal.org). The candidate genes introduced by the EC mechanism may already be known as cancers causing genes in databases or can be novel. The candidate genes have the potential to be presented as genomic risk biomarkers after completing the steps of clinical trials (34) and used for personalized targeted therapy (35). Furthermore, there is evidence that driver genes that effectively prognose cancers could be used in therapeutic applications to access more effective therapies (36). The proposed EC method combines decisions of three base classifiers, including non-linear Support Vector Machine (NLSVM) (37), Artificial Neural Network (ANN) (38), and Random Forest (RF) (39). The features for these three classifiers were extracted from four software tools: MutSigCV v.1.4 (40), OncodriveCLUST 0.4.1 (41), OncodriveFM (42), and NetBox 1.0 (43). Overall, we aim to focus on the findings in five steps of BRCA and MBCA prognosis and diagnosis. 1. A list of mutated genes ranked by four software tools based on p -value is presented as the features. 2. Driver and passenger genes predicted by three individual machine learning methods and EARN are introduced and compared. 3. Biological validation of predictions based on gene set enrichment analysis is done and discussed. Indeed, we evaluate the top genes predicted by EARN and three base classifiers for BRCA and MBCA by searching these genes in the list of cancer-associated genes in the public databases, including the Online Mendelian Inheritance in Man (OMIM), the Cancer Gene Census (CGC) (44), the Network of Cancer Genes (NCG) (45, 46), and the human cancer metastasis database (HCMDB) (47). 4. The performance of all machine learning methods is evaluated by monitoring some statistical metrics. 5. Finally, a targeted driver gene panel for MBCA diagnosis based on pathway enrichment analysis (PEA) of top 100 predicted by EARN (EARN 100 ) is proposed. 2. Methods In this study, an ensemble method as a synergistic combination of computational tools has been designed and proposed to find the putative cancer drivers. This fusion system can help to analyze the Whole-Exome Sequencing (WES) data. It consists of four steps: selection of dataset, feature extraction, feature integration, and decision integration (Fig. 1). We have also shared Python source code and other requirements for the implementation of the proposed ensemble machine learning algorithm as the protocol via GitHub (https://github.com/lmirsadeghi/EARN/). 2.1 Selection of dataset In this study, to identify candidate driver genes based on mutations that occur in genes, breast cancer primary and metastasis data have been analyzed. For primary breast cancer, an open-access mutation annotation format (.maf) file was downloaded from TCGA data set regarding BRCA (48). This file includes 90969 masked somatic mutations identified in 17990 genes from 983 tumor samples of BRCA patients that their whole exome had been sequenced by Illumina Genome Analyzer II [see this mutation file (.xlsx) in Additional file 1: Table S1]. Also, for processing of sequences, a bioinformatics pipeline framework called "MuSE Variant Aggregation and Masking" in TCGA has been used. For MBCA, two files (.txt) were downloaded from the cBio Cancer Genomics Portal (49, 50). The first mutation file includes WES of 213 tumor samples from 213 MBCA patients by Illumina HiSeq. It is associated with 22949 somatic mutation counts that occurred among 10791 genes (51). The second file consists of WES of 237 metastasis tumor samples by Illumina GAIIx from 180 patients regarding 24027 somatic mutations identified in 10273 genes (52). Clinical data shows that 86 samples were taken when patients were in the metastatic disease stage, and other samples had been taken less than 4 months prior to the metastatic disease is diagnosed (53). After selecting two initial datasets concerning MBCA [see these mutation files (.xlsx) in Additional file 2: Table S2 and S3], they augmented to build a comprehensive mutation data file, including 46928 somatic mutations identified among 14293 genes from 450 MBCA tumor samples (393 patients). 2.2 Selection of software tools for feature extraction After preparing mutation files, four software tools including MutSigCV v.1.4, OncodriveCLUST 0.4.1, OncodriveFM, and NetBox 1.0 were used to extract the convenient numerical features. The selection of tools for feature extraction was a crucial step to achieve better performance on the final algorithm of the proposed ensemble learning model. We select the four software tools based on evidences of a paper in 2015 on identification and ranking of plausible drivers for BRCA and ovarian (OV) cancer (16). It had been demonstrated that among ten tools for extracting features, OncodriveFM and NetBox generate high sensitivity, especially about BRCA. Also, the sensitivity of OncodriveCLUST tool is high concerning the OV cancer. On the other hand, in both cancers, it had been shown that the positive predictive value (PPV) for NetBox and OncodriveFM is high. MutsigCV was able to propose a large number of drivers in the top 50 genes for BRCA, where at least five other methods had also predicted them as top genes. These advantages led us to use these tools. Practically, these four tools evaluate original mutation files from different aspects and assign a score ( p- value) to genes to show their relevance to disease according to that software’s logic. MutsigCV gets data concerning point mutations and small insertions and deletions (INDELs) from the WES file. After analyzing and estimating mutation frequency, it can identify and introduce a significant list of mutated genes for cancers (40). OncodriveClust software tool is able to identify mutations that generate oncogenes and leads to changes in the function of the proteins. For this purpose, it analyzes synonymous mutations and protein-affecting mutations, including non-synonymous, stop, and splice-site mutations (41). Also, this tool uses data from the Cancer Gene Census (CGC) database (44). for selecting known drivers associated with cancers. OncodriveFM is our next tool which can detect driver genes across tumor samples, identify pathways in cancers, and discover gene modules by using information that is available in the WES file. This data is provided by three methods, including SIFT, PolyPhen2, and MutationAssessor (42). The fourth tool is NetBox, and it can detect driver mutations based on a network. First, a global human interaction network is constructed by this tool. Then, it finds the linker genes between mutated genes for module discovery and identification of candidate drivers (43). Indeed. the concept and criteria of selecting these four software tools are based on the study in 2015 where the performance metrics of ten methods for prediction of plausible driver genes of BRCA and ovarian OV were compared (16). 2.2.1 Feature extraction and feature vector construction In this step, the four software tools explained above are used for the extraction of features from primary and metastasis mutation data files. After running the tools, all genes are ranked based on p- value as output data, and each method assigns a number (0≤ p- value≤1) to genes as numerical features. Therefore, a four-dimensional feature vector is constructed for each gene (Fig. 2a). Since the genes with lower p- value play a more critical role in the development of cancer, we decided to use "1- p- value" as the final numerical feature for each gene. With this plan, the genes that are more important in the occurrence of BRCA and MBCA will also get higher feature values. Different and independent logics behind the ranking mechanisms in the exploited tools guarantee enough diversity between inputs of the ensemble system which is an essential property for efficient fusion methods. 2.3 Classifier model selection Three supervised machine learning methods, including non-linear SVM to learn non-linear functions to separate the classes, ANN, and RF are used as individual classifiers. For selecting these methods, the literature and previous studies were surveyed. We did a comprehensive review regarding fusion systems and the results showed that SVM has been used as a base classifier in many studies or applied as a baseline for comparison between the performance of the different machine learning methods (54–57). Since in this case, the positive and negative training gene set for the implementation of learning algorithms are highly imbalanced, 40 positive genes versus 2151 negative genes (refer to 2.4), a solution must be found. It has been demonstrated the SVM classifier can be a robust method for generating optimal results with imbalanced positive and negative datasets (58), especially when an Instance-weighted SVM algorithm is used (59). So, we weighed this algorithm to get better results. On the other hand, RF is an ensemble machine learning method used as one of the individual classifiers. This method can partially solve the problem of the unbalanced positive and negative training set by bootstrap sampling and can also improve performance, i.e., predictive accuracy reached 88.89% using RF for breast cancer risk prediction (60). The ANN classifier is another machine learning method with long-lasting profound literature. In 1990, Hansen and Salamon integrated multiple neural networks and improved results (61). This method is also widely used in biology studies and has achieved high performance. In 2017, it was shown that ANN could be used for the diagnosis of lung cancer (62). Meanwhile, in this study, the positive training set is small, and recent researches have revealed that ANN may improve performance for problems with small training set sizes and give better performance, especially for problems with time-series data category (63). All of these reasons and criteria led to the selection of these three machine learning methods as base classifiers of the final ensemble system. 2.4 Training and testing In this study, we train separate models for BRCA and MBCA. Some criteria for the selection of training data sets are described below and visualized in Fig. 3a. For testing the performance of models in terms of evaluation metrics (e.g. recall, precision, etc.), we average over 100 trials. In each trial, 3-fold cross-validation with random shuffles is used to calculate the metrics on all data. Finally, the mean and standard deviation of metrics over 100 trials are obtained. Average outputs for cross-validation of the estimator of each model on testing data based on some metrics, including precision, f1 score, recall, accuracy, and Receiver Operating Characteristic-Area under Curve (ROC-AUC) are presented in section 3. 2.4.1 Training data set selection The positive training set of genes for BRCA and MBCA were obtained from searching known genes and mentioned drivers concerning these cancers in several databases, including the OMIM, CGC, NCG, HCMDB, and the Human Protein Atlas (HPA) (https://www.proteinatlas.org/). Also, about selecting negative training gene set, we reviewed a comprehensive list of prior works. Since there is no gold and standard database for a negative set selection, most researchers have used the bootstrap method for resampling, and the negative training genes have been mostly selected randomly. In this study, negative data was selected by counting the occurrence of mutations across all samples in the initial mutation data file, and the genes with the lowest mutation count were used as negative training set (16). It is crucial to note that in both positive and negative training data, we only accepted protein-coding genes. [see further details for training genes in Additional file 3: Supplementary Methods and Additional file 4: Table S4-S7]. 2.5 Genome-wide screening For the genome-wide screening, 20208 homo sapiens genes annotated as protein-coding were downloaded from ftp://ftp.ncbi.nlm.nih.gov/gene/DATA/GENE_INFO/Mammalia/ on February 2019. The proposed ensemble model is applied to 18017 genes for BRCA and 16698 genes for MBCA, after excluding positive and negative training sets (Fig. 3b) [see Additional file 4: Table S8-S10]. 2.6 Implementation of three machine learning algorithms based on feature integration After adding features to the system, and training and testing of learning methods including non-linear SVM, ANN, and RF, they are applied to the protein-coding genes as the unseen data. Each of these methods integrates the features extracted from the initial mutation file (refer to section 2.2.1). We use scikit-learn package to implement our algorithms in python (64). Since this problem is a binary classification of genes based on drivers and passengers, they could label genes based on two indexes -1 and +1 (-1 means passenger genes and +1 means drivers), and also compute a score for each gene, independently (Fig. 2b). 2.7 Implementation of proposed ensemble machine based on decision integration Finally, the decision-making strategy for ensemble machine is based on aggregation of the predicted scores obtained from other machines. We call the proposed EC machine learning method EARN (ensemble of ANN, RF, and non-linear SVM). EARN uses the average of the scores of the outputs of the three basic classifiers to assign a new score (ranging from 0 to 1) to each gene. The genes with higher prediction scores (scores ≥0.5) are labeled as drivers (+1) while the other genes will be passengers (-1). This process has been illustrated in Fig. 2c. 2.8 Biological inferences At this step, all the driver genes introduced for BRCA and MBCA, as well as top genes predicted by learning machines, are searched in the public databases to determine which genes have been already known related to cancer and which ones are new. Pathway enrichment analysis is also performed using ReactomeFIVIz tool ( FDR <0.03) (65–67)to identify the biochemical pathways associated with the candidate genes and examine the biological role of them. It is applied to find biological pathways and patterns related to cancer and other complex diseases. 3. Results This study is an attempt to focus on the findings in five steps of BRCA and MBCA prognosis and diagnosis. 1. A list of mutated genes ranked by four software tools based on p- value is presented as the features. 2. Driver genes and passengers predicted by three individual machine learning methods, NLSVM, ANN, RF, and the proposed EC are introduced. 3. Biological validation of predictions based on gene set enrichment analysis is done 4. Statistical validation of all learning methods based on evaluation metrics is carried out. 5. A targeted gene panel for MBCA based on pathway enrichment analysis (PEA) is proposed. 3.1 BRCA The description of the results for BRCA is presented in Additional file 5: Supplementary Results and Table S11 and S12. However, the comparative results of each algorithm for BRCA and MBCA are illustrated in the next section. 3.2 MBCA 3.2.1 Investigation of the diversity of features extracted from the original mutation file Four software tools are used to extract and rank the list of mutated genes for MBCA as features based on p- value to be used for the machine learning implementation in the next step. The use of multiple tools for generating features creates an effective diverse committee for better classification. It is known that machine learning method can do better discrimination with higher-dimensional feature vectors and perform the classification with higher accuracy (26). To illustrate the existence of diversity in features and also for comparison between results of the tools, we plot the GeneVenn diagram (68) by setting p- value≤0.05 as the threshold. The plotting Venn diagram ( p- value≤0.05) shows that the results of four software tools in the ranking of mutated genes for BRCA and MBCA are varied (fig. 4a). It means that the extracted features by these tools from the original mutation file are sufficiently diverse and can be applied for machine learning implementation step. The comparison shows that five genes, C12orf29, OXCT1, PIK3CA, GCNT4, and C8orf44, are just common among the outputs. Also, PIK3CA has been selected by all software tools in both cases of BRCA and MBCA [see the outputs of software tools for BRCA and MBCA, and comparison among mutated genes ( p- value≤0.05) extracted by these tools for MBCA in Additional file 6: Table S13-S26]. 3.2.2 Outputs of three individual classifiers and EARN The three base classifiers and EARN predicted the labels and scores of 16698 protein-coding genes for MBCA. The percentage of the predicted driver and passenger genes using the four learning methods for BRCA and MBCA has been shown in Fig. 4b. These findings have been presented in an extra file [see Additional file 7: Table S27-S31]. 3.2.3 Investigation of top 100 genes predicted by the four machine learning methods The comparison of the top 100 genes predicted by the four methods using GeneVenn diagram tool shows that 16 genes are predicted by all four machines for MBCA (Fig. 4c). The results of the enrichment of these genes in public databases are considered in Table 2. Other common and unique driver genes predicted by methods are presented in the extra file [see Additional file 8: Table S32-S41]. Also, among the outputs of EARN 100 , BDNF, PRKCG, TH, PRKCD, and PIP5K1B are just predicted by this learning machine in the list of top 100 genes. Among these five genes, BDNF and PRKCG have been already introduced regarding metastatic cancers but the others are new. Table 2 The 16 common genes predicted by all machines in the top 100. The confirmed genes as the known genes related to different primary cancers or primary breast tumors in OMIM, CGC, and NCG databases have been marked in the last two columns Symbol NSCGMCH (#) NSCGMBH (#) PKGECC PKGEBC OXCT1* #N/A #N/A #N/A #N/A KDR** 7 2 ✓ #N/A APEX1* #N/A #N/A #N/A #N/A GCM2* #N/A #N/A #N/A #N/A UNC13D* #N/A #N/A #N/A #N/A NCOR1 #N/A #N/A ✓ ✓ KRAS 20 #N/A ✓ ✓ THAP3* #N/A #N/A #N/A #N/A SERPINE2 1 #N/A #N/A #N/A BATF* #N/A #N/A #N/A #N/A C8orf44* #N/A #N/A #N/A #N/A C12orf29* #N/A #N/A #N/A #N/A ZNF546* #N/A #N/A #N/A #N/A KDM6B 1 #N/A #N/A #N/A GCNT4* #N/A #N/A #N/A #N/A FOXA1 #N/A #N/A ✓ ✓ *Ten new genes that have not already been introduced in the databases. ** KDR is confirmed in HCMDB related to metastatic breast cancer in two studies. NSCGMCH: N umber of s tudies that have c ited g enes related to different m etastatic c ancers in the H CMDB NSCGMBH: N umber of s tudies that have c ited these g enes related to m etastatic b reast cancer in the H CMDB PKGECC: P redicted k nown g enes by E C associated with different c ancers that are c onfirmed in OMIM, CGC, and NCG PKGEBC: P redicted k nown g enes by E C associated with B reast cancer that are c onfirmed in OMIM, CGC, and NCG 3.2.4 Biological validation of predictions based on gene set enrichment analysis The biological analysis of genes predicted by EARN is performed based on two plans; (a) analysis of the results based on all predicted driver genes (labeled as +1) and (b) analysis of the findings based on the top-scoring genes. To investigate outputs of the EARN for MBCA from a biological point of view based on the label, we analyzed the results concerning the public databases. There is a gene-metastasis association data file (.xls) in the HCMDB that lists 2240 genes related to metastatic cancers based on experiments performed in various studies. 622 genes out of these genes were introduced for metastatic breast cancer specifically. It should be noted that all 37 genes in the positive training gene set have overlap with the gene list of HCMDB in relation to both of different metastatic cancers and metastatic breast cancer. These 37 genes must be excluded to analyze the results. Table 3 (a, b) present the frequency of driver genes enriched in the public databases for MBCA and BRCA. Table 3 The enrichment rate of driver genes predicted by EARN. (a) MBCA, (b) BRCA (a) MBCA All different cancers Metastatic breast cancer HCMDB HCMDB PGECCH (#) RGCHP (#) PGECH (%) PGEMCH (#) RGMHP (#) PGEMH (%) 292 2203 13.25% 73* 585 12.48% (b) BRCA All different cancers Breast cancer OMIM, CGC, and NCG OMIM, CGC, and NCG PKGECC (#) RKGCP (#) PKGECC (%) PKGEBC (#) RKGBP (#) PKGEBC (%) 1398 2403 58.18% 145 201 72.14% *These 73 genes have been also cited in 108 studies of HCMDB [see Additional file 9: S42] PGECCH: P redicted g enes by E C associated with different metastatic c ancers that are c onfirmed in H CMDB RGCHP: R emained g enes related to different metastatic c ancers in the H CMDB after excluding p ositive training set PGEMCH: P redicted g enes by E C associated with M etastatic breast cancer that are c onfirmed in H CMDB PKGECC: P redicted k nown g enes by E C associated with different c ancers that are c onfirmed in OMIM, CGC, and NCG RKGCP: R emained k nown g enes related to different c ancers in the p ublic databases after excluding positive training set PKGEBC: P redicted k nown g enes by E C associated with b reast cancer that are c onfirmed in OMIM, CGC, and NCG RKGBP: R emained k nown g enes related to b reast cancer in the p ublic databases after excluding positive training set Also, the top 50 genes predicted by all learning methods for MBCA are searched in the list of metastatic cancer-associated genes in the HCMDB. The comparison shows the enrichment score of 24%, 22%, and 16% for RF, ANN, and NLSVM compared to 24% for EARN. Although the value of enrichment in the top 50 is the same for EARN and RF, the number of studies that introduce these enriched genes is 59 for the EARN method compared to 22 for RF. Table 4 presents these genes and also provides more information about them. Table 4 12 driver genes predicted by EARN 50 which are confirmed for metastatic cancers in the HCMDB. Also, the rank number, score, and mutation count for these genes are provided in the table. The confirmed genes as the known genes related to any primary cancers or primary breast tumors in OMIM, CGC, and NCG databases have been marked in the last two columns Symbol Prediction score Rank PSMM (51) PSMM (52) NSCGMCH NSCGMBH MCMGM PKGECC PKGEBC APEX1 0.900511991 5 0.50% 1.70% 1 #N/A 5 #N/A #N/A ARID1A 0.895213526 11 2.40% 5.10% 2 #N/A 24 ✓ ✓ KDM6B 0.894029187 13 1.40% 4.60% 1 #N/A 16 #N/A #N/A TBX3 0.893837209 14 2.80% 5.10% 1 #N/A 21 ✓ ✓ KDR* 0.890079401 17 0.90% 1.70% 7 2 9 ✓ #N/A SERPINE2 0.889205475 19 0.90% 0.80% 1 #N/A 4 #N/A #N/A TBL1XR1 0.871240171 27 0.90% 0.80% 2 #N/A 4 ✓ ✓ KRAS 0.868267682 30 1.40% 1.70% 20 #N/A 7 ✓ ✓ NOS3 0.861560093 31 2.40% 2.10% 1 #N/A 12 #N/A #N/A RAPGEF3 0.851947423 42 #N/A 2.50% 2 #N/A 6 #N/A #N/A SELE* 0.847865292 49 0.90% 1.30% 12 1 5 #N/A #N/A MME* 0.847698297 50 0.90% 2.50% 9 1 9 #N/A #N/A * These genes have been specifically introduced concerning metastatic breast cancer. PSMM: P ercentage of s amples with one or more m utations based on initial m utation file NSCGMCH: N umber of s tudies that have c ited g enes related to different m etastatic c ancers in the H CMDB NSCGMBH: N umber of s tudies that have c ited g enes related to m etastatic b reast cancer in H CMDB MCMGM: M utation c ounts for m utated g enes across 450 metastasis tumor samples based on the initial m utation file PKGECC: P redicted k nown g enes by E C associated with different c ancers that are c onfirmed in OMIM, CGC, and NCG PKGEBC: P redicted k nown g enes by E C associated with B reast cancer that are c onfirmed in OMIM, CGC, and NCG Furthermore, 38 genes listed by EARN 50 have not been introduced in the HCMDB related to any metastatic cancers. So, these genes can be considered as new genes for more investigations [see Additional file 9: Table S43]. 3.2.5 Statistical validation of three individual classifiers and EARN based on evaluation measures For MBCA, a comparison of the metrics based on 3-fold cross-validation on the test data shows that EARN and ANN achieve the best precision with zero FPR. Also, accuracy, F1 score, average precision, and recall for EARN and ANN are better than the others, especially compared with NLSVM. It can be also observed that EARN has the best ROC-AUC (99.24%). Thus, in overall, the proposed EARN outperforms the other three learning methods. For comparison, evaluation metrics of learning methods for MBCA and BRCA are presented in table 5 (a, b). Table 5 Validation of four learning methods by some evaluation metrics. (a) MBCA, (b) BRCA Method name F1 score False Positive Rate Maximum Precision Average-Precision recall ROC-AUC* (a) MBCA EARN 0.7961 0 1.0 0.8266 0.6701 0.9924 RF 0.7560 0.0008 0.9069 0.7873 0.6603 0.9418 ANN 0.7990 0 1.0 0.8074 0.6733 0.9680 NLSVM 0.3972 0.0154 0.3092 0.5852 0.5885 0.9770 (b) BRCA EARN 0.9313 0 1.0 0.9585 0.8749 0.9979 RF 0.8864 0.0019 0.9061 0.9171 0.8774 0.9719 ANN 0.8996 0 1.0 0.9417 0.8225 0.9873 NLSVM 0.5441 0.0279 0.4460 0.8590 0.8422 0.9926 * Receiver Operating Characteristic-Area under Curve The comparative survey in table 5 shows when we use a larger mutation dataset (983 tumor samples for BRCA vs. 450 tumor samples for MBCA) for feature extraction, where positive set is larger (40 for BRCA vs. 37 for MBCA), and negative set is smaller (2151 for BRCA vs. 3473 for MBCA), EARN achieves better statistical results. Among all statistical validation metrics, F1 score as a measure of combining the precision and recall has been used to compare performance of the learning methods for both BRCA and MBCA (Fig. 4d). 3.3 BRCA and MBCA 3.3.1 Targeted gene panel discovery for MBCA based on pathway enrichment analysis (PEA) In this section, a pathway-based biological analysis is carried out by ReactomeFIVIz tool (65–67). For EARN 100 , we find 63 ( FDR <0.03) such pathways for BRCA and 42 ( FDR <0.03) such pathways for MBCA. It is observed that 14 ( FDR <0.03) enriched pathways are common among BRCA and MBCA (Fig. 5a), [see these specific and common pathways and the genes involved in each pathway in Additional file 10: Table S44]. These enriched pathways for BRCA are a subset of the other seven main pathways: Extracellular matrix organization, Signal Transduction, Gene expression (Transcription), Immune System, Hemostasis, Developmental Biology, and Metabolism of RNA. Also, the main pathways of MBCA include Gene expression (Transcription), Signal Transduction, Chromatin organization, Circadian Clock, Organelle biogenesis and maintenance, Neuronal System, and Metabolism. The common and specific main pathways ( FDR <0.03) of BRCA and MBCA, and the frequency of genes involved in these main pathways are compared in Fig. 5b and Table 6. Given this, it can be found two ( FDR <0.03) such common main pathways consist of Signal Transduction and Gene expression (Transcription) for BRCA and MBCA, and 5 ( FDR <0.03) such specific main pathways for each of them. Table 6 The common and specific main pathways for BRCA and MBCA Number Pathways BRCA MBCA Number of genes Name of genes Number of genes Name of genes The specific main pathways for BRCA 1 Extracellular matrix organization 8 DCN, FN1, ICAM1, ITGA4, ITGAM, ITGAV, ITGB3, ITGB5 0 None 2 Immune System 15 FN1, GAB2, ICAM1, IL1RAPL1, IL1RN, IL2RB, ITGAM, ITGAV, ITGB5, JAK1, MSN, POU2F1, PTPN11, SMARCA4, SYK 0 None 3 Hemostasis 12 EGF, FN1, GRB7, ITGA4, ITGAM, ITGAV, ITGB3, PIK3CG, PRKCZ, PTPN11, SERPINA1, SYK 0 None 4 Developmental Biology 9 ACVR1B, GAB1, GAB2, GRB7, PTPN11, RELN, SMAD2, SMAD4, VLDLR 0 None 5 Metabolism of RNA 6 CPSF1, CPSF3, PCF11, PRPF40A, SF3A1, SF3B1 0 None The specific main pathways for MBCA 6 Chromatin organization 0 None 7 TBL1XR1, NCOR1, HDAC3, GPS2, ACTB, KDM6B, PRMT1 7 Circadian Clock 0 None 2 NCOR1, HDAC3 8 Organelle biogenesis and maintenance 0 None 4 TBL1XR1, SIRT4, NCOR1, HDAC3 9 Neuronal System 0 None 7 ABAT, KPNA2, PRKCG, CACNA1E, PLCB1, GRIN1, KRAS 10 Metabolism 0 None 5 TBL1XR1, SIN3A, NCOR1, HDAC3, GPS2 The common main pathways for BRCA and MBCA 11 Signal Transduction 25 ACVR1B, EGF, ERBB3, FLT1, FN1, GAB1, GAB2, GRB7, ITGAV, ITGB3, JAK1, NOTCH4, NR4A1, PARD3, PPARG, PRKCZ, PTEN, PTPN11, RUNX1, SMAD2, SMAD4, SMURF1, SYK, TFDP1, TGFBR2 25 ACTB, AR, BDNF, BUB1B, CBFB, COL4A3, FOXA1, KDR, KPNA2, KRAS, NCOR1, NOS3, PDGFD, PIK3R1, PKN2, PLCB1, PRKCD, PRKCG, PRMT1, PTPRJ, RUNX1, STAG1, STAT1, WAS, YWHAE 12 Gene expression (Transcription) 19 ABL1, CBFB, CPSF1, CPSF3, MED23, NBN, NOTCH4, NR4A1, PCF11, POU2F1, PPARG, PTEN, PTPN11, RUNX1, SMAD2, SMAD4, SMARCA4, SMURF1, TFDP1 14 AR, BDNF, CBFB, GPS2, HDAC3, KLF4, KRAS, NCOR1, PRMT1, RUNX1, SIN3A, STAT1, TBL1XR1, YWHAE Further investigation in Table 6 shows that 16 genes contribute to five enriched specific main pathways of MBCA. Among them, four genes are involved in more than one main pathway. In particular, NCOR1 and HDAC3 are engaged in four pathways. In three out of five pathways TBL1XR1 is active, and GPS2 gets involved in two pathways. Table 7 introduces 16 genes that are enriched in these five main pathways and provides more information about them. Table 7 The plausible driver genes involved in the proposed main pathways related to MBCA Specific main pathways PPDMB KGCC KGBC IGMC IGMB Chromatin organization Circadian Clock Organelle biogenesis and maintenance Neuronal System Metabolism NCOR1 1 1 #N/A #N/A ✔ ✔ ✔ ✔ HDAC3* #N/A #N/A #N/A #N/A ✔ ✔ ✔ ✔ TBL1XR1 1 1 2 #N/A ✔ ✔ ✔ SIRT4 1 #N/A #N/A #N/A ✔ ABAT* #N/A #N/A #N/A #N/A ✔ KRAS 1 1 20 #N/A ✔ GRIN1* #N/A #N/A #N/A #N/A ✔ PLCB1* #N/A #N/A #N/A #N/A ✔ CACNA1E 1 #N/A #N/A #N/A ✔ PRKCG 1 #N/A 1 #N/A ✔ KPNA2* #N/A #N/A #N/A #N/A ✔ GPS2 1 1 #N/A #N/A ✔ ✔ SIN3A 1 #N/A #N/A #N/A ✔ ACTB 1 #N/A 1 #N/A ✔ KDM6B #N/A #N/A 1 #N/A ✔ PRMT1 #N/A #N/A 2 #N/A ✔ * Five new genes that have not been already introduced in the public databases PPDMB: P roposed p lausible d river related to m etastatic b reast cancer KGCC: K nown g enes related to c ancers that are c onfirmed in OMIM, CGC, and NCG KGBC: K nown g enes related to b reast cancer that are c onfirmed in OMIM, CGC, and NCG IGMC: I ntroduced g enes related to different m etastatic c ancers in HCMDB IGMB: I ntroduced g enes related to m etastatic b reast cancer in HCMDB This gene set can be considered as a targeted biomarker panel in the case of metastatic breast cancer to examine more in the next molecular and clinical analysis phase. More investigations on these genes can hopefully be helpful in MBCA prognosis and diagnosis. Table 7 shows that five genes, HDAC3, ABAT, GRIN1, PLCB1, and KPNA2 are new and not confirmed in the public databases for cancer prognosis. However, there is some evidence to suggest that these genes play a clinical role in cancer progression. HDAC3 contributes to four pathways alongside NCOR1. The other four genes engage in the Neuronal System pathway. The recent investigations on Basal-like breast cancer (BLBC), the most aggressive subtype of this cancer, have documented the expression of ABAT was considerably decreased in this cancer (69). Besides, alterations in the expression levels of ABAT have been reported in the promotion of breast cancer (70). ABAT was also identified as a biomarker for endocrine-responsiveness breast cancer patients (71). Furthermore, GRIN1 encodes GluN1 subunit of N-methyl-D-aspartate receptor (NMDAR). It has been shown that this subunit in more than 90% of all breast cancer subtypes is uniformly expressed to promote Breast-to-brain metastasis (B2BM) (72). Recently, the role of HDAC3 in the deregulation of P53 pathway in the aneuploid cancer cell lines has been analyzed (73). Also, HDAC3 is overexpressed in breast cancer patients. It has been illustrated that breast cancer stem cells, which are resistant to treatment and are responsible for metastasis, are the target of the histone deacetylase (HDAC) inhibitors (74). On the other, The results of enrichment in cBioPortal show that the above-mentioned 16 genes are altered in 243 (54%) of 450 MBCA samples in two studies performed in 2016 (51) and 2017 (52). Genomic alterations (Fig. 6) in these genes have been visualized using OncoPrint component (49, 50). Among them, the highest percentage of somatic mutation frequency (SMF) is observed in CACNA1E, NCOR1, KDM6B, and GPS2. Using the Needle Plot component (49, 50), we visualize SMF and can also map mutations on the linear protein and its domains for these four genes (Fig. 7). 4. Discussion In this work, we proposed an EC machine learning method called EARN, combining three base classifiers to predict and estimate the potential of plausible driver genes in BRCA and MBCA. Leveraged by both feature fusion and decision fusion, the proposed ensemble model made better decisions in comparison with base classifiers, especially in the list of the top genes. Although EARN uses the simple average operator for aggregating the decisions of the three base learners to predict the driver genes, it could find some new genes in the list of EARN 100 , which were not observed in the top 100 of the individual classifiers. It can be rational evidence for using the ensemble systems for gene prioritization. For biological validation of outputs and after the enrichment of EARN 50 in the public databases, where the ensemble learning method uses most of the power to discriminate and predict, we could obtain the enrichment rate of 52% for BRCA, which outperforms the three individual classifiers. For MBCA, the enrichment of EARN 50 in the HCMDB resulted in an enrichment rate of 24%, which is better than the two base classifiers, NLSVM and ANN, while being comparable to RF. The results are also analyzed using a statistical test with cross-validation. The evaluation of results showed that EARN performs well, especially for BRCA. In the case of BRCA, the open-access mutation annotation format (.maf) file is larger and the mutation data is obtained from more samples (983 BRCA tumor samples vs. 450 MBCA tumor samples). Thus, the proper features could be extracted. Finally, the performance of EARN for ranking human protein-coding genes is improved. Further, to evaluate the possibility of enhancement in the combination of the base classifiers results, we tried StackingCVClassifier, an effective ensemble-learning meta-classifier for stacking ( 75 , 76 ). For BRCA, there was no improvement in the results. It could be because the results of the originally proposed ensemble model were good enough. While the metrics such as F1 score (81.31% vs. 79.61%) and recall (69.62 vs. 67.02) were slightly improved for MBCA. Finally, the existence of specific enriched pathways by ReactomeFIVIz ( FDR < 0.03) for the top genes predicted by EARN for BRCA and MBCA led us to suggest a gene panel regarding metastatic breast cancer. In present study, we faced some limitations to find the appropriate drivers of MBCA. This fact that the original mutation datasets involved in the whole-exome sequencing of the tumor samples of the metastatic breast cancer patients are small. Also, the lack of definitive driver genes confirmed in the public databases for metastatic cancers makes it difficult to select a positive training set. These issues decreased the performance of EARN for MBCA in comparison with BRCA. Further, the result of enriching all predicted genes by EARN for BRCA in the OMIM, CGC, and NCG was encouraging (72.14%, Please refer to Table 3 (b)). But, the result of the enrichment of the predicted genes by EARN for MBCA was not satisfactory (12.48%, see Table 3 (a)). This may be due to the lack of sufficient studies on metastatic cancers, and particularly because of the limited databases regarding metastatic cancers to enrich driver genes. 5. Conclusions Since using computational methods such as ensemble machine learning approaches are less expensive than bio-molecular techniques, it can help to significantly reduce the search space for bio-molecular and medical science researchers in the identification of plausible driver genes to facilitate prognosis and diagnosis of complex diseases. In this work, we mainly focused on the use of genomics data. Meanwhile, the changes of epigenomic, genomic, transcriptional, and proteomic that occur during progression to metastatic encourage us to use multi-omics integration ( 77 ). It has been demonstrated that multi-Omics data integration can improve predictive performance ( 78 ) (e.g., it has been applied to predict robust biomarkers of drug efficacy for targeted therapies in triple-negative breast cancer ( 79 )). A direction of future research would be to apply a combination of different levels of data, including genomics, epigenomics, transcriptomics, proteomics, metabolomics, and microbiomics data to optimize the ensemble system for introducing Omics-driven markers. In the end, we emphasize this research needs clinical trials to be validated and to evaluate the potential of the proposed drivers for discrimination between different stages of cancers. The limited number of markers obtained from the trial validation assists precision oncologists to design compact targeted panels that eliminate the need for whole-genome/exome sequencing. Abbreviations B2BM: Breast-to-brain metastasis; BLBC:Basal-like breast cancer; BRCA:primary breast invasive carcinoma; CGC:Cancer Gene Census; DECORATE:Diverse Ensemble Creation by Oppositional Relabeling of Artificial Training Examples; EARN:Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine; EC:ensemble classifier; FPR:false-positive rate; GSEA:gene set enrichment analysis; HCMDB:human cancer metastasis database; HPA:Human Protein Atlas; HyDRA:Hybrid Distance-score Rank Aggregation; INDELs:insertions and deletions; maf:mutation annotation format; MBCA:metastatic breast cancer; NCG:Network of Cancer Genes; NGS:next-generation sequencing; NLSVM:non-linear Support Vector Machine; NMDAR:N-methyl-D-aspartate receptor; OMIM:Online Mendelian Inheritance in Man; OV:ovarian; PEA:pathway enrichment analysis; PPV:positive predictive value; RF:Random Forest; ROC-AUC:Receiver Operating Characteristic-Area under Curve; SMF:somatic mutation frequency; TCGA:The Cancer Genome Atlas; WES:Whole-Exome Sequencing Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials All required data is available in Additional files 1-3. We have also shared Python source code and other requirements for implementation of the proposed ensemble machine learning algorithm as the protocol via GitHub (https://github.com/lmirsadeghi/EARN/). Competing interests The authors declare that they have no competing interests. Funding This study was financially supported by grant No: 960903 of the Biotechnology Development Council of the Islamic Republic of Iran . Authors' contributions LM and KK developed the concept, designed the research. LM performed the required research works and developed the solution under the joint supervision of KK, RHH, and AMBM. LM wrote the manuscript and contributed to visualize results. LM and KK contributed to the interpretation of the data and discussion. KK, RHH, and AMBM gave final approval for publication. Acknowledgements The authors like to thank Dr. Hossein Hajimirsadeghi[1] for his useful advice and invaluable help during this research. Authors' information 1 Department of Biology, Faculty of Science, Payame Noor University, Tehran, Iran. 2 Laboratory of Genomics and Epigenomics (LGE), Department of Biochemistry, Institute of Biochemistry and Biophysics (IBB), University of Tehran, Tehran, Iran. 3 Laboratory of Complex Biological Systems and Bioinformatics (CBB), Department of Bioinformatics, Institute of Biochemistry and Biophysics (IBB), University of Tehran, Tehran, Iran. [1] https://hossein-h.github.io/ E-mail: [email protected] References Kumar A, Singla A. Epidemiology of Breast Cancer: Current Figures and Trends. In: Preventive Oncology for the Gynecologist. Springer; 2019. p. 335–9. Zhao D, Qiao J, He H, Song J, Zhao S, Yu J. TFPI2 suppresses breast cancer progression through inhibiting TWIST-integrin α5 pathway. Mol Med. 2020;26:1–10. Sheikine Y, Kuo FC, Lindeman NI. Clinical and technical aspects of genomic diagnostics for precision oncology. J Clin Oncol. 2017;35(9):929–33. Smith NG, Gyanchandani R, Shah OS, Gurda GT, Lucas PC, Hartmaier RJ, et al. Targeted mutation detection in breast cancer using MammaSeq TM . Breast Cancer Res. 2019;21(1):22. Kulasingam V, Diamandis EP. Strategies for discovering novel cancer biomarkers through utilization of emerging technologies. Nat Rev Clin Oncol. 2008;5(10):588. Rajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347–58. Baronti F, Micheli A, Passaro A, Starita A. Machine learning contribution to solve prognostic medical problems. Outcome Predict Cancer. 2006;261. Breiman L. Bagging predictors. Mach Learn. 1996;24(2):123–40. Hosni M, Abnane I, Idri A, de Gea JMC, Alemán JLF. Reviewing Ensemble Classification Methods in Breast Cancer. Comput Methods Programs Biomed. 2019; Mirsadeghi L, Banaei-Moghaddam AM, Beh-Afarin SR, Haji R. A post-method condition analysis of using ensemble machine learning for cancer prognosis and diagnosis: a systematic review. Gevaert O, De Smet F, Timmerman D, Moreau Y, De Moor B. Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks. Bioinformatics. 2006;22(14):e184–90. Moriyama T, Imoto S, Hayashi S, Shiraishi Y, Miyano S, Yamaguchi R. A Bayesian model integration for mutation calling through data partitioning. Bioinformatics. 2019; Cheriguene S, Azizi N, Zemmal N, Dey N, Djellali H, Farah N. Optimized Tumor Breast Cancer Classification Using Combining Random Subspace and Static Classifiers Selection Paradigms. In: Applications of Intelligent Optimization in Biology and Medicine. Springer; 2016. p. 289–307. Les T, Markiewicz T, Osowski S, Kozlowski W, Jesiotr M. Fusion of FISH image analysis methods of HER2 status determination in breast cancer. Expert Syst Appl. 2016;61:78–85. Zakeri P, Elshal S, Moreau Y. Gene prioritization through geometric-inspired kernel data fusion. In: Bioinformatics and Biomedicine (BIBM), 2015 IEEE International Conference on. IEEE; 2015. p. 1559–65. Liu Y, Tian F, Hu Z, DeLisi C. Evaluation and integration of cancer gene classifiers: identification and ranking of plausible drivers. Sci Rep. 2015;5. Kim M, Farnoud F, Milenkovic O. HyDRA: gene prioritization via hybrid distance-score rank aggregation. Bioinformatics. 2015;31(7):1034–43. Reboiro-Jato M, Díaz F, Glez-Peña D, Fdez-Riverola F. A novel ensemble of classifiers that use biological relevant gene sets for microarray classification. Appl Soft Comput. 2014;17:117–26. Kuncheva LI, Rodríguez JJ. A weighted voting framework for classifiers ensembles. Knowl Inf Syst. 2014;38(2):259–75. Janghel RR, Shukla A, Sharma S, Gnaneswar A V. Evolutionary Ensemble Model for Breast Cancer Classification. In: International Conference in Swarm Intelligence. Springer; 2014. p. 8–16. Cun Y, Fröhlich H. Network and data integration for biomarker signature discovery via network smoothed t-statistics. PLoS One. 2013;8(9):e73074. Azizi N, Tlili-Guiassa Y, Zemmal N. A computer-aided diagnosis system for breast cancer combining features complementarily and new scheme of SVM classifiers fusion. Int J Multimed Ubiquitous Eng. 2013;8(4):45–58. Yang R, Daigle BJ, Petzold LR, Doyle FJ. Core module biomarker identification with network exploration for breast cancer metastasis. BMC Bioinformatics. 2012;13(1):1. Glaab E, Bacardit J, Garibaldi JM, Krasnogor N. Using rule-based machine learning for candidate disease gene prioritization and sample classification of cancer gene expression data. PLoS One. 2012;7(7):e39932. Reboiro-Jato M, Glez-Peña D, Díaz F, Fdez-Riverola F. A novel ensemble approach for multicategory classification of DNA microarray data using biological relevant gene sets. Int J Data Min Bioinform. 2012;6(6):602–16. Lederman D, Wang X, Zheng B, Sumkin JH, Tublin M, Gur D. Fusion of classifiers for REIS-based detection of suspicious breast lesions. In: SPIE Medical Imaging. International Society for Optics and Photonics; 2011. p. 79661C-79661C. Zeng T, Liu J. Mixture classification model based on clinical markers for breast cancer prognosis. Artif Intell Med. 2010;48(2):129–37. Zhang X. Boosting twin support vector machine approach for MCs detection. In: Information Processing, 2009 APCIP 2009 Asia-Pacific Conference on. IEEE; 2009. p. 149–52. Zhang X, Gao X, Wang M. MCs detection approach using Bagging and Boosting based twin support vector machine. In: Systems, Man and Cybernetics, 2009 SMC 2009 IEEE International Conference on. IEEE; 2009. p. 5000–505. Djebbari A, Liu Z, Phan S, Famili F. An ensemble machine learning approach to predict survival in breast cancer. Int J Comput Biol Drug Des. 2008;1(3):275–94. Alam KMR, Islam MM. Combining boosting with negative correlation learning for training neural network ensembles. In: 2007 International Conference on Information and Communication Technology. IEEE; 2007. p. 68–71. Franke L, Bakel H Van, Fokkens L, Jong ED De, Egmont-petersen M, Wijmenga C. Reconstruction of a Functional Human Gene Network , with an Application for Prioritizing Positional Candidate Genes. Am J Hum Genet. 2006;78(June):1011–25. Peng Y. Integration of gene functional diversity for effective cancer detection. Int J Syst Sci. 2006;37(13):931–8. Matsui S. Genomic biomarkers for personalized medicine: development and validation in clinical studies. Comput Math Methods Med. 2013;2013. Huang L, Jiang X-L, Liang H-B, Li J-C, Chin L-H, Wei J-P, et al. Genetic profiling of primary and secondary tumors from patients with lung adenocarcinoma and bone metastases reveals targeted therapy options. Mol Med. 2020;26(1):1–11. Lan Y, Zhao E, Luo S, Xiao Y, Li X, Cheng S. Revealing clonality and subclonality of driver genes for clinical survival benefits in breast cancer. Breast Cancer Res Treat. 2019;175(1):91–104. Baesens B, Viaene S, Van Gestel T, Suykens J, Dedene G, De Moor B, et al. Least squares support vector machine classifiers: an empirical evaluation. DTEW Res Rep 0003. 2000;1–16. Maclin PS, Dempsey J, Brooks J, Rand J. Using neural networks to diagnose cancer. J Med Syst. 1991;15(1):11–9. Breiman L. Random forests. Mach Learn. 2001;45(1):5–32. Lawrence MS, Stojanov P, Polak P, Kryukov G V, Cibulskis K, Sivachenko A, et al. Mutational heterogeneity in cancer and the search for new cancer-associated genes. Nature. 2013;499(7457):214–8. Tamborero D, Gonzalez-Perez A, Lopez-Bigas N. OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes. Bioinformatics. 2013;29(18):2238–44. Gonzalez-Perez A, Lopez-Bigas N. Functional impact bias reveals cancer drivers. Nucleic Acids Res. 2012;40(21):e169–e169. Cerami E, Demir E, Schultz N, Taylor BS, Sander C. Automated network analysis identifies core pathways in glioblastoma. PLoS One. 2010;5(2):e8918. Futreal PA, Coin L, Marshall M, Down T, Hubbard T, Wooster R, et al. A census of human cancer genes. Nat Rev cancer. 2004;4(3):177. An O, Pendino V, D’Antonio M, Ratti E, Gentilini M, Ciccarelli FD. NCG 4.0: the network of cancer genes in the era of massive mutational screenings of cancer genomes. Database. 2014;2014:bau015. Repana D, Nulsen J, Dressler L, Bortolomeazzi M, Venkata SK, Tourna A, et al. The Network of Cancer Genes (NCG): a comprehensive catalogue of known and candidate cancer genes from cancer sequencing screens. Genome Biol. 2019;20(1):1. The experimentally supported gene-metastasis association data. 2017. https://hcmdb.isanger.com/images/hcmdb/gene_publication.xls. Accessed 22-Jun-2017. TCGA.BRCA.muse.b8ca5856-9819-459c-87c5-94e91aca4032.DR-10.0.somatic.maf.gz. 2018. https://portal.gdc.cancer.gov/files/b8ca5856-9819-459c-87c5-94e91aca4032. Accessed 23-Aug-2018. Cerami E, Gao J, Dogrusoz U, Gross BE, Sumer SO, Aksoy BA, et al. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. AACR; 2012. Gao J, Aksoy BA, Dogrusoz U, Dresdner G, Gross B, Sumer SO, et al. Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal. Sci Signal. 2013;6(269):pl1–pl1. Lefebvre C, Bachelot T, Filleron T, Pedrero M, Campone M, Soria J-C, et al. Mutational profile of metastatic breast cancers: a retrospective analysis. PLoS Med. 2016;13(12):e1002201. Wagle N, Painter C, Anastasio E, Dunphy M, McGillicuddy M, Kim D, et al. The Metastatic Breast Cancer (MBC) project: Accelerating translational research through direct patient engagement. American Society of Clinical Oncology; 2017. cBioPortal/datahub-study-curation-tools. 2019. https://github.com/cBioPortal/datahub-study-curationtools/tree/master/split_data_clinical_sample_patient. Accessed 11-Jan-2019. García-Díaz P, Sánchez-Berriel I, Martínez-Rojas JA, Diez-Pascual AM. Unsupervised feature selection algorithm for multiclass cancer classification of gene expression RNA-Seq data. Genomics. 2020;112(2):1916–25. Kim S, Park T, Kon M. Cancer survival classification using integrated data sets and intermediate information. Artif Intell Med. 2014;62(1):23–31. Dashtban M, Balafar M, Suravajhala P. Gene selection for tumor classification using a novel bio-inspired multi-objective approach. Genomics. 2018;110(1):10–7. Bhanot G, Alexe G, Venkataraghavan B, Levine AJ. A robust meta‐classification strategy for cancer detection from MS data. Proteomics. 2006;6(2):592–604. Palade V. Class imbalance learning methods for support vector machines. 2013; Wang X, Liu X, Matwin S. A distributed instance-weighted SVM algorithm on large-scale imbalanced datasets. Proc - 2014 IEEE Int Conf Big Data, IEEE Big Data 2014. 2015;45–51. Ming C, Viassolo V, Probst-Hensch N, Chappuis PO, Dinov ID, Katapodi MC. Machine learning techniques for personalized breast cancer risk prediction: comparison with the BCRAT and BOADICEA models. Breast Cancer Res. 2019;21(1):75. Polikar R. Ensemble based systems in decision making. Circuits Syst Mag IEEE. 2006;6(3):21–45. Duan X, Yang Y, Tan S, Wang S, Feng X, Cui L, et al. Application of artificial neural network model combined with four biomarkers in auxiliary diagnosis of lung cancer. Med Biol Eng Comput. 2017;55(8):1239–48. Walczak S. Artificial neural networks. In: Encyclopedia of Information Science and Technology, Fourth Edition. IGI Global; 2018. p. 120–31. Pedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine learning in Python. J Mach Learn Res. 2011;12(Oct):2825–30. Fabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, et al. The reactome pathway knowledgebase. Nucleic Acids Res. 2017;46(D1):D649–55. Wu G, Haw R. Functional Interaction Network Construction and Analysis for Disease Discovery. In: Protein Bioinformatics. Springer; 2017. p. 235–53. Fabregat A, Sidiropoulos K, Viteri G, Forner O, Marin-Garcia P, Arnau V, et al. Reactome pathway analysis: a high-performance in-memory approach. BMC Bioinformatics. 2017;18(1):142. Bioinformatics & Evolutionary Genomics. 2018. http://bioinformatics.psb.ugent.be/webtools/Venn/. Accessed 20 Nov 2018. Chen X, Cao Q, Liao R, Wu X, Xun S, Huang J, et al. Loss of ABAT-Mediated GABAergic System Promotes Basal-Like Breast Cancer Progression by Activating Ca2+-NFAT1 Axis. Theranostics. 2019;9(1):34. Zhao G, Li N, Li S, Wu W, Wang X, Gu J. High methylation of the 4-aminobutyrate aminotransferase gene predicts a poor prognosis in patients with myelodysplastic syndrome. Int J Oncol. 2019;54(2):491–504. Sas L, Lardon F, Vermeulen PB, Hauspy J, Van Dam P, Pauwels P, et al. The interaction between ER and NFκB in resistance to endocrine therapy. Breast Cancer Res. 2012;14(4):212. Zeng Q, Michael IP, Zhang P, Saghafinia S, Knott G, Jiao W, et al. Synaptic proximity enables NMDAR signalling to promote brain metastasis. Nature. 2019;573(7775):526–31. Cilluffo D, Barra V, Spatafora S, Coronnello C, Contino F, Bivona S, et al. Aneuploid IMR90 cells induced by depletion of pRB, DNMT1 and MAD2 show a common gene expression signature. Genomics. 2020; Hii L-W, Chung FF-L, Soo JS-S, Tan BS, Mai C-W, Leong C-O. Histone deacetylase (HDAC) inhibitors and doxorubicin combinations target both breast cancer stem cells and non-stem breast cancer cells simultaneously. Breast Cancer Res Treat. 2019;1–15. Tang J, Alelyani S, Liu H. Data classification: algorithms and applications. Data Min Knowl Discov Ser CRC Press. 2014;37–64. Wolpert DH. Stacked generalization. Neural networks. 1992;5(2):241–59. Griffith OL, Gray JW. ’Omic approaches to preventing or managing metastatic breast cancer. Breast Cancer Res. 2011;13(6):230. Rohart F, Gautier B, Singh A, Lê Cao K-A. mixOmics: An R package for ‘omics feature selection and multiple data integration. PLoS Comput Biol. 2017;13(11):e1005752. Merrill NM, Lachacz EJ, Vandecan NM, Ulintz PJ, Bao L, Lloyd JP, et al. Molecular determinants of drug response in TNBC cell lines. Breast Cancer Res Treat. 2020;179(2):337–47. Additional files Additional file 1 _ Supplementary Table S1 (.xlsx): The original mutation files for primary breast tumors Additional file 2 _ Supplementary Table S2 and S3 (.xlsx): The original mutation files for metastasis breast tumors Additional file 3 _ Supplementary Methods (.PDF): Selection of positive/negative training sets for BRCA and MBCA Additional file 4 _ Supplementary Table S4-S7 (.xls): List of positive/negative gene set for BRCA and MBCA Supplementary Table S8-S10 (.xlsx): Homo_sapiens genes for BRCA and MBCA Additional file 5 _ Supplementary Results (.PDF): Results for BRCA Table S11 (.xls): Unique driver genes predicted by EARN 100 for BRCA Table S12 (.xls): The list of enriched known genes of EARN 50 in the public databases for BRCA Additional file 6 _ Supplementary Table S13-S26 (.xls): The list of mutated genes extracted by software tools for BRCA and MBCA, and comparison among these genes (p-value≤0.05) for MBCA Additional file 7 _ Supplementary Table S27-S31 (.xls): The list of driver and passenger genes of four learning machines for MBCA Additional file 8 _ Supplementary Table S32-S41 (.xls): The comparison of drivers predicted by all machine learning methods for MBCA Additional file 9 _ Supplementary Table S42 (.xls): The list of driver genes of EARN for MBCA that have been cited in 108 studies of HCMDB Table S43 (.xls): The list of new predicted genes by EARN 50 Additional file 10_ Supplementary Table S44 (.xls): The common/specific enriched pathways for BRCA and MBCA using ReactomeFIVIz (FDR < 0.03) Supplementary Files GraphicalAbstract.png GraphicalAbstract.png GraphicalAbstract.png GraphicalAbstract.png GraphicalAbstract.png Additionalfile1.xlsx Additionalfile1.xlsx Additionalfile1.xlsx Additionalfile1.xlsx Additionalfile1.xlsx Additionalfile10.xls Additionalfile10.xls Additionalfile10.xls Additionalfile10.xls Additionalfile10.xls Additionalfile2.xlsx Additionalfile2.xlsx Additionalfile2.xlsx Additionalfile2.xlsx Additionalfile2.xlsx Additionalfile3SupplementaryMethods.pdf Additionalfile3SupplementaryMethods.pdf Additionalfile3SupplementaryMethods.pdf Additionalfile3SupplementaryMethods.pdf Additionalfile3SupplementaryMethods.pdf Additionalfile4.xls Additionalfile4.xls Additionalfile4.xls Additionalfile4.xls Additionalfile4.xls Additionalfile5SupplementaryResults.pdf Additionalfile5SupplementaryResults.pdf Additionalfile5SupplementaryResults.pdf Additionalfile5SupplementaryResults.pdf Additionalfile5SupplementaryResults.pdf Additionalfile6.xls Additionalfile6.xls Additionalfile6.xls Additionalfile6.xls Additionalfile6.xls Additionalfile7.xls Additionalfile7.xls Additionalfile7.xls Additionalfile7.xls Additionalfile7.xls Additionalfile8.xls Additionalfile8.xls Additionalfile8.xls Additionalfile8.xls Additionalfile8.xls Additionalfile9.xls Additionalfile9.xls Additionalfile9.xls Additionalfile9.xls Additionalfile9.xls Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About In Review Editorial Policies Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-113748","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research article","associatedPublications":[],"authors":[{"id":5145626,"identity":"71a3cf5a-f29e-4c4a-b10e-c86ee4c3e877","order_by":0,"name":"Leila Mirsadeghi","email":"","orcid":"","institution":"Department of Biology, Faculty of Science, Payame Noor University, Tehran, Iran","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Leila","middleName":"","lastName":"Mirsadeghi","suffix":""},{"id":5145627,"identity":"3dd3f523-0e74-43f5-b5e9-e084962f6bee","order_by":1,"name":"Reza Haji Hosseini","email":"","orcid":"","institution":"Department of Biology, Faculty of Science, Payame Noor University, Tehran, Iran","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Reza","middleName":"Haji","lastName":"Hosseini","suffix":""},{"id":5145628,"identity":"36631ba0-5f8f-4bdc-9bad-9dae2f1946fb","order_by":2,"name":"Ali Mohammad Banaei-Moghaddam","email":"","orcid":"","institution":"Laboratory of Genomics and Epigenomics, Department of Biochemistry, Institute of Biochemistry and Biophysics, University of Tehran, Tehran, Iran","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ali","middleName":"Mohammad","lastName":"Banaei-Moghaddam","suffix":""},{"id":5145629,"identity":"e6c605cc-68e7-4c0e-a502-e7bb1ed600cf","order_by":3,"name":"Kaveh Kavousi","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAwklEQVRIiWNgGAWjYBACCR4IzcMPFeAhXotkAzOJWhgMDjAT6TDJnsPPJD621ckY38g/wPCjhkHGvIGAFmneNjPJmW2HecxuJDMw9hxj4JE5QECLHD+DsTFv2wGwFgbeBgYeCUIOk+Nn/wzUUsdjPANoy19itEjz9hg+5m1j5jGQSGZgJsoWyZ4zhQ9nnDvMI3HmscFhmWMShLVInEnfcOBDWZ09f3viw4dvamzsCWpBAQeARpCkYRSMglEwCkYBDgAA9t0yKO3aljwAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0002-1906-3912","institution":"University of Tehran, Institute of Biochemistry and Biophysics","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Kaveh","middleName":"","lastName":"Kavousi","suffix":""}],"badges":[],"createdAt":"2020-11-22 15:05:54","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-113748/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-113748/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":3917685,"identity":"9ef77103-f13a-4683-b4dd-dd86ec3b0e7d","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":108780,"visible":true,"origin":"","legend":"The proposed fusion system workflow for prediction of driver genes in cancers","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/8d693138a9ceffe0b5a7ee17.png"},{"id":3917647,"identity":"0e8ee5a6-d478-4ebf-aba2-48b23970bf11","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":108780,"visible":true,"origin":"","legend":"The proposed fusion system workflow for prediction of driver genes in cancers","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/69a593e63b897fbd7c0d2c27.png"},{"id":3917665,"identity":"66e1eb0f-66b2-4a4b-841b-704a6ef3ebef","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":108780,"visible":true,"origin":"","legend":"The proposed fusion system workflow for prediction of driver genes in cancers","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/40a0f23babceaa97ce807e67.png"},{"id":3917625,"identity":"823b4bf5-030c-4cdd-be22-e338a0cc62a7","added_by":"auto","created_at":"2020-12-01 15:12:06","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":108780,"visible":true,"origin":"","legend":"The proposed fusion system workflow for prediction of driver genes in cancers","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/19fdf9308a34aa536c76ae71.png"},{"id":3917611,"identity":"9cc5966f-8d49-4d03-8e5b-8c832e6ad83b","added_by":"auto","created_at":"2020-12-01 15:12:05","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":108780,"visible":true,"origin":"","legend":"The proposed fusion system workflow for prediction of driver genes in cancers","description":"","filename":"Fig1.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/cbef6f55faf2fd56654ccd45.png"},{"id":3917686,"identity":"0874f50b-bcfc-49b7-b9e6-d0686ba45a40","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":147102,"visible":true,"origin":"","legend":"The Workflow for software tools and machine learning methods. (a) Feature extraction and Feature vector construction, (b) Feature integration, (c) Decision integration","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/7d21bedf819e3ad536a7c212.png"},{"id":3917648,"identity":"8474bfb0-50e5-4398-a33f-5d254637b224","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":147102,"visible":true,"origin":"","legend":"The Workflow for software tools and machine learning methods. (a) Feature extraction and Feature vector construction, (b) Feature integration, (c) Decision integration","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/ef30bdc251b2bb2eb25f7923.png"},{"id":3917666,"identity":"9913ed30-6a42-4b29-8f19-6a5234d25de3","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":147102,"visible":true,"origin":"","legend":"The Workflow for software tools and machine learning methods. (a) Feature extraction and Feature vector construction, (b) Feature integration, (c) Decision integration","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/773fce8f49f160987f4b172a.png"},{"id":3917627,"identity":"c111fb8c-86b1-42d3-8642-ef8aa19fbaae","added_by":"auto","created_at":"2020-12-01 15:12:07","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":147102,"visible":true,"origin":"","legend":"The Workflow for software tools and machine learning methods. (a) Feature extraction and Feature vector construction, (b) Feature integration, (c) Decision integration","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/1c0eb48b7034c0d885ed7651.png"},{"id":3917612,"identity":"322ec201-2874-43f9-835f-825a9a83775c","added_by":"auto","created_at":"2020-12-01 15:12:06","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":147102,"visible":true,"origin":"","legend":"The Workflow for software tools and machine learning methods. (a) Feature extraction and Feature vector construction, (b) Feature integration, (c) Decision integration","description":"","filename":"Fig2.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/ab4d1b5c1c461b3b313de76f.png"},{"id":3917688,"identity":"952de298-f069-459e-9a10-644b47298530","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":252809,"visible":true,"origin":"","legend":"The workflows for the selection of training data and unseen data for BRCA and MBCA. (a) Positive and negative training genes, (b) Genome-wide screening","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/e1b2ff3da58648f03dca9a71.png"},{"id":3917650,"identity":"d80c20c0-d1d3-42ca-a554-2005b002d901","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":252809,"visible":true,"origin":"","legend":"The workflows for the selection of training data and unseen data for BRCA and MBCA. (a) Positive and negative training genes, (b) Genome-wide screening","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/6a0014e3505a2ae8bb95b8ec.png"},{"id":3917668,"identity":"51c1b3f1-90b3-41eb-9240-80435e3b0927","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":252809,"visible":true,"origin":"","legend":"The workflows for the selection of training data and unseen data for BRCA and MBCA. (a) Positive and negative training genes, (b) Genome-wide screening","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/f164a22d028efe4e318fd8fc.png"},{"id":3917631,"identity":"87c4847c-2160-42b8-9c20-77cef80027b7","added_by":"auto","created_at":"2020-12-01 15:12:07","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":252809,"visible":true,"origin":"","legend":"The workflows for the selection of training data and unseen data for BRCA and MBCA. (a) Positive and negative training genes, (b) Genome-wide screening","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/f32277e283da717a44249dec.png"},{"id":3917614,"identity":"65f0a7fe-44ff-464c-a00a-3a9694ad94fd","added_by":"auto","created_at":"2020-12-01 15:12:06","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":252809,"visible":true,"origin":"","legend":"The workflows for the selection of training data and unseen data for BRCA and MBCA. (a) Positive and negative training genes, (b) Genome-wide screening","description":"","filename":"Fig3.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/5d0dc4bd16358fe0fb98fa33.png"},{"id":3917690,"identity":"0812e008-e206-4948-80a5-f734aaef5225","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":229964,"visible":true,"origin":"","legend":"Outputs for BRCA and MBCA. (a) The existence of diversity among features extracted from software tools after setting p-value≤0.05, (b) Frequency of predicted driver and passenger genes using four learning methods, (c) The comparison of driver genes predicted by four methods, (d) The comparison of F1 scores as an evaluation metric of methods","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/7665de4c05b6dcfa97901bff.png"},{"id":3917670,"identity":"b070ddd8-bdd5-4141-9038-9631f02764af","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":229964,"visible":true,"origin":"","legend":"Outputs for BRCA and MBCA. (a) The existence of diversity among features extracted from software tools after setting p-value≤0.05, (b) Frequency of predicted driver and passenger genes using four learning methods, (c) The comparison of driver genes predicted by four methods, (d) The comparison of F1 scores as an evaluation metric of methods","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/a38fae765b6ae844ef4c03dd.png"},{"id":3917652,"identity":"faed2bbd-bccb-4836-8646-c1ce6550655d","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":229964,"visible":true,"origin":"","legend":"Outputs for BRCA and MBCA. (a) The existence of diversity among features extracted from software tools after setting p-value≤0.05, (b) Frequency of predicted driver and passenger genes using four learning methods, (c) The comparison of driver genes predicted by four methods, (d) The comparison of F1 scores as an evaluation metric of methods","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/8158a0df8053d1d467cf1296.png"},{"id":3917634,"identity":"235de1e4-da30-4851-b3ad-7c957d4171ef","added_by":"auto","created_at":"2020-12-01 15:12:08","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":229964,"visible":true,"origin":"","legend":"Outputs for BRCA and MBCA. (a) The existence of diversity among features extracted from software tools after setting p-value≤0.05, (b) Frequency of predicted driver and passenger genes using four learning methods, (c) The comparison of driver genes predicted by four methods, (d) The comparison of F1 scores as an evaluation metric of methods","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/caf22267ad936ba15bb41242.png"},{"id":3917616,"identity":"e98f018a-83ee-40f6-b9a2-3f95c9547830","added_by":"auto","created_at":"2020-12-01 15:12:07","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":229964,"visible":true,"origin":"","legend":"Outputs for BRCA and MBCA. (a) The existence of diversity among features extracted from software tools after setting p-value≤0.05, (b) Frequency of predicted driver and passenger genes using four learning methods, (c) The comparison of driver genes predicted by four methods, (d) The comparison of F1 scores as an evaluation metric of methods","description":"","filename":"Fig4.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/e8f32372cd668d12452a4255.png"},{"id":3917692,"identity":"599c3422-ff56-4e58-abd9-a8b31ada282d","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":123293,"visible":true,"origin":"","legend":"PEA for BRCA and MBCA in top 100 of EARN. (a) The common enriched pathways and the comparison of frequency of top 100 genes predicted by EARN in these pathways. The pathways (1-14) are listed in the guideline box, (b) the common/specific enriched main pathways","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/12d0fa859573095930fa3d84.png"},{"id":3917654,"identity":"5b70ac83-9850-46ac-b6ad-38af871cf47c","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":123293,"visible":true,"origin":"","legend":"PEA for BRCA and MBCA in top 100 of EARN. (a) The common enriched pathways and the comparison of frequency of top 100 genes predicted by EARN in these pathways. The pathways (1-14) are listed in the guideline box, (b) the common/specific enriched main pathways","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/8cf161b5247a621dc2518f2c.png"},{"id":3917672,"identity":"8351c5bc-26f4-480c-b5c7-eb10c71a03a4","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":123293,"visible":true,"origin":"","legend":"PEA for BRCA and MBCA in top 100 of EARN. (a) The common enriched pathways and the comparison of frequency of top 100 genes predicted by EARN in these pathways. The pathways (1-14) are listed in the guideline box, (b) the common/specific enriched main pathways","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/9c8873767eda024c12ed23a5.png"},{"id":3917618,"identity":"47dd13a2-b9da-40ad-95bd-01ca8439a700","added_by":"auto","created_at":"2020-12-01 15:12:09","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":123293,"visible":true,"origin":"","legend":"PEA for BRCA and MBCA in top 100 of EARN. (a) The common enriched pathways and the comparison of frequency of top 100 genes predicted by EARN in these pathways. The pathways (1-14) are listed in the guideline box, (b) the common/specific enriched main pathways","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/5474f114ef153c237acb9e62.png"},{"id":3917636,"identity":"7ae591c2-e3c9-4866-b37e-eb24655dfbf7","added_by":"auto","created_at":"2020-12-01 15:12:08","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":123293,"visible":true,"origin":"","legend":"PEA for BRCA and MBCA in top 100 of EARN. (a) The common enriched pathways and the comparison of frequency of top 100 genes predicted by EARN in these pathways. The pathways (1-14) are listed in the guideline box, (b) the common/specific enriched main pathways","description":"","filename":"Fig5.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/04ee279fbae9197958e4d307.png"},{"id":3917694,"identity":"e4bf2af1-7e24-4208-b182-604691616b0a","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":172678,"visible":true,"origin":"","legend":"Analysis plot of genomic alterations in 16 proposed genes for MBCA using cBioPortal","description":"","filename":"Fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/ad7f5d0838b175f48af1f8e9.png"},{"id":3917674,"identity":"3b11e796-9ee6-4f8a-a8ae-698d6a3fe062","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":172678,"visible":true,"origin":"","legend":"Analysis plot of genomic alterations in 16 proposed genes for MBCA using cBioPortal","description":"","filename":"Fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/926efd6a5d9c30cc0e1ce048.png"},{"id":3917656,"identity":"ea233c8d-002e-4fd0-ba55-3854adfad7db","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":172678,"visible":true,"origin":"","legend":"Analysis plot of genomic alterations in 16 proposed genes for MBCA using cBioPortal","description":"","filename":"Fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/2b2d30b05aba35f197792c54.png"},{"id":3917620,"identity":"95a767a9-64e6-4cfd-a5b3-b56e12277c5d","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":172678,"visible":true,"origin":"","legend":"Analysis plot of genomic alterations in 16 proposed genes for MBCA using cBioPortal","description":"","filename":"Fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/7b10215f41d61e636ef37655.png"},{"id":3917638,"identity":"3a08686c-d195-45f9-a5da-73affa4fd9ca","added_by":"auto","created_at":"2020-12-01 15:12:09","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":172678,"visible":true,"origin":"","legend":"Analysis plot of genomic alterations in 16 proposed genes for MBCA using cBioPortal","description":"","filename":"Fig6.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/3edd3f00cce6e3aeabcd5a36.png"},{"id":3917696,"identity":"b940607b-8ebf-4690-b66e-f92a7c9723ac","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":100403,"visible":true,"origin":"","legend":"Mutations mapping on a linear protein and its domains using MutationMapper in cBioPortal. SMF (%) and the type of these somatic mutations for four genes are specified in the guideline box","description":"","filename":"Fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/00f3fbf3b46ed5295358ea36.png"},{"id":3917658,"identity":"4a6b7338-36a2-49ca-ac99-a39441403340","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":100403,"visible":true,"origin":"","legend":"Mutations mapping on a linear protein and its domains using MutationMapper in cBioPortal. SMF (%) and the type of these somatic mutations for four genes are specified in the guideline box","description":"","filename":"Fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/09137b1d5aaca25cb578d762.png"},{"id":3917676,"identity":"6b14a820-c402-41a2-bfe4-cf1ab65bf87d","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":100403,"visible":true,"origin":"","legend":"Mutations mapping on a linear protein and its domains using MutationMapper in cBioPortal. SMF (%) and the type of these somatic mutations for four genes are specified in the guideline box","description":"","filename":"Fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/0b375bdf1483dcca982823e1.png"},{"id":3917622,"identity":"11546072-cc32-4d08-8069-7582e7c22311","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":100403,"visible":true,"origin":"","legend":"Mutations mapping on a linear protein and its domains using MutationMapper in cBioPortal. SMF (%) and the type of these somatic mutations for four genes are specified in the guideline box","description":"","filename":"Fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/f31c27470029176b959834d2.png"},{"id":3917640,"identity":"43cb6aa4-9bcd-474c-ae8a-0c4e7be5e01f","added_by":"auto","created_at":"2020-12-01 15:12:09","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":100403,"visible":true,"origin":"","legend":"Mutations mapping on a linear protein and its domains using MutationMapper in cBioPortal. SMF (%) and the type of these somatic mutations for four genes are specified in the guideline box","description":"","filename":"Fig7.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/35821ee326a591dd17ab649d.png"},{"id":13620611,"identity":"76a0dffc-7abb-4e62-8246-69d5529f753e","added_by":"auto","created_at":"2021-09-17 07:06:33","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":7041442,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/a5256a27-1c71-4b4b-9967-ca83f83ae463.pdf"},{"id":3917684,"identity":"987fad68-36b4-4cf6-8c7f-251ee710ae63","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":334772,"visible":true,"origin":"","legend":"","description":"","filename":"GraphicalAbstract.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/19d22bbe46e717fe09374521.png"},{"id":3917646,"identity":"001e8dd2-1f12-4038-8797-670989f0be67","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":334772,"visible":true,"origin":"","legend":"","description":"","filename":"GraphicalAbstract.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/6202288920dbdaf421a191ee.png"},{"id":3917664,"identity":"5aaf7852-0674-460e-a0b0-76243f17cc59","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":334772,"visible":true,"origin":"","legend":"","description":"","filename":"GraphicalAbstract.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/5af2587e12b61e4346a95260.png"},{"id":3917623,"identity":"cf8a0f86-d642-44dc-a769-3bba3875d78a","added_by":"auto","created_at":"2020-12-01 15:12:06","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":334772,"visible":true,"origin":"","legend":"","description":"","filename":"GraphicalAbstract.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/4f5a6476fb3735054398e7d5.png"},{"id":3917610,"identity":"75c5c000-e3a8-4f88-a4c1-6fc15a16ed6c","added_by":"auto","created_at":"2020-12-01 15:12:05","extension":"png","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":334772,"visible":true,"origin":"","legend":"","description":"","filename":"GraphicalAbstract.png","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/26e16abe98e778cbd5cb13ca.png"},{"id":3917687,"identity":"10c55c78-6660-4a95-b727-27fab956858b","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":20873774,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/22b43a84f1c235aab067ac69.xlsx"},{"id":3917649,"identity":"8aeb803f-4a6d-466b-9de2-64830788ad88","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":20873774,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/c8a9db8937a67c0769192372.xlsx"},{"id":3917667,"identity":"66246c27-a7cb-45c5-bc71-f3f97de96e5b","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":20873774,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/a863b5c748b5089d7b0c09bf.xlsx"},{"id":3917628,"identity":"27b3cb08-1a7a-4f73-b4ff-10b3afd7fdde","added_by":"auto","created_at":"2020-12-01 15:12:07","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":20873774,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/b9bd2b8c48846aebb2453f4b.xlsx"},{"id":3917613,"identity":"5b6c637d-d77a-4fae-aea5-6540fd0a7f64","added_by":"auto","created_at":"2020-12-01 15:12:06","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":20873774,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/c293b76af4ec64bb7d60640f.xlsx"},{"id":3917689,"identity":"75057299-0fc8-43a9-955e-50ff023bc664","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xls","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":48128,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile10.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/6bdfd865b4c43d0b60869e5f.xls"},{"id":3917651,"identity":"b04a2aa9-65db-433c-9472-c78e08a0823d","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"xls","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":48128,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile10.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/bbefb2ff2794701a463a3d51.xls"},{"id":3917669,"identity":"07f4823c-3334-4c2f-83be-eb310d0b0f98","added_by":"auto","created_at":"2020-12-01 15:12:12","extension":"xls","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":48128,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile10.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/e071d05c00485dda5b8efa7c.xls"},{"id":3917633,"identity":"9400ace6-0bbb-47ea-b10d-9d96143b2413","added_by":"auto","created_at":"2020-12-01 15:12:08","extension":"xls","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":48128,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile10.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/88cbcbff26b8de80feff9560.xls"},{"id":3917615,"identity":"f8ef0e01-c848-42a2-a66c-749aa92979ab","added_by":"auto","created_at":"2020-12-01 15:12:07","extension":"xls","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":48128,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile10.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/1f9c393f8055447781a62959.xls"},{"id":3917691,"identity":"c9829ac7-b8c8-40a8-8557-bfeb2f4fe7aa","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":12215188,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/32a8815730dfee44c6f0f482.xlsx"},{"id":3917653,"identity":"6c142766-27dc-4806-b03a-5771ea8853e2","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":12215188,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/a2e1579524b0a27192697aa7.xlsx"},{"id":3917671,"identity":"8325d9ce-a846-45ac-8df4-f327851e1c28","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":12215188,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/77aa80617f8c039f7a6fb675.xlsx"},{"id":3917617,"identity":"7179dc86-974a-430f-bf77-a42d1d174c8f","added_by":"auto","created_at":"2020-12-01 15:12:09","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":12215188,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/97e68ca9239707f88409ab8e.xlsx"},{"id":3917635,"identity":"3bfa0b81-2cf4-4fc2-ae18-8f1388ca7596","added_by":"auto","created_at":"2020-12-01 15:12:08","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":12215188,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/d64653f5fdebade1fa446469.xlsx"},{"id":3917693,"identity":"751469b8-84e1-4d8e-b475-e4077f1c1280","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":164125,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile3SupplementaryMethods.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/8cf98e38a89bd80e6ef7b8a1.pdf"},{"id":3917655,"identity":"d2879278-b11f-4fb4-9b77-5e51bc420d8f","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":164125,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile3SupplementaryMethods.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/1e39b9ee5c2ad6a42041d97e.pdf"},{"id":3917673,"identity":"8358f026-579d-49e8-9e41-ae7e2fca6c85","added_by":"auto","created_at":"2020-12-01 15:12:13","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":164125,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile3SupplementaryMethods.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/e6f08e5d7b7f1a838210ceed.pdf"},{"id":3917619,"identity":"78d9a28f-e7cf-4af9-8c5f-05f795c6805b","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":164125,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile3SupplementaryMethods.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/511a4ec19f869d80cc7ab542.pdf"},{"id":3917637,"identity":"81454715-a27a-4646-b51a-aae6cad91f84","added_by":"auto","created_at":"2020-12-01 15:12:09","extension":"pdf","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":164125,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile3SupplementaryMethods.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/fa6f0ffb1fd966cf30bc915a.pdf"},{"id":3917695,"identity":"b74a273c-0df3-4a42-8dbf-892b5460ecbe","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":2435584,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile4.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/60fa96620e747d51f50d1b2f.xls"},{"id":3917657,"identity":"c3097829-acb7-46f4-a5e4-6beeac890664","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":2435584,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile4.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/8965868f242240aed5d7ee8d.xls"},{"id":3917675,"identity":"3a08d358-4585-41b7-a4d1-70d06875a112","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":2435584,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile4.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/4128626ee3d6f88ab5f49aa7.xls"},{"id":3917621,"identity":"372bff11-b20b-4e06-92c4-9eb473fb4658","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":2435584,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile4.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/0a4c7c2127bc3d91d4157ecd.xls"},{"id":3917639,"identity":"823201aa-95de-43cf-9a94-ce7f4e57632f","added_by":"auto","created_at":"2020-12-01 15:12:09","extension":"xls","order_by":6,"title":"","display":"","copyAsset":false,"role":"supplement","size":2435584,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile4.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/d5d8b0998d2b306681563179.xls"},{"id":3917697,"identity":"42b4c883-da76-4a05-9e02-99a86673263c","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":223037,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile5SupplementaryResults.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/2709e3ded173698f9794a97d.pdf"},{"id":3917659,"identity":"175c27bd-ca60-4617-8ef0-a6d5255e4b5f","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":223037,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile5SupplementaryResults.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/0ec20de2586a5d8eb6073bf3.pdf"},{"id":3917677,"identity":"0e71aaae-e6ac-46c8-a213-89e25d5122d4","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":223037,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile5SupplementaryResults.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/743234f5683e106bf604ce60.pdf"},{"id":3917624,"identity":"9d902e5a-cf09-443a-b3a1-d43422e99640","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":223037,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile5SupplementaryResults.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/7ae4db782bb6e7d31492ed45.pdf"},{"id":3917641,"identity":"4fb6cefa-2afb-44dd-b484-a6e3f12d0e5d","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"pdf","order_by":7,"title":"","display":"","copyAsset":false,"role":"supplement","size":223037,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile5SupplementaryResults.pdf","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/33b8e5e31dd9f74fd8cbba65.pdf"},{"id":3917698,"identity":"176657f3-9702-4332-b6fa-a3c2e4e9bd22","added_by":"auto","created_at":"2020-12-01 15:12:16","extension":"xls","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":3653632,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile6.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/aead2cc79b910fcfd3039564.xls"},{"id":3917660,"identity":"9f50eeca-841e-4b4e-938a-1c3bb4481e33","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xls","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":3653632,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile6.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/48ea9c7af2b46bf1f5812437.xls"},{"id":3917678,"identity":"4eab065e-699d-413d-bdf6-9b4b83c3c20e","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xls","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":3653632,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile6.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/a399389c434df4a084b82904.xls"},{"id":3917626,"identity":"cfbb6dbe-906d-4898-98ea-b169d4bf4764","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"xls","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":3653632,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile6.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/7c4f6fcc62152ce5cd4c1f50.xls"},{"id":3917642,"identity":"9746ea9e-8f27-4ea8-bab6-0df027ba22bc","added_by":"auto","created_at":"2020-12-01 15:12:10","extension":"xls","order_by":8,"title":"","display":"","copyAsset":false,"role":"supplement","size":3653632,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile6.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/df819f4840a60db70a93d37a.xls"},{"id":3917699,"identity":"8b137ebd-54c7-459b-8fb9-25f997a5d7f9","added_by":"auto","created_at":"2020-12-01 15:12:16","extension":"xls","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":6815744,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile7.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/8213dd458358dc97509ba6f8.xls"},{"id":3917679,"identity":"205c171f-652e-4b80-a319-547dae897e27","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"xls","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":6815744,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile7.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/eb0db662cee76d3d381f8685.xls"},{"id":3917661,"identity":"02529dac-96fa-4b36-a3b0-c921e252d976","added_by":"auto","created_at":"2020-12-01 15:12:14","extension":"xls","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":6815744,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile7.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/ec20959bb07ce637471eb0ae.xls"},{"id":3917629,"identity":"da6fe6f3-4d9f-4c67-b0b6-552df598ff1d","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"xls","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":6815744,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile7.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/aea40109bdcf69b20258753c.xls"},{"id":3917643,"identity":"dc27beef-5d08-4779-83d6-6fa3a07d7e83","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"xls","order_by":9,"title":"","display":"","copyAsset":false,"role":"supplement","size":6815744,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile7.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/eb5492c20489c262c476c988.xls"},{"id":3917700,"identity":"a778d422-86ac-4c16-ac07-f6bad863fb58","added_by":"auto","created_at":"2020-12-01 15:12:16","extension":"xls","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":45568,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile8.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/af5de6adb7926417831092cb.xls"},{"id":3917662,"identity":"6255fe69-75ff-4acb-b0cd-3b239c9f9090","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"xls","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":45568,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile8.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/0dde2405e4dbf9050cd90f7b.xls"},{"id":3917680,"identity":"f44201b0-5b12-420d-b517-12b2b79a64a9","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"xls","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":45568,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile8.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/df563386c67becd60f83997a.xls"},{"id":3917630,"identity":"bbb7af17-8612-4e80-8db8-b12f5c9e7eaa","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"xls","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":45568,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile8.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/177a3db2cad013e1ff439578.xls"},{"id":3917644,"identity":"f5a1a82f-23d4-4f94-b0a2-9584e2e00813","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"xls","order_by":10,"title":"","display":"","copyAsset":false,"role":"supplement","size":45568,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile8.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/6b7782003b9451adb728bcdc.xls"},{"id":3917701,"identity":"4bfc65af-91f1-411f-b2be-58bba378cf35","added_by":"auto","created_at":"2020-12-01 15:12:16","extension":"xls","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":51200,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile9.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/4c694daa0857c8118b8d0d5d.xls"},{"id":3917663,"identity":"a1f9d647-b4c8-40b2-8f75-989f3d215063","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"xls","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":51200,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile9.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/c402b646911671c14181749f.xls"},{"id":3917681,"identity":"ca88543e-874a-4b70-9bbf-1548d01fc6d5","added_by":"auto","created_at":"2020-12-01 15:12:15","extension":"xls","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":51200,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile9.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/39e1b69b696c08c0dcda6721.xls"},{"id":3917632,"identity":"fda990db-5675-45fa-9672-02f4459ee9d1","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"xls","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":51200,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile9.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/dc29859246ac6ab0ee16853a.xls"},{"id":3917645,"identity":"2a277584-253c-4880-802a-b39bd2d07b22","added_by":"auto","created_at":"2020-12-01 15:12:11","extension":"xls","order_by":11,"title":"","display":"","copyAsset":false,"role":"supplement","size":51200,"visible":true,"origin":"","legend":"","description":"","filename":"Additionalfile9.xls","url":"https://assets-eu.researchsquare.com/files/rs-113748/v1/990605e20a263510df90ce58.xls"}],"financialInterests":"","formattedTitle":"EARN: an ensemble machine learning algorithm to predict driver genes in metastatic breast cancer","fulltext":[{"header":"1. Background","content":"\u003cp\u003eThe mutations induce small changes to the genes. If they cause damage and remain untreated, it drives multifactorial anomalies which are called complex diseases. Cancers are one kind of these complex diseases which are induced by defective driver genes. Among cancers, primary breast cancer as a complex disease is the most commonly diagnosed carcinoma in women worldwide and will be fatal if it progresses towards the secondary-stage. It is the most common and the second most common cause of cancer death in women in developing regions and developed regions, respectively (1). Over the last 10 years, the incidence of breast cancer has increased almost 10 times (2). The concern about this growing trend has prompted oncologists to seek early detection. Nowadays, the molecular technique of next-generation sequencing (NGS), including whole-exome sequencing could generate a large amount of data related to mutated genes called mutation data (3). The analysis of this massive data requires the use of robust computational approaches to exploit the information effectively. Precision oncology by focusing on targeted clinical panel sequencing can be helpful, i.e., a breast cancer-specific NGS panel, including 79 genes has been validated for use in primary and metastatic breast cancer (4). In this way, the advent of bioinformatics tools in parallel with the development of molecular techniques could lead to discovering biomarkers that are efficient in cancer diagnosis and prognosis (5). The machine learning algorithms as one of the computational approaches can be trained with data from countless patients whereas it is too difficult for human physicians and biologists to gain such experience in an entire career or their researches. These models equip experts to make better decisions (6). Some of them are the ensemble classifier (EC) machine learning methods that combine two or several models to optimize the performance of the base components in order to improve data analysis. In previous studies, it has been mentioned that committee approaches can outperform even powerful individual models in many cases (7). Also, investigations show using ensemble models (a.k.a fusion systems) is widely increasing in many fields of inquiry, including the detection of cancers and their subtypes and especially in the area of breast cancer detection. In 1996, a breast cancer dataset, including 699 samples, were analyzed by bagging nearest neighbor classifiers as a fusion system (8). Since then, many ensemble classification methods have been applied to breast cancer prognosis (9). In this regard, we reviewed 42 ensemble methods related to 18 cancers (10). Among these, 22 approaches have been reported for analyzing breast cancer data in the literature (Table 1).\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eTable 1\u003c/strong\u003e 22 ensemble learning methods concerned with the detection of breast cancer\u003c/p\u003e\n\u003ctable border=\"1\" width=\"546\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003e\u003cstrong\u003eMethod name\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e\u003cstrong\u003ePublication year\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eBayesian networks-based model integration (11, 12)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2006 and 2019\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eRSS-SCS method (13)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2016\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eCollective approach (correlation, color palette, color proportion, and SVM) (14)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2016\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eKernel-based Data Fusion Method for Gene Prioritization (15)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2015\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eDECORATE method* (16)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2015\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eHyDRA method* (17)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2015\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eGenEnsemble method* (NBS-IB3-SVM-C4.5 DT) (18)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2014\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e8\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eNB (Na\u0026iuml;ve Bayes) combiner method (19)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2014\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eEvolutionary Ensemble Model (20)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2014\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e10\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003esmoothed t-statistic SVM (stSVM) (21)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2013\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eSVM Classifiers Fusion (three SVM) (22)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2013\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e12\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eCOMBINER (Core Module Biomarker Identification)* (23)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2012\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eEnsembles of BioHEL Rule Set (24)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2012\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eStacking IB3-NBS-RF-SVM method (25)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2012\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e15\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eREIS-based ensemble method (26)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2011\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e16\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eMRS method (27)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2010\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e17\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eBoosting-TWSVM method (28)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2009\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e18\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eBagging and boosting-based TWSVM (29)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2009\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e19\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eFeature Subsets Method (30)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2008\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eBNCE method (31)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2007\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e21\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eBayesian Network Classifier (32)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2006\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e22\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"421\"\u003e\n\u003cp\u003eenSVM (200 SVM) (33)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"83\"\u003e\n\u003cp\u003e2006\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e* Some proposed methods to discover genomic markers related to breast cancer.\u003c/p\u003e\u003cp\u003eIn some of these studies, ECs have been used for introducing driver genes associated with breast cancer and the evaluation of genomic biomarkers regarding this cancer. In this work, we propose the EC learning approach called EARN (Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine). It is used to find candidate drivers in primary breast invasive carcinoma (BRCA) and metastatic breast cancer (MBCA) samples from mutation data available in the Cancer Genome Atlas (TCGA) (https://portal.gdc.cancer.gov) and cBioPortal (http://cbioportal.org). The candidate genes introduced by the EC mechanism may already be known as cancers causing genes in databases or can be novel. The candidate genes have the potential to be presented as genomic risk biomarkers after completing the steps of clinical trials (34) and used for personalized targeted therapy (35). Furthermore, there is evidence that driver genes that effectively prognose cancers could be used in therapeutic applications to access more effective therapies (36). The proposed EC method combines decisions of three base classifiers, including non-linear Support Vector Machine (NLSVM) (37), Artificial Neural Network (ANN) (38), and Random Forest (RF) (39). The features for these three classifiers were extracted from four software tools: MutSigCV v.1.4 (40), OncodriveCLUST 0.4.1 (41), OncodriveFM (42), and NetBox 1.0 (43). Overall, we aim to focus on the findings in five steps of BRCA and MBCA prognosis and diagnosis. 1. A list of mutated genes ranked by four software tools based on \u003cem\u003ep\u003c/em\u003e-value is presented as the features. 2. Driver and passenger genes predicted by three individual machine learning methods and EARN are introduced and compared. 3. Biological validation of predictions based on gene set enrichment analysis is done and discussed. Indeed, we evaluate the top genes predicted by EARN and three base classifiers for BRCA and MBCA by searching these genes in the list of cancer-associated genes in the public databases, including the Online Mendelian Inheritance in Man (OMIM), the Cancer Gene Census (CGC) (44), the Network of Cancer Genes (NCG) (45, 46), and the human cancer metastasis database (HCMDB) (47). 4. The performance of all machine learning methods is evaluated by monitoring some statistical metrics. 5. Finally, a targeted driver gene panel for MBCA diagnosis based on pathway enrichment analysis (PEA) of top 100 predicted by EARN (EARN\u003csub\u003e100\u003c/sub\u003e) is proposed.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cp\u003eIn this study, an ensemble method as a synergistic combination of computational\u0026nbsp;tools has been designed and proposed to find the putative cancer drivers. This fusion system can help to analyze the Whole-Exome Sequencing (WES) data. It consists of four steps: selection of dataset, feature extraction, feature integration, and decision integration (Fig. 1). We have also shared Python source code and other requirements for the implementation of the proposed ensemble machine learning algorithm as the protocol via GitHub (https://github.com/lmirsadeghi/EARN/).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.1 Selection of dataset\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this study, to identify candidate driver genes based on mutations that occur in genes, breast cancer primary and metastasis data have been analyzed.\u003c/p\u003e\n\u003cp\u003eFor primary breast cancer, an open-access mutation annotation format (.maf) file was downloaded from TCGA data set regarding BRCA (48). This file includes 90969 masked somatic mutations identified in 17990 genes from 983 tumor samples of BRCA patients that their whole exome had been sequenced by Illumina Genome Analyzer II [see this mutation file (.xlsx) in Additional file 1: Table S1]. Also, for processing of sequences, a bioinformatics\u0026nbsp;pipeline\u0026nbsp;framework called \"MuSE Variant Aggregation and Masking\" in TCGA has been used. For MBCA, two files (.txt) were downloaded from the cBio Cancer Genomics Portal (49, 50). The first mutation file includes WES of 213 tumor samples from 213 MBCA patients by Illumina HiSeq. It is associated with 22949 somatic mutation counts that occurred among 10791 genes (51). The second file consists of WES of 237 metastasis tumor samples by Illumina GAIIx from 180 patients regarding 24027 somatic mutations identified in 10273 genes (52). Clinical data shows that 86 samples were taken when patients were in the metastatic disease stage, and other samples had been taken less than 4 months prior to the metastatic disease is diagnosed (53). After selecting two initial datasets concerning MBCA [see these mutation files (.xlsx) in Additional file 2: Table S2 and S3], they augmented to build a comprehensive mutation data file, including 46928 somatic mutations identified among 14293 genes from 450 MBCA tumor samples (393 patients).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.2 Selection of software tools for feature extraction\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAfter preparing mutation files, four software tools including MutSigCV v.1.4, OncodriveCLUST 0.4.1, OncodriveFM, and NetBox 1.0 were used to extract the convenient numerical features. The selection of tools for feature extraction was a crucial step to achieve better performance on the final algorithm of the proposed ensemble learning model. We select the four software tools based on evidences of a paper in 2015 on identification and ranking of plausible drivers for BRCA and ovarian (OV) cancer (16). It had been demonstrated that among ten tools for extracting features, OncodriveFM and NetBox generate high sensitivity, especially about BRCA. Also, the sensitivity of OncodriveCLUST tool is high concerning the OV cancer. On the other hand, in both cancers, it had been shown that the positive predictive value (PPV) for NetBox and OncodriveFM is high. MutsigCV was able to propose a large number of drivers in the top 50 genes for BRCA, where at least five other methods had also predicted them as top genes. These advantages led us to use these tools. Practically, these four tools evaluate original mutation files from different aspects and assign a score (\u003cem\u003ep-\u003c/em\u003evalue) to genes to show their relevance to disease according to that software\u0026rsquo;s logic. MutsigCV gets data concerning point mutations and small insertions and deletions (INDELs) from the WES file. After analyzing and estimating mutation frequency, it can identify and introduce a significant list of mutated genes for cancers (40). OncodriveClust software tool is able to identify mutations that generate oncogenes and leads to changes in the function of the proteins. For this purpose, it analyzes synonymous mutations and protein-affecting mutations, including non-synonymous, stop, and splice-site mutations (41). Also, this tool uses data from the Cancer Gene Census (CGC) database (44). for selecting known drivers associated with cancers. OncodriveFM is our next tool which can detect driver genes across tumor samples, identify pathways in cancers, and discover gene modules by using information that is available in the WES file. This data is provided by three methods, including SIFT, PolyPhen2, and MutationAssessor (42). The fourth tool is NetBox, and it can detect driver mutations based on a network. First, a global human interaction network is constructed by this tool. Then, it finds the linker genes between mutated genes for module discovery and identification of candidate drivers (43). Indeed. the concept and criteria of selecting these four software tools are based on the study in 2015 where the performance metrics of ten methods for prediction of plausible driver genes of BRCA and ovarian OV were compared (16).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.2.1 Feature extraction and feature vector construction\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this step, the four software tools explained above are used for the extraction of features from primary and metastasis mutation data files. After running the tools, all genes are ranked based on \u003cem\u003ep-\u003c/em\u003evalue as output data, and each method assigns a number (0\u0026le;\u003cem\u003ep-\u003c/em\u003evalue\u0026le;1) to genes as numerical features. Therefore, a four-dimensional feature vector is constructed for each gene (Fig. 2a). Since the genes with lower \u003cem\u003ep-\u003c/em\u003evalue play a more critical role in the development of cancer, we decided to use \"1- \u003cem\u003ep-\u003c/em\u003evalue\" as the final numerical feature for each gene. With this plan, the genes that are more important in the occurrence of BRCA and MBCA will also get higher feature values. Different and independent logics behind the ranking mechanisms in the exploited tools guarantee enough diversity between inputs of the ensemble system which is an essential property for efficient fusion methods.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.3 Classifier model selection \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThree supervised machine learning methods, including non-linear SVM to learn non-linear functions to separate the classes, ANN, and RF are used as individual classifiers. For selecting these methods, the literature and previous studies were surveyed. We did a comprehensive review regarding fusion systems and the results showed that SVM has been used as a base classifier in many studies or applied as a baseline for comparison between the performance of the different machine learning methods (54\u0026ndash;57). Since in this case, the positive and negative training gene set for the implementation of learning algorithms are highly imbalanced, 40 positive genes versus 2151 negative genes (refer to 2.4), a solution must be found. It has been demonstrated the SVM classifier can be a robust method for generating optimal results with imbalanced positive and negative datasets (58), especially when an Instance-weighted SVM algorithm is used (59). So, we weighed this algorithm to get better results. On the other hand, RF is an ensemble machine learning method used as one of the individual classifiers. This method can partially solve the problem of the unbalanced positive and negative training set by bootstrap sampling and can also improve performance, i.e., predictive accuracy reached 88.89% using RF for breast cancer risk prediction (60). The ANN classifier is another machine learning method with long-lasting profound literature. In 1990, Hansen and Salamon integrated multiple neural networks and improved results (61). This method is also widely used in biology studies and has achieved high performance. In 2017, it was shown that ANN could be used for the diagnosis of lung cancer (62). Meanwhile, in this study, the positive training set is small, and recent researches have revealed that ANN may improve performance for problems with small training set sizes and give better performance, especially for problems with time-series data category (63). All of these reasons and criteria led to the selection of these three machine learning methods as base classifiers of the final ensemble system.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.4 Training and testing \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this study, we train separate models for BRCA and MBCA. Some criteria for the selection of training data sets are described below and visualized in Fig. 3a. For testing the performance of models in terms of evaluation metrics (e.g. recall, precision, etc.), we average over 100 trials. In each trial, 3-fold cross-validation with random shuffles is used to calculate the metrics on all data. Finally, the mean and standard deviation of metrics over 100 trials are obtained. Average outputs for cross-validation of the estimator of each model on testing data based on some metrics, including precision, f1 score, recall, accuracy, and Receiver Operating Characteristic-Area under Curve (ROC-AUC) are presented in section 3.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.4.1 Training data set selection\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe positive training set of genes for BRCA and MBCA were obtained from searching known genes and mentioned drivers concerning these cancers in several databases, including the OMIM, CGC, NCG, HCMDB, and the Human Protein Atlas (HPA) (https://www.proteinatlas.org/). Also, about selecting negative training gene set, we reviewed a comprehensive list of prior works. Since there is no gold and standard database for a negative set selection, most researchers have used the bootstrap method for resampling, and the negative training genes have been mostly selected randomly. In this study, negative data was selected by counting the occurrence of mutations across all samples in the initial mutation data file, and the genes with the lowest mutation count were used as negative training set (16). It is crucial to note that in both positive and negative training data, we only accepted protein-coding genes. [see further details for training genes in Additional file 3: Supplementary Methods and Additional file 4: Table S4-S7].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.5 Genome-wide screening \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor the genome-wide screening, 20208 homo sapiens genes annotated as protein-coding were downloaded from ftp://ftp.ncbi.nlm.nih.gov/gene/DATA/GENE_INFO/Mammalia/ on February 2019. The proposed ensemble model is applied to 18017 genes for BRCA and 16698 genes for MBCA, after excluding positive and negative training sets (Fig. 3b) [see Additional file 4: Table S8-S10].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.6 Implementation of three machine learning algorithms based on feature integration \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAfter adding features to the system, and training and testing of learning methods including non-linear SVM, ANN, and RF, they are applied to the protein-coding genes as the unseen data. Each of these methods integrates the features extracted from the initial mutation file (refer to section 2.2.1). We use scikit-learn package to implement our algorithms in python (64). Since this problem is a binary classification of genes based on drivers and passengers, they could label genes based on two indexes -1 and +1 (-1 means passenger genes and +1 means drivers), and also compute a score for each gene, independently (Fig. 2b).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.7 Implementation of proposed ensemble machine based on decision integration\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFinally, the decision-making strategy for ensemble machine is based on aggregation of the predicted scores obtained from other machines. We call the proposed EC machine learning method EARN (ensemble of ANN, RF, and non-linear SVM). EARN uses the average of the scores of the outputs of the three basic classifiers to assign a new score (ranging from 0 to 1) to each gene. The genes with higher prediction scores (scores \u0026ge;0.5) are labeled as drivers (+1) while the other genes will be passengers (-1). This process has been illustrated in Fig. 2c.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2.8 Biological inferences\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAt this step, all the driver genes introduced for BRCA and MBCA, as well as top genes predicted by learning machines, are searched in the public databases to determine which genes have been already known related to cancer and which ones are new. Pathway enrichment analysis is also performed using ReactomeFIVIz tool (\u003cem\u003eFDR\u003c/em\u003e \u0026lt;0.03) (65\u0026ndash;67)to identify the biochemical pathways associated with the candidate genes and examine the biological role of them. It is applied to find biological pathways and patterns related to cancer and other complex diseases.\u003c/p\u003e"},{"header":"3. Results","content":"\u003cp\u003eThis study is an attempt to focus on the findings in five steps of BRCA and MBCA prognosis and diagnosis. 1. A list of mutated genes ranked by four software tools based on \u003cem\u003ep-\u003c/em\u003evalue is presented as the features. 2. Driver genes and passengers predicted by three individual machine learning methods, NLSVM, ANN, RF, and the proposed EC are introduced. 3. Biological validation of predictions based on gene set enrichment analysis is done 4. Statistical validation of all learning methods based on evaluation metrics is carried out. 5. A targeted gene panel for MBCA based on pathway enrichment analysis (PEA) is proposed.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.1 BRCA\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe description of the results for BRCA is presented in Additional file 5: Supplementary Results and Table S11 and S12. However, the comparative results of each algorithm for BRCA and MBCA are illustrated in the next section.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2 MBCA\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2.1\u003c/strong\u003e \u003cstrong\u003eInvestigation of the diversity of features extracted from the original mutation file\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFour software tools are used to extract and rank the list of mutated genes for MBCA as features based on \u003cem\u003ep-\u003c/em\u003evalue to be used for the machine learning implementation in the next step. The use of multiple tools for generating features creates an effective diverse committee for better classification. It is known that machine learning method can do better discrimination with higher-dimensional feature vectors and perform the classification with higher accuracy (26). To illustrate the existence of diversity in features and also for comparison between results of the tools, we plot the GeneVenn diagram (68) by setting \u003cem\u003ep-\u003c/em\u003evalue\u0026le;0.05 as the threshold. The plotting Venn diagram (\u003cem\u003ep-\u003c/em\u003evalue\u0026le;0.05) shows that the results of four software tools in the ranking of mutated genes for BRCA and MBCA are varied (fig. 4a). It means that the extracted features by these tools from the original mutation file are sufficiently diverse and can be applied for machine learning implementation step. The comparison shows that five genes, C12orf29, OXCT1, PIK3CA, GCNT4, and C8orf44, are just common among the outputs. Also, PIK3CA has been selected by all software tools in both cases of BRCA and MBCA [see the outputs of software tools for BRCA and MBCA, and comparison among mutated genes (\u003cem\u003ep-\u003c/em\u003evalue\u0026le;0.05) extracted by these tools for MBCA in Additional file 6: Table S13-S26].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2.2 Outputs of three individual classifiers and EARN\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe three base classifiers and EARN predicted the labels and scores of 16698 protein-coding genes for MBCA. The percentage of the predicted driver and passenger genes using the four learning methods for BRCA and MBCA has been shown in Fig. 4b. These findings have been presented in an extra file [see Additional file 7: Table S27-S31].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2.3 Investigation of top 100 genes predicted by the four machine learning methods\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe comparison of the top 100 genes predicted by the four methods using GeneVenn diagram tool shows that 16 genes are predicted by all four machines for MBCA (Fig. 4c). The results of the enrichment of these genes in public databases are considered in Table 2. Other common and unique driver genes predicted by methods are presented in the extra file [see Additional file 8: Table S32-S41]. Also, among the outputs of EARN\u003csub\u003e100\u003c/sub\u003e, BDNF, PRKCG, TH, PRKCD, and PIP5K1B are just predicted by this learning machine in the list of top 100 genes. Among these five genes, BDNF and PRKCG have been already introduced regarding metastatic cancers but the others are new.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 2\u003c/strong\u003e The 16 common genes predicted by all machines in the top 100. The confirmed genes as the known genes related to different primary cancers or primary breast tumors in OMIM, CGC, and NCG databases have been marked in the last two columns\u003c/p\u003e\n\u003ctable border=\"1\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003eSymbol\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003eNSCGMCH (#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"78\"\u003e\n\u003cp\u003e\u003cstrong\u003eNSCGMBH\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"71\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGECC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGEBC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eOXCT1*\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eKDR**\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eAPEX1*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eGCM2*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eUNC13D*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eNCOR1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eKRAS\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eTHAP3*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eSERPINE2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eBATF*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eC8orf44*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eC12orf29*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u0026nbsp;ZNF546*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eKDM6B\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eGCNT4*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003eFOXA1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"94\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"55\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"23\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"71\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"70\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e*Ten new genes that have not already been introduced in the databases.\u003c/p\u003e\n\u003cp\u003e** KDR is confirmed in HCMDB related to metastatic breast cancer in two studies.\u003c/p\u003e\n\u003cp\u003eNSCGMCH: \u003cstrong\u003eN\u003c/strong\u003eumber of \u003cstrong\u003es\u003c/strong\u003etudies that have \u003cstrong\u003ec\u003c/strong\u003eited \u003cstrong\u003eg\u003c/strong\u003eenes related to different \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003ec\u003c/strong\u003eancers in the \u003cstrong\u003eH\u003c/strong\u003eCMDB\u003c/p\u003e\n\u003cp\u003eNSCGMBH: \u003cstrong\u003eN\u003c/strong\u003eumber of \u003cstrong\u003es\u003c/strong\u003etudies that have \u003cstrong\u003ec\u003c/strong\u003eited these \u003cstrong\u003eg\u003c/strong\u003eenes related to \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003eb\u003c/strong\u003ereast cancer in the \u003cstrong\u003eH\u003c/strong\u003eCMDB\u003c/p\u003e\n\u003cp\u003ePKGECC: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with different \u003cstrong\u003ec\u003c/strong\u003eancers that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003ePKGEBC: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with \u003cstrong\u003eB\u003c/strong\u003ereast cancer that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2.4 Biological validation of predictions based on gene set enrichment analysis\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe biological analysis of genes predicted by EARN is performed based on two plans; (a) analysis of the results based on all predicted driver genes (labeled as +1) and (b) analysis of the findings based on the top-scoring genes. To investigate outputs of the EARN for MBCA from a biological point of view based on the label, we analyzed the results concerning the public databases. There is a gene-metastasis association data file (.xls) in the HCMDB that lists 2240 genes related to metastatic cancers based on experiments performed in various studies. 622 genes out of these genes were introduced for metastatic breast cancer specifically. It should be noted that all 37 genes in the positive training gene set have overlap with the gene list of HCMDB in relation to both of different metastatic cancers and metastatic breast cancer. These 37 genes must be excluded to analyze the results. Table 3 (a, b) present the frequency of driver genes enriched in the public databases for MBCA and BRCA.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 3\u003c/strong\u003e The enrichment rate of driver genes predicted by EARN. (a) MBCA, (b) BRCA\u003c/p\u003e\n\u003ctable border=\"1\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"4\" width=\"47\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(a)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMBCA\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"236\"\u003e\n\u003cp\u003e\u003cstrong\u003eAll different cancers\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"237\"\u003e\n\u003cp\u003e\u003cstrong\u003eMetastatic breast cancer\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"3\" width=\"236\"\u003e\n\u003cp\u003e\u003cstrong\u003eHCMDB\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"237\"\u003e\n\u003cp\u003e\u003cstrong\u003eHCMDB\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"82\"\u003e\n\u003cp\u003e\u003cstrong\u003ePGECCH \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"74\"\u003e\n\u003cp\u003e\u003cstrong\u003eRGCHP \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"80\"\u003e\n\u003cp\u003e\u003cstrong\u003ePGECH \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(%)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003ePGEMCH (#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003eRGMHP \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"86\"\u003e\n\u003cp\u003e\u003cstrong\u003ePGEMH \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(%)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"82\"\u003e\n\u003cp\u003e292\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"74\"\u003e\n\u003cp\u003e2203\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"80\"\u003e\n\u003cp\u003e13.25%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e73*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e585\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"86\"\u003e\n\u003cp\u003e12.48%\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"4\" width=\"47\"\u003e\n\u003cp\u003e\u003cstrong\u003e(b)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBRCA\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"236\"\u003e\n\u003cp\u003e\u003cstrong\u003eAll different cancers\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"237\"\u003e\n\u003cp\u003e\u003cstrong\u003eBreast cancer\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd colspan=\"3\" width=\"236\"\u003e\n\u003cp\u003e\u003cstrong\u003eOMIM, CGC, and NCG\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"3\" width=\"237\"\u003e\n\u003cp\u003e\u003cstrong\u003eOMIM, CGC, and NCG\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"82\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGECC \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"74\"\u003e\n\u003cp\u003e\u003cstrong\u003eRKGCP \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"80\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGECC \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(%)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGEBC \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003eRKGBP\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(#)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"86\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGEBC \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e(%)\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"82\"\u003e\n\u003cp\u003e1398\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"74\"\u003e\n\u003cp\u003e2403\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"80\"\u003e\n\u003cp\u003e58.18%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e145\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e201\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"86\"\u003e\n\u003cp\u003e72.14%\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003ctable border=\"1\" width=\"602\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"602\"\u003e\n\u003cp\u003e*These 73 genes have been also cited in 108 studies of HCMDB [see Additional file 9: S42]\u003c/p\u003e\n\u003cp\u003ePGECCH: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with different metastatic \u003cstrong\u003ec\u003c/strong\u003eancers that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in \u003cstrong\u003eH\u003c/strong\u003eCMDB\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003e\n\u003cp\u003eRGCHP: \u003cstrong\u003eR\u003c/strong\u003eemained \u003cstrong\u003eg\u003c/strong\u003eenes related to different metastatic \u003cstrong\u003ec\u003c/strong\u003eancers in the \u003cstrong\u003eH\u003c/strong\u003eCMDB after excluding \u003cstrong\u003ep\u003c/strong\u003eositive training set\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003e\n\u003cp\u003ePGEMCH: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with \u003cstrong\u003eM\u003c/strong\u003eetastatic breast cancer that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in \u003cstrong\u003eH\u003c/strong\u003eCMDB\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003e\n\u003cp\u003ePKGECC: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with different \u003cstrong\u003ec\u003c/strong\u003eancers that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd\u003e\n\u003cp\u003eRKGCP: \u003cstrong\u003eR\u003c/strong\u003eemained \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes related to different \u003cstrong\u003ec\u003c/strong\u003eancers in the \u003cstrong\u003ep\u003c/strong\u003eublic databases after excluding positive training set\u003c/p\u003e\n\u003cp\u003ePKGEBC: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with \u003cstrong\u003eb\u003c/strong\u003ereast cancer that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003eRKGBP: \u003cstrong\u003eR\u003c/strong\u003eemained \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes related to \u003cstrong\u003eb\u003c/strong\u003ereast cancer in the \u003cstrong\u003ep\u003c/strong\u003eublic databases after excluding positive training set\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eAlso, the top 50 genes predicted by all learning methods for MBCA are searched in the list of metastatic cancer-associated genes in the HCMDB. The comparison shows the enrichment score of 24%, 22%, and 16% for RF, ANN, and NLSVM compared to 24% for EARN. Although the value of enrichment in the top 50 is the same for EARN and RF, the number of studies that introduce these enriched genes is 59 for the EARN method compared to 22 for RF. Table 4 presents these genes and also provides more information about them.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 4\u003c/strong\u003e 12 driver genes predicted by EARN\u003csub\u003e50\u003c/sub\u003e which are confirmed for metastatic cancers in the HCMDB. Also, the rank number, score, and mutation count for these genes are provided in the table. The confirmed genes as the known genes related to any primary cancers or primary breast tumors in OMIM, CGC, and NCG databases have been marked in the last two columns\u003c/p\u003e\n\u003ctable border=\"1\" width=\"626\"\u003e\n\u003ctbody\u003e\n\u003ctr style=\"height: 48px;\"\u003e\n\u003ctd style=\"height: 48px;\" width=\"69\"\u003e\n\u003cp\u003e\u003cstrong\u003eSymbol\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"73\"\u003e\n\u003cp\u003e\u003cstrong\u003ePrediction score\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"47\"\u003e\n\u003cp\u003e\u003cstrong\u003eRank\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"76\"\u003e\n\u003cp\u003e\u003cstrong\u003ePSMM \u003cbr /\u003e \u003c/strong\u003e(51)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"65\"\u003e\n\u003cp\u003e\u003cstrong\u003ePSMM \u003cbr /\u003e \u003c/strong\u003e(52)\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"66\"\u003e\n\u003cp\u003e\u003cstrong\u003eNSCGMCH \u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"63\"\u003e\n\u003cp\u003e\u003cstrong\u003eNSCGMBH\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"60\"\u003e\n\u003cp\u003e\u003cstrong\u003eMCMGM\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"54\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGECC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 48px;\" width=\"54\"\u003e\n\u003cp\u003e\u003cstrong\u003ePKGEBC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eAPEX1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.900511991\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e0.50%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e1.70%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 37px;\"\u003e\n\u003ctd style=\"height: 37px;\" width=\"69\"\u003e\n\u003cp\u003eARID1A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"73\"\u003e\n\u003cp\u003e0.895213526\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"47\"\u003e\n\u003cp\u003e11\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"76\"\u003e\n\u003cp\u003e2.40%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"65\"\u003e\n\u003cp\u003e5.10%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"66\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"60\"\u003e\n\u003cp\u003e24\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eKDM6B\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.894029187\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e13\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e1.40%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e4.60%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e16\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 37px;\"\u003e\n\u003ctd style=\"height: 37px;\" width=\"69\"\u003e\n\u003cp\u003eTBX3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"73\"\u003e\n\u003cp\u003e0.893837209\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"47\"\u003e\n\u003cp\u003e14\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"76\"\u003e\n\u003cp\u003e2.80%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"65\"\u003e\n\u003cp\u003e5.10%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"66\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"60\"\u003e\n\u003cp\u003e21\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 37px;\"\u003e\n\u003ctd style=\"height: 37px;\" width=\"69\"\u003e\n\u003cp\u003eKDR*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"73\"\u003e\n\u003cp\u003e0.890079401\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"47\"\u003e\n\u003cp\u003e17\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"76\"\u003e\n\u003cp\u003e0.90%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"65\"\u003e\n\u003cp\u003e1.70%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"66\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"63\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"60\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eSERPINE2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.889205475\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e19\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e0.90%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e0.80%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 37px;\"\u003e\n\u003ctd style=\"height: 37px;\" width=\"69\"\u003e\n\u003cp\u003eTBL1XR1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"73\"\u003e\n\u003cp\u003e0.871240171\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"47\"\u003e\n\u003cp\u003e27\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"76\"\u003e\n\u003cp\u003e0.90%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"65\"\u003e\n\u003cp\u003e0.80%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"66\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"60\"\u003e\n\u003cp\u003e4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 37px;\"\u003e\n\u003ctd style=\"height: 37px;\" width=\"69\"\u003e\n\u003cp\u003eKRAS\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"73\"\u003e\n\u003cp\u003e0.868267682\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"47\"\u003e\n\u003cp\u003e30\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"76\"\u003e\n\u003cp\u003e1.40%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"65\"\u003e\n\u003cp\u003e1.70%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"66\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"60\"\u003e\n\u003cp\u003e7\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 37px;\" width=\"54\"\u003e\n\u003cp\u003e✓\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eNOS3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.861560093\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e31\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e2.40%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e2.10%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e12\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eRAPGEF3\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.851947423\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e42\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e2.50%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e6\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eSELE*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.847865292\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e49\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e0.90%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e1.30%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e12\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e5\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr style=\"height: 35px;\"\u003e\n\u003ctd style=\"height: 35px;\" width=\"69\"\u003e\n\u003cp\u003eMME*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"73\"\u003e\n\u003cp\u003e0.847698297\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"47\"\u003e\n\u003cp\u003e50\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"76\"\u003e\n\u003cp\u003e0.90%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"65\"\u003e\n\u003cp\u003e2.50%\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"66\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"63\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"60\"\u003e\n\u003cp\u003e9\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd style=\"height: 35px;\" width=\"54\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e* These genes have been specifically introduced concerning metastatic breast cancer.\u003c/p\u003e\n\u003cp\u003ePSMM: \u003cstrong\u003eP\u003c/strong\u003eercentage of \u003cstrong\u003es\u003c/strong\u003eamples with one or more \u003cstrong\u003em\u003c/strong\u003eutations based on initial \u003cstrong\u003em\u003c/strong\u003eutation file\u003c/p\u003e\n\u003cp\u003eNSCGMCH: \u003cstrong\u003eN\u003c/strong\u003eumber of \u003cstrong\u003es\u003c/strong\u003etudies that have \u003cstrong\u003ec\u003c/strong\u003eited \u003cstrong\u003eg\u003c/strong\u003eenes related to different \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003ec\u003c/strong\u003eancers in the \u003cstrong\u003eH\u003c/strong\u003eCMDB\u003c/p\u003e\n\u003cp\u003eNSCGMBH: \u003cstrong\u003eN\u003c/strong\u003eumber of \u003cstrong\u003es\u003c/strong\u003etudies that have \u003cstrong\u003ec\u003c/strong\u003eited \u003cstrong\u003eg\u003c/strong\u003eenes related to \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003eb\u003c/strong\u003ereast cancer in \u003cstrong\u003eH\u003c/strong\u003eCMDB\u003c/p\u003e\n\u003cp\u003eMCMGM: \u003cstrong\u003eM\u003c/strong\u003eutation \u003cstrong\u003ec\u003c/strong\u003eounts for \u003cstrong\u003em\u003c/strong\u003eutated \u003cstrong\u003eg\u003c/strong\u003eenes across 450 metastasis tumor samples based on the initial \u003cstrong\u003em\u003c/strong\u003eutation file\u003c/p\u003e\n\u003cp\u003ePKGECC: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with different \u003cstrong\u003ec\u003c/strong\u003eancers that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003ePKGEBC: \u003cstrong\u003eP\u003c/strong\u003eredicted \u003cstrong\u003ek\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes by \u003cstrong\u003eE\u003c/strong\u003eC associated with \u003cstrong\u003eB\u003c/strong\u003ereast cancer that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003eFurthermore, 38 genes listed by EARN\u003csub\u003e50\u003c/sub\u003e have not been introduced in the HCMDB related to any metastatic cancers. So, these genes can be considered as new genes for more investigations [see Additional file 9: Table S43].\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.2.5 Statistical validation of three individual classifiers and EARN based on evaluation measures\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eFor MBCA, a comparison of the metrics based on 3-fold cross-validation on the test data shows that EARN and ANN achieve the best precision with zero FPR. Also, accuracy, F1 score, average precision, and recall for EARN and ANN are better than the others, especially compared with NLSVM. It can be also observed that EARN has the best ROC-AUC (99.24%). Thus, in overall, the proposed EARN outperforms the other three learning methods. For comparison, evaluation metrics of learning methods for MBCA and BRCA are presented in table 5 (a, b).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 5\u003c/strong\u003e Validation of four learning methods by some evaluation metrics. (a) MBCA, (b) BRCA\u003c/p\u003e\n\u003ctable border=\"1\" width=\"86%\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e\u003cstrong\u003eMethod name\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e\u003cstrong\u003eF1 score\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e\u003cstrong\u003eFalse Positive Rate\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e\u003cstrong\u003eMaximum Precision\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e\u003cstrong\u003eAverage-Precision\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e\u003cstrong\u003erecall\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e\u003cstrong\u003eROC-AUC*\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"4\" width=\"12%\"\u003e\n\u003cp\u003e\u003cstrong\u003e(a)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMBCA\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eEARN\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.7961\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e1.0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.8266\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.6701\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9924\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eRF\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.7560\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.0008\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9069\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.7873\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.6603\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9418\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eANN\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.7990\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e1.0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.8074\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.6733\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9680\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eNLSVM\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.3972\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.0154\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.3092\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.5852\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.5885\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9770\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"4\" width=\"12%\"\u003e\n\u003cp\u003e\u003cstrong\u003e(b)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBRCA\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eEARN\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.9313\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e1.0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9585\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.8749\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9979\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eRF\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.8864\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.0019\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9061\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9171\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.8774\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9719\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eANN\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.8996\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e1.0\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9417\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.8225\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9873\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003eNLSVM\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"11%\"\u003e\n\u003cp\u003e0.5441\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.0279\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.4460\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.8590\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"12%\"\u003e\n\u003cp\u003e0.8422\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"13%\"\u003e\n\u003cp\u003e0.9926\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e* Receiver Operating Characteristic-Area under Curve\u003c/p\u003e\n\u003cp\u003eThe comparative survey in table 5 shows when we use a larger mutation dataset (983 tumor samples for BRCA vs. 450 tumor samples for MBCA) for feature extraction, where positive set is larger (40 for BRCA vs. 37 for MBCA), and negative set is smaller (2151 for BRCA vs. 3473 for MBCA), EARN achieves better statistical results. Among all statistical validation metrics, F1 score as a measure of combining the precision and recall has been used to compare performance of the learning methods for both BRCA and MBCA (Fig. 4d).\u003c/p\u003e \n\u003cp\u003e\u003cstrong\u003e3.3 BRCA and MBCA\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3.3.1 Targeted gene panel discovery for MBCA based on pathway enrichment analysis (PEA) \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this section, a pathway-based biological analysis is carried out by ReactomeFIVIz tool (65\u0026ndash;67). For EARN\u003csub\u003e100\u003c/sub\u003e, we find 63 (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) such pathways for BRCA and 42 (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) such pathways for MBCA. It is observed that 14 (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) enriched pathways are common among BRCA and MBCA (Fig. 5a), [see these specific and common pathways and the genes involved in each pathway in Additional file 10: Table S44]. These enriched pathways for BRCA are a subset of the other seven main pathways: Extracellular matrix organization, Signal Transduction, Gene expression (Transcription), Immune System, Hemostasis, Developmental Biology, and Metabolism of RNA. Also, the main pathways of MBCA include Gene expression (Transcription), Signal Transduction, Chromatin organization, Circadian Clock, Organelle biogenesis and maintenance, Neuronal System, and Metabolism. The common and specific main pathways (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) of BRCA and MBCA, and the frequency of genes involved in these main pathways are compared in Fig. 5b and Table 6. Given this, it can be found two (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) such common main pathways consist of Signal Transduction and Gene expression (Transcription) for BRCA and MBCA, and 5 (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) such specific main pathways for each of them.\u003c/p\u003e\n \n\u003cp\u003e\u0026nbsp;\u003cstrong\u003eTable 6\u003c/strong\u003e The common and specific main pathways for BRCA and MBCA\u003c/p\u003e\n\u003ctable border=\"1\" width=\"567\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"2\" width=\"95\"\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd rowspan=\"2\" width=\"43\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eNumber \u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd rowspan=\"2\" width=\"106\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003ePathways\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"173\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eBRCA\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"150\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eMBCA\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eNumber of genes\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eName of genes \u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eNumber of genes\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eName of genes \u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"5\" width=\"95\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eThe specific main pathways for BRCA\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e1\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eExtracellular matrix organization\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e8\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eDCN, FN1, ICAM1, ITGA4, ITGAM, ITGAV, ITGB3, ITGB5\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e2\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eImmune System\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e15\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eFN1, GAB2, ICAM1, IL1RAPL1, IL1RN, IL2RB, ITGAM, ITGAV, ITGB5, JAK1, MSN, POU2F1, PTPN11, SMARCA4, SYK\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e3\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eHemostasis\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e12\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eEGF, FN1, GRB7, ITGA4, ITGAM, ITGAV, ITGB3, PIK3CG, PRKCZ, PTPN11, SERPINA1, SYK\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e4\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eDevelopmental Biology\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e9\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eACVR1B, GAB1, GAB2, GRB7, PTPN11, RELN, SMAD2, SMAD4, VLDLR\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e5\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eMetabolism of RNA\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e6\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eCPSF1, CPSF3, PCF11, PRPF40A,\u0026nbsp; SF3A1, SF3B1\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"5\" width=\"95\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eThe specific main pathways for MBCA\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e6\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eChromatin organization\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e7\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eTBL1XR1, NCOR1, HDAC3, GPS2, ACTB, KDM6B, PRMT1\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e7\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eCircadian Clock\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e2\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eNCOR1, HDAC3\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e8\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eOrganelle biogenesis and maintenance\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e4\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eTBL1XR1, SIRT4, NCOR1, HDAC3\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e9\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eNeuronal System\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e7\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eABAT, KPNA2, PRKCG, CACNA1E, PLCB1, GRIN1, KRAS\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e10\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eMetabolism\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e0\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eNone\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e5\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eTBL1XR1, SIN3A, NCOR1, HDAC3, GPS2\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd rowspan=\"2\" width=\"95\"\u003e\n\u003cp\u003e\u003cstrong\u003e\u003csub\u003eThe common main pathways for BRCA and MBCA\u003c/sub\u003e\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e11\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eSignal Transduction\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e25\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003eACVR1B, EGF, ERBB3, FLT1, FN1, GAB1, GAB2, GRB7, ITGAV, ITGB3, JAK1, NOTCH4, NR4A1, PARD3, PPARG, PRKCZ, PTEN, PTPN11, RUNX1, SMAD2, SMAD4, SMURF1, SYK, TFDP1, TGFBR2 \u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e25\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eACTB, AR, BDNF, BUB1B, CBFB, COL4A3, FOXA1, KDR, KPNA2, KRAS, NCOR1, NOS3, PDGFD, PIK3R1, PKN2, PLCB1, PRKCD, PRKCG, PRMT1, PTPRJ, RUNX1, STAG1, STAT1, WAS, YWHAE\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e12\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"106\"\u003e\n\u003cp\u003e\u003csub\u003eGene expression (Transcription)\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"48\"\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e19\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"125\"\u003e\n\u003cp\u003e\u003csub\u003eABL1, CBFB, CPSF1, CPSF3, MED23, NBN, NOTCH4, NR4A1, PCF11, POU2F1, PPARG, PTEN, PTPN11, RUNX1, SMAD2, SMAD4, SMARCA4, SMURF1, TFDP1\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"43\"\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e\u0026nbsp;\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003csub\u003e14\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"107\"\u003e\n\u003cp\u003e\u003csub\u003eAR, BDNF, CBFB, GPS2, HDAC3, KLF4, KRAS, NCOR1, PRMT1, RUNX1, SIN3A, STAT1, TBL1XR1, YWHAE\u003c/sub\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003eFurther investigation in Table 6 shows that 16 genes contribute to five enriched specific main pathways of MBCA. Among them, four genes are involved in more than one main pathway. In particular, NCOR1 and HDAC3 are engaged in four pathways. In three out of five pathways TBL1XR1 is active, and GPS2 gets involved in two pathways. Table 7 introduces 16 genes that are enriched in these five main pathways and provides more information about them.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTable 7\u003c/strong\u003e The plausible driver genes involved in the proposed main pathways related to MBCA\u003c/p\u003e\n\u003ctable border=\"1\" width=\"594\"\u003e\n\u003ctbody\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"7\" width=\"366\"\u003e\n\u003cp\u003e\u003cstrong\u003eSpecific main pathways\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003e\u003cstrong\u003ePPDMB\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e\u003cstrong\u003eKGCC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e\u003cstrong\u003eKGBC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e\u003cstrong\u003eIGMC\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e\u003cstrong\u003eIGMB\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e\u003cstrong\u003eChromatin organization\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"68\"\u003e\n\u003cp\u003e\u003cstrong\u003eCircadian Clock\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"85\"\u003e\n\u003cp\u003e\u003cstrong\u003eOrganelle biogenesis and maintenance\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"66\"\u003e\n\u003cp\u003e\u003cstrong\u003eNeuronal System\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e\u003cstrong\u003eMetabolism\u003c/strong\u003e\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eNCOR1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eHDAC3*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eTBL1XR1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eSIRT4\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eABAT*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eKRAS\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e20\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eGRIN1*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003ePLCB1*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eCACNA1E\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003ePRKCG\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eKPNA2*\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eGPS2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eSIN3A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eACTB\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003eKDM6B\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003ctr\u003e\n\u003ctd width=\"60\"\u003e\n\u003cp\u003ePRMT1\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e2\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"42\"\u003e\n\u003cp\u003e#N/A\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"72\"\u003e\n\u003cp\u003e✔\u003c/p\u003e\n\u003c/td\u003e\n\u003ctd width=\"60\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"84\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd colspan=\"2\" width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003ctd width=\"75\"\u003e\u0026nbsp;\u003c/td\u003e\n\u003c/tr\u003e\n\u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e*\u0026nbsp; Five new genes that have not been already introduced in the public databases\u003c/p\u003e\n\u003cp\u003ePPDMB: \u003cstrong\u003eP\u003c/strong\u003eroposed \u003cstrong\u003ep\u003c/strong\u003elausible \u003cstrong\u003ed\u003c/strong\u003eriver related to \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003eb\u003c/strong\u003ereast cancer\u003c/p\u003e\n\u003cp\u003eKGCC: \u003cstrong\u003eK\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes related to \u003cstrong\u003ec\u003c/strong\u003eancers that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003eKGBC: \u003cstrong\u003eK\u003c/strong\u003enown \u003cstrong\u003eg\u003c/strong\u003eenes related to \u003cstrong\u003eb\u003c/strong\u003ereast cancer that are \u003cstrong\u003ec\u003c/strong\u003eonfirmed in OMIM, CGC, and NCG\u003c/p\u003e\n\u003cp\u003eIGMC: \u003cstrong\u003eI\u003c/strong\u003entroduced \u003cstrong\u003eg\u003c/strong\u003eenes related to different \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003ec\u003c/strong\u003eancers in HCMDB\u003c/p\u003e\n\u003cp\u003eIGMB: \u003cstrong\u003eI\u003c/strong\u003entroduced \u003cstrong\u003eg\u003c/strong\u003eenes related to \u003cstrong\u003em\u003c/strong\u003eetastatic \u003cstrong\u003eb\u003c/strong\u003ereast cancer in HCMDB\u003c/p\u003e\n\u003cp\u003eThis gene set can be considered as a targeted biomarker panel in the case of metastatic breast cancer to examine more in the next molecular and clinical analysis phase. More investigations on these genes can hopefully be helpful in MBCA prognosis and diagnosis. Table 7 shows that five genes, HDAC3, ABAT, GRIN1, PLCB1, and KPNA2 are new and not confirmed in the public databases for cancer prognosis. However, there is some evidence to suggest that these genes play a clinical role in cancer progression. HDAC3 contributes to four pathways alongside NCOR1. The other four genes engage in the Neuronal System pathway. The recent investigations on Basal-like breast cancer (BLBC), the most aggressive subtype of this cancer, have documented the expression of ABAT was considerably decreased in this cancer (69). Besides, alterations in the expression levels of ABAT have been reported in the promotion of breast cancer (70). ABAT was also identified as a biomarker for endocrine-responsiveness breast cancer patients (71). Furthermore, GRIN1 encodes GluN1 subunit of N-methyl-D-aspartate receptor (NMDAR). It has been shown that this subunit in more than 90% of all breast cancer subtypes is uniformly expressed to promote Breast-to-brain metastasis (B2BM) (72). Recently, the role of HDAC3 in the deregulation of P53 pathway in the aneuploid cancer cell lines has been analyzed (73). Also, HDAC3 is overexpressed in breast cancer patients.\u0026nbsp; It has been illustrated that breast cancer stem cells, which are resistant to treatment and are responsible for metastasis, are the target of the histone deacetylase (HDAC) inhibitors (74). On the other, The results of enrichment in cBioPortal show that the above-mentioned 16 genes are altered in 243 (54%) of 450 MBCA samples in two studies performed in 2016 (51) and 2017 (52). Genomic alterations (Fig. 6) in these genes have been visualized using OncoPrint component \u0026nbsp;(49, 50). Among them, the highest percentage of somatic mutation frequency (SMF) is observed in CACNA1E, NCOR1, KDM6B, and GPS2. Using the Needle Plot component (49, 50), we visualize SMF and can also map mutations on the linear protein and its domains for these four genes (Fig. 7).\u003c/p\u003e"},{"header":"4. Discussion","content":" \u003cp\u003eIn this work, we proposed an EC machine learning method called EARN, combining three base classifiers to predict and estimate the potential of plausible driver genes in BRCA and MBCA. Leveraged by both feature fusion and decision fusion, the proposed ensemble model made better decisions in comparison with base classifiers, especially in the list of the top genes. Although EARN uses the simple average operator for aggregating the decisions of the three base learners to predict the driver genes, it could find some new genes in the list of EARN\u003csub\u003e100\u003c/sub\u003e, which were not observed in the top\u003csub\u003e100\u003c/sub\u003e of the individual classifiers. It can be rational evidence for using the ensemble systems for gene prioritization.\u003c/p\u003e \u003cp\u003eFor biological validation of outputs and after the enrichment of EARN\u003csub\u003e50\u003c/sub\u003e in the public databases, where the ensemble learning method uses most of the power to discriminate and predict, we could obtain the enrichment rate of 52% for BRCA, which outperforms the three individual classifiers. For MBCA, the enrichment of EARN\u003csub\u003e50\u003c/sub\u003e in the HCMDB resulted in an enrichment rate of 24%, which is better than the two base classifiers, NLSVM and ANN, while being comparable to RF.\u003c/p\u003e \u003cp\u003eThe results are also analyzed using a statistical test with cross-validation. The evaluation of results showed that EARN performs well, especially for BRCA. In the case of BRCA, the open-access mutation annotation format (.maf) file is larger and the mutation data is obtained from more samples (983 BRCA tumor samples vs. 450 MBCA tumor samples). Thus, the proper features could be extracted. Finally, the performance of EARN for ranking human protein-coding genes is improved. Further, to evaluate the possibility of enhancement in the combination of the base classifiers results, we tried StackingCVClassifier, an effective ensemble-learning meta-classifier for stacking (\u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e75\u003c/span\u003e, \u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e76\u003c/span\u003e). For BRCA, there was no improvement in the results. It could be because the results of the originally proposed ensemble model were good enough. While the metrics such as F1 score (81.31% vs. 79.61%) and recall (69.62 vs. 67.02) were slightly improved for MBCA.\u003c/p\u003e \u003cp\u003eFinally, the existence of specific enriched pathways by ReactomeFIVIz (\u003cem\u003eFDR\u003c/em\u003e\u0026thinsp;\u0026lt;\u0026thinsp;0.03) for the top genes predicted by EARN for BRCA and MBCA led us to suggest a gene panel regarding metastatic breast cancer.\u003c/p\u003e \u003cp\u003eIn present study, we faced some limitations to find the appropriate drivers of MBCA. This fact that the original mutation datasets involved in the whole-exome sequencing of the tumor samples of the metastatic breast cancer patients are small. Also, the lack of definitive driver genes confirmed in the public databases for metastatic cancers makes it difficult to select a positive training set. These issues decreased the performance of EARN for MBCA in comparison with BRCA. Further, the result of enriching all predicted genes by EARN for BRCA in the OMIM, CGC, and NCG was encouraging (72.14%, Please refer to Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e (b)). But, the result of the enrichment of the predicted genes by EARN for MBCA was not satisfactory (12.48%, see Table \u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e (a)). This may be due to the lack of sufficient studies on metastatic cancers, and particularly because of the limited databases regarding metastatic cancers to enrich driver genes.\u003c/p\u003e "},{"header":"5. Conclusions","content":" \u003cp\u003eSince using computational methods such as ensemble machine learning approaches are less expensive than bio-molecular techniques, it can help to significantly reduce the search space for bio-molecular and medical science researchers in the identification of plausible driver genes to facilitate prognosis and diagnosis of complex diseases. In this work, we mainly focused on the use of genomics data. Meanwhile, the changes of epigenomic, genomic, transcriptional, and proteomic that occur during progression to metastatic encourage us to use multi-omics integration (\u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e77\u003c/span\u003e). It has been demonstrated that multi-Omics data integration can improve predictive performance (\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e78\u003c/span\u003e) (e.g., it has been applied to predict robust biomarkers of drug efficacy for targeted therapies in triple-negative breast cancer (\u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e79\u003c/span\u003e)). A direction of future research would be to apply a combination of different levels of data, including genomics, epigenomics, transcriptomics, proteomics, metabolomics, and microbiomics data to optimize the ensemble system for introducing Omics-driven markers. In the end, we emphasize this research needs clinical trials to be validated and to evaluate the potential of the proposed drivers for discrimination between different stages of cancers. The limited number of markers obtained from the trial validation assists precision oncologists to design compact targeted panels that eliminate the need for whole-genome/exome sequencing.\u003c/p\u003e "},{"header":"Abbreviations","content":" \u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003e\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eB2BM: Breast-to-brain metastasis; BLBC:Basal-like breast cancer; BRCA:primary breast invasive carcinoma; CGC:Cancer Gene Census; DECORATE:Diverse Ensemble Creation by Oppositional Relabeling of Artificial Training Examples; EARN:Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine; EC:ensemble classifier; FPR:false-positive rate; GSEA:gene set enrichment analysis; HCMDB:human cancer metastasis database; HPA:Human Protein Atlas; HyDRA:Hybrid Distance-score Rank Aggregation; INDELs:insertions and deletions; maf:mutation annotation format; MBCA:metastatic breast cancer; NCG:Network of Cancer Genes; NGS:next-generation sequencing; NLSVM:non-linear Support Vector Machine; NMDAR:N-methyl-D-aspartate receptor; OMIM:Online Mendelian Inheritance in Man; OV:ovarian; PEA:pathway enrichment analysis; PPV:positive predictive value; RF:Random Forest; ROC-AUC:Receiver Operating Characteristic-Area under Curve; SMF:somatic mutation frequency; TCGA:The Cancer Genome Atlas; WES:Whole-Exome Sequencing\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e "},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eEthics approval and consent to participate\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConsent for publication\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll required data is available in Additional files 1-3.\u003c/p\u003e\n\u003cp\u003eWe have also shared Python source code and other requirements for implementation of the proposed ensemble machine learning algorithm as the protocol via GitHub (https://github.com/lmirsadeghi/EARN/).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study was financially supported by grant No: 960903 of the Biotechnology Development Council of the Islamic Republic of Iran\u003cstrong\u003e.\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors' contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLM and KK developed the concept, designed the research. LM performed the required research works and developed the solution under the joint supervision of KK, RHH, and AMBM. LM wrote the manuscript and contributed to visualize results. LM and KK contributed to the interpretation of the data and discussion. KK, RHH, and AMBM gave final approval for publication.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors like to thank Dr. Hossein Hajimirsadeghi[1] for his useful advice and invaluable help during this research.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors' information\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e1\u003c/sup\u003e Department of Biology, Faculty of Science, Payame Noor University, Tehran, Iran.\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e2\u003c/sup\u003eLaboratory of Genomics and Epigenomics (LGE), Department of Biochemistry, Institute of Biochemistry and Biophysics (IBB), University of Tehran, Tehran, Iran.\u003c/p\u003e\n\u003cp\u003e\u003csup\u003e3\u003c/sup\u003eLaboratory of Complex Biological Systems and Bioinformatics (CBB), Department of Bioinformatics, Institute of Biochemistry and Biophysics (IBB), University of Tehran, Tehran, Iran.\u003c/p\u003e\n\u003cp\u003e[1] https://hossein-h.github.io/\u003c/p\u003e\n\u003cp\u003eE-mail: [email protected]\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eKumar A, Singla A. Epidemiology of Breast Cancer: Current Figures and Trends. In: Preventive Oncology for the Gynecologist. Springer; 2019. p. 335\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eZhao D, Qiao J, He H, Song J, Zhao S, Yu J. TFPI2 suppresses breast cancer progression through inhibiting TWIST-integrin \u0026alpha;5 pathway. Mol Med. 2020;26:1\u0026ndash;10.\u003c/li\u003e\n\u003cli\u003eSheikine Y, Kuo FC, Lindeman NI. Clinical and technical aspects of genomic diagnostics for precision oncology. J Clin Oncol. 2017;35(9):929\u0026ndash;33.\u003c/li\u003e\n\u003cli\u003eSmith NG, Gyanchandani R, Shah OS, Gurda GT, Lucas PC, Hartmaier RJ, et al. Targeted mutation detection in breast cancer using MammaSeq\u003csup\u003eTM\u003c/sup\u003e. Breast Cancer Res. 2019;21(1):22.\u003c/li\u003e\n\u003cli\u003eKulasingam V, Diamandis EP. Strategies for discovering novel cancer biomarkers through utilization of emerging technologies. Nat Rev Clin Oncol. 2008;5(10):588.\u003c/li\u003e\n\u003cli\u003eRajkomar A, Dean J, Kohane I. Machine learning in medicine. N Engl J Med. 2019;380(14):1347\u0026ndash;58.\u003c/li\u003e\n\u003cli\u003eBaronti F, Micheli A, Passaro A, Starita A. Machine learning contribution to solve prognostic medical problems. Outcome Predict Cancer. 2006;261.\u003c/li\u003e\n\u003cli\u003eBreiman L. Bagging predictors. Mach Learn. 1996;24(2):123\u0026ndash;40.\u003c/li\u003e\n\u003cli\u003eHosni M, Abnane I, Idri A, de Gea JMC, Alem\u0026aacute;n JLF. Reviewing Ensemble Classification Methods in Breast Cancer. Comput Methods Programs Biomed. 2019;\u003c/li\u003e\n\u003cli\u003eMirsadeghi L, Banaei-Moghaddam AM, Beh-Afarin SR, Haji R. A post-method condition analysis of using ensemble machine learning for cancer prognosis and diagnosis: a systematic review.\u003c/li\u003e\n\u003cli\u003eGevaert O, De Smet F, Timmerman D, Moreau Y, De Moor B. Predicting the prognosis of breast cancer by integrating clinical and microarray data with Bayesian networks. Bioinformatics. 2006;22(14):e184\u0026ndash;90.\u003c/li\u003e\n\u003cli\u003eMoriyama T, Imoto S, Hayashi S, Shiraishi Y, Miyano S, Yamaguchi R. A Bayesian model integration for mutation calling through data partitioning. Bioinformatics. 2019;\u003c/li\u003e\n\u003cli\u003eCheriguene S, Azizi N, Zemmal N, Dey N, Djellali H, Farah N. Optimized Tumor Breast Cancer Classification Using Combining Random Subspace and Static Classifiers Selection Paradigms. In: Applications of Intelligent Optimization in Biology and Medicine. Springer; 2016. p. 289\u0026ndash;307.\u003c/li\u003e\n\u003cli\u003eLes T, Markiewicz T, Osowski S, Kozlowski W, Jesiotr M. Fusion of FISH image analysis methods of HER2 status determination in breast cancer. Expert Syst Appl. 2016;61:78\u0026ndash;85.\u003c/li\u003e\n\u003cli\u003eZakeri P, Elshal S, Moreau Y. Gene prioritization through geometric-inspired kernel data fusion. In: Bioinformatics and Biomedicine (BIBM), 2015 IEEE International Conference on. IEEE; 2015. p. 1559\u0026ndash;65.\u003c/li\u003e\n\u003cli\u003eLiu Y, Tian F, Hu Z, DeLisi C. Evaluation and integration of cancer gene classifiers: identification and ranking of plausible drivers. Sci Rep. 2015;5.\u003c/li\u003e\n\u003cli\u003eKim M, Farnoud F, Milenkovic O. HyDRA: gene prioritization via hybrid distance-score rank aggregation. Bioinformatics. 2015;31(7):1034\u0026ndash;43.\u003c/li\u003e\n\u003cli\u003eReboiro-Jato M, D\u0026iacute;az F, Glez-Pe\u0026ntilde;a D, Fdez-Riverola F. A novel ensemble of classifiers that use biological relevant gene sets for microarray classification. Appl Soft Comput. 2014;17:117\u0026ndash;26.\u003c/li\u003e\n\u003cli\u003eKuncheva LI, Rodr\u0026iacute;guez JJ. A weighted voting framework for classifiers ensembles. Knowl Inf Syst. 2014;38(2):259\u0026ndash;75.\u003c/li\u003e\n\u003cli\u003eJanghel RR, Shukla A, Sharma S, Gnaneswar A V. Evolutionary Ensemble Model for Breast Cancer Classification. In: International Conference in Swarm Intelligence. Springer; 2014. p. 8\u0026ndash;16.\u003c/li\u003e\n\u003cli\u003eCun Y, Fr\u0026ouml;hlich H. Network and data integration for biomarker signature discovery via network smoothed t-statistics. PLoS One. 2013;8(9):e73074.\u003c/li\u003e\n\u003cli\u003eAzizi N, Tlili-Guiassa Y, Zemmal N. A computer-aided diagnosis system for breast cancer combining features complementarily and new scheme of SVM classifiers fusion. Int J Multimed Ubiquitous Eng. 2013;8(4):45\u0026ndash;58.\u003c/li\u003e\n\u003cli\u003eYang R, Daigle BJ, Petzold LR, Doyle FJ. Core module biomarker identification with network exploration for breast cancer metastasis. BMC Bioinformatics. 2012;13(1):1.\u003c/li\u003e\n\u003cli\u003eGlaab E, Bacardit J, Garibaldi JM, Krasnogor N. Using rule-based machine learning for candidate disease gene prioritization and sample classification of cancer gene expression data. PLoS One. 2012;7(7):e39932.\u003c/li\u003e\n\u003cli\u003eReboiro-Jato M, Glez-Pe\u0026ntilde;a D, D\u0026iacute;az F, Fdez-Riverola F. A novel ensemble approach for multicategory classification of DNA microarray data using biological relevant gene sets. Int J Data Min Bioinform. 2012;6(6):602\u0026ndash;16.\u003c/li\u003e\n\u003cli\u003eLederman D, Wang X, Zheng B, Sumkin JH, Tublin M, Gur D. Fusion of classifiers for REIS-based detection of suspicious breast lesions. In: SPIE Medical Imaging. International Society for Optics and Photonics; 2011. p. 79661C-79661C.\u003c/li\u003e\n\u003cli\u003eZeng T, Liu J. Mixture classification model based on clinical markers for breast cancer prognosis. Artif Intell Med. 2010;48(2):129\u0026ndash;37.\u003c/li\u003e\n\u003cli\u003eZhang X. Boosting twin support vector machine approach for MCs detection. In: Information Processing, 2009 APCIP 2009 Asia-Pacific Conference on. IEEE; 2009. p. 149\u0026ndash;52.\u003c/li\u003e\n\u003cli\u003eZhang X, Gao X, Wang M. MCs detection approach using Bagging and Boosting based twin support vector machine. In: Systems, Man and Cybernetics, 2009 SMC 2009 IEEE International Conference on. IEEE; 2009. p. 5000\u0026ndash;505.\u003c/li\u003e\n\u003cli\u003eDjebbari A, Liu Z, Phan S, Famili F. An ensemble machine learning approach to predict survival in breast cancer. Int J Comput Biol Drug Des. 2008;1(3):275\u0026ndash;94.\u003c/li\u003e\n\u003cli\u003eAlam KMR, Islam MM. Combining boosting with negative correlation learning for training neural network ensembles. In: 2007 International Conference on Information and Communication Technology. IEEE; 2007. p. 68\u0026ndash;71.\u003c/li\u003e\n\u003cli\u003eFranke L, Bakel H Van, Fokkens L, Jong ED De, Egmont-petersen M, Wijmenga C. Reconstruction of a Functional Human Gene Network , with an Application for Prioritizing Positional Candidate Genes. Am J Hum Genet. 2006;78(June):1011\u0026ndash;25.\u003c/li\u003e\n\u003cli\u003ePeng Y. Integration of gene functional diversity for effective cancer detection. Int J Syst Sci. 2006;37(13):931\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003eMatsui S. Genomic biomarkers for personalized medicine: development and validation in clinical studies. Comput Math Methods Med. 2013;2013.\u003c/li\u003e\n\u003cli\u003eHuang L, Jiang X-L, Liang H-B, Li J-C, Chin L-H, Wei J-P, et al. Genetic profiling of primary and secondary tumors from patients with lung adenocarcinoma and bone metastases reveals targeted therapy options. Mol Med. 2020;26(1):1\u0026ndash;11.\u003c/li\u003e\n\u003cli\u003eLan Y, Zhao E, Luo S, Xiao Y, Li X, Cheng S. Revealing clonality and subclonality of driver genes for clinical survival benefits in breast cancer. Breast Cancer Res Treat. 2019;175(1):91\u0026ndash;104.\u003c/li\u003e\n\u003cli\u003eBaesens B, Viaene S, Van Gestel T, Suykens J, Dedene G, De Moor B, et al. Least squares support vector machine classifiers: an empirical evaluation. DTEW Res Rep 0003. 2000;1\u0026ndash;16.\u003c/li\u003e\n\u003cli\u003eMaclin PS, Dempsey J, Brooks J, Rand J. Using neural networks to diagnose cancer. J Med Syst. 1991;15(1):11\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eBreiman L. Random forests. Mach Learn. 2001;45(1):5\u0026ndash;32.\u003c/li\u003e\n\u003cli\u003eLawrence MS, Stojanov P, Polak P, Kryukov G V, Cibulskis K, Sivachenko A, et al. Mutational heterogeneity in cancer and the search for new cancer-associated genes. Nature. 2013;499(7457):214\u0026ndash;8.\u003c/li\u003e\n\u003cli\u003eTamborero D, Gonzalez-Perez A, Lopez-Bigas N. OncodriveCLUST: exploiting the positional clustering of somatic mutations to identify cancer genes. Bioinformatics. 2013;29(18):2238\u0026ndash;44.\u003c/li\u003e\n\u003cli\u003eGonzalez-Perez A, Lopez-Bigas N. Functional impact bias reveals cancer drivers. Nucleic Acids Res. 2012;40(21):e169\u0026ndash;e169.\u003c/li\u003e\n\u003cli\u003eCerami E, Demir E, Schultz N, Taylor BS, Sander C. Automated network analysis identifies core pathways in glioblastoma. PLoS One. 2010;5(2):e8918.\u003c/li\u003e\n\u003cli\u003eFutreal PA, Coin L, Marshall M, Down T, Hubbard T, Wooster R, et al. A census of human cancer genes. Nat Rev cancer. 2004;4(3):177.\u003c/li\u003e\n\u003cli\u003eAn O, Pendino V, D\u0026rsquo;Antonio M, Ratti E, Gentilini M, Ciccarelli FD. NCG 4.0: the network of cancer genes in the era of massive mutational screenings of cancer genomes. Database. 2014;2014:bau015.\u003c/li\u003e\n\u003cli\u003eRepana D, Nulsen J, Dressler L, Bortolomeazzi M, Venkata SK, Tourna A, et al. The Network of Cancer Genes (NCG): a comprehensive catalogue of known and candidate cancer genes from cancer sequencing screens. Genome Biol. 2019;20(1):1.\u003c/li\u003e\n\u003cli\u003eThe experimentally supported gene-metastasis association data. 2017. https://hcmdb.isanger.com/images/hcmdb/gene_publication.xls. Accessed 22-Jun-2017.\u003c/li\u003e\n\u003cli\u003eTCGA.BRCA.muse.b8ca5856-9819-459c-87c5-94e91aca4032.DR-10.0.somatic.maf.gz. 2018. https://portal.gdc.cancer.gov/files/b8ca5856-9819-459c-87c5-94e91aca4032. Accessed 23-Aug-2018.\u003c/li\u003e\n\u003cli\u003eCerami E, Gao J, Dogrusoz U, Gross BE, Sumer SO, Aksoy BA, et al. The cBio cancer genomics portal: an open platform for exploring multidimensional cancer genomics data. AACR; 2012.\u003c/li\u003e\n\u003cli\u003eGao J, Aksoy BA, Dogrusoz U, Dresdner G, Gross B, Sumer SO, et al. Integrative analysis of complex cancer genomics and clinical profiles using the cBioPortal. Sci Signal. 2013;6(269):pl1\u0026ndash;pl1.\u003c/li\u003e\n\u003cli\u003eLefebvre C, Bachelot T, Filleron T, Pedrero M, Campone M, Soria J-C, et al. Mutational profile of metastatic breast cancers: a retrospective analysis. PLoS Med. 2016;13(12):e1002201.\u003c/li\u003e\n\u003cli\u003eWagle N, Painter C, Anastasio E, Dunphy M, McGillicuddy M, Kim D, et al. The Metastatic Breast Cancer (MBC) project: Accelerating translational research through direct patient engagement. American Society of Clinical Oncology; 2017.\u003c/li\u003e\n\u003cli\u003ecBioPortal/datahub-study-curation-tools. 2019. https://github.com/cBioPortal/datahub-study-curationtools/tree/master/split_data_clinical_sample_patient. Accessed 11-Jan-2019.\u003c/li\u003e\n\u003cli\u003eGarc\u0026iacute;a-D\u0026iacute;az P, S\u0026aacute;nchez-Berriel I, Mart\u0026iacute;nez-Rojas JA, Diez-Pascual AM. Unsupervised feature selection algorithm for multiclass cancer classification of gene expression RNA-Seq data. Genomics. 2020;112(2):1916\u0026ndash;25.\u003c/li\u003e\n\u003cli\u003eKim S, Park T, Kon M. Cancer survival classification using integrated data sets and intermediate information. Artif Intell Med. 2014;62(1):23\u0026ndash;31.\u003c/li\u003e\n\u003cli\u003eDashtban M, Balafar M, Suravajhala P. Gene selection for tumor classification using a novel bio-inspired multi-objective approach. Genomics. 2018;110(1):10\u0026ndash;7.\u003c/li\u003e\n\u003cli\u003eBhanot G, Alexe G, Venkataraghavan B, Levine AJ. A robust meta‐classification strategy for cancer detection from MS data. Proteomics. 2006;6(2):592\u0026ndash;604.\u003c/li\u003e\n\u003cli\u003ePalade V. Class imbalance learning methods for support vector machines. 2013;\u003c/li\u003e\n\u003cli\u003eWang X, Liu X, Matwin S. A distributed instance-weighted SVM algorithm on large-scale imbalanced datasets. Proc - 2014 IEEE Int Conf Big Data, IEEE Big Data 2014. 2015;45\u0026ndash;51.\u003c/li\u003e\n\u003cli\u003eMing C, Viassolo V, Probst-Hensch N, Chappuis PO, Dinov ID, Katapodi MC. Machine learning techniques for personalized breast cancer risk prediction: comparison with the BCRAT and BOADICEA models. Breast Cancer Res. 2019;21(1):75.\u003c/li\u003e\n\u003cli\u003ePolikar R. Ensemble based systems in decision making. Circuits Syst Mag IEEE. 2006;6(3):21\u0026ndash;45.\u003c/li\u003e\n\u003cli\u003eDuan X, Yang Y, Tan S, Wang S, Feng X, Cui L, et al. Application of artificial neural network model combined with four biomarkers in auxiliary diagnosis of lung cancer. Med Biol Eng Comput. 2017;55(8):1239\u0026ndash;48.\u003c/li\u003e\n\u003cli\u003eWalczak S. Artificial neural networks. In: Encyclopedia of Information Science and Technology, Fourth Edition. IGI Global; 2018. p. 120\u0026ndash;31.\u003c/li\u003e\n\u003cli\u003ePedregosa F, Varoquaux G, Gramfort A, Michel V, Thirion B, Grisel O, et al. Scikit-learn: Machine learning in Python. J Mach Learn Res. 2011;12(Oct):2825\u0026ndash;30.\u003c/li\u003e\n\u003cli\u003eFabregat A, Jupe S, Matthews L, Sidiropoulos K, Gillespie M, Garapati P, et al. The reactome pathway knowledgebase. Nucleic Acids Res. 2017;46(D1):D649\u0026ndash;55.\u003c/li\u003e\n\u003cli\u003eWu G, Haw R. Functional Interaction Network Construction and Analysis for Disease Discovery. In: Protein Bioinformatics. Springer; 2017. p. 235\u0026ndash;53.\u003c/li\u003e\n\u003cli\u003eFabregat A, Sidiropoulos K, Viteri G, Forner O, Marin-Garcia P, Arnau V, et al. Reactome pathway analysis: a high-performance in-memory approach. BMC Bioinformatics. 2017;18(1):142.\u003c/li\u003e\n\u003cli\u003eBioinformatics \u0026amp; Evolutionary Genomics. 2018. http://bioinformatics.psb.ugent.be/webtools/Venn/. Accessed 20 Nov 2018.\u003c/li\u003e\n\u003cli\u003eChen X, Cao Q, Liao R, Wu X, Xun S, Huang J, et al. Loss of ABAT-Mediated GABAergic System Promotes Basal-Like Breast Cancer Progression by Activating Ca2+-NFAT1 Axis. Theranostics. 2019;9(1):34.\u003c/li\u003e\n\u003cli\u003eZhao G, Li N, Li S, Wu W, Wang X, Gu J. High methylation of the 4-aminobutyrate aminotransferase gene predicts a poor prognosis in patients with myelodysplastic syndrome. Int J Oncol. 2019;54(2):491\u0026ndash;504.\u003c/li\u003e\n\u003cli\u003eSas L, Lardon F, Vermeulen PB, Hauspy J, Van Dam P, Pauwels P, et al. The interaction between ER and NF\u0026kappa;B in resistance to endocrine therapy. Breast Cancer Res. 2012;14(4):212.\u003c/li\u003e\n\u003cli\u003eZeng Q, Michael IP, Zhang P, Saghafinia S, Knott G, Jiao W, et al. Synaptic proximity enables NMDAR signalling to promote brain metastasis. Nature. 2019;573(7775):526\u0026ndash;31.\u003c/li\u003e\n\u003cli\u003eCilluffo D, Barra V, Spatafora S, Coronnello C, Contino F, Bivona S, et al. Aneuploid IMR90 cells induced by depletion of pRB, DNMT1 and MAD2 show a common gene expression signature. Genomics. 2020;\u003c/li\u003e\n\u003cli\u003eHii L-W, Chung FF-L, Soo JS-S, Tan BS, Mai C-W, Leong C-O. Histone deacetylase (HDAC) inhibitors and doxorubicin combinations target both breast cancer stem cells and non-stem breast cancer cells simultaneously. Breast Cancer Res Treat. 2019;1\u0026ndash;15.\u003c/li\u003e\n\u003cli\u003eTang J, Alelyani S, Liu H. Data classification: algorithms and applications. Data Min Knowl Discov Ser CRC Press. 2014;37\u0026ndash;64.\u003c/li\u003e\n\u003cli\u003eWolpert DH. Stacked generalization. Neural networks. 1992;5(2):241\u0026ndash;59.\u003c/li\u003e\n\u003cli\u003eGriffith OL, Gray JW. \u0026rsquo;Omic approaches to preventing or managing metastatic breast cancer. Breast Cancer Res. 2011;13(6):230.\u003c/li\u003e\n\u003cli\u003eRohart F, Gautier B, Singh A, L\u0026ecirc; Cao K-A. mixOmics: An R package for \u0026lsquo;omics feature selection and multiple data integration. PLoS Comput Biol. 2017;13(11):e1005752.\u003c/li\u003e\n\u003cli\u003eMerrill NM, Lachacz EJ, Vandecan NM, Ulintz PJ, Bao L, Lloyd JP, et al. Molecular determinants of drug response in TNBC cell lines. Breast Cancer Res Treat. 2020;179(2):337\u0026ndash;47.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Additional files","content":"\u003cp\u003e\u003cstrong\u003eAdditional file 1\u003c/strong\u003e_ Supplementary Table S1 (.xlsx): The original mutation files for primary breast tumors\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 2\u003c/strong\u003e_ Supplementary Table S2 and S3 (.xlsx): The original mutation files for metastasis breast tumors\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 3\u003c/strong\u003e_ Supplementary Methods (.PDF): Selection of positive/negative training sets for BRCA and MBCA\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 4\u003c/strong\u003e_ Supplementary Table S4-S7 (.xls): List of positive/negative gene set for BRCA and MBCA\u003c/p\u003e\n\u003cp\u003eSupplementary Table S8-S10 (.xlsx): Homo_sapiens genes for BRCA and MBCA\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 5\u003c/strong\u003e_ Supplementary Results (.PDF): Results for BRCA\u003c/p\u003e\n\u003cp\u003eTable S11 (.xls): Unique driver genes predicted by EARN\u003csub\u003e100\u003c/sub\u003e for BRCA\u003c/p\u003e\n\u003cp\u003eTable S12 (.xls): The list of enriched known genes of EARN\u003csub\u003e50\u003c/sub\u003e in the public databases for BRCA\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 6\u003c/strong\u003e_ Supplementary Table S13-S26 (.xls): The list of mutated genes extracted by software tools for BRCA and MBCA, and comparison among these genes (p-value\u0026le;0.05) for MBCA\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 7\u003c/strong\u003e_ Supplementary Table S27-S31 (.xls): The list of driver and passenger genes of four learning machines for MBCA\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 8\u003c/strong\u003e_ Supplementary Table S32-S41 (.xls): The comparison of drivers predicted by all machine learning methods for MBCA\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 9\u003c/strong\u003e_ Supplementary Table S42 (.xls): The list of driver genes of EARN for MBCA that have been cited in 108 studies of HCMDB\u003c/p\u003e\n\u003cp\u003eTable S43 (.xls): The list of new predicted genes by EARN\u003csub\u003e50\u003c/sub\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAdditional file 10_\u003c/strong\u003e Supplementary Table S44 (.xls): The common/specific enriched pathways for BRCA and MBCA using ReactomeFIVIz (FDR \u0026lt; 0.03)\u003c/p\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Metastasis breast tumor, Mutation data, Ensemble classifier, Plausible drivers, Targeted clinical panel sequencing ","lastPublishedDoi":"10.21203/rs.3.rs-113748/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-113748/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003e\u003cstrong\u003eBackground\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eToday, there are a lot of markers on the prognosis and diagnosis of complex diseases such as primary breast cancer. However, our understanding of the drivers that influence cancer aggression is limited.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eMethods\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eIn this work, we study somatic mutation data consists of 450 metastatic breast tumor samples from cBio Cancer Genomics Portal. We use four software tools to extract features from this data. Then, an ensemble classifier (EC) learning algorithm called EARN (Ensemble of Artificial Neural Network, Random Forest, and non-linear Support Vector Machine) is proposed to evaluate plausible driver genes for metastatic breast cancer (MBCA). \u003c/p\u003e\u003cp\u003e\u003cstrong\u003eResults\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eThis study is an attempt to focus on the findings in several aspects of MBCA prognosis and diagnosis. First, drivers and passengers predicted by SVM, ANN, RF, and EARN are introduced. Second, biological inferences of predictions based on gene set enrichment analysis are discussed. Third, statistical validation and comparison of all learning methods based on evaluation metrics are done. Finally, the pathway enrichment analysis (PEA) using ReactomeFIVIz tool (\u003cem\u003eFDR\u003c/em\u003e\u0026lt;0.03) for the top 100 genes predicted by EARN leads us to propose a new gene set panel for MBCA, including HDAC3, ABAT, GRIN1, PLCB1, and KPNA2 as well as NCOR1, TBL1XR1, SIRT4, KRAS, CACNA1E, PRKCG, GPS2, SIN3A, ACTB, KDM6B, and PRMT1. Furthermore, we compare results for MBCA to other outputs regarding 983 primary tumor samples of breast invasive carcinoma (BRCA) obtained from the Cancer Genome Atlas (TCGA). The comparison between outputs shows that ROC-AUC reached 99.24% using EARN for MBCA and 99.79% for BRCA. This statistical result is better than three individual classifiers in each case.\u003c/p\u003e\u003cp\u003e\u003cstrong\u003eConclusions\u003c/strong\u003e\u003c/p\u003e\u003cp\u003eThis research using an integrative approach assists precision oncologists to design compact targeted panels that eliminate the need for whole-genome/exome sequencing.\u003c/p\u003e","manuscriptTitle":"EARN: an ensemble machine learning algorithm to predict driver genes in metastatic breast cancer","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2020-12-01 15:12:03","doi":"10.21203/rs.3.rs-113748/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0df137a8-1a0f-4622-ae4e-d031007f4e8e","owner":[],"postedDate":"December 1st, 2020","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":1225047,"name":"Molecular Genetics"},{"id":1225048,"name":"Molecular Biology"},{"id":1225049,"name":"Epigenetics \u0026 Genomics"}],"tags":[],"updatedAt":"2020-12-01T15:12:05+00:00","versionOfRecord":[],"versionCreatedAt":"2020-12-01 15:12:03","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-113748","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-113748","identity":"rs-113748","version":["v1"]},"buildId":"re_ckhLnmML6MCF96OHNJ","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

⚙ Ask this paper AI returns verbatim quotes from the full text · source: preprint-html ⓘ

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00