Measuring Performance of Ensemble of Ensembles (EoE) Models

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract To improve the performance of machine-learning algorithms irrelevant and redundant features must be removed. Accordingly, many different feature selection methods have been proposed to analyze the high-dimensional datasets and determine subsets of relevant features. Ensemble feature selection methods have been proposed to integrate the advantages of single-feature selection methods. The ensemble feature selection concept is based on aggregating the outputs of the different feature selection methods into a single output. Many studies have shown that feature subsets generated by ensemble methods improve the performance of machine learning models more than single feature selection methods. In this study, we extended the concept of ensemble feature selection to the concept of ensemble of ensembles of feature selection, where the outputs of different ensembles are aggregated into a single output. Through comprehensive experimental tests, we study the effect of the new concept on the performance of machine-learning models. The results show that the ensemble of ensembles does not guarantee the improvement of the performance of the learning models in comparison to individual ensembles. Thus, the complexity of the structure of an ensemble of ensembles makes it an unworthy approach for improving model performance.
Full text 99,067 characters · extracted from preprint-html · click to expand
Measuring Performance of Ensemble of Ensembles (EoE) Models | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Measuring Performance of Ensemble of Ensembles (EoE) Models Mazza Yousuf Attia¹, Noureldien A. Noureldien This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3192966/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract To improve the performance of machine-learning algorithms irrelevant and redundant features must be removed. Accordingly, many different feature selection methods have been proposed to analyze the high-dimensional datasets and determine subsets of relevant features. Ensemble feature selection methods have been proposed to integrate the advantages of single-feature selection methods. The ensemble feature selection concept is based on aggregating the outputs of the different feature selection methods into a single output. Many studies have shown that feature subsets generated by ensemble methods improve the performance of machine learning models more than single feature selection methods. In this study, we extended the concept of ensemble feature selection to the concept of ensemble of ensembles of feature selection, where the outputs of different ensembles are aggregated into a single output. Through comprehensive experimental tests, we study the effect of the new concept on the performance of machine-learning models. The results show that the ensemble of ensembles does not guarantee the improvement of the performance of the learning models in comparison to individual ensembles. Thus, the complexity of the structure of an ensemble of ensembles makes it an unworthy approach for improving model performance. Feature Selection Feature Selection Methods Ensemble of Feature Selection Ensemble of Ensembles Figures Figure 1 1. Introduction Feature selection refers to the selection of a subset of features from the original features in a data set [ 1 ]. Feature selection is an important activity in the data preprocessing stage of big datasets, and has been widely studied by machine learning researchers. Feature selection is an integral part of data mining or machine learning to remove unwanted and redundant information from data. By selecting important features, the dimension of the data is reduced, which helps lower the computational cost of the learning process. Feature selection improves the classification accuracy, model generalization, and speed of learning processing [ 2 ] [ 3 ] [ 4 ]. Several feature-selection methods have been proposed. The focus of a feature selection method is to select a subset of features from the original set of features that can efficiently describe the original set while reducing the effects of noise or irrelevant features while still providing good prediction results [ 5 ]. Feature selection methods can be divided into three types: filter, wrapper, and embedded methods, based on the relationship between the feature selection and inductive learning methods [ 6 ]. The filter methods operate directly on the dataset and provide a feature ranking or subset as the output [ 7 ]. Filter methods rely on the general characteristics of the training data and carry out the feature selection process as a pre-processing step, independent of the induction algorithm [ 8 ]. The advantages of filter methods are its low computational cost and good generalization ability. Filter methods include correlation, information gain, relief scores, and consistency-based filters [ 6 ]. The optimal feature subset depends on specific biases and heuristics of the induction algorithm. Based on this assumption, wrapper methods utilize a specific classifier to evaluate the quality of selected features [ 9 ]. The wrapper approach depends on different search strategies such as exponential, sequential, and randomized searches. Embedded methods perform feature selection in the training process and are usually specific to a given machine-learning algorithm. Embedded methods learn the features that best contribute to the accuracy of the learning model while the model is being created. This approach can capture dependencies at a lower computational cost than the wrappers [ 6 ]. Common embedded methods are recursive feature elimination and feature selection perceptron [ 8 ]. The feature selection ensemble aims to combine the outputs of multiple feature selection methods, thereby producing a more robust feature subset for subsequent classifier learning tasks [ 10 ]. In the last decade, ensemble learning has become a recognized approach based on the assumption that the combination of the output of several features selection methods yields better output than the output of any individual method [ 11 ]. Many research studies, such as [ 10 ][ 11 ][ 12 ][ 13 ][ 14 ], have shown that ensemble feature selection subsets improve the performance of the corresponding machine learning models. In this study, we extended the concept of an ensemble of feature selection to the concept of an ensemble of ensembles (EoEs) of feature selection. As the output of each ensemble is a subset of the original set of features, a further ensemble can be constructed from these subsets to obtain a new subset. This defines what we call an ensemble of ensembles for feature selection. The question now is whether the output subset of the features of this new construct improves the machine learning models. To answer this question, we conducted comprehensive experimental tests on the performance of machine learning models that use the output subset features of EoEs over different datasets. The investigation was limited to the ensembles of filter feature selection and the combination methods Min and Median. The rest of this paper is organized as follows; in section 2 the material and methods used to conduct the research are detailed, and in section 3 results and discussion are given. 2. Material and Methods To investigate the performance of feature subsets generated by an ensemble of ensembles on different machine learning models, we must: 1. Select several different data sets that are needed to build the learning models. 2. Select a set of filter feature selection methods and determine their generated subsets out of the original set of data sets features. 3. Select a combination or aggregation methods and use them to combine outputs of single filter selection methods in step B to construct ensembles of feature selection. 4. Determine the output subsets of ensembles in step 3. 5. Use the combination methods to aggregate outputs of ensembles in step D to construct an ensemble of ensembles. 6. Determine the output subsets of an ensemble of ensembles in step E. 7. Select different machine learning algorithms (classifiers) and use them to build learning models using subsets of features obtained by ensembles in step D and ensembles in step F. 8. Compare the performance of the different learning models. Figure (1) shows our research methodology steps. 2.1 Datasets To evaluate the effectiveness of the ensemble of ensemble feature subsets, we selected four datasets from the UC Irvine Machine Learning Repository [ 15 ] with a variant number of features and instances. Table (1) shows a summary of the characteristics of the dataset, and datasets archive URL. Table 1 Characteristics of the Datasets Used in Experiments # Dataset Name Number of Features Number of Instances Data Set Archive URL 1 Credit-g 20 1000 https://archive.ics.uci.edu/dataset/144/statlog+german+credit+data 2 Ozone-level-8hr 72 2534 : https://archive.ics.uci.edu/dataset/172/ozone+level+detection 3 Madelon 500 2600 https://archive.ics.uci.edu/dataset/171/Madelon 4 Micro-mass 1300 571 https://archive.ics.uci.edu/dataset/253/micromass 2.2 Selection of Filter Feature Selection (FFS) Methods Four filter feature selection methods are selected and constructed using search methods and attribute evaluators. Table (2) shows the selected filter feature selection methods that were used. Table 2 Filter Feature Selection Methods Used in the Experiments Feature Selection Method Search Method Attribute Evaluator Filter Ranker CorrelationAttributeEval InfoGainAttributeEval GainRatioAttributeEval OneRAtrributeEval 2.2.1 Generation of FFS Subsets for each Method Since the output of each FFS is a ranked list of features, we heuristically set the following rules to specify the number of features that we have to chose to be in our testing subsets according to the number of total features ( N ). This hypothesis was adapted from [ 2 ]. - If N < = 50, select the top 10 features. - If 50 < N < = 100, select the top 15 features. - If 100 < N 500, select the top 25 features. 2.3 Aggregation of Combination Methods Because we focus on filter feature selection methods, that returns an ordered ranking of all features according to their relevance, we need an aggregation or combination methods that can receive several rankings obtained by the different weak selectors as input and combine them into a single final ranking. Examples of such combination methods are simple operations between the ranks (Min, Median, Arithmetic Mean, and Geometric Mean), the Stuart Aggregation Method, Robust Rank Aggregation, and support vector machine rank [ 11 ]. The easiest way to combine the rankings of features is to apply simple operations such as the median or mean. Some popular methods can be defined as follows [ 14 ]. - Min: assigns each element to be ranked in the minimum (best) position that it has achieved among all rankings. - Median: assigning the median of all the positions that it has achieved among all rankings to each element to be ranked. - Arithmetic Mean: assigning the mean of all the positions that it has achieved among all rankings to each element to be ranked. - Geom. Mean: assigning the geometric mean of all the positions that it has achieved among all rankings to each element to be ranked. 2.4 Ensemble and Ensemble of Ensembles (EoE) Feature Subsets In this study, we used only filter feature selection methods to generate ensemble feature subset lists from different filter feature selection methods. The combination methods used to construct ensembles and ensemble of ensembles were Min and Median. 2.5 The Selection of Classifiers The performance of the feature selection methods subsets was evaluated using classification. Classification is the process of predicting an unknown property based on a feature subset that corresponds to the class labels. This task was performed in the WEKA environment using the 10 folds cross-validation setting for testing. In our experiments, we use nine machine-learning algorithms from different classes (Bayes, rules, and trees), as shown in Table (3), to evaluate the performance of the ensemble and EoE subsets. Table (3): Machine Learning Classifiers Classifier Class Classifier Tree Trees.J48 Trees.RandomForest Trees.RandomTree Rules Rules.DecisionTable Rules.JRip Rules.OneR Byes Bayes.BayesNet Bayes.NaiveBayes Bayes.NaiveBayes- MultinominalText 2.5 Examination of the Efficiency of Ensemble and EOE Machine Learning Models on Different Datasets. To evaluate the efficiency of machine learning models that use the ensemble and ensemble of ensemble feature subsets constructed with the two combination methods Min and Median, we selected the subsets according to the hypothesis stated in Section (5). The performance of each subset of the ensemble subsets was evaluated using cross-validation (fold 10) with each classifier. In the 10-fold cross-validation, the available data were randomly divided into 10 disjoint subsets of approximately equal size. One of the subsets was used as the test set, and the remaining nine sets were used for training the classifier. 2.5.1 Examining Ensemble and EoE Models Using Credit-G Dataset This dataset contains 20 attributes and 1000 instances. Table (4) shows the ranked lists generated by applying each filter separately, and the lists generated by the ensemble models. Table (4): Ranked Lists Generated by Filter and Ensemble Subsets Using Credit-g Dataset Filter Ranked List CorrelationAttributeEval 1,2,5,6,15,14,13,3,20,4,8,9,12,7,16,19,17,10,18,11 InfoGainAttributeEval 1,3,2,6,4,5,12,7,15,13,14,9,20,10,17,19,18,8,11,16 GainRatioAttributeEval 1,20,3,2,5,6,13,15,14,4,10,12,7,9,19,17,16,8,11,18 OneRAttributeEval 3,2,20,7,8,4,6,19,9,10,11,18,16,12,15,14,17,13,1,5 Ensemble Combination method Ranked List Ensemble1 Min 1,3,2,20,5,6,7,4,8,15,14,12,13,19,9,10,11,18,16,17 Ensemble2 Median 1,2,3,5,6,20,4,13,15,7,14,9,10,12,8,19,16,17,18,11 EOE Min 1,2,3,5,20,6,4,7,13,8,15,14,9,12,10,19,11,16,17,18 EOE Median 1,2,3,5,6,20,4,7,15,13,14,8,9,12,10,19,11,16,18,17. Table (5) shows the top 50% listed features for the ensembles and ensemble of ensembles and Table (6) shows the classification accuracy achieved using classifiers listed in table (3). Table (5): The Selected Subsets for an Ensemble and EOE for Credit-g Dataset Ensemble Model Selected Features Ensemble1 1,3,2,20,5,6,7,4,8,15 Ensemble2 1,2,3,5,6,20,4,13,15,7 EOE (min) 1,2,3,5,20,6,4,7,13,8 EOE (median) 1,2,3,5,6,20,4,7,15,13 Table (6): Performance of Ensemble and EOE Models Using Credit-g Dataset Classifier Class Classifier Accuracy of Ensemble 1 Accuracy of Ensemble 2 Accuracy of EoE(Min) Accuracy of EoE(Median) Tree Trees.J48 72.8 % 72.1 % 73.1 % 72.1 % Trees.RandomForest 76.1 % 77.1 % 77 % 77.1 % Trees.RandomTree 68 % 71.2 % 67.5 % 71.2 % Rules Rules.DecisionTable 70.7 % 70.6 % 70.6 % 70.6 % Rules.JRip 72.3 % 71 % 72.3 % 71 % Rules.OneR 66.1 % 66.1 % 66.1 % 66.1 % Byes Bayes.BayesNet 75.1 % 74.6 % 74.6 % 74.6 % Bayes.NaiveBayes 75.5 % 75.4 % 76 % 75.4 % Bayes.NaiveBayes- MultinominalText 70 % 70 % 70 % 70 % 2.52 Examining Ensemble and EOE Models Using Ozone-Level-8hr Dataset This dataset contains 72 attributes and 2534 instances. Table (7) shows the ranked lists generated by applying each filter separately, and lists generated by the ensemble models. Table (7): Ranked Lists Generated by Filter and Ensemble Subsets Using Ozone-level-8hr Dataset Filter Ranked List CorrelationAttributeEval 42,43,41,44,40,51,45,39,38,46,13,12,11,37,47,48,52,14,36,26,49,10,3,53,50,35,2,1,4,60,15,65,5,6,25,7,34,61,…. InfoGainAttributeEval 41,42,40,51,43,44,39,45,38,46,47,48,12,37,13,52,49,56,36,11,26,3,2,50,14,4,1,35,53,10,….. GainRatioAttributeEval 60,42,50,41,43,40,45,44,51,39,1,12,46,47,52,37,48,38,13,5,36,26,34,3,49,30,53,2,11,31,…. OneRAttributeEval 48,50,41,52,47,37,38,43,46,36,40,5,10,24,23,22,4,25,11,28,29,26,3,2,6,7,8,14,15,13,21,12,….. Ensemble Model Combination method Ranked List Ensemble1 Min 41,42,48,60,43,50,40,44,51,52,47,37,38,39,45,46,36,1,13,5,12,10,11,24,23,22,4,… Ensemble2 Median 42,41,40,43,44,45,51,38,39,46,12,47,50,37,48,52,13,11,36,26,3,49,2,5,10,14,1,… EOE (min) Min 72,2,25,10,23,24,20,21,48,46,44,8,27,50,22,26,71,51,56,58,5,19,52,3,6,68,4,12,7,14,…. EOE(median) Median 23,25,72,10,21,71,51,27,5,24,52,6,20,22,26,3,4,19,1,7,69,8,58,56,11,9,28,60,29,… Table (8) shows the top 15 features for each filter and ensemble. Table (8), shows the classification accuracy achieved by using classifiers belonging to the classes Trees, Rules, and Bayes. Table (8): The Selected Subsets for an Ensemble and EOE Subsets for Ozone-level-8hrDataset Ensemble Model Selected Features Ensemble1 41,42,48,60,43,50,40,44,51,52,47,37,38,39,45 Ensemble2 42,41,40,43,44,45,51,38,39,46,12,47,50,37,48 EOE (min) 72,2,25,10,23,24,20,21,48,46,44,8,27,50,22 EOE (median) 23,25,72,10,21,71,51,27,5,24,52,6,20,22,26 Table (9): Performance of Ensemble and EOE Subsets Using Ozone-level-8hr Dataset Classifier Class Classifier Accuracy of Ensemble 1 Accuracy of Ensemble 2 Accuracy of EoE(Min) Accuracy of EoE(Median) Tree Trees.J48 93.2912% 93.015 % 92.3441% 92.6204 % Trees.RandomForest 93.6069% 93.9227% 94.041% 93.6464% Trees.RandomTree 90.2131% 91.397 % 91.2391% 90.5683% Rules Rules.DecisionTable 93.6859% 93.8832% 93.7253% 93.4886% Rules.JRip 92.9755% 93.7648% 93.4491% 93.0545% Rules.OneR 93.4886% 93.4886% 93.4491% 93.6069% Byes Bayes.BayesNet 67.9163% 67.7979% 82.0047% 81.8074% Bayes.NaiveBayes 64.1279% 65.3512% 73.9937% 73.0466% Bayes.NaiveBayes- MultinominalText 93.6859% 93.6859% 93.6859% 93.6859% 5.2.3 Examining Ensemble and EoE Models Using Madelon Dataset This dataset contains 500 attributes and 2600 instances. Table (10) shows the ranked lists generated by applying each filter separately, and lists generated by the ensemble models. Table (10): Ranked Lists Generated by Filter and Ensemble Subsets Using Madelon Dataset Filter Ranked List CorrelationAttributeEval 476,242,337,65,339,129,106,473,49,443,379,385,494,454,463,136,324,375,410,438 InfoGainAttributeEval 476,242,339,473,443,337,65,106,129,5,375,189,410,160,165,161,158,166,164,163 GainRatioAttributeEval 473,476,242,106,443,339,189,337,375,65,129,5,410,167,165,161,158,166,164,160 OneRAttributeEval 476,242,473,458,22,106,86,396,104,384,159,180,339,90,177,363,116,119,8,335 Ensemble Combination method Ranked List Ensemble1 Min 143,242,339,476,277,337,473,65,106,129,443,236,192,82,49,434,454,379,485,494 Ensemble2 Median 242,476,339,129,65,337,106,443,473,49,454,494,379,244,5,165,205,177,180,164 EOE (min) Min 286,484,483,129,211,258,257,130,212,472,384,164,97,410,409,9,98,10,487,111 EOE (Median) Median 483,484,129,211,257,258,130,97,212,487,9,329,409,354,488,360,270,284,318,327 Table (11) shows the top 20 features for each filter and ensemble. Table (12) shows the classification accuracy achieved using classifiers belonging to the classes Trees, Rules, and Bayes. Table (11): The Selected Subsets for an Ensemble and EOE Subsets for Madelon Dataset Ensemble Model Selected Features Ensemble1 143,242,339,476,277,337,473,65,106,129,443,236,192,82,49,434,454,379,485,494 Ensemble2 242,476,339,129,65,337,106,443,473,49,454,494,379,244,5,165,205,177,180,164 EOE (min) 286,484,483,129,211,258,257,130,212,472,384,164,97,410,409,9,98,10,487,111 EOE (median) 483,484,129,211,257,258,130,97,212,487,9,329,409,354,488,360,270,284,318,327 Table (12): Performance of the different classifiers over the Ensemble Subsets for Madelon Dataset Classifier Class Classifier Accuracy of Ensemble 1 Accuracy of Ensemble 2 Accuracy of EOE(Min) Accuracy of EOE(Median) Tree Trees.J48 80.1154% 78.2692% 53.7179% 46.6667% Trees.RandomForest 87.8462% 86.0385% 54.6154% 52.6923% Trees.RandomTree 77.7692% 77.2692% 49.6154% 51.5385% Rules Rules.DecisionTable 61.7949% 61.1538% 50.641 % 51.5385% Rules.JRip 73.9744% 74.7436% 48.4615% 53.2051% Rules.OneR 58.9744% 58.9744% 48.3333% 48.2051% Byes Bayes.BayesNet 62.1923% 62.1923% 50.3846% 50.3846% Bayes.NaiveBayes 59.5385% 60.0385% 55.3846% 54.6154% Bayes.NaiveBayes- MultinominalText 50.2564% 50.2564% 50.2564% 50.2564% 5.2.4. Examining Ensemble and EOE Models Using Micro-Mass Dataset This dataset contains 1300 attributes and 571 instances. Table (13) shows the ranked lists generated by applying each filter separately, and the lists generated by the ensemble models. Table (13): Ranked Lists Generated by Filter and Ensemble Subsets Using Micro-mass Dataset Filter Ranked List CorrelationAttributeEval 128,1261,1002,579,159,471,712,729,22,1008,620,1129,1215,243,512,1250,1181,70,937,787,422,414,343,1269,651,61,.. InfoGainAttributeEval 1261,620,1215,729,579,128,22,838,512,1129,471,937,1002,121,26,1008,418,1250,430,761,651,1163,167,1269,695,446,.. GainRatioAttributeEval 1169,368,1181,408,787,712,11,328,243,81,1172,1129,1046,521,22,416,896,1261,409,189,1215,216,1150,994,343,512,793,… OneRAttributeEval 1023,620,729,85,128,545,1181,472,70,343,570,446,1002,414,416,167,512,422,980,538,468,1129,810,579,1046,712,418,937,…. Ensemble Model Combination method Ranked List Ensemble1 Min 128,1023,1169,1261,368,620,729,1002,1181,1215,85,408,579,159,787,471,545,712,11,22,328,472,838,70,243,512,81,343,1008,1129,570,1172,… Ensemble2 Median 128,620,729,1261,22,1129,1181,1002,579,512,712,1215,243,937,343,1008,416,1046,1250,167,418,70,570,1269,85,838,368,409,1150,538,651,422,61,787, … EOE (min) Min 128,620,1023,729,1169,1261,22,368,1129,1181,1002,579,512,1215,85,712,408,243,159,937,343,787,471,1008,416,545,1046,11,1250,167,328,418,70,… EOE (median) Median 128,620,1261,729,1002,1181,579,1215,22,712,368,85,512,1129,243,343,1169,1008,70,787,838,937,408,1046,570,11,328,416,167,471,1250,1023,418,472, … Table (14) shows the top 25 listed features for each filter and ensemble. Table (15) shows the classification accuracy achieved using classifiers belonging to the classes Trees, Rules, and Bayes. Table (14): The Selected Subsets for an Ensemble and EOE Subsets for Micro-mass Dataset Ensemble Model Selected Features Ensemble1 128,1023,1169,1261,368,620,729,1002,1181,1215,85,408,579,159,787,471,545,712,11,22,328,472,838,70,243 Ensemble2 128,620,729,1261,22,1129,1181,1002,579,512,712,1215,243,937,343,1008,416,1046,1250,167,418,70,570,1269,85 EOE (min) 128,620,1023,729,1169,1261,22,368,1129,1181,1002,579,512,1215,85,712,408,243,159,937,343,787,471,1008,416 EOE (median) 128,620,1261,729,1002,1181,579,1215,22,712,368,85,512,1129,243,343,1169,1008,70,787,838,937,408,1046,570 Table (15): Performance of Ensemble and EOE Subsets Using Micro-mass Dataset Classifier Class Classifier Accuracy of Ensemble 1 Accuracy of Ensemble 2 Accuracy of EoE(Min) Accuracy of EoE(Median) Tree Trees.J48 56.7426% 63.5727% 59.3695% 61.1208% Trees.RandomForest 67.2504% 67.6007% 67.4256% 66.3748% Trees.RandomTree 58.8441% 58.669 % 61.4711% 60.4203 % Rules Rules.DecisionTable 49.0368% 55.5166% 50.0876% 50.4378% Rules.JRip 54.2907% 56.3923% 56.2172% 57.4431% Rules.OneR 17.338 % 17.338 % 17.338 % 17.338% Byes Bayes.BayesNet 64.0981% 66.1996% 61.8214% 65.324% Bayes.NaiveBayes 60.0701% 59.8949% 56.7426% 61.8214% Bayes.NaiveBayes- MultinominalText 10.5079% 10.5079% 10.5079% 10.5079 3. Results and Discussion Tables (6, 9, 12, and 15) above, show that the ensemble of ensembles achieved higher accuracy in nine out of the 36 cases, namely, EOE (Min) in six cases and EOE(median) in three cases. Thus, EOEs improve the learning accuracy in 25% of all cases. Out of these nine cases, 4 of them are in the Ozone-level-8hr dataset. This dataset contains 72 attributes and 2534 instances. For this specific dataset, EOEs improved learning accuracy in four out of nine cases, that is, in 44% of cases. Table (16) shows the achievement of EOEs in each dataset. Table (16): Achievement of EOEs per the number of features in the datasets # Dataset Name Number of Features Number of Instances Higher Accuracy % Achieved by EoE out of all Cases 1 Credit-g 20 1000 2/9 = 22% 2 Ozone-level-8hr 72 2534 4/9 = 44% 3 Madelon 500 2600 1/9 = 11% 4 Micro-mass 1300 571 2/9 = 22% Thus, generally, EOEs do not guarantee an improvement in learning accuracy, and in its best case, it improves the accuracy with the probability of 0.44 when the number of features is around 72 features. Thus, generally, EOEs do not grantees improve of learning accuracy. Declarations Funding The research is funded by the authors. Conflict of Interest Not Applicable Data Availability Data is acquired from the UC Irvine Machine Learning Repository, https://archive.ics.uci.edu/ Code Availability Not applicable Authors' Contributions This research paper is an output of a student research work for the M.Sc. degree in Information Security, at university of science and information technology. The first author Mazza is the student who carried out all practical work (data collection, experiments, and results). The second author Noureldien is the professor who supervised the student and guides the research through all phases (idea, objective, methodology). Both authors commented and improve previous versions of the manuscript. Both authors read and approve this last version of the manuscript. References Jie Cai, Jiawei Luo, Shulin Wang, and Sheng Yang, "Feature selection in machine learning: A new perspective", Neurocomputing, vol. 300, pp. 70-79, March 2018. Verónica Bolón-Canedo, Noelia Sánchez-Maroño, and Amparo Alonso Betanzos, "A review of feature selection methods on synthetic data", Knowledge and Information Systems, A Coruña, Spain, vol. 34, no. 3, pp. 483-519, March 2012. Atsushi Kawamura and Basabi Chakraborty,” A New Filter Evaluation Function for Feature Subset Selection with Evolutionary Computation”, 9th International Conference on Awareness Science and Technology (iCAST) JAPAN, 2018. Shilan S. Hameed, Olutomilayo Olayemi Petinrin1, Abdirahman Osman Hashi1, and Faisal Saeed, "Filter-Wrapper Combination and Embedded Feature Selection for Gene Expression Data", Int. J. Advance Soft Compu Appl, March 2018. Girish Chandrashekar and Ferat Sahin, “A Survey on Feature Selection Methods”, Electrical and Microelectronic Engineering, vol. 40, no. 1, pp. 16-28, December 2013. M. Dash and H. Liu," Feature Selection for Classification", Intelligent Data Analysis, March 1997. Kenji Kira and Larry A. Rendell, "The Feature Selection Problem: Traditional Methods and a New Algorithm", Proceedings of Ninth National Conference on Artificial Intelligence, 129- 134, 1992. Haq, N. F., et al,” An ensemble framework of anomaly detection using hybridized feature selection approach (HFSA)”, SAI Intelligent Systems Conference (IntelliSys), IEEE, 2015. Narendra and Fukunaga, "A Branch and Bound Algorithm for Feature Subset Selection", IEEE Trans. Computer, vol. 26, no. 9, pp. 917-922, Sept. 197\ Koller, and Sahami, "Toward Optimal Feature Selection", Proceedings of International Conference on Machine Learning, 1996. Yvan Saeys1, Inaki Inza, and Pedro Larranaga," A Review of Feature Selection Techniques in Bioinformatics”, Oxford University, vol. 23, no. 19, pp. 2507-2517, August 2007. Nazrul Hoque, Mihir Singh, Dhruba K. Bhattacharyya, "EFS-MI: An Ensemble Feature Selection, Method for Classification", Complex & Intelligent Systems, vol. 4, no. 2, pp. 105-118,2018. “Shodhganga : A Reservoir of Indian Theses.” [Online]. Available: https://shodhganga.inflibnet.ac.in/bitstream/10603/137885/3/12%20chapter%203.pdf. [Accessed: 10-Mar-2022]. Bolón-Canedo, V. and A. Alonso-Betanzos, “Recent Advances in Ensembles for Feature Selection”, Springer, 2018. UC Irvine Machine earning Repository, https://archive.ics.uci.edu/ Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3192966","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":225365247,"identity":"a6e24bdd-44dd-47a0-b88c-8999541ef54b","order_by":0,"name":"Mazza Yousuf Attia¹","email":"","orcid":"","institution":"University of Science and Technology (UST)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Mazza","middleName":"Yousuf","lastName":"Attia¹","suffix":""},{"id":225365248,"identity":"360a4acb-ed8b-4cb9-9345-fb820bb7d254","order_by":1,"name":"Noureldien A. Noureldien","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7klEQVRIiWNgGAWjYBACgwMgsoCZh4GB+RhYhI2dgBYLsBYDkBa2NAaGBCDFTECLDVQLkOAxA2thIKjleI/Zhw8G1jL80j3fHnz8sU2ej5mB8cPHHNxazM6cMZ45wyCdR3LO2e2GMxJuG7YxMzBLztyGR8uNHGNmHoPDPAY3crdJ8yTcZgRqYWPmxaPFGKTlD1CL/Y2cZyAt9gS1GIK0MIBskchhA2lJJKzlzLFixh6gXyRupJlJzki7ndzGzNiM1y8Gx5s3M/yosLbnn5H8TOKDzW3b+e3NBz98xKMFG2BsIE39KBgFo2AUjAIMAABpO0r0ETQL6AAAAABJRU5ErkJggg==","orcid":"","institution":"University of Science and Technology (UST)","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Noureldien","middleName":"A.","lastName":"Noureldien","suffix":""}],"badges":[],"createdAt":"2023-07-21 19:29:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3192966/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3192966/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":41639356,"identity":"d20db82b-4893-4a38-92b5-249e1a063852","added_by":"auto","created_at":"2023-08-16 14:08:44","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":41099,"visible":true,"origin":"","legend":"\u003cp\u003eResearch Methodology Steps\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-3192966/v1/830d2db4281434f63466654f.png"},{"id":43404325,"identity":"2d464c7e-6346-4c9e-8b54-244749f601da","added_by":"auto","created_at":"2023-09-20 06:52:43","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1368820,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3192966/v1/f848f1bb-6939-43eb-ad1c-f7a7016884f4.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Measuring Performance of Ensemble of Ensembles (EoE) Models","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eFeature selection refers to the selection of a subset of features from the original features in a data set [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Feature selection is an important activity in the data preprocessing stage of big datasets, and has been widely studied by machine learning researchers. Feature selection is an integral part of data mining or machine learning to remove unwanted and redundant information from data. By selecting important features, the dimension of the data is reduced, which helps lower the computational cost of the learning process. Feature selection improves the classification accuracy, model generalization, and speed of learning processing [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e] [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e] [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eSeveral feature-selection methods have been proposed. The focus of a feature selection method is to select a subset of features from the original set of features that can efficiently describe the original set while reducing the effects of noise or irrelevant features while still providing good prediction results [\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFeature selection methods can be divided into three types: filter, wrapper, and embedded methods, based on the relationship between the feature selection and inductive learning methods [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe filter methods operate directly on the dataset and provide a feature ranking or subset as the output [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Filter methods rely on the general characteristics of the training data and carry out the feature selection process as a pre-processing step, independent of the induction algorithm [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e]. The advantages of filter methods are its low computational cost and good generalization ability. Filter methods include correlation, information gain, relief scores, and consistency-based filters [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe optimal feature subset depends on specific biases and heuristics of the induction algorithm. Based on this assumption, wrapper methods utilize a specific classifier to evaluate the quality of selected features [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. The wrapper approach depends on different search strategies such as exponential, sequential, and randomized searches.\u003c/p\u003e \u003cp\u003eEmbedded methods perform feature selection in the training process and are usually specific to a given machine-learning algorithm. Embedded methods learn the features that best contribute to the accuracy of the learning model while the model is being created. This approach can capture dependencies at a lower computational cost than the wrappers [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. Common embedded methods are recursive feature elimination and feature selection perceptron [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eThe feature selection ensemble aims to combine the outputs of multiple feature selection methods, thereby producing a more robust feature subset for subsequent classifier learning tasks [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]. In the last decade, ensemble learning has become a recognized approach based on the assumption that the combination of the output of several features selection methods yields better output than the output of any individual method [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMany research studies, such as [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e][\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e][\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e][\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e][\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e], have shown that ensemble feature selection subsets improve the performance of the corresponding machine learning models.\u003c/p\u003e \u003cp\u003eIn this study, we extended the concept of an ensemble of feature selection to the concept of an ensemble of ensembles (EoEs) of feature selection. As the output of each ensemble is a subset of the original set of features, a further ensemble can be constructed from these subsets to obtain a new subset. This defines what we call an ensemble of ensembles for feature selection. The question now is whether the output subset of the features of this new construct improves the machine learning models. To answer this question, we conducted comprehensive experimental tests on the performance of machine learning models that use the output subset features of EoEs over different datasets. The investigation was limited to the ensembles of filter feature selection and the combination methods Min and Median.\u003c/p\u003e \u003cp\u003eThe rest of this paper is organized as follows; in section \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e the material and methods used to conduct the research are detailed, and in section \u003cspan refid=\"Sec14\" class=\"InternalRef\"\u003e3\u003c/span\u003e results and discussion are given.\u003c/p\u003e"},{"header":"2. Material and Methods","content":"\u003cp\u003eTo investigate the performance of feature subsets generated by an ensemble of ensembles on different machine learning models, we must:\u003c/p\u003e\n\u003cp\u003e1. Select several different data sets that are needed to build the learning models.\u003c/p\u003e\n\u003cp\u003e2. Select a set of filter feature selection methods and determine their generated subsets out of the original set of data sets features. \u003cspan\u003e3. Select a combination or aggregation methods and use them to combine outputs of single filter selection methods in step B to construct ensembles of feature selection.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e4. Determine the output subsets of ensembles in step 3.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e5. Use the combination methods to aggregate outputs of ensembles in step D to construct an ensemble of ensembles.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e6. Determine the output subsets of an ensemble of ensembles in step E.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e7. Select different machine learning algorithms (classifiers) and use them to build learning models using subsets of features obtained by ensembles in step D and ensembles in step F.\u003cbr\u003e\u003c/span\u003e\u003cspan\u003e8. Compare the performance of the different learning models.\u003cbr\u003e\u003c/span\u003e\u003c/p\u003e\n\u003cp\u003eFigure (1) shows our research methodology steps.\u003c/p\u003e\n\u003cp\u003e2.1 Datasets\u003c/p\u003e\n\u003cp\u003eTo evaluate the effectiveness of the ensemble of ensemble feature subsets, we selected four datasets from the UC Irvine Machine Learning Repository [\u003cspan class=\"CitationRef\"\u003e15\u003c/span\u003e] with a variant number of features and instances. Table (1) shows a summary of the characteristics of the dataset, and datasets archive URL.\u003c/p\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eCharacteristics of the Datasets Used in Experiments\u003c/p\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003e#\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eDataset Name\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eNumber of Features\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eNumber of Instances\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eData Set Archive URL\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003e1\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eCredit-g\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e20\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e1000\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://archive.ics.uci.edu/dataset/144/statlog+german+credit+data\u003c/span\u003e\u003c/span\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e2\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eOzone-level-8hr\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e72\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e2534\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e: \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://archive.ics.uci.edu/dataset/172/ozone+level+detection\u003c/span\u003e\u003c/span\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e3\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMadelon\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e500\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e2600\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://archive.ics.uci.edu/dataset/171/Madelon\u003c/span\u003e\u003c/span\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e4\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMicro-mass\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e1300\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e571\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://archive.ics.uci.edu/dataset/253/micromass\u003c/span\u003e\u003c/span\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n \u003ch2\u003e2.2 Selection of Filter Feature Selection (FFS) Methods\u003c/h2\u003e\n \u003cp\u003eFour filter feature selection methods are selected and constructed using search methods and attribute evaluators. Table (2) shows the selected filter feature selection methods that were used.\u003c/p\u003e\n \u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eFilter Feature Selection Methods Used in the Experiments\u003c/p\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eFeature Selection Method\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eSearch Method\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAttribute Evaluator\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003eFilter\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRanker\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eCorrelationAttributeEval\u003cbr\u003eInfoGainAttributeEval\u003cbr\u003eGainRatioAttributeEval\u003cbr\u003eOneRAtrributeEval\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003ch2\u003e2.2.1 Generation of FFS Subsets for each Method\u003c/h2\u003e\n \u003cp\u003eSince the output of each FFS is a ranked list of features, we heuristically set the following rules to specify the number of features that we have to chose to be in our testing subsets according to the number of total features (\u003cem\u003eN\u003c/em\u003e). This hypothesis was adapted from [\u003cspan class=\"CitationRef\"\u003e2\u003c/span\u003e].\u003c/p\u003e- If \u003cem\u003eN\u0026thinsp;\u0026lt;\u0026thinsp;=\u003c/em\u003e\u0026thinsp;50, select the top 10 features.\u003cp\u003e- If 50\u0026thinsp;\u0026lt;\u0026thinsp;\u003cem\u003eN\u0026thinsp;\u0026lt;\u0026thinsp;=\u003c/em\u003e\u0026thinsp;100, select the top 15 features.\u003c/p\u003e\n \u003cp\u003e- If 100\u0026thinsp;\u003cem\u003e\u0026lt;\u0026thinsp;N\u0026thinsp;\u0026lt;\u0026thinsp;=\u003c/em\u003e\u0026thinsp;500, select the top 20 features.\u003c/p\u003e\n \u003cp\u003e- If N\u0026thinsp;\u003cem\u003e\u0026gt;\u003c/em\u003e\u0026thinsp;500, select the top 25 features.\u003c/p\u003e\n \u003ch2\u003e2.3 Aggregation of Combination Methods\u003c/h2\u003e\n \u003cp\u003eBecause we focus on filter feature selection methods, that returns an ordered ranking of all features according to their relevance, we need an aggregation or combination methods that can receive several rankings obtained by the different weak selectors as input and combine them into a single final ranking. Examples of such combination methods are simple operations between the ranks (Min, Median, Arithmetic Mean, and Geometric Mean), the Stuart Aggregation Method, Robust Rank Aggregation, and support vector machine rank [\u003cspan class=\"CitationRef\"\u003e11\u003c/span\u003e].\u003c/p\u003e\n \u003cp\u003eThe easiest way to combine the rankings of features is to apply simple operations such as the median or mean. Some popular methods can be defined as follows [\u003cspan class=\"CitationRef\"\u003e14\u003c/span\u003e].\u003c/p\u003e- Min: assigns each element to be ranked in the minimum (best) position that it has achieved among all rankings.\u003cp\u003e- Median: assigning the median of all the positions that it has achieved among all rankings to each element to be ranked.\u003c/p\u003e\n \u003cp\u003e- Arithmetic Mean: assigning the mean of all the positions that it has achieved among all rankings to each element to be ranked.\u003c/p\u003e\n \u003cp\u003e- Geom. Mean: assigning the geometric mean of all the positions that it has achieved among all rankings to each element to be ranked.\u003c/p\u003e2.4 Ensemble and Ensemble of Ensembles (EoE) Feature Subsets\u003cp\u003eIn this study, we used only filter feature selection methods to generate ensemble feature subset lists from different filter feature selection methods. The combination methods used to construct ensembles and ensemble of ensembles were Min and Median.\u003c/p\u003e\n \u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n \u003ch2\u003e2.5 The Selection of Classifiers\u003c/h2\u003e\n \u003cp\u003eThe performance of the feature selection methods subsets was evaluated using classification. Classification is the process of predicting an unknown property based on a feature subset that corresponds to the class labels. This task was performed in the WEKA environment using the 10 folds cross-validation setting for testing.\u003c/p\u003e\n \u003cp\u003eIn our experiments, we use nine machine-learning algorithms from different classes (Bayes, rules, and trees), as shown in Table (3), to evaluate the performance of the ensemble and EoE subsets.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (3): Machine Learning Classifiers\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Taba\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eClassifier Class\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eClassifier\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eTrees.J48\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomForest\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomTree\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eRules\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRules.DecisionTable\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.JRip\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.OneR\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eByes\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eBayes.BayesNet\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes- MultinominalText\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e2.5 Examination of the Efficiency of Ensemble and EOE Machine Learning Models on Different Datasets.\u003cp\u003eTo evaluate the efficiency of machine learning models that use the ensemble and ensemble of ensemble feature subsets constructed with the two combination methods Min and Median, we selected the subsets according to the hypothesis stated in Section (5).\u003c/p\u003e\n \u003cp\u003eThe performance of each subset of the ensemble subsets was evaluated using cross-validation (fold 10) with each classifier. In the 10-fold cross-validation, the available data were randomly divided into 10 disjoint subsets of approximately equal size. One of the subsets was used as the test set, and the remaining nine sets were used for training the classifier.\u003c/p\u003e\n \u003cdiv id=\"Sec10\" class=\"Section3\"\u003e\n \u003ch2\u003e2.5.1 Examining Ensemble and EoE Models Using Credit-G Dataset\u003c/h2\u003e\n \u003cp\u003eThis dataset contains 20 attributes and 1000 instances. Table (4) shows the ranked lists generated by applying each filter separately, and the lists generated by the ensemble models.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (4): Ranked Lists Generated by Filter and Ensemble Subsets Using Credit-g Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabb\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth colspan=\"2\"\u003eFilter\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eRanked List\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eCorrelationAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,2,5,6,15,14,13,3,20,4,8,9,12,7,16,19,17,10,18,11\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eInfoGainAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,3,2,6,4,5,12,7,15,13,14,9,20,10,17,19,18,8,11,16\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eGainRatioAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,20,3,2,5,6,13,15,14,4,10,12,7,9,19,17,16,8,11,18\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eOneRAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e3,2,20,7,8,4,6,19,9,10,11,18,16,12,15,14,17,13,1,5\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cstrong\u003eEnsemble\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003eCombination method\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003eRanked List\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,3,2,20,5,6,7,4,8,15,14,12,13,19,9,10,11,18,16,17\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,2,3,5,6,20,4,13,15,7,14,9,10,12,8,19,16,17,18,11\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,2,3,5,20,6,4,7,13,8,15,14,9,12,10,19,11,16,17,18\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1,2,3,5,6,20,4,7,15,13,14,8,9,12,10,19,11,16,18,17.\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003eTable (5) shows the top 50% listed features for the ensembles and ensemble of ensembles and Table (6) shows the classification accuracy achieved using classifiers listed in table (3).\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (5): The Selected Subsets for an Ensemble and EOE for Credit-g Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabc\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eEnsemble Model\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eSelected Features\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e1,3,2,20,5,6,7,4,8,15\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e1,2,3,5,6,20,4,13,15,7\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e1,2,3,5,20,6,4,7,13,8\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e1,2,3,5,6,20,4,7,15,13\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eTable (6): Performance of Ensemble and EOE Models Using Credit-g Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabd\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eClassifier Class\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eClassifier\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 1\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 2\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEoE(Min)\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEoE(Median)\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eTrees.J48\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e72.8 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e72.1 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e73.1 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e72.1 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomForest\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e76.1 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e77.1 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e77 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e77.1 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e68 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e71.2 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e67.5 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e71.2 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eRules\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRules.DecisionTable\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e70.7 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e70.6 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e70.6 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e70.6 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.JRip\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e72.3 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e71 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e72.3 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e71 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.OneR\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e66.1 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e66.1 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e66.1 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e66.1 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eByes\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eBayes.BayesNet\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e75.1 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e74.6 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e74.6 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e74.6 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e75.5 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e75.4 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e76 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e75.4 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes- MultinominalText\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003e70 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e70 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e70 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e70 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003e2.52 Examining Ensemble and EOE Models Using Ozone-Level-8hr Dataset\u003c/h2\u003e\n \u003cp\u003eThis dataset contains 72 attributes and 2534 instances. Table (7) shows the ranked lists generated by applying each filter separately, and lists generated by the ensemble models.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (7): Ranked Lists Generated by Filter and Ensemble Subsets Using Ozone-level-8hr Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabe\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth colspan=\"2\"\u003eFilter\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eRanked List\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eCorrelationAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e42,43,41,44,40,51,45,39,38,46,13,12,11,37,47,48,52,14,36,26,49,10,3,53,50,35,2,1,4,60,15,65,5,6,25,7,34,61,\u0026hellip;.\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eInfoGainAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e41,42,40,51,43,44,39,45,38,46,47,48,12,37,13,52,49,56,36,11,26,3,2,50,14,4,1,35,53,10,\u0026hellip;..\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eGainRatioAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e60,42,50,41,43,40,45,44,51,39,1,12,46,47,52,37,48,38,13,5,36,26,34,3,49,30,53,2,11,31,\u0026hellip;.\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eOneRAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e48,50,41,52,47,37,38,43,46,36,40,5,10,24,23,22,4,25,11,28,29,26,3,2,6,7,8,14,15,13,21,12,\u0026hellip;..\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cstrong\u003eEnsemble Model\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003eCombination method\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003eRanked List\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e41,42,48,60,43,50,40,44,51,52,47,37,38,39,45,46,36,1,13,5,12,10,11,24,23,22,4,\u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e42,41,40,43,44,45,51,38,39,46,12,47,50,37,48,52,13,11,36,26,3,49,2,5,10,14,1,\u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e72,2,25,10,23,24,20,21,48,46,44,8,27,50,22,26,71,51,56,58,5,19,52,3,6,68,4,12,7,14,\u0026hellip;.\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE(median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e23,25,72,10,21,71,51,27,5,24,52,6,20,22,26,3,4,19,1,7,69,8,58,56,11,9,28,60,29,\u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003eTable (8) shows the top 15 features for each filter and ensemble. Table (8), shows the classification accuracy achieved by using classifiers belonging to the classes Trees, Rules, and Bayes.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (8): The Selected Subsets for an Ensemble and EOE Subsets for Ozone-level-8hrDataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabf\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eEnsemble Model\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eSelected Features\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e41,42,48,60,43,50,40,44,51,52,47,37,38,39,45\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e42,41,40,43,44,45,51,38,39,46,12,47,50,37,48\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e72,2,25,10,23,24,20,21,48,46,44,8,27,50,22\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e23,25,72,10,21,71,51,27,5,24,52,6,20,22,26\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003e\u003cstrong\u003eTable (9): Performance of Ensemble and EOE Subsets Using Ozone-level-8hr Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabg\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eClassifier Class\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eClassifier\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 1\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 2\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEoE(Min)\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEoE(Median)\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eTrees.J48\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.2912%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.015 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e92.3441%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e92.6204 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomForest\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6069%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.9227%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e94.041%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6464%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e90.2131%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e91.397 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e91.2391%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e90.5683%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eRules\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRules.DecisionTable\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6859%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.8832%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.7253%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.4886%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.JRip\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e92.9755%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.7648%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.4491%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.0545%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.OneR\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.4886%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.4886%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.4491%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6069%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eByes\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eBayes.BayesNet\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e67.9163%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e67.7979%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e82.0047%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e81.8074%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e64.1279%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e65.3512%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e73.9937%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e73.0466%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes- MultinominalText\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6859%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6859%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6859%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e93.6859%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\n \u003ch2\u003e5.2.3 Examining Ensemble and EoE Models Using Madelon Dataset\u003c/h2\u003e\n \u003cp\u003eThis dataset contains 500 attributes and 2600 instances. Table (10) shows the ranked lists generated by applying each filter separately, and lists generated by the ensemble models.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (10): Ranked Lists Generated by Filter and Ensemble Subsets Using Madelon Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabh\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth colspan=\"2\"\u003eFilter\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eRanked List\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eCorrelationAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e476,242,337,65,339,129,106,473,49,443,379,385,494,454,463,136,324,375,410,438\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eInfoGainAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e476,242,339,473,443,337,65,106,129,5,375,189,410,160,165,161,158,166,164,163\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eGainRatioAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e473,476,242,106,443,339,189,337,375,65,129,5,410,167,165,161,158,166,164,160\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eOneRAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e476,242,473,458,22,106,86,396,104,384,159,180,339,90,177,363,116,119,8,335\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003e\u003cstrong\u003eEnsemble\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003eCombination method\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e\u003cstrong\u003eRanked List\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e143,242,339,476,277,337,473,65,106,129,443,236,192,82,49,434,454,379,485,494\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e242,476,339,129,65,337,106,443,473,49,454,494,379,244,5,165,205,177,180,164\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e286,484,483,129,211,258,257,130,212,472,384,164,97,410,409,9,98,10,487,111\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (Median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e483,484,129,211,257,258,130,97,212,487,9,329,409,354,488,360,270,284,318,327\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003cp\u003eTable (11) shows the top 20 features for each filter and ensemble. Table (12) shows the classification accuracy achieved using classifiers belonging to the classes Trees, Rules, and Bayes.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (11): The Selected Subsets for an Ensemble and EOE Subsets for Madelon Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003ctable id=\"Tabi\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eEnsemble Model\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eSelected Features\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e143,242,339,476,277,337,473,65,106,129,443,236,192,82,49,434,454,379,485,494\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e242,476,339,129,65,337,106,443,473,49,454,494,379,244,5,165,205,177,180,164\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e286,484,483,129,211,258,257,130,212,472,384,164,97,410,409,9,98,10,487,111\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e483,484,129,211,257,258,130,97,212,487,9,329,409,354,488,360,270,284,318,327\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\u003cstrong\u003e\u003cstrong\u003eTable (12): Performance of the different classifiers over the Ensemble Subsets for Madelon Dataset\u003ctable id=\"Tabj\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eClassifier Class\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eClassifier\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 1\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 2\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEOE(Min)\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEOE(Median)\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eTrees.J48\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e80.1154%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e78.2692%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e53.7179%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e46.6667%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomForest\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e87.8462%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e86.0385%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e54.6154%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e52.6923%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e77.7692%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e77.2692%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e49.6154%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e51.5385%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eRules\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRules.DecisionTable\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e61.7949%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e61.1538%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.641 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e51.5385%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.JRip\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e73.9744%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e74.7436%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e48.4615%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e53.2051%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.OneR\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e58.9744%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e58.9744%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e48.3333%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e48.2051%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eByes\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eBayes.BayesNet\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e62.1923%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e62.1923%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.3846%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.3846%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e59.5385%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e60.0385%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e55.3846%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e54.6154%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes- MultinominalText\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.2564%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.2564%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.2564%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.2564%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003ch2\u003e5.2.4. Examining Ensemble and EOE Models Using Micro-Mass Dataset\u003c/h2\u003e\n \u003cp\u003eThis dataset contains 1300 attributes and 571 instances. Table (13) shows the ranked lists generated by applying each filter separately, and the lists generated by the ensemble models.\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eTable (13): Ranked Lists Generated by Filter and Ensemble Subsets Using Micro-mass Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003cdiv class=\"colspec\"\u003e\u003cstrong\u003e\u003cbr\u003e\n \u003ctable id=\"Tabk\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth colspan=\"2\"\u003eFilter\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eRanked List\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eCorrelationAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e128,1261,1002,579,159,471,712,729,22,1008,620,1129,1215,243,512,1250,1181,70,937,787,422,414,343,1269,651,61,..\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eInfoGainAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1261,620,1215,729,579,128,22,838,512,1129,471,937,1002,121,26,1008,418,1250,430,761,651,1163,167,1269,695,446,..\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eGainRatioAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1169,368,1181,408,787,712,11,328,243,81,1172,1129,1046,521,22,416,896,1261,409,189,1215,216,1150,994,343,512,793,\u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd colspan=\"2\"\u003eOneRAttributeEval\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e1023,620,729,85,128,545,1181,472,70,343,570,446,1002,414,416,167,512,422,980,538,468,1129,810,579,1046,712,418,937,\u0026hellip;.\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble Model\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eCombination method\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRanked List\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e128,1023,1169,1261,368,620,729,1002,1181,1215,85,408,579,159,787,471,545,712,11,22,328,472,838,70,243,512,81,343,1008,1129,570,1172,\u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e128,620,729,1261,22,1129,1181,1002,579,512,712,1215,243,937,343,1008,416,1046,1250,167,418,70,570,1269,85,838,368,409,1150,538,651,422,61,787, \u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMin\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e128,620,1023,729,1169,1261,22,368,1129,1181,1002,579,512,1215,85,712,408,243,159,937,343,787,471,1008,416,545,1046,11,1250,167,328,418,70,\u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eMedian\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e128,620,1261,729,1002,1181,579,1215,22,712,368,85,512,1129,243,343,1169,1008,70,787,838,937,408,1046,570,11,328,416,167,471,1250,1023,418,472, \u0026hellip;\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003eTable (14) shows the top 25 listed features for each filter and ensemble. Table (15) shows the classification accuracy achieved using classifiers belonging to the classes Trees, Rules, and Bayes.\u003cp\u003e\u003cstrong\u003eTable (14): The Selected Subsets for an Ensemble and EOE Subsets for Micro-mass Dataset\u003c/strong\u003e\u003c/p\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\n \u003ctable id=\"Tabl\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eEnsemble Model\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eSelected Features\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble1\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e128,1023,1169,1261,368,620,729,1002,1181,1215,85,408,579,159,787,471,545,712,11,22,328,472,838,70,243\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEnsemble2\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e128,620,729,1261,22,1129,1181,1002,579,512,712,1215,243,937,343,1008,416,1046,1250,167,418,70,570,1269,85\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (min)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e128,620,1023,729,1169,1261,22,368,1129,1181,1002,579,512,1215,85,712,408,243,159,937,343,787,471,1008,416\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eEOE (median)\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e128,620,1261,729,1002,1181,579,1215,22,712,368,85,512,1129,243,343,1169,1008,70,787,838,937,408,1046,570\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\u003cstrong\u003eTable (15): Performance of Ensemble and EOE Subsets Using Micro-mass Dataset\u003c/strong\u003e\u003cbr\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\"\u003e\u003cbr\u003e\n \u003cdiv class=\"colspec\"\u003e\n \u003ctable id=\"Tabm\" border=\"1\"\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth\u003eClassifier Class\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eClassifier\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 1\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEnsemble 2\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEoE(Min)\u003cbr\u003e\u003c/th\u003e\n \u003cth\u003eAccuracy of\u003cbr\u003eEoE(Median)\u003cbr\u003e\u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eTrees.J48\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e56.7426%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e63.5727%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e59.3695%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e61.1208%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomForest\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e67.2504%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e67.6007%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e67.4256%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e66.3748%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eTrees.RandomTree\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e58.8441%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e58.669 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e61.4711%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e60.4203 %\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eRules\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eRules.DecisionTable\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e49.0368%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e55.5166%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.0876%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e50.4378%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.JRip\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e54.2907%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e56.3923%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e56.2172%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e57.4431%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eRules.OneR\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e17.338 %\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e17.338 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e17.338 %\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e17.338%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd rowspan=\"3\"\u003eByes\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003eBayes.BayesNet\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e64.0981%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e66.1996%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e61.8214%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e65.324%\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e60.0701%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e59.8949%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e56.7426%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e61.8214%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd\u003eBayes.NaiveBayes- MultinominalText\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e\u003cstrong\u003e10.5079%\u003c/strong\u003e\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e10.5079%\u003cbr\u003e\u003c/td\u003e\n \u003ctd align=\"char\"\u003e10.5079%\u003cbr\u003e\u003c/td\u003e\n \u003ctd\u003e10.5079\u003cbr\u003e\u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003cp\u003e\u003c/p\u003e\n \u003c/strong\u003e\u003c/div\u003e\n \u003cp\u003e\u003c/p\u003e\n \u003c/strong\u003e\u003c/strong\u003e\n \u003cp\u003e\u003c/p\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n \u003c/div\u003e\n\u003c/div\u003e"},{"header":"3. Results and Discussion","content":"\u003cp\u003eTables\u0026nbsp;(6, 9, 12, and 15) above, show that the ensemble of ensembles achieved higher accuracy in nine out of the 36 cases, namely, EOE (Min) in six cases and EOE(median) in three cases. Thus, EOEs improve the learning accuracy in 25% of all cases. Out of these nine cases, 4 of them are in the Ozone-level-8hr dataset. This dataset contains 72 attributes and 2534 instances. For this specific dataset, EOEs improved learning accuracy in four out of nine cases, that is, in 44% of cases. Table\u0026nbsp;(16) shows the achievement of EOEs in each dataset.\u003c/p\u003e \u003cp\u003e \u003cb\u003eTable\u0026nbsp;(16): Achievement of EOEs per the number of features in the datasets\u003c/b\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabn\" border=\"1\"\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003e#\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDataset Name\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNumber of Features\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNumber of Instances\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eHigher Accuracy % Achieved by EoE out of all Cases\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCredit-g\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2/9\u0026thinsp;=\u0026thinsp;22%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOzone-level-8hr\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2534\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e4/9\u0026thinsp;=\u0026thinsp;44%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMadelon\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e500\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2600\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1/9\u0026thinsp;=\u0026thinsp;11%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMicro-mass\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1300\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e571\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e2/9\u0026thinsp;=\u0026thinsp;22%\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThus, generally, EOEs do not guarantee an improvement in learning accuracy, and in its best case, it improves the accuracy with the probability of 0.44 when the number of features is around 72 features. Thus, generally, EOEs do not grantees improve of learning accuracy.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe research is funded by the authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of Interest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot Applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData Availability\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eData is acquired from the UC Irvine Machine Learning Repository, https://archive.ics.uci.edu/\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e\u0026nbsp;\u003c/strong\u003e\u003cstrong\u003eCode Availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthors\u0026apos; Contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research paper is an output of a student research work for the M.Sc. degree in Information Security, at university of science and information technology. The first author Mazza is the student who carried out all practical work (data collection, experiments, and results). The second author Noureldien is the professor who supervised the student and guides the research through all phases (idea, objective, methodology). Both authors commented and improve previous versions of the manuscript. Both authors read and approve this last version of the manuscript.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eJie Cai, Jiawei Luo, Shulin Wang, and Sheng Yang, \u0026quot;Feature selection in machine learning: A new perspective\u0026quot;, Neurocomputing, vol. 300, pp. 70-79, March 2018.\u003c/li\u003e\n\u003cli\u003eVer\u0026oacute;nica Bol\u0026oacute;n-Canedo, Noelia S\u0026aacute;nchez-Maro\u0026ntilde;o, and Amparo Alonso Betanzos, \u0026quot;A review of feature selection methods on synthetic data\u0026quot;, Knowledge and Information Systems, A Coru\u0026ntilde;a, Spain, vol. 34, no. 3, pp. 483-519, March 2012.\u003c/li\u003e\n\u003cli\u003eAtsushi Kawamura and Basabi Chakraborty,\u0026rdquo; A New Filter Evaluation Function for Feature Subset Selection with Evolutionary Computation\u0026rdquo;, 9th International Conference on Awareness Science and Technology (iCAST) JAPAN, 2018.\u003c/li\u003e\n\u003cli\u003eShilan S. Hameed, Olutomilayo Olayemi Petinrin1, Abdirahman Osman Hashi1, and Faisal Saeed, \u0026quot;Filter-Wrapper Combination and Embedded Feature Selection for Gene Expression Data\u0026quot;, Int. J. Advance Soft Compu Appl, March 2018.\u003c/li\u003e\n\u003cli\u003eGirish Chandrashekar and Ferat Sahin, \u0026ldquo;A Survey on Feature Selection Methods\u0026rdquo;, Electrical and Microelectronic Engineering, vol. 40, no. 1, pp. 16-28, December 2013.\u003c/li\u003e\n\u003cli\u003eM. Dash and H. Liu,\u0026quot; Feature Selection for Classification\u0026quot;, Intelligent Data Analysis, March 1997.\u003c/li\u003e\n\u003cli\u003eKenji Kira and Larry A. Rendell, \u0026quot;The Feature Selection Problem: Traditional Methods and a New Algorithm\u0026quot;, Proceedings of Ninth National Conference on Artificial Intelligence, 129- 134, 1992.\u003c/li\u003e\n\u003cli\u003eHaq, N. F., et al,\u0026rdquo; An ensemble framework of anomaly detection using hybridized feature selection approach (HFSA)\u0026rdquo;, SAI Intelligent Systems Conference (IntelliSys), IEEE, 2015. \u003c/li\u003e\n\u003cli\u003eNarendra and Fukunaga, \u0026quot;A Branch and Bound Algorithm for Feature Subset Selection\u0026quot;, IEEE Trans. Computer, vol. 26, no. 9, pp. 917-922, Sept. 197\\\u003c/li\u003e\n\u003cli\u003eKoller, and Sahami, \u0026quot;Toward Optimal Feature Selection\u0026quot;, Proceedings of International Conference on Machine Learning, 1996.\u003c/li\u003e\n\u003cli\u003eYvan Saeys1, Inaki Inza, and Pedro Larranaga,\u0026quot; A Review of Feature Selection Techniques in Bioinformatics\u0026rdquo;, Oxford University, vol. 23, no. 19, pp. 2507-2517, August 2007.\u003c/li\u003e\n\u003cli\u003eNazrul Hoque, Mihir Singh, Dhruba K. Bhattacharyya, \u0026quot;EFS-MI: An Ensemble Feature Selection, Method for Classification\u0026quot;, Complex \u0026amp; Intelligent Systems, vol. 4, no. 2, pp. 105-118,2018.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Shodhganga : A Reservoir of Indian Theses.\u0026rdquo; [Online]. Available: https://shodhganga.inflibnet.ac.in/bitstream/10603/137885/3/12%20chapter%203.pdf. [Accessed: 10-Mar-2022].\u003c/li\u003e\n\u003cli\u003eBol\u0026oacute;n-Canedo, V. and A. Alonso-Betanzos, \u0026ldquo;Recent Advances in Ensembles for Feature Selection\u0026rdquo;, Springer, 2018.\u003c/li\u003e\n\u003cli\u003eUC Irvine Machine earning Repository, https://archive.ics.uci.edu/\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Feature Selection, Feature Selection Methods, Ensemble of Feature Selection, Ensemble of Ensembles","lastPublishedDoi":"10.21203/rs.3.rs-3192966/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3192966/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eTo improve the performance of machine-learning algorithms irrelevant and redundant features must be removed. Accordingly, many different feature selection methods have been proposed to analyze the high-dimensional datasets and determine subsets of relevant features. Ensemble feature selection methods have been proposed to integrate the advantages of single-feature selection methods. The ensemble feature selection concept is based on aggregating the outputs of the different feature selection methods into a single output. Many studies have shown that feature subsets generated by ensemble methods improve the performance of machine learning models more than single feature selection methods.\u003c/p\u003e \u003cp\u003eIn this study, we extended the concept of ensemble feature selection to the concept of ensemble of ensembles of feature selection, where the outputs of different ensembles are aggregated into a single output. Through comprehensive experimental tests, we study the effect of the new concept on the performance of machine-learning models.\u003c/p\u003e \u003cp\u003eThe results show that the ensemble of ensembles does not guarantee the improvement of the performance of the learning models in comparison to individual ensembles. Thus, the complexity of the structure of an ensemble of ensembles makes it an unworthy approach for improving model performance.\u003c/p\u003e","manuscriptTitle":"Measuring Performance of Ensemble of Ensembles (EoE) Models","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-08-16 14:08:39","doi":"10.21203/rs.3.rs-3192966/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"a620591c-a95a-4b98-88b4-5bd3e2a2d0e1","owner":[],"postedDate":"August 16th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2023-09-20T06:44:35+00:00","versionOfRecord":[],"versionCreatedAt":"2023-08-16 14:08:39","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3192966","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3192966","identity":"rs-3192966","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. The paper's references may be in our DB but unresolved to ``paper_id`` (resolution happens at ingest when the cited DOI matches a row we already have). Run the cross-source citation reconcile pass to retry.

Source provenance

europepmc
last seen: 2026-05-19T01:45:01.086888+00:00