An effective heuristic for developing hybrid feature selection in high dimensional and low sample size datasets

preprint OA: closed
Full text JSON View at publisher
AI-generated summary by claude@2026-07, 2026-07-17

This study introduces a hybrid feature selection method combining gradual permutation filtering and heuristic tribrid search, which effectively identifies informative features and improves prediction models in high-dimensional, low-sample-size datasets.

One-sentence paraphrase of the abstract; not a substitute for reading it. No clinical advice. How this works

AI-generated deep summary by claude@2026-07, 2026-07-17 · read from full text

The paper studies how to perform effective feature selection for high-dimensional, low-sample-size (HDLSS) biological datasets, motivated by tasks like gene/microarray classification. Using a two-phase methodology, the authors combine gradual permutation filtering (which iteratively ranks and removes low permutation-importance features while recalculating importance to reduce bias) with a heuristic tribrid search strategy that integrates forward search, “consolation matches,” and backward elimination while leveraging pre-ranked features to shrink runtime and avoid local optima. Across benchmark comparisons, the proposed method reduced the average number of selected features from 37.8 to 5.5 and improved prediction performance from 0.855 to 0.927, and it introduces an HDLSS-focused metric that jointly evaluates the number and quality of selected features. A key caveat is that the study frames its approach as a preprint and reports results primarily via benchmark datasets rather than disease-specific validation. This paper does not explicitly discuss endometriosis or adenomyosis; it was included in the corpus via a keyword match in the upstream search index.

Read from the paper's body, not the abstract. Not a substitute for reading the paper. No clinical advice. How this works

Abstract

Abstract Background. High-dimensional datasets with low sample sizes (HDLSS) are pivotal in the fields of biology and bioinformatics. One of core objective of HDLSS is to select most informative features and discarding redundant or irrelevant features. This is particularly crucial in bioinformatics, where accurate feature (gene) selection can lead to breakthroughs in drug development and provide insights into disease diagnostics. Despite its importance, identifying optimal features is still a significant challenge in HDLSS. Results. To address this challenge, we propose an effective feature selection method that combines gradual permutation filtering with a heuristic tribrid search strategy, specifically tailored for HDLSS contexts. The proposed method considers inter-feature interactions and leverages feature rankings during the search process. In addition, a new performance metric for the HDLSS that evaluates both the number and quality of selected features is suggested. Through the comparison of the benchmark dataset with existing methods, the proposed method reduced the average number of selected features from 37.8 to 5.5 and improved the performance of the prediction model, based on the selected features, from 0.855 to 0.927. Conclusions. The proposed method effectively selects a small number of important features and achieves high prediction performance.
Full text 168,841 characters · extracted from preprint-html · click to expand
An effective heuristic for developing hybrid feature selection in high dimensional and low sample size datasets | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article An effective heuristic for developing hybrid feature selection in high dimensional and low sample size datasets Hyunseok Shin, Sejong Oh This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-5260669/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 26 Dec, 2024 Read the published version in BMC Bioinformatics → Version 1 posted 12 You are reading this latest preprint version Abstract Background. High-dimensional datasets with low sample sizes (HDLSS) are pivotal in the fields of biology and bioinformatics. One of core objective of HDLSS is to select most informative features and discarding redundant or irrelevant features. This is particularly crucial in bioinformatics, where accurate feature (gene) selection can lead to breakthroughs in drug development and provide insights into disease diagnostics. Despite its importance, identifying optimal features is still a significant challenge in HDLSS. Results. To address this challenge, we propose an effective feature selection method that combines gradual permutation filtering with a heuristic tribrid search strategy, specifically tailored for HDLSS contexts. The proposed method considers inter-feature interactions and leverages feature rankings during the search process. In addition, a new performance metric for the HDLSS that evaluates both the number and quality of selected features is suggested. Through the comparison of the benchmark dataset with existing methods, the proposed method reduced the average number of selected features from 37.8 to 5.5 and improved the performance of the prediction model, based on the selected features, from 0.855 to 0.927. Conclusions. The proposed method effectively selects a small number of important features and achieves high prediction performance. HDLSS feature selection machine learning filter method wrapper method Figures Figure 1 Figure 2 Figure 3 BACKGROUND In the current data era, information is recorded intricately across various dimensions. Consequently, the dimensionality of the data increases rapidly. High-dimensional datasets refer to data where the number of features ( p ) is considerably larger than the sample size ( n ). High-dimensional datasets have driven significant scientific discoveries in fields such as biology and bioinformatics by revealing previously unknown complex patterns. However, analyzing and utilizing high dimensional and low sample size (HDLSS) datasets remain a challenge for researchers [ 1 – 6 ]. Microarray data are prime examples of HDLSS datasets. With the capacity to capture the expression of tens of thousands of genes, the core aim is to identify genetic variations in specific cancers or other genetically-related diseases [ 3 ]. However, a notable proportion of these genes either lacks direct relevance to the disease or exhibits substantial overlap in feature information, leading to redundancy [ 7 ]. Furthermore, owing to the patient-centric nature of these data, sample size was inherently limited. These aspects amplify risks such as overfitting and the curse of dimensionality [ 3 , 4 ]. Consequently, developing cancer classification models and analyzing HDLSS datasets such as microarrays, pose intricate challenges for computer science researchers [ 7 ]. A fundamental approach for addressing these challenges in machine learning is dimensionality reduction via feature selection. The primary objective of feature selection is to filter out irrelevant and redundant features, thereby honing the most pertinent features relevant to the subject of interest. This enhances the clarity, generalizability, and predictive accuracy of classification models, concurrently reducing computational demands and boosting efficiency [ 1 ]. In bioinformatics, specifically, feature selection plays a critical role in uncovering potential avenues for drug development and offering insights into disease diagnostics and causation [ 1 – 3 , 8 ]. Identifying the optimal features is an NP-hard problem [ 3 , 6 , 8 , 9 ]. Numerous feature-selection methods have been proposed to address this challenge. These methods are typically classified into three categories based on their interactions with the classification model: filter, wrapper, and embedded methods [ 3 , 4 ]. The filter method takes a model-independent stance, assigns rankings to features based on metrics such as statistical or information-theoretic properties, and subsequently selects those that exceed a predetermined threshold. In contrast, the wrapper method incorporates a model-dependent approach, by selecting feature subsets based on the performance of a specific model. The embedded method differs from the wrapper method by integrating feature selection directly into the model-training process. Recent advancements in feature selection include methods based on graphs and deep learning frameworks. These methods aim to capture the intricate relationships among features. In the graph-based approach, features are visualized as nodes and their interrelationships as edges. This method determines the importance of features by analyzing the structural properties of a graph [ 6 , 10 , 11 , 12 ]. Deep learning techniques for feature selection utilize neural networks with features as inputs [ 6 – 9 , 12 – 14 ]. The significance of each feature is gauged by the magnitude of the gradients during network training. For HDLSS data, several studies on feature selection have adopted a hybrid approach involving two stages, filtering and searching, as illustrated in Fig. 1 [ 7 ]. This approach seeks to harness the computational efficiency of filter methods and the high performance of wrapper methods. The filter method can be viewed as a preprocessing step that significantly reduces the search scope by eliminating unnecessary features. By contrast, the wrapper method refines the search within this narrow scope, aiming to identify the most optimal features more precisely [ 3 , 7 ]. From the collective insights of previous studies on feature selection, we can draw the following conclusions: Understanding intricate relationships, such as the interactions between features, is pivotal. Finding the optimal features for HDLSS is realistically difficult. An alternative is to find a better solution among several suboptimal features. Therefore, it is necessary to avoid the local optima. This can be achieved by balancing the diversity and focus in the feature search space, introducing randomness during the search process, and employing a hybrid method that leverages both filter and wrapper techniques. In this study, we propose a new metaheuristic method [ 15 – 19 ] to improve upon previous feature selection approaches for HDLSS datasets, aiming to achieve near-optimal performance with minimal features. By integrating gradual permutation filtering with a diverse search strategy, we designed our approach based on core principles. The highlights of our method include. The details are outlined in the next section. Filtering: Leveraging permutation importance, filtering accounts for feature interactions and establishes a correct threshold. Search: Through “consolation matches,” the nesting effect can be overcome, wherein a once-selected (or excluded) feature remains unaltered. Additionally, by varying “first-choice features,” the search range is expanded to approach a more global optimum. Runtime optimization: Using ranked features, the method efficiently reduces the search space. Assessing the fitness of the selected features: We crafted a unique performance metric for the HDLSS, providing a holistic perspective on feature quantity and quality. METHODS Overall procedure The proposed method also follows the general procedure of feature selection for HDLSS datasets, as shown in Fig. 1 . As illustrated in Fig. 2 and elaborated in Algorithm 1, the proposed method comprises two primary phases: filtering and searching. Gradual Permutation Filtering (GPF): This phase ingests all the HDLSS data features. It ranks features based on their permutation importance and subsequently eliminates irrelevant features (Algorithm 2). Heuristic Tribrid Search (HTS): By leveraging the features ranked by the GPF, HTS employs a heuristic blend of forward search, consolation match, and backward elimination. This approach aims to identify the near-optimal feature set by considering both the number of features and classification performance using the features (see Algorithms 3–6). The subsequent sections will delve into the specifics of each phase. Algorithm 1. Overall procedure Input : WF – whole features CL - class label vector Output : SFL - selected feature list // Gradual Permutation Filtering FILTERED ← WF FOR RATIO IN [0.1, 0.2, ..., 0.9] DO RANKED ← gradual_permutation_filtering(FILTERED, CL, RATIO) // Algorithm 2 FILTERED ← FILTERED [RANKED] END FOR // Heuristic Tribrid Search: This section can be processed in parallel. N ← Length of WF FOR i FROM 1 TO 10 SELECTED[i], max_fit[i] = heuristic_tribird_search(FILTERED, CL, i, N) // Algorithm 3 END FOR SFL ← Choose the SELECTED with the highest max_fit from the 10 results RETURN SFL Gradual permutation filtering The GPF stage prioritizes features based on their permutation importance and eliminates unimportant features. This method chooses features with an importance value greater than zero because values below this threshold often indicate redundancy or suggest potential noise. This method offers a refined approach for evaluating feature importance by adopting a gradual process of eliminating noise and recalculating the importance, thereby minimizing the potential biases associated with single-step elimination. Considering that the permutation importance is evaluated within the evaluation group (features), the approach emphasizes filtering out irrelevant features and reevaluating the importance of enhancing purity, which is defined solely based on performance-impacting features. Notably, a gradual feature filtering method, as opposed to filtering all at once, provides a more precise selection of features with importance on the verge of 0. To ensure robust feature selection, the GPF measures the permutation importance of each feature multiple times; i.e., 50 times in our case. From these measurements, the features that exceed the importance of zero for a specific threshold number were selected. For the selected features, the permutation importance was recalculated by applying a progressively higher threshold each time. This iterative process, detailed in Algorithm 2, refines the ranking of significant features. The final ranking of the features was determined by averaging the last measured importance values of the features selected until the end. Algorithm 2. gradual_permutation_filtering(FILTERED, CL, RATIO) Input : FILTERED - filtered features CL - class label vector RATIO - threshold ratio Output : RANKED - ranked feature list CONST M ← 50 THRESHOLD ← M * RATIO PI ← matrix of dimensions M by length of FILTERED FOR i FROM 1 TO M PI[i] ← permutation importance based on FILTERED, varying the seed each time END FOR CNT ← Count of positive importance values in each column of PI MPI ← the mean of PI across columns, resulting in a vector of the same length as FILTERED MASK ← Set to 1 for entries in CNT that are ≥ THRESHOLD, otherwise set to 0 RANKED ← Selected features where the elementwise product of MPI and MASK is greater than 0, then sort them in descending order based on their importance RETURN RANKED Heuristic tribrid search While the GPF is designed to filter out unimportant features, HTS focuses on identifying the most informative features from a pool refined by the GPF. HTS operates in three distinct stages, forward search , consolation match , and backward elimination , as outlined in Algorithm 3. We slightly modified the original forward search technique. The original technique began with an empty feature set and determined its first feature. In contrast, the proposed method uses “first-choice feature” that is chosen from ranked feature list of GPF. This procedure incrementally adds to the feature set, guided by the performance increments detailed in Algorithm 4. The performance was measured using a classification metric based on the selected features. When the performance increase caused by the selected features stops during the forward search, the process shifts to “consolation match.” This stage aims to enhance the performance by swapping a single feature between selected and unselected feature pools, as depicted in Algorithm 5. This strategy offers an escape route from potential local optima. Notably, the duration of this stage depends on the volume of features initially filtered by the GPF. Techniques for expediting this stage are explored in subsequent sections. If the consolation match yields improved performance, the forward search resumes further enhancements. However, if no such gains are evident, HTS transitions to a backward elimination phase. This final stage discards any remaining unimportant features, thereby ensuring the relevance of the end feature set. Importantly, at each stage of the HTS (Algorithms 4–6), the performance is gauged using a novel metric introduced in this study. This metric considers both the classification performance of the model and feature count. An elaboration of this metric is presented in the following section. Algorithm 3. heuristic_tribird_search(FILTERED, CL, i, N) Input : FILTERED – selected features sorted by ranking CL - class label vector i – index number N - Number of columns in original dataset Output : SELECTED - selected feature list max_fit - maximum fit value SELECTED ← FILTERED[i] CANDIDATE ← FILTERED – {SELECTED} max_fit ← 0 max_update ← TRUE WHILE max_update DO max_update ← FALSE SELECTED, max_update, max_fit ← forward_search(FILTERED, CL, SELECTED, max_fit, N) // Algorithm 4 IF NOT max_update THEN SELECTED, max_update, max_fit ← consolation_match(FILTERED CL, SELECTED, max_fit, N) // Algorithm 5 END IF END WHILE SELECTED ← Features obtained from SELECTED through backward elimination with stopping criteria based on Algorithm 6 RETURN SELECTED, max_fit Algorithm 4. forward_search(FILTERED CL, SELECTED, max_fit, N) Input : FILTERED – selected features sorted by ranking CL - class label vector SELECTED - selected feature list max_fit - maximum fit value N - Number of columns in original dataset Output : SELECTED - selected feature list max_update - max fit update status max_fit - maximum fit value DO max_update ← FALSE CANDIDATE ← {columns of FILTERED } – SELECTED FOR each feature F IN CANDIDATE DO TEMP ← SELECTED ∪ {F} current_fit ← fitness(CL, TEMP, N) IF current_fit > max_fit THEN max_fs ← TEMP max_fit ← current_fit max_update ← TRUE END IF END FOR IF max_update THEN SELECTED ← max_fs END IF WHILE max_update RETURN SELECTED, max_update, max_fit Algorithm 5. consolation_match(FILTERED, CL, SELECTED, max_fit, N) Input : FILTERED – selected features sorted by ranking CL - class label vector SELECTED - selected feature list max_fit - maximum fit value N - Number of columns in original dataset Output : max_fs - selected feature list with max fit value max_update - max fit update status max_fit - maximum fit value max_update ← FALSE CANDIDATE ← {columns of FILTERED } – SELECTED // Iterate over each feature in SELECTED and CANDIDATE to find an optimal swap FOR X IN SELECTED DO FOR Y IN CANDIDATE DO TEMP ← (SELECTED - {X}) U {Y} current_fit ← fitness(CL, TEMP, N) IF current_fit > max_fit THEN max_fs ← TEMP max_fit ← current_fit max_update ← TRUE END IF END FOR END FOR RETURN max_fs, max_update, max_fit Log comprehensive metric for HDLSS datasets Balancing the number of selected features with the classification performance is imperative when working with HDLSS datasets. Therefore, this study proposes log-comprehensive metric (LCM). This metric, which is a refined iteration of the conventional fitness function, was meticulously designed for the HDLSS datasets, as elaborated in Algorithm 6. Algorithm 6. log_comprehensive_metric(CL, SELECTED, N) Input : CL - class label vector SELECTED - selected feature list N - Number of columns in original dataset Output : LCM - log comprehensive metric value // Calculate the LCR metric using a weighted average of the LRR (Log Reduction Rate) derived from performance and selected feature count. CONST C ← 0.005 // C is trade-off constant for feature count and performance. X ← length of SELECTED// Number of selected features Y ← log(2)/log(N) THETA ← Y/(C + Y) PERFORMANCE ← the average performance obtained with 5-fold and XGBoost on SELECTED and CL LRR ← 1-log(X)/log(N) LCM ← THETA*PERFORMANCE + (1-THETA)*LRR RETURN LCM The widely used fitness function is described by Eq. ( 1 ). It constitutes a weighted sum of the model error and the ratio of the selected features to the total. The aim is to minimize the following function: $$\:fitness=\theta\:*Error+\left(1-\theta\:\right)*\left(\frac{Number\:of\:selected\:features}{Total\:EquationNumber\:of\:features}\right),\:\:\:where\:0\le\:\theta\:\le\:1$$ 1 By contrast, the LCM, expressed in Eq. ( 2 ), fuses the model’s classification success with the log reduction rate (LRR) using weighted factors. The objective is to maximize. $$\:LCM=\theta\:*Performance+\left(1-\theta\:\right)*LRR,\:\:where\:\:0\le\:\theta\:\le\:1$$ 2 LRR is further explained by Eq. ( 3 ). Its value fluctuates between 0 and 1 and approaches 1 as the number of selected features decreases. $$\:LRR=1-{\text{log}}_{a}b=1-\text{log}b/\text{log}a,$$ $$\:where:0\le\:LRR\le\:1\:,\:\:$$ $$\:a=Total\:EquationNumber\:of\:features,\:b=\:Number\:of\:selected\:features\:$$ 3 In contrast to the reduction rate (RR), which is formulated as RR = 1 - (Number of selected features / Total number of features) , the advantage of LRR lies in its enhanced sensitivity, especially when the number of selected features is small. Table 1 demonstrates the effectiveness of the LCM measure. By focusing solely on the AUC performance, one might opt for the set with 50 features that presents the highest AUC. However, when evaluating based on the LCM, the most optimal feature set is that with only 10 features. This LCM-based selection results in a minor AUC performance drop of 0.002, but with the benefit of reducing the selected feature count by 40. Table 1 Comparison of two measures, LRR and LCM. Number of selected features Performance (AUC) LRR LCM 5 0.852 0.811 0.850 10 0.858 0.730 0.851 25 0.859 0.622 0.845 50 0.860 0.541 0.842 EXPERIMENTAL PROCEDURE Benchmark datasets To evaluate the previous and proposed feature selection methods, we utilized 24 publicly available datasets designed for binary classification, as presented in Table 2 . Of these, No. 1 to 17 are cancer-related microarray datasets. The features within these microarray datasets represent gene expression profiles, which are often referred to as probes. These datasets were acquired from a widely-used R package, “datamicroarray” [ 20 ]. We also included additional datasets (No. 18 to 24) obtained from public databases [ 21 , 22 ] to demonstrate that the proposed method is effective for non-microarray or non- HDLSS data. No. 1 to 18 are HDLSS datasets, while No. 19 to 24 are not. Table 2 List of benchmark datasets. No. Dataset Disease Instances Features Dimensionality Index* Balance Ratio** 1 ALLAML Leukemia 72 7,129 2.07 0.347 2 alon Colon Cancer 62 2,000 1.84 0.355 3 borovecki Huntington 31 22,283 2.92 0.452 4 chiaretti Leukemia 128 12,625 1.95 0.422 5 chin Breast Cancer 118 22,215 2.1 0.364 6 chowdary Breast Cancer 104 22,283 2.16 0.404 7 GLI_85 Gliomas 85 22,283 2.25 0.306 8 gordon Lung Cancer 181 12,533 1.82 0.171 9 gravier Breast Cancer 168 2,905 1.56 0.339 10 pomeroy CNS Tumor 60 7,128 2.17 0.35 11 Prostate_GE Prostate Cancer 102 5,966 1.88 0.49 12 shipp Lymphoma 77 7,129 2.04 0.247 13 singh Prostate Cancer 102 12,600 2.04 0.49 14 SMK_CAN_187 Lung cancer 187 19,993 1.89 0.481 15 subramanian N/A 50 10,100 2.36 0.34 16 tian Myeloma 173 12,625 1.83 0.208 17 west Breast Cancer 49 7,129 2.28 0.49 18 arcene 200 10,000 1.74 0.44 19 gisette 7,000 5,000 0.96 0.5 20 Hill_valley 1,212 100 0.65 0.495 21 ionosphere 351 33 0.6 0.359 22 madelon 2,600 500 0.79 0.5 23 sonar 208 60 0.77 0.466 24 wdbc 569 30 0.54 0.373 * Dimensionality index = log( \(\:Number\:of\:features\) )/log( \(\:Number\:of\:instances\) ). It is a measure of how high-dimensional a given dataset is. ** Balanced ratio: the proportion of the samples in lower class of the entire dataset. A value 0.5 is ideal for a binary class dataset. Compared feature selection methods and experimental environment To assess the efficacy of the proposed method, we compared it with various established feature selection methods, as presented in Table 3 . Table 3 List of benchmark feature selection methods. No. Feature selection method Description 1 F-statistic Filter method. 2 mRMR Filter method that incorporates redundancy measurement 3 Permutation importance Filter method that considers feature interactions 4 Infinite feature selection (INF) Graph-based filtering method [ 10 ] 5 GRACES Graph convolutional network-based feature selection [ 6 ] Among these, the first four adopted a filter-based approach. For each technique, feature rankings were computed and the top 100 features were selected. In the case of GRACES, a fixed number of features were produced. Hence, to maintain parity with the number of features selected by the other methods, we set this count to 100 and used the resulting rankings. The rankings for both INF and GRACES were determined after parameter optimization using 5-fold cross validation. Direct comparisons between any filter method and the proposed approach are challenging because filter methods do not explicitly indicate a subset of features that yield optimal performance. Consequently, it is imperative to identify the optimal feature subsets for benchmarking methods. For this purpose, we employed two distinct strategies: ‘simple sequential search’ and ‘forward search.’ In simple sequential search, the performance of the classification model was evaluated at each step while increasing the number of features from 1 to 100 based on the ranking of each feature, and the subset of features with the highest performance was selected. Forward search is similar to simple sequential search, but it continues selecting good features until no further improvement in performance is observed. To evaluate the quality of the selected features, we applied the eXtreme Gradient Boosting (XGB) classifier, equipped with specific parameters, nrounds = 5, objective = "multi:softmax", while retaining default configurations for the remaining parameters. Owing to its established potency and computational efficiency, XGB is an appropriate choice for evaluating feature selection methods. For the performance evaluation, we incorporated two metrics: Area Under the Curve (AUC) and LCM. AUC is a useful measure for comparing the performance of prediction models, especially when the classes of the dataset are imbalanced. As explained previously, LCM is suitable for evaluating both the classification performance and the number of selected features in the HDLSS dataset. For our analysis, a θ value of 0.8 was adopted as the reference point. Figure 3 illustrates the overall process of the comparative experiments. To orchestrate the efficacious experiments, we created two computational environments, R and Python, as shown in Table 4 . Although the rankings for GRACE and INF were procured within a Python framework, the search and evaluation phases for all the techniques were uniformly executed in the R environment. The associated experimental code is available at https://bitldku.github.io/home/sw/heuristic_search.html . Table 4 Experimental environment. Task Hardware OS Language Package (Library) - F-statistics - mRMR -Permutation importance - Proposed Performance evaluation Processor:AMD Ryzen 9 5900X 12-Core RAM:16GB, GPU:RTX3090 Windows 11 R 4.3.1 mlr 2.19.1 ranger 0.15.1 mRMRe 2.1.2.1 Parallel (base) xgboost 1.7.5.1 caret 6.0–94 - GRACES - INF Processor:Intel i7 = 10700 16-Core RAM:32.0GB GPU:RTX3080 Ubuntu 20.04.6 LTS Python 3.8.10 pytorch 2.0.1 torch_geometric 2.3 sklearn 1.2.2 xgboost 1.7.6 numpy 1.24.3 pandas 2.0.2 scipy 1.10.1 INF( https://pypi.org/project/PyIFS/ ) GRACES( https://github.com/canc1993/graces ) RESULTS Number of selected features through GPF and HTS The number of selected features is an important criterion for feature selection in the HDLSS datasets. The proposed method comprised two procedures for feature selection: GPF and HTS. The GPF filters useless features, and the HTS chooses informative features from GPF results. Table 5 summarizes the results of the GPF and HTS analyses. Details are provided in Supplementary Table S1 . During the GPF process, approximately 98.2% of the total features were eliminated from the 18 HDLSS datasets. Conversely, in the six non-HDLSS datasets, the average feature reduction rate was 61.2%, indicating considerable variability among the datasets. Notably, datasets with fewer than 100 features were not filtered. In the HTS phase, an additional 97.6% of the previously reduced features in the HDLSS dataset were excluded. The non-HDLSS datasets also exhibited a significant elimination rate, with an average of 98.3%. In conclusion, after the integrated application of both the GPF and HTS methodologies, the HDLSS datasets consistently selected 12 or fewer features, representing an average of merely 0.07% of the initial feature set. Table 5 Average number of features determined by proposed GPF and HTS procedures. Dataset Original(A) GPF(B) HTS(C) Reduction Rate (1-B/A) Selection Rate (C/A) HDLSS (1–18) 12,163.6 222.6 5.3 0.982 0.0004 Non-HDLSS (19–24) 954.8 370.8 6.3 0.612 0.0066 Performance comparison based on simple sequential search In this section, we compare the classification performance the previous and proposed methods using the selected features. To determine the performance of the filter method, a simple sequential search was performed as described in the “Compared feature selection methods” section. AUC and LCM were used as performance criteria. A summary of the comparison results is presented in Table 6 . Details are provided in Supplementary Table S2 and S3. Table 6 Comparison of the number of selected features, AUC, and LCM between the proposed method and previous works using simple sequential search. Method Selected AUC LCM In AUC In LCM features Top Ranking Top Ranking F-statistics (FS) 32.9 0.891 0.833 2 3.2 ± 1.2 1 3.5 ± 1.0 mRMR (MR) 27.8 0.874 0.824 3 2.9 ± 1.3 3 3.0 ± 1.4 Permutation (PE) 21.7 0.905 0.855 5 2.3 ± 0.9 1 2.5 ± 0.7 INF 55.3 0.796 0.731 0 5.3 ± 0.8 0 5.2 ± 0.5 GRACES (GE) 51.3 0.816 0.748 0 5.1 ± 0.7 0 5.3 ± 0.7 Proposed (PRO) 5.5 0.927 0.900 16 1.7 ± 1.3 20 1.3 ± 0.7 When compared using the AUC performance metric, the proposed method exhibited equivalent or superior performance in 16 out of 24 datasets, while maintaining the lowest performance variability among all methods, as detailed in Supplementary Table S2 and S3. The AUC of our method exceeded that of comparative approaches, with margins ranging from 2.2–13.1%. However, the number of selected features was markedly reduced, averaging only 5.5 features, which is between 1/4 and 1/10 that of the other methods selected. Utilizing the LCM performance metric with a θ value of 0.8, our method displayed enhanced performance in 20 out of 24 datasets, representing 83.3%. This indicates that when considering both the performance and influence of the number of selected features, the proposed method excels, highlighting its ability to achieve comparable or better results with fewer features. Among the traditional methods, permutation exhibits the best performance, followed by F-statistics, mRMR, GRACES, and INF. Notably, GRACES and INF exhibited a tendency to select a greater number of features than the other methods. Furthermore, INF was unable to produce results for datasets exceeding 20,000 features, and GRACES also encountered difficulties in generating results for specific datasets. Performance comparison based on forward search In this section, we compare the classification performances of previous methods and the proposed method. To determine the performance of the filter method, a forward and a simple sequential search was used as described in the “Compared feature selection methods” section. A summary of the results is presented in Table 7 , and the complete experimental results are detailed in Supplementary Table S4 and S5. Table 7 Comparison of the number of selected features, AUC, and LCM between the proposed method and previous works using forward search. Method Selectedfeatures AUC LCM In AUC In LCM Top Ranking Top Ranking F-statistics (FS) 6.3 0.905 0.876 2 2.7 ± 1.0 2 2.7 ± 1.0 mRMR (MR) 5.8 0.868 0.849 1 3.4 ± 1.3 1 3.3 ± 1.4 Permutation (PE) 5.8 0.905 0.877 2 2.7 ± 1.2 2 2.7 ± 1.2 INF 6.7 0.819 0.806 1 4.1 ± 1.4 0 4.4 ± 1.2 GRACES (GE) 5.3 0.816 0.811 1 4.5 ± 1.5 0 4.7 ± 1.3 Proposed (PRO) 5.5 0.927 0.900 20 1.3 ± 0.7 21 1.3 ± 0.7 Compared to the results in previous section, most filter methods showed an improvement in the AUC and significantly reduced the number of selected features. This trend suggests that the top 100 features chosen using traditional filtering methods exhibit high redundancy. A forward search helped remove many overlapping features. In specific datasets subh as Alon, both the feature count and AUC decreased, which might indicate a fall into local optima. Using the AUC metric, the proposed method matched or exceeded the performance of the other methods on 20 of the 24 datasets (83.3%). It also exhibited the lowest variability in performance. The AUC of the proposed method was between 2.2% and 11.1% higher than those of the other methods. Although the comparative methods selected fewer features on average than the sequence results, the proposed method consistently selected the least number of features. Furthermore, when using the LCM with θ at 0.8, the proposed method outperformed other methods in 21 of the 24 datasets (87.5%). Execution time of the proposed method Supplementary Table S6 lists the average runtimes of our method over 30 iterations for the HDLSS datasets. The Arcene dataset required approximately 12 min, the longest duration, whereas 16 out of the 18 datasets completed the feature selection within 5 min. Notably, the HTS phase constituted 77% of the total duration, with 94% spiking in some cases. The execution times for both GPF and HTS increased proportionally with instance size. Conversely, apart from SMK_CAN_187, the GPF duration remained relatively consistent irrespective of the feature count, whereas HTS showed a pronounced rise as the features increased. In essence, the duration of the HTS phase significantly shapes the total runtime, with the efficiency of the GPF in transmitting a reduced feature set being pivotal to the HTS timeframe. In conclusion, the average execution time of the proposed method is 169 seconds (2.8 minutes), which is reasonable. DISCUSSION AND CONCLUSIONS The essence of feature selection lies in effectively finding a subset that is close to the global optimum, where the relevance to the class is high and the redundancy between features is low. We proposed two methods, GPF for filtering and HTS for searching and verified their superior performance compared to those of other methods. Particularly, for microarray datasets with HDLSS properties, we selected a minimal subset of core features within a reasonable time. The advantages of the proposed method are summarized as follows: Incorporating feature interactions for relevance measurement in the filtering stage Oh [ 23 ] experimentally and theoretically proved that the importance of a feature can be decomposed into its intrinsic predictive power (feature power) and the effect arising when combined with other features (interaction). The importance of a feature not only lies in its intrinsic power but also varies depending on the surrounding features. Therefore, the relevance between the subset and class can be evaluated more accurately by considering feature interactions. As we progressively filtered out features with low importance in the filtering stage and then measured their importance within the refined subset, we were able to accurately measure the interactions based on this subset. Although the permutation importance method we used does not explicitly distinguish between feature power and interaction, it is clear that the interaction effect is reflected within the importance. Elimination of redundancy among features during the search stage While basic permutation importance can account for interactions between features, it does not capture redundancy. Thus, there is a possibility that highly redundant features exist within the filtered subset. By contrast, the proposed HTS can exclude redundancy among the features. To incorporate a new feature into the existing subset in the HTS process, the performance should be improved. If a perfectly redundant feature exists, no performance improvement observed, which leads to its exclusion. Therefore, HTS can play a complementary role in removing redundant features present in GPF. Clear threshold criteria in the filtering stage Traditional filter methods do not consider the best candidate features for the next selection stage. Therefore, the users randomly select the best n features. In other methods, the candidate features produced depend on threshold. Therefore, users must tune these parameters to obtain better candidate features. By contrast, the proposed method produces effective candidate features without parameter settings. For our filtering process, we utilized a fixed threshold value based on theoretical evidence. A permutation importance of zero indicates that the feature has no effect on the performance improvement [ 24 ]. If a feature degrades the performance or causes a negative interaction, its importance becomes less than 0. Hence, setting the threshold to zero allows for a clear and intuitive filtering of features that contribute to performance enhancement. Improvement of computational efficiency using feature ranking In the filtering stage, existing methods utilize feature evaluation values exclusively for the removal of weak features, whereas our proposed method harnesses these values to optimize subsequent computations. Specifically, SFFS [ 25 ] thoroughly explores the whole feature space, encompassing both selected and unselected features, to identify pairs that can overcome the nesting effect. However, by leveraging the feature-ranking information already obtained during the filtering stage, our approach constrained the search space. This not only enhances the computational efficiency during the search phase but also avoids performance degradation. Expended search space to get better features Beyond “consolation match”, we incorporated a semi-random approach to pinpoint the global optimum. At the start of the search, we considered multiple candidate features with high importance to alter the first-choice feature and selected the feature subset with the best performance. From our observations, the optimal feature set frequently emerged when the first-choice feature was not ranked as the top feature. Selection of a sufficiently good features for microarray datasets To understand or diagnose complex diseases, it is essential to not only achieve high predictive (classification) performance but also identify pivotal biomarkers (features) [ 24 ]. For example, a predictive accuracy of 0.82 with 20 features is more meaningful than an accuracy of 0.85 with 120 features. Although the proposed method does not always result in a better predictive performance than the other methods, the derived features are very compact with a decent prediction performance. Therefore, it would be helpful to identify novel genes associated with diseases. In our proposed method, during the consolation match, we specifically operated on the latter half of the selected features and the former half of the non-selected features. Unexpectedly, when we expanded this search range during our experiments, a decline in the performance was observed. This phenomenon raises questions that require further exploration, particularly concerning feature interactions. This is a topic for further research. Declarations Ethics approval and consent to participate Not applicable. Consent for publication Not applicable. Availability of data and materials Code about the results of this study are publicly shared at: https://bitldku.github.io/home/sw/heuristic_search.html. Competing interests The authors declare that they have no competing interests. Funding This work was supported by the Institute for Information & Communications Technology Planning & Evaluation (IITP) grand funded by the Ministry of Science, ICT (MSIT), Korea (No. RS-2023-00222191, Development of data fabric technology to support logical data integration and compound analysis of distributed data) Authors' contributions Shin and Oh conceptualized the research. Shin contributed to development of the model. Shin collected and annotated data. All authors read and approved the final manuscript. Acknowledgements None References Ang JC, Mirzal A, Haron H, Hamed HNA. Supervised, Unsupervised, and Semi-Supervised Feature Selection: A Review on Gene Selection. IEEE/ACM Trans Comput Biol Bioinform. 2016;13(5):971–89. 10.1109/TCBB.2015.2478454 . Lu H, Chen J, Yan K, Jin Q, Xue Y, Gao Z. A hybrid feature selection algorithm for gene expression data classification. Neurocomputing. 2017;256:56–62. 10.1016/j.neucom.2016.07.080 . Almugren N, Alshamlan H. A Survey on Hybrid Feature Selection Methods in Microarray Gene Expression Data for Cancer Classification. IEEE Access. 2019;7:78533–78548 doi:0.1109/ACCESS.2019.2922987. Bommert A, Sun X, Bischl B, Rahnenführer J, Lang M. Benchmark for filter methods for feature selection in high-dimensional classification data. Comput Stat Data Anal. 2020;143:106839. 10.1016/j.csda.2019.106839 . Manikandan G, Abirami S. Feature Selection Is Important: State-of-the-Art Methods and Application Domains of Feature Selection on High-Dimensional Data. In: Kumar R, Paiva S, editors. Applications in Ubiquitous Computing. Cham: Springer; 2021. pp. 177–96. Chen C, Weiss ST, Liu YY. Graph convolutional network-based feature selection for high-dimensional and low-sample size data. Bioinformatics. 2023;39(4):btad135. 10.1093/bioinformatics/btad135 . Alhenawi E, Al-Sayyed R, Hudaib A, Mirjalili S. Feature selection methods on gene expression microarray data for cancer classification: A systematic review. Comput Biol Med. 2022;140:105051. 10.1016/j.compbiomed.2021.105051 . Liu B, Wei Y, Zhang Y, Yang Q. Deep neural networks for high dimension, low sample size data. In: Proceedings of the 26th International Joint Conference on Artificial Intelligence. Melbourne: AAAI, 2017;2287–2293. Li K, Wang F, Yang L, Liu R. Deep feature screening: Feature selection for ultra high-dimensional data via deep neural networks. Neurocomputing. 2023;538:126186. 10.1016/j.neucom.2023.03.047 . Roffo G, Melzi S, Castellani U, Vinciarelli A, Cristani M. Infinite Feature Selection: A Graph-based Feature Filtering Approach. IEEE Trans Pattern Anal Mach Intell. 2020;43(12):4396–410. 10.1109/TPAMI.2020.3002843 . Rostami M, Forouzandeh S, Berahmand K, Soltani M, Shahsavari M, Oussalah M. Gene selection for microarray data classification via multi-objective graph theoretic-based method. Artif Intell Med. 2022;123:102228. 10.1016/j.artmed.2021.102228 . Pati SK, Banerjee A, Manna S. Gene selection of microarray data using Heatmap Analysis and Graph Neural Network. Appl Soft Comput. 2023;135:110034. 10.1016/j.asoc.2023.110034 . Li Y, Chen CY, Wasserman WW. Deep Feature Selection: Theory and Application to Identify Enhancers and Promoters. J Comput Biol. 2016;23(5):322–36. 10.1089/cmb.2015.0189 . Chowdhury S, Dong X, Li X. Recurrent Neural Network Based Feature Selection for High Dimensional and Low Sample Size Micro-array Data. In: Proceedings of 2019 IEEE International Conference on Big Data. Piscataway: IEEE, 2019;4823–4828. Agrawal P, Abutarboush HF, Ganesh T, Mohamed AW. Metaheuristic Algorithms on Feature Selection: A Survey of One Decade of Research (2009–2019). IEEE Access. 2021;9:26766–91. 10.1109/ACCESS.2021.3056407 . Dokeroglu T, Deniz A, Kiziloz HE. A comprehensive survey on recent metaheuristics for feature selection. Neurocomputing. 2022;494:269–96. 10.1016/j.neucom.2022.04.083 . Ferri FJ, Pudil P, Hatef M, Kittler J. Comparative study of techniques for large-scale feature selection. Mach Intell Pattern Recognit. 1994;16:403–13. 10.1016/B978-0-444-81892-8.50040-7 . Wang L, Wang Y, Chang Q. Feature selection methods for big data bioinformatics: A survey from the search perspective. Methods. 2016;111:21–31. 10.1016/j.ymeth.2016.08.014 . Hsu HH, Hsieh CW, Lu MD. Hybrid feature selection by combining filters and wrappers. Expert Syst Appl. 2011;38(7):8144–50. 10.1016/j.eswa.2010.12.156 . Ramey J, Datamicroarray. Jan. Available at https://github.com/ramhiser/datamicroarray (accessed 19 2024). Arizona State University (ASU). Feature selection datasets. Available at https://jundongl.github.io/scikit-feature/datasets.html (accessed 19 Jan. 2024). OpenML. A worldwide machine learning lab. Available at https://www.openml.org/ (accessed 19 May 2024). Oh S. Predictive case-based feature importance and interaction. Inf Sci. 2022;593:155–76. 10.1016/j.ins.2022.02.003 . Mi X, Zou B, Zou F, Hu J. Permutation-based identification of important biomarkers for complex diseases via machine learning models. Nat Commun. 2021;12(1):3008. 10.1038/s41467-021-22756-2 . Pudil P, Novovičová J, Kittler J. Floating search methods in feature selection. Pattern Recognit Lett. 1994;15(11):1119–25. 10.1016/0167-8655(94)90127-9 . Additional Declarations No competing interests reported. Supplementary Files Supplementarysubmit.docx Cite Share Download PDF Status: Published Journal Publication published 26 Dec, 2024 Read the published version in BMC Bioinformatics → Version 1 posted Editorial decision: Revision requested 29 Nov, 2024 Reviews received at journal 31 Oct, 2024 Reviews received at journal 23 Oct, 2024 Reviewers agreed at journal 23 Oct, 2024 Reviewers agreed at journal 22 Oct, 2024 Reviewers agreed at journal 22 Oct, 2024 Reviewers agreed at journal 22 Oct, 2024 Reviewers invited by journal 22 Oct, 2024 Editor invited by journal 17 Oct, 2024 Editor assigned by journal 16 Oct, 2024 Submission checks completed at journal 15 Oct, 2024 First submitted to journal 14 Oct, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-5260669","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":387094987,"identity":"9d1f6dcc-9752-4731-b012-062e24db47b1","order_by":0,"name":"Hyunseok Shin","email":"","orcid":"","institution":"Dankook University","correspondingAuthor":false,"prefix":"","firstName":"Hyunseok","middleName":"","lastName":"Shin","suffix":""},{"id":387094988,"identity":"46535cc2-b5df-4057-a5a3-2899f336d68c","order_by":1,"name":"Sejong Oh","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA+0lEQVRIiWNgGAWjYHCCBDDBD6QZGxgYDEA8EIOwFskGErRA9BkcIFaLfHvDw8cFv2zyjI8nP3s4o8LGmIH98APGmXtwazE4cyDZeGZfWrHZmWfmhhvOpJkx8KQZMG54hkeLREKaNG/P4cRtNxLMJB+2HbZhYMhhYHxwAI/D5j+AaNk8I/0bRAv/G/xaGG4wpEnz/DicuEEix0xyY9thMwYJoC0b8GgxOJOQbMzbkJY448ybMskZZ9KM2SSeGRycgc9h7WcSH/P8sUnsb0/fJtlTYWPYz5/88GEPPocx8ADjow2JzwbEeDUwMLAD5f/gVzIKRsEoGAUjHAAAPetZMwSrmjsAAAAASUVORK5CYII=","orcid":"","institution":"Dankook University","correspondingAuthor":true,"prefix":"","firstName":"Sejong","middleName":"","lastName":"Oh","suffix":""}],"badges":[],"createdAt":"2024-10-14 11:23:03","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-5260669/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-5260669/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1186/s12859-024-06017-9","type":"published","date":"2024-12-26T15:57:13+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":71503244,"identity":"2b12ba64-9650-40f3-9e70-077cbda33709","added_by":"auto","created_at":"2024-12-16 09:24:30","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":11546,"visible":true,"origin":"","legend":"\u003cp\u003eGeneral process of feature selection for HDLSS datasets.\u003c/p\u003e","description":"","filename":"floatimage18.png","url":"https://assets-eu.researchsquare.com/files/rs-5260669/v1/75a0fbbf2b64fd3b78abcf2c.png"},{"id":71503246,"identity":"cd932ef2-f9b4-488a-af5c-4f51c73ca4c5","added_by":"auto","created_at":"2024-12-16 09:24:31","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":19266,"visible":true,"origin":"","legend":"\u003cp\u003eProcedure of the proposed method.\u003c/p\u003e","description":"","filename":"OnlineFig2.png","url":"https://assets-eu.researchsquare.com/files/rs-5260669/v1/a07a251f0bc43a60ea100b26.png"},{"id":71503260,"identity":"f16e3c48-58e8-4752-9b12-1508bdb22da3","added_by":"auto","created_at":"2024-12-16 09:24:39","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":84583,"visible":true,"origin":"","legend":"\u003cp\u003eOverall procedure of the comparative experiment.\u003c/p\u003e","description":"","filename":"OnlineFig3.png","url":"https://assets-eu.researchsquare.com/files/rs-5260669/v1/f74699221f5b1494e13b7eea.png"},{"id":72640749,"identity":"a1aabe8b-8fa1-4818-8ecc-4a2460eac02e","added_by":"auto","created_at":"2024-12-30 16:09:24","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1157591,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-5260669/v1/2aed4fdc-780b-4044-ae8e-10a9f6bfb8ec.pdf"},{"id":71503258,"identity":"3131e142-dc31-455e-acf9-a70d3c71efab","added_by":"auto","created_at":"2024-12-16 09:24:36","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":62754,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarysubmit.docx","url":"https://assets-eu.researchsquare.com/files/rs-5260669/v1/3651b565ea821ad441b1d62c.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"An effective heuristic for developing hybrid feature selection in high dimensional and low sample size datasets","fulltext":[{"header":"BACKGROUND","content":"\u003cp\u003eIn the current data era, information is recorded intricately across various dimensions. Consequently, the dimensionality of the data increases rapidly. High-dimensional datasets refer to data where the number of features (\u003cem\u003ep\u003c/em\u003e) is considerably larger than the sample size (\u003cem\u003en\u003c/em\u003e). High-dimensional datasets have driven significant scientific discoveries in fields such as biology and bioinformatics by revealing previously unknown complex patterns. However, analyzing and utilizing high dimensional and low sample size (HDLSS) datasets remain a challenge for researchers [\u003cspan additionalcitationids=\"CR2 CR3 CR4 CR5\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMicroarray data are prime examples of HDLSS datasets. With the capacity to capture the expression of tens of thousands of genes, the core aim is to identify genetic variations in specific cancers or other genetically-related diseases [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e]. However, a notable proportion of these genes either lacks direct relevance to the disease or exhibits substantial overlap in feature information, leading to redundancy [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. Furthermore, owing to the patient-centric nature of these data, sample size was inherently limited. These aspects amplify risks such as overfitting and the curse of dimensionality [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. Consequently, developing cancer classification models and analyzing HDLSS datasets such as microarrays, pose intricate challenges for computer science researchers [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eA fundamental approach for addressing these challenges in machine learning is dimensionality reduction via feature selection. The primary objective of feature selection is to filter out irrelevant and redundant features, thereby honing the most pertinent features relevant to the subject of interest. This enhances the clarity, generalizability, and predictive accuracy of classification models, concurrently reducing computational demands and boosting efficiency [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. In bioinformatics, specifically, feature selection plays a critical role in uncovering potential avenues for drug development and offering insights into disease diagnostics and causation [\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIdentifying the optimal features is an NP-hard problem [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e, \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Numerous feature-selection methods have been proposed to address this challenge. These methods are typically classified into three categories based on their interactions with the classification model: filter, wrapper, and embedded methods [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. The filter method takes a model-independent stance, assigns rankings to features based on metrics such as statistical or information-theoretic properties, and subsequently selects those that exceed a predetermined threshold. In contrast, the wrapper method incorporates a model-dependent approach, by selecting feature subsets based on the performance of a specific model. The embedded method differs from the wrapper method by integrating feature selection directly into the model-training process.\u003c/p\u003e \u003cp\u003eRecent advancements in feature selection include methods based on graphs and deep learning frameworks. These methods aim to capture the intricate relationships among features. In the graph-based approach, features are visualized as nodes and their interrelationships as edges. This method determines the importance of features by analyzing the structural properties of a graph [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e, \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e, \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Deep learning techniques for feature selection utilize neural networks with features as inputs [\u003cspan additionalcitationids=\"CR7 CR8\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e, \u003cspan additionalcitationids=\"CR13\" citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e]. The significance of each feature is gauged by the magnitude of the gradients during network training.\u003c/p\u003e \u003cp\u003eFor HDLSS data, several studies on feature selection have adopted a hybrid approach involving two stages, filtering and searching, as illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e]. This approach seeks to harness the computational efficiency of filter methods and the high performance of wrapper methods. The filter method can be viewed as a preprocessing step that significantly reduces the search scope by eliminating unnecessary features. By contrast, the wrapper method refines the search within this narrow scope, aiming to identify the most optimal features more precisely [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e, \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eFrom the collective insights of previous studies on feature selection, we can draw the following conclusions:\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eUnderstanding intricate relationships, such as the interactions between features, is pivotal.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eFinding the optimal features for HDLSS is realistically difficult. An alternative is to find a better solution among several suboptimal features. Therefore, it is necessary to avoid the local optima.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eThis can be achieved by balancing the diversity and focus in the feature search space, introducing randomness during the search process, and employing a hybrid method that leverages both filter and wrapper techniques.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eIn this study, we propose a new metaheuristic method [\u003cspan additionalcitationids=\"CR16 CR17 CR18\" citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e] to improve upon previous feature selection approaches for HDLSS datasets, aiming to achieve near-optimal performance with minimal features. By integrating gradual permutation filtering with a diverse search strategy, we designed our approach based on core principles. The highlights of our method include. The details are outlined in the next section.\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eFiltering: Leveraging permutation importance, filtering accounts for feature interactions and establishes a correct threshold.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eSearch: Through \u0026ldquo;consolation matches,\u0026rdquo; the nesting effect can be overcome, wherein a once-selected (or excluded) feature remains unaltered. Additionally, by varying \u0026ldquo;first-choice features,\u0026rdquo; the search range is expanded to approach a more global optimum.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eRuntime optimization: Using ranked features, the method efficiently reduces the search space.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eAssessing the fitness of the selected features: We crafted a unique performance metric for the HDLSS, providing a holistic perspective on feature quantity and quality.\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e"},{"header":"METHODS","content":"\u003cp\u003eOverall procedure\u003c/p\u003e \u003cp\u003eThe proposed method also follows the general procedure of feature selection for HDLSS datasets, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. As illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and elaborated in Algorithm 1, the proposed method comprises two primary phases: filtering and searching.\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eGradual Permutation Filtering (GPF): This phase ingests all the HDLSS data features. It ranks features based on their permutation importance and subsequently eliminates irrelevant features (Algorithm 2).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003eHeuristic Tribrid Search (HTS): By leveraging the features ranked by the GPF, HTS employs a heuristic blend of forward search, consolation match, and backward elimination. This approach aims to identify the near-optimal feature set by considering both the number of features and classification performance using the features (see Algorithms 3\u0026ndash;6).\u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eThe subsequent sections will delve into the specifics of each phase.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Taba\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm 1. Overall procedure\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eWF \u0026ndash; whole features\u003c/p\u003e \u003cp\u003eCL - class label vector\u003c/p\u003e \u003cp\u003e\u003cb\u003eOutput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eSFL - selected feature list\u003c/p\u003e \u003cp\u003e// Gradual Permutation Filtering\u003c/p\u003e \u003cp\u003eFILTERED \u0026larr; WF\u003c/p\u003e \u003cp\u003e\u003cb\u003eFOR\u003c/b\u003e RATIO \u003cb\u003eIN\u003c/b\u003e [0.1, 0.2, ..., 0.9] \u003cb\u003eDO\u003c/b\u003e\u003c/p\u003e \u003cp\u003eRANKED \u0026larr; gradual_permutation_filtering(FILTERED, CL, RATIO) // Algorithm 2\u003c/p\u003e \u003cp\u003eFILTERED \u0026larr; FILTERED [RANKED]\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND FOR\u003c/b\u003e\u003c/p\u003e \u003cp\u003e// Heuristic Tribrid Search: This section can be processed in parallel.\u003c/p\u003e \u003cp\u003eN \u0026larr; Length of WF\u003c/p\u003e \u003cp\u003e\u003cb\u003eFOR\u003c/b\u003e i \u003cb\u003eFROM\u003c/b\u003e 1 \u003cb\u003eTO\u003c/b\u003e 10\u003c/p\u003e \u003cp\u003eSELECTED[i], max_fit[i]\u0026thinsp;=\u0026thinsp;heuristic_tribird_search(FILTERED, CL, i, N) // Algorithm 3\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND FOR\u003c/b\u003e\u003c/p\u003e \u003cp\u003eSFL \u0026larr; Choose the SELECTED with the highest max_fit from the 10 results\u003c/p\u003e \u003cp\u003e\u003cb\u003eRETURN\u003c/b\u003e SFL\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eGradual permutation filtering\u003c/p\u003e \u003cp\u003eThe GPF stage prioritizes features based on their permutation importance and eliminates unimportant features. This method chooses features with an importance value greater than zero because values below this threshold often indicate redundancy or suggest potential noise.\u003c/p\u003e \u003cp\u003eThis method offers a refined approach for evaluating feature importance by adopting a gradual process of eliminating noise and recalculating the importance, thereby minimizing the potential biases associated with single-step elimination. Considering that the permutation importance is evaluated within the evaluation group (features), the approach emphasizes filtering out irrelevant features and reevaluating the importance of enhancing purity, which is defined solely based on performance-impacting features. Notably, a gradual feature filtering method, as opposed to filtering all at once, provides a more precise selection of features with importance on the verge of 0.\u003c/p\u003e \u003cp\u003eTo ensure robust feature selection, the GPF measures the permutation importance of each feature multiple times; i.e., 50 times in our case. From these measurements, the features that exceed the importance of zero for a specific threshold number were selected. For the selected features, the permutation importance was recalculated by applying a progressively higher threshold each time. This iterative process, detailed in Algorithm 2, refines the ranking of significant features. The final ranking of the features was determined by averaging the last measured importance values of the features selected until the end.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabb\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm 2. gradual_permutation_filtering(FILTERED, CL, RATIO)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eFILTERED - filtered features\u003c/p\u003e \u003cp\u003eCL - class label vector\u003c/p\u003e \u003cp\u003eRATIO - threshold ratio\u003c/p\u003e \u003cp\u003e\u003cb\u003eOutput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eRANKED - ranked feature list\u003c/p\u003e \u003cp\u003eCONST M \u0026larr; 50\u003c/p\u003e \u003cp\u003eTHRESHOLD \u0026larr; M * RATIO\u003c/p\u003e \u003cp\u003ePI \u0026larr; matrix of dimensions M by length of FILTERED\u003c/p\u003e \u003cp\u003e\u003cb\u003eFOR\u003c/b\u003e i \u003cb\u003eFROM\u003c/b\u003e 1 \u003cb\u003eTO\u003c/b\u003e M\u003c/p\u003e \u003cp\u003ePI[i] \u0026larr; permutation importance based on FILTERED, varying the seed each time\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND FOR\u003c/b\u003e\u003c/p\u003e \u003cp\u003eCNT \u0026larr; Count of positive importance values in each column of PI\u003c/p\u003e \u003cp\u003eMPI \u0026larr; the mean of PI across columns, resulting in a vector of the same length as FILTERED\u003c/p\u003e \u003cp\u003eMASK \u0026larr; Set to 1 for entries in CNT that are \u0026ge;\u0026thinsp;THRESHOLD, otherwise set to 0\u003c/p\u003e \u003cp\u003eRANKED \u0026larr; Selected features where the elementwise product of MPI and MASK is greater than 0, then sort them in descending order based on their importance\u003c/p\u003e \u003cp\u003e\u003cb\u003eRETURN\u003c/b\u003e RANKED\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eHeuristic tribrid search\u003c/p\u003e \u003cp\u003eWhile the GPF is designed to filter out unimportant features, HTS focuses on identifying the most informative features from a pool refined by the GPF. HTS operates in three distinct stages, \u003cb\u003eforward search\u003c/b\u003e, \u003cb\u003econsolation match\u003c/b\u003e, and \u003cb\u003ebackward elimination\u003c/b\u003e, as outlined in Algorithm 3.\u003c/p\u003e \u003cp\u003eWe slightly modified the original forward search technique. The original technique began with an empty feature set and determined its first feature. In contrast, the proposed method uses \u0026ldquo;first-choice feature\u0026rdquo; that is chosen from ranked feature list of GPF. This procedure incrementally adds to the feature set, guided by the performance increments detailed in Algorithm 4. The performance was measured using a classification metric based on the selected features.\u003c/p\u003e \u003cp\u003eWhen the performance increase caused by the selected features stops during the forward search, the process shifts to \u0026ldquo;consolation match.\u0026rdquo; This stage aims to enhance the performance by swapping a single feature between selected and unselected feature pools, as depicted in Algorithm 5. This strategy offers an escape route from potential local optima. Notably, the duration of this stage depends on the volume of features initially filtered by the GPF. Techniques for expediting this stage are explored in subsequent sections.\u003c/p\u003e \u003cp\u003eIf the consolation match yields improved performance, the forward search resumes further enhancements. However, if no such gains are evident, HTS transitions to a backward elimination phase. This final stage discards any remaining unimportant features, thereby ensuring the relevance of the end feature set.\u003c/p\u003e \u003cp\u003eImportantly, at each stage of the HTS (Algorithms 4\u0026ndash;6), the performance is gauged using a novel metric introduced in this study. This metric considers both the classification performance of the model and feature count. An elaboration of this metric is presented in the following section.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabc\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm 3. heuristic_tribird_search(FILTERED, CL, i, N)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eFILTERED \u0026ndash; selected features sorted by ranking\u003c/p\u003e \u003cp\u003eCL - class label vector\u003c/p\u003e \u003cp\u003ei \u0026ndash; index number\u003c/p\u003e \u003cp\u003eN - Number of columns in original dataset\u003c/p\u003e \u003cp\u003e\u003cb\u003eOutput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eSELECTED - selected feature list\u003c/p\u003e \u003cp\u003emax_fit - maximum fit value\u003c/p\u003e \u003cp\u003eSELECTED \u0026larr; FILTERED[i]\u003c/p\u003e \u003cp\u003eCANDIDATE \u0026larr; FILTERED \u0026ndash; {SELECTED}\u003c/p\u003e \u003cp\u003emax_fit \u0026larr; 0\u003c/p\u003e \u003cp\u003emax_update \u0026larr; TRUE\u003c/p\u003e \u003cp\u003e\u003cb\u003eWHILE\u003c/b\u003e max_update \u003cb\u003eDO\u003c/b\u003e\u003c/p\u003e \u003cp\u003emax_update \u0026larr; FALSE\u003c/p\u003e \u003cp\u003eSELECTED, max_update, max_fit \u0026larr; forward_search(FILTERED, CL, SELECTED, max_fit, N) // Algorithm 4\u003c/p\u003e \u003cp\u003e\u003cb\u003eIF NOT\u003c/b\u003e max_update \u003cb\u003eTHEN\u003c/b\u003e\u003c/p\u003e \u003cp\u003eSELECTED, max_update, max_fit \u0026larr; consolation_match(FILTERED CL, SELECTED, max_fit, N) // Algorithm 5\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND IF\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND WHILE\u003c/b\u003e\u003c/p\u003e \u003cp\u003eSELECTED \u0026larr; Features obtained from SELECTED through backward elimination with stopping criteria based on Algorithm 6\u003c/p\u003e \u003cp\u003e\u003cb\u003eRETURN\u003c/b\u003e SELECTED, max_fit\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabd\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm 4. forward_search(FILTERED CL, SELECTED, max_fit, N)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eFILTERED \u0026ndash; selected features sorted by ranking\u003c/p\u003e \u003cp\u003eCL - class label vector\u003c/p\u003e \u003cp\u003eSELECTED - selected feature list\u003c/p\u003e \u003cp\u003emax_fit - maximum fit value\u003c/p\u003e \u003cp\u003eN - Number of columns in original dataset\u003c/p\u003e \u003cp\u003e\u003cb\u003eOutput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eSELECTED - selected feature list\u003c/p\u003e \u003cp\u003emax_update - max fit update status\u003c/p\u003e \u003cp\u003emax_fit - maximum fit value\u003c/p\u003e \u003cp\u003e\u003cb\u003eDO\u003c/b\u003e\u003c/p\u003e \u003cp\u003emax_update \u0026larr; FALSE\u003c/p\u003e \u003cp\u003eCANDIDATE \u0026larr; {columns of FILTERED } \u0026ndash; SELECTED\u003c/p\u003e \u003cp\u003e\u003cb\u003eFOR\u003c/b\u003e each feature F \u003cb\u003eIN\u003c/b\u003e CANDIDATE \u003cb\u003eDO\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTEMP \u0026larr; SELECTED \u0026cup; {F}\u003c/p\u003e \u003cp\u003ecurrent_fit \u0026larr; fitness(CL, TEMP, N)\u003c/p\u003e \u003cp\u003e\u003cb\u003eIF\u003c/b\u003e current_fit\u0026thinsp;\u0026gt;\u0026thinsp;max_fit \u003cb\u003eTHEN\u003c/b\u003e\u003c/p\u003e \u003cp\u003emax_fs \u0026larr; TEMP\u003c/p\u003e \u003cp\u003emax_fit \u0026larr; current_fit\u003c/p\u003e \u003cp\u003emax_update \u0026larr; TRUE\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND IF\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND FOR\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eIF max_update THEN\u003c/b\u003e\u003c/p\u003e \u003cp\u003eSELECTED \u0026larr; max_fs\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND IF\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eWHILE\u003c/b\u003e max_update\u003c/p\u003e \u003cp\u003e\u003cb\u003eRETURN\u003c/b\u003e SELECTED, max_update, max_fit\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabe\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm 5. consolation_match(FILTERED, CL, SELECTED, max_fit, N)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eFILTERED \u0026ndash; selected features sorted by ranking\u003c/p\u003e \u003cp\u003eCL - class label vector\u003c/p\u003e \u003cp\u003eSELECTED - selected feature list\u003c/p\u003e \u003cp\u003emax_fit - maximum fit value\u003c/p\u003e \u003cp\u003eN - Number of columns in original dataset\u003c/p\u003e \u003cp\u003e\u003cb\u003eOutput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003emax_fs - selected feature list with max fit value\u003c/p\u003e \u003cp\u003emax_update - max fit update status\u003c/p\u003e \u003cp\u003emax_fit - maximum fit value\u003c/p\u003e \u003cp\u003emax_update \u0026larr; FALSE\u003c/p\u003e \u003cp\u003eCANDIDATE \u0026larr; {columns of FILTERED } \u0026ndash; SELECTED\u003c/p\u003e \u003cp\u003e// Iterate over each feature in SELECTED and CANDIDATE to find an optimal swap\u003c/p\u003e \u003cp\u003e\u003cb\u003eFOR\u003c/b\u003e X \u003cb\u003eIN\u003c/b\u003e SELECTED \u003cb\u003eDO\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eFOR\u003c/b\u003e Y \u003cb\u003eIN\u003c/b\u003e CANDIDATE \u003cb\u003eDO\u003c/b\u003e\u003c/p\u003e \u003cp\u003eTEMP \u0026larr; (SELECTED - {X}) U {Y}\u003c/p\u003e \u003cp\u003ecurrent_fit \u0026larr; fitness(CL, TEMP, N)\u003c/p\u003e \u003cp\u003e\u003cb\u003eIF\u003c/b\u003e current_fit\u0026thinsp;\u0026gt;\u0026thinsp;max_fit \u003cb\u003eTHEN\u003c/b\u003e\u003c/p\u003e \u003cp\u003emax_fs \u0026larr; TEMP\u003c/p\u003e \u003cp\u003emax_fit \u0026larr; current_fit\u003c/p\u003e \u003cp\u003emax_update \u0026larr; TRUE\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND IF\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND FOR\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eEND FOR\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003eRETURN\u003c/b\u003e max_fs, max_update, max_fit\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eLog comprehensive metric for HDLSS datasets\u003c/p\u003e \u003cp\u003eBalancing the number of selected features with the classification performance is imperative when working with HDLSS datasets. Therefore, this study proposes log-comprehensive metric (LCM). This metric, which is a refined iteration of the conventional fitness function, was meticulously designed for the HDLSS datasets, as elaborated in Algorithm 6.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"No\" id=\"Tabf\" border=\"1\"\u003e \u003ccolgroup cols=\"1\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm 6. log_comprehensive_metric(CL, SELECTED, N)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eInput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eCL - class label vector\u003c/p\u003e \u003cp\u003eSELECTED - selected feature list\u003c/p\u003e \u003cp\u003eN - Number of columns in original dataset\u003c/p\u003e \u003cp\u003e\u003cb\u003eOutput\u003c/b\u003e:\u003c/p\u003e \u003cp\u003eLCM - log comprehensive metric value\u003c/p\u003e \u003cp\u003e// Calculate the LCR metric using a weighted average of the LRR (Log Reduction Rate) derived from performance and selected feature count.\u003c/p\u003e \u003cp\u003eCONST C \u0026larr; 0.005 // C is trade-off constant for feature count and performance.\u003c/p\u003e \u003cp\u003eX \u0026larr; length of SELECTED// Number of selected features\u003c/p\u003e \u003cp\u003eY \u0026larr; log(2)/log(N)\u003c/p\u003e \u003cp\u003eTHETA \u0026larr; Y/(C\u0026thinsp;+\u0026thinsp;Y)\u003c/p\u003e \u003cp\u003ePERFORMANCE \u0026larr; the average performance obtained with 5-fold and XGBoost on SELECTED and CL\u003c/p\u003e \u003cp\u003eLRR \u0026larr; 1-log(X)/log(N)\u003c/p\u003e \u003cp\u003eLCM \u0026larr; THETA*PERFORMANCE + (1-THETA)*LRR\u003c/p\u003e \u003cp\u003e\u003cb\u003eRETURN\u003c/b\u003e LCM\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eThe widely used fitness function is described by Eq.\u0026nbsp;(\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). It constitutes a weighted sum of the model error and the ratio of the selected features to the total. The aim is to minimize the following function:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:fitness=\\theta\\:*Error+\\left(1-\\theta\\:\\right)*\\left(\\frac{Number\\:of\\:selected\\:features}{Total\\:EquationNumber\\:of\\:features}\\right),\\:\\:\\:where\\:0\\le\\:\\theta\\:\\le\\:1$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eBy contrast, the LCM, expressed in Eq.\u0026nbsp;(\u003cspan refid=\"Equ2\" class=\"InternalRef\"\u003e2\u003c/span\u003e), fuses the model\u0026rsquo;s classification success with the log reduction rate (LRR) using weighted factors. The objective is to maximize.\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:LCM=\\theta\\:*Performance+\\left(1-\\theta\\:\\right)*LRR,\\:\\:where\\:\\:0\\le\\:\\theta\\:\\le\\:1$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eLRR is further explained by Eq.\u0026nbsp;(\u003cspan refid=\"Equ3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). Its value fluctuates between 0 and 1 and approaches 1 as the number of selected features decreases.\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:LRR=1-{\\text{log}}_{a}b=1-\\text{log}b/\\text{log}a,$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:where:0\\le\\:LRR\\le\\:1\\:,\\:\\:$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\:a=Total\\:EquationNumber\\:of\\:features,\\:b=\\:Number\\:of\\:selected\\:features\\:$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eIn contrast to the reduction rate (RR), which is formulated \u003cem\u003eas RR\u0026thinsp;=\u0026thinsp;1 - (Number of selected features / Total number of features)\u003c/em\u003e, the advantage of LRR lies in its enhanced sensitivity, especially when the number of selected features is small.\u003c/p\u003e \u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e demonstrates the effectiveness of the LCM measure. By focusing solely on the AUC performance, one might opt for the set with 50 features that presents the highest AUC. However, when evaluating based on the LCM, the most optimal feature set is that with only 10 features. This LCM-based selection results in a minor AUC performance drop of 0.002, but with the benefit of reducing the selected feature count by 40.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of two measures, LRR and LCM.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNumber of selected features\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePerformance (AUC)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLRR\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLCM\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.852\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.811\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.850\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.858\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.730\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.851\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.859\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.622\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.845\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.860\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.541\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.842\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003eEXPERIMENTAL PROCEDURE\u003c/h2\u003e \u003cp\u003eBenchmark datasets\u003c/p\u003e \u003cp\u003eTo evaluate the previous and proposed feature selection methods, we utilized 24 publicly available datasets designed for binary classification, as presented in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Of these, No. 1 to 17 are cancer-related microarray datasets. The features within these microarray datasets represent gene expression profiles, which are often referred to as probes. These datasets were acquired from a widely-used R package, \u0026ldquo;datamicroarray\u0026rdquo; [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. We also included additional datasets (No. 18 to 24) obtained from public databases [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e, \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e] to demonstrate that the proposed method is effective for non-microarray or non- HDLSS data. No. 1 to 18 are HDLSS datasets, while No. 19 to 24 are not.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of benchmark datasets.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"7\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDisease\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eInstances\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eFeatures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eDimensionality Index*\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eBalance Ratio**\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eALLAML\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLeukemia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e72\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7,129\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.07\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.347\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ealon\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eColon Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e62\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.355\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eborovecki\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eHuntington\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e22,283\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.452\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003echiaretti\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLeukemia\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e128\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12,625\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.422\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003echin\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBreast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e118\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e22,215\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.364\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003echowdary\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBreast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e104\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e22,283\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.404\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGLI_85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGliomas\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e22,283\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.25\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.306\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003egordon\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLung Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e181\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12,533\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.82\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.171\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003egravier\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBreast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e168\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2,905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.339\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003epomeroy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eCNS Tumor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7,128\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.35\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eProstate_GE\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eProstate Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e102\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5,966\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e12\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eshipp\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLymphoma\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7,129\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.247\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003esingh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eProstate Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e102\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12,600\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e14\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSMK_CAN_187\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eLung cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e187\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e19,993\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.89\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.481\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e15\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003esubramanian\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eN/A\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10,100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.34\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e16\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003etian\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMyeloma\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e173\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e12,625\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.208\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ewest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eBreast Cancer\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e7,129\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e2.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003earcene\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e200\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e10,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e1.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.44\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003egisette\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e7,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5,000\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e20\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHill_valley\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1,212\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e100\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.65\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.495\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eionosphere\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e351\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e33\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.359\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emadelon\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e2,600\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e500\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e23\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003esonar\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e208\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.77\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.466\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ewdbc\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e569\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e30\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.54\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0.373\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e* Dimensionality index\u0026thinsp;=\u0026thinsp;log(\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Number\\:of\\:features\\)\u003c/span\u003e\u003c/span\u003e)/log(\u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:Number\\:of\\:instances\\)\u003c/span\u003e\u003c/span\u003e). It is a measure of how high-dimensional a given dataset is.\u003c/p\u003e \u003cp\u003e** Balanced ratio: the proportion of the samples in lower class of the entire dataset. A value 0.5 is ideal for a binary class dataset.\u003c/p\u003e \u003cp\u003eCompared feature selection methods and experimental environment\u003c/p\u003e \u003cp\u003eTo assess the efficacy of the proposed method, we compared it with various established feature selection methods, as presented in Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eList of benchmark feature selection methods.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNo.\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFeature selection method\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eF-statistic\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFilter method.\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emRMR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFilter method that incorporates redundancy measurement\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePermutation importance\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eFilter method that considers feature interactions\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eInfinite feature selection (INF)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGraph-based filtering method [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eGRACES\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGraph convolutional network-based feature selection [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eAmong these, the first four adopted a filter-based approach. For each technique, feature rankings were computed and the top 100 features were selected. In the case of GRACES, a fixed number of features were produced. Hence, to maintain parity with the number of features selected by the other methods, we set this count to 100 and used the resulting rankings. The rankings for both INF and GRACES were determined after parameter optimization using 5-fold cross validation.\u003c/p\u003e \u003cp\u003eDirect comparisons between any filter method and the proposed approach are challenging because filter methods do not explicitly indicate a subset of features that yield optimal performance. Consequently, it is imperative to identify the optimal feature subsets for benchmarking methods. For this purpose, we employed two distinct strategies: \u0026lsquo;simple sequential search\u0026rsquo; and \u0026lsquo;forward search.\u0026rsquo; In simple sequential search, the performance of the classification model was evaluated at each step while increasing the number of features from 1 to 100 based on the ranking of each feature, and the subset of features with the highest performance was selected. Forward search is similar to simple sequential search, but it continues selecting good features until no further improvement in performance is observed.\u003c/p\u003e \u003cp\u003eTo evaluate the quality of the selected features, we applied the eXtreme Gradient Boosting (XGB) classifier, equipped with specific parameters, nrounds\u0026thinsp;=\u0026thinsp;5, objective = \"multi:softmax\", while retaining default configurations for the remaining parameters. Owing to its established potency and computational efficiency, XGB is an appropriate choice for evaluating feature selection methods.\u003c/p\u003e \u003cp\u003eFor the performance evaluation, we incorporated two metrics: Area Under the Curve (AUC) and LCM. AUC is a useful measure for comparing the performance of prediction models, especially when the classes of the dataset are imbalanced. As explained previously, LCM is suitable for evaluating both the classification performance and the number of selected features in the HDLSS dataset. For our analysis, a θ value of 0.8 was adopted as the reference point. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e illustrates the overall process of the comparative experiments.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eTo orchestrate the efficacious experiments, we created two computational environments, R and Python, as shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. Although the rankings for GRACE and INF were procured within a Python framework, the search and evaluation phases for all the techniques were uniformly executed in the R environment. The associated experimental code is available at \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://bitldku.github.io/home/sw/heuristic_search.html\u003c/span\u003e\u003cspan address=\"https://bitldku.github.io/home/sw/heuristic_search.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eExperimental environment.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eTask\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHardware\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eOS\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eLanguage\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003ePackage (Library)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003e- F-statistics\u003c/p\u003e \u003cp\u003e- mRMR\u003c/p\u003e \u003cp\u003e-Permutation\u003c/p\u003e \u003cp\u003eimportance\u003c/p\u003e \u003cp\u003e- Proposed\u003c/p\u003e \u003cp\u003ePerformance\u003c/p\u003e \u003cp\u003eevaluation\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eProcessor:AMD Ryzen 9 5900X 12-Core\u003c/p\u003e \u003cp\u003eRAM:16GB,\u003c/p\u003e \u003cp\u003eGPU:RTX3090\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eWindows 11\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"5\" rowspan=\"6\"\u003e \u003cp\u003eR 4.3.1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003emlr 2.19.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eranger 0.15.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003emRMRe 2.1.2.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eParallel (base)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003exgboost 1.7.5.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003ecaret 6.0\u0026ndash;94\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"8\" rowspan=\"9\"\u003e \u003cp\u003e- GRACES\u003c/p\u003e \u003cp\u003e- INF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\" morerows=\"8\" rowspan=\"9\"\u003e \u003cp\u003eProcessor:Intel i7\u0026thinsp;=\u0026thinsp;10700 16-Core\u003c/p\u003e \u003cp\u003eRAM:32.0GB GPU:RTX3080\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\" morerows=\"8\" rowspan=\"9\"\u003e \u003cp\u003eUbuntu 20.04.6 LTS\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\" morerows=\"8\" rowspan=\"9\"\u003e \u003cp\u003ePython 3.8.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003epytorch 2.0.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003etorch_geometric 2.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003esklearn 1.2.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003exgboost 1.7.6\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003enumpy 1.24.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003epandas 2.0.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003escipy 1.10.1\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eINF(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pypi.org/project/PyIFS/\u003c/span\u003e\u003cspan address=\"https://pypi.org/project/PyIFS/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003eGRACES(\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/canc1993/graces\u003c/span\u003e\u003cspan address=\"https://github.com/canc1993/graces\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS","content":"\u003cp\u003eNumber of selected features through GPF and HTS\u003c/p\u003e \u003cp\u003eThe number of selected features is an important criterion for feature selection in the HDLSS datasets. The proposed method comprised two procedures for feature selection: GPF and HTS. The GPF filters useless features, and the HTS chooses informative features from GPF results. Table\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e summarizes the results of the GPF and HTS analyses. Details are provided in Supplementary Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. During the GPF process, approximately 98.2% of the total features were eliminated from the 18 HDLSS datasets. Conversely, in the six non-HDLSS datasets, the average feature reduction rate was 61.2%, indicating considerable variability among the datasets. Notably, datasets with fewer than 100 features were not filtered. In the HTS phase, an additional 97.6% of the previously reduced features in the HDLSS dataset were excluded. The non-HDLSS datasets also exhibited a significant elimination rate, with an average of 98.3%. In conclusion, after the integrated application of both the GPF and HTS methodologies, the HDLSS datasets consistently selected 12 or fewer features, representing an average of merely 0.07% of the initial feature set.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eAverage number of features determined by proposed GPF and HTS procedures.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"6\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDataset\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eOriginal(A)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eGPF(B)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHTS(C)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eReduction Rate\u003c/p\u003e \u003cp\u003e(1-B/A)\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eSelection Rate\u003c/p\u003e \u003cp\u003e(C/A)\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eHDLSS (1\u0026ndash;18)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e12,163.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e222.6\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.982\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0004\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eNon-HDLSS (19\u0026ndash;24)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e954.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e370.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.612\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e \u003cp\u003e0.0066\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003ePerformance comparison based on simple sequential search\u003c/p\u003e \u003cp\u003eIn this section, we compare the classification performance the previous and proposed methods using the selected features. To determine the performance of the filter method, a simple sequential search was performed as described in the \u0026ldquo;Compared feature selection methods\u0026rdquo; section. AUC and LCM were used as performance criteria. A summary of the comparison results is presented in Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e. Details are provided in Supplementary Table S2 and S3.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the number of selected features, AUC, and LCM between the proposed method and previous works using simple sequential search.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSelected\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLCM\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e \u003cp\u003eIn AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c8\" namest=\"c7\"\u003e \u003cp\u003eIn LCM\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003efeatures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTop\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eRanking\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eTop\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRanking\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eF-statistics (FS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e32.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.891\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.833\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e3.2\u0026thinsp;\u0026plusmn;\u0026thinsp;1.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e3.5\u0026thinsp;\u0026plusmn;\u0026thinsp;1.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emRMR (MR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e27.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.874\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.824\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e2.9\u0026thinsp;\u0026plusmn;\u0026thinsp;1.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e3.0\u0026thinsp;\u0026plusmn;\u0026thinsp;1.4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePermutation (PE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e21.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.855\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e2.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.9\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e2.5\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eINF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e55.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.796\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.731\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e5.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e5.2\u0026thinsp;\u0026plusmn;\u0026thinsp;0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGRACES (GE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e51.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.816\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.748\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e5.1\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e5.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eProposed (PRO)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e5.5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.927\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.900\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e16\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e1.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.3\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e\u003cb\u003e20\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e1.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eWhen compared using the AUC performance metric, the proposed method exhibited equivalent or superior performance in 16 out of 24 datasets, while maintaining the lowest performance variability among all methods, as detailed in Supplementary Table S2 and S3. The AUC of our method exceeded that of comparative approaches, with margins ranging from 2.2\u0026ndash;13.1%. However, the number of selected features was markedly reduced, averaging only 5.5 features, which is between 1/4 and 1/10 that of the other methods selected.\u003c/p\u003e \u003cp\u003eUtilizing the LCM performance metric with a θ value of 0.8, our method displayed enhanced performance in 20 out of 24 datasets, representing 83.3%. This indicates that when considering both the performance and influence of the number of selected features, the proposed method excels, highlighting its ability to achieve comparable or better results with fewer features.\u003c/p\u003e \u003cp\u003eAmong the traditional methods, permutation exhibits the best performance, followed by F-statistics, mRMR, GRACES, and INF. Notably, GRACES and INF exhibited a tendency to select a greater number of features than the other methods. Furthermore, INF was unable to produce results for datasets exceeding 20,000 features, and GRACES also encountered difficulties in generating results for specific datasets.\u003c/p\u003e \u003cp\u003ePerformance comparison based on forward search\u003c/p\u003e \u003cp\u003eIn this section, we compare the classification performances of previous methods and the proposed method. To determine the performance of the filter method, a forward and a simple sequential search was used as described in the \u0026ldquo;Compared feature selection methods\u0026rdquo; section. A summary of the results is presented in Table\u0026nbsp;\u003cspan refid=\"Tab7\" class=\"InternalRef\"\u003e7\u003c/span\u003e, and the complete experimental results are detailed in Supplementary Table S4 and S5.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab7\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eComparison of the number of selected features, AUC, and LCM between the proposed method and previous works using forward search.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"8\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\"\u0026plusmn;\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eMethod\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eSelectedfeatures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eAUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eLCM\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e \u003cp\u003eIn AUC\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c8\" namest=\"c7\"\u003e \u003cp\u003eIn LCM\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eTop\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003eRanking\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003eTop\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003eRanking\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eF-statistics (FS)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.876\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e2.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e2.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003emRMR (MR)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.868\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.849\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e3.4\u0026thinsp;\u0026plusmn;\u0026thinsp;1.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e3.3\u0026thinsp;\u0026plusmn;\u0026thinsp;1.4\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePermutation (PE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5.8\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.905\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.877\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e2.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e2\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e2.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eINF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e6.7\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.819\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.806\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e4.1\u0026thinsp;\u0026plusmn;\u0026thinsp;1.4\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e4.4\u0026thinsp;\u0026plusmn;\u0026thinsp;1.2\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGRACES (GE)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e5.3\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.816\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.811\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e1\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e4.5\u0026thinsp;\u0026plusmn;\u0026thinsp;1.5\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e4.7\u0026thinsp;\u0026plusmn;\u0026thinsp;1.3\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eProposed (PRO)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e\u003cb\u003e5.5\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e\u003cb\u003e0.927\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e\u003cb\u003e0.900\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e\u003cb\u003e20\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c6\"\u003e \u003cp\u003e\u003cb\u003e1.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e \u003cp\u003e21\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\"\u0026plusmn;\" colname=\"c8\"\u003e \u003cp\u003e\u003cb\u003e1.3\u0026thinsp;\u0026plusmn;\u0026thinsp;0.7\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eCompared to the results in previous section, most filter methods showed an improvement in the AUC and significantly reduced the number of selected features. This trend suggests that the top 100 features chosen using traditional filtering methods exhibit high redundancy. A forward search helped remove many overlapping features. In specific datasets subh as Alon, both the feature count and AUC decreased, which might indicate a fall into local optima.\u003c/p\u003e \u003cp\u003eUsing the AUC metric, the proposed method matched or exceeded the performance of the other methods on 20 of the 24 datasets (83.3%). It also exhibited the lowest variability in performance. The AUC of the proposed method was between 2.2% and 11.1% higher than those of the other methods. Although the comparative methods selected fewer features on average than the sequence results, the proposed method consistently selected the least number of features. Furthermore, when using the LCM with θ at 0.8, the proposed method outperformed other methods in 21 of the 24 datasets (87.5%).\u003c/p\u003e \u003cp\u003eExecution time of the proposed method\u003c/p\u003e \u003cp\u003eSupplementary Table S6 lists the average runtimes of our method over 30 iterations for the HDLSS datasets. The Arcene dataset required approximately 12 min, the longest duration, whereas 16 out of the 18 datasets completed the feature selection within 5 min. Notably, the HTS phase constituted 77% of the total duration, with 94% spiking in some cases. The execution times for both GPF and HTS increased proportionally with instance size. Conversely, apart from SMK_CAN_187, the GPF duration remained relatively consistent irrespective of the feature count, whereas HTS showed a pronounced rise as the features increased. In essence, the duration of the HTS phase significantly shapes the total runtime, with the efficiency of the GPF in transmitting a reduced feature set being pivotal to the HTS timeframe. In conclusion, the average execution time of the proposed method is 169 seconds (2.8 minutes), which is reasonable.\u003c/p\u003e"},{"header":"DISCUSSION AND CONCLUSIONS","content":"\u003cp\u003eThe essence of feature selection lies in effectively finding a subset that is close to the global optimum, where the relevance to the class is high and the redundancy between features is low. We proposed two methods, GPF for filtering and HTS for searching and verified their superior performance compared to those of other methods. Particularly, for microarray datasets with HDLSS properties, we selected a minimal subset of core features within a reasonable time. The advantages of the proposed method are summarized as follows:\u003c/p\u003e \u003cp\u003e \u003cstrong\u003eIncorporating feature interactions for relevance measurement in the filtering stage\u003c/strong\u003e \u003cp\u003eOh [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e] experimentally and theoretically proved that the importance of a feature can be decomposed into its intrinsic predictive power (feature power) and the effect arising when combined with other features (interaction). The importance of a feature not only lies in its intrinsic power but also varies depending on the surrounding features. Therefore, the relevance between the subset and class can be evaluated more accurately by considering feature interactions. As we progressively filtered out features with low importance in the filtering stage and then measured their importance within the refined subset, we were able to accurately measure the interactions based on this subset. Although the permutation importance method we used does not explicitly distinguish between feature power and interaction, it is clear that the interaction effect is reflected within the importance.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eElimination of redundancy among features during the search stage\u003c/strong\u003e \u003cp\u003eWhile basic permutation importance can account for interactions between features, it does not capture redundancy. Thus, there is a possibility that highly redundant features exist within the filtered subset. By contrast, the proposed HTS can exclude redundancy among the features. To incorporate a new feature into the existing subset in the HTS process, the performance should be improved. If a perfectly redundant feature exists, no performance improvement observed, which leads to its exclusion. Therefore, HTS can play a complementary role in removing redundant features present in GPF.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eClear threshold criteria in the filtering stage\u003c/strong\u003e \u003cp\u003eTraditional filter methods do not consider the best candidate features for the next selection stage. Therefore, the users randomly select the best \u003cem\u003en\u003c/em\u003e features. In other methods, the candidate features produced depend on threshold. Therefore, users must tune these parameters to obtain better candidate features. By contrast, the proposed method produces effective candidate features without parameter settings. For our filtering process, we utilized a fixed threshold value based on theoretical evidence. A permutation importance of zero indicates that the feature has no effect on the performance improvement [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. If a feature degrades the performance or causes a negative interaction, its importance becomes less than 0. Hence, setting the threshold to zero allows for a clear and intuitive filtering of features that contribute to performance enhancement.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eImprovement of computational efficiency using feature ranking\u003c/strong\u003e \u003cp\u003eIn the filtering stage, existing methods utilize feature evaluation values exclusively for the removal of weak features, whereas our proposed method harnesses these values to optimize subsequent computations. Specifically, SFFS [\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e] thoroughly explores the whole feature space, encompassing both selected and unselected features, to identify pairs that can overcome the nesting effect. However, by leveraging the feature-ranking information already obtained during the filtering stage, our approach constrained the search space. This not only enhances the computational efficiency during the search phase but also avoids performance degradation.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eExpended search space to get better features\u003c/strong\u003e \u003cp\u003eBeyond \u0026ldquo;consolation match\u0026rdquo;, we incorporated a semi-random approach to pinpoint the global optimum. At the start of the search, we considered multiple candidate features with high importance to alter the first-choice feature and selected the feature subset with the best performance. From our observations, the optimal feature set frequently emerged when the first-choice feature was not ranked as the top feature.\u003c/p\u003e \u003c/p\u003e \u003cp\u003e \u003cstrong\u003eSelection of a sufficiently good features for microarray datasets\u003c/strong\u003e \u003cp\u003eTo understand or diagnose complex diseases, it is essential to not only achieve high predictive (classification) performance but also identify pivotal biomarkers (features) [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. For example, a predictive accuracy of 0.82 with 20 features is more meaningful than an accuracy of 0.85 with 120 features. Although the proposed method does not always result in a better predictive performance than the other methods, the derived features are very compact with a decent prediction performance. Therefore, it would be helpful to identify novel genes associated with diseases.\u003c/p\u003e \u003c/p\u003e \u003cp\u003eIn our proposed method, during the consolation match, we specifically operated on the latter half of the selected features and the former half of the non-selected features. Unexpectedly, when we expanded this search range during our experiments, a decline in the performance was observed. This phenomenon raises questions that require further exploration, particularly concerning feature interactions. This is a topic for further research.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eEthics approval and consent to participate\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003eConsent for publication\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003eAvailability of data and materials\u003c/p\u003e\n\u003cp\u003eCode about the results of this study are publicly shared at: https://bitldku.github.io/home/sw/heuristic_search.html.\u003c/p\u003e\n\u003cp\u003eCompeting interests\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003eFunding\u003c/p\u003e\n\u003cp\u003eThis work was supported by the Institute for Information \u0026amp; Communications Technology Planning \u0026amp; Evaluation (IITP) grand funded by the Ministry of Science, ICT (MSIT), Korea (No. RS-2023-00222191, Development of data fabric technology to support logical data integration and compound analysis of distributed data)\u003c/p\u003e\n\u003cp\u003eAuthors\u0026apos; contributions\u003c/p\u003e\n\u003cp\u003eShin and Oh conceptualized the research. Shin contributed to development of the model. Shin collected and annotated data. All authors read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003eAcknowledgements\u003c/p\u003e\n\u003cp\u003eNone\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAng JC, Mirzal A, Haron H, Hamed HNA. Supervised, Unsupervised, and Semi-Supervised Feature Selection: A Review on Gene Selection. IEEE/ACM Trans Comput Biol Bioinform. 2016;13(5):971\u0026ndash;89. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TCBB.2015.2478454\u003c/span\u003e\u003cspan address=\"10.1109/TCBB.2015.2478454\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLu H, Chen J, Yan K, Jin Q, Xue Y, Gao Z. A hybrid feature selection algorithm for gene expression data classification. Neurocomputing. 2017;256:56\u0026ndash;62. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.neucom.2016.07.080\u003c/span\u003e\u003cspan address=\"10.1016/j.neucom.2016.07.080\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlmugren N, Alshamlan H. A Survey on Hybrid Feature Selection Methods in Microarray Gene Expression Data for Cancer Classification. IEEE Access. 2019;7:78533\u0026ndash;78548 doi:0.1109/ACCESS.2019.2922987.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBommert A, Sun X, Bischl B, Rahnenf\u0026uuml;hrer J, Lang M. Benchmark for filter methods for feature selection in high-dimensional classification data. Comput Stat Data Anal. 2020;143:106839. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.csda.2019.106839\u003c/span\u003e\u003cspan address=\"10.1016/j.csda.2019.106839\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eManikandan G, Abirami S. Feature Selection Is Important: State-of-the-Art Methods and Application Domains of Feature Selection on High-Dimensional Data. In: Kumar R, Paiva S, editors. Applications in Ubiquitous Computing. Cham: Springer; 2021. pp. 177\u0026ndash;96.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen C, Weiss ST, Liu YY. Graph convolutional network-based feature selection for high-dimensional and low-sample size data. Bioinformatics. 2023;39(4):btad135. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1093/bioinformatics/btad135\u003c/span\u003e\u003cspan address=\"10.1093/bioinformatics/btad135\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlhenawi E, Al-Sayyed R, Hudaib A, Mirjalili S. Feature selection methods on gene expression microarray data for cancer classification: A systematic review. Comput Biol Med. 2022;140:105051. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.compbiomed.2021.105051\u003c/span\u003e\u003cspan address=\"10.1016/j.compbiomed.2021.105051\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu B, Wei Y, Zhang Y, Yang Q. Deep neural networks for high dimension, low sample size data. In: Proceedings of the 26th International Joint Conference on Artificial Intelligence. Melbourne: AAAI, 2017;2287\u0026ndash;2293.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi K, Wang F, Yang L, Liu R. Deep feature screening: Feature selection for ultra high-dimensional data via deep neural networks. Neurocomputing. 2023;538:126186. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.neucom.2023.03.047\u003c/span\u003e\u003cspan address=\"10.1016/j.neucom.2023.03.047\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRoffo G, Melzi S, Castellani U, Vinciarelli A, Cristani M. Infinite Feature Selection: A Graph-based Feature Filtering Approach. IEEE Trans Pattern Anal Mach Intell. 2020;43(12):4396\u0026ndash;410. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/TPAMI.2020.3002843\u003c/span\u003e\u003cspan address=\"10.1109/TPAMI.2020.3002843\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRostami M, Forouzandeh S, Berahmand K, Soltani M, Shahsavari M, Oussalah M. Gene selection for microarray data classification via multi-objective graph theoretic-based method. Artif Intell Med. 2022;123:102228. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.artmed.2021.102228\u003c/span\u003e\u003cspan address=\"10.1016/j.artmed.2021.102228\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePati SK, Banerjee A, Manna S. Gene selection of microarray data using Heatmap Analysis and Graph Neural Network. Appl Soft Comput. 2023;135:110034. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.asoc.2023.110034\u003c/span\u003e\u003cspan address=\"10.1016/j.asoc.2023.110034\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi Y, Chen CY, Wasserman WW. Deep Feature Selection: Theory and Application to Identify Enhancers and Promoters. J Comput Biol. 2016;23(5):322\u0026ndash;36. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1089/cmb.2015.0189\u003c/span\u003e\u003cspan address=\"10.1089/cmb.2015.0189\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChowdhury S, Dong X, Li X. Recurrent Neural Network Based Feature Selection for High Dimensional and Low Sample Size Micro-array Data. In: Proceedings of 2019 IEEE International Conference on Big Data. Piscataway: IEEE, 2019;4823\u0026ndash;4828.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAgrawal P, Abutarboush HF, Ganesh T, Mohamed AW. Metaheuristic Algorithms on Feature Selection: A Survey of One Decade of Research (2009\u0026ndash;2019). IEEE Access. 2021;9:26766\u0026ndash;91. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1109/ACCESS.2021.3056407\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2021.3056407\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDokeroglu T, Deniz A, Kiziloz HE. A comprehensive survey on recent metaheuristics for feature selection. Neurocomputing. 2022;494:269\u0026ndash;96. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.neucom.2022.04.083\u003c/span\u003e\u003cspan address=\"10.1016/j.neucom.2022.04.083\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFerri FJ, Pudil P, Hatef M, Kittler J. Comparative study of techniques for large-scale feature selection. Mach Intell Pattern Recognit. 1994;16:403\u0026ndash;13. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/B978-0-444-81892-8.50040-7\u003c/span\u003e\u003cspan address=\"10.1016/B978-0-444-81892-8.50040-7\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang L, Wang Y, Chang Q. Feature selection methods for big data bioinformatics: A survey from the search perspective. Methods. 2016;111:21\u0026ndash;31. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ymeth.2016.08.014\u003c/span\u003e\u003cspan address=\"10.1016/j.ymeth.2016.08.014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHsu HH, Hsieh CW, Lu MD. Hybrid feature selection by combining filters and wrappers. Expert Syst Appl. 2011;38(7):8144\u0026ndash;50. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.eswa.2010.12.156\u003c/span\u003e\u003cspan address=\"10.1016/j.eswa.2010.12.156\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRamey J, Datamicroarray. Jan. \u003cem\u003eAvailable at\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/ramhiser/datamicroarray\u003c/span\u003e\u003cspan address=\"https://github.com/ramhiser/datamicroarray\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 19 2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArizona State University (ASU). Feature selection datasets. \u003cem\u003eAvailable at\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://jundongl.github.io/scikit-feature/datasets.html\u003c/span\u003e\u003cspan address=\"https://jundongl.github.io/scikit-feature/datasets.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 19 Jan. 2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOpenML. A worldwide machine learning lab. \u003cem\u003eAvailable at\u003c/em\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.openml.org/\u003c/span\u003e\u003cspan address=\"https://www.openml.org/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e (accessed 19 May 2024).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOh S. Predictive case-based feature importance and interaction. Inf Sci. 2022;593:155\u0026ndash;76. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/j.ins.2022.02.003\u003c/span\u003e\u003cspan address=\"10.1016/j.ins.2022.02.003\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMi X, Zou B, Zou F, Hu J. Permutation-based identification of important biomarkers for complex diseases via machine learning models. Nat Commun. 2021;12(1):3008. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1038/s41467-021-22756-2\u003c/span\u003e\u003cspan address=\"10.1038/s41467-021-22756-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePudil P, Novovičov\u0026aacute; J, Kittler J. Floating search methods in feature selection. Pattern Recognit Lett. 1994;15(11):1119\u0026ndash;25. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003e10.1016/0167-8655(94)90127-9\u003c/span\u003e\u003cspan address=\"10.1016/0167-8655(94)90127-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"bmc-bioinformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"binf","sideBox":"Learn more about [BMC Bioinformatics](http://bmcbioinformatics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/binf","title":"BMC Bioinformatics","twitterHandle":"@BMC_Bioinformatics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"HDLSS, feature selection, machine learning, filter method, wrapper method","lastPublishedDoi":"10.21203/rs.3.rs-5260669/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-5260669/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003ch2\u003eBackground.\u003c/h2\u003e \u003cp\u003eHigh-dimensional datasets with low sample sizes (HDLSS) are pivotal in the fields of biology and bioinformatics. One of core objective of HDLSS is to select most informative features and discarding redundant or irrelevant features. This is particularly crucial in bioinformatics, where accurate feature (gene) selection can lead to breakthroughs in drug development and provide insights into disease diagnostics. Despite its importance, identifying optimal features is still a significant challenge in HDLSS.\u003c/p\u003e\u003ch2\u003eResults.\u003c/h2\u003e \u003cp\u003eTo address this challenge, we propose an effective feature selection method that combines gradual permutation filtering with a heuristic tribrid search strategy, specifically tailored for HDLSS contexts. The proposed method considers inter-feature interactions and leverages feature rankings during the search process. In addition, a new performance metric for the HDLSS that evaluates both the number and quality of selected features is suggested. Through the comparison of the benchmark dataset with existing methods, the proposed method reduced the average number of selected features from 37.8 to 5.5 and improved the performance of the prediction model, based on the selected features, from 0.855 to 0.927.\u003c/p\u003e\u003ch2\u003eConclusions.\u003c/h2\u003e \u003cp\u003eThe proposed method effectively selects a small number of important features and achieves high prediction performance.\u003c/p\u003e","manuscriptTitle":"An effective heuristic for developing hybrid feature selection in high dimensional and low sample size datasets","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-12-16 09:24:14","doi":"10.21203/rs.3.rs-5260669/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2024-11-29T10:19:56+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-11-01T03:16:53+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2024-10-23T08:47:07+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"52851016529810545650952971982095646875","date":"2024-10-23T07:46:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"64317500677041489081341771275450051647","date":"2024-10-23T02:26:18+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"161468431402112266793845386208741663778","date":"2024-10-22T12:22:47+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"39655419753631800884119271283256597890","date":"2024-10-22T12:05:16+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-10-22T11:22:44+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2024-10-17T10:55:06+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-10-16T08:54:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2024-10-16T00:13:52+00:00","index":"","fulltext":""},{"type":"submitted","content":"BMC Bioinformatics","date":"2024-10-14T11:07:03+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"bmc-bioinformatics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"binf","sideBox":"Learn more about [BMC Bioinformatics](http://bmcbioinformatics.biomedcentral.com/)","snPcode":"","submissionUrl":"https://www.editorialmanager.com/binf","title":"BMC Bioinformatics","twitterHandle":"@BMC_Bioinformatics","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"em","reportingPortfolio":"BMC Series","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"d68ad2cd-cff1-40bb-be23-dd495a017239","owner":[],"postedDate":"December 16th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-12-30T16:04:12+00:00","versionOfRecord":{"articleIdentity":"rs-5260669","link":"https://doi.org/10.1186/s12859-024-06017-9","journal":{"identity":"bmc-bioinformatics","isVorOnly":false,"title":"BMC Bioinformatics"},"publishedOn":"2024-12-26 15:57:13","publishedOnDateReadable":"December 26th, 2024"},"versionCreatedAt":"2024-12-16 09:24:14","video":"","vorDoi":"10.1186/s12859-024-06017-9","vorDoiUrl":"https://doi.org/10.1186/s12859-024-06017-9","workflowStages":[]},"version":"v1","identity":"rs-5260669","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-5260669","identity":"rs-5260669","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2024) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00