Comparison of dimensionality reduction and feature selection for cognitive task decoding in functional magnetic resonance imaging

preprint OA: closed
📄 Open PDF Full text JSON View at publisher

Abstract

Background Advances in functional magnetic resonance imaging (fMRI) have led to the ability to study the brain across many contexts. However, the large number of features generated by functional connectivity approaches may overfit the data. These problems can be overcome with either feature selection (FS) or dimensionality reduction (DR), which can be applied to less complex models. We utilize two open-source datasets to compare the performance of DR/FS methods on cognitive task decoding using a suite of ML classifiers. New Method While DR and FS methods have been used previously in decoding research, no systematic comparison of their performance has been undertaken. Here, we compare available methods using commonly utilized machine learning libraries to establish which methods provide the best predictive performance. We then conduct statistical tests to examine the relative contributions of DR and FS methods and classifiers on decoding accuracy. Results Neither DR or FS was found to be superior. However, differences were identified across datasets and tasks. In the majority of methods and datasets, a peak in predictive performance was found using a small percentage of the total number of original features. Comparison with existing methods Some methods perform better than the baseline method of prediction with all available features or selecting features randomly. Decoding performance utilizing the HCP datasets with certain DR/FS methods exceeds that of deep learning approaches. Conclusions Simple machine learning models with DR/FS have competitive decoding performance. These results suggest a “sweet spot” for the tradeoff between the retention of features and predictive accuracy.
Full text 64,257 characters · extracted from preprint-html · click to expand
Comparison of dimensionality reduction and feature selection for cognitive task decoding in functional magnetic resonance imaging | bioRxiv /* */ /* */ <!-- <!-- /*! * yepnope1.5.4 * (c) WTFPL, GPLv2 */ (function(a,b,c){function d(a){return"[object Function]"==o.call(a)}function e(a){return"string"==typeof a}function f(){}function g(a){return!a||"loaded"==a||"complete"==a||"uninitialized"==a}function h(){var a=p.shift();q=1,a?a.t?m(function(){("c"==a.t?B.injectCss:B.injectJs)(a.s,0,a.a,a.x,a.e,1)},0):(a(),h()):q=0}function i(a,c,d,e,f,i,j){function k(b){if(!o&&g(l.readyState)&&(u.r=o=1,!q&&h(),l.onload=l.onreadystatechange=null,b)){"img"!=a&&m(function(){t.removeChild(l)},50);for(var d in y[c])y[c].hasOwnProperty(d)&&y[c][d].onload()}}var j=j||B.errorTimeout,l=b.createElement(a),o=0,r=0,u={t:d,s:c,e:f,a:i,x:j};1===y[c]&&(r=1,y[c]=[]),"object"==a?l.data=c:(l.src=c,l.type=a),l.width=l.height="0",l.onerror=l.onload=l.onreadystatechange=function(){k.call(this,r)},p.splice(e,0,u),"img"!=a&&(r||2===y[c]?(t.insertBefore(l,s?null:n),m(k,j)):y[c].push(l))}function j(a,b,c,d,f){return q=0,b=b||"j",e(a)?i("c"==b?v:u,a,b,this.i++,c,d,f):(p.splice(this.i++,0,a),1==p.length&&h()),this}function k(){var a=B;return a.loader={load:j,i:0},a}var l=b.documentElement,m=a.setTimeout,n=b.getElementsByTagName("script")[0],o={}.toString,p=[],q=0,r="MozAppearance"in l.style,s=r&&!!b.createRange().compareNode,t=s?l:n.parentNode,l=a.opera&&"[object Opera]"==o.call(a.opera),l=!!b.attachEvent&&!l,u=r?"object":l?"script":"img",v=l?"script":u,w=Array.isArray||function(a){return"[object Array]"==o.call(a)},x=[],y={},z={timeout:function(a,b){return b.length&&(a.timeout=b[0]),a}},A,B;B=function(a){function b(a){var a=a.split("!"),b=x.length,c=a.pop(),d=a.length,c={url:c,origUrl:c,prefixes:a},e,f,g;for(f=0;f<d;f++)g=a[f].split("="),(e=z[g.shift()])&&(c=e(c,g));for(f=0;f<b;f++)c=x[f](c);return c}function g(a,e,f,g,h){var i=b(a),j=i.autoCallback;i.url.split(".").pop().split("?").shift(),i.bypass||(e&&(e=d(e)?e:e[a]||e[g]||e[a.split("/").pop().split("?")[0]]),i.instead?i.instead(a,e,f,g,h):(y[i.url]?i.noexec=!0:y[i.url]=1,f.load(i.url,i.forceCSS||!i.forceJS&&"css"==i.url.split(".").pop().split("?").shift()?"c":c,i.noexec,i.attrs,i.timeout),(d(e)||d(j))&&f.load(function(){k(),e&&e(i.origUrl,h,g),j&&j(i.origUrl,h,g),y[i.url]=2})))}function h(a,b){function c(a,c){if(a){if(e(a))c||(j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}),g(a,j,b,0,h);else if(Object(a)===a)for(n in m=function(){var b=0,c;for(c in a)a.hasOwnProperty(c)&&b++;return b}(),a)a.hasOwnProperty(n)&&(!c&&!--m&&(d(j)?j=function(){var a=[].slice.call(arguments);k.apply(this,a),l()}:j[n]=function(a){return function(){var b=[].slice.call(arguments);a&&a.apply(this,b),l()}}(k[n])),g(a[n],j,b,n,h))}else!c&&l()}var h=!!a.test,i=a.load||a.both,j=a.callback||f,k=j,l=a.complete||f,m,n;c(h?a.yep:a.nope,!!i),i&&c(i)}var i,j,l=this.yepnope.loader;if(e(a))g(a,0,l,0);else if(w(a))for(i=0;i (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0];var j=d.createElement(s);var dl=l!='dataLayer'?'&l='+l:'';j.src='//www.googletagmanager.com/gtm.js?id='+i+dl;j.type='text/javascript';j.async=true;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-M677548'); Skip to main content Home About Submit ALERTS / RSS Search for this keyword Advanced Search New Results Comparison of dimensionality reduction and feature selection for cognitive task decoding in functional magnetic resonance imaging View ORCID Profile Corey J. Richier , View ORCID Profile Kyle A. Baacke , View ORCID Profile Sarah M. Olshan , View ORCID Profile Wendy Heller doi: https://doi.org/10.1101/2025.10.27.684660 Corey J. Richier a University of Illinois Urbana-Champaign, Department of Psychology , 603 E Daniel St, Champaign, IL, 61820 Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Corey J. Richier For correspondence: coreyjr2{at}illinois.edu kbaacke{at}uwm.edu Kyle A. Baacke b University of Wisconsin-Milwaukee, Department of Psychology , 2441 E. Hartford Ave., Milwaukee, WI 53211 Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Kyle A. Baacke For correspondence: coreyjr2{at}illinois.edu kbaacke{at}uwm.edu Sarah M. Olshan a University of Illinois Urbana-Champaign, Department of Psychology , 603 E Daniel St, Champaign, IL, 61820 Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Sarah M. Olshan Wendy Heller a University of Illinois Urbana-Champaign, Department of Psychology , 603 E Daniel St, Champaign, IL, 61820 Find this author on Google Scholar Find this author on PubMed Search for this author on this site ORCID record for Wendy Heller Abstract Full Text Info/History Metrics Supplementary material Preview PDF Abstract Background Advances in functional magnetic resonance imaging (fMRI) have led to the ability to study the brain across many contexts. However, the large number of features generated by functional connectivity approaches may overfit the data. These problems can be overcome with either feature selection (FS) or dimensionality reduction (DR), which can be applied to less complex models. We utilize two open-source datasets to compare the performance of DR/FS methods on cognitive task decoding using a suite of ML classifiers. New Method While DR and FS methods have been used previously in decoding research, no systematic comparison of their performance has been undertaken. Here, we compare available methods using commonly utilized machine learning libraries to establish which methods provide the best predictive performance. We then conduct statistical tests to examine the relative contributions of DR and FS methods and classifiers on decoding accuracy. Results Neither DR or FS was found to be superior. However, differences were identified across datasets and tasks. In the majority of methods and datasets, a peak in predictive performance was found using a small percentage of the total number of original features. Comparison with existing methods Some methods perform better than the baseline method of prediction with all available features or selecting features randomly. Decoding performance utilizing the HCP datasets with certain DR/FS methods exceeds that of deep learning approaches. Conclusions Simple machine learning models with DR/FS have competitive decoding performance. These results suggest a “sweet spot” for the tradeoff between the retention of features and predictive accuracy. 1. Introduction Advances in the analysis of functional magnetic resonance imaging (fMRI) data have led to the ability to densely sample the human brain during a variety of contexts. Functional connectivity (FC) methods have shown promise for establishing linkages between brain function, behaviors, and cognition ( Rogers et al., 2007 ). However, the large number of features generated by FC approaches (on the order of 10^3 to 10^7; Brown & Hamarneh, 2016 ) results in high dimensionality (known colloquially as the “curse of dimensionality”; Mwangi et al., 2014 ). A large number of features is undesirable due to the fact that the predictor variables can have high multicollinearity, thus reducing the interpretability of any given predictor ( Alin, 2010 ). In predictive modeling contexts, a large number of features that exceed the number of observations may lead to models overfitting the data ( Hua et al., 2009 ; Whelan & Garavan, 2014 ). This problem is generally overcome by employing either feature selection (FS) or dimensionality reduction (DR) methods. FS methods reduce the number of features by selecting particularly important or representative variables ( Hauskrecht et al., 2007 ; Li et al., 2018 ), whereas DR methods transform existing variables into fewer variables that capture the original data’s variance ( van der Maaten et al., 2009 ). Arguably, one of the most foundational goals of neuroscience is examining how patterns of neural activity relate to mental processes. Decoding approaches have become particularly useful for examining relationships between cognitive states and representations of neural activity associated with those states ( Poldrack, 2011 ), as they can allow for reverse inference (i.e., what is the behavioral state associated with a given pattern of neural activity). Decoding is flexible in its ability to be applied at various scales, from the level of cellular activity ( Glaser et al., 2020 ) to whole brain recordings ( Haynes & Rees, 2006 ), and is applicable for basic science ( Haxby et al., 2001 ; LaConte, 2011 ) and clinical research ( Rashid & Calhoun, 2020 ). Additionally, such methods have implications for brain computer interfaces ( Rao, 2019 ). Recent initiatives in decoding have sought to make models as flexible as possible ( Rubin et al., 2017 ) due to the fact that the nature of mental faculties are highly dimensional. Humans experience different stimuli and recruit a variety of different cognitive mechanisms to navigate the world. Effective decoding schemes therefore should be able to distinguish between as many different contexts and conditions as possible. Decoding schemes can be considered multi-class classification problems, in which the goal is to correctly classify the task associated with accompanying neural data. Given a large number of features in neural decoding approaches, methods to reduce the dimensionality of the input data may make these models more tractable. Thus, DR and FS approaches are suitable tools to improve classification accuracy. Some DR methods and FS methods have been previously applied to decoding analyses ( Mwangi et al., 2014 ). Among FS methods, random forest-based methods for feature selection have been found to outperform univariate methods (such as F-test score) for decoding approaches such as an object recognition task ( Genuer et al., 2010 ). For DR methods, principal component analysis (PCA) has been found to be effective for decoding hand trajectory using electroencephalography and found it to be a more efficient method compared to using the mean value for all regions of interest at a given time point ( Srisrisawang & Müller-Putz, 2022 ). DR and FS have been more formally compared in non-neuroimaging datasets. Anowar et al. (2021) compared multiple DR and FS methods across three different types of data modalities. In an electrocardiogram dataset, kernel PCA was found to perform best. However, in a mass spectroscopy dataset, it was found that laplacian eigenmaps were found to be the optimal method. Lastly, in a dataset for decoding mobility status from smartphone data, latent discriminant analysis was found to perform best. This emphasizes the need to explore the merits of different DR and FS methods in specific contexts, as the best- performing method depends on the structure of the input data. Although deep learning methods can be considered the state of the art for machine learning approaches in neuroscience, they are not always the most advantageous method, depending on the circumstances. Linear models, by merit of their structure, can offer advantages with respect to interpretability (Varoquax et al., 2017). Deep learning models are often critiqued for being “black box”, because it is not straightforward to interpret which features are used for classification. Knowing which brain features are implicated in certain cognitive processes, emotions, or disease states contributes not only to understanding basic mechanisms but has implications for applied utility (e.g., in the case of clinical disorders). Furthermore, more complex models like deep learning do not always perform better in supervised learning contexts, where simpler classifiers can perform comparably ( Rudin, 2019 ). While there have been great large-scale advances in deep learning (for example, commercial LLMs), there have been criticisms of the environmental impact of large-scale models ( Selvan et al., 2022 ). Some machine learning researchers have advocated that the efficiency of models is also a valid performance metric ( Schwartz et al., 2020 ). While more classic machine learning models are less complex, they are more computationally efficient than some of the latest large-scale architectures. Thus, it can be easier to run many simpler models at scale to test hyperparameters or research questions when many permutations of the data or models must be examined. Additionally, the scaling of deep learning research has resulted in certain institutions, including private corporations, having access to methods that are simply intractable for many researchers ( Besiroglu et al., 2024 ; Shu et al., 2024). A lack of access to computational resources that can employ such large-scale deep learning models on imaging data can lead to scientific and research inequity ( Nielsen & Andersen, 2021 ). Thus, determining inclusive machine learning methods for neuroscience, which operates using highly dimensional data, will be important for the continued participation of researchers around the globe, who may not always have access to state of the art computational resources ( Ahmed & Wahed, 2020 ; Probst, 2023 ). While DR and FS have been used in analyses of fMRI data, no formal comparisons have taken place to examine their efficacy in a decoding task. Taking inspiration from previous research in other fields ( Anowar et al., 2021 ), the present analysis sought to compare numerous DR and FS methods on task-based fMRI for how well they may improve classification accuracy. We compared decoding accuracy of cognitive tasks in multiple datasets using various DR and FS methods. We used seven DR/FS methods to reduce the input space generated by functional connectivity values from parcellated fMRI sessions in two publicly available datasets. We then used classifier models to decode tasks from the resulting input feature sets. The cross-validation accuracy of the resulting models was then used to evaluate their efficacy. Stepwise hierarchical regressions were then used to evaluate whether the classifier, DR/FS method, and interactions between classifier and method impacted model performance after accounting for the impact of the number of features used in each model. 2. Method 2.1 Datasets Data used in the present work were obtained from the Human Connectome Project (HCP) database ( Van Essen et al., 2012 ; https://db.humanconnectome.org/ ) and the UCLA Consortium for Neuropsychiatric Phenomics dataset ( Poldrack et al., 2016 ; obtained from https://openneuro.org/datasets/ds000030 , version 1.0). Data from the HCP project were preprocessed using the HCPPipelines BIDS App wrapper ( https://github.com/BIDS-Apps/HCPPipelines ; Glasser et al., 2013 ; Smith, 2018 ). Preprocessing for the UCLA dataset relied on fMRIPrep 20.2.0 ( Esteban et al., 2019 , 2023 ), which is based on Nipype 1.5.1 ( Esteban et al., 2022 ; Gorgolewski et al., 2011 ); RRID:SCR_002502). Further details about preprocessing methods are provided in Supplementary Materials. In HCP, the subjects completed tasks as follows: 1) a working memory task, as measured via an N-back paradigm, 2) a motor task, where participants moved their fingers, toes, or tongue, 3) a language processing task, where participants completed basic arithmetic and story comprehension questions, 4) a social cognition tasks, where participants were to infer theory of mind in a series of inanimate objects moving in a way that suggested interaction, 5) a relational processing task, where they were presented with different images and told to compare and contrast the characteristics of these images, 6) an emotional processing task, where they were asked to compare facial expressions across images and indicated whether a series of faces were the same or not, and 7) an incentive processing task, where they were presented with a risk/reward paradigm with potential to gain or lose money through a series of decisions. In the UCLA study, participants engaged in a 1) balloon analog risk task, where participants had to risk maximizing the amount of money acquired while risking losing what they had earned, a paired association memory task, with 2) an encoding phase and 3) a retrieval phase, 4) a spatial working memory task, 5) a stop signal task, where participants underwent a go no-go paradigm, 6) a task-switching task, where participants were asked to respond to different stimuli in changing sequences, 7) a breath-holding task, and 8) a resting state scan with eyes open. Further details of these tasks can be obtained in Barch et al. (2013) for HCP and Poldrack et al. (2016) for UCLA. 2.2 Data processing The non-smoothed task-fMRI data for each participant were parcellated using the NiftiLabelsMasker function in NiLearn v0.9.0 ( Abraham et al., 2014 ; RRID:SCR_001362), referencing the 200 parcel 7 network atlas generated by Schaefer et al. (2018) . The parcellation combines voxels that belong to specific brain regions. The parcellated timeseries were then concatenated for each subject for each task (RL and LR phase encoding). The mean of all voxel BOLD signals in each parcel was then subtracted from the values in the timeseries to center values for each parcel around 0. These values in each parcel were then Fisher-z transformed. After this, every parcel was correlated with one another using the Pearson correlation to generate functional connectivity matrices for each task session for each subject. The unique values from this matrix (i.e., the lower or upper diagonal) were extracted to serve as input data for the classifiers. This process was repeated for every subject’s available task data. Thus, the total number of observations from each data set represented the total number of unique task runs across subjects (such that each subject was represented in the data the number of times they completed a unique task). The final matrix was made up of columns representing the different connectivity values for each pair of parcels, and rows representing a unique task session. Each column of this data matrix was then standardized with a z - score transformation. 2.3 Feature selection and dimensionality reduction methods All dimensionality reduction and prediction generation steps were conducted with Python 3.6.15. The present DR and FS methods were selected because they can be fitted to a training set of data and then fitted on a separate test set. The full feature set (19,900) was used to compare the various DR and FS as a baseline. For another baseline comparison, random feature sets were selected from the full feature set. For each of the feature set sizes generated using the other DR/FS methods, 10 sets of features were randomly selected from the initial 19,900 features. 2.3.1 Feature selection methods 2.3.1.1 Hierarchical Clustering Hierarchical clustering (HC) was performed as follows using scipy v1.3.3 ( Virtanen et al., 2020 ; RRID:SCR_008058). Spearman correlations between variables were used as an input euclidean distance matrix to perform Ward’s linkage to identify the hierarchical structure of clusters by minimizing variance within each cluster. The resulting linkage matrix was then used to form flat clusters for each within-cluster cophenetic distance cap between 1 and 250. Unique cluster structures were saved for use in the prediction generation phase. 2.3.1.2 Permutation Importance Permutation importance methods constitute a broad class of techniques that randomly shuffle feature values to disrupt any systematic association between predictors and the outcome, using this permuted data to generate a null distribution of test statistics or p -values for comparison( Winkler et al., 2014 ). The permutation importance method employed in scikit-learn 0.23.2 (( Pedregosa et al., 2011 ); RRID:SCR_002577) permutes each variable in the dataset and evaluates the change in overall accuracy when the feature is permuted. Features that have a greater reduction in accuracy when they are permuted are deemed to be more important for successful prediction. 2.3.1.3 Select from model The select from model (SFM) method retrieves features from the model training process based on different importance metrics. For the random forest, the Gini impurity decrease is the metric used for establishing feature importance. In the support vector machine, model coefficients are selected by their magnitude. SFM was performed using scikit-learn v0.23.2. 2.3.2 Dimensionality reduction methods 2.3.2.1 Principal Component Analysis Principal component analysis (PCA) is a commonly used dimensionality reduction method that identifies orthogonal axes in the data, known as principal components. These axes capture the maximal amount of variance in the data. PCA projects data onto these axes, and these are used as new predictor variables. Therefore, fewer variables can be used to capture a greater amount of variation in the data. All components were retained for prediction generation. For all feature set sizes generated using hierarchical clustering, the same number of components was retained as input in the prediction generation step. PCA was performed using scikit- learn v0.23.2. 2.3.2.2 Kernel Principal Component Analysis Kernel PCA (kPCA) is a similar method to PCA that projects data to higher dimensional space to better account for potential non-linear relationships in the data. The present analysis utilized kPCA with linear and radial basis function kernels. kPCA was performed using scikit-learn 0.23.2. 2.3.2.3 Truncated Singular Value Decomposition Singular value decomposition (SVD) is a dimensionality reduction method that decomposes any matrix into three matrices that contain information about variation in the data. In truncated SVD (tSVD) only a specified number of singular values and singular vectors, which can be used as new predictor variables, are retained. tSVD was performed using scikit-learn 0.23.2. 2.3.2.4 Linear Discriminant Analysis Linear discriminant analysis (LDA) is a supervised learning method utilizing DR to find the optimal linear combination between components of the predictor. LDA will generate an equal number of predictor variables to the number of classes to predict. These variables attempt to maximize the variation between classes while minimizing the variation within the classes. 2.3.3 Classifiers Research has demonstrated that the choice of model can have an effect on performance in machine learning decoding analyses with fMRI data ( Jollans et al., 2019 ). Thus, the present study utilized three types of models with different architectures. A random forest classifier (RFC), a support vector machine (SVM), and ridge regression (RR) were utilized. All three models were given data from all of the aforementioned DR and FS methods. A grid search was utilized to examine the best combination of hyperparameters for each of the different classifiers. For the SVM, all models were instantiated with a linear kernel, a squared hinge loss function, and 100,000 max iterations. Values of C were iterated over in the grid search (0.1, 1, 10). For the RFC, the number of estimators were grid searched over 100 and 200, and the minimal samples for leafing were set at 2 and 5. For the RR, alphas were set to .001, .01, .1, 1, and 10. 2.3.4 Training and testing procedure Ninety percent of the data was initially selected for the training procedure. A group k-fold method was used such that no subjects appeared in both the training and testing data, as this might result in inflated estimates of testing set accuracy (Varoquax et al., 2017). This procedure entailed a further 90/10 split ten-fold cross validation to obtain the best features for predictive accuracy. These features were then applied to the held out 10% of the data and evaluated for accuracy. Each classifier was cross validated with each feature set size and DR/FS. A group k-fold with five splits was used for training (roughly 18% was used as the testing data). 3. Results 3.1 Performance by dataset, classifier and DR/FS method Figure 1 demonstrates a performance curve of each method in the HCP dataset, and Figure 2 demonstrates the same in the UCLA dataset. The curves were generated by comparing each method to selecting features at random of the same distribution of features sizes (shown as grey dots) to serve as a benchmark for each method. Several trends were evident based on the figure. Many methods were at their highest classification accuracy between the order of 10-1,000 features. These methods seem to taper with respect to classification performance when the number of features becomes large enough. However, some methods did not perform well at any size. Performance of the SVM in the HCP dataset was poor over most of the possible feature sizes, except for a larger number of features in SFM, RFPI, and HC. For tSVD, PCA, and kPCA, most numbers of features result in a very low degree of accuracy, suggesting that these methods did not capture meaningful task-related variance in the data. However, this effect was not noted in the UCLA dataset, in which these methods are substantively more effective in classification of tasks than randomly selecting features. Additionally, in the HCP dataset, hierarchical clustering did not perform better than randomly selecting features. This effect was also noted in the UCLA dataset. Download figure Open in new tab Figure 1. Performance Curve of Each DR and FS Method Across Feature Sizes in HCP Note: Red dots represent accuracy of that classifier at a certain number of features. Grey dots are randomly selected features to serve as a comparison to the method. LDA is not displayed because only one set of features was utilized. Numbers in parentheses represent how many observations were used to train the model. RFPI: Random Forest Permutation Importance, tSVD: Truncated Singular Value Decomposition, PCA: Principal Component Analysis, kPCA: Kernel Principal Component Analysis Download figure Open in new tab Figure 2. Performance Curve of Each DR and FS Method Across Feature Sizes in UCLA Note: Red dots represent accuracy of that classifier at a certain number of features. Grey dots are randomly selected features to serve as a comparison to the method. LDA is not displayed because only one set of features was utilized. Numbers in parentheses represent how many observations were used to train the model. RFPI: Random Forest Permutation Importance, tSVD: Truncated Singular Value Decomposition, PCA: Principal Component Analysis, kPCA: Kernel Principal Component Analysis 3.2 Best performing classifier and DR/FS method by dataset Table 1 reports the best performing classifier for each dataset. Across all methods and feature sizes, utilizing ridge regression with 1,959 features derived from tSVD was the most accurate classification method in the HCP dataset. In UCLA, the most effective method for ridge regression used all 19,900 features, whereas the second most effective method that reduced the number of features was again tSVD using a support vector machine. View this table: View inline View popup Download powerpoint Table 1. Best accuracy of feature selection/dimensionality reduction method by classifier and dataset Table 2 lists the overall accuracy of each type of DR/FS method. Of note, ridge regression was found to be the method which gave rise to the most accurate DR/FS results for each method. In HCP, the most accurate method was kPCA, although most methods were closely accurate (roughly 97-98%). In UCLA, the two most accurate methods were selected using the entire feature set or close to it (HC and RFPI). The next most accurate methods that did not default to the entire (or close to) the entire feature set were kPCA, PCA, and tSVD, which were 80% accurate and used 948 features for tSVD, and 1,065 features for PCA and kPCA. Of note, these results do not suggest an obvious performative edge in DR or FS methods in either dataset. View this table: View inline View popup Download powerpoint Table 2. Overall accuracy of feature selection/dimensionality reduction method by dataset across all feature To examine overall trends in the classification accuracy, we chose to examine the top 10% highest testing accuracy DR/FS methods for each dataset. Overall, the average testing accuracy in the top 10% of the HCP dataset was 96.24% (SD = 1.53%, range = 93.58-99.06%) and 71.79% (SD = 4.07%, range = 66.84-82.11%) in the UCLA dataset, indicating that decoding was more accurate in the HCP dataset. The results of the top 10% best performing classifiers with respect to decoding accuracy are reported in table 3 . The best performing overall classifier was ridge regression. Random forest and ridge regression performed better in HCP than UCLA; however, the support vector machine performed better in the UCLA dataset. View this table: View inline View popup Download powerpoint Table 3. Decoding accuracy by classifier in top 10% best performing DR/FS methods for each classifier Next, the top 10% of best performing DR/FS methods were examined. These results are shown in table 4 . In addition to the individual DR/FS methods, both the average of all three classifiers’ accuracy using the full set of features (i.e., no DR/FS) and randomly selected features are reported as a baseline comparison. In HCP, all methods performed better than the random selection or full feature set. In the UCLA dataset, the full feature set performed better than SFM and HC, and HC was comparable to random feature selection. Across both datasets, tSVD was the best performing DR method, whereas RFPI was the best performing FS method. View this table: View inline View popup Table 4. Overall accuracy of feature selection/dimensionality reduction method by dataset across all feature sizes 3.3 Task-level results In addition to overall performance across DR/FS methods and classifiers, we examined differences in decoding accuracy across specific tasks in the top 10% of DR/FS sets ( Tables 5 and 6 ). It is worth noting that for the HCP dataset, the support vector machine struggled to classify several tasks, such as the emotion, gambling, and working memory tasks. This explains the relatively lower degree of decoding accuracy seen in the SVM relative to other models. The social processing and relational task was the most accurately classified. In the UCLA dataset, the breath holding task was the most accurately classified, whereas the task switching task was the least accurately classified. View this table: View inline View popup Download powerpoint Table 5. Accuracy of prediction of each class in the HCP Dataset by classifier in top 10% decoding accuracy View this table: View inline View popup Download powerpoint Table 6. Accuracy of prediction of each class in the UCLA Dataset by classifier Tables 7 and 8 examine the relative contributions of each DR and FS method to overall task decoding accuracy in the top 10% of decoding accuracy for each method. The overall task level decoding accuracy in HCP was quite high, and comparable to using the full feature sets. In UCLA, some DR/FS methods were not able to achieve greater classification accuracy in some of the tasks relative to using the full feature set, or the top performing groups of randomly selected features. View this table: View inline View popup Download powerpoint Table 7. Accuracy of prediction of each class in the HCP Dataset by DR/FS method View this table: View inline View popup Download powerpoint Table 8. Accuracy of prediction of each class in the UCLA Dataset by DR/FS method Some tasks, such as the social processing task, had high classification accuracy regardless of method employed. The UCLA dataset overall had lower task-based decoding accuracy than HCP. The breath holding task in particular had quite discrepant accuracies with respect to different methods (i.e., 92.00% with LDA, but 73% with SFM). 3.5 Stepwise hierarchical regression For each of the two datasets, a stepwise hierarchical regression was performed to evaluate the impact of feature set size, classifier type, DR/FS method, and interactions between DR/FS methods and classifier types on classifier performance. In the first step, only the number of features and three transformations of the number of features (log, square root, and cubed root) were included in the models. For the UCLA dataset, 54.29% of variance in model accuracy was accounted for by these four variables ( F [4, 5911]=1757.44, p < 0.001). For the HCP dataset, 15.93% of variance in model accuracy was accounted for by these four variables ( F [4, 5818]=276.77, p < 0.001). The second step added dummy coded variables to indicate which classifiers were used (with SVM being the implicit baseline). Adding factors as classifier resulted in a significant increase in variance accounted for in both the UCLA dataset ( F [6, 5909]=1260.82, p < 0.001 R 2 = 0.56; Δ R 2 = 0.02, F [2,5909] = 122.77, p <0.001) and HCP dataset ( F [6, 5818]=276.77, p < 0.001 R 2 = 0.853; Δ R 2 = 0.69, F [2,5816] = 13692, p <0.001). The third step added additional dummy coded variables to represent which DR/FS methods were used with the implicit baseline of randomly selected features. Adding predictors indicating DR/FS method further increased the amount of variance accounted for in model accuracy for both datasets (UCLA: F [12, 5903]= 2550.96, p < 0.001 R 2 = 0.84; Δ R 2 = 0.27, F [6,5903] = 1685, p <0.001; HCP: F [12, 5810]= 3593.78, p < 0.001 R 2 = 0.88; Δ R 2 = 0.03, F [6,5810] = 232, p <0.001). In the fourth and final step, interaction terms between classifier and method were included. Adding the interaction terms resulted in small but significant increases in R 2 in both datasets (UCLA: F [24, 5891]= 1358.07, p < 0.001 R 2 = 0.85; Δ R 2 = 0.008, F [12,5891] = 26.89, p <0.001; HCP: F [24, 5798]= 2333.05, p < 0.001 R 2 = 0.91; Δ R 2 = 0.02, F [12,5798] = 128.2, p <0.001). Overall model fit indices are reported in table 9 , and coefficients from the final models are reported in table 10 . View this table: View inline View popup Download powerpoint Table 9. Stepwise Hierarchical Regression Overview View this table: View inline View popup Table 10. Final Model Regression Results Due to the exceptionally low performance of the SVMs in the HCP dataset, the entire stepwise regression procedure was repeated without SCM results for this dataset ( Table 11 ). The results from the regressions in HCP without data from SVMs more closely mirrored the results from the UCLA dataset. In the first step, 66.2% of the variance in model accuracy was accounted for by the number of features and the three transformations thereof ( F [4, 3877] = 1897.80, p < 0.001). The second step, which included only one dummy-coded variable to indicate whether RR was used instead of RF, accounted for a significantly greater portion of variance in overall model accuracy ( F [5, 3876] = 1592.54, p < 0.001, R 2 = 0.672; Δ R 2 = 0.01, F [1,3876] = 126.25, p <0.001). The addition of variables indicating which DR/FS method was used further improved model performance to a significant degree ( F [11, 3870] = 2486.92, p < 0.001, R 2 Adj = 0.876; Δ R 2 = 0.20, F [6,3870] = 1058.90, p <0.001). The final step, which included interaction terms between classifier and DR/FS method accounted for a small but statistically significant portion of variance in model accuracy ( F [17, 3864] = 1621.99, p < 0.001, R 2 = 0.877; Δ R 2 = 0.001, F [6, 3864] = 5.37, p <0.001). View this table: View inline View popup Download powerpoint Table 11. Stepwise Hierarchical Regression Overview View this table: View inline View popup Table 12. Final Model Regression Results without the SVC classifier for HCP 4. Discussion The present study examined the effects of utilizing a suite of DR and FS methods on cognitive task decoding in two publicly available fMRI datasets. At a broad level, both DR and FS methods performed better than simply providing the classifiers the full input data (19,990 features) or by randomly selecting features. This suggests that such methods can capture meaningful information in functional connectivity features regarding cognitive task decoding. Figures 1 and 2 demonstrate the overall trends of each method and classifier with respect to the full number of possible features selected, from small to large. In several configurations of classifier and DR/FS method, an inverted parabolic function in the number of features is noted, where the optimal decoding accuracy is between roughly 100-2,000 features. These methods often perform better than randomly selecting features, suggesting that useful information is contained in the specific smaller number of features selected, or combined through a DR approach. This “sweet spot” for features has pragmatic computational implications. Considering a baseline total feature set of 19,990 features, this suggests that optimal classification performance can be achieved with many of these methods with only .005-.10% of the total feature space. Thus, a large reduction in the total computational complexity of predictive models can result in very good classification performance. Higher accuracy was obtained in the HCP dataset relative to the UCLA dataset across methods. Some potential reasons for this discrepancy may be scanning parameters, data quality, or the nature of the tasks themselves. However, the support vector machine had poor performance in the HCP dataset overall. Task-level analysis demonstrated that the SVC in HCP in the top 10% of performing SVCs suggested that certain tasks were more difficult to decode for the SVC. This difficulty was not replicated in the other classifiers, nor was this effect demonstrated in the UCLA dataset. In the UCLA dataset, some tasks were not decoded at a rate better than baseline methods in the top 10% of performing models. However, this was not the case in HCP, where there was better task-level decoding. Additionally, in the UCLA dataset, there was more variability in the decoding of tasks, where some tasks, such as the breath-holding tasks, were classified more accurately relative to other tasks. Given the overall greater degree of decoding accuracy in HCP, these results suggest caution in the selection of classifiers in future research, such that certain datasets may have certain characteristics that will cause some classifiers to reach failure modes. When considering all the classifiers and DR/FS methods, our hierarchical regression approach found that DR/FS approaches factored into greater classifier performance in the UCLA than HCP dataset. However, when the poorly performing SVC was removed from the model, this difference disappeared. In both datasets, DR/FS methods accounted for a large Δ R 2 , supporting the idea that DR/FS method matters above and beyond classifier and feature set size. While the present study focuses on decoding accuracy as the optimal outcome, decoding accuracy may not be the sole metric to optimize. A relative advantage of the classifiers used in the present study, and by extension FS methods relative to DR approaches, is the fact that they are more interpretable. In the context of more complex deep learning approaches, which can often be “black box”, these methods can identify particular features associated with the outcome variable of interest. In the HCP dataset in particular, rates of decoding accuracy were often over 95%. The use of more complex deep learning models would not offer substantive gains in accuracy and would be at the expense of interpretability ( Rudin, 2019 ). As an example, the decoding accuracy obtained in the present study in the HCP dataset exceeds that of recent research utilizing more complex modeling architectures ( Nishimura et al., 2024 ). Additionally, the usage of simpler classifiers like the ones in the present study allow for large scale hyperparameter tuning and cross validation, which may not be tractable with deep learning models. In addition to interpretability, some researchers have argued that environmental impact should be considered part of what makes a classifier the “best” performing (Schwartz et al., 2019). In Figure 3 , we plot the runtime of each set of features for each of the different classifiers. This visually represents the time it takes for each classifier with each DR/FS method to be run through the entire cross validated training and testing process. At around 1,000 features, the duration of the runtime dramatically increases. However, this threshold falls within the identified window of higher decoding accuracies in figures 1 and 2. This suggests that an optimal tradeoff between runtime and decoding accuracy can be achieved. This effect was noted in both HCP and UCLA datasets. It is possible that this tradeoff can inform future DR/FS implementations with other fMRI decoding datasets, especially as the energy footprint of developing, training, and testing models will grow to be a larger consideration of ML research in neuroscience. Download figure Open in new tab Figure 3. Classifier runtime as a function of feature sizes and dataset Our task-level analysis suggests that some DR/FS methods and classifiers were better at decoding certain tasks relative to others. However, some tasks, in particular the social processing task in HCP, had a high degree of decoding accuracy regardless of classifier or DR/FS method. However, more variance was noted in the UCLA dataset, where the breath holding task (which was the most accurately decoded task) was not as consistently classified across all DR/FS methods. This suggests that some cognitive tasks have high levels of signal-to-noise ratios relative to other tasks. These results pose several potentially exciting future directions regarding the decodable information contained in task-based fMRI. In particular, it will be useful to examine what makes some tasks robust to methods for a high level of decoding accuracy across classifiers. Potentially, task design or cognitive mechanisms might explain why some tasks are more decodable than others. Future research can perhaps replicate these task designs to better examine factors that contribute to their decodability. 4.1 Limitations A key advantage of this study is the reliance upon methods that can be fit on a separate sample of data than the test set. While many DR or FS methods can be applied regardless of dataset, applying methods that can be tested on data not in the original DR or FS step prevents the possibility of any data leakage from the test set in evaluation of classification accuracy. Leakage of information from the training set into the testing set can be a source of inflated estimates of accuracy ( Rosenblatt et al., 2023 ). The usage of these methods that can be fit on training data and applied to testing data provide unbiased estimates of their performance. We selected methods that operated in this fashion in the present study. Many dimensionality reduction methods exist that do not have the same ability as the ones we have employed here. Thus, while we acknowledge that the methods used in the study are not comprehensive, they utilize best practices in machine learning to reduce the likelihood of overfitting, leakage, and other methodological problems that can cause inflated estimates of accuracy, which can be a problem in the machine learning and brain imaging literature (Varoquax et al., 2017). Additionally, the number of potential classifiers to choose from are quite large, especially with respect to traditional machine learning classifiers. Our choice of classifiers was motivated by frequently employed methods currently used in the fMRI literature that were not wrapped methods that conduct inherent DR or FS (i.e., LASSO or elastic net regression). Additionally, given the large number of classifiers with a large number of DR and FS methods being subject to a grid search for optimal hyperparameters and cross validation, some classifiers (e.g., XGBoost, k-nearest neighbors) can have longer run times at scale that make them prohibitively costly to test at the scope that were tested in the present study ( Bhatia, 2010 ; Bentéjac et al., 2021 ). 4.2 Future Directions The results of the present study suggest that in some datasets, near-perfect decoding accuracy can be achieved using traditional machine learning classifiers with DR and FS methods. The present study utilized static functional connectivity for decoding. While static connectivity is relatively lower in terms of dimensionality, method comparisons such as the ones performed in the present study should be undertaken with dynamic functional connectivity ( Spencer & Goodfellow, 2022 ) or other imaging modalities such as electroencephalography (EEG; García-Laencina et al., 2014 ). Future research should also examine the potential drivers of the “peak” in performance around 100-2,000 features. This work could explore whether these results hold with other methods as well as other datasets. These findings can then inform future decoding research. It would also be of note to explore whether this effect holds of a certain peak of features results in optimal classification in fMRI research outside of cognitive task decoding schemes. Additionally, the context of decoding in the present study was the overall task level. Future research could examine whether this decoding holds at either the block or the trial level within a task, or across broad types of tasks. While each study contained a working memory task, they are difficult to directly compare due to differences in the type of contents in working memory. More finely grained cognitive task phenotyping could be a useful methodological approach to the type of comparative data analysis utilized in the present study ( Poldrack and Yarkoni, 2016 ). Another potential avenue is to explore whether these methods may benefit the decoding of phenotypic information, such as age, sex, psychiatric disorders, or neurological diseases. Finally, while test and train splits were utilized to test the fit of the DR and FS methods utilized in the present study, future work should combine samples of data together where possible to examine if DR or FS methods may still be effective in identical tasks collected in different samples, which can elucidate the contribution of differences in scanner and other protocol steps involved in neuroimaging research. 4.3 Conclusion The present study examined the effectiveness of various DR and FS methods in the context of cognitive task decoding in two open source datasets. These methods performed better than baseline methods, suggesting that DR/FS help improve the signal to noise ratio with respect to decoding. At the dataset level, neither DR or FS appeared to be superior to one another. However, specific DR and FS methods varied in effectiveness with type of task being decoded, with some methods not performing better than baseline for some tasks. An important advantage of FS over DR methods is that features used for classification can be identified. Trends in the size of features were also evident. The number of features can be reduced by a large degree to still achieve at times near-perfect decoding accuracy using traditional machine learning models. This can be of benefit to researchers who are interested in neuroimaging research but may not have access to large scale computing resources. Furthermore, this accuracy can be achieved at low computational cost, thereby demonstrating a route forward for optimizing decoding performance while also minimizing environmental impact. The current trend in neuroimaging is for the size of datasets to increase. As the volume of neuroimaging data and tasks in decoding analysis grows, these findings will have implications for research or design of brain machine interfaces that maximize predictive accuracy while accounting for a large number of predictor variables and while minimizing computational and environmental cost. Code Scripts used to generate the results in this manuscript are available at https://github.com/coreyjr2/HCP-Analyses . Acknowledgments We would like to thank Greg Miller for his support providing computational resources to conduct the analyses in the present paper. References ↵ Abraham , A. , Pedregosa , F. , Eickenberg , M. , Gervais , P. , Mueller , A. , Kossaifi , J. , Gramfort , A. , Thirion , B. , & Varoquaux , G . ( 2014 ). Machine learning for neuroimaging with scikit-learn . Frontiers in Neuroinformatics , 8 . doi: 10.3389/fninf.2014.00014 OpenUrl CrossRef PubMed ↵ Ahmed , N. , & Wahed , M. ( 2020 ). The De-democratization of AI: Deep Learning and the Compute Divide in Artificial Intelligence Research . arXiv preprint arXiv:2010.15581. https://arxiv.org/abs/2010.15581 ↵ Alin , A . ( 2010 ). Multicollinearity . WIREs Computational Statistics , 2 ( 3 ), 370 – 374 . doi: 10.1002/wics.84 OpenUrl CrossRef ↵ Anowar , F. , Sadaoui , S. , & Selim , B . ( 2021 ). Conceptual and empirical comparison of dimensionality reduction algorithms (PCA, KPCA, LDA, MDS, SVD, LLE, ISOMAP, LE, ICA, t-SNE) . Computer Science Review , 40 , 100378 . doi: 10.1016/j.cosrev.2021.100378 OpenUrl CrossRef ↵ Barch , D. M. , Burgess , G. C. , Harms , M. P. , Petersen , S. E. , Schlaggar , B. L. , Corbetta , M. , Glasser , M. F. , Curtiss , S. , Dixit , S. , Feldt , C. , Nolan , D. , Bryant , E. , Hartley , T. , Footer , O. , Bjork , J. M. , Poldrack , R. , Smith , S. , Johansen-Berg , H. , Snyder , A. Z ., … WU-Minn HCP Consortium . ( 2013 ). Function in the human connectome: Task-fMRI and individual differences in behavior . NeuroImage , 80 , 169 – 189 . doi: 10.1016/j.neuroimage.2013.05.033 OpenUrl CrossRef PubMed Web of Science ↵ Bentéjac , C. , Csörgő , A. , & Martínez-Muñoz , G . ( 2021 ). A comparative analysis of gradient boosting algorithms . Artificial Intelligence Review , 54 ( 3 ), 1937 – 1967 . OpenUrl ↵ Besiroglu , T. , Bergerson , S. A. , Michael , A. , Heim , L. , Luo , X. , & Thompson , N . ( 2024 ). The compute divide in machine learning: A threat to academic contribution and scrutiny? . arXiv preprint arXiv : 2401 . 02452 . OpenUrl ↵ Bhatia , N . ( 2010 ). Survey of nearest neighbor techniques . arXiv preprint arXiv : 1007 . 0085 . OpenUrl ↵ Brown , C. J. , & Hamarneh , G. ( 2016 ). Machine Learning on Human Connectome Data from MRI . arXiv Preprint. doi: 10.48550/ARXIV.1611.08699 OpenUrl CrossRef ↵ Esteban , O. , Markiewicz , C. J. , Blair , R. W. , Moodie , C. A. , Isik , A. I. , Erramuzpe , A. , Kent , J. D. , Goncalves , M. , DuPre , E. , Snyder , M. , Oya , H. , Ghosh , S. S. , Wright , J. , Durnez , J. , Poldrack , R. A. , & Gorgolewski , K. J . ( 2019 ). fMRIPrep: A robust preprocessing pipeline for functional MRI . Nature Methods , 16 ( 1 ), 111 – 116 . doi: 10.1038/s41592-018-0235-4 OpenUrl CrossRef PubMed ↵ Esteban , O. , Markiewicz , C. J. , Burns , C. , Goncalves , M. , Jarecka , D. , Ziegler , E. , Berleant , S. , Ellis , D. G. , Pinsard , B. , Madison , C. , Waskom , M. , Notter , M. P. , Clark , D. , Manhães-Savio , A. , Clark , D. , Jordan , K. , Dayan , M. , Halchenko , Y. O. , Loney , F. , … Ghosh , S. ( 2022 ). nipy/nipype: 1.8.3 (Version 1.8.3) [Computer software] . Zenodo . doi: 10.5281/ZENODO.596855 OpenUrl CrossRef ↵ Esteban , O. , Markiewicz , C. J. , Goncalves , M. , Provins , C. , Kent , J. D. , DuPre , E. , Salo , T. , Ciric , R. , Pinsard , B. , Blair , R. W. , Poldrack , R. A. , & Gorgolewski , K. J . ( 2023 ). fMRIPrep: A robust preprocessing pipeline for functional MRI (Version 23.1.2) [Computer software] . Zenodo . doi: 10.5281/ZENODO.852659 OpenUrl CrossRef ↵ García-Laencina , P. J. , Rodríguez-Bermudez , G. , & Roca-Dorda , J . ( 2014 ). Exploring dimensionality reduction of EEG features in motor imagery task classification . Expert Systems with Applications , 41 ( 11 ), 5285 – 5295 . OpenUrl ↵ Genuer , R. , Michel , V. , & Eger , E. ( 2010 ). Random Forests Based Feature Selection for Decoding fMRI Data . ↵ Glaser , J. I. , Benjamin , A. S. , Chowdhury , R. H. , Perich , M. G. , Miller , L. E. , & Kording , K. P . ( 2020 ). Machine Learning for Neural Decoding . Eneuro , 7 ( 4 ), ENEURO.0506-19.2020. doi: 10.1523/ENEURO.0506-19.2020 OpenUrl Abstract / FREE Full Text ↵ Glasser , M. F. , Sotiropoulos , S. N. , Wilson , J. A. , Coalson , T. S. , Fischl , B. , Andersson , J. L. , Xu , J. , Jbabdi , S. , Webster , M. , Polimeni , J. R. , Van Essen , D. C. , & Jenkinson , M. ( 2013 ). The minimal preprocessing pipelines for the Human Connectome Project . NeuroImage , 80 , 105 – 124 . doi: 10.1016/j.neuroimage.2013.04.127 OpenUrl CrossRef PubMed Web of Science ↵ Gorgolewski , K. , Burns , C. D. , Madison , C. , Clark , D. , Halchenko , Y. O. , Waskom , M. L. , & Ghosh , S. S . ( 2011 ). Nipype: A flexible, lightweight and extensible neuroimaging data processing framework in python . Frontiers in Neuroinformatics , 5 , 13 . doi: 10.3389/fninf.2011.00013 OpenUrl CrossRef PubMed ↵ Hauskrecht , M. , Pelikan , R. , Valko , M. , & Lyons-Weiler , J . ( 2007 ). Feature Selection and Dimensionality Reduction in Genomics and Proteomics . In Fundamentals of Data Mining in Genomics and Proteomics (pp. 149 – 172 ). Springer , Boston, MA . doi: 10.1007/978-0-387-47509-7_7 OpenUrl CrossRef ↵ Haxby , J. V. , Gobbini , M. I. , Furey , M. L. , Ishai , A. , Schouten , J. L. , & Pietrini , P . ( 2001 ). Distributed and Overlapping Representations of Faces and Objects in Ventral Temporal Cortex . Science , 293 ( 5539 ), 2425 – 2430 . doi: 10.1126/science.1063736 OpenUrl Abstract / FREE Full Text ↵ Haynes , J.-D. , & Rees , G . ( 2006 ). Decoding mental states from brain activity in humans . Nature Reviews Neuroscience , 7 ( 7 ), 523 – 534 . doi: 10.1038/nrn1931 OpenUrl CrossRef PubMed Web of Science ↵ Hua , J. , Tembe , W. D. , & Dougherty , E. R . ( 2009 ). Performance of feature-selection methods in the classification of high-dimension data . Pattern Recognition , 42 ( 3 ), 409 – 424 . doi: 10.1016/j.patcog.2008.08.001 OpenUrl CrossRef ↵ Jollans , L. , Boyle , R. , Artiges , E. , Banaschewski , T. , Desrivieres , S. , Grigis , A. , … & Whelan , R. ( 2019 ). Quantifying performance of machine learning methods for neuroimaging data . NeuroImage , 199 , 351 – 365 . OpenUrl CrossRef PubMed ↵ LaConte , S. M . ( 2011 ). Decoding fMRI brain states in real-time . NeuroImage , 56 ( 2 ), 440 – 454 . doi: 10.1016/j.neuroimage.2010.06.052 OpenUrl CrossRef PubMed Web of Science ↵ Li , J. , Cheng , K. , Wang , S. , Morstatter , F. , Trevino , R. P. , Tang , J. , & Liu , H . ( 2018 ). Feature Selection: A Data Perspective . ACM Computing Surveys , 50 ( 6 ), 1 – 45 . doi: 10.1145/3136625 OpenUrl CrossRef ↵ Mwangi , B. , Tian , T. S. , & Soares , J. C . ( 2014 ). A Review of Feature Reduction Techniques in Neuroimaging . Neuroinformatics , 12 ( 2 ), 229 – 244 . doi: 10.1007/s12021-013-9204-3 OpenUrl CrossRef PubMed ↵ Nielsen , M. W. , & Andersen , J. P . ( 2021 ). Global citation inequality is on the rise . Proceedings of the National Academy of Sciences , 118 ( 7 ), e2012208118 . OpenUrl Abstract / FREE Full Text ↵ Nishimura , Y. , Sawayama , M. , Yamashita , A. , Nakayama , H. , & Amano , K . ( 2024 ). BrainCodec: Neural fMRI codec for the decoding of cognitive brain states . arXiv preprint arXiv : 2410 . 04383 . OpenUrl ↵ Pedregosa , F. , Varoquaux , G. , Gramfort , A. , Michel , V. , Thirion , B. , Grisel , O. , Blondel , M. , Prettenhofer , P. , Weiss , R. , Dubourg , V. , Vanderplas , J. , Passos , A. , Cournapeau , D. , Brucher , M. , Perrot , M. , & Duchesnay , É . ( 2011 ). Scikit-learn: Machine Learning in Python . J. Mach. Learn. Res ., 12 ( null ), 2825 – 2830 . OpenUrl CrossRef PubMed ↵ Poldrack , R. A . ( 2011 ). Inferring Mental States from Neuroimaging Data: From Reverse Inference to Large-Scale Decoding . Neuron , 72 ( 5 ), 692 – 697 . doi: 10.1016/j.neuron.2011.11.001 OpenUrl CrossRef PubMed Web of Science ↵ Poldrack , R. A. , Congdon , E. , Triplett , W. , Gorgolewski , K. J. , Karlsgodt , K. H. , Mumford , J. A. , Sabb , F. W. , Freimer , N. B. , London , E. D. , Cannon , T. D. , & Bilder , R. M . ( 2016 ). A phenome-wide examination of neural and cognitive function . [dataset] Scientific Data , 3 ( 1 ), 160110 . doi: 10.1038/sdata.2016.110 OpenUrl CrossRef PubMed ↵ Poldrack , R. A. , & Yarkoni , T . ( 2016 ). From brain maps to cognitive ontologies: informatics and the search for mental structure . Annual review of psychology , 67 ( 1 ), 587 – 612 . OpenUrl CrossRef PubMed ↵ Probst D . ( 2023 ). The Societal and Scientific Importance of Inclusivity, Diversity, and Equity in Machine Learning for Chemistry . Chimia , 77 ( 1-2 ), 56 – 61 . doi: 10.2533/chimia.2023.56 OpenUrl CrossRef PubMed ↵ Rao , R. P . ( 2019 ). Towards neural co-processors for the brain: Combining decoding and encoding in brain–computer interfaces . Current Opinion in Neurobiology , 55 , 142 – 151 . doi: 10.1016/j.conb.2019.03.008 OpenUrl CrossRef PubMed ↵ Rashid , B. , & Calhoun , V . ( 2020 ). Towards a brain-based predictome of mental illness . Human Brain Mapping , 41 ( 12 ), 3468 – 3535 . doi: 10.1002/hbm.25013 OpenUrl CrossRef PubMed ↵ Rogers , B. P. , Morgan , V. L. , Newton , A. T. , & Gore , J. C . ( 2007 ). Assessing functional connectivity in the human brain by fMRI . Magnetic Resonance Imaging , 25 ( 10 ), 1347 – 1357 . doi: 10.1016/j.mri.2007.03.007 OpenUrl CrossRef PubMed Web of Science ↵ Rosenblatt , M. , Tejavibulya , L. , Jiang , R. , Noble , S. , & Scheinost , D . ( 2023 ). The effects of data leakage on connectome-based machine learning models . Neuroscience . doi: 10.1101/2023.06.09.544383 OpenUrl Abstract / FREE Full Text ↵ Rubin , T. N. , Koyejo , O. , Gorgolewski , K. J. , Jones , M. N. , Poldrack , R. A. , & Yarkoni , T . ( 2017 ). Decoding brain activity using a large-scale probabilistic functional-anatomical atlas of human cognition . PLOS Computational Biology , 13 ( 10 ), e1005649 . doi: 10.1371/journal.pcbi.1005649 OpenUrl CrossRef PubMed ↵ Rudin , C . ( 2019 ). Stop explaining black box machine learning models for high stakes decisions and use interpretable models instead . Nature machine intelligence , 1 ( 5 ), 206 – 215 . OpenUrl PubMed ↵ Schaefer , A. , Kong , R. , Gordon , E. M. , Laumann , T. O. , Zuo , X.-N. , Holmes , A. J. , Eickhoff , S. B. , & Yeo , B. T. T . ( 2018 ). Local-Global Parcellation of the Human Cerebral Cortex from Intrinsic Functional Connectivity MRI . Cerebral Cortex , 28 ( 9 ), 3095 – 3114 . doi: 10.1093/cercor/bhx179 OpenUrl CrossRef PubMed ↵ Schwartz , R. , Dodge , J. , Smith , N. A. , & Etzioni , O . ( 2020 ). Green ai . Communications of the ACM , 63 ( 12 ), 54 – 63 . OpenUrl CrossRef Shu , P. , Chen , J. , Liu , Z. , Zhao , H. , Li , X. , & Liu , T. ( 2025 ). Survey of HPC in US Research Institutions . arXiv preprint arXiv:2506.19019. ↵ Selvan , R. , Bhagwat , N. , Wolff Anthony , L. F. , Kanding , B. , & Dam , E. B. ( 2022 ). Carbon Footprint of Selecting and Training Deep Learning Models for Medical Image Analysis . In International Conference on Medical Image Computing and Computer-Assisted Intervention (pp. 506-516). ↵ Smith , G . ( 2018 ). Step away from stepwise . Journal of Big Data , 5 ( 1 ), 32 . doi: 10.1186/s40537-018-0143-6 OpenUrl CrossRef ↵ Spencer , A. P. , & Goodfellow , M . ( 2022 ). Using deep clustering to improve fMRI dynamic functional connectivity analysis . NeuroImage , 257 , 119288 . OpenUrl CrossRef PubMed ↵ Srisrisawang , N. , & Müller-Putz , G. R . ( 2022 ). Applying Dimensionality Reduction Techniques in Source-Space Electroencephalography via Template and Magnetic Resonance Imaging- Derived Head Models to Continuously Decode Hand Trajectories . Frontiers in Human Neuroscience , 16 , 830221 . doi: 10.3389/fnhum.2022.830221 OpenUrl CrossRef PubMed ↵ van der Maaten , L. J. P. , Postma , E. O. , & van den Herik , H. J. ( 2009 ). Dimensionality_Reduction_A_Comparative_Review . Journal of Machine Learning Research . https://members.loria.fr/moberger/Enseignement/AVR/Exposes/TR_Dimensiereductie.pdf ↵ Van Essen , D. C. , Ugurbil , K. , Auerbach , E. , Barch , D. , Behrens , T. E. J. , Bucholz , R. , Chang , A. , Chen , L. , Corbetta , M. , Curtiss , S. W. , Della Penna , S. , Feinberg , D. , Glasser , M. F. , Harel , N. , Heath , A. C. , Larson-Prior , L. , Marcus , D. , Michalareas , G. , Moeller , S. , … Yacoub , E. ( 2012 ). The Human Connectome Project: A data acquisition perspective . [ dataset] NeuroImage , 62 ( 4 ), 2222 – 2231 . doi: 10.1016/j.neuroimage.2012.02.018 OpenUrl CrossRef PubMed Varoquaux , G. , Raamana , P. R. , Engemann , D. A. , Hoyos-Idrobo , A. , Schwartz , Y. , & Thirion , B . ( 2017 ). Assessing and tuning brain decoders: cross-validation, caveats, and guidelines . NeuroImage , 145 , 166 – 179 . OpenUrl CrossRef PubMed ↵ Virtanen , P. , Gommers , R. , Oliphant , T. E. , Haberland , M. , Reddy , T. , Cournapeau , D. , Burovski , E. , Peterson , P. , Weckesser , W. , Bright , J. , Van Der Walt , S. J. , Brett , M. , Wilson , J. , Millman , K. J. , Mayorov , N. , Nelson , A. R. J. , Jones , E. , Kern , R. , Larson , E. , … Vázquez-Baeza , Y. ( 2020 ). SciPy 1.0: Fundamental algorithms for scientific computing in Python . Nature Methods , 17 ( 3 ), 261 – 272 . doi: 10.1038/s41592-019-0686-2 OpenUrl CrossRef PubMed ↵ Winkler , A. M. , Ridgway , G. R. , Webster , M. A. , Smith , S. M. , & Nichols , T. E . ( 2014 ). Permutation inference for the general linear model . NeuroImage , 92 , 381 – 397 . doi: 10.1016/j.neuroimage.2014.01.060 OpenUrl CrossRef PubMed Web of Science ↵ Whelan , R. , & Garavan , H . ( 2014 ). When optimism hurts: inflated predictions in psychiatric neuroimaging . Biological psychiatry , 75 ( 9 ), 746 – 748 . OpenUrl CrossRef PubMed View the discussion thread. Back to top Previous Next Posted October 28, 2025. Download PDF Supplementary Material Email Thank you for your interest in spreading the word about bioRxiv. NOTE: Your email address is requested solely to identify you as the sender of this article. Your Email * Your Name * Send To * Enter multiple addresses on separate lines or separate them with commas. You are going to email the following Comparison of dimensionality reduction and feature selection for cognitive task decoding in functional magnetic resonance imaging Message Subject (Your Name) has forwarded a page to you from bioRxiv Message Body (Your Name) thought you would like to see this page from the bioRxiv website. Your Personal Message CAPTCHA This question is for testing whether or not you are a human visitor and to prevent automated spam submissions. Share Comparison of dimensionality reduction and feature selection for cognitive task decoding in functional magnetic resonance imaging Corey J. Richier , Kyle A. Baacke , Sarah M. Olshan , Wendy Heller bioRxiv 2025.10.27.684660; doi: https://doi.org/10.1101/2025.10.27.684660 Share This Article: Copy Citation Tools Comparison of dimensionality reduction and feature selection for cognitive task decoding in functional magnetic resonance imaging Corey J. Richier , Kyle A. Baacke , Sarah M. Olshan , Wendy Heller bioRxiv 2025.10.27.684660; doi: https://doi.org/10.1101/2025.10.27.684660 Citation Manager Formats BibTeX Bookends EasyBib EndNote (tagged) EndNote 8 (xml) Medlars Mendeley Papers RefWorks Tagged Ref Manager RIS Zotero Tweet Widget Facebook Like Google Plus One Subject Area Neuroscience Subject Areas All Articles Animal Behavior and Cognition (7618) Biochemistry (17636) Bioengineering (13860) Bioinformatics (41847) Biophysics (21401) Cancer Biology (18536) Cell Biology (25424) Clinical Trials (138) Developmental Biology (13353) Ecology (19860) Epidemiology (2067) Evolutionary Biology (24287) Genetics (15583) Genomics (22463) Immunology (17701) Microbiology (40300) Molecular Biology (17141) Neuroscience (88434) Paleontology (666) Pathology (2825) Pharmacology and Toxicology (4813) Physiology (7633) Plant Biology (15107) Scientific Communication and Education (2042) Synthetic Biology (4285) Systems Biology (9808) Zoology (2268)

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00