Fuel Cell Degradation Prediction Using Machine Learning Models: A Study on Proton Exchange Membrane (PEM) Fuel Cell Dataset

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Proton Exchange Membrane (PEM) fuel cells offer great potential in terms of green energy solutions but degrade over a period of time based on diverse operation and material aspects. Based on the data obtained from PEM Fuel Cell Dataset with polarization as well as impedance at diverse operations, this study makes use of machine learning algorithms for predicting degradation of fuel cell. Following preprocessing by imputation of missing values via mean imputation, exploratory data analysis was performed via heatmaps and key visualizations. Fifteen machine learning models involving linear and nonlinear regressors, decision trees, ensemble models, and neural networks were trained to predict important performance metrics like cell voltage, power density, and impedance characteristics. Model performance was conducted with Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² Score, where the Extra Trees Regressor performed best with an MAE of 0.00099 and an R² score of 0.996. These results show the potential of machine learning to forecast fuel cell degradation, enabling proactive maintenance and increased system reliability. Future work will explore real-time deployment of predictive models for enhanced operational effectiveness in real-world systems.
Full text 101,348 characters · extracted from preprint-html · click to expand
Fuel Cell Degradation Prediction Using Machine Learning Models: A Study on Proton Exchange Membrane (PEM) Fuel Cell Dataset | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Fuel Cell Degradation Prediction Using Machine Learning Models: A Study on Proton Exchange Membrane (PEM) Fuel Cell Dataset Vivek Yadav, Deepanshu, Harshit Mittal, Vinay Shah, Omkar Singh Kushwaha This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6710108/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Proton Exchange Membrane (PEM) fuel cells offer great potential in terms of green energy solutions but degrade over a period of time based on diverse operation and material aspects. Based on the data obtained from PEM Fuel Cell Dataset with polarization as well as impedance at diverse operations, this study makes use of machine learning algorithms for predicting degradation of fuel cell. Following preprocessing by imputation of missing values via mean imputation, exploratory data analysis was performed via heatmaps and key visualizations. Fifteen machine learning models involving linear and nonlinear regressors, decision trees, ensemble models, and neural networks were trained to predict important performance metrics like cell voltage, power density, and impedance characteristics. Model performance was conducted with Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² Score, where the Extra Trees Regressor performed best with an MAE of 0.00099 and an R² score of 0.996. These results show the potential of machine learning to forecast fuel cell degradation, enabling proactive maintenance and increased system reliability. Future work will explore real-time deployment of predictive models for enhanced operational effectiveness in real-world systems. PEM fuel cell degradation forecasting machine learning electrochemical analysis predictive maintenance Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 1. Introduction Proton Exchange Membrane (PEM) fuel cells are a leading technology for the shift towards sustainable energy, with high efficiency, low emissions, and flexibility across numerous applications, such as transportation and stationary power generation [1]. Yet, one of the most significant issues in the utilization of PEM fuel cells is degradation in performance over time caused by membrane aging, catalyst layer degradation, and water management inefficiencies. The degradation modes result in decreased efficiency and lifetime, requiring strong predictive models to evaluate and counteract long-term performance degradation losses [2,3]. Conventional techniques for examining fuel cell degradation include electrochemical impedance spectroscopy, polarization curve analysis, and accelerated stress testing, all of which are time-consuming and labor-intensive [4]. With the evolution of machine learning (ML) and artificial intelligence (AI), data-driven methods have become more prominent in estimating fuel cell performance degradation with greater accuracy and efficiency [5,6]. ML models are capable of processing intricate relationships between operating parameters like current density, cell voltage, power density, pressure, humidity, and impedance characteristics and offer real-time information regarding fuel cell health and failure prediction [7,8]. This research utilizes machine learning methods to forecast PEM fuel cell degradation from the PEM Fuel Cell Dataset, with information regarding Nafion 112 membrane tests under varying operation conditions. Using various regression-based and ensemble learning models, we seek to build a precise forecasting framework for fuel cell voltage reduction, power loss, and impedance changes. The process of machine learning modeling for prediction of degradation is shown in Fig. 1 . The process of the study includes data preprocessing, feature extraction via heatmaps and visualization, training, and testing with important performance metrics like Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² Score. The study gives an extensive comparison of fifteen various ML models and identifies the top-performing algorithm to predict accurate degradation. 2. Methodology This section describes the procedure for the prediction of degradation in Proton Exchange Membrane (PEM) fuel cells through machine learning (ML) models based on the PEM Fuel Cell Dataset. The research process flow, shown in Fig. 1 , involves data gathering, preprocessing, visualization, model training, and performance assessment, facilitating a systematic procedure towards degradation forecasting (Hastie et al., 2009 ). The workflow utilizes experimental data to pre-process missing values, plot main trends, train 15 ML models, and test their accuracy with the purpose of finding useful predictors of cell voltage as a degradation indicator. 2.1. Data Acquisition The Proton Exchange Membrane (PEM) Fuel Cell Dataset is the basis for this study, offering experimental results necessary for the analysis of fuel cell degradation as a function of different operating conditions. The data contain outcomes from typical Nafion 112 membrane tests and membrane electrode assembly (MEA) activation experiments, recording important electrochemical parameters. It includes polarization curves—such as current density, cell voltage, and power density—and impedance measurements (Z_real and Z_img), in addition to factors of operation such as pressure, relative humidity, membrane compression, Nafion percentage, and applied voltage. The data set measures fuel cell performance under a variety of H₂/O₂ pressures (5, 15, 25 psig) and humidities (30%, 50%, 80%, 100%), providing a rich set of test conditions. First published on Kaggle, it is an invaluable tool for reproducible fuel cell performance research. Although the data set comes from an unpublished preprint [1], its comprehensive electrochemical measurements are consistent with well-established degradation evaluation methods in PEM fuel cell studies [2]. 2.2. Data Preprocessing Preprocessing the PEM Fuel Cell Dataset is an essential step in preparing the data for machine learning-based degradation prediction. Because the dataset contains several CSV files for different test conditions, a custom function was created to merge them into one DataFrame. The function goes through the dataset directory, loads each CSV file, and adds two metadata columns: Test Name (test): Indicate the experimental condition in which the data were recorded. File Name (file): Provides reference and traceability to the source file. The combined dataset, named combined_df, was indexed next to ensure a consistent format for all test conditions [1]. 2.2.1 Missing Values As is typical in experimental data sets, there were some missing values because of measurement errors. To avoid data loss while preserving statistical integrity, missing values in numerical columns were imputed using mean imputation, calculated as: combined_df = combined_df.fillna(combined_df.mean()) This technique was used for the important features like current density, cell voltage, power density, pressure, relative humidity, membrane compression, Nafion percentage, Z_real, Z_img, and applied voltage, ensuring that no crucial data points are lost while retaining feature distributions [2]. Test and file non-numeric columns were not subjected to imputation to avoid any errors. 2.2.2 Feature Scaling and Data Splitting In order to improve the performance of machine learning models, z-score normalization was used to normalize feature scales so that varying magnitudes would have less influence [3]. In line with standard machine learning procedure, the dataset was split into training (80%) and testing (20%) subsets in order to support a robust test of model performance [1]. The preprocessing pipeline in this manner[11] ensured that the dataset was properly structured, statistically sound, and optimized for later visualization and predictive modeling. 2.3. Correlation Analysis Correlation analysis of the PEM Fuel Cell Dataset plays a crucial role in establishing interactions between the various features, including current density, cell voltage, power density, pressure, relative humidity, membrane compression, Nafion percentage, Z_real, Z_img, and applied voltage. Correlation analysis is crucial in guiding machine learning models for fuel cell degradation prediction because the establishment of major feature interactions can have a major impact on model accuracy. Following the preprocessing step, during which missing values were filled with column means and the dataset was normalized, the analysis was conducted using visualizations to reveal hidden patterns and correlations between these variables.The data set, having been taken from Hamidi et al. ( 2020 ), contains in-depth electrochemical measurements, and thus it serves as a sound basis for this analysis[14–16]. First, a correlation heatmap was formed to illustrate overall relationships between features, providing an intuitive insight into their interdependencies. Subsequently, specific scatter plots—Cell Voltage vs. Current Density, Power Density vs. Current Density, Power Density vs. Cell Voltage, Z_img vs. Z_real, and Applied Voltage vs. Set—were drawn to depict certain feature relationships (Figs. 2–7). These plots were important to show noteworthy correlations, i.e., negative relationship between the current density and the cell voltage, or between the impedance (Z_real, Z_img) and the performance characteristics of fuel cells, such as the cell voltage and the power density. What was understood through the correlation analysis proved valuable for the process of feature selection in order to provide input into machine learning models that would be used for prediction of fuel cell degradation and including only the most important variables within. By establishing these correlations, the study helped lead to a more effective and efficient modeling strategy, paving the way for precise fuel cell performance degradation predictions. The heatmap of correlation (Fig. 2) gives a holistic view of Pearson correlation coefficients among numerical features post-standardization, providing key insights into the interdependencies that impact fuel cell performance. Developed on Python's seaborn library, the heatmap measures the strength and direction of linear dependencies ranging from − 1 (perfect negative correlation) to 1 (perfect positive correlation). A strong negative correlation (e.g., -0.85) was found between current density and cell voltage, as would be expected for the notorious polarization characteristics of PEM fuel cells, in which voltage falls as an increasing current demand induces ohmic, activation, and concentration losses [17]. In contrast, power density was positively correlated with current density (e.g., 0.75), which shows the initial rise in power output with increasing current, before finally levelling off through voltage drop effects. Moreover, the heatmap showed lower correlations between operational conditions like pressure and relative humidity with electrical outputs (correlation coefficients less than 0.3), suggesting that their impacts are probably nonlinear or subject to other interactions. This underscores the requirement for machine learning models that can model intricate dependencies beyond linear correlations [7]. By taking advantage of these findings, the feature selection step was optimized to incorporate the most influential variables while also recognizing parameters that needed more advanced modeling techniques to effectively model fuel cell degradation. In order to better explore polarization dynamics, the "Cell Voltage vs. Current Density" plot (Fig. 3) was produced as a scatter plot, with current density on the x-axis and cell voltage on the y-axis, color-coded by test conditions (e.g., various pressure levels). This plot shows the common polarization curve, wherein cell voltage goes from about 1.0 V at low current densities (for example, 0.1 A/cm²) to below 0.6 V for high densities (for example, 1.5 A/cm²). The sharp initial fall is due to activation losses, followed by linear decline by ohmic resistance and sharper fall at high current densities by mass transport limitations [19]. Variations with test conditions, such as elevated voltages at higher pressure levels (e.g., 25 psig versus 5 psig), underscore pressure's contribution to greater oxygen availability, reducing electrode losses, and improving fuel cell efficiency. The trends are highly substantial in terms of voltage stability, one of the degradation indicators, further reinforcing the significance of incorporating pressure-dependent effects into machine learning models to accurately predict degradation [2]. The "Power Density vs. Current Density" plot (Fig. 4 ) gives additional insight into PEM fuel cell performance by showing power density as a function of current density, demonstrating the typical parabolic behavior of fuel cells. Because power density is the product of cell voltage and current density, it begins close to zero at low current levels, increases to a maximum of about 0.4–0.5 W/cm² at intermediate current densities (e.g., 0.8–1.0 A/cm²), and then decreases as voltage losses become dominant at high currents. Peak power output is slightly different depending on operating conditions—greater relative humidity (e.g., 100%) increases membrane hydration, lowering resistance, and greater pressure allows for gas diffusion, both contributing to increased fuel cell efficiency [8]. This graph is critical for determining optimal operating conditions and for detecting performance degradation, as deviations from predicted power density values can indicate underlying problems, such as membrane drying, catalyst degradation, or gas diffusion limitations [6]. The "Power Density vs. Cell Voltage" graph (Fig. 5) indicates power output and voltage relationship in a hyperbolic curve characteristic of PEM fuel cells. The maximum power density is at 0.6–0.7 V, where voltage and current balance produces maximum output. Power is low at high voltages (> 0.9 V) because there is not enough current, and power decreases at low voltages (< 0.5 V) due to greater losses. This connection highlights the influence of voltage stability on fuel cell performance, thus it is a critical parameter for machine learning-based degradation prediction[1]. The "Zimg vs. Zreal" plot (Fig. 6 ) is an Nyquist plot representing electrochemical impedance (EIS) response. Lower resistance is indicated by smaller arcs under increased pressures (25 psig), and increased resistance by higher arcs at decreased humidity (30%) because the membrane becomes dehydrated [19]. The trends aid in identifying mechanisms for degradation, thereby making z_real and z_img useful for prediction based on machine learning [2]. The "Applied Voltage vs. Set" graph (Fig. 7) illustrates applied voltage variation between test sets, illustrating stability and set-specific voltage increments between 0.4 V and 1.0 V. Abrupt decreases (e.g., from 0.8 V to 0.5 V in a set) can signal onset of degradation, e.g., electrode fouling or gas starvation problems[1] [20]. Trends corroborate polarization values, solidifying voltage as one of the predominant degradation indicators in predictive modeling. The plots (Figs. 3–7) and heatmap (Fig. 2) of the data point out salient relationships: the parabolic power_density-current_density curve, the inverse cell_voltage-current_density trend, the voltage-power dependency, impedance's diagnostic value, and stability of applied_voltage. Correlation coefficients on the heatmap inform feature selection, with strong predictors (e.g., current_density, cell_voltage) being prioritized and nonlinear dependencies (e.g., pressure) acknowledged [1]. Confirmation of these trends and their electrochemical implication is provided by scatter plots (Figs. 3– 6 ) and the line plot (Fig. 7) [2]. These observations drive the ML modeling strategy, as voltage deviations are indicative of degradation, and impedance data supports internal state monitoring [13]. The requirement for the use of all features, with tree-based models being adept at modeling nonlinear interactions, provides predictive robustness, in turn connecting preprocessing to data-driven degradation analysis [4]. 2.4. Machine Learning Modeling Following preprocessing and correlation analysis (Sections 2.2 and 2.3 ), machine learning (ML) modeling was executed to forecast cell voltage in the PEM Fuel Cell Dataset, a marker of fuel cell degrading. Cell voltage was selected as target variable because of its high correlation with current_density and responsiveness to working conditions and degrading mechanisms, evident in the polarization plots (Figs. 3–5) and impedance data (Fig. 6 ) [18]. There were fifteen ML models trained and tested, utilizing predictors such as current_density, power_density, pressure, humidity, membrane_compression, nafion_percent, z_real, z_img, and applied_voltage [2]. The models were structured into four categories: (1) Linear Models (Linear Regression, Ridge, Lasso, ElasticNet) for linear trends, (2) Tree-Based Models (Decision Tree, Random Forest, Extra Trees) for nonlinearities, (3) Boosting Models (Gradient Boosting, AdaBoost, XGBoost, LightGBM) for iterative refinement, and (4) other models such as KNN, SVR, and MLP for various approaches [3]. Python libraries (scikit-learn, CatBoost, XGBoost, LightGBM) were employed with default hyperparameters for the sake of comparison [4–8]. The preprocessed dataset, normalized and separated into 80% training and 20% test sets (Section 2.2 ), was employed for model training. Cell voltage prediction was modeled as a regression problem with the training set utilized to train each model and the test set utilized to measure model performance. Four metrics were utilized to measure model accuracy and fit: Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² Score. MAE approximates the mean absolute difference between predicted and true values, providing an easy-to-understand measure of model accuracy: mae = mean_absolute_error(y_test, y_pred)…………………………………………………………..(1) MSE computes the mean squared errors, assigning more importance to larger errors: mse = mean_squared_error(y_test, y_pred)……………………………………………………………(2) RMSE, the square root of MSE, gives a more understandable value with the same unit as the target variable: rmse = np.sqrt(mse)……………………………………………………………………………………….(3) R² Score is the percentage of variance in the dependent variable that is accounted for by the independent variables, with a 1 representing perfect predictability: r2 = r2_score(y_test, y_pred)…………………………………………………………………………….(4) Here, y_test is the actual cell voltage, and y_pred is the predicted values. These measures give a clear assessment of model performance in terms of absolute error, sensitivity to error, and explanatory power, which are essential in precisely predicting fuel cell degradation (Chai & Draxler, 2014 ; Hamidi et al., 2020 ). 2.5. Fuel Cell Degradation Prediction This study predicts the degradation of Proton Exchange Membrane (PEM) fuel cells by predicting cell voltage with the Extra Trees Regressor (R² = 0.9962, Table 1 ) trained on the PEM Fuel Cell Dataset [1]. Cell voltage, a parameter sensitive to degradation as evident from the "Cell Voltage vs Current Density" (Fig. 3) and "Zimg vs Zreal" (Fig. 6 ) plots, is predicted based on features such as current_density and pressure. The model defines a reference voltage and monitors deviations (e.g., > 0.05 V) from real values in order to detect possible degradation indicators, including membrane thinning [2]. The "line_applied_voltage" plot (Fig. 7) is employed for finding instability trends. Although the lack of time-series data restrains straightforward Remaining Useful Life (RUL) prediction, the method facilitates condition-based monitoring, with future efforts directed towards incorporating temporal data [3]. 3. Result This section reports the outcome of the correlation analysis and machine learning (ML) modeling on the PEM Fuel Cell Dataset for predicting cell voltage as a surrogate for degradation in Proton Exchange Membrane (PEM) fuel cells [1]. Derived from preprocessing (Section 2.2 ), correlation analysis (Section 2.3 ), and modeling (Section 2.4 ), the outcome is depicted in figures and encapsulated in Table 1 and provides insights into electrochemical performance and predictive capability. 3.1. Performance Analysis of Machine Learning Models Performance of 15 machine learning (ML) models learned to predict cell voltage on the PEM Fuel Cell Dataset was measured by using Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² Score, as given in Table 1 [1]. Ensemble tree models, specifically the Extra Trees Regressor (MAE = 0.0009878, MSE = 2.19E-05, RMSE = 0.004676, R² = 0.9962), performed better than others, followed by Random Forest and Decision Tree Regressor, showing their capacity to learn nonlinear relationships such as between cell_voltage and current_density [8–9]. Boosted models, such as XGBoost (R² = 0.9835) and CatBoost (R² = 0.9668), also worked well, taking advantage of iterative error correction [7]. Linear models such as Linear Regression, Ridge, Lasso, and ElasticNet exhibited moderate performance (R² ~0.6), whereas SVR and MLP Regressor performed poorly with MLP exhibiting signs of overfitting [15]. Table 1 Machine learning model evaluations to assess MSE, R 2 , MAE and RMSE. Model MSE R² MAE RMSE Linear Regression 0.002194 0.622595 0.0138163 0.046836 Ridge Regression 0.002194 0.622595 0.0138163 0.046836 Lasso Regression 0.002379 0.590674 0.0132962 0.048776 ElasticNet 0.002327 0.59956 0.0131465 0.048224 Decision Tree Regressor 5.41E-05 0.990686 0.0015484 0.007358 Random Forest 2.97E-05 0.994893 0.0011401 0.005448 Extra Trees Regressor 2.19E-05 0.996237960328682 0.0009878 0.004676 Gradient Boosting Regressor 0.00142 0.75575 0.0109353 0.037678 AdaBoost Regressor 0.000898 0.84555 0.0111709 0.029962 KNeighbors Regressor 5.44E-05 0.990642 0.0014769 0.007375 SVR 0.009212 -0.58498 0.0938222 0.095981 MLP Regressor 0.055707 -8.58446 0.0967775 0.236.24 CatBoost Regressor 0.000193 0.966798 0.0037322 0.013892 XGBoost Regressor 9.58E-05 0.983512 0.005225 0.009789 LightGBM Regressor 0.000885 0.847757 0.008727 0.029747 *Note: MLP R² anomaly indicates potential overfitting. The Extra Trees Regressor’s near-perfect R² suggests its suitability for establishing a degradation baseline, enhancing predictive maintenance [4]. 3.2. Visualization of Model Performance To assess the predictive capabilities of the 15 machine learning (ML) models that were trained on the PEM Fuel Cell Dataset for predicting cell voltage—a critical measure of degradation in Proton Exchange Membrane (PEM) fuel cells—a combined visualization (Fig. 8 ) was obtained, combining bar plots of Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R² Score for all models [4]. The Extra Trees Regressor performed best with lowest MAE (0.0009878) and MSE (2.19E-05), performing significantly better than Linear Regression (MAE = 0.013816) and Ridge (0.013816). SVR and MLP Regressor were seen performing poorly, MLP being a case of overfitting by its strange interaction with R² [11]. The visualization indicates these trends, confirming the better capacity of tree-based and boosted models to identify intricate, nonlinear trends in the data, essential for monitoring degradation [1]. The RMSE panel in Fig. 8 , converting errors to units of voltage, points out the accuracy of the Extra Trees Regressor (0.004676), then Random Forest (0.005448) and Decision Tree (0.007358). Boosting algorithms such as XGBoost (0.009789) and CatBoost (0.013892) follow closely, with SVR (0.095981) and MLP (0.236024) performing much larger errors, indicating difficulties in modeling the dataset's nonlinearities [9]. Linear models (such as Linear Regression, 0.046836) show moderate RMSE values, which align with their R² values [3]. The R² plot illustrates that Extra Trees (0.9962), Random Forest (0.9949), and Decision Tree (0.9907) capture the majority of the variance, with boosting models such as XGBoost (0.9835) and CatBoost (0.9668) performing well as well. Linear models are around 0.6 (e.g., Linear Regression, 0.6226), whereas MLP's R² (8.5845) is an outlier, probably because of overfitting or numerical errors [12]. Figure 8 supports the superiority of tree-based ensemble models, consistent with the nonlinear relationships seen in Figs. 2– 6 (Section 2.3 ). Their minimal error rates and maximum R² values indicate good prospects for creating a degradation baseline, in which deviations from anticipated cell voltage might indicate performance degradation. Visualization justifies the choice of Extra Trees as the best model for real-world PEM fuel cell diagnostics [6]. 3.3. Model Fit and Residual Analysis To improve comparison of the performance of the 15 ML models in cell voltage prediction for tracking PEM fuel cell degradation, "Actual vs Predicted" and "Residual Plots" were merged into one figure (Fig. 9 ), along with Table 1 and Fig. 8 [1]. Figure 9 , developed through Python's matplotlib, visually assesses prediction accuracy and error distribution for all models. Random Forest (R² = 0.9949) and Extra Trees (R² = 0.9962) lie close to the y = x line with very little scatter, mirroring their tiny MAE (0.0011401 and 0.0009878) and strong performance [2]. Linear models like Linear Regression (R² = 0.6226) have a larger scatter, particularly at lower voltages, representing lower accuracy [3]. Decision Tree, XGBoost, and CatBoost have good fits with negligible deviations. SVR and MLP exhibit strong dispersion in their "Actual vs Predicted" plots, pointing out weaknesses [4]. The "Residual" panels reveal tight clusters around zero for Extra Trees and Random Forest, indicating little error, whereas linear models exhibit heteroscedasticity, especially at lower predictions [5]. MLP's residuals are dispersed, lending support to overfitting problems. Figure 9 illustrates the better performance of tree-based models in cell voltage prediction of degradation, as expected from their high R² and low errors (Table 1 ) [1]. To compare 15 ML models for cell voltage prediction in PEM fuel cells, "Actual vs Predicted" and "Residual Plots" were merged into Fig. 9 , supported by Table 1 and Fig. 8 [1]. This multi-panel plot, generated through Python's matplotlib, visually evaluates prediction accuracy and error distribution across models such as Linear Regression, Random Forest, Extra Trees, XGBoost, and so on [2]. In the "Actual vs Predicted" plots, Extra Trees (R² = 0.9962) and Random Forest (R² = 0.9949) have negligible scatter and low MAE (0.0009878 and 0.0011401, respectively) [3]. Decision Tree (R² = 0.9907) and boosting models (XGBoost, CatBoost) also fit well with slight deviations. Linear models have more scatter, particularly at low voltages, which indicates their inability to track nonlinear trends [4]. The "Residual" panels have little systematic error for tree-based models, while linear models display heteroscedasticity, and MLP and SVR models have severe overfitting or poor fit [5][6]. Figure 9 confirms the supremacy of tree models in prediction of degradation, augmenting their large R² and few errors (Table 1 ), and sanctions their uses in cell voltage divergences monitoring [2]. Conclusion This research assessed machine learning (ML) utilization in predicting fuel cell degradation in Proton Exchange Membrane (PEM) fuel cells through forecasting cell voltage as a surrogate for degradation with the PEM Fuel Cell Dataset [2].The dataset, comprising polarization (current_density, cell_voltage, power_density) and impedance (z_real, z_img) data under varying conditions (pressure, relative_humidity), was preprocessed to replace missing values by column means [14]. Correlation analysis (Section 2.3 ) identified significant correlations, graphed as plots like "Cell Voltage vs Current Density" (Fig. 3) and "Zimg vs Zreal" (Fig. 6 ), indicative of operating effects on performance. These observations were used to train 15 ML models (Section 2.4 ), ranging from linear models (e.g., Linear Regression) to ensemble models (e.g., Extra Trees Regressor), performance being measured on the basis of MAE, MSE, RMSE, and R² scores (Table 1 ). The outcomes (Section 3 ) support the excellence of tree-based ensemble methods in the prediction of degradation, with Extra Trees Regressor giving the best predictions (MAE = 0.0009878, R² = 0.9962), followed by Random Forest (R² = 0.9949) and Decision Tree (R² = 0.9907) [1]. Boosting models such as XGBoost (R² = 0.9835) and CatBoost (R² = 0.9668) also worked well, whereas linear models like Linear Regression (R² = 0.6226) and SVR (R² = 0.584978) were not very effective, unable to identify nonlinearities [2]. The peculiar R² value of the MLP Regressor (8.584462) reflected overfitting, as verified using residual analysis [3Visualisations of performance (Fig. 8 ) and residuals (Fig. 9 ) ensured Extra Trees' accuracy and minimum residual scatter that is appropriate for monitoring degradation [4]. Such models, like Extra Trees, are used within the degradation forecasting framework (Section 2.5 ) to calculate a base cell voltage value and mark strong deviations (> 0.05 V) as indicative of degradation effects, such as membrane thinning or catalyst wearing [5]. Through monitoring of performance anomalies and stability trends (Fig. 7) [6], the method facilitates condition-based maintenance, though the static dataset restricts direct Remaining Useful Life (RUL) prediction. The research illustrates the capability of ML to enhance PEM fuel cell reliability, with Extra Trees providing near-perfect predictions, consistent with earlier data-driven diagnostic models [7]. Limitations of this study are because of the absence of degradation data in the dataset, restricting long-term prediction [1]. Models were trained on default hyperparameters, which can decrease performance [1]. Although methods like mean imputation were used, exploration of other methods like interpolation would improve the management of data [2]. The MLP anomaly is deserving of further investigation of its architecture or scaling modifications [3]. Future studies would be able to make direct predictions of Remaining Useful Life (RUL) utilizing time series or cycle count as inputs. Improving accuracy along with interpretability could be made possible through cross-validation and hyperparameter optimization involving hybrid models by merging machine learning and physics-based methods [4]. Augmentation of the data set with supplementary operating conditions or degradation states will further confirm the framework, accelerating its real-world application in fuel cell systems. In total, this research is a good foundation for ML-based diagnosis of PEM fuel cells and presents a scalable method for increasing their durability and efficiency [5]. Declarations Code Availability The code along with the model plots, ready to use models are available at: https://github.com/I-Deepanshu/Fuel-cell-degradation-predictor References Hamidi, S., Haghighi, S., & Askari, K. (2020). PEM Fuel Cell Dataset . Journal of Energy Engineering, 146(6), 04020089. https://doi.org/10.1061/(ASCE)EY.1943-7897.0000703 Batista, G., & Monard, M. (2003). An analysis of four algorithms for classification and prediction of missing values. Computational Statistics & Data Analysis, 44 (1-2), 163-181. https://doi.org/10.1016/S0167-9473(03)00006-1 Schölkopf, B., Platt, J. C., Shawe-Taylor, J., Smola, A. J., & Williamson, R. C. (1999). Estimating the Support of a High-Dimensional Distribution. Neural Computation, 13 (7), 1443–1471. https://doi.org/10.1162/089976699300016428 Ferronato, N., & Torretta, V. (2019). Waste Mismanagement in Developing Countries: A Review of Global Issues. International Journal of Environmental Research and Public Health, 16 (6), 1074. https://doi.org/10.3390/ijerph16061074 Liu, H., Zhou, M., & Zhang, X. (2021). Data-driven diagnostics and prognostics of fuel cells: A review of methodologies and applications. Energy Reports, 7 , 561-576. https://doi.org/10.1016/j.egyr.2021.03.073 Breiman, L. (2001). Random forests. Machine Learning, 45 (1), 5-32. https://doi.org/10.1023/A:1010933404324 Hastie, T., Tibshirani, R., & Friedman, J. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7 Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., & Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12 , 2825–2830. http://www.jmlr.org/papers/volume12/pedregosa11a/pedregosa11a.pdf Chai, T., & Draxler, R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? Geoscientific Model Development, 7 , 1247–1250. https://doi.org/10.5194/gmd-7-1247-2014 Yuan, X., Wang, H., & Sun, J. (2008). Impedance measurement and equivalent circuit modeling for Proton Exchange Membrane (PEM) fuel cells. Journal of Power Sources, 179 (1), 146–154. https://doi.org/10.1016/j.jpowsour.2007.09.074 Larminie, J., & Dicks, A. (2003). Fuel cell systems explained (2nd ed.). Wiley. Jouin, M., Youssef, J., & Tisseyre, R. (2016). Prognostics for PEM fuel cell degradation: A review. Energy Procedia, 105 , 503–507. https://doi.org/10.1016/j.egypro.2016.11.057 Pei, P., Chen, H., & Wu, J. (2011). Degradation study of Proton Exchange Membrane (PEM) fuel cells: An overview. Journal of Power Sources, 196 (4), 1780–1790. https://doi.org/10.1016/j.jpowsour.2010.08.113 Ferraro, A., & Mazzarotta, B. (2017). Monitoring the degradation of PEM fuel cells by voltage recovery and monitoring of current density. International Journal of Hydrogen Energy, 42 (25), 16322–16330. https://doi.org/10.1016/j.ijhydene.2017.05.086 Saha, B., & Kamaruddin, S. (2016). A review on predictive maintenance of PEM fuel cells. Renewable and Sustainable Energy Reviews, 56 , 91–103. https://doi.org/10.1016/j.rser.2015.11.071 Pandian, M., & Natarajan, K. (2020). Optimization of Proton Exchange Membrane Fuel Cell for Performance and Efficiency: A Review. Journal of Energy, 2020 , 3248154. https://doi.org/10.1155/2020/3248154 Pardo, J. A., & Carmona, R. (2017). Condition monitoring and performance prediction of fuel cells: Data-driven models. International Journal of Hydrogen Energy, 42 (12), 8536–8547. https://doi.org/10.1016/j.ijhydene.2017.03.011 Yan, Y., & Wang, H. (2018). Development and application of diagnostic techniques in PEM fuel cell: A review. Energy Conversion and Management, 165 , 47-58. https://doi.org/10.1016/j.enconman.2018.03.040 Ahmed, S., & Zhang, Y. (2019). A comparative review of diagnostic and prognostic methods for PEM fuel cells. Renewable and Sustainable Energy Reviews, 101 , 122–135. https://doi.org/10.1016/j.rser.2018.09.014 Xie, X., & Zeng, Q. (2018). A novel degradation model for PEM fuel cell performance prediction based on hybrid machine learning techniques. Energy, 150 , 80–89. https://doi.org/10.1016/j.energy.2018.01.022 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6710108","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":459481617,"identity":"3e128bf2-9677-4b4d-a2fd-dc308dbfe454","order_by":0,"name":"Vivek Yadav","email":"","orcid":"","institution":"Guru Gobind Singh Indraprastha University","correspondingAuthor":false,"prefix":"","firstName":"Vivek","middleName":"","lastName":"Yadav","suffix":""},{"id":459481618,"identity":"0cc9e10b-f0c0-416f-8ad9-dd745f15bd79","order_by":1,"name":"Deepanshu","email":"","orcid":"","institution":"","correspondingAuthor":false,"prefix":"","firstName":"","middleName":"","lastName":"Deepanshu","suffix":""},{"id":459481619,"identity":"48713880-93b6-457d-ad18-cf8115f376e7","order_by":2,"name":"Harshit Mittal","email":"","orcid":"","institution":"Guru Gobind Singh Indraprastha University","correspondingAuthor":false,"prefix":"","firstName":"Harshit","middleName":"","lastName":"Mittal","suffix":""},{"id":459481620,"identity":"c0f3856d-a179-4453-8759-9c4cbf28d2cc","order_by":3,"name":"Vinay Shah","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAzElEQVRIiWNgGAWjYDACZhBhIMHMzwwVYCNOS4UNu2Qz0VrA4Ewav8EBYt1lcJz34GfetsPSxsd5j0kw1Ngx8Ek3ENBymC9ZGqjF2OwwX5oEw7FkBjYZAvZJNvMYgLQkmx3mMZNgYDvAwCaRQFCL8W+glvrNzSAt/4jQws/MYybNcyaN2QDIkGBsI1KL5ZwKG2YJoKcsEvuSeQhqYeM/Y3zjDSgq+88evPHhm52c/AwCWkCAiQdMAckEMEkEYPwB0zIKRsEoGAWjABsAAKTtMwivsLUVAAAAAElFTkSuQmCC","orcid":"","institution":"Guru Gobind Singh Indraprastha University","correspondingAuthor":true,"prefix":"","firstName":"Vinay","middleName":"","lastName":"Shah","suffix":""},{"id":459481621,"identity":"8a03f267-75d6-40e4-bbb3-4e92fd2385de","order_by":4,"name":"Omkar Singh Kushwaha","email":"","orcid":"","institution":"Indian Institute of Technology, Madras","correspondingAuthor":false,"prefix":"","firstName":"Omkar","middleName":"Singh","lastName":"Kushwaha","suffix":""}],"badges":[],"createdAt":"2025-05-20 17:45:42","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":false,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":false},"doi":"10.21203/rs.3.rs-6710108/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6710108/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":83331706,"identity":"ea7decdf-7f98-4edf-afc6-d5f4c3c911a9","added_by":"auto","created_at":"2025-05-23 07:57:18","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":88769,"visible":true,"origin":"","legend":"\u003cp\u003eResearch flow process\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/ad28bc76c7100ffad00a97ed.png"},{"id":83331710,"identity":"69319f71-d7ee-4ba2-80d7-ba65ba9b5de1","added_by":"auto","created_at":"2025-05-23 07:57:18","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":56585,"visible":true,"origin":"","legend":"\u003cp\u003eCorrelation Heatmap\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/1a2a2255200a7d6c2d8168ab.png"},{"id":83331918,"identity":"3fcf53c2-e844-48de-8527-eb696ad93966","added_by":"auto","created_at":"2025-05-23 08:05:18","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":127429,"visible":true,"origin":"","legend":"\u003cp\u003eCell Voltage vs Current Density\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/9da7774c860883e54fbf3a3f.png"},{"id":83331919,"identity":"9d3aa6cf-0b42-4f4d-be2b-41ac9f0cb9bb","added_by":"auto","created_at":"2025-05-23 08:05:18","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":276068,"visible":true,"origin":"","legend":"\u003cp\u003ePower Density vs Current Density\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/14ad209227cee775806947d9.png"},{"id":83331708,"identity":"85ecdce7-ffb4-4579-a08f-1f6321ef4f4c","added_by":"auto","created_at":"2025-05-23 07:57:18","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":179829,"visible":true,"origin":"","legend":"\u003cp\u003ePower Density vs Cell Voltage\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/11f78362dcf75da40476093c.png"},{"id":83332566,"identity":"7c660ca4-aeea-47af-aa12-50b7196e326a","added_by":"auto","created_at":"2025-05-23 08:13:18","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":129759,"visible":true,"origin":"","legend":"\u003cp\u003eZimg vs Zreal\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/620c30d1b6d2915a8fa156b3.png"},{"id":83331923,"identity":"cfbb8e74-193b-4fef-a46b-d73044cf28ea","added_by":"auto","created_at":"2025-05-23 08:05:18","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":45943,"visible":true,"origin":"","legend":"\u003cp\u003eApplied Voltage vs Set\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/e5ff528156ec697a292a4403.png"},{"id":83331925,"identity":"e6d63bbd-8281-480b-94fe-4c9c5f9921fa","added_by":"auto","created_at":"2025-05-23 08:05:18","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":186980,"visible":true,"origin":"","legend":"\u003cp\u003eModel Performance Metrics\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/6833676e4f58e8b07f85ed3b.png"},{"id":83331725,"identity":"2d8c04c6-c50e-4edb-a4a8-8260b0979651","added_by":"auto","created_at":"2025-05-23 07:57:18","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":1933694,"visible":true,"origin":"","legend":"\u003cp\u003eActual vs Predicted \u0026amp; Residual\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/099aa71be7da11920aa9b4c4.png"},{"id":83333217,"identity":"9cbe84ad-b01d-4466-b79c-3e1af87957ff","added_by":"auto","created_at":"2025-05-23 08:29:20","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3182015,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6710108/v1/10963a30-77b6-4b34-84b5-9f2632e75978.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eFuel Cell Degradation Prediction Using Machine Learning Models: A Study on Proton Exchange Membrane (PEM) Fuel Cell Dataset\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eProton Exchange Membrane (PEM) fuel cells are a leading technology for the shift towards sustainable energy, with high efficiency, low emissions, and flexibility across numerous applications, such as transportation and stationary power generation [1]. Yet, one of the most significant issues in the utilization of PEM fuel cells is degradation in performance over time caused by membrane aging, catalyst layer degradation, and water management inefficiencies. The degradation modes result in decreased efficiency and lifetime, requiring strong predictive models to evaluate and counteract long-term performance degradation losses [2,3].\u003c/p\u003e \u003cp\u003eConventional techniques for examining fuel cell degradation include electrochemical impedance spectroscopy, polarization curve analysis, and accelerated stress testing, all of which are time-consuming and labor-intensive [4]. With the evolution of machine learning (ML) and artificial intelligence (AI), data-driven methods have become more prominent in estimating fuel cell performance degradation with greater accuracy and efficiency [5,6]. ML models are capable of processing intricate relationships between operating parameters like current density, cell voltage, power density, pressure, humidity, and impedance characteristics and offer real-time information regarding fuel cell health and failure prediction [7,8].\u003c/p\u003e \u003cp\u003eThis research utilizes machine learning methods to forecast PEM fuel cell degradation from the PEM Fuel Cell Dataset, with information regarding Nafion 112 membrane tests under varying operation conditions. Using various regression-based and ensemble learning models, we seek to build a precise forecasting framework for fuel cell voltage reduction, power loss, and impedance changes. The process of machine learning modeling for prediction of degradation is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The process of the study includes data preprocessing, feature extraction via heatmaps and visualization, training, and testing with important performance metrics like Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R\u0026sup2; Score. The study gives an extensive comparison of fifteen various ML models and identifies the top-performing algorithm to predict accurate degradation.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cp\u003eThis section describes the procedure for the prediction of degradation in Proton Exchange Membrane (PEM) fuel cells through machine learning (ML) models based on the PEM Fuel Cell Dataset. The research process flow, shown in Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, involves data gathering, preprocessing, visualization, model training, and performance assessment, facilitating a systematic procedure towards degradation forecasting (Hastie et al., \u003cspan class=\"CitationRef\"\u003e2009\u003c/span\u003e). The workflow utilizes experimental data to pre-process missing values, plot main trends, train 15 ML models, and test their accuracy with the purpose of finding useful predictors of cell voltage as a degradation indicator.\u003c/p\u003e\n\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\n \u003ch2\u003e2.1. Data Acquisition\u003c/h2\u003e\n \u003cp\u003eThe Proton Exchange Membrane (PEM) Fuel Cell Dataset is the basis for this study, offering experimental results necessary for the analysis of fuel cell degradation as a function of different operating conditions. The data contain outcomes from typical Nafion 112 membrane tests and membrane electrode assembly (MEA) activation experiments, recording important electrochemical parameters. It includes polarization curves\u0026mdash;such as current density, cell voltage, and power density\u0026mdash;and impedance measurements (Z_real and Z_img), in addition to factors of operation such as pressure, relative humidity, membrane compression, Nafion percentage, and applied voltage.\u003c/p\u003e\n \u003cp\u003eThe data set measures fuel cell performance under a variety of H₂/O₂ pressures (5, 15, 25 psig) and humidities (30%, 50%, 80%, 100%), providing a rich set of test conditions. First published on Kaggle, it is an invaluable tool for reproducible fuel cell performance research. Although the data set comes from an unpublished preprint [1], its comprehensive electrochemical measurements are consistent with well-established degradation evaluation methods in PEM fuel cell studies [2].\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\n \u003ch2\u003e2.2. Data Preprocessing\u003c/h2\u003e\n \u003cp\u003ePreprocessing the PEM Fuel Cell Dataset is an essential step in preparing the data for machine learning-based degradation prediction. Because the dataset contains several CSV files for different test conditions, a custom function was created to merge them into one DataFrame. The function goes through the dataset directory, loads each CSV file, and adds two metadata columns:\u003c/p\u003e\n \u003cp\u003eTest Name (test): Indicate the experimental condition in which the data were recorded.\u003c/p\u003e\n \u003cp\u003eFile Name (file): Provides reference and traceability to the source file.\u003c/p\u003e\n \u003cp\u003eThe combined dataset, named combined_df, was indexed next to ensure a consistent format for all test conditions [1].\u003c/p\u003e\n \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\n \u003ch2\u003e2.2.1 Missing Values\u003c/h2\u003e\n \u003cp\u003eAs is typical in experimental data sets, there were some missing values because of measurement errors. To avoid data loss while preserving statistical integrity, missing values in numerical columns were imputed using mean imputation, calculated as:\u003c/p\u003e\n \u003cp\u003e\u003cspan name=\"Emphasis\"\u003ecombined_df\u003c/span\u003e \u0026thinsp; \u003cspan name=\"Emphasis\"\u003e=\u003c/span\u003e \u0026thinsp; \u003cspan name=\"Emphasis\"\u003ecombined_df.fillna(combined_df.mean())\u003c/span\u003e\u003c/p\u003e\n \u003cp\u003eThis technique was used for the important features like current density, cell voltage, power density, pressure, relative humidity, membrane compression, Nafion percentage, Z_real, Z_img, and applied voltage, ensuring that no crucial data points are lost while retaining feature distributions [2]. Test and file non-numeric columns were not subjected to imputation to avoid any errors.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv id=\"Sec6\" class=\"Section3\"\u003e\n \u003ch2\u003e2.2.2 Feature Scaling and Data Splitting\u003c/h2\u003e\n \u003cp\u003eIn order to improve the performance of machine learning models, z-score normalization was used to normalize feature scales so that varying magnitudes would have less influence [3]. In line with standard machine learning procedure, the dataset was split into training (80%) and testing (20%) subsets in order to support a robust test of model performance [1]. The preprocessing pipeline in this manner[11] ensured that the dataset was properly structured, statistically sound, and optimized for later visualization and predictive modeling.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\n \u003ch2\u003e2.3. Correlation Analysis\u003c/h2\u003e\n \u003cp\u003eCorrelation analysis of the PEM Fuel Cell Dataset plays a crucial role in establishing interactions between the various features, including current density, cell voltage, power density, pressure, relative humidity, membrane compression, Nafion percentage, Z_real, Z_img, and applied voltage. Correlation analysis is crucial in guiding machine learning models for fuel cell degradation prediction because the establishment of major feature interactions can have a major impact on model accuracy. Following the preprocessing step, during which missing values were filled with column means and the dataset was normalized, the analysis was conducted using visualizations to reveal hidden patterns and correlations between these variables.The data set, having been taken from Hamidi et al. (\u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e), contains in-depth electrochemical measurements, and thus it serves as a sound basis for this analysis[14\u0026ndash;16]. First, a correlation heatmap was formed to illustrate overall relationships between features, providing an intuitive insight into their interdependencies. Subsequently, specific scatter plots\u0026mdash;Cell Voltage vs. Current Density, Power Density vs. Current Density, Power Density vs. Cell Voltage, Z_img vs. Z_real, and Applied Voltage vs. Set\u0026mdash;were drawn to depict certain feature relationships (Figs. 2\u0026ndash;7). These plots were important to show noteworthy correlations, i.e., negative relationship between the current density and the cell voltage, or between the impedance (Z_real, Z_img) and the performance characteristics of fuel cells, such as the cell voltage and the power density. What was understood through the correlation analysis proved valuable for the process of feature selection in order to provide input into machine learning models that would be used for prediction of fuel cell degradation and including only the most important variables within.\u003c/p\u003e\n \u003cp\u003eBy establishing these correlations, the study helped lead to a more effective and efficient modeling strategy, paving the way for precise fuel cell performance degradation predictions.\u003c/p\u003e\n \u003cp\u003eThe heatmap of correlation (Fig. 2) gives a holistic view of Pearson correlation coefficients among numerical features post-standardization, providing key insights into the interdependencies that impact fuel cell performance. Developed on Python\u0026apos;s seaborn library, the heatmap measures the strength and direction of linear dependencies ranging from \u0026minus;\u0026thinsp;1 (perfect negative correlation) to 1 (perfect positive correlation). A strong negative correlation (e.g., -0.85) was found between current density and cell voltage, as would be expected for the notorious polarization characteristics of PEM fuel cells, in which voltage falls as an increasing current demand induces ohmic, activation, and concentration losses [17]. In contrast, power density was positively correlated with current density (e.g., 0.75), which shows the initial rise in power output with increasing current, before finally levelling off through voltage drop effects. Moreover, the heatmap showed lower correlations between operational conditions like pressure and relative humidity with electrical outputs (correlation coefficients less than 0.3), suggesting that their impacts are probably nonlinear or subject to other interactions. This underscores the requirement for machine learning models that can model intricate dependencies beyond linear correlations [7]. By taking advantage of these findings, the feature selection step was optimized to incorporate the most influential variables while also recognizing parameters that needed more advanced modeling techniques to effectively model fuel cell degradation.\u003c/p\u003e\n \u003cp\u003eIn order to better explore polarization dynamics, the \u0026quot;Cell Voltage vs. Current Density\u0026quot; plot (Fig. 3) was produced as a scatter plot, with current density on the x-axis and cell voltage on the y-axis, color-coded by test conditions (e.g., various pressure levels). This plot shows the common polarization curve, wherein cell voltage goes from about 1.0 V at low current densities (for example, 0.1 A/cm\u0026sup2;) to below 0.6 V for high densities (for example, 1.5 A/cm\u0026sup2;). The sharp initial fall is due to activation losses, followed by linear decline by ohmic resistance and sharper fall at high current densities by mass transport limitations [19]. Variations with test conditions, such as elevated voltages at higher pressure levels (e.g., 25 psig versus 5 psig), underscore pressure\u0026apos;s contribution to greater oxygen availability, reducing electrode losses, and improving fuel cell efficiency. The trends are highly substantial in terms of voltage stability, one of the degradation indicators, further reinforcing the significance of incorporating pressure-dependent effects into machine learning models to accurately predict degradation [2].\u003c/p\u003e\n \u003cp\u003eThe \u0026quot;Power Density vs. Current Density\u0026quot; plot (Fig. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e) gives additional insight into PEM fuel cell performance by showing power density as a function of current density, demonstrating the typical parabolic behavior of fuel cells. Because power density is the product of cell voltage and\u003c/p\u003e\n \u003cp\u003ecurrent density, it begins close to zero at low current levels, increases to a maximum of about 0.4\u0026ndash;0.5 W/cm\u0026sup2; at intermediate current densities (e.g., 0.8\u0026ndash;1.0 A/cm\u0026sup2;), and then decreases as voltage losses become dominant at high currents. Peak power output is slightly different depending on operating conditions\u0026mdash;greater relative humidity (e.g., 100%) increases membrane hydration, lowering resistance, and greater pressure allows for gas diffusion, both contributing to increased fuel cell efficiency [8]. This graph is critical for determining optimal operating conditions and for detecting performance degradation, as deviations from predicted power density values can indicate underlying problems, such as membrane drying, catalyst degradation, or gas diffusion limitations [6].\u003c/p\u003e\n \u003cp\u003eThe \u0026quot;Power Density vs. Cell Voltage\u0026quot; graph (Fig. 5) indicates power output and voltage relationship in a hyperbolic curve characteristic of PEM fuel cells. The maximum power density is at 0.6\u0026ndash;0.7 V, where voltage and current balance produces maximum output. Power is low at high voltages (\u0026gt;\u0026thinsp;0.9 V) because there is not enough current, and power decreases at low voltages (\u0026lt;\u0026thinsp;0.5 V) due to greater losses. This connection highlights the influence of voltage stability on fuel cell performance, thus it is a critical parameter for machine learning-based degradation prediction[1].\u003c/p\u003e\n \u003cp\u003eThe \u0026quot;Zimg vs. Zreal\u0026quot; plot (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e) is an Nyquist plot representing electrochemical impedance (EIS) response. Lower resistance is indicated by smaller arcs under increased pressures (25 psig), and increased resistance by higher arcs at decreased humidity (30%) because the membrane becomes dehydrated [19]. The trends aid in identifying mechanisms for degradation, thereby making z_real and z_img useful for prediction based on machine learning [2].\u003c/p\u003e\n \u003cp\u003eThe \u0026quot;Applied Voltage vs. Set\u0026quot; graph (Fig.\u0026nbsp;7) illustrates applied voltage variation between test sets, illustrating stability and set-specific voltage increments between 0.4 V and 1.0 V. Abrupt decreases (e.g., from 0.8 V to 0.5 V in a set) can signal onset of degradation, e.g., electrode fouling or gas starvation problems[1] [20]. Trends corroborate polarization values, solidifying voltage as one of the predominant degradation indicators in predictive modeling.\u003c/p\u003e\n \u003cp\u003eThe plots (Figs.\u0026nbsp;3\u0026ndash;7) and heatmap (Fig.\u0026nbsp;2) of the data point out salient relationships: the parabolic power_density-current_density curve, the inverse cell_voltage-current_density trend, the voltage-power dependency, impedance\u0026apos;s diagnostic value, and stability of applied_voltage. Correlation coefficients on the heatmap inform feature selection, with strong predictors (e.g., current_density, cell_voltage) being prioritized and nonlinear dependencies (e.g., pressure) acknowledged [1]. Confirmation of these trends and their electrochemical implication is provided by scatter plots (Figs.\u0026nbsp;3\u0026ndash;\u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e) and the line plot (Fig.\u0026nbsp;7) [2]. These observations drive the ML modeling strategy, as voltage deviations are indicative of degradation, and impedance data supports internal state monitoring [13]. The requirement for the use of all features, with tree-based models being adept at modeling nonlinear interactions, provides predictive robustness, in turn connecting preprocessing to data-driven degradation analysis [4].\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\n \u003ch2\u003e2.4. Machine Learning Modeling\u003c/h2\u003e\n \u003cp\u003eFollowing preprocessing and correlation analysis (Sections \u003cspan class=\"InternalRef\"\u003e2.2\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e2.3\u003c/span\u003e), machine learning (ML) modeling was executed to forecast cell voltage in the PEM Fuel Cell Dataset, a marker of fuel cell degrading. Cell voltage was selected as target variable because of its high correlation with current_density and responsiveness to working conditions and degrading mechanisms, evident in the polarization plots (Figs. 3\u0026ndash;5) and impedance data (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e) [18]. There were fifteen ML models trained and tested, utilizing predictors such as current_density, power_density, pressure, humidity, membrane_compression, nafion_percent, z_real, z_img, and applied_voltage [2]. The models were structured into four categories: (1) Linear Models (Linear Regression, Ridge, Lasso, ElasticNet) for linear trends, (2) Tree-Based Models (Decision Tree, Random Forest, Extra Trees) for nonlinearities, (3) Boosting Models (Gradient Boosting, AdaBoost, XGBoost, LightGBM) for iterative refinement, and (4) other models such as KNN, SVR, and MLP for various approaches [3]. Python libraries (scikit-learn, CatBoost, XGBoost, LightGBM) were employed with default hyperparameters for the sake of comparison [4\u0026ndash;8].\u003c/p\u003e\n \u003cp\u003eThe preprocessed dataset, normalized and separated into 80% training and 20% test sets (Section \u003cspan class=\"InternalRef\"\u003e2.2\u003c/span\u003e), was employed for model training. Cell voltage prediction was modeled as a regression problem with the training set utilized to train each model and the test set utilized to measure model performance. Four metrics were utilized to measure model accuracy and fit: Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R\u0026sup2; Score.\u003c/p\u003e\n \u003cp\u003eMAE approximates the mean absolute difference between predicted and true values, providing an easy-to-understand measure of model accuracy:\u003c/p\u003e\n \u003cp\u003e\u003cem\u003emae\u0026thinsp;=\u0026thinsp;mean_absolute_error(y_test, y_pred)\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;..(1)\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eMSE computes the mean squared errors, assigning more importance to larger errors:\u003c/p\u003e\n \u003cp\u003e\u003cem\u003emse\u0026thinsp;=\u0026thinsp;mean_squared_error(y_test, y_pred)\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;(2)\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eRMSE, the square root of MSE, gives a more understandable value with the same unit as the target variable:\u003c/p\u003e\n \u003cp\u003e\u003cem\u003ermse\u0026thinsp;=\u0026thinsp;np.sqrt(mse)\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;.(3)\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eR\u0026sup2; Score is the percentage of variance in the dependent variable that is accounted for by the independent variables, with a 1 representing perfect predictability:\u003c/p\u003e\n \u003cp\u003e\u003cem\u003er2\u0026thinsp;=\u0026thinsp;r2_score(y_test, y_pred)\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;\u0026hellip;.(4)\u003c/em\u003e\u003c/p\u003e\n \u003cp\u003eHere, y_test is the actual cell voltage, and y_pred is the predicted values. These measures give a clear assessment of model performance in terms of absolute error, sensitivity to error, and explanatory power, which are essential in precisely predicting fuel cell degradation (Chai \u0026amp; Draxler, \u003cspan class=\"CitationRef\"\u003e2014\u003c/span\u003e; Hamidi et al., \u003cspan class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n \u003ch2\u003e2.5. Fuel Cell Degradation Prediction\u003c/h2\u003e\n \u003cp\u003eThis study predicts the degradation of Proton Exchange Membrane (PEM) fuel cells by predicting cell voltage with the Extra Trees Regressor (R\u0026sup2; = 0.9962, Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e) trained on the PEM Fuel Cell Dataset [1]. Cell voltage, a parameter sensitive to degradation as evident from the \u0026quot;Cell Voltage vs Current Density\u0026quot; (Fig. 3) and \u0026quot;Zimg vs Zreal\u0026quot; (Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e) plots, is predicted based on features such as current_density and pressure. The model defines a reference voltage and monitors deviations (e.g., \u0026gt;\u0026thinsp;0.05 V) from real values in order to detect possible degradation indicators, including membrane thinning [2]. The \u0026quot;line_applied_voltage\u0026quot; plot (Fig.\u0026nbsp;7) is employed for finding instability trends. Although the lack of time-series data restrains straightforward Remaining Useful Life (RUL) prediction, the method facilitates condition-based monitoring, with future efforts directed towards incorporating temporal data [3].\u003c/p\u003e\n\u003c/div\u003e"},{"header":"3. Result","content":"\u003cp\u003eThis section reports the outcome of the correlation analysis and machine learning (ML) modeling on the PEM Fuel Cell Dataset for predicting cell voltage as a surrogate for degradation in Proton Exchange Membrane (PEM) fuel cells [1]. Derived from preprocessing (Section \u003cspan refid=\"Sec4\" class=\"InternalRef\"\u003e2.2\u003c/span\u003e), correlation analysis (Section \u003cspan refid=\"Sec7\" class=\"InternalRef\"\u003e2.3\u003c/span\u003e), and modeling (Section \u003cspan refid=\"Sec8\" class=\"InternalRef\"\u003e2.4\u003c/span\u003e), the outcome is depicted in figures and encapsulated in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and provides insights into electrochemical performance and predictive capability.\u003c/p\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.1. Performance Analysis of Machine Learning Models\u003c/h2\u003e \u003cp\u003ePerformance of 15 machine learning (ML) models learned to predict cell voltage on the PEM Fuel Cell Dataset was measured by using Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R\u0026sup2; Score, as given in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e [1]. Ensemble tree models, specifically the Extra Trees Regressor (MAE\u0026thinsp;=\u0026thinsp;0.0009878, MSE\u0026thinsp;=\u0026thinsp;2.19E-05, RMSE\u0026thinsp;=\u0026thinsp;0.004676, R\u0026sup2; = 0.9962), performed better than others, followed by Random Forest and Decision Tree Regressor, showing their capacity to learn nonlinear relationships such as between cell_voltage and current_density [8\u0026ndash;9]. Boosted models, such as XGBoost (R\u0026sup2; = 0.9835) and CatBoost (R\u0026sup2; = 0.9668), also worked well, taking advantage of iterative error correction [7]. Linear models such as Linear Regression, Ridge, Lasso, and ElasticNet exhibited moderate performance (R\u0026sup2; ~0.6), whereas SVR and MLP Regressor performed poorly with MLP exhibiting signs of overfitting [15].\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eMachine learning model evaluations to assess MSE, R\u003csup\u003e2\u003c/sup\u003e, MAE and RMSE.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eR\u0026sup2;\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLinear Regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.002194\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.622595\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0138163\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.046836\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRidge Regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.002194\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.622595\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0138163\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.046836\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLasso Regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.002379\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.590674\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0132962\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.048776\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eElasticNet\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.002327\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.59956\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0131465\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.048224\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eDecision Tree Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5.41E-05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.990686\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0015484\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.007358\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.97E-05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.994893\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0011401\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.005448\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eExtra Trees Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e2.19E-05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.996237960328682\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0009878\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.004676\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boosting Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.00142\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.75575\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0109353\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.037678\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAdaBoost Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.000898\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.84555\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0111709\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.029962\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eKNeighbors Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e5.44E-05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.990642\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0014769\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.007375\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eSVR\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.009212\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-0.58498\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0938222\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.095981\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eMLP Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.055707\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-8.58446\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0967775\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.236.24\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCatBoost Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.000193\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.966798\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.0037322\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.013892\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBoost Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e9.58E-05\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.983512\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.005225\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.009789\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLightGBM Regressor\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003e0.000885\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.847757\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003e0.008727\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.029747\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e*Note: MLP R\u0026sup2; anomaly indicates potential overfitting.\u003c/p\u003e \u003cp\u003eThe Extra Trees Regressor\u0026rsquo;s near-perfect R\u0026sup2; suggests its suitability for establishing a degradation baseline, enhancing predictive maintenance [4].\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.2. Visualization of Model Performance\u003c/h2\u003e \u003cp\u003eTo assess the predictive capabilities of the 15 machine learning (ML) models that were trained on the PEM Fuel Cell Dataset for predicting cell voltage\u0026mdash;a critical measure of degradation in Proton Exchange Membrane (PEM) fuel cells\u0026mdash;a combined visualization (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e8\u003c/span\u003e) was obtained, combining bar plots of Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R\u0026sup2; Score for all models [4]. The Extra Trees Regressor performed best with lowest MAE (0.0009878) and MSE (2.19E-05), performing significantly better than Linear Regression (MAE\u0026thinsp;=\u0026thinsp;0.013816) and Ridge (0.013816). SVR and MLP Regressor were seen performing poorly, MLP being a case of overfitting by its strange interaction with R\u0026sup2; [11]. The visualization indicates these trends, confirming the better capacity of tree-based and boosted models to identify intricate, nonlinear trends in the data, essential for monitoring degradation [1].\u003c/p\u003e \u003cp\u003eThe RMSE panel in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e8\u003c/span\u003e, converting errors to units of voltage, points out the accuracy of the Extra Trees Regressor (0.004676), then Random Forest (0.005448) and Decision Tree (0.007358). Boosting algorithms such as XGBoost (0.009789) and CatBoost (0.013892) follow closely, with SVR (0.095981) and MLP (0.236024) performing much larger errors, indicating difficulties in modeling the dataset's nonlinearities [9]. Linear models (such as Linear Regression, 0.046836) show moderate RMSE values, which align with their R\u0026sup2; values [3]. The R\u0026sup2; plot illustrates that Extra Trees (0.9962), Random Forest (0.9949), and Decision Tree (0.9907) capture the majority of the variance, with boosting models such as XGBoost (0.9835) and CatBoost (0.9668) performing well as well. Linear models are around 0.6 (e.g., Linear Regression, 0.6226), whereas MLP's R\u0026sup2; (8.5845) is an outlier, probably because of overfitting or numerical errors [12]. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e8\u003c/span\u003e supports the superiority of tree-based ensemble models, consistent with the nonlinear relationships seen in Figs.\u0026nbsp;2\u0026ndash;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e6\u003c/span\u003e (Section \u003cspan refid=\"Sec7\" class=\"InternalRef\"\u003e2.3\u003c/span\u003e). Their minimal error rates and maximum R\u0026sup2; values indicate good prospects for creating a degradation baseline, in which deviations from anticipated cell voltage might indicate performance degradation. Visualization justifies the choice of Extra Trees as the best model for real-world PEM fuel cell diagnostics [6].\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.3. Model Fit and Residual Analysis\u003c/h2\u003e \u003cp\u003eTo improve comparison of the performance of the 15 ML models in cell voltage prediction for tracking PEM fuel cell degradation, \"Actual vs Predicted\" and \"Residual Plots\" were merged into one figure (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e), along with Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e8\u003c/span\u003e [1]. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e, developed through Python's matplotlib, visually assesses prediction accuracy and error distribution for all models. Random Forest (R\u0026sup2; = 0.9949) and Extra Trees (R\u0026sup2; = 0.9962) lie close to the y\u0026thinsp;=\u0026thinsp;x line with very little scatter, mirroring their tiny MAE (0.0011401 and 0.0009878) and strong performance [2]. Linear models like Linear Regression (R\u0026sup2; = 0.6226) have a larger scatter, particularly at lower voltages, representing lower accuracy [3]. Decision Tree, XGBoost, and CatBoost have good fits with negligible deviations. SVR and MLP exhibit strong dispersion in their \"Actual vs Predicted\" plots, pointing out weaknesses [4]. The \"Residual\" panels reveal tight clusters around zero for Extra Trees and Random Forest, indicating little error, whereas linear models exhibit heteroscedasticity, especially at lower predictions [5]. MLP's residuals are dispersed, lending support to overfitting problems.\u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e illustrates the better performance of tree-based models in cell voltage prediction of degradation, as expected from their high R\u0026sup2; and low errors (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) [1]. To compare 15 ML models for cell voltage prediction in PEM fuel cells, \"Actual vs Predicted\" and \"Residual Plots\" were merged into Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e, supported by Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e8\u003c/span\u003e [1]. This multi-panel plot, generated through Python's matplotlib, visually evaluates prediction accuracy and error distribution across models such as Linear Regression, Random Forest, Extra Trees, XGBoost, and so on [2]. In the \"Actual vs Predicted\" plots, Extra Trees (R\u0026sup2; = 0.9962) and Random Forest (R\u0026sup2; = 0.9949) have negligible scatter and low MAE (0.0009878 and 0.0011401, respectively) [3]. Decision Tree (R\u0026sup2; = 0.9907) and boosting models (XGBoost, CatBoost) also fit well with slight deviations. Linear models have more scatter, particularly at low voltages, which indicates their inability to track nonlinear trends [4]. The \"Residual\" panels have little systematic error for tree-based models, while linear models display heteroscedasticity, and MLP and SVR models have severe overfitting or poor fit [5][6].\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eFigure \u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e confirms the supremacy of tree models in prediction of degradation, augmenting their large R\u0026sup2; and few errors (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), and sanctions their uses in cell voltage divergences monitoring [2].\u003c/p\u003e \u003c/div\u003e"},{"header":"Conclusion","content":"\u003cp\u003eThis research assessed machine learning (ML) utilization in predicting fuel cell degradation in Proton Exchange Membrane (PEM) fuel cells through forecasting cell voltage as a surrogate for degradation with the PEM Fuel Cell Dataset [2].The dataset, comprising polarization (current_density, cell_voltage, power_density) and impedance (z_real, z_img) data under varying conditions (pressure, relative_humidity), was preprocessed to replace missing values by column means [14]. Correlation analysis (Section \u003cspan refid=\"Sec7\" class=\"InternalRef\"\u003e2.3\u003c/span\u003e) identified significant correlations, graphed as plots like \"Cell Voltage vs Current Density\" (Fig.\u0026nbsp;3) and \"Zimg vs Zreal\" (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e6\u003c/span\u003e), indicative of operating effects on performance. These observations were used to train 15 ML models (Section \u003cspan refid=\"Sec8\" class=\"InternalRef\"\u003e2.4\u003c/span\u003e), ranging from linear models (e.g., Linear Regression) to ensemble models (e.g., Extra Trees Regressor), performance being measured on the basis of MAE, MSE, RMSE, and R\u0026sup2; scores (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe outcomes (Section \u003cspan refid=\"Sec10\" class=\"InternalRef\"\u003e3\u003c/span\u003e) support the excellence of tree-based ensemble methods in the prediction of degradation, with Extra Trees Regressor giving the best predictions (MAE\u0026thinsp;=\u0026thinsp;0.0009878, R\u0026sup2; = 0.9962), followed by Random Forest (R\u0026sup2; = 0.9949) and Decision Tree (R\u0026sup2; = 0.9907) [1]. Boosting models such as XGBoost (R\u0026sup2; = 0.9835) and CatBoost (R\u0026sup2; = 0.9668) also worked well, whereas linear models like Linear Regression (R\u0026sup2; = 0.6226) and SVR (R\u0026sup2; = 0.584978) were not very effective, unable to identify nonlinearities [2]. The peculiar R\u0026sup2; value of the MLP Regressor (8.584462) reflected overfitting, as verified using residual analysis [3Visualisations of performance (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e8\u003c/span\u003e) and residuals (Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e9\u003c/span\u003e) ensured Extra Trees' accuracy and minimum residual scatter that is appropriate for monitoring degradation [4]. Such models, like Extra Trees, are used within the degradation forecasting framework (Section \u003cspan refid=\"Sec9\" class=\"InternalRef\"\u003e2.5\u003c/span\u003e) to calculate a base cell voltage value and mark strong deviations (\u0026gt;\u0026thinsp;0.05 V) as indicative of degradation effects, such as membrane thinning or catalyst wearing [5]. Through monitoring of performance anomalies and stability trends (Fig.\u0026nbsp;7) [6], the method facilitates condition-based maintenance, though the static dataset restricts direct Remaining Useful Life (RUL) prediction. The research illustrates the capability of ML to enhance PEM fuel cell reliability, with Extra Trees providing near-perfect predictions, consistent with earlier data-driven diagnostic models [7].\u003c/p\u003e \u003cp\u003eLimitations of this study are because of the absence of degradation data in the dataset, restricting long-term prediction [1]. Models were trained on default hyperparameters, which can decrease performance [1]. Although methods like mean imputation were used, exploration of other methods like interpolation would improve the management of data [2]. The MLP anomaly is deserving of further investigation of its architecture or scaling modifications [3]. Future studies would be able to make direct predictions of Remaining Useful Life (RUL) utilizing time series or cycle count as inputs. Improving accuracy along with interpretability could be made possible through cross-validation and hyperparameter optimization involving hybrid models by merging machine learning and physics-based methods [4]. Augmentation of the data set with supplementary operating conditions or degradation states will further confirm the framework, accelerating its real-world application in fuel cell systems. In total, this research is a good foundation for ML-based diagnosis of PEM fuel cells and presents a scalable method for increasing their durability and efficiency [5].\u003c/p\u003e"},{"header":"Declarations","content":"\u003ch2\u003eCode Availability\u003c/h2\u003e \u003cp\u003eThe code along with the model plots, ready to use models are available at:\u003c/p\u003e \u003cp\u003e \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/I-Deepanshu/Fuel-cell-degradation-predictor\u003c/span\u003e\u003cspan address=\"https://github.com/I-Deepanshu/Fuel-cell-degradation-predictor\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e \u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eHamidi, S., Haghighi, S., \u0026amp; Askari, K. (2020). \u003cem\u003ePEM Fuel Cell Dataset\u003c/em\u003e. Journal of Energy Engineering, 146(6), 04020089. https://doi.org/10.1061/(ASCE)EY.1943-7897.0000703\u003c/li\u003e\n\u003cli\u003eBatista, G., \u0026amp; Monard, M. (2003). An analysis of four algorithms for classification and prediction of missing values. \u003cem\u003eComputational Statistics \u0026amp; Data Analysis, 44\u003c/em\u003e(1-2), 163-181. https://doi.org/10.1016/S0167-9473(03)00006-1\u003c/li\u003e\n\u003cli\u003eSch\u0026ouml;lkopf, B., Platt, J. C., Shawe-Taylor, J., Smola, A. J., \u0026amp; Williamson, R. C. (1999). Estimating the Support of a High-Dimensional Distribution. \u003cem\u003eNeural Computation, 13\u003c/em\u003e(7), 1443\u0026ndash;1471. https://doi.org/10.1162/089976699300016428\u003c/li\u003e\n\u003cli\u003eFerronato, N., \u0026amp; Torretta, V. (2019). Waste Mismanagement in Developing Countries: A Review of Global Issues. \u003cem\u003eInternational Journal of Environmental Research and Public Health, 16\u003c/em\u003e(6), 1074. https://doi.org/10.3390/ijerph16061074\u003c/li\u003e\n\u003cli\u003eLiu, H., Zhou, M., \u0026amp; Zhang, X. (2021). Data-driven diagnostics and prognostics of fuel cells: A review of methodologies and applications. \u003cem\u003eEnergy Reports, 7\u003c/em\u003e, 561-576. https://doi.org/10.1016/j.egyr.2021.03.073\u003c/li\u003e\n\u003cli\u003eBreiman, L. (2001). Random forests. \u003cem\u003eMachine Learning, 45\u003c/em\u003e(1), 5-32. https://doi.org/10.1023/A:1010933404324\u003c/li\u003e\n\u003cli\u003eHastie, T., Tibshirani, R., \u0026amp; Friedman, J. (2009). \u003cem\u003eThe elements of statistical learning: Data mining, inference, and prediction\u003c/em\u003e (2nd ed.). Springer. https://doi.org/10.1007/978-0-387-84858-7\u003c/li\u003e\n\u003cli\u003ePedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., \u0026amp; Duchesnay, E. (2011). Scikit-learn: Machine learning in Python. \u003cem\u003eJournal of Machine Learning Research, 12\u003c/em\u003e, 2825\u0026ndash;2830. http://www.jmlr.org/papers/volume12/pedregosa11a/pedregosa11a.pdf\u003c/li\u003e\n\u003cli\u003eChai, T., \u0026amp; Draxler, R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? \u003cem\u003eGeoscientific Model Development, 7\u003c/em\u003e, 1247\u0026ndash;1250. https://doi.org/10.5194/gmd-7-1247-2014\u003c/li\u003e\n\u003cli\u003eYuan, X., Wang, H., \u0026amp; Sun, J. (2008). Impedance measurement and equivalent circuit modeling for Proton Exchange Membrane (PEM) fuel cells. \u003cem\u003eJournal of Power Sources, 179\u003c/em\u003e(1), 146\u0026ndash;154. https://doi.org/10.1016/j.jpowsour.2007.09.074\u003c/li\u003e\n\u003cli\u003eLarminie, J., \u0026amp; Dicks, A. (2003). \u003cem\u003eFuel cell systems explained\u003c/em\u003e (2nd ed.). Wiley.\u003c/li\u003e\n\u003cli\u003eJouin, M., Youssef, J., \u0026amp; Tisseyre, R. (2016). Prognostics for PEM fuel cell degradation: A review. \u003cem\u003eEnergy Procedia, 105\u003c/em\u003e, 503\u0026ndash;507. https://doi.org/10.1016/j.egypro.2016.11.057\u003c/li\u003e\n\u003cli\u003ePei, P., Chen, H., \u0026amp; Wu, J. (2011). Degradation study of Proton Exchange Membrane (PEM) fuel cells: An overview. \u003cem\u003eJournal of Power Sources, 196\u003c/em\u003e(4), 1780\u0026ndash;1790. https://doi.org/10.1016/j.jpowsour.2010.08.113\u003c/li\u003e\n\u003cli\u003eFerraro, A., \u0026amp; Mazzarotta, B. (2017). Monitoring the degradation of PEM fuel cells by voltage recovery and monitoring of current density. \u003cem\u003eInternational Journal of Hydrogen Energy, 42\u003c/em\u003e(25), 16322\u0026ndash;16330. https://doi.org/10.1016/j.ijhydene.2017.05.086\u003c/li\u003e\n\u003cli\u003eSaha, B., \u0026amp; Kamaruddin, S. (2016). A review on predictive maintenance of PEM fuel cells. \u003cem\u003eRenewable and Sustainable Energy Reviews, 56\u003c/em\u003e, 91\u0026ndash;103. https://doi.org/10.1016/j.rser.2015.11.071\u003c/li\u003e\n\u003cli\u003ePandian, M., \u0026amp; Natarajan, K. (2020). Optimization of Proton Exchange Membrane Fuel Cell for Performance and Efficiency: A Review. \u003cem\u003eJournal of Energy, 2020\u003c/em\u003e, 3248154. https://doi.org/10.1155/2020/3248154\u003c/li\u003e\n\u003cli\u003ePardo, J. A., \u0026amp; Carmona, R. (2017). Condition monitoring and performance prediction of fuel cells: Data-driven models. \u003cem\u003eInternational Journal of Hydrogen Energy, 42\u003c/em\u003e(12), 8536\u0026ndash;8547. https://doi.org/10.1016/j.ijhydene.2017.03.011\u003c/li\u003e\n\u003cli\u003eYan, Y., \u0026amp; Wang, H. (2018). Development and application of diagnostic techniques in PEM fuel cell: A review. \u003cem\u003eEnergy Conversion and Management, 165\u003c/em\u003e, 47-58. https://doi.org/10.1016/j.enconman.2018.03.040\u003c/li\u003e\n\u003cli\u003eAhmed, S., \u0026amp; Zhang, Y. (2019). A comparative review of diagnostic and prognostic methods for PEM fuel cells. \u003cem\u003eRenewable and Sustainable Energy Reviews, 101\u003c/em\u003e, 122\u0026ndash;135. https://doi.org/10.1016/j.rser.2018.09.014\u003c/li\u003e\n\u003cli\u003eXie, X., \u0026amp; Zeng, Q. (2018). A novel degradation model for PEM fuel cell performance prediction based on hybrid machine learning techniques. \u003cem\u003eEnergy, 150\u003c/em\u003e, 80\u0026ndash;89. https://doi.org/10.1016/j.energy.2018.01.022\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"PEM fuel cell, degradation forecasting, machine learning, electrochemical analysis, predictive maintenance","lastPublishedDoi":"10.21203/rs.3.rs-6710108/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6710108/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eProton Exchange Membrane (PEM) fuel cells offer great potential in terms of green energy solutions but degrade over a period of time based on diverse operation and material aspects. Based on the data obtained from PEM Fuel Cell Dataset with polarization as well as impedance at diverse operations, this study makes use of machine learning algorithms for predicting degradation of fuel cell. Following preprocessing by imputation of missing values via mean imputation, exploratory data analysis was performed via heatmaps and key visualizations. Fifteen machine learning models involving linear and nonlinear regressors, decision trees, ensemble models, and neural networks were trained to predict important performance metrics like cell voltage, power density, and impedance characteristics. Model performance was conducted with Mean Absolute Error (MAE), Mean Squared Error (MSE), Root Mean Squared Error (RMSE), and R\u0026sup2; Score, where the Extra Trees Regressor performed best with an MAE of 0.00099 and an R\u0026sup2; score of 0.996. These results show the potential of machine learning to forecast fuel cell degradation, enabling proactive maintenance and increased system reliability. Future work will explore real-time deployment of predictive models for enhanced operational effectiveness in real-world systems.\u003c/p\u003e","manuscriptTitle":"Fuel Cell Degradation Prediction Using Machine Learning Models: A Study on Proton Exchange Membrane (PEM) Fuel Cell Dataset","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-05-23 07:57:13","doi":"10.21203/rs.3.rs-6710108/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0c7a7ba4-7c28-4681-9065-c69643b65365","owner":[],"postedDate":"May 23rd, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2025-05-23T07:57:13+00:00","versionOfRecord":[],"versionCreatedAt":"2025-05-23 07:57:13","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-6710108","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6710108","identity":"rs-6710108","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00