A novel forecasting model for crude oil prices that integrates CEEMDAN-VMD multiscale decomposition with an Attention-based Bidirectional LSTM network | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article A novel forecasting model for crude oil prices that integrates CEEMDAN-VMD multiscale decomposition with an Attention-based Bidirectional LSTM network Haotian Song, Hengrui Zhang, Donghe Li This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8242442/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Accurate prediction of WTI crude oil prices is of great significance for the business decision-making of oil and gas enterprises, the formulation of national energy strategies, and the risk management of the global financial market. However, current traditional prediction methods have some limitations in predicting WTI crude oil prices. For instance, traditional methods are insufficient in integrating complex factors that affect international oil prices, such as geopolitical conflicts, global supply and demand imbalances, and financial speculation. They also have limited ability to extract nonlinear and time-varying correlation features between oil prices and multiple influencing factors, and do not adequately consider data noise. As a result, it is difficult to clearly explain the contribution mechanism of key influencing factors to the prediction results of oil prices, and the model's interpretability is insufficient. To resolve these challenges, this paper integrates Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN), Variational Mode Decomposition (VMD), Attention Mechanism (Attention), and Bidirectional Long Short-Term Memory (BiLSTM), and proposes a deep learning-based hybrid prediction model (CEEMDAN-VMD-Attention-BiLSTM). Specifically, the non-stationary and non-linear characteristics of the WTI oil price series are decomposed into several stationary sub-series by using CEEMDAN-VMD, reducing data noise and capturing the non-linear relationship between crude oil prices and macroeconomic variables. An Attention-BiLSTM model is constructed to predict the sub-series of oil prices decomposed by CEEMDAN-VMD, and the predicted values of these sub-series are summed to reconstruct the final predicted value. In order to augment the interpretability of the model's forecast results, the SHAP method is adopted to quantify the contribution of different input parameters to the model's prediction results. Based on 28 years of time series data, the study shows that the MAPE of the proposed hybrid prediction model is 7.66%, and the R² is 0.9665. The proposed model demonstrates superior predictive accuracy and notably robust performance in comparative analysis. Through SHAP analysis, the top 5 key factors influencing international oil prices are Brent Crude Oil Price, LBMA Gold Price, Federal Funds Effective Rate, RMB-USD Exchange Rate, and Henry Hub Natural Gas Spot Price. The proposed model helps countries grasp the trend of the crude oil market and provides scientific basis for the formulation of energy policies. Physical sciences/Engineering Physical sciences/Mathematics and computing WTI crude oil price forecast Deep learning CEEMDAN-VMD Attention-BiLSTM Interpretability Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 1. Introduction WTI crude oil, as one of the core benchmarks of the global crude oil pricing system, plays a significant role in various fields such as energy trade, financial markets, macroeconomic regulation, and industrial production. It not only reflects the supply and demand situation of the global energy market but is also an important investment asset in the financial market. As a commodity, the supply and demand dynamics and price changes of WTI crude oil have a significant impact on the economic security, monetary policy formulation, and geopolitical landscape of various countries around the world [ 1 ]. As one of the most actively traded commodities globally, the formation mechanism of WTI crude oil prices is complex and variable. Taking the recent market as an example, the price fluctuations are influenced by multiple factors such as geopolitical conflicts, OPEC + production policies, the trend of the US dollar exchange rate, and global macroeconomic expectations. The relevant data exhibit characteristics such as information redundancy, noise interference, and non-linearity, which greatly increase the difficulty of oil price prediction [ 2 ]. For oil-importing countries, every increase of $ 10 per barrel in WTI oil prices may mean that the country needs to pay several billion more dollars in import costs each year. By accurately predicting oil prices, market participants can make more informed investment decisions and risk management strategies, such as optimizing inventory management, formulating hedging strategies, and adjusting energy policies, etc. [ 3 ]. In recent years, with the rapid development of artificial intelligence technology, machine learning methods have been widely applied in the field of oil price prediction due to their powerful feature extraction capabilities and outstanding predictive performance [ 4 – 5 ]. Figure 1 shows the trend of the number of papers on machine learning methods published in the field of oil price prediction from 2005 to 2024, which is derived from the Web of Science database. Through statistical analysis, it can be seen that the research on machine learning methods in the international oil price prediction field has shown an explosive growth trend. With the development of quantitative investment and intelligent decision-making systems, the application and development of data-driven machine learning methods in the crude oil market have become inevitable. Given the intrinsic non-stationary and non-linear features of the WTI crude oil price time series, the accurate forecasting of crude oil prices and their volatility is recognized as a prominent challenge. Currently, the methods for predicting crude oil prices mainly include traditional statistical-mathematical methods and artificial intelligence methods. Commonly used statistical-mathematical methods include Autoregressive Integrated Moving Average (ARIMA) [ 6 ], Exponential Smoothing (ETS) [ 7 ], Vector Autoregression (VAR) [ 8 ], and Generalized ARCH (GARCH) [ 9 ]. Although these conventional econometric approaches can produce precise prediction outcomes under the assumption of approximately linear stationarity of time series, the actual crude oil price series are non-linear and non-stationary. Therefore, traditional regression methods perform poorly in predicting non-stationary and non-linear time series data. Compared with traditional methods, machine learning methods can adapt to complex nonlinear relationships and have advantages such as high computational efficiency, suitability for real-time prediction, strong adaptability and continuous learning ability. Many scholars have shifted the focus of crude oil price prediction research to artificial intelligence methods. Support Vector Regression (SVR) [ 10 ], Extreme Gradient Boosting (XGBoost) [ 11 ], and various other fundamental machine learning methods [ 12 – 13 ] have been developed. However, the aforementioned machine learning methods still have problems such as local optimization, limited generalization ability, and noise interference. Compared with conventional machine learning methods, deep learning, as a promising field of artificial intelligence methods, has garnered widespread attention among researchers in recent years. It has obvious advantages in extracting nonlinear features of time series and fitting generalization ability. Among them, the Long Short-Term Memory Neural Network (LSTM), as a refined variant of the Recurrent Neural Network (RNN), is capable of effectively tackling the problems of gradient explosion or gradient disappearance that occur during the practical implementation of RNNs. Many researchers have adopted this method for predictive applications [ 14 – 15 ]. However, a single LSTM can only capture temporal dependencies in a single direction and is unable to effectively separate noise and interpret its intricate multi-scale features internally. Therefore, BiLSTM is adopted as the core for prediction, synchronously integrating historical and future dependency features to generate a more comprehensive bidirectional hidden state, and comprehensively capturing the long-term evolution patterns and dynamic dependency relationships of the sequence [ 16 ]. Additionally, at the prediction mechanism level, considering that BiLSTM cannot distinguish the importance differences of different historical input pairs for the current prediction, the Attention mechanism is introduced. It can automatically assign higher weights to key time points (such as policy changes, geopolitical conflict periods) and core features, enabling the model to autonomously focus on the most relevant key historical information related to the current prediction, effectively suppressing the interference of irrelevant noise[ 17 ]. Recent application studies have demonstrated that the data decomposition algorithm, by decomposing the original noisy data sequence into multiple sub-sequences, can effectively improve the prediction accuracy of the model. Therefore, in order to reduce the impact of data noise on the prediction results, A large number of researchers have put forward diverse data decomposition techniques. In the hybrid model, the main features of the data sequence are identified and extracted by combining some data decomposition techniques. These methods include Wavelet Transform (WT) [ 18 ], Empirical Mode Decomposition (EMD) [ 19 ], and Ensemble Empirical Mode Decomposition (EEMD) [ 20 ]. By means of data decomposition methods, the key features of the original data sequence are extracted, which in turn improves the prediction accuracy of the model. However, the above-mentioned methods still cannot efficiently reduce the impact of data noise on the model's predictive performance. Therefore, a two-level data decomposition strategy using CEEMDAN and VMD was employed to conduct a deep decomposition of the data. Firstly, CEEMDAN was utilized to adaptively decompose the original signal into a series of intrinsic mode functions, initially eliminating modal aliasing. Subsequently, VMD was applied to the high-frequency subsequence components for a secondary optimization decomposition. This processing significantly reduced the non-stationarity and complexity of the original data, transforming a complex prediction problem into multiple relatively simple and stable subsequence prediction problems, and greatly reducing the impact of noise interference on prediction accuracy [ 21 ]. Therefore, this study employs the "two-level decomposition and denoising + bidirectional dependency modeling + attentional precise focusing" collaborative mechanism of the CEEMDAN-VMD-BiLSTM-Attention model to overcome the inherent limitations of the LSTM model. This provides a comprehensive technical solution for achieving higher accuracy and stronger robustness in predictions. However, in current engineering applications, the research results of many machine learning methods often present in a "black box" form, with the mapping relationship between input and output lacking interpretability. Although these models can achieve accurate predictions, on account of the inadequate interpretability of the prediction results, their application in some high-risk fields is limited. Therefore, strengthening the research on the interpretability of AI technology is of great significance. In this field, interpretability research is relatively scarce. For the purpose of enhancing the model’s interpretability, the Shapley Additive Explanations (SHAP) method was applied to examine the model’s interpretability, analyzing the significance of input features for the prediction results [ 22 ]. This study aims to construct a WTI crude oil price prediction model using the CEEMDAN-VMD-BiLSTM-Attention hybrid model. To enhance the interpretability and novelty of the research, the SHAP values are calculated to analyze the importance of input features to the prediction results, thus enhancing the model’s interpretability. The organization of the subsequent sections in this paper is outlined as follows. Section 2 presents an overview of the employed methods, and details the key aspects of data collection and model construction. Section 3 assesses the predictive performance of the proposed model. Section 4 outlines the conclusions, underscoring the key findings derived from the research. 2. Individual methods This section presents a concise overview of the multiple machine learning methods utilized for constructing the hybrid model, and elaborates on the key specifics during the processes of data collection, data cleaning, and model construction. 2.1. CEEMDAN-VMD data decomposition algorithm CEEMDAN-VMD is a dual decomposition technique for non-stationary time series. Its core principle is based on the cascaded application of complementary ensemble empirical mode decomposition (CEEMDAN) and variational mode decomposition (VMD). As shown in Fig. 2 , CEEMDAN effectively suppresses the mode aliasing problem of traditional empirical mode decomposition (EMD) by introducing adaptive Gaussian white noise into the original signal and performing multiple ensemble averaging. It decomposes the signal into multiple intrinsic mode functions (IMF) with different frequency characteristics and a residual component [ 23 ]. This step can initially separate the high-frequency fluctuations, periodic terms and trend terms in the signal, diminishing the non-stationarity present in the raw data. Subsequently, VMD constructs a variational optimization model to further decompose each IMF component obtained by CEEMDAN into several modal components with limited bandwidth. Figure 3 shows the subsequence obtained by VMD decomposing IMF1. The core idea of VMD is to minimize the sum of the estimated bandwidths of each mode, adaptively determine the center frequency and bandwidth of each mode, and thereby achieve a refined division of the signal in the frequency domain. This secondary decomposition can more accurately extract the complex frequency components hidden in the IMF, such as further separating the noise in the high-frequency IMF from the effective information, or refining the periodic characteristics of the intermediate frequency components [ 24 – 25 ]. In the process of crude oil price prediction, the dual decomposition strategy of CEEMDAN-VMD can effectively decompose the nonlinear and multi-scale fluctuations in the oil price sequence, providing purer and more feature-specific input subsequences for the subsequent Attention-BiLSTM model, thereby improving the prediction accuracy. This combined method, through multi-level signal decomposition, not only retains the robustness of CEEMDAN against noise but also leverages the advantages of VMD in frequency domain optimization, enabling it to serve as an effective tool for addressing complex time-series data. 2.2. Attention-BiLSTM After the signal decomposition is completed by CEEMDAN-VMD, Attention-BiLSTM further refines the modeling of multi-scale subsequences. Attention-BiLSTM is a deep learning architecture that integrates bidirectional temporal modeling and dynamic feature selection capabilities. Among them, BiLSTM processes time series data in parallel through two LSTM layers, forward and backward, capable of simultaneously capturing the bidirectional dependencies of past and future in the oil price sequence. This bidirectional information integration can more comprehensively depict the temporal dynamics of oil price fluctuations, especially suitable for complex scenarios influenced by long-cycle factors such as geopolitics and economic policies. Attention, on the other hand, calculates the "importance score" at each time step to dynamically adjust the model's focus on information at different moments [ 26 ]. 2.3. Construction of residual strength prediction model This study proposes a hybrid CEEMDAN-VMD-Attention-BiLSTM model designed to predict WTI crude oil prices. The prediction framework is shown in Fig. 4 . This framework includes data collection, preprocessing, model construction, prediction accuracy evaluation and model interpretability analysis. In terms of data collection, time series data from 1996 to 2024 were collected on a daily basis. Eight relevant factors were selected as input parameters, including Brent Crude Oil Price, CBOT Volatility Index, Economic Policy Uncertainty Index, Federal Funds Effective Rate, Henry Hub Natural Gas Spot Price, LBMA Gold Price, NASDAQ-100 Index, and RMB-USD Exchange Rate data. Take the price of WIT Crude oil as the output parameter of the model. In terms of data processing, the non-stationary and non-linear characteristics of the WTI oil price sequence were decomposed into several stationary sub-sequences using CEEMDAN-VMD, thereby reducing data noise and capturing the nonlinear relationship between crude oil prices and relevant variables. Since different variables have different magnitudes and measurement units, Min-Max normalization was adopted to normalize the data to the range of [-1, 1], in order to improve the training efficiency of the model [ 27 ]. In terms of model construction, due to the advantages of grid search (GS) in finding optimal solutions, stable and reliable results, and easy to understand search results, the GS method is used to optimize the hyperparameters of the model [ 28 ]. Using Attention BiLSTM to construct a bidirectional temporal and attention weighted prediction model, sum up the predicted values of the oil price subsequence obtained by CEEMDAN-VMD decomposition, and reconstruct the final predicted value. In terms of model performance evaluation, five mainstream models including RE, R 2 , MAPE, RMSE, and MAE are employed to evaluate the model's forecasting accuracy [ 29 – 30 ]. Finally, SHAP is used to study the interpretability of the model, quantify the impact of input variables on the output results of the model, and improve the interpretability of the model. 3. Case study 3.1. Evaluation of predictive performance To evaluate the predictive performance of the proposed hybrid model, the CEEMDAN-VMD-Attention-BiLSTM hybrid model was compared with 5 other models. Figure 5 shows the histogram of the prediction error distribution for the six models. The analysis indicates that the mean and standard deviation of the prediction errors of the LSTM model are the highest, followed by BiLSTM, Attention-BiLSTM, VMD-Attention-BiLSTM, and CEEMDAN-Attention-BiLSTM. Compared with the other five models, the CEEMDAN-VMD-Attention-BiLSTM hybrid model has the smallest mean and standard deviation of the prediction errors and exhibits better predictive performance. To further assess the predictive capability of the hybrid model and examine the fitting effects of different models, the research results are shown in Figs. 6 – 7 . Through Fig. 6 , it can be observed that the predicted values of the CEEMDAN-VMD-Attention-BiLSTM model are the closest to the actual data, but the fitting effect between the predicted results and the actual data cannot be directly quantified. Therefore, a further evaluation analysis of the predictive performance was conducted by combining R 2 and MAPE, and the results are shown in Fig. 7 . The R 2 of the LSTM model is 0.5383, and the MAPE is 14.5%. The R 2 of the BiLSTM model is 0.6441, and the MAPE is 13.76%. The R 2 of the Attention-BiLSTM model is 0.7365, and the MAPE is 11.8%. The R 2 of the VMD-Attention-BiLSTM model is 0.9457, and the MAPE is 9.93%. The R 2 of the CEEMDAN-Attention-BiLSTM model is 0.96, and the MAPE is 9.27%. The R 2 of the CEEMDAN-VMD-Attention-BiLSTM model is 0.9665, and the MAPE is 7.66%. The study found that the prediction effect of LSTM is the worst, while the prediction effects of BiLSTM and Attention-BiLSTM gradually improve. The prediction performance of VMD-Attention-BiLSTM, CEEMDAN-Attention-BiLSTM, and CEEMDAN-VMD-Attention-BiLSTM has been further improved. The research proves that the selected modules in the constructed hybrid model CEEMDAN-VMD-Attention-BiLSTM are all reasonable, and each module is very important for the purpose of improving the predictive performance of the model. Taking the RMSE evaluation index as an example, comparing the prediction results of LSTM and BiLSTM, it is found that the RMSE of BiLSTM is 11.7161%, and its prediction performance is better than that of LSTM, proving that BiLSTM can learn long-term bidirectional dependency relationships and effectively handle time series data, improving the prediction accuracy of the model. Comparing the prediction results of BiLSTM and Attention-BiLSTM, the latter has an RMSE of 11.1824, and its prediction performance is better than that of BiLSTM, proving that the attention mechanism can dynamically and selectively utilize historical information, enhancing the influence of important information on the output of BiLSTM, and improving the prediction accuracy. Comparing the prediction results of Attention-BiLSTM, VMD-Attention-BiLSTM, CEEMDAN-Attention-BiLSTM and CEEMDAN-VMD-Attention-BiLSTM, the prediction effect of BiLSTM is the worst, while the RMSE of CEEMDAN-VMD-Attention-BiLSTM model is 5.103, which has the best prediction performance. This further proves that the data decomposition algorithm is crucial in improving the prediction performance of time series. The CEEMDAN and VMD signal decomposition stages can effectively handle the non-linearity and non-stationarity of the crude oil price sequence, diminish the impact of noise interference, and the CEEMDAN-VMD double-layer data decomposition algorithm brings the best effect. For the MAE evaluation index, similar conclusions can also be found. Overall, the CEEMDAN-VMD-Attention-BiLSTM model advances predictive capability by integrating cutting-edge signal decomposition methods with deep learning frameworks. Compared with existing models, the established hybrid model has the best prediction performance. 3.2. SHAP analysis At present, with the increasing complexity of machine learning models, the problem of the lack of interpretability caused by the "black box" feature of models has become more and more prominent. Conducting SHAP (SHapley Additive exPlanations) analysis is precisely the key link to solve this predicament and ensure the reliable application of models. SHAP is based on the Shapley principle in game theory and can assign fair and consistent importance weights to each feature[ 31 ]. It not only quantifies the contribution of individual features to the model's prediction results but also avoids the limitations of traditional feature importance analysis methods that only reflect global contributions and cannot explain local prediction logic. Enable developers and users to clearly understand "why the model makes such predictions", thereby building trust in the model and reducing the risk of decision-making errors caused by the model's "black box"[ 32 ]. The absolute value of SHAP is used to quantify the influence of various input parameters on the price of WTI crude oil. The greater the absolute value of SHAP, the greater the impact of the parameter on the WTI crude oil price. By means of correlation analysis, there are 5 factors with a correlation coefficient greater than 0.3. Therefore, SHAP analysis is carried out based on the above five types of parameters. As shown in Fig. 9 , The top five factors influencing the WTI Crude Oil Price, in descending order of their influence, are Brent Crude Oil Price, LBMA Gold Price, Federal Funds Effective Rate, RMB-USD Exchange Rate, and Henry Hub Natural Gas Spot Price. Figure 10 further illustrates the relationship between the SHAP values of WTI crude oil prices and various influencing factors. As shown in Fig. 10 , the WTI Crude Oil Price and the Brent Crude Oil Price show a strong positive correlation. This is mainly because both have the same commodity attributes, similar uses and qualities, and are affected by the common supply and demand fundamentals in the global crude oil market. Brent crude oil serves as the benchmark for the global crude oil market, and its price changes reflect the dynamics of global oil supply and demand. Although WTI crude oil mainly reflects the domestic market situation in the United States, as a major global oil producer and consumer, the United States is also affected by changes in international market supply and demand, and usually follows the price trend of Brent crude oil. Furthermore, when there is a significant price difference between the two, arbitrage trading will become active. The buying and selling operations of traders will make the prices of the two tend to balance, which further strengthens the positive correlation between them. The WTI crude oil Price shows a strong positive correlation with the LBMA Gold Price, mainly driven by the macro environment. When global economic expectations are optimistic or inflation heats up, the outlook for crude oil demand strengthens, while gold, as an anti-inflation asset, is favored. The demand for both rises simultaneously. Conversely, when risk events (such as geopolitical conflicts) trigger market risk aversion, oil prices may rise due to geopolitical supply risks, and the safe-haven nature of gold also pushes up its price, resulting in a same-direction fluctuation. Therefore, although the attributes of the commodities are different, the common macro logic often makes the two show a strong positive correlation. The WTI crude oil price often shows a positive correlation with the Federal Funds Effective Rate, which is mainly due to the interaction between the economic cycle and monetary policy. When the economy expands, industrial and consumer demands push up oil prices, while inflation heats up. The Federal Reserve raises interest rates to control inflation, which in turn drives up interest rates, and the two occur simultaneously. The rise in oil prices will also pass on inflation and further force interest rate hikes. Although there may be temporary differentiations in the short term due to suppressed demand and other factors, in the long run, policy adjustments driven by demand and inflation keep the two strongly positively correlated. The WTI crude oil price often shows a negative correlation with the RMB-USD Exchange Rate. The core mechanism lies in the US dollar pricing and trade flows. International crude oil is priced in US dollars. When the US dollar appreciates against the RMB, the cost of crude oil is higher for Chinese importers holding RMB. This may suppress demand and thus exert downward pressure on WTI oil prices. Conversely, the depreciation of the US dollar has reduced China's import costs, which may stimulate demand and support the rise in oil prices. Therefore, the fluctuation of exchange rates directly affects the purchasing power of the world's largest crude oil importer, forming a negative correlation between the two. The WTI crude oil Price shows a positive correlation with the Henry Hub Natural Gas Spot Price, mainly because there is a close connection between the two at the supply and demand level. From the demand side, global economic growth or recession will simultaneously affect the demand for crude oil and natural gas. From the supply side, oil and gas producers will comprehensively consider the development of both when making investment decisions. When crude oil prices rise, it may affect the production input of natural gas, and thereby influence the supply and price of natural gas. In addition, there is a certain substitution relationship between crude oil and natural gas in areas such as power generation. When the price of crude oil is high, some users may switch to natural gas, thereby pushing up its price. 4. Conclusions (1) The BiLSTM is adopted to capture long-term bidirectional dependency relationships, process time series data, and solve the problem that a single LSTM can only capture temporal dependencies in a single direction and has insufficient utilization of information in the sequence before and after. The Attention mechanism is used to dynamically calculate and allocate the weights of input features at different time steps, enabling the model to autonomously focus on the key historical information most relevant to the current prediction and suppressing the interference of irrelevant noise. The CEEMDAN-VMD is used to obtain sub-sequences with more physical significance and more thorough band separation, diminishing the non-stationarity and inherent complexity of the raw data. (2) An innovative CEEMDAN-VMD-Attention-BiLSTM hybrid model is proposed, and the prediction results are evaluated using RE, R 2 , MAPE, RMSE, and MAE. The research shows that the MAPE of the proposed hybrid prediction model is 7.66%, and R² is 0.9665. The constructed hybrid model improves the prediction by combining advanced signal decomposition and deep learning techniques. Compared with 5 mainstream models, the established hybrid model has the best prediction performance. (3) Through SHAP analysis, the interpretability of the model has been enhanced. The top 5 key factors influencing international oil prices are Brent Crude Oil Price, LBMA Gold Price, Federal Funds Effective Rate, RMB-USD Exchange Rate, and Henry Hub Natural Gas Spot Price. The proposed model helps countries capture the dynamic trends of the crude oil market and provides scientific basis for the formulation of energy policies. Declarations Conflicts of interest There are no conflicts of interest to declare. Recommended Data Availability Statements The raw data supporting the conclusions of this article will be made available by the authors on request. Funding information This research received no funding. Author Contribution Haotian Song was responsible for data collection and the writing of the paper. Hengrui Zhang was responsible for the construction and comparative analysis of the prediction model. Donghe Li was responsible for the overall technical review of the paper and the drawing of charts. Data Availability The raw data supporting the conclusions of this article will be made available by the authors on request. References Simsek, A. I. et al. A novel approach to Predict WTI crude spot oil price: LSTM-based feature extraction with Xgboost Regressor. Energy 309 , 133102 (2024). Dong, Y. et al. A novel crude oil price forecasting model using decomposition and deep learning networks. Eng. Appl. Artif. Intell. 133 , 108111 (2024). Liu, J. P. et al. A novel link prediction model for interval-valued crude oil prices based on complex network and multi-source information. Appl. Energy . 376 , 124261 (2024). Mukhaninga, M., Ravele, T. & Sigauke, C. Short-Term Forecasting of the JSE All-Share Index Using Gradient Boosting Machines. Economies 13 , 219 (2025). Xu, Y. et al. Crude oil price forecasting with multivariate selection, machine learning, and a nonlinear combination strategy. Eng. Appl. Artif. Intell. 139 , 109510 (2025). Moreno, P. et al. Forecasting Oil Prices with Non-Linear Dynamic Regression Modeling. Energies 17 (9), 2182 (2024). Ahmar, A. S., Alfairus, M. Q. & Nursya, N. Sustainable energy risk management: An integrated exponential smoothing and ARCH-GARCH framework for probabilistic forecasting. Dev. Sustain. Econ. Finance . 8 , 100087 (2025). Rao, A. et al. Crude oil Price forecasting: Leveraging machine learning for global economic stability. Technol. Forecast. Soc. Chang. 216 , 124133 (2025). Kristjanpoller, W. & Minutlol, M. C. Forecasting volatility of oil price using an artificial neural network-GARCH model. Expert Syst. Appl. 65 (15), 233–241 (2016). Yu, L., Zhang, X. & Wang, S. Assessing potentiality of support vector ma-chine method in crude oil price forecasting. EURASIA J. Mathemat-ics Sci. Technol. Educ. 13 (12), 7893–7904 (2017). Tissaoui, K. et al. Do gas price and uncertainty indices forecast crude oil prices? Fresh evidence through XGBoost modeling. Comput. Econ. 62 (2), 663–687 (2023). Liu, L. L. et al. A robust time-varying weight combined model for crude oil price forecasting. Energy 299 (15), 131352 (2024). Lin, S. C. et al. Hybrid Method for Oil Price Prediction Based on Feature Selection and XGBOOST-LSTM. Energies 18 , 2246 (2025). Aditya, P. et al. Enhanced household energy consumption forecasting using multivariate long short-term memory (LSTM) networks with weather data integration. Results Eng. 27 , 106512 (2025). Celalettin, K. & Omer, I. Temperature Prediction Using Transformer–LSTM Deep Learning Models and Sarimax from a Signal Processing Perspective. Appl. Sci. 15 (17), 9372 (2025). Jin, C. et al. Remaining Useful Life Prediction of Rolling Bearings Based on Empirical Mode Decomposition and Transformer Bi-LSTM Network. Appl. Sci. 15 (17), 9529 (2025). Zhao, J. H. et al. Fusion of KANO theory and Attention-BiLSTM models for user demand analysis and trend prediction. Inform. Fusion . 122 , 103210 (2025). Li, J. et al. Effect of microstructure on the corrosion resistance of 2205 duplex stainless steel. Part 2: Electrochemical noise analysis of corrosion behaviors of different microstructures based on wavelet transform. Constr. Build. Mater. 189 , 1294–1302 (2018). Li, Z. et al. New corrosion rate prediction method for oil and gas pipelines based on EMD and modified GM (1,N) model. Hot Working Technol. 52 (10), 35–42 (2023). Ning, F. L. et al. A framework combining acoustic features extraction method and random forest algorithm for gas pipeline leak detection and classification. Appl. Acoust. 182 , 108255 (2021). Xu, L. et al. Research and Application for Corrosion Rate Prediction of Natural Gas Pipelines Based on a Novel Hybrid Machine Learning Approach. Coatings 13 , 856 (2023). Xu, L. et al. Corrosion failure prediction in natural gas pipelines using an interpretable XGBoost model: Insights and applications. Energy 325 , 136157 (2025). Huang, N. E. et al. The empirical mode decomposition and the Hilbert spectrum for nonlinear and non-stationary time series analysis. Proceedings A. ; 454(1971): 903–995. (1998). Torres, M. E. et al. A complete ensemble empirical mode decomposition with adaptive noise. In: Acoustics, speech and signal processing (ICASSP), IEEE international conference on. IEEE; 2011. (2011). Dragomiretskiy, K. & Zosso, D. Variational mode decomposition. IEEE Trans. Signal. Process. 62 (3), 531–544 (2014). Wang, X. Y. et al. Predicting abrupt depletion of dissolved oxygen in Chaohu lake using CNN-BiLSTM with improved attention mechanism. Water Res. 261 (1), 122027 (2024). Zhang, B. D. et al. Sulfur dioxide emissions predictive model of tail gas treatment unit based on ensemble learning algorithm. Chem. Eng. Oil Gas . 54 (1), 9–17 (2025). Dang, V. T. et al. An integrated framework for bio-hydrogen production optimization using novel metaheuristic algorithms and explainable machine learning tuned via grid search. Int. J. Hydrog. Energy . 177 (13), 151555 (2025). Xu, L. et al. The research progress and prospect of data mining methods on corrosion prediction of oil and gas pipelines. Eng. Fail. Anal. 144 , 106951 (2023). Wang, S. H. et al. Multi-objective predictive modeling of natural gas desulfurization process based on deep learning[J] Vol. 54, 1–11 (Chemical Engineering of Oil & Gas, 2025). 4. Zhang, C. C. & Lin, B. Q. Assessing and interpreting carbon market efficiency based on an interpretable machine learning. Process Saf. Environ. Prot. ; 179822–179834. (2023). Wen, Z. P. et al. Explainable machine learning rapid approach to evaluate coal ash content based on X-ray fluorescence. Fuel ; 332125991. (2023). Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8242442","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":556550429,"identity":"f0d1ae76-5494-461a-ae5a-c8b4f6a8ea76","order_by":0,"name":"Haotian Song","email":"","orcid":"","institution":"Xi'an Jiaotong University","correspondingAuthor":false,"prefix":"","firstName":"Haotian","middleName":"","lastName":"Song","suffix":""},{"id":556550430,"identity":"df5b64b2-bc8e-42a2-b817-4c712e959362","order_by":1,"name":"Hengrui Zhang","email":"","orcid":"","institution":"Xi'an University of Posts \u0026 Telecommunications","correspondingAuthor":false,"prefix":"","firstName":"Hengrui","middleName":"","lastName":"Zhang","suffix":""},{"id":556550431,"identity":"c4f892b0-65f8-4326-b49d-8683cfb49b41","order_by":2,"name":"Donghe Li","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABCElEQVRIiWNgGAWjYBACPmYGAyB1gIGfmfnAgQ8MbGBRCXxa2GBaJNvbEh/OYGCTIKyFAarF4MwZY2MeqGr8WtiZNz4u+HVHjuFGgpm0bRtfncEB5oO3eRjs8nA7jK3YeGbfM2PGGQlp0jln2CQMDrAlW/MwJBfj1sJjJs3bczixWSLhmHROBUgLUISH4UBiAwEt9W0SiW3SFgYgLfzfCGvh+XE4gYfnMLMxA8QWNgJagH7hbThsOIO9jfFhzxk2yZmH2Ywt5xgk49TCz39442OeP4fl7Q/zfzjws+0YP9/x5oc33lTY4dQCBoxtcOYxBgZmEG2ATz0I/IGzaggpHQWjYBSMghEIAEnwUEU4Spo9AAAAAElFTkSuQmCC","orcid":"","institution":"Xi'an Jiaotong University","correspondingAuthor":true,"prefix":"","firstName":"Donghe","middleName":"","lastName":"Li","suffix":""}],"badges":[],"createdAt":"2025-11-30 13:53:36","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-8242442/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8242442/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":98426534,"identity":"33985f62-2bcc-4230-88d5-bc05daa4cc57","added_by":"auto","created_at":"2025-12-17 16:36:37","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1480705,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.docx","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/b4260db6e658662459e935eb.docx"},{"id":98008371,"identity":"5cf9cbd2-a82f-438a-b3ec-b6bf31a5e8e1","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":6210,"visible":true,"origin":"","legend":"","description":"","filename":"e87ac5a4267f4b1089728a86cc196d41.json","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/ef52eb620add58738e44c113.json"},{"id":98425757,"identity":"03d084ec-af85-4aee-b54c-c091fe045862","added_by":"auto","created_at":"2025-12-17 16:35:11","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":76170,"visible":true,"origin":"","legend":"","description":"","filename":"e87ac5a4267f4b1089728a86cc196d411enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/573cf90598078e289c43fe8e.xml"},{"id":98426464,"identity":"90c804c3-cd73-4d82-b39c-fc8db1a0a983","added_by":"auto","created_at":"2025-12-17 16:36:25","extension":"emf","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":136684,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.emf","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/a86b4cae601aacd33ea8b0d7.emf"},{"id":98426189,"identity":"7f2becf7-99b0-4da8-bca2-55d8a56ead2f","added_by":"auto","created_at":"2025-12-17 16:35:50","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":93549,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/35867a7c3f4543256653f231.png"},{"id":98425992,"identity":"6367d140-8ad2-4d68-b4eb-c72bc9c62c4c","added_by":"auto","created_at":"2025-12-17 16:35:28","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":288667,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/84188ac9ca187ae9e710e6d6.png"},{"id":98008380,"identity":"e47e7eba-fe28-4d03-b1f9-f63744886344","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":285781,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/8e67416f0df048d4683c3d36.png"},{"id":98425405,"identity":"80a8a285-c337-47d5-9880-c935a72e2d61","added_by":"auto","created_at":"2025-12-17 16:34:43","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":139066,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/3238c1dc0e21bddcc7561e5e.png"},{"id":98426581,"identity":"7a87db60-bb72-42ca-8acf-0fb8f24f53c5","added_by":"auto","created_at":"2025-12-17 16:36:59","extension":"jpeg","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":687651,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/5996ee0ddef88c3a862eedb5.jpeg"},{"id":98425177,"identity":"307ed4c9-4fb1-44fe-8478-11045a394cb0","added_by":"auto","created_at":"2025-12-17 16:34:29","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":85646,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/7794b1d4140f895c55021fe4.png"},{"id":98424757,"identity":"5697e04d-00b5-4ef4-a556-86362a968818","added_by":"auto","created_at":"2025-12-17 16:33:48","extension":"jpeg","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1076358,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/9fd71a42284cda4fb900899d.jpeg"},{"id":98008391,"identity":"e7ce9da0-1680-4a67-aef9-bca3c23337c5","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":37258,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/2c85d54a407fd13c83f64ac7.png"},{"id":98426455,"identity":"80a99a75-a7d3-4278-8a5a-abc2076da5ec","added_by":"auto","created_at":"2025-12-17 16:36:24","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":48323,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/7d12fce4adee74424fdf23e6.png"},{"id":98008387,"identity":"8cca914c-3302-4852-9332-8a3dad139151","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":63249,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/1187ab71c1dc9d503da97859.png"},{"id":98008388,"identity":"d0e8c717-e814-4d4b-a7e4-6684bc94d529","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":17379,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/231f12ba3d924967c792a5df.png"},{"id":98424569,"identity":"2349d9f5-229e-482a-85ce-f7178e121d03","added_by":"auto","created_at":"2025-12-17 16:33:29","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":51989,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/a891e76e927b666458b3f31b.png"},{"id":98426676,"identity":"f70c2865-dfc0-4882-841c-7f58e5ac9fa3","added_by":"auto","created_at":"2025-12-17 16:38:10","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":54054,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/02bb91c9385e8b23365942fa.png"},{"id":98425390,"identity":"b6dcce0d-184f-4ba1-be0c-75687a820c92","added_by":"auto","created_at":"2025-12-17 16:34:42","extension":"png","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":40827,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/84c2aa1bf83d2a23eac00bf5.png"},{"id":98008397,"identity":"61aef605-2410-4348-8b70-0f2f76110e81","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":135868,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/7fcea2a41ebac3ea9347995f.png"},{"id":98008395,"identity":"1b110fbc-895a-4dea-9aac-eb5e2f00cc2a","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":19,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":60798,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/d2999272c1b1542919d4d9a2.png"},{"id":98008398,"identity":"ee5050f2-3334-4df0-963c-0bfc12275808","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":20,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":259960,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/419857f137900c8cfa2fac7c.png"},{"id":98425202,"identity":"c6056982-3795-47fd-b8a2-396d9bff4f72","added_by":"auto","created_at":"2025-12-17 16:34:31","extension":"png","order_by":21,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":33036,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/d67441f40de168ccfa99ecfd.png"},{"id":98426609,"identity":"4b2f07ff-1b9f-4fc4-a2ca-9b16ccc69df5","added_by":"auto","created_at":"2025-12-17 16:37:32","extension":"png","order_by":22,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":11826,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/5dd5b8dbd6d4e5839d448c1a.png"},{"id":98008400,"identity":"cb9a45ea-e693-4b8f-9600-e31be467a7d6","added_by":"auto","created_at":"2025-12-11 17:37:50","extension":"xml","order_by":23,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":74991,"visible":true,"origin":"","legend":"","description":"","filename":"e87ac5a4267f4b1089728a86cc196d411structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/6af74c6ed8b5c48da2af9296.xml"},{"id":98008399,"identity":"5b33bfe2-d858-4a61-a0cb-14ddaca639ac","added_by":"auto","created_at":"2025-12-11 17:37:50","extension":"html","order_by":24,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":83094,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/e88b537172f5fe41e73f7331.html"},{"id":98008366,"identity":"6fc96ce7-2fc0-47a3-9ec3-a4b8643e5490","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":63249,"visible":true,"origin":"","legend":"\u003cp\u003epublication situation of papers in the field of \"oil price prediction\" from 2005 to 2024\u003c/p\u003e","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/8ba21d462e5185aa98717e0d.png"},{"id":98425870,"identity":"c6d39333-59c1-497a-9b0e-86fee56f77c8","added_by":"auto","created_at":"2025-12-17 16:35:19","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":288667,"visible":true,"origin":"","legend":"\u003cp\u003eThe subsequences obtained through the CEEMDAN decomposition\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/abfdb130bc395e7c21ce1455.png"},{"id":98008367,"identity":"fc6531b4-2358-4b29-9dd3-c4560b253583","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":285781,"visible":true,"origin":"","legend":"\u003cp\u003eThe subsequence obtained by decomposing IMF1 based on VMD\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/31b7d8a84c1263a6f6358691.png"},{"id":98008368,"identity":"ac337005-efe4-44a9-b49c-85ad1ea9e5ae","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":139066,"visible":true,"origin":"","legend":"\u003cp\u003ePredictive framework process\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/7f8e7f6f95c06c46309f8b43.png"},{"id":98425856,"identity":"2894bb7a-522d-4d41-b922-66f0d18f1e6e","added_by":"auto","created_at":"2025-12-17 16:35:18","extension":"jpeg","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":687651,"visible":true,"origin":"","legend":"\u003cp\u003eHistograms of the prediction error distribution of the six differentmodels.\u003c/p\u003e","description":"","filename":"floatimage5.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/6f18e95bed662fcbd46b981c.jpeg"},{"id":98008374,"identity":"2f5b044c-f2ce-4285-b20a-ec75afe355c2","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":85646,"visible":true,"origin":"","legend":"\u003cp\u003eThe comparison curve graph between the true value and the predicted values of different models.\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/6a80a764b62b6d3ff54340d2.png"},{"id":98008376,"identity":"18cc98d2-677d-4df9-916f-5f6b7a40b944","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"jpeg","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":1076358,"visible":true,"origin":"","legend":"\u003cp\u003ePredictive performance of six different models in terms of R²and MAPE.\u003c/p\u003e","description":"","filename":"floatimage7.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/59b41a5807d9966c64981845.jpeg"},{"id":98426517,"identity":"90a0acb3-a5fb-4fcb-8d70-26c2563a1e3d","added_by":"auto","created_at":"2025-12-17 16:36:33","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":37258,"visible":true,"origin":"","legend":"\u003cp\u003ePredictive performance of six different models in terms of RMSE and MAE.\u003c/p\u003e","description":"","filename":"floatimage8.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/72935af868e80088e0fd5e9b.png"},{"id":98425915,"identity":"52d05298-3e2a-45fd-aee8-1e191495a5ba","added_by":"auto","created_at":"2025-12-17 16:35:21","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":48323,"visible":true,"origin":"","legend":"\u003cp\u003eContribution ranking of factors affecting WTI crude oil prices\u003c/p\u003e","description":"","filename":"floatimage9.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/ebd848d1b00db3daa21aa2ae.png"},{"id":98008384,"identity":"b7916421-2f6c-4707-9b3d-e79d8149323a","added_by":"auto","created_at":"2025-12-11 17:37:49","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":93549,"visible":true,"origin":"","legend":"\u003cp\u003eThe relative impact of varying input factors on predictive modeling.\u003c/p\u003e","description":"","filename":"floatimage10.png","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/3de96900decba0ea5949c624.png"},{"id":100549060,"identity":"c5159cf1-0d80-4af0-b174-00989ac9bbca","added_by":"auto","created_at":"2026-01-19 08:22:15","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3295468,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8242442/v1/508c1545-851e-4406-afde-80d3bdfe0290.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"\u003cp\u003eA novel forecasting model for crude oil prices that integrates CEEMDAN-VMD multiscale decomposition with an Attention-based Bidirectional LSTM network\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eWTI crude oil, as one of the core benchmarks of the global crude oil pricing system, plays a significant role in various fields such as energy trade, financial markets, macroeconomic regulation, and industrial production. It not only reflects the supply and demand situation of the global energy market but is also an important investment asset in the financial market. As a commodity, the supply and demand dynamics and price changes of WTI crude oil have a significant impact on the economic security, monetary policy formulation, and geopolitical landscape of various countries around the world [\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. As one of the most actively traded commodities globally, the formation mechanism of WTI crude oil prices is complex and variable. Taking the recent market as an example, the price fluctuations are influenced by multiple factors such as geopolitical conflicts, OPEC\u0026thinsp;+\u0026thinsp;production policies, the trend of the US dollar exchange rate, and global macroeconomic expectations. The relevant data exhibit characteristics such as information redundancy, noise interference, and non-linearity, which greatly increase the difficulty of oil price prediction [\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. For oil-importing countries, every increase of \u003cspan\u003e$\u003c/span\u003e10 per barrel in WTI oil prices may mean that the country needs to pay several billion more dollars in import costs each year. By accurately predicting oil prices, market participants can make more informed investment decisions and risk management strategies, such as optimizing inventory management, formulating hedging strategies, and adjusting energy policies, etc. [\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eIn recent years, with the rapid development of artificial intelligence technology, machine learning methods have been widely applied in the field of oil price prediction due to their powerful feature extraction capabilities and outstanding predictive performance [\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]. Figure\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the trend of the number of papers on machine learning methods published in the field of oil price prediction from 2005 to 2024, which is derived from the Web of Science database. Through statistical analysis, it can be seen that the research on machine learning methods in the international oil price prediction field has shown an explosive growth trend. With the development of quantitative investment and intelligent decision-making systems, the application and development of data-driven machine learning methods in the crude oil market have become inevitable.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eGiven the intrinsic non-stationary and non-linear features of the WTI crude oil price time series, the accurate forecasting of crude oil prices and their volatility is recognized as a prominent challenge. Currently, the methods for predicting crude oil prices mainly include traditional statistical-mathematical methods and artificial intelligence methods. Commonly used statistical-mathematical methods include Autoregressive Integrated Moving Average (ARIMA) [\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e], Exponential Smoothing (ETS) [\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], Vector Autoregression (VAR) [\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], and Generalized ARCH (GARCH) [\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e]. Although these conventional econometric approaches can produce precise prediction outcomes under the assumption of approximately linear stationarity of time series, the actual crude oil price series are non-linear and non-stationary. Therefore, traditional regression methods perform poorly in predicting non-stationary and non-linear time series data. Compared with traditional methods, machine learning methods can adapt to complex nonlinear relationships and have advantages such as high computational efficiency, suitability for real-time prediction, strong adaptability and continuous learning ability. Many scholars have shifted the focus of crude oil price prediction research to artificial intelligence methods. Support Vector Regression (SVR) [\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], Extreme Gradient Boosting (XGBoost) [\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], and various other fundamental machine learning methods [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] have been developed.\u003c/p\u003e\u003cp\u003eHowever, the aforementioned machine learning methods still have problems such as local optimization, limited generalization ability, and noise interference. Compared with conventional machine learning methods, deep learning, as a promising field of artificial intelligence methods, has garnered widespread attention among researchers in recent years. It has obvious advantages in extracting nonlinear features of time series and fitting generalization ability. Among them, the Long Short-Term Memory Neural Network (LSTM), as a refined variant of the Recurrent Neural Network (RNN), is capable of effectively tackling the problems of gradient explosion or gradient disappearance that occur during the practical implementation of RNNs. Many researchers have adopted this method for predictive applications [\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e]. However, a single LSTM can only capture temporal dependencies in a single direction and is unable to effectively separate noise and interpret its intricate multi-scale features internally. Therefore, BiLSTM is adopted as the core for prediction, synchronously integrating historical and future dependency features to generate a more comprehensive bidirectional hidden state, and comprehensively capturing the long-term evolution patterns and dynamic dependency relationships of the sequence [\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Additionally, at the prediction mechanism level, considering that BiLSTM cannot distinguish the importance differences of different historical input pairs for the current prediction, the Attention mechanism is introduced. It can automatically assign higher weights to key time points (such as policy changes, geopolitical conflict periods) and core features, enabling the model to autonomously focus on the most relevant key historical information related to the current prediction, effectively suppressing the interference of irrelevant noise[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eRecent application studies have demonstrated that the data decomposition algorithm, by decomposing the original noisy data sequence into multiple sub-sequences, can effectively improve the prediction accuracy of the model. Therefore, in order to reduce the impact of data noise on the prediction results, A large number of researchers have put forward diverse data decomposition techniques. In the hybrid model, the main features of the data sequence are identified and extracted by combining some data decomposition techniques. These methods include Wavelet Transform (WT) [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e], Empirical Mode Decomposition (EMD) [\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e], and Ensemble Empirical Mode Decomposition (EEMD) [\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e]. By means of data decomposition methods, the key features of the original data sequence are extracted, which in turn improves the prediction accuracy of the model. However, the above-mentioned methods still cannot efficiently reduce the impact of data noise on the model's predictive performance. Therefore, a two-level data decomposition strategy using CEEMDAN and VMD was employed to conduct a deep decomposition of the data. Firstly, CEEMDAN was utilized to adaptively decompose the original signal into a series of intrinsic mode functions, initially eliminating modal aliasing. Subsequently, VMD was applied to the high-frequency subsequence components for a secondary optimization decomposition. This processing significantly reduced the non-stationarity and complexity of the original data, transforming a complex prediction problem into multiple relatively simple and stable subsequence prediction problems, and greatly reducing the impact of noise interference on prediction accuracy [\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. Therefore, this study employs the \"two-level decomposition and denoising\u0026thinsp;+\u0026thinsp;bidirectional dependency modeling\u0026thinsp;+\u0026thinsp;attentional precise focusing\" collaborative mechanism of the CEEMDAN-VMD-BiLSTM-Attention model to overcome the inherent limitations of the LSTM model. This provides a comprehensive technical solution for achieving higher accuracy and stronger robustness in predictions.\u003c/p\u003e\u003cp\u003eHowever, in current engineering applications, the research results of many machine learning methods often present in a \"black box\" form, with the mapping relationship between input and output lacking interpretability. Although these models can achieve accurate predictions, on account of the inadequate interpretability of the prediction results, their application in some high-risk fields is limited. Therefore, strengthening the research on the interpretability of AI technology is of great significance. In this field, interpretability research is relatively scarce. For the purpose of enhancing the model\u0026rsquo;s interpretability, the Shapley Additive Explanations (SHAP) method was applied to examine the model\u0026rsquo;s interpretability, analyzing the significance of input features for the prediction results [\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eThis study aims to construct a WTI crude oil price prediction model using the CEEMDAN-VMD-BiLSTM-Attention hybrid model. To enhance the interpretability and novelty of the research, the SHAP values are calculated to analyze the importance of input features to the prediction results, thus enhancing the model\u0026rsquo;s interpretability. The organization of the subsequent sections in this paper is outlined as follows. Section \u003cspan refid=\"Sec2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents an overview of the employed methods, and details the key aspects of data collection and model construction. Section \u003cspan refid=\"Sec6\" class=\"InternalRef\"\u003e3\u003c/span\u003e assesses the predictive performance of the proposed model. Section \u003cspan refid=\"Sec9\" class=\"InternalRef\"\u003e4\u003c/span\u003e outlines the conclusions, underscoring the key findings derived from the research.\u003c/p\u003e"},{"header":"2. Individual methods","content":"\u003cp\u003eThis section presents a concise overview of the multiple machine learning methods utilized for constructing the hybrid model, and elaborates on the key specifics during the processes of data collection, data cleaning, and model construction.\u003c/p\u003e\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1. CEEMDAN-VMD data decomposition algorithm\u003c/h2\u003e\u003cp\u003eCEEMDAN-VMD is a dual decomposition technique for non-stationary time series. Its core principle is based on the cascaded application of complementary ensemble empirical mode decomposition (CEEMDAN) and variational mode decomposition (VMD). As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e, CEEMDAN effectively suppresses the mode aliasing problem of traditional empirical mode decomposition (EMD) by introducing adaptive Gaussian white noise into the original signal and performing multiple ensemble averaging. It decomposes the signal into multiple intrinsic mode functions (IMF) with different frequency characteristics and a residual component [\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. This step can initially separate the high-frequency fluctuations, periodic terms and trend terms in the signal, diminishing the non-stationarity present in the raw data. Subsequently, VMD constructs a variational optimization model to further decompose each IMF component obtained by CEEMDAN into several modal components with limited bandwidth. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e shows the subsequence obtained by VMD decomposing IMF1. The core idea of VMD is to minimize the sum of the estimated bandwidths of each mode, adaptively determine the center frequency and bandwidth of each mode, and thereby achieve a refined division of the signal in the frequency domain. This secondary decomposition can more accurately extract the complex frequency components hidden in the IMF, such as further separating the noise in the high-frequency IMF from the effective information, or refining the periodic characteristics of the intermediate frequency components [\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eIn the process of crude oil price prediction, the dual decomposition strategy of CEEMDAN-VMD can effectively decompose the nonlinear and multi-scale fluctuations in the oil price sequence, providing purer and more feature-specific input subsequences for the subsequent Attention-BiLSTM model, thereby improving the prediction accuracy. This combined method, through multi-level signal decomposition, not only retains the robustness of CEEMDAN against noise but also leverages the advantages of VMD in frequency domain optimization, enabling it to serve as an effective tool for addressing complex time-series data.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2. Attention-BiLSTM\u003c/h2\u003e\u003cp\u003eAfter the signal decomposition is completed by CEEMDAN-VMD, Attention-BiLSTM further refines the modeling of multi-scale subsequences. Attention-BiLSTM is a deep learning architecture that integrates bidirectional temporal modeling and dynamic feature selection capabilities. Among them, BiLSTM processes time series data in parallel through two LSTM layers, forward and backward, capable of simultaneously capturing the bidirectional dependencies of past and future in the oil price sequence. This bidirectional information integration can more comprehensively depict the temporal dynamics of oil price fluctuations, especially suitable for complex scenarios influenced by long-cycle factors such as geopolitics and economic policies. Attention, on the other hand, calculates the \"importance score\" at each time step to dynamically adjust the model's focus on information at different moments [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e].\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec5\" class=\"Section2\"\u003e\u003ch2\u003e2.3. Construction of residual strength prediction model\u003c/h2\u003e\u003cp\u003eThis study proposes a hybrid CEEMDAN-VMD-Attention-BiLSTM model designed to predict WTI crude oil prices. The prediction framework is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e. This framework includes data collection, preprocessing, model construction, prediction accuracy evaluation and model interpretability analysis. In terms of data collection, time series data from 1996 to 2024 were collected on a daily basis. Eight relevant factors were selected as input parameters, including Brent Crude Oil Price, CBOT Volatility Index, Economic Policy Uncertainty Index, Federal Funds Effective Rate, Henry Hub Natural Gas Spot Price, LBMA Gold Price, NASDAQ-100 Index, and RMB-USD Exchange Rate data. Take the price of WIT Crude oil as the output parameter of the model. In terms of data processing, the non-stationary and non-linear characteristics of the WTI oil price sequence were decomposed into several stationary sub-sequences using CEEMDAN-VMD, thereby reducing data noise and capturing the nonlinear relationship between crude oil prices and relevant variables. Since different variables have different magnitudes and measurement units, Min-Max normalization was adopted to normalize the data to the range of [-1, 1], in order to improve the training efficiency of the model [\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. In terms of model construction, due to the advantages of grid search (GS) in finding optimal solutions, stable and reliable results, and easy to understand search results, the GS method is used to optimize the hyperparameters of the model [\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e]. Using Attention BiLSTM to construct a bidirectional temporal and attention weighted prediction model, sum up the predicted values of the oil price subsequence obtained by CEEMDAN-VMD decomposition, and reconstruct the final predicted value. In terms of model performance evaluation, five mainstream models including RE, R\u003csup\u003e2\u003c/sup\u003e, MAPE, RMSE, and MAE are employed to evaluate the model's forecasting accuracy [\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Finally, SHAP is used to study the interpretability of the model, quantify the impact of input variables on the output results of the model, and improve the interpretability of the model.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"3. Case study","content":"\u003cdiv id=\"Sec7\" class=\"Section2\"\u003e\u003ch2\u003e3.1. Evaluation of predictive performance\u003c/h2\u003e\u003cp\u003eTo evaluate the predictive performance of the proposed hybrid model, the CEEMDAN-VMD-Attention-BiLSTM hybrid model was compared with 5 other models. Figure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e shows the histogram of the prediction error distribution for the six models. The analysis indicates that the mean and standard deviation of the prediction errors of the LSTM model are the highest, followed by BiLSTM, Attention-BiLSTM, VMD-Attention-BiLSTM, and CEEMDAN-Attention-BiLSTM. Compared with the other five models, the CEEMDAN-VMD-Attention-BiLSTM hybrid model has the smallest mean and standard deviation of the prediction errors and exhibits better predictive performance.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTo further assess the predictive capability of the hybrid model and examine the fitting effects of different models, the research results are shown in Figs.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e. Through Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e, it can be observed that the predicted values of the CEEMDAN-VMD-Attention-BiLSTM model are the closest to the actual data, but the fitting effect between the predicted results and the actual data cannot be directly quantified. Therefore, a further evaluation analysis of the predictive performance was conducted by combining R\u003csup\u003e2\u003c/sup\u003e and MAPE, and the results are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig7\" class=\"InternalRef\"\u003e7\u003c/span\u003e. The R\u003csup\u003e2\u003c/sup\u003e of the LSTM model is 0.5383, and the MAPE is 14.5%. The R\u003csup\u003e2\u003c/sup\u003e of the BiLSTM model is 0.6441, and the MAPE is 13.76%. The R\u003csup\u003e2\u003c/sup\u003e of the Attention-BiLSTM model is 0.7365, and the MAPE is 11.8%. The R\u003csup\u003e2\u003c/sup\u003e of the VMD-Attention-BiLSTM model is 0.9457, and the MAPE is 9.93%. The R\u003csup\u003e2\u003c/sup\u003e of the CEEMDAN-Attention-BiLSTM model is 0.96, and the MAPE is 9.27%. The R\u003csup\u003e2\u003c/sup\u003e of the CEEMDAN-VMD-Attention-BiLSTM model is 0.9665, and the MAPE is 7.66%. The study found that the prediction effect of LSTM is the worst, while the prediction effects of BiLSTM and Attention-BiLSTM gradually improve. The prediction performance of VMD-Attention-BiLSTM, CEEMDAN-Attention-BiLSTM, and CEEMDAN-VMD-Attention-BiLSTM has been further improved. The research proves that the selected modules in the constructed hybrid model CEEMDAN-VMD-Attention-BiLSTM are all reasonable, and each module is very important for the purpose of improving the predictive performance of the model.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTaking the RMSE evaluation index as an example, comparing the prediction results of LSTM and BiLSTM, it is found that the RMSE of BiLSTM is 11.7161%, and its prediction performance is better than that of LSTM, proving that BiLSTM can learn long-term bidirectional dependency relationships and effectively handle time series data, improving the prediction accuracy of the model. Comparing the prediction results of BiLSTM and Attention-BiLSTM, the latter has an RMSE of 11.1824, and its prediction performance is better than that of BiLSTM, proving that the attention mechanism can dynamically and selectively utilize historical information, enhancing the influence of important information on the output of BiLSTM, and improving the prediction accuracy. Comparing the prediction results of Attention-BiLSTM, VMD-Attention-BiLSTM, CEEMDAN-Attention-BiLSTM and CEEMDAN-VMD-Attention-BiLSTM, the prediction effect of BiLSTM is the worst, while the RMSE of CEEMDAN-VMD-Attention-BiLSTM model is 5.103, which has the best prediction performance. This further proves that the data decomposition algorithm is crucial in improving the prediction performance of time series. The CEEMDAN and VMD signal decomposition stages can effectively handle the non-linearity and non-stationarity of the crude oil price sequence, diminish the impact of noise interference, and the CEEMDAN-VMD double-layer data decomposition algorithm brings the best effect. For the MAE evaluation index, similar conclusions can also be found. Overall, the CEEMDAN-VMD-Attention-BiLSTM model advances predictive capability by integrating cutting-edge signal decomposition methods with deep learning frameworks. Compared with existing models, the established hybrid model has the best prediction performance.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section2\"\u003e\u003ch2\u003e3.2. SHAP analysis\u003c/h2\u003e\u003cp\u003eAt present, with the increasing complexity of machine learning models, the problem of the lack of interpretability caused by the \"black box\" feature of models has become more and more prominent. Conducting SHAP (SHapley Additive exPlanations) analysis is precisely the key link to solve this predicament and ensure the reliable application of models. SHAP is based on the Shapley principle in game theory and can assign fair and consistent importance weights to each feature[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. It not only quantifies the contribution of individual features to the model's prediction results but also avoids the limitations of traditional feature importance analysis methods that only reflect global contributions and cannot explain local prediction logic. Enable developers and users to clearly understand \"why the model makes such predictions\", thereby building trust in the model and reducing the risk of decision-making errors caused by the model's \"black box\"[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e].\u003c/p\u003e\u003cp\u003eThe absolute value of SHAP is used to quantify the influence of various input parameters on the price of WTI crude oil. The greater the absolute value of SHAP, the greater the impact of the parameter on the WTI crude oil price. By means of correlation analysis, there are 5 factors with a correlation coefficient greater than 0.3. Therefore, SHAP analysis is carried out based on the above five types of parameters. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig9\" class=\"InternalRef\"\u003e9\u003c/span\u003e, The top five factors influencing the WTI Crude Oil Price, in descending order of their influence, are Brent Crude Oil Price, LBMA Gold Price, Federal Funds Effective Rate, RMB-USD Exchange Rate, and Henry Hub Natural Gas Spot Price.\u003c/p\u003e\u003cp\u003eFigure \u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e further illustrates the relationship between the SHAP values of WTI crude oil prices and various influencing factors. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig10\" class=\"InternalRef\"\u003e10\u003c/span\u003e, the WTI Crude Oil Price and the Brent Crude Oil Price show a strong positive correlation. This is mainly because both have the same commodity attributes, similar uses and qualities, and are affected by the common supply and demand fundamentals in the global crude oil market. Brent crude oil serves as the benchmark for the global crude oil market, and its price changes reflect the dynamics of global oil supply and demand. Although WTI crude oil mainly reflects the domestic market situation in the United States, as a major global oil producer and consumer, the United States is also affected by changes in international market supply and demand, and usually follows the price trend of Brent crude oil. Furthermore, when there is a significant price difference between the two, arbitrage trading will become active. The buying and selling operations of traders will make the prices of the two tend to balance, which further strengthens the positive correlation between them. The WTI crude oil Price shows a strong positive correlation with the LBMA Gold Price, mainly driven by the macro environment. When global economic expectations are optimistic or inflation heats up, the outlook for crude oil demand strengthens, while gold, as an anti-inflation asset, is favored. The demand for both rises simultaneously. Conversely, when risk events (such as geopolitical conflicts) trigger market risk aversion, oil prices may rise due to geopolitical supply risks, and the safe-haven nature of gold also pushes up its price, resulting in a same-direction fluctuation. Therefore, although the attributes of the commodities are different, the common macro logic often makes the two show a strong positive correlation. The WTI crude oil price often shows a positive correlation with the Federal Funds Effective Rate, which is mainly due to the interaction between the economic cycle and monetary policy. When the economy expands, industrial and consumer demands push up oil prices, while inflation heats up. The Federal Reserve raises interest rates to control inflation, which in turn drives up interest rates, and the two occur simultaneously. The rise in oil prices will also pass on inflation and further force interest rate hikes. Although there may be temporary differentiations in the short term due to suppressed demand and other factors, in the long run, policy adjustments driven by demand and inflation keep the two strongly positively correlated. The WTI crude oil price often shows a negative correlation with the RMB-USD Exchange Rate. The core mechanism lies in the US dollar pricing and trade flows. International crude oil is priced in US dollars. When the US dollar appreciates against the RMB, the cost of crude oil is higher for Chinese importers holding RMB. This may suppress demand and thus exert downward pressure on WTI oil prices. Conversely, the depreciation of the US dollar has reduced China's import costs, which may stimulate demand and support the rise in oil prices. Therefore, the fluctuation of exchange rates directly affects the purchasing power of the world's largest crude oil importer, forming a negative correlation between the two. The WTI crude oil Price shows a positive correlation with the Henry Hub Natural Gas Spot Price, mainly because there is a close connection between the two at the supply and demand level. From the demand side, global economic growth or recession will simultaneously affect the demand for crude oil and natural gas. From the supply side, oil and gas producers will comprehensively consider the development of both when making investment decisions. When crude oil prices rise, it may affect the production input of natural gas, and thereby influence the supply and price of natural gas. In addition, there is a certain substitution relationship between crude oil and natural gas in areas such as power generation. When the price of crude oil is high, some users may switch to natural gas, thereby pushing up its price.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Conclusions","content":"\u003cp\u003e(1) The BiLSTM is adopted to capture long-term bidirectional dependency relationships, process time series data, and solve the problem that a single LSTM can only capture temporal dependencies in a single direction and has insufficient utilization of information in the sequence before and after. The Attention mechanism is used to dynamically calculate and allocate the weights of input features at different time steps, enabling the model to autonomously focus on the key historical information most relevant to the current prediction and suppressing the interference of irrelevant noise. The CEEMDAN-VMD is used to obtain sub-sequences with more physical significance and more thorough band separation, diminishing the non-stationarity and inherent complexity of the raw data.\u003c/p\u003e\u003cp\u003e(2) An innovative CEEMDAN-VMD-Attention-BiLSTM hybrid model is proposed, and the prediction results are evaluated using RE, R\u003csup\u003e2\u003c/sup\u003e, MAPE, RMSE, and MAE. The research shows that the MAPE of the proposed hybrid prediction model is 7.66%, and R\u0026sup2; is 0.9665. The constructed hybrid model improves the prediction by combining advanced signal decomposition and deep learning techniques. Compared with 5 mainstream models, the established hybrid model has the best prediction performance.\u003c/p\u003e\u003cp\u003e(3) Through SHAP analysis, the interpretability of the model has been enhanced. The top 5 key factors influencing international oil prices are Brent Crude Oil Price, LBMA Gold Price, Federal Funds Effective Rate, RMB-USD Exchange Rate, and Henry Hub Natural Gas Spot Price. The proposed model helps countries capture the dynamic trends of the crude oil market and provides scientific basis for the formulation of energy policies.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003ch2\u003eConflicts of interest\u003c/h2\u003e\u003cp\u003eThere are no conflicts of interest to declare.\u003c/p\u003e\u003c/p\u003e\u003cp\u003e\u003ch2\u003eRecommended Data Availability Statements\u003c/h2\u003e\u003cp\u003eThe raw data supporting the conclusions of this article will be made available by the authors on request.\u003c/p\u003e\u003c/p\u003e\u003ch2\u003eFunding information\u003c/h2\u003e\u003cp\u003eThis research received no funding.\u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eHaotian Song was responsible for data collection and the writing of the paper. Hengrui Zhang was responsible for the construction and comparative analysis of the prediction model. Donghe Li was responsible for the overall technical review of the paper and the drawing of charts.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eThe raw data supporting the conclusions of this article will be made available by the authors on request.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eSimsek, A. I. et al. A novel approach to Predict WTI crude spot oil price: LSTM-based feature extraction with Xgboost Regressor. \u003cem\u003eEnergy\u003c/em\u003e \u003cb\u003e309\u003c/b\u003e, 133102 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDong, Y. et al. A novel crude oil price forecasting model using decomposition and deep learning networks. \u003cem\u003eEng. Appl. Artif. Intell.\u003c/em\u003e \u003cb\u003e133\u003c/b\u003e, 108111 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu, J. P. et al. A novel link prediction model for interval-valued crude oil prices based on complex network and multi-source information. \u003cem\u003eAppl. Energy\u003c/em\u003e. \u003cb\u003e376\u003c/b\u003e, 124261 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMukhaninga, M., Ravele, T. \u0026amp; Sigauke, C. Short-Term Forecasting of the JSE All-Share Index Using Gradient Boosting Machines. \u003cem\u003eEconomies\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e, 219 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu, Y. et al. Crude oil price forecasting with multivariate selection, machine learning, and a nonlinear combination strategy. \u003cem\u003eEng. Appl. Artif. Intell.\u003c/em\u003e \u003cb\u003e139\u003c/b\u003e, 109510 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eMoreno, P. et al. Forecasting Oil Prices with Non-Linear Dynamic Regression Modeling. \u003cem\u003eEnergies\u003c/em\u003e \u003cb\u003e17\u003c/b\u003e (9), 2182 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAhmar, A. S., Alfairus, M. Q. \u0026amp; Nursya, N. Sustainable energy risk management: An integrated exponential smoothing and ARCH-GARCH framework for probabilistic forecasting. \u003cem\u003eDev. Sustain. Econ. Finance\u003c/em\u003e. \u003cb\u003e8\u003c/b\u003e, 100087 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eRao, A. et al. Crude oil Price forecasting: Leveraging machine learning for global economic stability. \u003cem\u003eTechnol. Forecast. Soc. Chang.\u003c/em\u003e \u003cb\u003e216\u003c/b\u003e, 124133 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eKristjanpoller, W. \u0026amp; Minutlol, M. C. Forecasting volatility of oil price using an artificial neural network-GARCH model. \u003cem\u003eExpert Syst. Appl.\u003c/em\u003e \u003cb\u003e65\u003c/b\u003e (15), 233\u0026ndash;241 (2016).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eYu, L., Zhang, X. \u0026amp; Wang, S. Assessing potentiality of support vector ma-chine method in crude oil price forecasting. \u003cem\u003eEURASIA J. Mathemat-ics Sci. Technol. Educ.\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e (12), 7893\u0026ndash;7904 (2017).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTissaoui, K. et al. Do gas price and uncertainty indices forecast crude oil prices? Fresh evidence through XGBoost modeling. \u003cem\u003eComput. Econ.\u003c/em\u003e \u003cb\u003e62\u003c/b\u003e (2), 663\u0026ndash;687 (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLiu, L. L. et al. A robust time-varying weight combined model for crude oil price forecasting. \u003cem\u003eEnergy\u003c/em\u003e \u003cb\u003e299\u003c/b\u003e (15), 131352 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLin, S. C. et al. Hybrid Method for Oil Price Prediction Based on Feature Selection and XGBOOST-LSTM. \u003cem\u003eEnergies\u003c/em\u003e \u003cb\u003e18\u003c/b\u003e, 2246 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eAditya, P. et al. Enhanced household energy consumption forecasting using multivariate long short-term memory (LSTM) networks with weather data integration. \u003cem\u003eResults Eng.\u003c/em\u003e \u003cb\u003e27\u003c/b\u003e, 106512 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eCelalettin, K. \u0026amp; Omer, I. Temperature Prediction Using Transformer\u0026ndash;LSTM Deep Learning Models and Sarimax from a Signal Processing Perspective. \u003cem\u003eAppl. Sci.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e (17), 9372 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eJin, C. et al. Remaining Useful Life Prediction of Rolling Bearings Based on Empirical Mode Decomposition and Transformer Bi-LSTM Network. \u003cem\u003eAppl. Sci.\u003c/em\u003e \u003cb\u003e15\u003c/b\u003e (17), 9529 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhao, J. H. et al. Fusion of KANO theory and Attention-BiLSTM models for user demand analysis and trend prediction. \u003cem\u003eInform. Fusion\u003c/em\u003e. \u003cb\u003e122\u003c/b\u003e, 103210 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi, J. et al. Effect of microstructure on the corrosion resistance of 2205 duplex stainless steel. Part 2: Electrochemical noise analysis of corrosion behaviors of different microstructures based on wavelet transform. \u003cem\u003eConstr. Build. Mater.\u003c/em\u003e \u003cb\u003e189\u003c/b\u003e, 1294\u0026ndash;1302 (2018).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eLi, Z. et al. New corrosion rate prediction method for oil and gas pipelines based on EMD and modified GM (1,N) model. \u003cem\u003eHot Working Technol.\u003c/em\u003e \u003cb\u003e52\u003c/b\u003e (10), 35\u0026ndash;42 (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eNing, F. L. et al. A framework combining acoustic features extraction method and random forest algorithm for gas pipeline leak detection and classification. \u003cem\u003eAppl. Acoust.\u003c/em\u003e \u003cb\u003e182\u003c/b\u003e, 108255 (2021).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu, L. et al. Research and Application for Corrosion Rate Prediction of Natural Gas Pipelines Based on a Novel Hybrid Machine Learning Approach. \u003cem\u003eCoatings\u003c/em\u003e \u003cb\u003e13\u003c/b\u003e, 856 (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu, L. et al. Corrosion failure prediction in natural gas pipelines using an interpretable XGBoost model: Insights and applications. \u003cem\u003eEnergy\u003c/em\u003e \u003cb\u003e325\u003c/b\u003e, 136157 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eHuang, N. E. et al. The empirical mode decomposition and the Hilbert spectrum for nonlinear and non-stationary time series analysis. Proceedings A. ; 454(1971): 903\u0026ndash;995. (1998).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eTorres, M. E. et al. A complete ensemble empirical mode decomposition with adaptive noise. In: Acoustics, speech and signal processing (ICASSP), IEEE international conference on. IEEE; 2011. (2011).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDragomiretskiy, K. \u0026amp; Zosso, D. Variational mode decomposition. \u003cem\u003eIEEE Trans. Signal. Process.\u003c/em\u003e \u003cb\u003e62\u003c/b\u003e (3), 531\u0026ndash;544 (2014).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWang, X. Y. et al. Predicting abrupt depletion of dissolved oxygen in Chaohu lake using CNN-BiLSTM with improved attention mechanism. \u003cem\u003eWater Res.\u003c/em\u003e \u003cb\u003e261\u003c/b\u003e (1), 122027 (2024).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang, B. D. et al. Sulfur dioxide emissions predictive model of tail gas treatment unit based on ensemble learning algorithm. \u003cem\u003eChem. Eng. Oil Gas\u003c/em\u003e. \u003cb\u003e54\u003c/b\u003e (1), 9\u0026ndash;17 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eDang, V. T. et al. An integrated framework for bio-hydrogen production optimization using novel metaheuristic algorithms and explainable machine learning tuned via grid search. \u003cem\u003eInt. J. Hydrog. Energy\u003c/em\u003e. \u003cb\u003e177\u003c/b\u003e (13), 151555 (2025).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eXu, L. et al. The research progress and prospect of data mining methods on corrosion prediction of oil and gas pipelines. \u003cem\u003eEng. Fail. Anal.\u003c/em\u003e \u003cb\u003e144\u003c/b\u003e, 106951 (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWang, S. H. et al. \u003cem\u003eMulti-objective predictive modeling of natural gas desulfurization process based on deep learning[J]\u003c/em\u003e Vol. 54, 1\u0026ndash;11 (Chemical Engineering of Oil \u0026amp; Gas, 2025). 4.\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eZhang, C. C. \u0026amp; Lin, B. Q. Assessing and interpreting carbon market efficiency based on an interpretable machine learning. \u003cem\u003eProcess Saf. Environ. Prot.\u003c/em\u003e ; 179822\u0026ndash;179834. (2023).\u003c/span\u003e\u003c/li\u003e\u003cli\u003e\u003cspan\u003eWen, Z. P. et al. Explainable machine learning rapid approach to evaluate coal ash content based on X-ray fluorescence. \u003cem\u003eFuel\u003c/em\u003e ; 332125991. (2023).\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"WTI crude oil price forecast, Deep learning, CEEMDAN-VMD, Attention-BiLSTM, Interpretability","lastPublishedDoi":"10.21203/rs.3.rs-8242442/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8242442/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccurate prediction of WTI crude oil prices is of great significance for the business decision-making of oil and gas enterprises, the formulation of national energy strategies, and the risk management of the global financial market. However, current traditional prediction methods have some limitations in predicting WTI crude oil prices. For instance, traditional methods are insufficient in integrating complex factors that affect international oil prices, such as geopolitical conflicts, global supply and demand imbalances, and financial speculation. They also have limited ability to extract nonlinear and time-varying correlation features between oil prices and multiple influencing factors, and do not adequately consider data noise. As a result, it is difficult to clearly explain the contribution mechanism of key influencing factors to the prediction results of oil prices, and the model's interpretability is insufficient. To resolve these challenges, this paper integrates Complete Ensemble Empirical Mode Decomposition with Adaptive Noise (CEEMDAN), Variational Mode Decomposition (VMD), Attention Mechanism (Attention), and Bidirectional Long Short-Term Memory (BiLSTM), and proposes a deep learning-based hybrid prediction model (CEEMDAN-VMD-Attention-BiLSTM). Specifically, the non-stationary and non-linear characteristics of the WTI oil price series are decomposed into several stationary sub-series by using CEEMDAN-VMD, reducing data noise and capturing the non-linear relationship between crude oil prices and macroeconomic variables. An Attention-BiLSTM model is constructed to predict the sub-series of oil prices decomposed by CEEMDAN-VMD, and the predicted values of these sub-series are summed to reconstruct the final predicted value. In order to augment the interpretability of the model's forecast results, the SHAP method is adopted to quantify the contribution of different input parameters to the model's prediction results. Based on 28 years of time series data, the study shows that the MAPE of the proposed hybrid prediction model is 7.66%, and the R\u0026sup2; is 0.9665. The proposed model demonstrates superior predictive accuracy and notably robust performance in comparative analysis. Through SHAP analysis, the top 5 key factors influencing international oil prices are Brent Crude Oil Price, LBMA Gold Price, Federal Funds Effective Rate, RMB-USD Exchange Rate, and Henry Hub Natural Gas Spot Price. The proposed model helps countries grasp the trend of the crude oil market and provides scientific basis for the formulation of energy policies.\u003c/p\u003e","manuscriptTitle":"A novel forecasting model for crude oil prices that integrates CEEMDAN-VMD multiscale decomposition with an Attention-based Bidirectional LSTM network","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-11 17:37:40","doi":"10.21203/rs.3.rs-8242442/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"83cfc5e9-b494-4b8b-a8e5-baa719876a1d","owner":[],"postedDate":"December 11th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":59509220,"name":"Physical sciences/Engineering"},{"id":59509221,"name":"Physical sciences/Mathematics and computing"}],"tags":[],"updatedAt":"2026-01-19T05:39:20+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-11 17:37:40","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8242442","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8242442","identity":"rs-8242442","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.