Forecasting Crude Oil Prices: Insights from Machine Learning Approaches

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract This study investigates the efficacy of machine learning (ML) models in forecasting crude oil prices, a critical factor influencing economic stability. Given the inherent volatility and complexity of oil markets, accurate prediction is essential for mitigating risks and informing strategic decisions. Employing an autoregressive framework, the research utilizes daily oil price data from March 2013 to February 2024 to train and evaluate a diverse set of ML algorithms, including linear regression, Random Forest, SVR, XGBoost, Gradient Boosting Machine (GBM), LSTM, and CNN. The analysis reveals that linear regression demonstrates exceptional performance, effectively capturing linear trends, while tree-based models like XGBoost and GBM also exhibit high predictive accuracy. LSTM proves adept at capturing long-term dependencies, though CNN shows superior agility in short-term forecasting. Overall, linear regression in machine learning, and LSTM from deep learning are best models in forecasting. The study underscores the importance of rigorous hyperparameter tuning and model selection based on specific forecasting objectives. However, limitations stemming from reliance on historical data and the univariate approach are acknowledged. Future research should explore incorporating external factors and hybrid model architectures to enhance forecasting robustness.
Full text 201,959 characters · extracted from preprint-html · click to expand
Forecasting Crude Oil Prices: Insights from Machine Learning Approaches | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Forecasting Crude Oil Prices: Insights from Machine Learning Approaches Haseen This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-8354755/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract This study investigates the efficacy of machine learning (ML) models in forecasting crude oil prices, a critical factor influencing economic stability. Given the inherent volatility and complexity of oil markets, accurate prediction is essential for mitigating risks and informing strategic decisions. Employing an autoregressive framework, the research utilizes daily oil price data from March 2013 to February 2024 to train and evaluate a diverse set of ML algorithms, including linear regression, Random Forest, SVR, XGBoost, Gradient Boosting Machine (GBM), LSTM, and CNN. The analysis reveals that linear regression demonstrates exceptional performance, effectively capturing linear trends, while tree-based models like XGBoost and GBM also exhibit high predictive accuracy. LSTM proves adept at capturing long-term dependencies, though CNN shows superior agility in short-term forecasting. Overall, linear regression in machine learning, and LSTM from deep learning are best models in forecasting. The study underscores the importance of rigorous hyperparameter tuning and model selection based on specific forecasting objectives. However, limitations stemming from reliance on historical data and the univariate approach are acknowledged. Future research should explore incorporating external factors and hybrid model architectures to enhance forecasting robustness. Finance s – Q43 C63 C53 C52 C45 Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 1. Introduction The role of oil prices in influencing macroeconomic variables is well-established in economic research (Barsky & Kilian 2002 ; Baumeister & Kilian 2016 ; Hamilton & Herrera 2004 ; Ahmed et al. 2025 ; Siddiqui et al. 2023 ). Crude oil prices hold a crucial position in oil-dependent economies, particularly those where exports and imports constitute a significant portion of the GDP (Gross Domestic Product) (Mohsin & Jamaani 2023 ). In such economies, variations in crude oil prices have critical financial effetcs. For instance, rising oil prices can lead to inflation and reduce an economy's competitiveness (Jahan-Parvar & Mohammadi, 2009 ). On contrary, declined oil prices can result in social unrest and political instability in oil-producing countries (Chen & Hsu 2012 ; Gholamian et al. 2005 ). Therefore, understanding and managing the impact of oil price volatility is essential for maintaining economic stability and growth in these regions. Since the 1970s, oil prices have become increasingly unstable, impacting both economic growth and investor confidence (Xu et al. 2023 ). This instability is not only due to the inherent non-linearity of oil prices but also because of extreme and irregular events. While oil prices are fundamentally governed by the principles of demand and supply, external shocks such as geopolitical events and pandemic of COVID-19 have significant influences as well (Bernabe et al. 2004 ; Hagen 1994 ; Stevens 1995 ; Tang et al. 2015 , Atif et al., 2022 ; Ahmed and Kaur, 2025 ). Additionally, oil is a finite natural resource and a major contributor to environmental pollution and global warming. Worldwide efforts to shift toward sustainable energy sources, spurred by international pledges to cut carbon emissions and address climate change through mitigation and adaptation strategies, also plays a crucial role in shaping crude oil prices. This shift not only affects the demand for oil but also creates uncertainties in the market, as the world navigates towards a more sustainable energy future. As highlighted by Fang et al. ( 2023 ), forecasting oil prices is crucial for both producers and consumers. Accurate predictions can significantly contribute to rapid economic growth, higher production levels, and reduced production costs, fostering a stable macroeconomic environment (Yu et al. 2017 ). However, because of unpredictable nature of oil prices, accurate forecasting remains challenging. Brent crude prices are affected by a complex interplay of endogenous and exogenous factors. Endogenous factors include the inventory levels of oil, consumption rates, and supply dynamics. On the other hand, exogenous factors encompass geopolitical events, conflicts, and broader economic developments, all of which can drastically impact oil prices. This high volatility generates uncertainty and fear within both the public and private sectors. Additionally, the decisions made by OPEC have a substantial effect on determining the prices of brent crude in the global market (Kaufmann et al. 2004 ). Over the past decade, energy-based commodities have experienced exponential growth, evolving into an indispensable asset class for investment purposes. This growth underscores the importance of understanding and forecasting oil prices to mitigate risks and capitalize on investment opportunities. In forecasting, researchers have historically relied on the ARIMA model. More recently, studies have begun to employ Artificial Neural Networks (ANN) for this purpose. Traditional methods rely on mostly on one type of model, time series models or regression models. The complexity of the oil prices is not captured in the findings. Further, most of the studies have applied the single machine learning models to forecast the oil prices. Although research utilizing ML techniques such as ANN and SVM., are able to forecast the oil prices, under the complexity and uncertainty, they lack in the analysis of the comparison among various ML models. It is in this context this study forecast the oil prices based on the autoregressive approach, using Linear Regression, SVM, Random Forest, Extreme Gradient Boosting (XGBoost), and GBM based ML models. This research aims to accurately forecast Brent crude oil prices by evaluating the predictive performance of various machine learning models. To ensure a thorough assessment, the study utilizes six key metrics: MSE, RMSE, MAE, MAPE, R-Squared, and the Diebold-Mariano (DM) statistic. Taking into account the necessity and complexity of the oil prices forecasting, this focus upon the evaluating best machine learning model in terms of prediction accuracy, in forecasting of the oil prices. Further, in this paper section 2 provides a brief review of past studies, followed by the Methodology. The section 4 discusses the results and findings followed by the conclusion and implications. 2. Literature Review The literature on the machine learning based forecast has grown many folds in the recent. Application of ML in forecasting the time series are applied in predicting the time series. Many researchers have employed individual ML techniques to forecast oil price trends. Ahmed et al. ( 2022 ) provides comprehensive studies exploring the use of AI and ML in the field of finance. Study found the application of AI and ML in financial sectors, including forecasting bankruptcy, predicting stock values, managing investment portfolios, and detecting money laundering activities., to be higher post 2015, and continues to rise. However, the AI and ML based research is localized in United States, China, and the United Kingdom. Among the notable earliest studies predicting the oil prices are (Gabralla & Abraham 2015 ; Khashman & Nwulu 2011 ; Panella et al. 2012 ; Shao et al. 2014 ; Tebyanian & Hedayati 2014 ,Ahmed, 2025 ). Among the recent studies, Naeem et al. ( 2024 ) applied the ARIMA, SVM and LSTM to check the prediction accuracy of the oil prices. Diebold – Mariano is used to compare the robustness and forecasting methods. The empirical findings suggest the higher accuracy of the hybrid models over the alternative approaches. Further, Liang et al. ( 2023 ) developed reinforcement learning based algorithm for forecasting the oil prices on three major commodity exchanges. Authors designed the mechanism based on stochastic process for generalization and accuracy, along with the learning efficiency. The study found the algorithm to be better predictor in accuracy, of the three major crude oil price benchmarks. Further, the algorithm can be extended to predict the other variables related to the natural resources. Numerous studies applied the machine learning models, along with the traditional methods such as ARIMA, GARCH etc. ( Cheng et al. 2022 ; Nanthiya et al. 2023 ; Tami & Owda 2024 ; Wei et al. 2023 ; Weng et al. 2020 ; Yu et al. 2017 ). Aldabagh et al. ( 2023 ) proposed the model based on the deep learning and traditional ARIMA, to forecast the oil prices. The author used the CNN with LSTM. Further, the CNN-LSTM model is compared with LSTM, CNN, SVM, and the ARIMA model. The suggested model demonstrates superior accuracy compared to the current model for both single-step and multiple-step forecasts. Das and Das ( 2024 ) employed a forecasting model that incorporated text-based sentiment related to inflation as additional inputs into an artificial neural network, which outperformed all other models in terms of forecast accuracy. Jha et al. ( 2024 ) applied the SVM to estimate the spot prices by utilizing the multiple variables. The proposed model is evaluated against the four benchmark models (Ordinary Least Squared, GARCH, and ANN), at different steps. Author found the LASSO model to be better predictor. Zhang and Zheng (2023) explored the oil price prediction by leveraging traditional models of time series, and the machine learning models. The Ordinary Least Square is used for the variable selection, and machine leaning models for the best overall performance. LSTM model is found to be the best in overall performance. Cheng et al. ( 2022 ) built framework to estimate the variability of brent crude prices by taking into account the structural breaks (Extreme Events). The oil prices are decomposed to find the extreme events or structural break. Further, the ARIMA and SVM are combined to predict the oil prices. The empirical findings found that combined models perform better than other models. Tissaoui et al. ( 2023 ) applied the XGBoost ML technique to estimate the oil prices, and compared the SVM, and the ARIMAX models to check the relationship with the forecasters. The empirical findings of the study indicate the dominance of ML models in forecasting the oil prices. Kakade et al. ( 2023 ) proposed model to estimate the crude oil prices based on the hybrid ensemble learning approach. It includes the ARIMA, PCA, and the LSTM model in ensemble approach, and is compared with the LSTM. Findings indicate that oil ensemble LSTM is more accurate than the traditional LSTM. Liu et al. ( 2024 ) forecasted the crude oil prices by employing the ARIMA, LSTM. Neural Networks based on extreme learning, and back propagation neural network. The proposed model’s accuracy is satisfactory, even during the COVID-19 pandemic. The oil prices are influenced by various factors, particularly in an oil dependent economy. In this regard, Albahooth ( 2024 ) predicted the oil prices in the Saudi Arabia, by employing the interest rates, currency rates, and stock markets from financial variables, and energy consumption, inflation rate, and GDP growth from macroeconomic variables. The study used various ML technique such as linear regression and SVM, and models are evaluated by examining the RMSE, MSE, and R-square. The study explores the significance of various financial and macro-economic variables. Similarly, Guliyev & Mustafayev ( 2022 ) predicted the West Texas Intermediate (WTI) oil prices using United States financial and macro-economic factors. The researchers employed a range of ML techniques, including Decision Tree, Random Forest, Logistic Regression, AdaBoost, and XGBoost algorithms. Authors found that Random Forest model and XGBoost model to be outperforming the traditional models. Moreover, the study used Shapley Additive exPlanations values for model evaluation and interpretation. Similarly, Jahandoost et al. ( 2023 ) proposed hybrid architecture for DNNs, and used 39 features for the prediction of the oil price. The proposed architecture enhances the accuracy of previous models. (Sen et al. 2023 ) proposed LSTM to estimate the oil prices by using the financial variables as predictor. In machine learning, ensemble learning is a technique that integrates predictions from multiple models to enhance the overall model's precision. This method combines various models' outputs to achieve improved accuracy in predictions. The objective is to reduce the errors in the predictions from other models. In ensemble approach, Hasan et al. ( 2023 ) combined ML and ensemble learning to predict the crude oil prices. Based on the plotted graph of actual vs predicted values, Ada Boost algorithm predicts oil prices with high accuracy. Further, the MSE, RMSE, MAE, MAPE, R^2, and function variance scores are used in validating the results. Tami & Owda ( 2024 ) developed the novel LSTM model to estimate the commodity prices. The model used features of moving average, price volatility, past prices. The finding based on the RMSE, MAPE, and R-squared conclude that proposed model is better than the traditional model. Additionally, the LSTM architecture significantly reduces computation time, allowing training to be completed in minutes rather than hours. Sajid et al. ( 2023 ) forecasted the brent oil prices by using the ML and the ensemble learning methodology. Study applies the Light Gradient Bosting, Random Forest ensemble ML algorithm, Lasso regression, and Decision tree ML algorithm. Out of the applied ML algorithms, Light GBM is found to be most accurate. Kim & Jang ( 2023 ) employed hybrid model to improve estimated abilities of the other hybrid models. The author used CNN, and RNN. Based on the evaluation metrics, the findings indicate the proposed model to be the better predictor as compared to the GRU and LSTM based models. Cheng et al. ( 2024 ) studied efficacy of combined estimate, and ML models in prediction of brent oil price volatility. Author found the ML to be more promising than the forecasting combination methods. Similarly, Hao ( 2024 ) proposed novel ML based model in estimation of the oil prices. The author combined the transformer algorithms and LSTM in forecasting the brent crude oil prices. The experimental results confirm the effectiveness of the examined method. The review of the literature reveals a significant surge in the application of ML for time series forecasting, particularly in oil price prediction. While studies have explored various ML techniques, including ARIMA, SVM, LSTM, and ensemble methods, several limitations persist. Many studies focus on individual ML models or combinations of traditional econometric methods with singular ML approaches, often neglecting a comprehensive comparative analysis of diverse ML algorithms. Furthermore, the geographical focus of research is skewed towards the United States, China, and the United Kingdom, potentially limiting the global applicability of findings. Despite the increasing complexity of hybrid models, such as CNN-LSTM and ensemble methods, there remains a gap in understanding the relative performance of a wide array of ML models specifically tailored for crude oil price forecasting. Additionally, while some studies incorporate external financial and macroeconomic variables, a systematic investigation into the optimal selection and integration of these features across various ML architectures is lacking. Therefore, a research gap exists in providing new insights into forecasting crude oil prices by conducting a thorough, comparative analysis of a broad spectrum of machine learning approaches, and systematically investigating the impact of feature selection and model combination, to identify the most effective and robust methods for this critical forecasting task. 3. Methodology and Data To capture the inherent non-linearity and non-stationarity of crude oil prices, as evidenced by data outliers, this study forgoes traditional pre-processing. Data, obtained from the IMF (IMF, 2024), utilizes lagged values ( lag = 1 ) as predictors. Following established practices in the literature (Kandil et al., 2023 ; Nguyen et al., 2023 ), an 80 − 20 train-test split is employed. 3.1 Machine Learning Models There are various type of ML models used in the prediction. The most common and popular among them are Linear Regression, SVM, Random Forest, Xgboost, and Gradient Boosting Machine (Adeniyi et al. 2022 ; Adenusi et al. 2022 ; Albahooth 2024 ; Bakshi et al. 2021 ; Boussatta et al. 2023 ; Ding et al. 2022 ; Drucker et al. 1996 ; Gao et al. 2022 ; Gong & Zhang 2022 ; Guliyev & Mustafayev 2022 ; Jahanshahi et al. 2022 ; Lakshminarayanan & McCrae 2019; Li et al. 2024 ; Manjula & Karthikeyan 2019 ; Murugesan et al. 2023 ; Nayak et al. 2021 ; Okechukwu Ajakwe et al. 2020 ; Sajid et al. 2023 ; Sulaiman et al. 2023 ; Tang et al. 2015 ; Tissaoui et al. 2023 ; Vapnik 1998 ; Zaidi & Oussalah 2018 ). Linear regression seeks the straight line that most closely aligns with this overall trend. The data is trained in linear regression, and this trained data helps the model discover the ideal coefficients for the straight-line equation that best captures the trend in the data. The equation of Linear Regression can be written as Hope ( 2020 ). $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:y=C+\beta\:x+ϵ$$ 1 Equation ( 1 ) represents how the target variable \(\:y\) is predicted based on the input \(\:x\) . Here, \(\:C\) is the intercept, \(\:\beta\:\) is the coefficient showing the impact of \(\:x\) on \(\:y\) , and \(\:ϵ\) is the error term accounting for unexplained variation. The model learns the best values for \(\:C\) and \(\:\beta\:\) to minimize prediction error and accurately estimate \(\:y\) . The SVM model, though applied mostly in classification problem, is commonly used model in forecasting. In regression, SVM introduces an ε-insensitive loss function which creates a hyperplane which minimize the difference among the estimated values of the training sample, and observed value of the response. In regression, a hyperplane along with ε creates an ε-insensitive tube for creating generalization bunds for the regression. The optimization is done by minimizing the ε-insensitive tube to as flat as possible, and simultaneously retaining the maximum training set. Common SVR Equation is - $$\:\:\:\:F\left(x\right){=W}^{T}x+b$$ 2 Where \(\:F\left(x\right)\) is product of vector weight \(\:W\) and input vector \(\:x\) . The tolerance limit is determined by the ε, also known as loss function. The SVR can be written as Vapnik ( 1998 ). $$\:mi{n}_{w,b,{{\xi\:}}_{i},{{\xi\:}}_{i}^{*}}\left(\frac{1}{2}|w{|}^{2}+C{\sum\:}_{i=1}^{n}\left({{\xi\:}}_{i}+{{\xi\:}}_{i}^{*}\right)\right)$$ 3 Subject to – $$\:{y}_{i}-⟨w,{x}_{i}⟩-b\le\:\epsilon\:+{\xi\:}_{i}$$ $$\:⟨w,{x}_{i}⟩+b-{y}_{i}\le\:\epsilon\:+{\xi\:}_{i}^{*}$$ $$\:-{\xi\:}_{i},\hspace{0.25em}{\xi\:}_{i}^{*}\ge\:0,\hspace{1em}i=1,\dots\:,N$$ \(\:\frac{1}{2}{\left|\left|W\right|\right|}^{2}\:\) is the regularization term that penalizes complex models, by controlling the magnitude of the weighted vector. Further, \(\:C*\sum\:{(\xi\:}_{n}+{{\xi\:}_{n}}^{*})\) is empirical loss function, where \(\:C\) determines the weight of the errors. A large \(\:C\) gives higher weight to the minimize the prediction errors, whereas low \(\:C\) value give higher weight to the minimize the flatness. The slack variables \(\:{\xi\:}_{n}\) , and \(\:{{\xi\:}_{n}}^{*}\) determine the points which can be tolerated outside the ε-insensitive tube. As per In a Random Forest regression model, the final prediction for a given input \(\:x\) is obtained by averaging the predictions of all individual decision trees in the ensemble. If \(\:{T}_{b}\left(x\right)\) denotes the prediction from the b th tree and there are B such trees in total, then the overall Random Forest prediction \(\:\widehat{f}\left(x\right)\) is given by the equation: $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\widehat{f}\left(x\right)=\frac{1}{B}{\sum\:}_{b=1}^{B}{T}_{b}\left(x\right)$$ 4 This averaging process helps reduce variance and improves the model’s generalization ability, making Random Forests robust and effective for regression task (Breiman, 2001 ; Dudek, 2015 ). XGboost combines the points which were flawed in decision tree. XGBoost involves summation of multiple trees as different stages in a boosting process. The objective function measures the difference among the estimated and the observed values. The function is - $$\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:\:L({Y}_{i},\widehat{f}\left(x\right))$$ 5 loss function for the \(\:{i}^{th}\) data point, where \(\:{Y}_{i}\) is the observed value, and \(\:\widehat{f}\left(x\right)\) is the forecasted value. In gradient boosting framework, focus is on errors, and each stage tries to better predict the previous stage. $$\:\widehat{f}\left(x\right)=\:\sum\:{F}_{t}\left(x\right)$$ 6 Where, \(\:\widehat{f}\left(x\right)\) is the final prediction, and \(\:{F}_{t}\left(x\right)\) is final prediction from the \(\:{t}^{th}\:\) tree in ensemble. Gradient Boosting Machine, abbreviated as GBM, is a type of ensemble learning technique. This method works by combining predictions from several less robust models to produce a single, more powerful model. These weaker models are typically decision trees. 3.2 Common Regression Evaluation Metrics Evaluating models goes beyond simply checking their accuracy; it's about assessing their reliability in making predictions. The accuracy and reliability are crucial in making forecast. There are various metrics which evaluate the models, including MSE, RMSE, and R-squared. Various studies have used the MSE, RMSE, MAE, MAPE and R-Square for the evaluation of forecasting accuracy of the machine learning models (Alqahtani & Abdelhafez 2023 ; Aziz et al. 2022 ; Chen 2023 ; Kandil et al. 2023 ; Muganda & Kasamani 2023 ; Nanthiya et al. 2023 ; Prakash & Singh 2022 ; Wang et al. 2022 ; Chicco et al., 2021 ). In machine learning, MSE and RMSE gauge how accurate a model's predictions are by capturing the average difference among forecasted and observed values. MSE acts like a trainer during the learning process, guiding the model to minimize this error. It excels at penalizing the large mistakes because it penalizes huge differences more heavily. However, MSE is measured in squared units, different from the target variable's unit, making interpretation a bit complex. To address this, RMSE simply takes the root square of MSE, presenting the error in the same units as the target variable for easier understanding. Both metrics are popular for their focus on large errors but are also sensitive to outliers, so keep the target variable's unit in mind when evaluating the error values. The R squared focuses on the part of variance explained in the dependent variable by the model. It essentially quantifies how well the model fits the data, indicating the degree to which it can explain the target variable's variance. R-squared is an easy-to-understand measure of model fit, and varies from zero to one, where one signifies a perfect fit. While it's scale-independent, adding more predictors can artificially inflate R-squared. Importantly, R-squared about the model's predictive power; it simply helps us understand the proportion of variance explained by the factors considered in the model (Kadri and Abdennbi, 2020 ; Kumar et al. 2020 ). MAE and MAPE are common metrics used to evaluate the accuracy of regression models. MAE measures the average absolute difference between predicted and actual values, providing a clear and interpretable indication of prediction error in the same units as the data. In contrast, MAPE expresses this error as a percentage of the actual values, making it useful for comparing performance across different datasets or scales. While MAE is unaffected by the scale of the data, MAPE can be sensitive to very small actual values, which may inflate the error percentage. The MSE, can be written as- $$\:MSE=\:\frac{1}{m}\sum\:_{m}^{i}({{X}_{i}-{Y}_{i})}^{2}$$ 7 Where, \(\:{X}_{i}\:\) is predicted value, and \(\:{Y}_{i}\) are the actual values. Whereas, the formula for RMSE can be written as- RMSE = \(\:\sqrt{\frac{1}{m}\sum\:_{m}^{i}({{X}_{i}-{Y}_{i})}^{2}}\:\:\:\:\:\) (8) and R-Square formula is $$\:{R}^{2}=1-\frac{{\sum\:}_{i=1}^{m}{\left({Y}_{i}-{X}_{i}\right)}^{2}}{{\sum\:}_{i=1}^{m}{\left({Y}_{i}-\stackrel{-}{{Y}_{i}}\right)}^{2}}$$ 9 Where \(\:\stackrel{-}{{Y}_{i}}\) is the mean value of the observed values. The MAE and MAPE can be written as. $$\:\text{}\text{MAE}=\frac{1}{m}{\sum\:}_{i=1}^{m}\left|{X}_{i}-{Y}_{i}\right|$$ 10 $$\:\text{MAPE}=\frac{1}{m}{\sum\:}_{i=1}^{m}\left|\frac{{Y}_{i}-{X}_{i}}{{Y}_{i}}\right|$$ 11 Hyperparameter Tuning Processes The hyperparameter tuning is crucial in improving forecast accuracy in machine learning and deep learning, which is to select optimum parameters based on specific model. There are various approaches of hyperparameter tuning, such as Grid Search, Bayesian Optimisation, and Random search. The approach depends upon the model selected, and as per the literature, Grid Search is highlighted for the linear regression, whereas for Random Forest, Random Search is preferred. Bayesian optimisation is recommended in SVR, XGBoost, and Gradient Boosting Model. (Hoque and Aljamaan, 2021 ; Rimal et al., 2024 ; Dhilsath and Samuel, 2021 ; Singh et al., 2023; Zhu et al.,2022). For the Linear Regression model, hyperparameter tuning was conducted using Grid Search with Cross-Validation (GridSearchCV) to identify the optimal value of the regularization parameter, alpha (λ). Given that alpha controls the L2 penalty, which discourages large coefficients and mitigates multicollinearity, its careful selection is critical. Six different alpha values—0.1, 1, 10, 100, 500, and 1000—were evaluated using 5-fold cross-validation, where model performance was assessed using negative MSE. The best alpha value was found to be 1000, indicating that a higher degree of regularization was beneficial. This suggests the presence of noise or multicollinearity in the dataset and implies that stronger penalization of coefficients led to better generalization. By effectively shrinking the coefficients, the model avoided overfitting and demonstrated robust performance across training and test sets, striking an optimal bias-variance trade-off. In contrast, the Random Forest model employed RandomizedSearchCV for hyperparameter tuning, a more computationally efficient alternative to grid search. It explored 30 random combinations across a specified hyperparameter space, including the number of estimators, maximum tree depth, minimum samples required for node splits and leaf nodes, and bootstrap sampling. Using 5-fold cross-validation and negative MSE as the scoring metric, the optimal configuration was identified as n_estimators = 200, max_depth = 10, and min_samples_split = 10. These values suggest a deliberate attempt to prevent overfitting by limiting model complexity while maintaining high predictive accuracy. The use of parallel computation (n_jobs=-1) enhanced efficiency. This strategy not only balanced model accuracy and generalization but also demonstrated computational prudence. The SVR model utilized Optuna, a state-of-the-art optimization framework, to fine-tune three critical hyperparameters: C, epsilon, and gamma. After running 30 optimization trials, the optimal configuration—C = 67.8443, epsilon = 0.7399, and gamma = 0.0009586—was found to offer a well-balanced fit. A high C indicates that the model imposed strong penalties on errors, making it more precise but potentially prone to overfitting. However, the moderately high epsilon introduced a tolerance band within which errors were not penalized, adding robustness. A small gamma extended the reach of the radial basis kernel, smoothing the decision boundary and improving generalization. This combination demonstrated that fine-tuned SVR could effectively balance complexity and flexibility to reduce prediction error. Similarly, the XGBoost Regressor (XGBRegressor) underwent tuning through Optuna over 30 trials. The search space spanned several impactful hyperparameters, including n_estimators, max_depth, learning_rate, subsample, colsample_bytree, gamma, reg_alpha, and reg_lambda. The best-performing configuration—colsample_bytree = 0.7432, gamma = 0.3972, learning_rate = 0.0126, max_depth = 4, and n_estimators = 193—indicates a highly conservative learning strategy. The relatively shallow tree depth and low learning rate ensured that the model trained gradually, avoiding abrupt jumps in optimization, while the moderate gamma discouraged overly aggressive tree splits. The balance of these hyperparameters suggests an emphasis on slow, regularized learning that ensures model robustness and strong generalization on unseen data. The GBM model also benefitted from Optuna-based hyperparameter optimization. The tuning process explored key parameters such as the learning rate, minimum samples per leaf, and the minimum number of samples required for a split. After 30 trials, the optimal configuration included a learning rate of 0.23, a minimum of 5 samples per leaf, and 8 samples required for node splitting, along with 72 boosting rounds. These settings reflect a model designed to balance performance and complexity. The relatively higher learning rate compared to XGBoost implies faster learning, while constraints on sample size at leaves and splits help regulate model complexity, reducing the risk of overfitting. The hyperparameter tuning processes across these models demonstrate a consistent theme of balancing model complexity and predictive accuracy. Each approach—be it GridSearchCV, RandomizedSearchCV, or Optuna—was strategically selected based on the model architecture and computational constraints. The success of these tuning strategies reinforces the notion that model performance is not merely a function of algorithm choice, but critically dependent on the thoughtful calibration of hyperparameters tailored to data characteristics and forecasting objectives. 4. Results and Discussion The results presented in Table 1 indicate a strong correlation between lagged oil prices and the estimator, evidenced by a high coefficient and statistically significant p-value. This supports the inclusion of lagged oil prices as a reliable predictor within the linear regression-based machine learning model. Table 1 – Predictor Coefficients Estimate Std. Error T-value Pr(>|t|) Intercept 0.141933 0.085230 1.665 0.096 Lagged_Oil 0.997636 0.001217 819.579 < .001 Multiple R -squared 0.9967 Adjusted R-Squared 0.9967 F Statistics 6.717e + 05 P value < .001 Source – Author’s Work The Linear Regression model exhibits an excellent fit to the oil price data, as evidenced by both the robust evaluation metrics and the visual representation in Fig. 1 , where the predicted and actual prices nearly overlap during the training period. Similarly, the Random Forest model demonstrates strong predictive capability, with Fig. 2 showing a close alignment between predicted and actual values in the training dataset. The SVR model also performs commendably, as shown in Fig. 3 , effectively capturing underlying patterns in both training and test datasets. The XGBoost model, illustrated in Fig. 4 , delivers performance comparable to linear regression, with a strong alignment between predicted and actual oil prices across both datasets. Likewise, the GBM model, depicted in Fig. 5 , shows a high degree of accuracy, as indicated by the close match between predicted and actual values, further confirming its effectiveness in modeling oil price dynamics. The Linear Regression model demonstrates outstanding performance, with exceptionally high R-squared values nearing 0.995 for both training and test sets, alongside consistently low error metrics (MSE, RMSE, MAE, MAPE), indicating strong predictive accuracy and excellent generalization to unseen data, as the predicted test values closely follow the actual ones without signs of overfitting. The Random Forest model also shows impressive performance, with a very low training MSE of 0.7216 and R-squared of 0.9987, though its test performance shows slightly higher error metrics (MSE: 3.3101, RMSE: 1.8194, MAE: 1.3601, MAPE: 2.0399), suggesting a modest decline in generalization while still maintaining a high R-squared of 0.9939. The SVR model offers reliable predictions with R-squared values around 0.99 for both datasets and relatively low error metrics, showing consistent performance and robust generalization, though it performs slightly below the Linear Regression model on this dataset. Similarly, the XGBoost model achieves high predictive accuracy with R-squared values near 0.99 and low error metrics across training and test datasets, confirming strong generalization and a performance level comparable to Linear Regression and better than SVR. The GBM model also yields high R-squared values (≈ 0.99) and low error metrics across both datasets, reinforcing its predictive strength and generalization ability, placing its performance on par with XGBoost and Linear Regression, and ahead of SVR for this dataset. In essence, all five models demonstrate strong predictive capabilities for oil price forecasting, with notable differences in generalization and robustness. The Linear Regression model stands out as a highly accurate and reliable predictor, exhibiting excellent generalization and strong statistical significance, as evidenced by consistently large negative Diebold-Mariano (DM) statistics of -18.548678, indicating a substantial improvement over naive forecasts. The Random Forest model also delivers accurate predictions, but the noticeable disparity between training and test performance suggests potential overfitting, warranting further refinement through regularization, feature engineering, and cross-validation, despite its statistically significant DM value of -18.536257. The SVR model offers solid performance with good accuracy and generalization, supported by highly negative DM statistics of -18.552112, confirming its superiority over naive models. Similarly, the XGBoost model proves to be a robust and accurate predictor, with strong generalization and consistent statistical significance reflected in its DM value of -18.548297. The GBM model mirrors the performance of XGBoost, showing strong predictive accuracy, solid generalization, and a statistically significant improvement over naive forecasts, as indicated by its DM statistic of -18.545874. In conclusion, all five models—Linear Regression, Random Forest, SVR, XGBoost, and GBM—demonstrate strong potential for forecasting oil prices, each exhibiting solid predictive accuracy and statistical significance over naive models. However, based on the evaluation metrics, visual alignment, generalization capability, and consistently strong Diebold-Mariano statistics, the Linear Regression model emerges as the best overall performer. It combines simplicity with high accuracy, excellent generalization, and minimal signs of overfitting, making it a robust and reliable choice for oil price prediction within this dataset. Table 2 –ML Model Performance Comparison Model Dataset MSE RMSE MAE MAPE R² DM Statistics Linear Regression Train 2.471180 1.571999 1.064543 1.673229 0.995626 -18.548678 Test 2.409677 1.552313 1.086515 1.662010 0.995546 Random Forest Train 0.721557 0.849445 0.605434 0.937417 0.998723 -18.536257 Test 3.310095 1.819367 1.360060 2.039885 0.993881 SVR Train 5.575188 2.361184 1.276119 2.470044 0.990131 -18.552112 Test 5.573400 2.360805 1.293814 2.298823 0.989698 XGBoost Train 2.213671 1.487841 1.043201 1.650890 0.996081 -18.548297 Test 2.815016 1.677801 1.186548 1.801803 0.994796 GBM Train 1.830616 1.353002 0.972099 1.488672 0.996760 -18.545874 Test 2.592087 1.609996 1.144261 1.739305 0.995209 Source – Author’s Work 4.1 LSTM and CNN Based Results The Optuna-tuned LSTM and CNN models were optimized with carefully selected hyperparameters to enhance their performance in forecasting oil prices from univariate time series data. The LSTM model employed 50 memory units, striking a balance between learning temporal dependencies and avoiding overfitting, supported by a dropout rate of 0.2 to introduce moderate regularization. A learning rate of 0.001 ensured steady convergence, while a batch size of 32 balanced training stability and generalization. In contrast, the CNN model, optimized for a lag-1 setup, utilized 92 filters in its Conv1D layer to capture fine-grained patterns from recent data, with a kernel size of 1 aligning perfectly with the one-step lag structure. A minimal dropout rate of 0.043 indicated low overfitting risk, while a small learning rate of 0.00027 facilitated cautious weight updates, crucial for handling noisy financial data. Additionally, the batch size of 9 introduced beneficial stochasticity during training, helping the model generalize well. Together, these hyperparameter configurations reflect robust and well-regularized architectures, capable of effectively modeling the complex, nonlinear behavior of oil price movements with precision and stability. The LSTM-based graph illustrates the model's ability to capture long-term dependencies in oil price movements. During the training phase (2013–2022), the predicted values (in cyan) closely align with the actual prices (in blue), indicating that the model has effectively learned historical patterns and underlying seasonality. While minor deviations are visible, they reflect typical behavior in financial time series due to noise, and overall, the model demonstrates strong in-sample learning. In the test period (2023–2024), the predicted values (red dashed line) track the actual oil prices (black line) reasonably well, though there is a slight lag in response to abrupt fluctuations—an expected characteristic of memory-based models like LSTM. This suggests that while the model generalizes effectively, it may underperform when faced with rapid, short-term volatility in oil prices. In contrast, the CNN-based graph reflects a model that is more agile and responsive to local fluctuations in the data. The training predictions (cyan) match the actual prices (blue) closely, albeit with more jaggedness compared to LSTM, which indicates the CNN's sensitivity to short-term variations. This responsiveness becomes even more evident during the test period (2023–2024), where the predicted values (red dashed) closely follow the actual prices (black), demonstrating the CNN’s strength in adapting to recent and abrupt changes in the oil price series. The tighter alignment of predictions in the test phase suggests that CNN may slightly outperform LSTM in short-horizon forecasting, particularly under conditions of high volatility. Overall, while LSTM provides stable and trend-following forecasts, CNN offers more reactive and precise short-term predictions, making each model uniquely suited to different forecasting objectives within financial time series modeling. Based on the evaluation metrics and model configurations, a comparative analysis of the Optuna-tuned LSTM and CNN models for oil price forecasting reveals several important insights regarding their performance and suitability for time series prediction. The LSTM model achieved superior performance on the training dataset, with an MSE of 2.2788, RMSE of 1.5096, and R² of 0.9958, indicating an excellent fit to the training data. It also produced relatively low training errors in terms of MAE (1.1783) and MAPE (2%). On the test dataset, while there was a slight drop in performance, the model still performed well, with a test MSE of 7.0434, RMSE of 2.6539, and R² of 0.9634. The MAE and MAPE on the test set were 2.0139 and 2.17%, respectively. Notably, the Diebold-Mariano (DM) test statistic of 3.8630 (p-value = 0.0001) suggests that the LSTM model's forecasts are statistically better than those of the naive model, with a significant improvement in predictive accuracy. In contrast, the CNN model showed a slightly higher training error (MSE: 5.9133, RMSE: 2.4317), though still strong, with an R² of 0.9892, and a MAE comparable to LSTM (1.1887). Interestingly, on the test set, the CNN outperformed the LSTM in terms of MSE (5.8695 vs. 7.0434), RMSE (2.4227 vs. 2.6539), MAE (1.7265 vs. 2.0139), and MAPE (1.86% vs. 2.17%), with a slightly better R² of 0.9695. These results indicate that the CNN model generalized marginally better than the LSTM model, especially in terms of absolute and percentage error measures. However, the DM test statistic of -0.5173 (p = 0.6051) shows no statistically significant difference between the CNN and naive forecasts, suggesting that the performance gains, while numerically present, are not statistically conclusive. From a modeling perspective, the LSTM’s architecture—optimized with 50 units, a 0.2 dropout rate, and a conservative learning rate—was effective at capturing long-term temporal dependencies, which are crucial in time series forecasting. However, the slightly higher test error and the statistically significant DM result imply that the model may have overfit the training data slightly more than desired. On the other hand, the CNN’s architecture—with 92 filters, a kernel size of 1, low dropout (≈ 0.043), a very small learning rate (0.00027), and a small batch size (9)—provided a robust and precise framework that balanced responsiveness to noise with a capacity to generalize. Its slightly better test metrics and balanced regularization suggest it handled the volatility in oil prices slightly more robustly, although its predictions were not statistically superior to the naive baseline. In conclusion, while both models performed impressively, the LSTM model demonstrated statistically significant forecasting superiority, especially when compared to naive predictions, making it a reliable choice when precision is paramount. The CNN model, although marginally better in test accuracy metrics, lacked statistical significance over naive forecasting but offered excellent generalization and computational efficiency, making it well-suited for applications requiring fast, interpretable models with limited data history ( lag = 1 ). The final choice between the two models could depend on the specific operational or strategic priorities—statistical significance and deeper sequence learning (LSTM) vs. faster generalization with simpler architectures (CNN). Table 3 LSTM and CNN Model Performance Model Dataset MSE RMSE MAE MAPE R² DM Statistics LSTM Train 2.2788 1.5096 1.1783 0.0200 0.9958 3.8630 Test 7.0434 2.6539 2.0139 0.0217 0.9634 CNN Train 5.9133 2.4317 1.1887 0.0301 0.9892 -0.5173 Test 5.8695 2.4227 1.7265 0.0186 0.9695 Source – Author’s work 5. Conclusion and Implications The analysis of various machine learning models for oil price forecasting reveals several key insights. Linear regression stands out for its exceptional performance, achieving near-perfect fits and robust generalization, indicating it effectively captures the underlying linear trends in the data. Tree-based models like Random Forest, XGBoost, and GBM also demonstrate high predictive accuracy, with XGBoost and GBM performing comparably to linear regression, while Random Forest suggests a need for further refinement to mitigate potential overfitting. SVR provides reliable predictions, though slightly less effective than the leading models. In the deep learning domain, LSTM excels in capturing long-term dependencies, showing statistically significant improvements over naive predictions, albeit with a slight lag in responding to abrupt fluctuations. Conversely, CNN demonstrates superior short-term forecasting capabilities, exhibiting agility and responsiveness to local fluctuations, though its statistical significance over naive predictions is less pronounced. The critical role of hyperparameter tuning, using techniques like Grid Search, RandomizedSearchCV, and Optuna, is emphasized, highlighting its importance in optimizing model performance and balancing complexity with generalization. This study, while effective, is limited by its reliance on historical data, neglecting sudden market changes. Model generalization to other datasets is uncertain, and the univariate approach omits potentially valuable external factors. Also, some statistical significance results require further validation. These findings underscore the need for model selection based on specific forecasting objectives, the importance of evaluating generalization, and the necessity of rigorous hyperparameter tuning. The statistical significance of the models predictions, is also a very important factor. Practically, these models have significant implications for industries reliant on accurate oil price predictions, informing decision-making and risk mitigation. Future research should explore integrating external factors, investigating hybrid model architectures, and expanding the dataset to enhance forecasting accuracy and robustness. Abbreviations CNN Convolutional Neural Network SVM Support Vector Machine DNN Deep Neural Network LSTM Long Short-Term Memory MSE Mean Squared Error RMSE Root Mean Squared Error GRU Gated Recurrent Unit GBM Gradient Boosting Machine ML Machine Learning AI Artificial Intelligence ARIMA Autoregressive Integrated Moving Average ANN Artificial Neural Network MAE Mean Absolute Error Declarations The manuscript is original and has not been published previously in any form. The research work is conducted ethically and responsibly. There is no conflict of interest related to this study. All sources of data and materials used in the study have been properly acknowledged. References Adeniyi EA, Gbadamosi B, Awotunde JB, Misra S, Sharma MM, Oluranti J (2022) Crude Oil Price Prediction Using Particle Swarm Optimization and Classification Algorithms . 418 LNNS , 1384–1394. Scopus. https://doi.org/10.1007/978-3-030-96308-8_128 Adenusi CA, Vincent OR, Bolarinwa J, Oluwakemi O (2022) Predicting the Pandemic Effect of COVID 19 on the Nigeria Economic, Crude Oil as a Measure Parameter Using Machine Learning. Int Ser Oper Res Manage Sci 320:79–93. https://doi.org/10.1007/978-3-030-87019-5_5 . Scopus Ahmed S, Alshater MM, Ammari AE, Hammami H (2022) Artificial intelligence and machine learning in finance: A bibliometric review. Research in International Business and Finance , 61 . Scopus. https://doi.org/10.1016/j.ribaf.2022.101646 Ahmed H, Siddiqui TA, Naushad M (2025) Navigating global economic turmoil: The dynamics of oil prices, exchange rates, and stock markets in BRICS. Invest Manage Financial Innovations 22(1):94 Ahmed H, Kaur R (2025) Dynamic Connectedness Among Oil Prices, Exchange Rate and Consumer Inflation: New Evidence from India. IIM Kozhikode Soc Manage Rev, 22779752251346247 Ahmed H (2025) Leveraging Machine Learning for Exchange Rate Prediction: Fresh Insights from BRICS Economies. Int J Econ Financial Issues 15(4):72 Albahooth B (2024) Oil price movements predictions in Kingdom of Saudi Arabia using financial and macro-economic variables. Energy Explor Exploit 42(2):747–771. https://doi.org/10.1177/01445987231206897 . Scopus Aldabagh H, Zheng X, Mukkamala R (2023) A Hybrid Deep Learning Approach for Crude Oil Price Prediction. J Risk Financial Manage 16(12). https://doi.org/10.3390/jrfm16120503 . Scopus Alqahtani MG, Abdelhafez HA (2023) STOCK MARKET PREDICITION USING STATISTICAL & DEEP LEARNING TECHNIQUES. J Theoretical Appl Inform Technol 101(23):7808–7825 Scopus Atif M, Rabbani MR, Jreisat A, Al-Mohamad S, Siddiqui TA, Hussain H, Ahmed H (2022) Time Varying Impact of Oil Prices on Stock Returns: Evidence from Developing Markets. Int J Sustainable Dev Plann, 17 (2) Aziz MIA, Barawi MH, Shahiri H (2022) Is Facebook PROPHET Superior than Hybrid ARIMA Model to Forecast Crude Oil Price? Sains Malaysiana 51(8):2633–2643. https://doi.org/10.17576/jsm-2022-5108-22 . Scopus Bakshi SS, Jaiswal RK, Jaiswal R (2021) Efficiency Check Using Cointegration and Machine Learning Approach: Crude Oil Futures Markets . 191 , 304–311. Scopus. https://doi.org/10.1016/j.procs.2021.07.038 Barsky RB, Kilian L (2002) Oil and the macroeconomy since the 1970s. J Economic Perspect 18(4):115–134 Baumeister C, Kilian L (2016) Forty years of oil price fluctuations: Why the price of oil may still surprise us. J Economic Perspect 30(1):139–160 Bernabe A, Martina E, Alvarez-Ramirez J, Ibarra-Valdez C (2004) A multi-model approach for describing crude oil price dynamics. Physica A 338(3–4):567–584 Breiman L (2001) Random Forests. Springer, New York. https://doi.org/10.1007/978-1-4757-3264-1 Boussatta H, Chihab M, Chihab Y, Chiny M (2023) Enhancing Oil Price Forecasting Through an Intelligent Hybridized Approach. Int J Adv Comput Sci Appl 14(9):115–125. https://doi.org/10.14569/IJACSA.2023.0140913 . Scopus Chen J (2023) Analysis of Bitcoin Price Prediction Using Machine Learning. J Risk Financial Manage 16(1). https://doi.org/10.3390/jrfm16010051 . Scopus Chen S-S, Hsu K-W (2012) Reverse globalization: Does high oil price volatility discourage international trade? Energy Econ 34(5):1634–1643 Cheng W, Ming K, Ullah M (2024) Oil price volatility prediction using out-of-sample analysis – Prediction efficiency of individual models, combination methods, and machine learning based shrinkage methods. Energy , 300 . Scopus. https://doi.org/10.1016/j.energy.2024.131496 Cheng Y, Yi J, Yang X, Lai KK, Seco L (2022) A CEEMD-ARIMA-SVM model with structural breaks to forecast the crude oil prices linked with extreme events. Soft Comput 26(17):8537–8551. https://doi.org/10.1007/s00500-022-07276-5 . Scopus Chicco D, Warrens MJ, Jurman G (2021) The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. Peerj Comput Sci 7:e623 Das PK, Das A (2020) Application of nonlinear stochastic single source of error state space models in the forecasting of mobile subscribers in India. Int J Data Sci 5(4):333–357 Das PK, Das PK (2024) Improvement in Inflation Forecasting: Ensembling Text Mining with Macro Data in Machine Learning Models. Int J Econ Finance 16(6):1–92 Dhilsath FM, Samuel SJ (2021) Hyperparameter tuning of ensemble classifiers using grid search and random search for prediction of heart disease. Comput Intell Healthc Inf, 139–158 Ding X, Fu L, Ding Y, Wang Y (2022) A novel hybrid method for oil price forecasting with ensemble thought. Energy Rep 8:15365–15376. https://doi.org/10.1016/j.egyr.2022.11.061 . Scopus Drucker H, Burges CJ, Kaufman L, Smola A, Vapnik V (1996) Support vector regression machines. Advances in Neural Information Processing Systems , 9 . https://proceedings.neurips.cc/paper/1996/hash/d38901788c533e8286cb6400b40b386d-Abstract.html Dudek G (2015) Short-term load forecasting using random forests. In Intelligent Systems' 2014: Proceedings of the 7th IEEE International Conference Intelligent Systems IS’2014, September 24-26, 2014, Warsaw, Poland, Volume 2: Tools, Architectures, Systems, Applications (pp. 821–828). Springer International Publishing Fang T, Zheng C, Wang D (2023) Forecasting the crude oil prices with an EMD-ISBM-FNN model. Energy 263:125407 Gabralla LA, Abraham A (2015) Comparison of hybrid intelligent approaches for prediction of crude oil price. Int J Comput Inform Syst Industrial Manage Appl 7(1):53–65 Scopus Gao X, Wang J, Yang L (2022) An Explainable Machine Learning Framework for Forecasting Crude Oil Price during the COVID-19 Pandemic. Axioms 11(8). https://doi.org/10.3390/axioms11080374 . Scopus Gholamian MR, Ghomi SF, Ghazanfari M (2005) A hybrid systematic design for multiobjective market problems: A case study in crude oil markets. Eng Appl Artif Intell 18(4):495–509 Gong Y, Zhang P (2022) Research and Implementation of Oil Trading Data Analysis and Prediction based on Random Forest Regression Algorithm . 14–22. Scopus. https://doi.org/10.1109/AIAM57466.2022.00011 Guliyev H, Mustafayev E (2022) Predicting the changes in the WTI crude oil price dynamics using machine learning models. Resources Policy , 77 . Scopus. https://doi.org/10.1016/j.resourpol.2022.102664 Hagen R (1994) How is the international price of a particular crude determined? OPEC Rev 18(1):127–135 Hamilton JD, Herrera AM (2004) Comment: Oil shocks and aggregate macroeconomic behavior: the role of monetary policy. J Money Credit Bank, 265–286 Hao X (2024) Crude Oil Prediction Based on Multi-Factor LSTM-Transformer Algorithm . 51 , 507–514. Scopus. https://doi.org/10.3233/ATDE240114 Hasan M, Das U, Datta RK, Abedin MZ (2023) Model Development for Predicting the Crude Oil Price: Comparative Evaluation of Ensemble and Machine Learning Methods. In International Series in Operations Research and Management Science (Vol. 336, pp. 167–179). Scopus. https://doi.org/10.1007/978-3-031-18552-6_10 Hope TM (2020) Linear regression. Machine learning. Academic, pp 67–81 Hoque KE, Aljamaan H (2021) Impact of hyperparameter tuning on machine learning models in stock price forecasting. IEEE Access 9:163815–163830 International Monetary Fund (2024) Global price of Brent Crude [POILBREUSDM], retrieved from FRED, Federal Reserve Bank of St. Louis; https://fred.stlouisfed.org/series/POILBREUSDM , June 7 Jahandoost A, Houshmand M, Hosseini SA (2023) Prediction of West Texas Intermediate Crude-oil Price Using Hybrid Attention-based Deep Neural Networks: A Comparative Study . 240–245. Scopus. https://doi.org/10.1109/ICCKE60553.2023.10326291 Jahan-Parvar MR, Mohammadi H (2009) Oil prices and competitiveness: Time series evidence from six oil-producing countries. J Economic Stud 36(1):98–118 Jahanshahi H, Uzun S, Kaçar S, Yao Q, Alassafi MO (2022) Artificial Intelligence-Based Prediction of Crude Oil Prices Using Multiple Features under the Effect of Russia–Ukraine War and COVID-19 Pandemic. Mathematics 10(22). https://doi.org/10.3390/math10224361 . Scopus Jha N, Tanneru K, Palla H, S., Hussain Mafat I (2024) Multivariate analysis and forecasting of the crude oil prices: Part I – Classical machine learning approaches. Energy , 296 . Scopus. https://doi.org/10.1016/j.energy.2024.131185 Kadri F, Abdennbi K (2020) Rnn-based deep-learning approach to forecasting hospital system demands: application to an emergency department. Int J Data Sci 5(1):1–25 Kakade KA, Ghate KS, Jaiswal RK, Jaiswal R (2023) A Novel Approach to Forecast Crude Oil Prices Using Machine Learning and Technical Indicators. J Adv Inform Technol 14(2):302–310. https://doi.org/10.12720/jait.14.2.302-310 . Scopus Kandil A, Khaled S, Elfakharany T (2023) Prediction of the equivalent circulation density using machine learning algorithms based on real-time data. AIMS Energy 11(3):425–453. https://doi.org/10.3934/energy.2023023 . Scopus Kaufmann RK, Dees S, Karadeloglou P, Sanchez M (2004) Does OPEC Matter? An Econometric Analysis of Oil Prices. Energy J 25(4):67–90. https://doi.org/10.5547/ISSN0195-6574-EJ-Vol25-No4-4 Keerthan JS, Nagasai Y, Shaik S (2019) Machine learning algorithms for oil price prediction. Int J Innovative Technol Exploring Eng 8(8):958–963 Scopus Khashman A, Nwulu NI (2011) Intelligent prediction of crude oil price using Support Vector Machines . 165–169. Scopus. https://doi.org/10.1109/SAMI.2011.5738868 Kim GI, Jang B (2023) Petroleum Price Prediction with CNN-LSTM and CNN-GRU Using Skip-Connection. Mathematics 11(3). https://doi.org/10.3390/math11030547 . Scopus Kumar R, Kumar P, Kumar Y (2020) Time series data prediction using IoT and machine learning technique. Procedia Comput Sci 167:373–381 Lakshminarayanan SK, McCrae J (2019) A comparative study of SVM and LSTM deep learning algorithms for stock market prediction . 2563 , 446–457. Scopus. https://www.scopus.com/inward/record.uri?eid=2-s2.0-85081603042&partnerID=40&md5=e3cf7e4ee4ff1913e56939f57e0b3843 Li N, Li J, Wang Q, Yan D, Wang L, Jia M (2024) A novel copper price forecasting ensemble method using adversarial interpretive structural model and sparrow search algorithm. Resources Policy , 91 . Scopus. https://doi.org/10.1016/j.resourpol.2024.104892 Liang X, Luo P, Li X, Wang X, Shu L (2023) Crude oil price prediction using deep reinforcement learning. Resources Policy , 81 . Scopus. https://doi.org/10.1016/j.resourpol.2023.103363 Liu L, Zhou S, Jie Q, Du P, Xu Y, Wang J (2024) A robust time-varying weight combined model for crude oil price forecasting. Energy , 299 . Scopus. https://doi.org/10.1016/j.energy.2024.131352 Manjula KA, Karthikeyan P (2019) Gold price prediction using ensemble based machine learning techniques . 2019-April , 1360–1364. Scopus. https://doi.org/10.1109/icoei.2019.8862557 Mohsin M, Jamaani F (2023) Green finance and the socio-politico-economic factors’ impact on the future oil prices: Evidence from machine learning. Resources Policy , 85 . Scopus. https://doi.org/10.1016/j.resourpol.2023.103780 Muganda BW, Kasamani BS (2023) Parallel Programming for Portfolio Optimization: A Robo-Advisor Prototype using Genetic Algorithms with Recurrent Neural Networks . 167–176. Scopus. https://doi.org/10.1109/ICCNS58795.2023.10193396 Müller KR, Smola AJ, Rätsch G, Schölkopf B, Kohlmorgen J, Vapnik V (1998) Using support vector machines for time series prediction Murugesan R, Azhaganathan B, Maitra S (2023) Commodities vs. S&P 500: Causal interaction, temporal analysis and predictive modelling using econometric approach, machine learning, and deep learning. Int J Bus Inform Syst 42(3–4):429–457. https://doi.org/10.1504/IJBIS.2023.129725 . Scopus Naeem M, Aamir M, Yu J, Albalawi O (2024) A Novel Approach for Reconstruction of IMFs of Decomposition and Ensemble Model for Forecasting of Crude Oil Prices. IEEE Access 12:34192–34207. https://doi.org/10.1109/ACCESS.2024.3370440 . Scopus Nanthiya D, Gopal SB, Balakumar S, Harisankar M, Midhun SP (2023) Gold Price Prediction using ARIMA model . ViTECoN 2023–2nd IEEE International Conference on Vision Towards Emerging Trends in Communication and Networking Technologies, Proceedings. Scopus. https://doi.org/10.1109/ViTECoN58111.2023.10157017 Nayak RK, Tripathy R, Mishra D, Burugari VK, Selvaraj P, Sethy A, Jena B (2021) Indian stock market prediction based rough set support vector Mach approach 153:345–355. https://doi.org/10.1007/978-981-15-6202-0_35 . Scopus Nguyen TT, Nguyen HG, Lee JY, Wang YL, Tsai CS (2023) The consumer price index prediction using machine learning approaches: Evidence from the United States. Heliyon, 9 (10) Okechukwu Ajakwe S, Nwakanma I, Lee C, J.-M., Kim D-S (2020) Machine Learning Algorithm for Intelligent Prediction for Military Logistics and Planning . 2020-October , 417–419. Scopus. https://doi.org/10.1109/ICTC49870.2020.9289286 Pai PF, Lin KP, Lin CS, Chang PT (2010) Time series forecasting by a seasonal support vector regression model. Expert Syst Appl 37(6):4261–4265 Panella M, Barcellona F, D’Ecclesia RL (2012) Forecasting energy commodity prices using neural networks. Advances in Decision Sciences , 2012 . Scopus. https://doi.org/10.1155/2012/289810 Prakash A, Singh SK (2022) A Comparative Study of Time Series, Machine Learning, and Ensemble Models for Crude Oil Price Prediction. 914:157–171. https://doi.org/10.1007/978-981-19-2980-9_13 . Scopus Rimal Y, Sharma N, Alsadoon A (2024) The accuracy of machine learning models relies on hyperparameter tuning: student result classification using random forest, randomized search, grid search, bayesian, genetic, and optuna algorithms. Multimedia Tools Appl 83(30):74349–74364 Siddiqui TA, Ahmed H, Naushad M, Khan U (2023) The relationship between oil prices and exchange rate: a systematic literature review. Int J Energy Econ Policy 13(3):566–578 Sajid SW, Hasan M, Rabbi MF, Abedin MZ (2023) An Ensemble LGBM (Light Gradient Boosting Machine) Approach for Crude Oil Price Prediction. In International Series in Operations Research and Management Science (Vol. 336, pp. 153–165). Scopus. https://doi.org/10.1007/978-3-031-18552-6_9 Sen D, Hamurcuoglu KI, Ersoy MZ, Tunç KMM, Günay ME (2023) Forecasting long-term world annual natural gas production by machine learning. Resources Policy , 80 . Scopus. https://doi.org/10.1016/j.resourpol.2022.103224 Shao YE, Lu C-J, Hou C-D (2014) Hybrid soft computing schemes for the prediction of import demand of crude oil in Taiwan. Mathematical Problems in Engineering , 2014 . Scopus. https://doi.org/10.1155/2014/257947 Stevens P (1995) Understanding the Oil Industry: Economics as a Help or a Hindrance. Energy J 16(3):125–139. https://doi.org/10.5547/ISSN0195-6574-EJ-Vol16-No3-6 Sulaiman A, Mahmood FE, Majeed SA (2023) Long-Term Solar Irradiance Forecasting Using Multilinear Predictors. Int J Electr Electron Eng Telecommunications 12(2):134–141. https://doi.org/10.18178/ijeetc.12.2.134-141 . Scopus Tami M, Owda AY (2024) Efficient commodity price forecasting using long short-term memory model. IAES Int J Artif Intell 13(1):994–1004. https://doi.org/10.11591/ijai.v13.i1.pp994-1004 . Scopus Tang L, Dai W, Yu L, Wang S (2015) A novel CEEMD-based eelm ensemble learning paradigm for crude oil price forecasting. Int J Inform Technol Decis Mak 14(1):141–169. https://doi.org/10.1142/S0219622015400015 . Scopus Tebyanian A, Hedayati F (2014) Intelligent crude oil price forecaster . 453–455. Scopus. https://doi.org/10.1109/ICMLA.2014.79 Tissaoui K, Zaghdoudi T, Hakimi A, Nsaibi M (2023) Do Gas Price and Uncertainty Indices Forecast Crude Oil Prices? Fresh Evidence Through XGBoost Modeling. Comput Econ 62(2):663–687. https://doi.org/10.1007/s10614-022-10305-y . Scopus Vapnik V (1998) Statistical learning theory Wiley. New York 1(624):2 Wang L, Xia Y, Lu Y (2022) A Novel Forecasting Approach by the GA-SVR-GRNN Hybrid Deep Learning Algorithm for Oil Future Prices. Computational Intelligence and Neuroscience , 2022 . Scopus. https://doi.org/10.1155/2022/4952215 Wei MX, Kee OL, Musa S (2023) Seasonal versus non-seasonal trends in stock market Malaysia . 389 . Scopus. https://doi.org/10.1051/e3sconf/202338909039 Weng F, Chen Y, Wang Z, Hou M, Luo J, Tian Z (2020) Gold price forecasting research based on an improved online extreme learning machine algorithm. J Ambient Intell Humaniz Comput 11(10):4101–4111. https://doi.org/10.1007/s12652-020-01682-z . Scopus Xu Z, Mohsin M, Ullah K, Ma X (2023) Using econometric and machine learning models to forecast crude oil prices: Insights from economic history. Resources Policy , 83 . Scopus. https://doi.org/10.1016/j.resourpol.2023.103614 Yu L, Zhang X, Wang S (2017) Assessing potentiality of support vector machine method in crude oil price forecasting. Eurasia J Math Sci Technol Educ 13(12):7893–7904. https://doi.org/10.12973/ejmste/77926 . Scopus Zaidi A, Oussalah M (2018) Forecasting weekly crude oil using twitter sentiment of U.S. foreign policy and oil companies data . 201–208. Scopus. https://doi.org/10.1109/IRI.2018.00037 Zhu N, Zhu C, Zhou L, Zhu Y, Zhang X (2022) Optimization of the random forest hyperparameters for power industrial control systems intrusion detection using an improved grid search algorithm. Appl Sci 12(20):10456 Additional Declarations The authors declare no competing interests. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-8354755","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":559894968,"identity":"11eb5133-e155-4f57-b502-1a55a291cae6","order_by":0,"name":"Haseen","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABBUlEQVRIiWNgGAWjYDACCSDmbQAzGR8kGNiA6MYDxGphNnhQkQbS0kC0FjbJB2cOg1l4tfDP7k588HaHnT0/e+8BicS283Zr2w8DbamxicZpyZ2zmw3nnklOnNlzLsEgse128rYziUAtx9JyG3DpuZG7TZq3jTnB4EaOQQJIi9kBoBbGhsM4tcjfyN3+m7et3t7+/huDA4lt55LNzj/Er8UAaAszb9thxg0SPIYNCWcO2JndIGCL4Y3czZJz244nzjiTY8yQUJGcYHYDaEsCHr/I3cjd+OFtW7U9f/sZ858/DOzszc6nP3zwocYGt/fRQSJYZQKxykHAnhTFo2AUjIJRMDIAAKuea8RUYZtmAAAAAElFTkSuQmCC","orcid":"https://orcid.org/0009-0001-1323-5162","institution":"New Delhi Institute of Management","correspondingAuthor":true,"prefix":"","firstName":"","middleName":"","lastName":"Haseen","suffix":""}],"badges":[],"createdAt":"2025-12-13 19:47:18","currentVersionCode":1,"declarations":{"humanSubjects":false,"vertebrateSubjects":true,"conflictsOfInterestStatement":false,"humanSubjectEthicalGuidelines":false,"humanSubjectConsent":false,"humanSubjectClinicalTrial":false,"humanSubjectCaseReport":false,"vertebrateSubjectEthicalGuidelines":true},"doi":"10.21203/rs.3.rs-8354755/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-8354755/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":98437088,"identity":"5f0aeb1a-c49f-4d96-be4b-5b2bf6735c68","added_by":"auto","created_at":"2025-12-17 16:56:55","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":1056747,"visible":true,"origin":"","legend":"","description":"","filename":"ML1Revised.docx","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/6425c75a5bdaa0e9ff5e7f07.docx"},{"id":98436119,"identity":"8f9d04a9-6ca2-4ac6-9314-f214d6f4e4f0","added_by":"auto","created_at":"2025-12-17 16:54:57","extension":"json","order_by":1,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":342,"visible":true,"origin":"","legend":"","description":"","filename":"rs8354755.json","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/1e02902e0460ed6e518716be.json"},{"id":98436033,"identity":"8c3029e6-e174-42ff-af35-70690291b677","added_by":"auto","created_at":"2025-12-17 16:54:46","extension":"xml","order_by":2,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":189183,"visible":true,"origin":"","legend":"","description":"","filename":"rs83547550enriched.xml","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/cb0d6fd04e70c697ee998db7.xml"},{"id":98436722,"identity":"86909d58-7b60-43ad-85fd-dd8bd3bc8682","added_by":"auto","created_at":"2025-12-17 16:56:11","extension":"png","order_by":3,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":131561,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/cc6e0c20bc428daeb87d0fa2.png"},{"id":98437488,"identity":"88bb39da-f1f6-4d35-9398-7fb9853d5dc2","added_by":"auto","created_at":"2025-12-17 16:57:24","extension":"png","order_by":4,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":135877,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/e23d1a2c1bbd6e29c9351431.png"},{"id":98436126,"identity":"bc82cd31-8ad9-40f5-98e0-3b5bd76cd9bf","added_by":"auto","created_at":"2025-12-17 16:54:57","extension":"png","order_by":5,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":134969,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/deb5a80aabe9dca4cba30181.png"},{"id":98435739,"identity":"201c3b28-d689-4463-88d1-822b5f83fa0f","added_by":"auto","created_at":"2025-12-17 16:54:20","extension":"png","order_by":6,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":139231,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/d9fed7fa14ad090be8589ebb.png"},{"id":98314212,"identity":"8637dd44-e8b2-4fc4-8b61-9f39c6197913","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"png","order_by":7,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":134106,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/0b86fb2d4249f5f2730da5e5.png"},{"id":98436318,"identity":"d5c5d728-1c0c-404f-8ac6-2c9b56add161","added_by":"auto","created_at":"2025-12-17 16:55:22","extension":"png","order_by":8,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":137897,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/b7090d5577be76c10b063421.png"},{"id":98437173,"identity":"d3acc3a0-b91d-4978-a15e-ab90b42899ba","added_by":"auto","created_at":"2025-12-17 16:57:05","extension":"png","order_by":9,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":135987,"visible":true,"origin":"","legend":"","description":"","filename":"floatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/75e608315b97d16f1ec7638c.png"},{"id":98314222,"identity":"88517a90-eaab-4dec-b031-cc7bebec16d3","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"png","order_by":10,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45129,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/b1df51e7bd6372de4a85e863.png"},{"id":98435479,"identity":"8ab0e65b-9ed7-492d-b203-9d5e9bdf7c99","added_by":"auto","created_at":"2025-12-17 16:53:54","extension":"png","order_by":11,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":46071,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/fd8bdd92adaae7cfe6324602.png"},{"id":98314217,"identity":"bd9e30df-5ed9-4aeb-8967-e75e0d05f779","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"png","order_by":12,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45839,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/4cf88494e69c70743f3b3ea0.png"},{"id":98435477,"identity":"67998578-5b3a-4472-9515-e740d51c90b9","added_by":"auto","created_at":"2025-12-17 16:53:54","extension":"png","order_by":13,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":47368,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/da6a10ef30c42256ac92b929.png"},{"id":98314220,"identity":"b31fbccb-cc65-4bc9-a1c4-57afc1943dd1","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"png","order_by":14,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":45586,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/6cc8920088947f6e67d0dff3.png"},{"id":98435832,"identity":"73841862-ea8a-4d5e-aab8-9fafb75ad6f3","added_by":"auto","created_at":"2025-12-17 16:54:28","extension":"png","order_by":15,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":46820,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/35e26e0a74405835d81e0a29.png"},{"id":98436036,"identity":"8d82bff2-e75d-4314-848b-6d255b4e565e","added_by":"auto","created_at":"2025-12-17 16:54:46","extension":"png","order_by":16,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":46178,"visible":true,"origin":"","legend":"","description":"","filename":"Onlinefloatimage7.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/57af7ebfb2439044b1052b86.png"},{"id":98314226,"identity":"f9b7db6d-45ed-44d7-b393-06f8c904d22e","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"xml","order_by":17,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":187376,"visible":true,"origin":"","legend":"","description":"","filename":"rs83547550structuring.xml","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/cef7a67f2387fd9decc5dcc2.xml"},{"id":98314227,"identity":"350f4f9f-bf2b-4869-8056-49415115d19f","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"html","order_by":18,"title":"","display":"","copyAsset":false,"role":"acdc-reference","size":202906,"visible":true,"origin":"","legend":"","description":"","filename":"earlyproof.html","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/772734a28d7a2d66c40ec628.html"},{"id":98314201,"identity":"5bb22065-e623-4e12-8cc6-ed0a66d8c9d1","added_by":"auto","created_at":"2025-12-16 12:56:04","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":131561,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLinear Regression\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource - Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image1.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/6c7020ee20e7dcecfbd5bb0c.png"},{"id":98437166,"identity":"116aa902-e289-4dc5-8fe2-c77aaeb71bfb","added_by":"auto","created_at":"2025-12-17 16:57:05","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":135877,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRandom Forest\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource - Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image2.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/7369c394ac3b3e5df0e33bcb.png"},{"id":98314207,"identity":"ab32862a-be1a-48a2-8d2e-e4df85459be0","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":134969,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSupport Vector Regression\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource - Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image3.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/5751d6de3ce163e14e053881.png"},{"id":98314206,"identity":"23cec1d8-f23c-4611-9052-88151be7212f","added_by":"auto","created_at":"2025-12-16 12:56:04","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":139231,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eXGBoost\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource - Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image4.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/30ebcf48c48982679fca3012.png"},{"id":98435929,"identity":"1936775a-4848-4ac1-b2de-283ab65bfd06","added_by":"auto","created_at":"2025-12-17 16:54:38","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":134106,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eGBM\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource - Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image5.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/2d5dd1948e77123cb85a4439.png"},{"id":98437558,"identity":"001c2cbc-a6b4-4927-8b7d-9aef9ff63744","added_by":"auto","created_at":"2025-12-17 16:57:28","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":137897,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eLSTM\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource – Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image6.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/e503a87771530797306319ab.png"},{"id":98314210,"identity":"92b2710d-2325-4931-8b76-aafc1a220523","added_by":"auto","created_at":"2025-12-16 12:56:05","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":135987,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cem\u003e\u003cstrong\u003eCNN\u003c/strong\u003e\u003c/em\u003e\u003c/p\u003e\n\u003cp\u003e\u003cem\u003eSource – Author’s Work\u003c/em\u003e\u003c/p\u003e","description":"","filename":"image7.png","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/371c71b340bf2eab2d47b64b.png"},{"id":98774943,"identity":"8a6b30d7-21b2-4ba4-932c-c19cd3bca8ef","added_by":"auto","created_at":"2025-12-22 12:17:21","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1629552,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-8354755/v1/1d96cd57-63b1-4976-af59-ed28e6efabc4.pdf"}],"financialInterests":"The authors declare no competing interests.","formattedTitle":"\u003cp\u003e\u003cstrong\u003eForecasting Crude Oil Prices: Insights from Machine Learning Approaches\u003c/strong\u003e\u003c/p\u003e","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe role of oil prices in influencing macroeconomic variables is well-established in economic research (Barsky \u0026amp; Kilian \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2002\u003c/span\u003e; Baumeister \u0026amp; Kilian \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Hamilton \u0026amp; Herrera \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2004\u003c/span\u003e; Ahmed et al. \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2025\u003c/span\u003e; Siddiqui et al. \u003cspan citationid=\"CR72\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Crude oil prices hold a crucial position in oil-dependent economies, particularly those where exports and imports constitute a significant portion of the GDP (Gross Domestic Product) (Mohsin \u0026amp; Jamaani \u003cspan citationid=\"CR59\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). In such economies, variations in crude oil prices have critical financial effetcs. For instance, rising oil prices can lead to inflation and reduce an economy's competitiveness (Jahan-Parvar \u0026amp; Mohammadi, \u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). On contrary, declined oil prices can result in social unrest and political instability in oil-producing countries (Chen \u0026amp; Hsu \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2012\u003c/span\u003e; Gholamian et al. \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2005\u003c/span\u003e). Therefore, understanding and managing the impact of oil price volatility is essential for maintaining economic stability and growth in these regions.\u003c/p\u003e \u003cp\u003eSince the 1970s, oil prices have become increasingly unstable, impacting both economic growth and investor confidence (Xu et al. \u003cspan citationid=\"CR86\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). This instability is not only due to the inherent non-linearity of oil prices but also because of extreme and irregular events. While oil prices are fundamentally governed by the principles of demand and supply, external shocks such as geopolitical events and pandemic of COVID-19 have significant influences as well (Bernabe et al. \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2004\u003c/span\u003e; Hagen \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e1994\u003c/span\u003e; Stevens \u003cspan citationid=\"CR76\" class=\"CitationRef\"\u003e1995\u003c/span\u003e; Tang et al. \u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e2015\u003c/span\u003e, Atif et al., \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Ahmed and Kaur, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2025\u003c/span\u003e). Additionally, oil is a finite natural resource and a major contributor to environmental pollution and global warming. Worldwide efforts to shift toward sustainable energy sources, spurred by international pledges to cut carbon emissions and address climate change through mitigation and adaptation strategies, also plays a crucial role in shaping crude oil prices. This shift not only affects the demand for oil but also creates uncertainties in the market, as the world navigates towards a more sustainable energy future.\u003c/p\u003e \u003cp\u003eAs highlighted by Fang et al. (\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), forecasting oil prices is crucial for both producers and consumers. Accurate predictions can significantly contribute to rapid economic growth, higher production levels, and reduced production costs, fostering a stable macroeconomic environment (Yu et al. \u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). However, because of unpredictable nature of oil prices, accurate forecasting remains challenging. Brent crude prices are affected by a complex interplay of endogenous and exogenous factors. Endogenous factors include the inventory levels of oil, consumption rates, and supply dynamics. On the other hand, exogenous factors encompass geopolitical events, conflicts, and broader economic developments, all of which can drastically impact oil prices. This high volatility generates uncertainty and fear within both the public and private sectors. Additionally, the decisions made by OPEC have a substantial effect on determining the prices of brent crude in the global market (Kaufmann et al. \u003cspan citationid=\"CR49\" class=\"CitationRef\"\u003e2004\u003c/span\u003e). Over the past decade, energy-based commodities have experienced exponential growth, evolving into an indispensable asset class for investment purposes. This growth underscores the importance of understanding and forecasting oil prices to mitigate risks and capitalize on investment opportunities.\u003c/p\u003e \u003cp\u003eIn forecasting, researchers have historically relied on the ARIMA model. More recently, studies have begun to employ Artificial Neural Networks (ANN) for this purpose. Traditional methods rely on mostly on one type of model, time series models or regression models. The complexity of the oil prices is not captured in the findings. Further, most of the studies have applied the single machine learning models to forecast the oil prices. Although research utilizing ML techniques such as ANN and SVM., are able to forecast the oil prices, under the complexity and uncertainty, they lack in the analysis of the comparison among various ML models. It is in this context this study forecast the oil prices based on the autoregressive approach, using Linear Regression, SVM, Random Forest, Extreme Gradient Boosting (XGBoost), and GBM based ML models. This research aims to accurately forecast Brent crude oil prices by evaluating the predictive performance of various machine learning models. To ensure a thorough assessment, the study utilizes six key metrics: MSE, RMSE, MAE, MAPE, R-Squared, and the Diebold-Mariano (DM) statistic.\u003c/p\u003e \u003cp\u003eTaking into account the necessity and complexity of the oil prices forecasting, this focus upon the evaluating best machine learning model in terms of prediction accuracy, in forecasting of the oil prices. Further, in this paper section 2 provides a brief review of past studies, followed by the Methodology. The section 4 discusses the results and findings followed by the conclusion and implications.\u003c/p\u003e"},{"header":"2. Literature Review","content":"\u003cp\u003eThe literature on the machine learning based forecast has grown many folds in the recent. Application of ML in forecasting the time series are applied in predicting the time series. Many researchers have employed individual ML techniques to forecast oil price trends. Ahmed et al. (\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) provides comprehensive studies exploring the use of AI and ML in the field of finance. Study found the application of AI and ML in financial sectors, including forecasting bankruptcy, predicting stock values, managing investment portfolios, and detecting money laundering activities., to be higher post 2015, and continues to rise. However, the AI and ML based research is localized in United States, China, and the United Kingdom. Among the notable earliest studies predicting the oil prices are (Gabralla \u0026amp; Abraham \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Khashman \u0026amp; Nwulu \u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e2011\u003c/span\u003e; Panella et al. \u003cspan citationid=\"CR69\" class=\"CitationRef\"\u003e2012\u003c/span\u003e; Shao et al. \u003cspan citationid=\"CR75\" class=\"CitationRef\"\u003e2014\u003c/span\u003e; Tebyanian \u0026amp; Hedayati \u003cspan citationid=\"CR80\" class=\"CitationRef\"\u003e2014\u003c/span\u003e,Ahmed, \u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eAmong the recent studies, Naeem et al. (\u003cspan citationid=\"CR63\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) applied the ARIMA, SVM and LSTM to check the prediction accuracy of the oil prices. Diebold \u0026ndash; Mariano is used to compare the robustness and forecasting methods. The empirical findings suggest the higher accuracy of the hybrid models over the alternative approaches. Further, Liang et al. (\u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) developed reinforcement learning based algorithm for forecasting the oil prices on three major commodity exchanges. Authors designed the mechanism based on stochastic process for generalization and accuracy, along with the learning efficiency. The study found the algorithm to be better predictor in accuracy, of the three major crude oil price benchmarks. Further, the algorithm can be extended to predict the other variables related to the natural resources.\u003c/p\u003e \u003cp\u003eNumerous studies applied the machine learning models, along with the traditional methods such as ARIMA, GARCH etc. ( Cheng et al. \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Nanthiya et al. \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Tami \u0026amp; Owda \u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Wei et al. \u003cspan citationid=\"CR84\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Weng et al. \u003cspan citationid=\"CR85\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Yu et al. \u003cspan citationid=\"CR87\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Aldabagh et al. (\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) proposed the model based on the deep learning and traditional ARIMA, to forecast the oil prices. The author used the CNN with LSTM. Further, the CNN-LSTM model is compared with LSTM, CNN, SVM, and the ARIMA model. The suggested model demonstrates superior accuracy compared to the current model for both single-step and multiple-step forecasts. Das and Das (\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) employed a forecasting model that incorporated text-based sentiment related to inflation as additional inputs into an artificial neural network, which outperformed all other models in terms of forecast accuracy. Jha et al. (\u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) applied the SVM to estimate the spot prices by utilizing the multiple variables. The proposed model is evaluated against the four benchmark models (Ordinary Least Squared, GARCH, and ANN), at different steps. Author found the LASSO model to be better predictor. Zhang and Zheng (2023) explored the oil price prediction by leveraging traditional models of time series, and the machine learning models. The Ordinary Least Square is used for the variable selection, and machine leaning models for the best overall performance. LSTM model is found to be the best in overall performance. Cheng et al. (\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) built framework to estimate the variability of brent crude prices by taking into account the structural breaks (Extreme Events). The oil prices are decomposed to find the extreme events or structural break. Further, the ARIMA and SVM are combined to predict the oil prices. The empirical findings found that combined models perform better than other models.\u003c/p\u003e \u003cp\u003eTissaoui et al. (\u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) applied the XGBoost ML technique to estimate the oil prices, and compared the SVM, and the ARIMAX models to check the relationship with the forecasters. The empirical findings of the study indicate the dominance of ML models in forecasting the oil prices. Kakade et al. (\u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) proposed model to estimate the crude oil prices based on the hybrid ensemble learning approach. It includes the ARIMA, PCA, and the LSTM model in ensemble approach, and is compared with the LSTM. Findings indicate that oil ensemble LSTM is more accurate than the traditional LSTM. Liu et al. (\u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) forecasted the crude oil prices by employing the ARIMA, LSTM. Neural Networks based on extreme learning, and back propagation neural network. The proposed model\u0026rsquo;s accuracy is satisfactory, even during the COVID-19 pandemic.\u003c/p\u003e \u003cp\u003eThe oil prices are influenced by various factors, particularly in an oil dependent economy. In this regard, Albahooth (\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) predicted the oil prices in the Saudi Arabia, by employing the interest rates, currency rates, and stock markets from financial variables, and energy consumption, inflation rate, and GDP growth from macroeconomic variables. The study used various ML technique such as linear regression and SVM, and models are evaluated by examining the RMSE, MSE, and R-square. The study explores the significance of various financial and macro-economic variables. Similarly, Guliyev \u0026amp; Mustafayev (\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2022\u003c/span\u003e) predicted the West Texas Intermediate (WTI) oil prices using United States financial and macro-economic factors. The researchers employed a range of ML techniques, including Decision Tree, Random Forest, Logistic Regression, AdaBoost, and XGBoost algorithms. Authors found that Random Forest model and XGBoost model to be outperforming the traditional models. Moreover, the study used Shapley Additive exPlanations values for model evaluation and interpretation. Similarly, Jahandoost et al. (\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) proposed hybrid architecture for DNNs, and used 39 features for the prediction of the oil price. The proposed architecture enhances the accuracy of previous models. (Sen et al. \u003cspan citationid=\"CR74\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) proposed LSTM to estimate the oil prices by using the financial variables as predictor.\u003c/p\u003e \u003cp\u003eIn machine learning, ensemble learning is a technique that integrates predictions from multiple models to enhance the overall model's precision. This method combines various models' outputs to achieve improved accuracy in predictions. The objective is to reduce the errors in the predictions from other models. In ensemble approach, Hasan et al. (\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) combined ML and ensemble learning to predict the crude oil prices. Based on the plotted graph of actual vs predicted values, Ada Boost algorithm predicts oil prices with high accuracy. Further, the MSE, RMSE, MAE, MAPE, R^2, and function variance scores are used in validating the results. Tami \u0026amp; Owda (\u003cspan citationid=\"CR78\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) developed the novel LSTM model to estimate the commodity prices. The model used features of moving average, price volatility, past prices. The finding based on the RMSE, MAPE, and R-squared conclude that proposed model is better than the traditional model. Additionally, the LSTM architecture significantly reduces computation time, allowing training to be completed in minutes rather than hours. Sajid et al. (\u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) forecasted the brent oil prices by using the ML and the ensemble learning methodology. Study applies the Light Gradient Bosting, Random Forest ensemble ML algorithm, Lasso regression, and Decision tree ML algorithm. Out of the applied ML algorithms, Light GBM is found to be most accurate. Kim \u0026amp; Jang (\u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) employed hybrid model to improve estimated abilities of the other hybrid models. The author used CNN, and RNN. Based on the evaluation metrics, the findings indicate the proposed model to be the better predictor as compared to the GRU and LSTM based models. Cheng et al. (\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) studied efficacy of combined estimate, and ML models in prediction of brent oil price volatility. Author found the ML to be more promising than the forecasting combination methods. Similarly, Hao (\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2024\u003c/span\u003e) proposed novel ML based model in estimation of the oil prices. The author combined the transformer algorithms and LSTM in forecasting the brent crude oil prices. The experimental results confirm the effectiveness of the examined method.\u003c/p\u003e \u003cp\u003eThe review of the literature reveals a significant surge in the application of ML for time series forecasting, particularly in oil price prediction. While studies have explored various ML techniques, including ARIMA, SVM, LSTM, and ensemble methods, several limitations persist. Many studies focus on individual ML models or combinations of traditional econometric methods with singular ML approaches, often neglecting a comprehensive comparative analysis of diverse ML algorithms. Furthermore, the geographical focus of research is skewed towards the United States, China, and the United Kingdom, potentially limiting the global applicability of findings. Despite the increasing complexity of hybrid models, such as CNN-LSTM and ensemble methods, there remains a gap in understanding the relative performance of a wide array of ML models specifically tailored for crude oil price forecasting. Additionally, while some studies incorporate external financial and macroeconomic variables, a systematic investigation into the optimal selection and integration of these features across various ML architectures is lacking. Therefore, a research gap exists in providing new insights into forecasting crude oil prices by conducting a thorough, comparative analysis of a broad spectrum of machine learning approaches, and systematically investigating the impact of feature selection and model combination, to identify the most effective and robust methods for this critical forecasting task.\u003c/p\u003e"},{"header":"3. Methodology and Data","content":"\u003cp\u003eTo capture the inherent non-linearity and non-stationarity of crude oil prices, as evidenced by data outliers, this study forgoes traditional pre-processing. Data, obtained from the IMF (IMF, 2024), utilizes lagged values (\u003cem\u003elag\u0026thinsp;=\u0026thinsp;1\u003c/em\u003e) as predictors. Following established practices in the literature (Kandil et al., \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Nguyen et al., \u003cspan citationid=\"CR66\" class=\"CitationRef\"\u003e2023\u003c/span\u003e), an 80\u0026thinsp;\u0026minus;\u0026thinsp;20 train-test split is employed.\u003c/p\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Machine Learning Models\u003c/h2\u003e \u003cp\u003eThere are various type of ML models used in the prediction. The most common and popular among them are Linear Regression, SVM, Random Forest, Xgboost, and Gradient Boosting Machine (Adeniyi et al. \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Adenusi et al. \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Albahooth \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Bakshi et al. \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Boussatta et al. \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Ding et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Drucker et al. \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e1996\u003c/span\u003e; Gao et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Gong \u0026amp; Zhang \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Guliyev \u0026amp; Mustafayev \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Jahanshahi et al. \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Lakshminarayanan \u0026amp; McCrae 2019; Li et al. \u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Manjula \u0026amp; Karthikeyan \u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Murugesan et al. \u003cspan citationid=\"CR62\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Nayak et al. \u003cspan citationid=\"CR65\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Okechukwu Ajakwe et al. \u003cspan citationid=\"CR67\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Sajid et al. \u003cspan citationid=\"CR73\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Sulaiman et al. \u003cspan citationid=\"CR77\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Tang et al. \u003cspan citationid=\"CR79\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Tissaoui et al. \u003cspan citationid=\"CR81\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Vapnik \u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e1998\u003c/span\u003e; Zaidi \u0026amp; Oussalah \u003cspan citationid=\"CR88\" class=\"CitationRef\"\u003e2018\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eLinear regression seeks the straight line that most closely aligns with this overall trend. The data is trained in linear regression, and this trained data helps the model discover the ideal coefficients for the straight-line equation that best captures the trend in the data. The equation of Linear Regression can be written as Hope (\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:y=C+\\beta\\:x+ϵ$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eEquation (\u003cspan refid=\"Equ1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) represents how the target variable \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y\\)\u003c/span\u003e\u003c/span\u003e is predicted based on the input \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e. Here, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:C\\)\u003c/span\u003e\u003c/span\u003e is the intercept, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\beta\\:\\)\u003c/span\u003e\u003c/span\u003e is the coefficient showing the impact of \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e on \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:ϵ\\)\u003c/span\u003e\u003c/span\u003e is the error term accounting for unexplained variation. The model learns the best values for \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:C\\)\u003c/span\u003e\u003c/span\u003e and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\beta\\:\\)\u003c/span\u003e\u003c/span\u003e to minimize prediction error and accurately estimate \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:y\\)\u003c/span\u003e\u003c/span\u003e.\u003c/p\u003e \u003cp\u003eThe SVM model, though applied mostly in classification problem, is commonly used model in forecasting. In regression, SVM introduces an ε-insensitive loss function which creates a hyperplane which minimize the difference among the estimated values of the training sample, and observed value of the response. In regression, a hyperplane along with ε creates an ε-insensitive tube for creating generalization bunds for the regression. The optimization is done by minimizing the ε-insensitive tube to as flat as possible, and simultaneously retaining the maximum training set. Common SVR Equation is -\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:F\\left(x\\right){=W}^{T}x+b$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:F\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e is product of vector weight \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:W\\)\u003c/span\u003e\u003c/span\u003e and input vector \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e. The tolerance limit is determined by the ε, also known as loss function. The SVR can be written as Vapnik (\u003cspan citationid=\"CR82\" class=\"CitationRef\"\u003e1998\u003c/span\u003e).\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\:mi{n}_{w,b,{{\\xi\\:}}_{i},{{\\xi\\:}}_{i}^{*}}\\left(\\frac{1}{2}|w{|}^{2}+C{\\sum\\:}_{i=1}^{n}\\left({{\\xi\\:}}_{i}+{{\\xi\\:}}_{i}^{*}\\right)\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eSubject to \u0026ndash;\u003cdiv id=\"Equa\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equa\" name=\"EquationSource\"\u003e\n$$\\:{y}_{i}-⟨w,{x}_{i}⟩-b\\le\\:\\epsilon\\:+{\\xi\\:}_{i}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equb\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equb\" name=\"EquationSource\"\u003e\n$$\\:⟨w,{x}_{i}⟩+b-{y}_{i}\\le\\:\\epsilon\\:+{\\xi\\:}_{i}^{*}$$\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equc\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equc\" name=\"EquationSource\"\u003e\n$$\\:-{\\xi\\:}_{i},\\hspace{0.25em}{\\xi\\:}_{i}^{*}\\ge\\:0,\\hspace{1em}i=1,\\dots\\:,N$$\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e \u003cspan class=\"InlineEquation\"\u003e \u003cspan class=\"mathinline\"\u003e\\(\\:\\frac{1}{2}{\\left|\\left|W\\right|\\right|}^{2}\\:\\)\u003c/span\u003e \u003c/span\u003e is the regularization term that penalizes complex models, by controlling the magnitude of the weighted vector. Further, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:C*\\sum\\:{(\\xi\\:}_{n}+{{\\xi\\:}_{n}}^{*})\\)\u003c/span\u003e\u003c/span\u003e is empirical loss function, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:C\\)\u003c/span\u003e\u003c/span\u003e determines the weight of the errors. A large \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:C\\)\u003c/span\u003e\u003c/span\u003e gives higher weight to the minimize the prediction errors, whereas low \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:C\\)\u003c/span\u003e\u003c/span\u003e value give higher weight to the minimize the flatness. The slack variables \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\xi\\:}_{n}\\)\u003c/span\u003e\u003c/span\u003e, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{{\\xi\\:}_{n}}^{*}\\)\u003c/span\u003e\u003c/span\u003e determine the points which can be tolerated outside the ε-insensitive tube.\u003c/p\u003e \u003cp\u003eAs per In a Random Forest regression model, the final prediction for a given input \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:x\\)\u003c/span\u003e\u003c/span\u003e is obtained by averaging the predictions of all individual decision trees in the ensemble. If \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{T}_{b}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e denotes the prediction from the \u003cem\u003eb\u003c/em\u003e\u003csup\u003e\u003cem\u003eth\u003c/em\u003e\u003c/sup\u003e tree and there are B such trees in total, then the overall Random Forest prediction \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{f}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e is given by the equation:\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\widehat{f}\\left(x\\right)=\\frac{1}{B}{\\sum\\:}_{b=1}^{B}{T}_{b}\\left(x\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eThis averaging process helps reduce variance and improves the model\u0026rsquo;s generalization ability, making Random Forests robust and effective for regression task (Breiman, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2001\u003c/span\u003e; Dudek, \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2015\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eXGboost combines the points which were flawed in decision tree. XGBoost involves summation of multiple trees as different stages in a boosting process. The objective function measures the difference among the estimated and the observed values. The function is -\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:\\:L({Y}_{i},\\widehat{f}\\left(x\\right))$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eloss function for the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{i}^{th}\\)\u003c/span\u003e\u003c/span\u003e data point, where \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Y}_{i}\\)\u003c/span\u003e\u003c/span\u003e is the observed value, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{f}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e is the forecasted value. In gradient boosting framework, focus is on errors, and each stage tries to better predict the previous stage.\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$\\:\\widehat{f}\\left(x\\right)=\\:\\sum\\:{F}_{t}\\left(x\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\widehat{f}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e is the final prediction, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{F}_{t}\\left(x\\right)\\)\u003c/span\u003e\u003c/span\u003e is final prediction from the \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{t}^{th}\\:\\)\u003c/span\u003e\u003c/span\u003e tree in ensemble.\u003c/p\u003e \u003cp\u003eGradient Boosting Machine, abbreviated as GBM, is a type of ensemble learning technique. This method works by combining predictions from several less robust models to produce a single, more powerful model. These weaker models are typically decision trees.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Common Regression Evaluation Metrics\u003c/h2\u003e \u003cp\u003eEvaluating models goes beyond simply checking their accuracy; it's about assessing their reliability in making predictions. The accuracy and reliability are crucial in making forecast. There are various metrics which evaluate the models, including MSE, RMSE, and R-squared. Various studies have used the MSE, RMSE, MAE, MAPE and R-Square for the evaluation of forecasting accuracy of the machine learning models (Alqahtani \u0026amp; Abdelhafez \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Aziz et al. \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Chen \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Kandil et al. \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Muganda \u0026amp; Kasamani \u003cspan citationid=\"CR60\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Nanthiya et al. \u003cspan citationid=\"CR64\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Prakash \u0026amp; Singh \u003cspan citationid=\"CR70\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Wang et al. \u003cspan citationid=\"CR83\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Chicco et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn machine learning, MSE and RMSE gauge how accurate a model's predictions are by capturing the average difference among forecasted and observed values. MSE acts like a trainer during the learning process, guiding the model to minimize this error. It excels at penalizing the large mistakes because it penalizes huge differences more heavily. However, MSE is measured in squared units, different from the target variable's unit, making interpretation a bit complex. To address this, RMSE simply takes the root square of MSE, presenting the error in the same units as the target variable for easier understanding. Both metrics are popular for their focus on large errors but are also sensitive to outliers, so keep the target variable's unit in mind when evaluating the error values. The R squared focuses on the part of variance explained in the dependent variable by the model. It essentially quantifies how well the model fits the data, indicating the degree to which it can explain the target variable's variance. R-squared is an easy-to-understand measure of model fit, and varies from zero to one, where one signifies a perfect fit. While it's scale-independent, adding more predictors can artificially inflate R-squared. Importantly, R-squared about the model's predictive power; it simply helps us understand the proportion of variance explained by the factors considered in the model (Kadri and Abdennbi, \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Kumar et al. \u003cspan citationid=\"CR53\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). MAE and MAPE are common metrics used to evaluate the accuracy of regression models. MAE measures the average absolute difference between predicted and actual values, providing a clear and interpretable indication of prediction error in the same units as the data. In contrast, MAPE expresses this error as a percentage of the actual values, making it useful for comparing performance across different datasets or scales. While MAE is unaffected by the scale of the data, MAPE can be sensitive to very small actual values, which may inflate the error percentage.\u003c/p\u003e \u003cp\u003eThe MSE, can be written as-\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$\\:MSE=\\:\\frac{1}{m}\\sum\\:_{m}^{i}({{X}_{i}-{Y}_{i})}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{X}_{i}\\:\\)\u003c/span\u003e\u003c/span\u003eis predicted value, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{Y}_{i}\\)\u003c/span\u003e\u003c/span\u003e are the actual values. Whereas, the formula for\u003c/p\u003e \u003cp\u003eRMSE can be written as-\u003c/p\u003e \u003cp\u003e \u003cem\u003eRMSE =\u003c/em\u003e \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\sqrt{\\frac{1}{m}\\sum\\:_{m}^{i}({{X}_{i}-{Y}_{i})}^{2}}\\:\\:\\:\\:\\:\\)\u003c/span\u003e\u003c/span\u003e (8)\u003c/p\u003e \u003cp\u003eand R-Square formula is\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$\\:{R}^{2}=1-\\frac{{\\sum\\:}_{i=1}^{m}{\\left({Y}_{i}-{X}_{i}\\right)}^{2}}{{\\sum\\:}_{i=1}^{m}{\\left({Y}_{i}-\\stackrel{-}{{Y}_{i}}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003eWhere \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{{Y}_{i}}\\)\u003c/span\u003e\u003c/span\u003e is the mean value of the observed values. The MAE and MAPE can be written as.\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$$\\:\\text{}\\text{MAE}=\\frac{1}{m}{\\sum\\:}_{i=1}^{m}\\left|{X}_{i}-{Y}_{i}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ10\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ10\" name=\"EquationSource\"\u003e\n$$\\:\\text{MAPE}=\\frac{1}{m}{\\sum\\:}_{i=1}^{m}\\left|\\frac{{Y}_{i}-{X}_{i}}{{Y}_{i}}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e11\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003e \u003col\u003e \u003cspan\u003e \u003cli\u003e \u003cp\u003e \u003cb\u003eHyperparameter Tuning Processes\u003c/b\u003e \u003c/p\u003e \u003c/li\u003e \u003c/span\u003e \u003c/ol\u003e \u003c/p\u003e \u003cp\u003eThe hyperparameter tuning is crucial in improving forecast accuracy in machine learning and deep learning, which is to select optimum parameters based on specific model. There are various approaches of hyperparameter tuning, such as Grid Search, Bayesian Optimisation, and Random search. The approach depends upon the model selected, and as per the literature, Grid Search is highlighted for the linear regression, whereas for Random Forest, Random Search is preferred. Bayesian optimisation is recommended in SVR, XGBoost, and Gradient Boosting Model. (Hoque and Aljamaan, \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Rimal et al., \u003cspan citationid=\"CR71\" class=\"CitationRef\"\u003e2024\u003c/span\u003e; Dhilsath and Samuel, \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Singh et al., 2023; Zhu et al.,2022).\u003c/p\u003e \u003cp\u003eFor the Linear Regression model, hyperparameter tuning was conducted using Grid Search with Cross-Validation (GridSearchCV) to identify the optimal value of the regularization parameter, alpha (λ). Given that alpha controls the L2 penalty, which discourages large coefficients and mitigates multicollinearity, its careful selection is critical. Six different alpha values\u0026mdash;0.1, 1, 10, 100, 500, and 1000\u0026mdash;were evaluated using 5-fold cross-validation, where model performance was assessed using negative MSE. The best alpha value was found to be 1000, indicating that a higher degree of regularization was beneficial. This suggests the presence of noise or multicollinearity in the dataset and implies that stronger penalization of coefficients led to better generalization. By effectively shrinking the coefficients, the model avoided overfitting and demonstrated robust performance across training and test sets, striking an optimal bias-variance trade-off. In contrast, the Random Forest model employed RandomizedSearchCV for hyperparameter tuning, a more computationally efficient alternative to grid search. It explored 30 random combinations across a specified hyperparameter space, including the number of estimators, maximum tree depth, minimum samples required for node splits and leaf nodes, and bootstrap sampling. Using 5-fold cross-validation and negative MSE as the scoring metric, the optimal configuration was identified as n_estimators\u0026thinsp;=\u0026thinsp;200, max_depth\u0026thinsp;=\u0026thinsp;10, and min_samples_split\u0026thinsp;=\u0026thinsp;10. These values suggest a deliberate attempt to prevent overfitting by limiting model complexity while maintaining high predictive accuracy. The use of parallel computation (n_jobs=-1) enhanced efficiency. This strategy not only balanced model accuracy and generalization but also demonstrated computational prudence.\u003c/p\u003e \u003cp\u003eThe SVR model utilized Optuna, a state-of-the-art optimization framework, to fine-tune three critical hyperparameters: C, epsilon, and gamma. After running 30 optimization trials, the optimal configuration\u0026mdash;C\u0026thinsp;=\u0026thinsp;67.8443, epsilon\u0026thinsp;=\u0026thinsp;0.7399, and gamma\u0026thinsp;=\u0026thinsp;0.0009586\u0026mdash;was found to offer a well-balanced fit. A high C indicates that the model imposed strong penalties on errors, making it more precise but potentially prone to overfitting. However, the moderately high epsilon introduced a tolerance band within which errors were not penalized, adding robustness. A small gamma extended the reach of the radial basis kernel, smoothing the decision boundary and improving generalization. This combination demonstrated that fine-tuned SVR could effectively balance complexity and flexibility to reduce prediction error. Similarly, the XGBoost Regressor (XGBRegressor) underwent tuning through Optuna over 30 trials. The search space spanned several impactful hyperparameters, including n_estimators, max_depth, learning_rate, subsample, colsample_bytree, gamma, reg_alpha, and reg_lambda. The best-performing configuration\u0026mdash;colsample_bytree\u0026thinsp;=\u0026thinsp;0.7432, gamma\u0026thinsp;=\u0026thinsp;0.3972, learning_rate\u0026thinsp;=\u0026thinsp;0.0126, max_depth\u0026thinsp;=\u0026thinsp;4, and n_estimators\u0026thinsp;=\u0026thinsp;193\u0026mdash;indicates a highly conservative learning strategy. The relatively shallow tree depth and low learning rate ensured that the model trained gradually, avoiding abrupt jumps in optimization, while the moderate gamma discouraged overly aggressive tree splits. The balance of these hyperparameters suggests an emphasis on slow, regularized learning that ensures model robustness and strong generalization on unseen data.\u003c/p\u003e \u003cp\u003eThe GBM model also benefitted from Optuna-based hyperparameter optimization. The tuning process explored key parameters such as the learning rate, minimum samples per leaf, and the minimum number of samples required for a split. After 30 trials, the optimal configuration included a learning rate of 0.23, a minimum of 5 samples per leaf, and 8 samples required for node splitting, along with 72 boosting rounds. These settings reflect a model designed to balance performance and complexity. The relatively higher learning rate compared to XGBoost implies faster learning, while constraints on sample size at leaves and splits help regulate model complexity, reducing the risk of overfitting. The hyperparameter tuning processes across these models demonstrate a consistent theme of balancing model complexity and predictive accuracy. Each approach\u0026mdash;be it GridSearchCV, RandomizedSearchCV, or Optuna\u0026mdash;was strategically selected based on the model architecture and computational constraints. The success of these tuning strategies reinforces the notion that model performance is not merely a function of algorithm choice, but critically dependent on the thoughtful calibration of hyperparameters tailored to data characteristics and forecasting objectives.\u003c/p\u003e \u003c/div\u003e"},{"header":"4. Results and Discussion","content":"\u003cp\u003eThe results presented in Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e indicate a strong correlation between lagged oil prices and the estimator, evidenced by a high coefficient and statistically significant p-value. This supports the inclusion of lagged oil prices as a reliable predictor within the linear regression-based machine learning model.\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003ctable id=\"Tab1\" border=\"1\"\u003e\u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003e– Predictor Coefficients\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eEstimate\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eStd. Error\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eT-value\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003ePr(\u0026gt;|t|)\u003c/p\u003e\n \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eIntercept\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003e0.141933\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.085230\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.665\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.096\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eLagged_Oil\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003e0.997636\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.001217\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e819.579\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eMultiple R -squared\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003e0.9967\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eAdjusted R-Squared\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003e0.9967\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eF Statistics\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003e6.717e + 05\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eP value\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lt; .001\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"5\"\u003e\u003cem\u003eSource – Author’s Work\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003eThe Linear Regression model exhibits an excellent fit to the oil price data, as evidenced by both the robust evaluation metrics and the visual representation in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e, where the predicted and actual prices nearly overlap during the training period. Similarly, the Random Forest model demonstrates strong predictive capability, with Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e showing a close alignment between predicted and actual values in the training dataset. The SVR model also performs commendably, as shown in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e, effectively capturing underlying patterns in both training and test datasets. The XGBoost model, illustrated in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e, delivers performance comparable to linear regression, with a strong alignment between predicted and actual oil prices across both datasets. Likewise, the GBM model, depicted in Fig.\u0026nbsp;\u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e, shows a high degree of accuracy, as indicated by the close match between predicted and actual values, further confirming its effectiveness in modeling oil price dynamics.\u003c/p\u003e\n\u003cp\u003eThe Linear Regression model demonstrates outstanding performance, with exceptionally high R-squared values nearing 0.995 for both training and test sets, alongside consistently low error metrics (MSE, RMSE, MAE, MAPE), indicating strong predictive accuracy and excellent generalization to unseen data, as the predicted test values closely follow the actual ones without signs of overfitting. The Random Forest model also shows impressive performance, with a very low training MSE of 0.7216 and R-squared of 0.9987, though its test performance shows slightly higher error metrics (MSE: 3.3101, RMSE: 1.8194, MAE: 1.3601, MAPE: 2.0399), suggesting a modest decline in generalization while still maintaining a high R-squared of 0.9939. The SVR model offers reliable predictions with R-squared values around 0.99 for both datasets and relatively low error metrics, showing consistent performance and robust generalization, though it performs slightly below the Linear Regression model on this dataset. Similarly, the XGBoost model achieves high predictive accuracy with R-squared values near 0.99 and low error metrics across training and test datasets, confirming strong generalization and a performance level comparable to Linear Regression and better than SVR. The GBM model also yields high R-squared values (≈ 0.99) and low error metrics across both datasets, reinforcing its predictive strength and generalization ability, placing its performance on par with XGBoost and Linear Regression, and ahead of SVR for this dataset.\u003c/p\u003e\n\u003cp\u003eIn essence, all five models demonstrate strong predictive capabilities for oil price forecasting, with notable differences in generalization and robustness. The Linear Regression model stands out as a highly accurate and reliable predictor, exhibiting excellent generalization and strong statistical significance, as evidenced by consistently large negative Diebold-Mariano (DM) statistics of -18.548678, indicating a substantial improvement over naive forecasts. The Random Forest model also delivers accurate predictions, but the noticeable disparity between training and test performance suggests potential overfitting, warranting further refinement through regularization, feature engineering, and cross-validation, despite its statistically significant DM value of -18.536257. The SVR model offers solid performance with good accuracy and generalization, supported by highly negative DM statistics of -18.552112, confirming its superiority over naive models. Similarly, the XGBoost model proves to be a robust and accurate predictor, with strong generalization and consistent statistical significance reflected in its DM value of -18.548297. The GBM model mirrors the performance of XGBoost, showing strong predictive accuracy, solid generalization, and a statistically significant improvement over naive forecasts, as indicated by its DM statistic of -18.545874.\u003c/p\u003e\n\u003cp\u003eIn conclusion, all five models—Linear Regression, Random Forest, SVR, XGBoost, and GBM—demonstrate strong potential for forecasting oil prices, each exhibiting solid predictive accuracy and statistical significance over naive models. However, based on the evaluation metrics, visual alignment, generalization capability, and consistently strong Diebold-Mariano statistics, the Linear Regression model emerges as the best overall performer. It combines simplicity with high accuracy, excellent generalization, and minimal signs of overfitting, making it a robust and reliable choice for oil price prediction within this dataset.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"char\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003ctable id=\"Tab2\" border=\"1\"\u003e\u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003e–ML Model Performance Comparison\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eModel\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eDataset\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eMSE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eMAE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eMAPE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eR²\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\"\u003e\n \u003cp\u003eDM Statistics\u003c/p\u003e\n \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eLinear Regression\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.471180\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.571999\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.064543\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.673229\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.995626\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e-18.548678\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.409677\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.552313\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.086515\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.662010\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.995546\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.721557\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.849445\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.605434\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.937417\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.998723\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e-18.536257\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e3.310095\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.819367\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.360060\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.039885\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.993881\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eSVR\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e5.575188\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.361184\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.276119\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.470044\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.990131\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e-18.552112\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e5.573400\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.360805\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.293814\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.298823\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.989698\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.213671\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.487841\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.043201\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.650890\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.996081\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e-18.548297\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.815016\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.677801\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.186548\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.801803\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.994796\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eGBM\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.830616\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.353002\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.972099\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.488672\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.996760\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e-18.545874\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e2.592087\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.609996\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.144261\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e1.739305\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\"\u003e\n \u003cp\u003e0.995209\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003ctfoot\u003e\u003ctr\u003e\u003ctd colspan=\"8\"\u003e\u003cem\u003eSource – Author’s Work\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e4.1 \u003cstrong\u003eLSTM and CNN Based Results\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Optuna-tuned LSTM and CNN models were optimized with carefully selected hyperparameters to enhance their performance in forecasting oil prices from univariate time series data. The LSTM model employed 50 memory units, striking a balance between learning temporal dependencies and avoiding overfitting, supported by a dropout rate of 0.2 to introduce moderate regularization. A learning rate of 0.001 ensured steady convergence, while a batch size of 32 balanced training stability and generalization. In contrast, the CNN model, optimized for a lag-1 setup, utilized 92 filters in its Conv1D layer to capture fine-grained patterns from recent data, with a kernel size of 1 aligning perfectly with the one-step lag structure. A minimal dropout rate of 0.043 indicated low overfitting risk, while a small learning rate of 0.00027 facilitated cautious weight updates, crucial for handling noisy financial data. Additionally, the batch size of 9 introduced beneficial stochasticity during training, helping the model generalize well. Together, these hyperparameter configurations reflect robust and well-regularized architectures, capable of effectively modeling the complex, nonlinear behavior of oil price movements with precision and stability.\u003c/p\u003e\n\u003cp\u003eThe LSTM-based graph illustrates the model's ability to capture long-term dependencies in oil price movements. During the training phase (2013–2022), the predicted values (in cyan) closely align with the actual prices (in blue), indicating that the model has effectively learned historical patterns and underlying seasonality. While minor deviations are visible, they reflect typical behavior in financial time series due to noise, and overall, the model demonstrates strong in-sample learning. In the test period (2023–2024), the predicted values (red dashed line) track the actual oil prices (black line) reasonably well, though there is a slight lag in response to abrupt fluctuations—an expected characteristic of memory-based models like LSTM. This suggests that while the model generalizes effectively, it may underperform when faced with rapid, short-term volatility in oil prices.\u003c/p\u003e\n\u003cp\u003eIn contrast, the CNN-based graph reflects a model that is more agile and responsive to local fluctuations in the data. The training predictions (cyan) match the actual prices (blue) closely, albeit with more jaggedness compared to LSTM, which indicates the CNN's sensitivity to short-term variations. This responsiveness becomes even more evident during the test period (2023–2024), where the predicted values (red dashed) closely follow the actual prices (black), demonstrating the CNN’s strength in adapting to recent and abrupt changes in the oil price series. The tighter alignment of predictions in the test phase suggests that CNN may slightly outperform LSTM in short-horizon forecasting, particularly under conditions of high volatility. Overall, while LSTM provides stable and trend-following forecasts, CNN offers more reactive and precise short-term predictions, making each model uniquely suited to different forecasting objectives within financial time series modeling.\u003c/p\u003e\n\u003cp\u003e\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eBased on the evaluation metrics and model configurations, a comparative analysis of the Optuna-tuned LSTM and CNN models for oil price forecasting reveals several important insights regarding their performance and suitability for time series prediction.\u003c/p\u003e\n\u003cp\u003eThe LSTM model achieved superior performance on the training dataset, with an MSE of 2.2788, RMSE of 1.5096, and R² of 0.9958, indicating an excellent fit to the training data. It also produced relatively low training errors in terms of MAE (1.1783) and MAPE (2%). On the test dataset, while there was a slight drop in performance, the model still performed well, with a test MSE of 7.0434, RMSE of 2.6539, and R² of 0.9634. The MAE and MAPE on the test set were 2.0139 and 2.17%, respectively. Notably, the Diebold-Mariano (DM) test statistic of 3.8630 (p-value = 0.0001) suggests that the LSTM model's forecasts are statistically better than those of the naive model, with a significant improvement in predictive accuracy.\u003c/p\u003e\n\u003cp\u003eIn contrast, the CNN model showed a slightly higher training error (MSE: 5.9133, RMSE: 2.4317), though still strong, with an R² of 0.9892, and a MAE comparable to LSTM (1.1887). Interestingly, on the test set, the CNN outperformed the LSTM in terms of MSE (5.8695 vs. 7.0434), RMSE (2.4227 vs. 2.6539), MAE (1.7265 vs. 2.0139), and MAPE (1.86% vs. 2.17%), with a slightly better R² of 0.9695. These results indicate that the CNN model generalized marginally better than the LSTM model, especially in terms of absolute and percentage error measures. However, the DM test statistic of -0.5173 (p = 0.6051) shows no statistically significant difference between the CNN and naive forecasts, suggesting that the performance gains, while numerically present, are not statistically conclusive.\u003c/p\u003e\n\u003cp\u003eFrom a modeling perspective, the LSTM’s architecture—optimized with 50 units, a 0.2 dropout rate, and a conservative learning rate—was effective at capturing long-term temporal dependencies, which are crucial in time series forecasting. However, the slightly higher test error and the statistically significant DM result imply that the model may have overfit the training data slightly more than desired. On the other hand, the CNN’s architecture—with 92 filters, a kernel size of 1, low dropout (≈ 0.043), a very small learning rate (0.00027), and a small batch size (9)—provided a robust and precise framework that balanced responsiveness to noise with a capacity to generalize. Its slightly better test metrics and balanced regularization suggest it handled the volatility in oil prices slightly more robustly, although its predictions were not statistically superior to the naive baseline.\u003c/p\u003e\n\u003cp\u003eIn conclusion, while both models performed impressively, the LSTM model demonstrated statistically significant forecasting superiority, especially when compared to naive predictions, making it a reliable choice when precision is paramount. The CNN model, although marginally better in test accuracy metrics, lacked statistical significance over naive forecasting but offered excellent generalization and computational efficiency, making it well-suited for applications requiring fast, interpretable models with limited data history (\u003cem\u003elag = 1\u003c/em\u003e). The final choice between the two models could depend on the specific operational or strategic priorities—statistical significance and deeper sequence learning (LSTM) vs. faster generalization with simpler architectures (CNN).\u003c/p\u003e\n\u003cdiv class=\"gridtable\"\u003e\n \u003cdiv class=\"colspec\" align=\"left\"\u003e\u0026nbsp;\u003c/div\u003e\n \u003ctable id=\"Tab3\" border=\"1\"\u003e\u003ccaption\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eLSTM and CNN Model Performance\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\u003cthead\u003e\u003ctr style=\"height: 35px;\"\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eModel\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eDataset\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eMSE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eMAE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eMAPE\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eR²\u003c/p\u003e\n \u003c/th\u003e\u003cth align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eDM Statistics\u003c/p\u003e\n \u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr style=\"height: 35px;\"\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eLSTM\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e2.2788\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e1.5096\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e1.1783\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.0200\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.9958\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e3.8630\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr style=\"height: 35px;\"\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e7.0434\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e2.6539\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e2.0139\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.0217\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.9634\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003ctr style=\"height: 35px;\"\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eCNN\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eTrain\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e5.9133\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e2.4317\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e1.1887\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.0301\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.9892\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e-0.5173\u003c/p\u003e\n \u003c/td\u003e\u003c/tr\u003e\u003ctr style=\"height: 35px;\"\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\u0026nbsp;\u003c/td\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\n \u003cp\u003eTest\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e5.8695\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e2.4227\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e1.7265\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.0186\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"char\" style=\"height: 35px;\"\u003e\n \u003cp\u003e0.9695\u003c/p\u003e\n \u003c/td\u003e\u003ctd align=\"left\" style=\"height: 35px;\"\u003e\u0026nbsp;\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003ctfoot\u003e\u003ctr style=\"height: 13.9562px;\"\u003e\u003ctd colspan=\"8\" style=\"height: 13.9562px;\"\u003e\u003cem\u003eSource – Author’s work\u003c/em\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tfoot\u003e\u003c/table\u003e\n\u003c/div\u003e\n\n"},{"header":"5. Conclusion and Implications","content":"\u003cp\u003eThe analysis of various machine learning models for oil price forecasting reveals several key insights. Linear regression stands out for its exceptional performance, achieving near-perfect fits and robust generalization, indicating it effectively captures the underlying linear trends in the data. Tree-based models like Random Forest, XGBoost, and GBM also demonstrate high predictive accuracy, with XGBoost and GBM performing comparably to linear regression, while Random Forest suggests a need for further refinement to mitigate potential overfitting. SVR provides reliable predictions, though slightly less effective than the leading models. In the deep learning domain, LSTM excels in capturing long-term dependencies, showing statistically significant improvements over naive predictions, albeit with a slight lag in responding to abrupt fluctuations. Conversely, CNN demonstrates superior short-term forecasting capabilities, exhibiting agility and responsiveness to local fluctuations, though its statistical significance over naive predictions is less pronounced. The critical role of hyperparameter tuning, using techniques like Grid Search, RandomizedSearchCV, and Optuna, is emphasized, highlighting its importance in optimizing model performance and balancing complexity with generalization. This study, while effective, is limited by its reliance on historical data, neglecting sudden market changes. Model generalization to other datasets is uncertain, and the univariate approach omits potentially valuable external factors. Also, some statistical significance results require further validation. These findings underscore the need for model selection based on specific forecasting objectives, the importance of evaluating generalization, and the necessity of rigorous hyperparameter tuning. The statistical significance of the models predictions, is also a very important factor. Practically, these models have significant implications for industries reliant on accurate oil price predictions, informing decision-making and risk mitigation. Future research should explore integrating external factors, investigating hybrid model architectures, and expanding the dataset to enhance forecasting accuracy and robustness.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eCNN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eConvolutional Neural Network\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eSVM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eSupport Vector Machine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDNN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDeep Neural Network\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLSTM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLong Short-Term Memory\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMSE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Squared Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRMSE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRoot Mean Squared Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGRU\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGated Recurrent Unit\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGBM\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGradient Boosting Machine\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eML\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMachine Learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eAI\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial Intelligence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eARIMA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAutoregressive Integrated Moving Average\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eANN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eArtificial Neural Network\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMAE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Absolute Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cul\u003e\n \u003cli\u003eThe manuscript is original and has not been published previously in any form.\u003c/li\u003e\n \u003cli\u003eThe research work is conducted ethically and responsibly.\u003c/li\u003e\n \u003cli\u003eThere is no conflict of interest related to this study.\u003c/li\u003e\n \u003cli\u003eAll sources of data and materials used in the study have been properly acknowledged.\u003c/li\u003e\n\u003c/ul\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAdeniyi EA, Gbadamosi B, Awotunde JB, Misra S, Sharma MM, Oluranti J (2022) \u003cem\u003eCrude Oil Price Prediction Using Particle Swarm Optimization and Classification Algorithms\u003c/em\u003e. \u003cem\u003e418 LNNS\u003c/em\u003e, 1384\u0026ndash;1394. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-3-030-96308-8_128\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-96308-8_128\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAdenusi CA, Vincent OR, Bolarinwa J, Oluwakemi O (2022) Predicting the Pandemic Effect of COVID 19 on the Nigeria Economic, Crude Oil as a Measure Parameter Using Machine Learning. Int Ser Oper Res Manage Sci 320:79\u0026ndash;93. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-3-030-87019-5_5\u003c/span\u003e\u003cspan address=\"10.1007/978-3-030-87019-5_5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmed S, Alshater MM, Ammari AE, Hammami H (2022) Artificial intelligence and machine learning in finance: A bibliometric review. \u003cem\u003eResearch in International Business and Finance\u003c/em\u003e, \u003cem\u003e61\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.ribaf.2022.101646\u003c/span\u003e\u003cspan address=\"10.1016/j.ribaf.2022.101646\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmed H, Siddiqui TA, Naushad M (2025) Navigating global economic turmoil: The dynamics of oil prices, exchange rates, and stock markets in BRICS. Invest Manage Financial Innovations 22(1):94\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmed H, Kaur R (2025) Dynamic Connectedness Among Oil Prices, Exchange Rate and Consumer Inflation: New Evidence from India. IIM Kozhikode Soc Manage Rev, 22779752251346247\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAhmed H (2025) Leveraging Machine Learning for Exchange Rate Prediction: Fresh Insights from BRICS Economies. Int J Econ Financial Issues 15(4):72\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlbahooth B (2024) Oil price movements predictions in Kingdom of Saudi Arabia using financial and macro-economic variables. Energy Explor Exploit 42(2):747\u0026ndash;771. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1177/01445987231206897\u003c/span\u003e\u003cspan address=\"10.1177/01445987231206897\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAldabagh H, Zheng X, Mukkamala R (2023) A Hybrid Deep Learning Approach for Crude Oil Price Prediction. J Risk Financial Manage 16(12). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/jrfm16120503\u003c/span\u003e\u003cspan address=\"10.3390/jrfm16120503\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlqahtani MG, Abdelhafez HA (2023) STOCK MARKET PREDICITION USING STATISTICAL \u0026amp; DEEP LEARNING TECHNIQUES. J Theoretical Appl Inform Technol 101(23):7808\u0026ndash;7825 Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAtif M, Rabbani MR, Jreisat A, Al-Mohamad S, Siddiqui TA, Hussain H, Ahmed H (2022) Time Varying Impact of Oil Prices on Stock Returns: Evidence from Developing Markets. Int J Sustainable Dev Plann, \u003cem\u003e17\u003c/em\u003e(2)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAziz MIA, Barawi MH, Shahiri H (2022) Is Facebook PROPHET Superior than Hybrid ARIMA Model to Forecast Crude Oil Price? Sains Malaysiana 51(8):2633\u0026ndash;2643. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.17576/jsm-2022-5108-22\u003c/span\u003e\u003cspan address=\"10.17576/jsm-2022-5108-22\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBakshi SS, Jaiswal RK, Jaiswal R (2021) \u003cem\u003eEfficiency Check Using Cointegration and Machine Learning Approach: Crude Oil Futures Markets\u003c/em\u003e. \u003cem\u003e191\u003c/em\u003e, 304\u0026ndash;311. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.procs.2021.07.038\u003c/span\u003e\u003cspan address=\"10.1016/j.procs.2021.07.038\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBarsky RB, Kilian L (2002) Oil and the macroeconomy since the 1970s. J Economic Perspect 18(4):115\u0026ndash;134\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBaumeister C, Kilian L (2016) Forty years of oil price fluctuations: Why the price of oil may still surprise us. J Economic Perspect 30(1):139\u0026ndash;160\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBernabe A, Martina E, Alvarez-Ramirez J, Ibarra-Valdez C (2004) A multi-model approach for describing crude oil price dynamics. Physica A 338(3\u0026ndash;4):567\u0026ndash;584\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBreiman L (2001) Random Forests. Springer, New York. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-1-4757-3264-1\u003c/span\u003e\u003cspan address=\"10.1007/978-1-4757-3264-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBoussatta H, Chihab M, Chihab Y, Chiny M (2023) Enhancing Oil Price Forecasting Through an Intelligent Hybridized Approach. Int J Adv Comput Sci Appl 14(9):115\u0026ndash;125. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.14569/IJACSA.2023.0140913\u003c/span\u003e\u003cspan address=\"10.14569/IJACSA.2023.0140913\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen J (2023) Analysis of Bitcoin Price Prediction Using Machine Learning. J Risk Financial Manage 16(1). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/jrfm16010051\u003c/span\u003e\u003cspan address=\"10.3390/jrfm16010051\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChen S-S, Hsu K-W (2012) Reverse globalization: Does high oil price volatility discourage international trade? Energy Econ 34(5):1634\u0026ndash;1643\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng W, Ming K, Ullah M (2024) Oil price volatility prediction using out-of-sample analysis \u0026ndash; Prediction efficiency of individual models, combination methods, and machine learning based shrinkage methods. \u003cem\u003eEnergy\u003c/em\u003e, \u003cem\u003e300\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.energy.2024.131496\u003c/span\u003e\u003cspan address=\"10.1016/j.energy.2024.131496\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCheng Y, Yi J, Yang X, Lai KK, Seco L (2022) A CEEMD-ARIMA-SVM model with structural breaks to forecast the crude oil prices linked with extreme events. Soft Comput 26(17):8537\u0026ndash;8551. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s00500-022-07276-5\u003c/span\u003e\u003cspan address=\"10.1007/s00500-022-07276-5\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChicco D, Warrens MJ, Jurman G (2021) The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. Peerj Comput Sci 7:e623\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDas PK, Das A (2020) Application of nonlinear stochastic single source of error state space models in the forecasting of mobile subscribers in India. Int J Data Sci 5(4):333\u0026ndash;357\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDas PK, Das PK (2024) Improvement in Inflation Forecasting: Ensembling Text Mining with Macro Data in Machine Learning Models. Int J Econ Finance 16(6):1\u0026ndash;92\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDhilsath FM, Samuel SJ (2021) Hyperparameter tuning of ensemble classifiers using grid search and random search for prediction of heart disease. Comput Intell Healthc Inf, 139\u0026ndash;158\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDing X, Fu L, Ding Y, Wang Y (2022) A novel hybrid method for oil price forecasting with ensemble thought. Energy Rep 8:15365\u0026ndash;15376. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.egyr.2022.11.061\u003c/span\u003e\u003cspan address=\"10.1016/j.egyr.2022.11.061\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDrucker H, Burges CJ, Kaufman L, Smola A, Vapnik V (1996) Support vector regression machines. \u003cem\u003eAdvances in Neural Information Processing Systems\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://proceedings.neurips.cc/paper/1996/hash/d38901788c533e8286cb6400b40b386d-Abstract.html\u003c/span\u003e\u003cspan address=\"https://proceedings.neurips.cc/paper/1996/hash/d38901788c533e8286cb6400b40b386d-Abstract.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDudek G (2015) Short-term load forecasting using random forests. In \u003cem\u003eIntelligent Systems' 2014: Proceedings of the 7th IEEE International Conference Intelligent Systems IS\u0026rsquo;2014, September 24-26, 2014, Warsaw, Poland, Volume 2: Tools, Architectures, Systems, Applications\u003c/em\u003e (pp. 821\u0026ndash;828). Springer International Publishing\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFang T, Zheng C, Wang D (2023) Forecasting the crude oil prices with an EMD-ISBM-FNN model. Energy 263:125407\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGabralla LA, Abraham A (2015) Comparison of hybrid intelligent approaches for prediction of crude oil price. Int J Comput Inform Syst Industrial Manage Appl 7(1):53\u0026ndash;65 Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGao X, Wang J, Yang L (2022) An Explainable Machine Learning Framework for Forecasting Crude Oil Price during the COVID-19 Pandemic. Axioms 11(8). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/axioms11080374\u003c/span\u003e\u003cspan address=\"10.3390/axioms11080374\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGholamian MR, Ghomi SF, Ghazanfari M (2005) A hybrid systematic design for multiobjective market problems: A case study in crude oil markets. Eng Appl Artif Intell 18(4):495\u0026ndash;509\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGong Y, Zhang P (2022) \u003cem\u003eResearch and Implementation of Oil Trading Data Analysis and Prediction based on Random Forest Regression Algorithm\u003c/em\u003e. 14\u0026ndash;22. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/AIAM57466.2022.00011\u003c/span\u003e\u003cspan address=\"10.1109/AIAM57466.2022.00011\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGuliyev H, Mustafayev E (2022) Predicting the changes in the WTI crude oil price dynamics using machine learning models. \u003cem\u003eResources Policy\u003c/em\u003e, \u003cem\u003e77\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.resourpol.2022.102664\u003c/span\u003e\u003cspan address=\"10.1016/j.resourpol.2022.102664\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHagen R (1994) How is the international price of a particular crude determined? OPEC Rev 18(1):127\u0026ndash;135\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHamilton JD, Herrera AM (2004) Comment: Oil shocks and aggregate macroeconomic behavior: the role of monetary policy. J Money Credit Bank, 265\u0026ndash;286\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHao X (2024) \u003cem\u003eCrude Oil Prediction Based on Multi-Factor LSTM-Transformer Algorithm\u003c/em\u003e. \u003cem\u003e51\u003c/em\u003e, 507\u0026ndash;514. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3233/ATDE240114\u003c/span\u003e\u003cspan address=\"10.3233/ATDE240114\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHasan M, Das U, Datta RK, Abedin MZ (2023) Model Development for Predicting the Crude Oil Price: Comparative Evaluation of Ensemble and Machine Learning Methods. In \u003cem\u003eInternational Series in Operations Research and Management Science\u003c/em\u003e (Vol. 336, pp. 167\u0026ndash;179). Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-3-031-18552-6_10\u003c/span\u003e\u003cspan address=\"10.1007/978-3-031-18552-6_10\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHope TM (2020) Linear regression. Machine learning. Academic, pp 67\u0026ndash;81\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoque KE, Aljamaan H (2021) Impact of hyperparameter tuning on machine learning models in stock price forecasting. IEEE Access 9:163815\u0026ndash;163830\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eInternational Monetary Fund (2024) Global price of Brent Crude [POILBREUSDM], retrieved from FRED, Federal Reserve Bank of St. Louis; \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://fred.stlouisfed.org/series/POILBREUSDM\u003c/span\u003e\u003cspan address=\"https://fred.stlouisfed.org/series/POILBREUSDM\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, June 7\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJahandoost A, Houshmand M, Hosseini SA (2023) \u003cem\u003ePrediction of West Texas Intermediate Crude-oil Price Using Hybrid Attention-based Deep Neural Networks: A Comparative Study\u003c/em\u003e. 240\u0026ndash;245. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ICCKE60553.2023.10326291\u003c/span\u003e\u003cspan address=\"10.1109/ICCKE60553.2023.10326291\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJahan-Parvar MR, Mohammadi H (2009) Oil prices and competitiveness: Time series evidence from six oil-producing countries. J Economic Stud 36(1):98\u0026ndash;118\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJahanshahi H, Uzun S, Ka\u0026ccedil;ar S, Yao Q, Alassafi MO (2022) Artificial Intelligence-Based Prediction of Crude Oil Prices Using Multiple Features under the Effect of Russia\u0026ndash;Ukraine War and COVID-19 Pandemic. Mathematics 10(22). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/math10224361\u003c/span\u003e\u003cspan address=\"10.3390/math10224361\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJha N, Tanneru K, Palla H, S., Hussain Mafat I (2024) Multivariate analysis and forecasting of the crude oil prices: Part I \u0026ndash; Classical machine learning approaches. \u003cem\u003eEnergy\u003c/em\u003e, \u003cem\u003e296\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.energy.2024.131185\u003c/span\u003e\u003cspan address=\"10.1016/j.energy.2024.131185\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKadri F, Abdennbi K (2020) Rnn-based deep-learning approach to forecasting hospital system demands: application to an emergency department. Int J Data Sci 5(1):1\u0026ndash;25\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKakade KA, Ghate KS, Jaiswal RK, Jaiswal R (2023) A Novel Approach to Forecast Crude Oil Prices Using Machine Learning and Technical Indicators. J Adv Inform Technol 14(2):302\u0026ndash;310. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.12720/jait.14.2.302-310\u003c/span\u003e\u003cspan address=\"10.12720/jait.14.2.302-310\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKandil A, Khaled S, Elfakharany T (2023) Prediction of the equivalent circulation density using machine learning algorithms based on real-time data. AIMS Energy 11(3):425\u0026ndash;453. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3934/energy.2023023\u003c/span\u003e\u003cspan address=\"10.3934/energy.2023023\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKaufmann RK, Dees S, Karadeloglou P, Sanchez M (2004) Does OPEC Matter? An Econometric Analysis of Oil Prices. Energy J 25(4):67\u0026ndash;90. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5547/ISSN0195-6574-EJ-Vol25-No4-4\u003c/span\u003e\u003cspan address=\"10.5547/ISSN0195-6574-EJ-Vol25-No4-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKeerthan JS, Nagasai Y, Shaik S (2019) Machine learning algorithms for oil price prediction. Int J Innovative Technol Exploring Eng 8(8):958\u0026ndash;963 Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhashman A, Nwulu NI (2011) \u003cem\u003eIntelligent prediction of crude oil price using Support Vector Machines\u003c/em\u003e. 165\u0026ndash;169. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/SAMI.2011.5738868\u003c/span\u003e\u003cspan address=\"10.1109/SAMI.2011.5738868\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKim GI, Jang B (2023) Petroleum Price Prediction with CNN-LSTM and CNN-GRU Using Skip-Connection. Mathematics 11(3). \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/math11030547\u003c/span\u003e\u003cspan address=\"10.3390/math11030547\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKumar R, Kumar P, Kumar Y (2020) Time series data prediction using IoT and machine learning technique. Procedia Comput Sci 167:373\u0026ndash;381\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLakshminarayanan SK, McCrae J (2019) \u003cem\u003eA comparative study of SVM and LSTM deep learning algorithms for stock market prediction\u003c/em\u003e. \u003cem\u003e2563\u003c/em\u003e, 446\u0026ndash;457. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://www.scopus.com/inward/record.uri?eid=2-s2.0-85081603042\u0026amp;partnerID=40\u0026amp;md5=e3cf7e4ee4ff1913e56939f57e0b3843\u003c/span\u003e\u003cspan address=\"https://www.scopus.com/inward/record.uri?eid=2-s2.0-85081603042\u0026amp;partnerID=40\u0026amp;md5=e3cf7e4ee4ff1913e56939f57e0b3843\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLi N, Li J, Wang Q, Yan D, Wang L, Jia M (2024) A novel copper price forecasting ensemble method using adversarial interpretive structural model and sparrow search algorithm. \u003cem\u003eResources Policy\u003c/em\u003e, \u003cem\u003e91\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.resourpol.2024.104892\u003c/span\u003e\u003cspan address=\"10.1016/j.resourpol.2024.104892\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiang X, Luo P, Li X, Wang X, Shu L (2023) Crude oil price prediction using deep reinforcement learning. \u003cem\u003eResources Policy\u003c/em\u003e, \u003cem\u003e81\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.resourpol.2023.103363\u003c/span\u003e\u003cspan address=\"10.1016/j.resourpol.2023.103363\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu L, Zhou S, Jie Q, Du P, Xu Y, Wang J (2024) A robust time-varying weight combined model for crude oil price forecasting. \u003cem\u003eEnergy\u003c/em\u003e, \u003cem\u003e299\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.energy.2024.131352\u003c/span\u003e\u003cspan address=\"10.1016/j.energy.2024.131352\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eManjula KA, Karthikeyan P (2019) \u003cem\u003eGold price prediction using ensemble based machine learning techniques\u003c/em\u003e. \u003cem\u003e2019-April\u003c/em\u003e, 1360\u0026ndash;1364. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/icoei.2019.8862557\u003c/span\u003e\u003cspan address=\"10.1109/icoei.2019.8862557\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMohsin M, Jamaani F (2023) Green finance and the socio-politico-economic factors\u0026rsquo; impact on the future oil prices: Evidence from machine learning. \u003cem\u003eResources Policy\u003c/em\u003e, \u003cem\u003e85\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.resourpol.2023.103780\u003c/span\u003e\u003cspan address=\"10.1016/j.resourpol.2023.103780\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMuganda BW, Kasamani BS (2023) \u003cem\u003eParallel Programming for Portfolio Optimization: A Robo-Advisor Prototype using Genetic Algorithms with Recurrent Neural Networks\u003c/em\u003e. 167\u0026ndash;176. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ICCNS58795.2023.10193396\u003c/span\u003e\u003cspan address=\"10.1109/ICCNS58795.2023.10193396\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eM\u0026uuml;ller KR, Smola AJ, R\u0026auml;tsch G, Sch\u0026ouml;lkopf B, Kohlmorgen J, Vapnik V (1998) Using support vector machines for time series prediction\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMurugesan R, Azhaganathan B, Maitra S (2023) Commodities vs. S\u0026amp;P 500: Causal interaction, temporal analysis and predictive modelling using econometric approach, machine learning, and deep learning. Int J Bus Inform Syst 42(3\u0026ndash;4):429\u0026ndash;457. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1504/IJBIS.2023.129725\u003c/span\u003e\u003cspan address=\"10.1504/IJBIS.2023.129725\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNaeem M, Aamir M, Yu J, Albalawi O (2024) A Novel Approach for Reconstruction of IMFs of Decomposition and Ensemble Model for Forecasting of Crude Oil Prices. IEEE Access 12:34192\u0026ndash;34207. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ACCESS.2024.3370440\u003c/span\u003e\u003cspan address=\"10.1109/ACCESS.2024.3370440\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNanthiya D, Gopal SB, Balakumar S, Harisankar M, Midhun SP (2023) \u003cem\u003eGold Price Prediction using ARIMA model\u003c/em\u003e. ViTECoN 2023\u0026ndash;2nd IEEE International Conference on Vision Towards Emerging Trends in Communication and Networking Technologies, Proceedings. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ViTECoN58111.2023.10157017\u003c/span\u003e\u003cspan address=\"10.1109/ViTECoN58111.2023.10157017\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNayak RK, Tripathy R, Mishra D, Burugari VK, Selvaraj P, Sethy A, Jena B (2021) Indian stock market prediction based rough set support vector Mach approach 153:345\u0026ndash;355. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-981-15-6202-0_35\u003c/span\u003e\u003cspan address=\"10.1007/978-981-15-6202-0_35\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNguyen TT, Nguyen HG, Lee JY, Wang YL, Tsai CS (2023) The consumer price index prediction using machine learning approaches: Evidence from the United States. Heliyon, \u003cem\u003e9\u003c/em\u003e(10)\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eOkechukwu Ajakwe S, Nwakanma I, Lee C, J.-M., Kim D-S (2020) \u003cem\u003eMachine Learning Algorithm for Intelligent Prediction for Military Logistics and Planning\u003c/em\u003e. \u003cem\u003e2020-October\u003c/em\u003e, 417\u0026ndash;419. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ICTC49870.2020.9289286\u003c/span\u003e\u003cspan address=\"10.1109/ICTC49870.2020.9289286\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePai PF, Lin KP, Lin CS, Chang PT (2010) Time series forecasting by a seasonal support vector regression model. Expert Syst Appl 37(6):4261\u0026ndash;4265\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePanella M, Barcellona F, D\u0026rsquo;Ecclesia RL (2012) Forecasting energy commodity prices using neural networks. \u003cem\u003eAdvances in Decision Sciences\u003c/em\u003e, \u003cem\u003e2012\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1155/2012/289810\u003c/span\u003e\u003cspan address=\"10.1155/2012/289810\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePrakash A, Singh SK (2022) A Comparative Study of Time Series, Machine Learning, and Ensemble Models for Crude Oil Price Prediction. 914:157\u0026ndash;171. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-981-19-2980-9_13\u003c/span\u003e\u003cspan address=\"10.1007/978-981-19-2980-9_13\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRimal Y, Sharma N, Alsadoon A (2024) The accuracy of machine learning models relies on hyperparameter tuning: student result classification using random forest, randomized search, grid search, bayesian, genetic, and optuna algorithms. Multimedia Tools Appl 83(30):74349\u0026ndash;74364\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSiddiqui TA, Ahmed H, Naushad M, Khan U (2023) The relationship between oil prices and exchange rate: a systematic literature review. Int J Energy Econ Policy 13(3):566\u0026ndash;578\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSajid SW, Hasan M, Rabbi MF, Abedin MZ (2023) An Ensemble LGBM (Light Gradient Boosting Machine) Approach for Crude Oil Price Prediction. In \u003cem\u003eInternational Series in Operations Research and Management Science\u003c/em\u003e (Vol. 336, pp. 153\u0026ndash;165). Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-3-031-18552-6_9\u003c/span\u003e\u003cspan address=\"10.1007/978-3-031-18552-6_9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSen D, Hamurcuoglu KI, Ersoy MZ, Tun\u0026ccedil; KMM, G\u0026uuml;nay ME (2023) Forecasting long-term world annual natural gas production by machine learning. \u003cem\u003eResources Policy\u003c/em\u003e, \u003cem\u003e80\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.resourpol.2022.103224\u003c/span\u003e\u003cspan address=\"10.1016/j.resourpol.2022.103224\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eShao YE, Lu C-J, Hou C-D (2014) Hybrid soft computing schemes for the prediction of import demand of crude oil in Taiwan. \u003cem\u003eMathematical Problems in Engineering\u003c/em\u003e, \u003cem\u003e2014\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1155/2014/257947\u003c/span\u003e\u003cspan address=\"10.1155/2014/257947\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eStevens P (1995) Understanding the Oil Industry: Economics as a Help or a Hindrance. Energy J 16(3):125\u0026ndash;139. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5547/ISSN0195-6574-EJ-Vol16-No3-6\u003c/span\u003e\u003cspan address=\"10.5547/ISSN0195-6574-EJ-Vol16-No3-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSulaiman A, Mahmood FE, Majeed SA (2023) Long-Term Solar Irradiance Forecasting Using Multilinear Predictors. Int J Electr Electron Eng Telecommunications 12(2):134\u0026ndash;141. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.18178/ijeetc.12.2.134-141\u003c/span\u003e\u003cspan address=\"10.18178/ijeetc.12.2.134-141\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTami M, Owda AY (2024) Efficient commodity price forecasting using long short-term memory model. IAES Int J Artif Intell 13(1):994\u0026ndash;1004. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.11591/ijai.v13.i1.pp994-1004\u003c/span\u003e\u003cspan address=\"10.11591/ijai.v13.i1.pp994-1004\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTang L, Dai W, Yu L, Wang S (2015) A novel CEEMD-based eelm ensemble learning paradigm for crude oil price forecasting. Int J Inform Technol Decis Mak 14(1):141\u0026ndash;169. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1142/S0219622015400015\u003c/span\u003e\u003cspan address=\"10.1142/S0219622015400015\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTebyanian A, Hedayati F (2014) \u003cem\u003eIntelligent crude oil price forecaster\u003c/em\u003e. 453\u0026ndash;455. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ICMLA.2014.79\u003c/span\u003e\u003cspan address=\"10.1109/ICMLA.2014.79\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eTissaoui K, Zaghdoudi T, Hakimi A, Nsaibi M (2023) Do Gas Price and Uncertainty Indices Forecast Crude Oil Prices? Fresh Evidence Through XGBoost Modeling. Comput Econ 62(2):663\u0026ndash;687. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s10614-022-10305-y\u003c/span\u003e\u003cspan address=\"10.1007/s10614-022-10305-y\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVapnik V (1998) Statistical learning theory Wiley. New York 1(624):2\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWang L, Xia Y, Lu Y (2022) A Novel Forecasting Approach by the GA-SVR-GRNN Hybrid Deep Learning Algorithm for Oil Future Prices. \u003cem\u003eComputational Intelligence and Neuroscience\u003c/em\u003e, \u003cem\u003e2022\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1155/2022/4952215\u003c/span\u003e\u003cspan address=\"10.1155/2022/4952215\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWei MX, Kee OL, Musa S (2023) \u003cem\u003eSeasonal versus non-seasonal trends in stock market Malaysia\u003c/em\u003e. \u003cem\u003e389\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1051/e3sconf/202338909039\u003c/span\u003e\u003cspan address=\"10.1051/e3sconf/202338909039\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWeng F, Chen Y, Wang Z, Hou M, Luo J, Tian Z (2020) Gold price forecasting research based on an improved online extreme learning machine algorithm. J Ambient Intell Humaniz Comput 11(10):4101\u0026ndash;4111. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s12652-020-01682-z\u003c/span\u003e\u003cspan address=\"10.1007/s12652-020-01682-z\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXu Z, Mohsin M, Ullah K, Ma X (2023) Using econometric and machine learning models to forecast crude oil prices: Insights from economic history. \u003cem\u003eResources Policy\u003c/em\u003e, \u003cem\u003e83\u003c/em\u003e. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.resourpol.2023.103614\u003c/span\u003e\u003cspan address=\"10.1016/j.resourpol.2023.103614\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYu L, Zhang X, Wang S (2017) Assessing potentiality of support vector machine method in crude oil price forecasting. Eurasia J Math Sci Technol Educ 13(12):7893\u0026ndash;7904. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.12973/ejmste/77926\u003c/span\u003e\u003cspan address=\"10.12973/ejmste/77926\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e. Scopus\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZaidi A, Oussalah M (2018) \u003cem\u003eForecasting weekly crude oil using twitter sentiment of U.S. foreign policy and oil companies data\u003c/em\u003e. 201\u0026ndash;208. Scopus. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/IRI.2018.00037\u003c/span\u003e\u003cspan address=\"10.1109/IRI.2018.00037\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhu N, Zhu C, Zhou L, Zhu Y, Zhang X (2022) Optimization of the random forest hyperparameters for power industrial control systems intrusion detection using an improved grid search algorithm. Appl Sci 12(20):10456\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":true,"hideJournal":true,"highlight":"","institution":"New Delhi Institute of Management","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"s – Q43, C63, C53, C52, C45","lastPublishedDoi":"10.21203/rs.3.rs-8354755/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-8354755/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThis study investigates the efficacy of machine learning (ML) models in forecasting crude oil prices, a critical factor influencing economic stability. Given the inherent volatility and complexity of oil markets, accurate prediction is essential for mitigating risks and informing strategic decisions. Employing an autoregressive framework, the research utilizes daily oil price data from March 2013 to February 2024 to train and evaluate a diverse set of ML algorithms, including linear regression, Random Forest, SVR, XGBoost, Gradient Boosting Machine (GBM), LSTM, and CNN. The analysis reveals that linear regression demonstrates exceptional performance, effectively capturing linear trends, while tree-based models like XGBoost and GBM also exhibit high predictive accuracy. LSTM proves adept at capturing long-term dependencies, though CNN shows superior agility in short-term forecasting. Overall, linear regression in machine learning, and LSTM from deep learning are best models in forecasting. The study underscores the importance of rigorous hyperparameter tuning and model selection based on specific forecasting objectives. However, limitations stemming from reliance on historical data and the univariate approach are acknowledged. Future research should explore incorporating external factors and hybrid model architectures to enhance forecasting robustness.\u003c/p\u003e","manuscriptTitle":"Forecasting Crude Oil Prices: Insights from Machine Learning Approaches","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-12-16 12:56:00","doi":"10.21203/rs.3.rs-8354755/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"7eb6b22f-9ad5-4e41-8471-8d5ea800b395","owner":[],"postedDate":"December 16th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[{"id":59608635,"name":"Finance"}],"tags":[],"updatedAt":"2025-12-16T12:56:00+00:00","versionOfRecord":[],"versionCreatedAt":"2025-12-16 12:56:00","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-8354755","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-8354755","identity":"rs-8354755","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Outcome instruments

MUSA

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00