Application of Machine Learning Algorithms for Groundwater Level Prediction in the Najafabad Plain

preprint OA: closed
Full text JSON View at publisher

Abstract

Abstract Accurate groundwater level prediction is vital for sustainable water management, particularly in arid and semi-arid regions under climatic and human-induced stress. This study investigates the performance of three machine learning algorithms—Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Support Vector Machine (SVM)—to forecast groundwater levels in five hydrogeological zones of the Najafabad plain, Iran. Models were trained using climatic (precipitation, temperature), hydrological (previous groundwater level), and anthropogenic (irrigation, pumping) inputs. Performance was evaluated using R², RMSE, and MAE during training and testing phases. XGBoost outperformed others with a mean testing R² of 0.848 and RMSE as low as 1.554, showing high robustness, especially in zones with stable to moderately varying water tables. RF achieved moderate accuracy (mean R² = 0.748) with consistent error margins. SVM showed high training performance (R² = 0.918) but reduced generalization (testing R² = 0.822), indicating overfitting tendencies. Results highlight the effectiveness of ensemble models like XGBoost for heterogeneous aquifer systems. However, limited resolution of anthropogenic data remains a challenge. This study reinforces the value of machine learning—particularly gradient boosting—as a reliable, scalable approach for groundwater monitoring in data-scarce, complex environments.
Full text 171,702 characters · extracted from preprint-html · click to expand
Application of Machine Learning Algorithms for Groundwater Level Prediction in the Najafabad Plain | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Article Application of Machine Learning Algorithms for Groundwater Level Prediction in the Najafabad Plain Sahar Davari, Saeid Eslamian, Mohammad Jamali, Hamid Reza Safavi This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-6992628/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 10 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted 12 You are reading this latest preprint version Abstract Accurate groundwater level prediction is vital for sustainable water management, particularly in arid and semi-arid regions under climatic and human-induced stress. This study investigates the performance of three machine learning algorithms—Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Support Vector Machine (SVM)—to forecast groundwater levels in five hydrogeological zones of the Najafabad plain, Iran. Models were trained using climatic (precipitation, temperature), hydrological (previous groundwater level), and anthropogenic (irrigation, pumping) inputs. Performance was evaluated using R², RMSE, and MAE during training and testing phases. XGBoost outperformed others with a mean testing R² of 0.848 and RMSE as low as 1.554, showing high robustness, especially in zones with stable to moderately varying water tables. RF achieved moderate accuracy (mean R² = 0.748) with consistent error margins. SVM showed high training performance (R² = 0.918) but reduced generalization (testing R² = 0.822), indicating overfitting tendencies. Results highlight the effectiveness of ensemble models like XGBoost for heterogeneous aquifer systems. However, limited resolution of anthropogenic data remains a challenge. This study reinforces the value of machine learning—particularly gradient boosting—as a reliable, scalable approach for groundwater monitoring in data-scarce, complex environments. Earth and environmental sciences/Environmental sciences Earth and environmental sciences/Hydrology Groundwater level prediction Machine Learning Najafabad Plain Random Forest Extreme Gradient Boosting Support Vector Machine Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 1. Introduction Groundwater plays a vital role in sustaining agricultural productivity, urban development, and ecosystem services, particularly in arid and semi-arid regions such as central Iran. In these environments, where surface water resources are limited or highly variable, aquifers serve as the primary source of freshwater for both irrigation and domestic use. However, overexploitation of groundwater due to increasing population, agricultural expansion, and declining precipitation under climate change has led to alarming consequences, including groundwater depletion, land subsidence, reduced water quality, and diminished long-term water security 1 – 3 . This escalating stress highlights the urgent need for robust groundwater monitoring and predictive tools to support sustainable resource management. Traditional physically-based models such as MODFLOW have long been used for simulating groundwater flow and levels 4 . While these models offer valuable process-based insights, their application is often constrained by the requirement for extensive input data (e.g., aquifer geometry, hydraulic conductivity, recharge rates), which are frequently sparse or uncertain—particularly in data-scarce regions like Iran 5 . Additionally, these models can be computationally intensive and difficult to calibrate under dynamically changing human-water interactions. In response to these challenges, data-driven models—especially machine learning (ML) algorithms—have emerged as powerful alternatives for groundwater prediction. Algorithms such as Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost) have shown promising results in capturing nonlinear relationships between groundwater levels and influencing factors such as precipitation, temperature, and human activities 6 – 8 . These models offer greater flexibility, faster computation, and reduced dependency on physical assumptions, making them suitable for complex or data-limited environments 9 . Several studies have successfully applied ML methods to groundwater level forecasting in various regions of Iran and beyond 10 – 13 . For example, Vafadar, et al. 14 applied RF, SVM, and XGBoost for groundwater potential mapping in the Tehran and Karaj plains and found XGBoost to outperform the other models. Similarly, Zarafshan, et al. 15 evaluated various ML models for groundwater prediction in the Najafabad region and confirmed the potential of support vector-based approaches. Ibrahem Ahmed Osman, et al. 16 , who applied XGBoost to predict groundwater depth in Selangor, Malaysia, and reported high predictive accuracy. However, many of these studies focused on either a single model or an aggregated spatial scale, often overlooking the spatial heterogeneity of hydrogeological conditions and the importance of local calibration. This study addresses this gap by conducting a comprehensive comparative analysis of RF, SVM, and XGBoost for predicting monthly groundwater levels in the Najafabad plain, Isfahan province. A novel contribution of this work is the spatial disaggregation of the study area into five distinct hydrogeological zones, enabling more detailed insight into the model performance under varying aquifer conditions. This zonal framework enhances the interpretability and applicability of the models for groundwater management in heterogeneous settings. The innovation of this research is threefold: It introduces a zone-based modeling approach to assess the spatial variability of ML model performance, rarely considered in previous Iranian studies. It compares three state-of-the-art ML models under consistent settings, evaluating their accuracy, robustness, and generalization ability. It provides a feature importance analysis to identify the dominant climatic and anthropogenic drivers of groundwater fluctuations across zones. The outcomes of this research are expected to support targeted groundwater management strategies, inform adaptive policy-making, and contribute to the growing body of literature on machine learning applications in groundwater science. However, the study is not without limitations. The models rely solely on historical time series data, which may not fully capture future changes driven by unexpected anthropogenic or climatic shifts. Additionally, the spatial resolution of input data and the availability of high-quality monitoring records may influence the reliability of predictions, particularly in areas with sparse observational coverage. Finally, while machine learning models are powerful in capturing correlations, they do not inherently provide mechanistic insights into subsurface hydrological processes, and thus should be complemented by physical understanding for integrated groundwater management. 2. Methods 2.1. Study area The Najafabad Plain, located in Isfahan Province in central Iran, is a key agricultural region within the Zayandeh Rud Basin. It spans approximately 1,712 km², comprising a central alluvial aquifer of about 940.9 km², with the remaining 772.5 km² consisting of surrounding highlands. Geographically, the plain lies between longitudes 50°57′ and 51°44′26″ E and latitudes 32°20′13″ to 32°49′21″ N (Fig. 1 ). The aquifer is hydrologically isolated from adjacent groundwater systems and functions as the principal reservoir for subsurface water in the area. The elevation of the plain ranges from roughly 2,906 meters above sea level in the northern highlands to approximately 1,538 meters in the southern lowlands, near the Zayandeh Rud River, which provides intermittent surface flow. The region experiences a semi-arid climate, with mean annual precipitation ranging from 158 to 172 mm, while the potential evapotranspiration exceeds 1,500 mm. The average annual temperature is approximately 15°C. Water usage in the region is dominated by agriculture, accounting for about 93.3% of the total water consumption. Domestic and industrial uses represent around 4.5% and 2.2%, respectively. Groundwater supplies nearly 72% of the total water demand, primarily through wells. In total, 15,933 groundwater extraction points have been identified, including 15,753 wells, 151 qanats, and 29 springs. Among them, 15,673 wells are located within the alluvial aquifer, collectively extracting approximately 881 million cubic meters of water annually 17 . In recent years, declining precipitation and increasing water demand have exerted considerable pressure on groundwater resources. Although the development of new wells is legally restricted, unauthorized wells—frequently reported by the Isfahan Regional Water Authority—pose a significant challenge to sustainable groundwater management in the plain. 2.2. Data Collection and Preprocessing 2.2.1. Meteorological, Surface Water, and Groundwater Data To investigate groundwater level fluctuations in the Najafabad Plain, comprehensive hydrological and meteorological datasets were collected and preprocessed. Meteorological variables, including precipitation and temperature, were obtained from the Najafabad synoptic station, which served as the primary source for representing regional climatic conditions. Surface water data were acquired from three active hydrometric stations—Lenj, Mousian, and Zefreh—located along the Zayandeh Rud River, the principal surface water source in the region. These stations provide valuable insights into river flow dynamics and their interactions with the underlying groundwater system. The total annual surface inflow into the plain, including contributions from the Lenjanat, Mahyar-e-Shomali, and Karon sub-basins, was estimated at approximately 557.4 million cubic meters. These representative stations (see Table 1 and Fig. 1 ) were selected based on their strategic locations and the availability of long-term, reliable data, ensuring adequate spatial and temporal coverage of the key hydrological and climatic variables influencing groundwater levels across the study area. During the study period (2011–2021), the average annual precipitation was approximately 195 mm in the upland regions and 153.7 mm in the plain. March was identified as the wettest month, with average rainfall values of 31.6 mm in the uplands and 29.3 mm in the plain. Mean annual temperatures were estimated at around 13.8°C in the uplands and 15.4°C in the plain. Table 1 Selected Stations for the Research Station Name Station Type Longitude (°E) Latitude (°N) Elevation (m) Start Year End Year Najafabad Synoptic 51.389 32.604 1634 1970 2022 Zefreh Hydrometric 51.4986 32.5019 1623 1965 2022 Mousian Hydrometric 51.5261 32.5772 1600 1995 2022 Lenej Hydrometric 51.5575 39.3236 1646 1980 2022 Groundwater level data were collected from 53 observation wells distributed across the Najafabad Plain, corresponding to a monitoring density of approximately 1.4 wells per 25 km² (see Fig. 1 ). These wells provided continuous records of groundwater level fluctuations throughout the aquifer system, which extends beneath approximately 87% of the study area and constitutes the primary source of groundwater extraction. For spatial analysis, the study area was divided into two major zones: the uplands and the plain. The uplands, primarily located in the northern and western parts of the region, cover an area of about 679.2 km² and are predominantly composed of Jurassic and Cretaceous limestone formations. In contrast, the plain—spanning approximately 932.2 km²—is characterized by Quaternary alluvial deposits with variable thicknesses and hydraulic conductivities. This zoning approach enabled differentiation of hydroclimatic and geological conditions across the region, thereby enhancing the accuracy of groundwater level modeling in response to meteorological and surface water drivers. 2.2.2. Outlier Detection and Treatment To enhance the quality and reliability of the dataset, outlier detection was performed using the boxplot method—a widely adopted graphical technique in hydrological studies for identifying anomalous values. This approach effectively visualizes the distribution and dispersion of data, highlighting values that deviate significantly from the typical range. Given that the primary objective of this study is not the analysis of extreme events, the identified outliers—presumed to be due to measurement or recording errors—were excluded from further analysis. Subsequently, the resulting gaps in the time series were filled using interpolation techniques, ensuring data continuity and temporal consistency across all variables. 2.2.3. Regional and Aquifer Zoning The Najafabad aquifer, like many complex groundwater systems, exhibits significant spatial heterogeneity due to the combined influence of natural and anthropogenic factors such as groundwater abstraction, precipitation, riverbed infiltration, and temperature variability. Owing to the physiographic diversity of the plain, these factors impact different regions of the aquifer to varying degrees. For example, in the western parts of the aquifer—located at a considerable distance from the Zayandeh-Rud River—the influence of river discharge on groundwater levels is minimal to negligible. To improve the spatial resolution and accuracy of groundwater level modeling, the aquifer was subdivided into five distinct zones. This zoning was based on a conceptual understanding of the system, as well as the physical and geographical characteristics of the area, particularly in relation to the presence of the river and irrigation canals. The spatial layout and boundaries of these zones are depicted in Fig. 1 , and their key characteristics are summarized in Table 2 . Table 2 Description of groundwater zones based on dominant influencing factors in the Najafabad aquifer. Zone Name Main Influential Factor Description Z1 River Corridor River discharge Influenced by proximity to river Z2 Right Side of Nekouabad Canal Irrigation canal (Nekouabad) Affected by controlled irrigation system Z3 Left Side of Nekouabad canal Irrigation canal (Nekouabad) Similar to Z2 but spatially distinct Z4 Khamiran Irrigation Area Groundwater abstraction Affected by Khamiran irrigation scheme Z5 Western Highlands Precipitation and natural recharge Mostly natural influence, limited access Zone 1 – River Corridor Zone: This zone includes areas adjacent to the Zayandeh-Rud River, where groundwater levels are directly influenced by river discharge. Piezometric data from this region reflect the dynamic interaction between surface water and groundwater. Zones 2 and 3 – Nekouabad Irrigation Zones: These zones encompass agricultural areas affected by the Nekouabad irrigation network. Based on their geographical orientation, they are divided into: Zone 2: Right side of the Nekouabad irrigation system, and Zone 3: Left side of the Nekouabad irrigation system. Zone 4 – Khamiran Irrigation Zone: Located in the western part of the aquifer, this zone is mainly influenced by groundwater abstraction for agriculture under the Khamiran irrigation scheme. Zone 5 – Western Highland Zone: This zone includes piezometers situated in the elevated western highlands. Here, groundwater dynamics are primarily governed by natural recharge from precipitation, with limited anthropogenic impact. 2.2.4. Exploratory Data Analysis of Groundwater Level Fluctuations To gain deeper insights into the temporal behavior of groundwater levels across the Najafabad Plain, an exploratory data analysis (EDA) was conducted using violin plots. These plots effectively combine boxplots and kernel density estimates to depict the distribution and variability of groundwater levels on both monthly and seasonal scales, across the five delineated observation zones (Z1 to Z5). The monthly and seasonal distributions are presented in Fig. 2 and Fig. 3 , respectively. The monthly violin plots (Fig. 2 ) reveal distinct fluctuation patterns among the zones. Zone 1 (Z1), located near the river corridor, exhibits pronounced variability during both summer and winter months, as indicated by the wider spread of values. Zone 2 (Z2) shows a gradual rise in groundwater levels from April to December, along with considerable dispersion during the transitional months, potentially due to irrigation cycles. Zone 3 (Z3) maintains relatively stable levels, with the most notable fluctuations occurring between June and October—likely reflecting seasonal agricultural demand. In Zone 4 (Z4), groundwater levels remain consistent throughout the year except for December, which displays a sharp increase in variability. Zone 5 (Z5), located in the western highlands, demonstrates a gradual rise in groundwater levels beginning in June and peaking around October, possibly due to delayed recharge from precipitation or irrigation return flows. Seasonal violin plots (Fig. 3 ) provide a broader view of temporal groundwater dynamics. Zones Z1, Z2, and Z3 show wider distributions during spring and summer, suggesting higher variability linked to increased evapotranspiration or agricultural withdrawals. In contrast, Z4 and Z5 exhibit narrower, more stable distributions across seasons, with a slight increase in variability during fall and winter. Median groundwater levels are generally higher during spring and summer, particularly in Z1 and Z2, which may reflect the influence of seasonal recharge or irrigation inputs. These visual analyses offer critical insights into the temporal variability of groundwater across different zones, highlighting areas that are more sensitive to environmental and anthropogenic influences. The outcomes of this analysis not only enhance conceptual understanding but also inform the subsequent stages of feature engineering, model training, and evaluation in the groundwater prediction framework. 2.3. Groundwater Level Modeling Approach 2.3.1. Machine Learning Algorithms In this study, three state-of-the-art supervised machine learning algorithms—Random Forest (RF), Gradient Boosting (GB), and Support Vector Machine (SVM)—were implemented to model and forecast groundwater level variations across the Najafabad aquifer. These algorithms are widely recognized for their robustness in handling nonlinear relationships, high-dimensional input spaces, and missing or noisy data, making them well-suited for hydrological applications 12 , 18 . Random Forest (RF): This ensemble learning method is based on the bagging (bootstrap aggregating) technique. Multiple regression trees are constructed using bootstrapped samples of the training data. Each tree h i (x) makes a prediction, and the final RF output is the average of predictions from all 𝑀 trees: $$\:\widehat{y}=\frac{1}{M}\sum\:_{i=1}^{M}{h}_{i}\left(x\right)$$ 1 RF minimizes the mean squared error (MSE) across trees and uses random subsets of features at each split to reduce correlation among trees. The objective function focuses on reducing variance while maintaining low bias. Feature importance is computed based on the mean decrease in impurity or prediction accuracy across all trees. In this study, the RF model was implemented using the scikit-learn library with hyperparameter tuning (e.g., number of trees, max depth) via grid search 19 , 20 . Gradient Boosting (GB): Gradient Boosting builds models sequentially by fitting each new model to the residual errors of the ensemble prediction so far. The objective function combines a loss function \(\:L(y,\widehat{y})\) (e.g., squared error loss) and a regularization term \(\:{\Omega\:}\left({f}_{t}\right)\) to penalize model complexity: $$\:{L}^{\left(t\right)}=\sum\:L\left({y}_{i},{\widehat{y}}_{i}^{(t-1)}+{f}_{t}\left({x}_{i}\right)\right)+{\Omega\:}\left({f}_{t}\right)$$ 2 In this study, the eXtreme Gradient Boosting (XGBoost) version was used, which improves efficiency and regularization. XGBoost uses second-order Taylor approximation of the loss function for optimization and supports shrinkage (learning rate) and column subsampling to reduce overfitting. The additive model is updated as: $$\:{\widehat{y}}^{\left(t\right)}={\widehat{y}}^{(t-1)}+\eta\:{f}_{t}\left(x\right)$$ 3 where 𝜂 is the learning rate. The model was implemented using the XGBoost library with hyperparameter tuning on parameters such as the number of estimators, learning rate, and maximum tree depth 21 , 22 . Support Vector Machine (SVM): Support Vector Regression (SVR) aims to find a function that approximates the underlying relationship between input features and target values while balancing model complexity and prediction accuracy. The SVR function is expressed as: $$\:f\left(x\right)=⟨w,\varphi\:\left(x\right)⟩+b$$ 4 where: 𝜙 (𝑥) maps the input vector 𝑥 into a high-dimensional feature space, and 𝑤 and 𝑏 are the model parameters. The SVR formulation attempts to minimize the model complexity (by keeping ∥𝑤∥ small) while ensuring that the predictions are within an ε-insensitive margin. The optimization problem is defined as: $$\:{min}_{w,b,{\xi\:}_{i},{\xi\:}_{i}^{*}}=\frac{1}{2}{‖w‖}^{2}+C\sum\:_{i=1}^{n}\left({\xi\:}_{i}+{\xi\:}_{i}^{*}\right)$$ 5 Subject to: $$\:\left\{\begin{array}{c}{y}_{i}-⟨w,\varphi\:\left({x}_{i}\right)⟩-b\le\:\epsilon\:+{\xi\:}_{i}\\\:⟨w,\varphi\:({x}_{i}⟩+b-{y}_{i}\le\:\epsilon\:+{\xi\:}_{i}^{*}\\\:{\xi\:}_{i},{\xi\:}_{i}^{*}\ge\:0\end{array}\right.$$ 6 where: 𝐶 is the regularization parameter that controls the trade-off between the flatness of the function and the amount up to which deviations larger than ε are tolerated. \(\:{\xi\:}_{i},{\xi\:}_{i}^{*}\) are slack variables that allow violations of the ε margin. To model nonlinear relationships, a kernel function \(\:K\left({x}_{i},{x}_{j}\right)=⟨\varphi\:\left({x}_{i}\right),\varphi\:\left({x}_{j}\right)⟩\) is employed. In this study, the Radial Basis Function (RBF) kernel was used: $$\:K\left({x}_{i},{x}_{j}\right)=\text{e}\text{x}\text{p}(-\gamma\:{‖{x}_{i}-{x}_{j}‖}^{2})$$ 7 The RBF kernel enables the model to learn complex, nonlinear patterns in the data. The SVR model was implemented using the scikit-learn library in Python. The key hyperparameters—𝐶 (penalty), 𝜀 (margin width), and 𝛾 (kernel coefficient)—were optimized through a grid search approach combined with cross-validation to ensure robust performance 23 , 24 . All models were trained using the historical time series of meteorological variables (e.g., precipitation, temperature), surface water inflow, and groundwater levels across five hydrogeological zones (Section 2.3.2 ). A combination of grid search and five-fold cross-validation was applied to fine-tune model hyperparameters and ensure robust generalization. Model performance was evaluated using standard statistical metrics, as detailed in Section 2.3.4 . The complete implementation codes, including data preprocessing, model training, and evaluation scripts, are provided in Supplementary Material (S2) for reference and reproducibility. 2.3.2. Input Variables per Zone To account for the spatial heterogeneity of hydrological and anthropogenic influences across the Najafabad aquifer, groundwater level modeling was conducted separately for five hydrogeological distinct zones (Z1–Z5). This zonal approach allows for a more accurate representation of local dynamics by tailoring model inputs to the specific conditions of each area. Table 3 Input Variables for Groundwater Level Change Modeling in the Defined Zones. Zone Inputs Output Meteorological Station Precipitation (P) Temperature (T) River Discharge (RD) Irrigation Volume (IV) Groundwater Abstraction (GwA) Groundwater Level (GwL) Z1 Zefreh ✓ ✓ ✓ - ✓ ✓ Z2 Zefreh ✓ ✓ - ✓ ✓ ✓ Z3 Zefreh ✓ ✓ - ✓ ✓ ✓ Z4 Najafabad ✓ ✓ - ✓ ✓ ✓ Z5 Najafabad ✓ ✓ - - - ✓ ✓: Variable included as model input, –: Variable excluded due to unavailability or irrelevance. The selection of input variables in each zone was guided by both data availability and their relevance to groundwater level fluctuations. The monthly change in groundwater level (GwL) was considered the output variable across all zones, while the predictor variables included monthly precipitation (P), average temperature (T), river discharge (RD), irrigation volume (IV), and groundwater abstraction (GwA). Table 3 provides an overview of the input variables used in each zone. The inclusion or exclusion of specific variables was influenced by each zone’s proximity to surface water resources, irrigation infrastructure, and elevation. For example, zones located downstream of the Nekouabad canal included irrigation volume as a model input, while zones without reliable abstraction data excluded that variable. To further substantiate the selection of input variables and enhance the understanding of interrelationships among the influencing factors, Pearson correlation heatmaps were generated separately for each of the five hydrogeological zones (Fig. 4 ). These heatmaps provide a clear visualization of the strength and direction of linear relationships between groundwater level change and its potential predictors. The results revealed spatial variability in correlation patterns across the zones, emphasizing the importance of localized analysis In Zone Z1, a moderate positive correlation was observed between groundwater level change and groundwater abstraction (r = 0.44), indicating a significant influence of human withdrawals on the aquifer’s behavior. Zone Z2 exhibited a strong negative correlation between temperature and groundwater level change (r = − 0.51), suggesting that increased evapotranspiration during warmer periods may accelerate groundwater depletion. In Zone Z3, the groundwater system appeared to be influenced by a balanced combination of climatic variables (such as precipitation and temperature) and anthropogenic drivers (including irrigation and abstraction). In contrast, Zones Z4 and Z5 showed generally weaker correlations among the variables, potentially indicating more stable hydrological conditions or the presence of other unmeasured influencing factors. These statistical insights, derived from correlation analysis, reinforce the appropriateness of the zonal input configurations and highlight the necessity of incorporating both spatial and causal heterogeneity in the modeling of groundwater dynamics within the Najafabad aquifer. 2.3.3. Model Training and Validation Strategy To ensure robust and generalizable model performance, the available dataset for each hydrogeological zone was split into training and testing subsets. Seventy percent (70%) of the data was used for training the machine learning models, while the remaining 30% was reserved for testing and evaluation. The time series were divided chronologically to preserve the temporal structure and prevent data leakage. To fine-tune the model parameters and enhance prediction accuracy, a five-fold cross-validation (CV) strategy was employed during the training phase. In this method, the training data was partitioned into five equal subsets, and each subset was used once for validation while the remaining four subsets were used for training. The final model performance was averaged across the five folds, ensuring a balanced evaluation of bias and variance. Hyperparameter optimization was performed for each algorithm using a grid search approach, where multiple combinations of key parameters were systematically tested. The optimal hyperparameter set was selected based on minimizing the validation error, specifically the Root Mean Square Error (RMSE). To further prevent overfitting, each model incorporated algorithm-specific regularization mechanisms. For example, the Random Forest model utilized a maximum tree depth constraint and minimum sample split size, the Gradient Boosting model implemented learning rate and subsampling adjustments, and the Support Vector Machine model optimized the kernel function and regularization coefficient (C). This training-validation strategy ensured that the models were both accurate and resilient to noise and overfitting, and it enabled fair comparisons between algorithms across different zones of the Najafabad Plain. 2.3.4. Model Evaluation Metrics To evaluate the predictive accuracy of the developed machine learning models, three standard performance metrics were utilized: Coefficient of Determination (R²), Root Mean Square Error (RMSE), and Mean Absolute Error (MAE). These metrics provide complementary insights into model performance in terms of accuracy, bias, and error distribution. Coefficient of Determination (R²) assesses how well the predicted values replicate the observed values. It is defined as: $$\:{R}^{2}=1-\frac{\sum\:_{i=1}^{n}{\left({O}_{i}-{P}_{i}\right)}^{2}}{\sum\:_{i=1}^{n}{\left({O}_{i}-\stackrel{-}{O}\right)}^{2}}$$ 8 Root Mean Square Error (RMSE) measures the square root of the average squared differences between predicted and observed values: $$\:RMSE=\sqrt{\frac{1}{n}\sum\:_{i=1}^{n}{\left({O}_{i}-{P}_{i}\right)}^{2}}$$ 9 Mean Absolute Error (MAE) quantifies the average magnitude of absolute differences between predictions and actual observations: $$\:MAE=\frac{1}{n}\sum\:_{i=1}^{n}\left|{O}_{i}-{P}_{i}\right|$$ 10 Where O i and P i represent the observed and predicted values, respectively, and \(\:\stackrel{-}{O}\) is the mean of observed values. In summary, a higher R 2 value closer to 1 indicates better model fit, while lower values of RMSE and MAE, approaching zero, reflect higher prediction accuracy and lower error magnitude. These metrics were computed for each model (RF, GB, SVM) and each zone to enable detailed comparison and identify the best-performing algorithm in each hydrogeological context. The selected metrics have been widely applied in hydrological and groundwater modeling studies 25 , 26 . For a detailed overview of the methodological framework including data preprocessing, zone delineation, model training, and evaluation steps, refer to the Methodology Flowchart provided in Supplementary Material (S1). 3. Results and Discussion This section presents the outcomes of the groundwater level modeling across the five defined zones of the Najafabad aquifer using three machine learning algorithms: Random Forest (RF), XGBoost (GB), and Support Vector Machine (SVM). The results are analyzed in terms of model accuracy, spatial variability, and the relative importance of input variables. Furthermore, the performance of the models is compared with findings from previous studies, and implications for regional groundwater management are discussed. The aim is to provide an in-depth understanding of how different factors affect groundwater level fluctuations and how data-driven models can support sustainable water resource planning in semi-arid regions. 3.1. Model Performance per Zone The performance of the three machine learning models—XGBoost, Random Forest (RF), and Support Vector Machine (SVM)—was evaluated across five hydrogeological zones using RMSE, MAE, and R² during both training and testing phases (Table 4 ). The scatter plots of predicted versus observed groundwater levels for each model and zone are illustrated in Fig. 5 , visually reinforcing the numerical results. Table 4 Performance metrics (RMSE, MAE, R²) of XGBoost, RF, and SVM models across hydrogeological zones Z1–Z5 during training and testing. Model Index Training Testing Z1 Z2 Z3 Z4 Z5 Z1 Z2 Z3 Z4 Z5 XGBoost RMSE 0.82 0.81 1.64 1.09 0.35 0.85 2.37 2.87 1.25 0.43 R 2 0.93 0.95 0.94 0.90 0.95 0.91 0.80 0.85 0.76 0.92 MAE 0.21 1.01 1.5 0.38 0.15 0.32 1.8 2.4 0.67 0.21 RF RMSE 0.74 1.43 2.14 1.08 0.68 0.77 2.77 2.93 1.23 0.83 R 2 0.82 0.88 0.86 0.70 0.9 0.78 0.75 0.77 0.59 0.85 MAE 0.42 1 1.7 0.41 0.2 0.72 1.5 2.8 0.66 0.48 SVM RMSE 0.79 1.23 2.34 1.13 0.23 0.87 2.05 4.16 1.26 0.44 R 2 0.91 0.94 0.92 0.85 0. 97 0.86 0.82 0.80 0.70 0.93 MAE 0.40 1.4 2 0.51 0.12 0.51 2.5 3.2 0.75 0.19 Overall, XGBoost outperformed the other models in most zones, demonstrating both accuracy and stability. It achieved the highest R² values during testing in Z1 (0.91), Z3 (0.85), Z4 (0.76), and Z5 (0.92), with relatively low RMSEs. For instance, in Z5, XGBoost yielded a remarkably low RMSE of 0.43 m and MAE of 0.21 m, while maintaining a high R² of 0.92, indicating excellent generalization and prediction capability in this zone. The SVM model, while slightly more variable, performed surprisingly well in Z5 (R² = 0.93, RMSE = 0.44 m, MAE = 0.19 m), and also showed competitive accuracy in Z1 and Z2. However, its performance dropped significantly in Z3, where it recorded the highest RMSE (4.16 m) and MAE (3.2 m) among all models and zones during testing. This can also be observed in Fig. 5 , where scatter points in Z3-SVM show a noticeable deviation from the 1:1 line, highlighting poor prediction consistency. These findings are consistent with those of Ibrahem Ahmed Osman, et al. 16 and Vafadar, et al. 14 , who also reported the superior performance of XGBoost in groundwater level prediction tasks, particularly in heterogeneous and data-limited environments. Random Forest (RF) offered moderate results across the zones, with its best performance in Z5 (R² = 0.85, RMSE = 0.83 m, MAE = 0.48 m). However, in Z2 and Z3, its testing RMSE exceeded 2.7 m, and R² dropped to 0.75 and 0.77, respectively. These outcomes align with the findings of Abbas, et al. 27 , who recognized RF as a dependable model for groundwater-related predictions, while also highlighting its smoothing effect on sudden variations due to the ensemble averaging process. The scatter plots in these zones reflect the relatively dispersed nature of the predictions, especially when compared with XGBoost’s tighter clustering along the 1:1 line. A notable observation is that Zone 5 (Z5), despite its high elevation and potentially complex hydrogeological features, was modeled more successfully than Zones 2 and 3. All models achieved R² values above 0.85 in Z5 during testing, with particularly low MAEs and RMSEs, suggesting more consistent and learnable patterns in the groundwater data of this zone. In contrast, Z3 posed a challenge for all models—particularly for SVM—likely due to higher variability or the presence of outliers. This is evident both in statistical metrics and in the wider spread of data points away from the 1:1 line in Fig. 5 . These results are consistent with Dong, et al. 28 , who emphasized the importance of appropriate kernel selection and hyperparameter tuning when using SVM in regions with high variability. Furthermore, Zarafshan, et al. 15 found that although SVM and its variants like SVR can perform acceptably under various groundwater regimes, their performance is highly sensitive to data stability and spatial heterogeneity. In summary, XGBoost demonstrated the most balanced and robust performance across all zones, while SVM and RF showed strong localized performance, especially in zones with less variability. These findings emphasize the importance of considering local hydrogeological characteristics in selecting the optimal modeling approach for groundwater prediction. 3.2. Importance of Input Variables To assess the influence of various input variables on groundwater level prediction, the relative importance of each variable was evaluated using three machine learning algorithms: Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). Figure 6 presents the comparative importance values across the five hydrogeological zones (Z1 to Z5). Across all zones, precipitation (P) and temperature (T) emerged as the most influential variables, though their relative dominance varied by zone and algorithm. In Zone 1, temperature had a slightly higher influence than precipitation, especially under RF and XGBoost, while SVM showed nearly equal importance for recharge depth (RD) and T. In Zone 2, SVM gave the highest importance to temperature, exceeding 65%, while the other algorithms showed a more balanced influence between temperature and precipitation. Zone 3 showed similar trends, with temperature again dominating in all models, particularly in RF and XGBoost, while groundwater abstraction (GwA) had minimal contribution. In Zone 4, XGBoost highlighted temperature as the most dominant variable (~ 47%), followed by precipitation and recharge depth, whereas SVM and RF placed more emphasis on precipitation. Interestingly, Zone 5, which includes only two variables (P and T), demonstrated the clearest consensus among models: temperature was consistently the most influential, with values close to or above 65% in all models. Overall, temperature and precipitation consistently ranked as the top predictors, aligning with the hydro-climatic nature of groundwater dynamics in the region. In contrast, groundwater abstraction and irrigation volume showed lower relative importance in most zones, possibly due to weaker correlations or limited temporal resolution. These findings underline the spatial variability in predictor importance, highlighting the need to tailor groundwater modeling approaches to regional characteristics. Our results align closely with those of Nourani, et al. 29 and Costantini, et al. 30 , who emphasized the predominant role of climatic drivers—particularly precipitation and temperature—over anthropogenic factors in arid and semi-arid aquifer systems. This consistency reinforces the conclusion that, in data-scarce regions like Najafabad, the predictive power of climatic variables outweighs that of human-induced parameters such as irrigation and pumping, whose impact may be masked by coarse temporal or spatial data resolution. In conclusion, this comparative analysis reaffirms the robustness of ensemble learning methods—particularly XGBoost—in predicting groundwater levels under complex hydrogeological conditions. It also underscores the importance of zone-based model calibration and input sensitivity analysis to account for local hydro-climatic variability and ensure the development of context-appropriate forecasting tools. 3.3. Model Robustness and Generalization Robustness and generalization are critical attributes of any predictive model, particularly in groundwater studies where data quality, completeness, and variability can significantly influence model outcomes. In the present study, all three machine learning models—Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost)—were evaluated across five hydrogeological zones with varying characteristics to assess their stability and transferability (Table 5 ). Table 5 Comparison of model performance metrics in training and testing phases across all zones. Model Phase Mean R² Mean RMSE ΔR² (Train - Test) ΔRMSE (Test - Train) XGBoost Training 0.934 0.942 0.086 0.612 Testing 0.848 1.554 RF Training 0.832 1.214 0.084 0.492 Testing 0.748 1.706 SVM Training 0.918 1.144 0.096 0.612 Testing 0.822 1.756 Among the models, XGBoost demonstrated the highest predictive performance during both training (mean R² = 0.934, RMSE = 0.942) and testing (mean R² = 0.848, RMSE = 1.554), with a moderate drop in performance (ΔR² = 0.086, ΔRMSE = 0.612). This relatively small gap highlights XGBoost’s robust generalization ability, especially considering its ability to capture nonlinear groundwater dynamics in both stable (e.g., Z1, Z5) and hydrologically complex regions (e.g., Z3). This robustness can be attributed to its embedded regularization mechanisms, capacity to manage multicollinearity, and iterative bias-variance reduction during training. The Random Forest (RF) model achieved lower overall predictive accuracy (R² = 0.832 in training and 0.748 in testing) but exhibited the smallest RMSE difference (ΔRMSE = 0.492), indicating good stability and a low tendency to overfit. While RF may underperform in zones with sharp hydrological fluctuations, it is relatively resilient to noisy or incomplete data, making it useful for preliminary analysis or areas with limited observations. The Support Vector Machine (SVM) model produced high accuracy during training (R² = 0.918) but exhibited the largest generalization gap (ΔR² = 0.096), indicating greater sensitivity to overfitting, especially in zones with high groundwater variability. Its testing performance dropped more sharply than the other models, which can be attributed to SVM's reliance on kernel function selection and sensitivity to feature distribution. These limitations necessitate careful preprocessing, normalization, and kernel tuning for improved robustness. These comparative results reinforce the importance of selecting models that balance accuracy with generalization, especially in regions such as the Najafabad plain where hydrogeological heterogeneity and climatic variability present significant modeling challenges. Among the three models, XGBoost emerged as the most robust and generalizable, offering both predictive strength and stability across diverse zones. Future studies should consider hybrid modeling strategies that integrate multiple machine learning models or external data sources (e.g., remote sensing, irrigation patterns, or socio-economic data) to enhance predictive robustness and transferability in similar complex groundwater systems. 3.4. Implications for Groundwater Management and Future Directions The findings of this study hold significant implications for groundwater management in data-scarce and hydrogeological diverse regions such as the Najafabad plain. By evaluating the performance of three machine learning algorithms—XGBoost, Random Forest (RF), and Support Vector Machine (SVM)—across five distinct zones, the research highlights the potential of data-driven approaches to enhance decision-making under conditions of uncertainty and limited observational data. The strong performance of XGBoost, particularly in terms of generalization and robustness, demonstrates its suitability for capturing complex and nonlinear groundwater dynamics. This makes it a valuable tool for supporting real-time groundwater level forecasting, resource allocation, and early warning systems. Moreover, the ability to quantify feature importance offers a transparent mechanism for identifying key drivers of groundwater fluctuations, such as precipitation and temperature, which can aid in the design of adaptive climate-resilient policies. From a management perspective, the zone-specific variability in model accuracy underscores the need for localized modeling frameworks. Zones like Z3, which exhibited more complex and unstable groundwater behavior, demand tailored calibration efforts and possibly the integration of external hydrological or anthropogenic indicators (e.g., land use, irrigation practices, pumping intensity) to improve predictive reliability. However, several limitations of this study should be acknowledged. First, the accuracy of model predictions is inherently tied to the quality and resolution of input data. In this study, limited availability of high-resolution spatiotemporal data on irrigation and pumping rates may have influenced model performance, particularly in zones with high human activity. Second, while the machine learning models demonstrated strong predictive ability, they operate as black-box systems with limited capacity for causal interpretation, which may hinder their adoption in regulatory or policy contexts that require process transparency. Future research should explore hybrid modeling approaches that combine the strengths of data-driven methods with physically-based models to enhance both prediction and interpretability. Additionally, the integration of remotely sensed data, socio-economic indicators, and long-term groundwater abstraction records could help to further refine model inputs and improve spatial transferability. Developing dynamic models that incorporate feedback mechanisms from land use changes and climate scenarios could also strengthen the applicability of machine learning in strategic groundwater resource planning. In summary, this study affirms the promise of machine learning models—particularly ensemble-based approaches like XGBoost—for operational groundwater management. With continued refinement and integration of interdisciplinary data sources, these models can play a vital role in addressing the growing challenges of groundwater sustainability in semi-arid and arid environments. 4. Conclusion This study evaluated the predictive performance of three machine learning algorithms—Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Support Vector Machine (SVM)—in forecasting groundwater level fluctuations across five hydrogeological zones in the Najafabad plain, a semi-arid region in central Iran. Using climatic, hydrological, and anthropogenic input variables, the models were trained and tested to assess their accuracy, robustness, and generalization under varying groundwater conditions. Among the models, XGBoost consistently outperformed the others, achieving the highest average coefficient of determination during testing (mean R² = 0.848) and the lowest root mean square error (RMSE = 1.554). It also demonstrated a moderate performance gap between training and testing phases (ΔR² = 0.086), indicating strong generalization ability and limited overfitting. This superior performance is attributed to its boosting mechanism, embedded regularization, and capacity to capture nonlinear groundwater dynamics even in complex zones (e.g., Z3). Random Forest (RF) showed moderate accuracy with mean R² values of 0.832 in training and 0.748 in testing, and a relatively small RMSE difference (ΔRMSE = 0.492), reflecting stable predictions with low overfitting risk. However, its performance declined notably in zones with highly variable groundwater levels. Support Vector Machine (SVM) achieved high training accuracy (mean R² = 0.918) but exhibited the largest generalization gap (ΔR² = 0.096) and higher testing RMSE (1.756), revealing sensitivity to data scaling and kernel selection, particularly in hydrologically volatile zones. These results highlight the trade-off between accuracy and robustness across models and emphasize the suitability of XGBoost for groundwater forecasting in semi-arid, hydrogeological heterogeneous regions like Najafabad. The ability of XGBoost to maintain high accuracy with moderate overfitting risk makes it an excellent tool for operational groundwater monitoring, resource allocation, and early warning systems. Nonetheless, limitations such as the lack of high-resolution anthropogenic data (e.g., pumping and irrigation rates) likely constrained model performance in certain zones. Furthermore, the “black-box” nature of these models limits direct causal interpretation, which can restrict their policy adoption. Future research should explore hybrid approaches combining data-driven models with physically-based groundwater models, and integrate diverse data sources such as remote sensing and socio-economic indicators. Additionally, dynamic modeling that incorporates land use and climate change scenarios could further improve forecasting reliability and sustainability in groundwater management. In conclusion, this study confirms that ensemble machine learning methods—particularly XGBoost—offer powerful, flexible, and scalable solutions for groundwater level prediction in complex semi-arid environments, supporting sustainable water resource management amid increasing climatic and anthropogenic pressures. Declarations All authors have read, understood, and have complied as applicable with the statement on "Ethical responsibilities of Authors" as found in the Instructions for Authors. Competing Interests The authors declare that they have no competing interests. Financial Disclosure statement The author(s) received no specific funding for this work. Author contributions: All authors contributed to the conceptualization and design of the study. Material preparation, data collection, and analysis were performed by Sahar Davari, Saeid Eslamian, and Mohammad Jamali. Code development was carried out by Mohammad Jamali and Sahar Davari. The first draft of the manuscript was written by Saeid Eslamian and Mohammad Jamali. Review and editing were conducted by Saeid Eslamian and Hamid Reza Safavi. All authors read and approved the final manuscript and provided comments on previous versions. Data availability The dataset generated and/or analyzed during the current study is available from the corresponding author upon reasonable request. The Python codes used for model development and evaluation are publicly available at the following GitHub repository: https://github.com/mohammadj74/Groundwater-ML-Najafabad Ethics approval This study is original research and has not been previously published or submitted elsewhere. The research does not involve any experiments on animals or human subjects. References Famiglietti, J. S. The global groundwater crisis. Nature Climate Change 4 , 945-948, doi:10.1038/nclimate2425 (2014). Jamali, M., Yazdian, H., Bahman, G. & Eslamian, S. Water agriculture nexus a system dynamics approach for the next three decades. Scientific Reports 15 , 5946, doi:10.1038/s41598-025-90728-3 (2025). Wada, Y. et al. Global depletion of groundwater resources. Geophysical Research Letters 37 , doi:https://doi.org/10.1029/2010GL044571 (2010). Harbaugh, A. W. MODFLOW-2005, the US Geological Survey modular ground-water model: the ground-water flow process . Vol. 6 (US Department of the Interior, US Geological Survey Reston, VA, USA, 2005). Jahanshahi, A. et al. Dependence of rainfall-runoff model transferability on climate conditions in Iran. Hydrological Sciences Journal 67 , 564-587, doi:10.1080/02626667.2022.2030867 (2022). Xu, T. & Liang, F. Machine learning for hydrologic sciences: An introductory overview. WIREs Water 8 , e1533, doi:https://doi.org/10.1002/wat2.1533 (2021). Lange, H. & Sippel, S. in Forest-Water Interactions (eds Delphis F. Levia et al. ) 233-257 (Springer International Publishing, 2020). Papacharalampous, G. & Tyralis, H. A review of machine learning concepts and methods for addressing challenges in probabilistic hydrological post-processing and forecasting. Frontiers in Water Volume 4 - 2022 , doi:10.3389/frwa.2022.961954 (2022). Eslamian, S. & Eslamian, F. Handbook of hydroinformatics: volume i: classic soft-computing techniques . (Elsevier, 2022). Tao, H. et al. Groundwater level prediction using machine learning models: A comprehensive review. Neurocomputing 489 , 271-308, doi:https://doi.org/10.1016/j.neucom.2022.03.014 (2022). Sahoo, S., Russo, T. A., Elliott, J. & Foster, I. Machine learning algorithms for modeling groundwater level changes in agricultural regions of the U.S. Water Resources Research 53 , 3878-3895, doi:https://doi.org/10.1002/2016WR019933 (2017). Pham, Q. B. et al. Groundwater level prediction using machine learning algorithms in a drought-prone area. Neural Computing and Applications 34 , 10751-10773, doi:10.1007/s00521-022-07009-7 (2022). Raturi, M., Khare, D. & Patidar, N. in Applications of Machine Learning in Hydroclimatology (eds Roshan Srivastav & Purna C. Nayak) 57-71 (Springer Nature Switzerland, 2025). Vafadar, S., Rahimzadegan, M. & Asadi, R. Evaluating the performance of machine learning methods and Geographic Information System (GIS) in identifying groundwater potential zones in Tehran-Karaj plain, Iran. Journal of Hydrology 624 , 129952, doi:https://doi.org/10.1016/j.jhydrol.2023.129952 (2023). Zarafshan, P. et al. Comparison of machine learning models for predicting groundwater level, case study: Najafabad region. Acta Geophysica 71 , 1817-1830, doi:10.1007/s11600-022-00948-8 (2023). Ibrahem Ahmed Osman, A., Najah Ahmed, A., Chow, M. F., Feng Huang, Y. & El-Shafie, A. Extreme gradient boosting (Xgboost) model to predict the groundwater levels in Selangor Malaysia. Ain Shams Engineering Journal 12 , 1545-1556, doi:https://doi.org/10.1016/j.asej.2020.11.011 (2021). Naderi, M. & Hajiketabi, M. Quantification of normal and sustainable management practices for groundwater resources: example of the arid Najafabad alluvial aquifer in Isfahan Province, Iran. Hydrogeology Journal 31 , 195-218, doi:10.1007/s10040-023-02596-8 (2023). Karimi, H. et al. Enhancing groundwater quality prediction through ensemble machine learning techniques. Environmental Monitoring and Assessment 197 , 21, doi:10.1007/s10661-024-13506-0 (2024). Breiman, L. Random Forests. Machine Learning 45 , 5-32, doi:10.1023/A:1010933404324 (2001). Pedregosa, F. et al. Scikit-learn: Machine learning in Python. the Journal of machine Learning research 12 , 2825-2830 (2011). Chen, T. & Guestrin, C. in Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining 785–794 (Association for Computing Machinery, San Francisco, California, USA, 2016). Kashani, A. & Safavi, H. R. Assessing groundwater drought in Iran using GRACE data and machine learning. Scientific Reports 15 , 14671, doi:10.1038/s41598-025-99342-9 (2025). Vapnik, V. The nature of statistical learning theory . (Springer science & business media, 2013). Smola, A. J. & Schölkopf, B. A tutorial on support vector regression. Statistics and Computing 14 , 199-222, doi:10.1023/B:STCO.0000035301.49549.88 (2004). Yadav, A., Raj, A. & Yadav, B. Enhancing local-scale groundwater quality predictions using advanced machine learning approaches. Journal of Environmental Management 370 , 122903, doi:https://doi.org/10.1016/j.jenvman.2024.122903 (2024). Thakur, S. & Karmakar, S. A Comparative Analysis of ANN, LSTM and Hybrid PSO-LSTM Algorithms for Groundwater Level Prediction. Transactions of the Indian National Academy of Engineering 10 , 101-108, doi:10.1007/s41403-024-00505-3 (2025). Abbas, F. et al. (Water 2024, 16, 941., 2024). Dong, J., Tsai, G. & Olivares, C. I. Prediction of 35 Target Per- and Polyfluoroalkyl Substances (PFASs) in California Groundwater Using Multilabel Semisupervised Machine Learning. ACS ES&T Water 4 , 969-981, doi:10.1021/acsestwater.3c00134 (2024). Nourani, V., Ghareh Tapeh, A. H., Khodkar, K. & Huang, J. J. Assessing long-term climate change impact on spatiotemporal changes of groundwater level using autoregressive-based and ensemble machine learning models. Journal of Environmental Management 336 , 117653, doi:https://doi.org/10.1016/j.jenvman.2023.117653 (2023). Costantini, M., Colin, J. & Decharme, B. Projected Climate‐Driven Changes of Water Table Depth in the World's Major Groundwater Basins. Earth's Future 11 , e2022EF003068 (2023). Additional Declarations No competing interests reported. Supplementary Files SupplementaryMaterial.pdf Cite Share Download PDF Status: Published Journal Publication published 10 Dec, 2025 Read the published version in Scientific Reports → Version 1 posted Editorial decision: Revision requested 04 Aug, 2025 Reviews received at journal 22 Jul, 2025 Reviewers agreed at journal 21 Jul, 2025 Reviews received at journal 20 Jul, 2025 Reviewers agreed at journal 16 Jul, 2025 Reviewers agreed at journal 16 Jul, 2025 Reviewers agreed at journal 16 Jul, 2025 Reviewers invited by journal 16 Jul, 2025 Editor assigned by journal 09 Jul, 2025 Editor invited by journal 08 Jul, 2025 Submission checks completed at journal 03 Jul, 2025 First submitted to journal 03 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-6992628","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Article","associatedPublications":[],"authors":[{"id":486993296,"identity":"e4d506a2-d91e-4b79-a6bb-9bbf680dd574","order_by":0,"name":"Sahar Davari","email":"","orcid":"","institution":"Isfahan University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Sahar","middleName":"","lastName":"Davari","suffix":""},{"id":486993297,"identity":"94be0b31-326d-43f9-bde5-7950d6aedb34","order_by":1,"name":"Saeid Eslamian","email":"","orcid":"","institution":"Isfahan University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Saeid","middleName":"","lastName":"Eslamian","suffix":""},{"id":486993299,"identity":"3728ad4d-f104-4709-acac-5499c8433293","order_by":2,"name":"Mohammad Jamali","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABJ0lEQVRIiWNgGAWjYFACHjDJ2MBwAEgVSMiBeAcYYIKEtRhIGEO0JBClBQQMGBIhjATcztJt7z346WaOnWwD4+GnG34YWKT39y9gPMz7g0Gev4G57QMWLWZnziVL525LNm5gOGZ2s8dAInfGjQcMh3kSGAxnHGBsnoFNy40cA6AWZqB7Dpjd4AFqabhxAKyFcQMDYzM2hwG1GP/O3VYP1HL8280/BhLp8lAt9ni0mAFtOQzUcsbsNtCWBIPzDWAtiTi1nDljZp277bhxG8OZstsyBhKGG28wNhyckyaRPOMwDi3He4xv526rlu2XOL7t5puKOnm584cPf3hjY2Pb397+GFdAgwGbxAEoSwIcNRIMDMx4NQABfwOMcQCPqlEwCkbBKBiJAABnjGsguwqRaAAAAABJRU5ErkJggg==","orcid":"","institution":"Isfahan University of Technology","correspondingAuthor":true,"prefix":"","firstName":"Mohammad","middleName":"","lastName":"Jamali","suffix":""},{"id":486993300,"identity":"b20f0269-5dfc-466b-9dc1-7a3d2cad1968","order_by":3,"name":"Hamid Reza Safavi","email":"","orcid":"","institution":"Isfahan University of Technology","correspondingAuthor":false,"prefix":"","firstName":"Hamid","middleName":"Reza","lastName":"Safavi","suffix":""}],"badges":[],"createdAt":"2025-06-27 14:38:27","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-6992628/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-6992628/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1038/s41598-025-32376-1","type":"published","date":"2025-12-10T15:57:17+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":87026623,"identity":"1a188993-9ff4-4b41-8c3b-3ec64e0f44af","added_by":"auto","created_at":"2025-07-18 12:12:23","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":9333953,"visible":true,"origin":"","legend":"\u003cp\u003eThe study area, including the location of Isfahan Province in Iran, the location of the Najafabad Plain within Isfahan Province, the zoning of the Najafabad Plain, the locations of the utilized stations, and the positions of existing wells.\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/80034df3ad150c1ca6084e7c.jpeg"},{"id":87026619,"identity":"246c60d1-ca7b-459b-ae57-d370e133e98c","added_by":"auto","created_at":"2025-07-18 12:12:23","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":3059276,"visible":true,"origin":"","legend":"\u003cp\u003eMonthly distribution of groundwater level fluctuations (m) across five observation zones (Z1 to Z5) in the Najafabad Plain.\u003c/p\u003e","description":"","filename":"floatimage2.png","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/27e65546486d8e534e66d28b.png"},{"id":87027065,"identity":"3dfa3a0d-d6ec-45fe-a176-d64be0a56254","added_by":"auto","created_at":"2025-07-18 12:20:23","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":2231432,"visible":true,"origin":"","legend":"\u003cp\u003eSeasonal distribution of groundwater level (m) across five observation zones (Z1 to Z5) in the Najafabad Plain.\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/c9f9543ef91c6f3b04e28869.png"},{"id":87026625,"identity":"c5f8dd38-c3d2-4574-b0a0-b90a3d6a8569","added_by":"auto","created_at":"2025-07-18 12:12:23","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1374051,"visible":true,"origin":"","legend":"\u003cp\u003ePearson correlation heatmaps between groundwater level, groundwater level decline and input variables across hydrogeological zones Z1 to Z5.\u003c/p\u003e","description":"","filename":"floatimage4.png","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/9842d26170f5402c01422329.png"},{"id":87027069,"identity":"4a4b9ff6-fb84-4942-a047-5556552c10d0","added_by":"auto","created_at":"2025-07-18 12:20:23","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":5316706,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of observed vs. predicted groundwater levels for SVM, RF, and XGBoost models across zones Z1–Z5, with the 1:1 line indicating perfect agreement.\u003c/p\u003e","description":"","filename":"floatimage5.png","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/bffe281c5e618bfa57925643.png"},{"id":87027066,"identity":"9faf04d6-e633-40ac-9be8-c3aeb7e54733","added_by":"auto","created_at":"2025-07-18 12:20:23","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":114067,"visible":true,"origin":"","legend":"\u003cp\u003eRelative Importance of Input Variables Across Zones (Z1–Z5)\u003c/p\u003e","description":"","filename":"floatimage6.png","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/f18af501be660348e496e539.png"},{"id":98245046,"identity":"b87eeca8-2396-468f-a201-a29811c7d33d","added_by":"auto","created_at":"2025-12-15 16:16:25","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":21965293,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/bbcad2ca-cdb1-4800-8d22-ca3add74fd49.pdf"},{"id":87027926,"identity":"38e1ac3b-7c28-44f4-a72b-07b6611debe5","added_by":"auto","created_at":"2025-07-18 12:28:23","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":1467823,"visible":true,"origin":"","legend":"","description":"","filename":"SupplementaryMaterial.pdf","url":"https://assets-eu.researchsquare.com/files/rs-6992628/v1/e084c64313d77c4b55c96807.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Application of Machine Learning Algorithms for Groundwater Level Prediction in the Najafabad Plain","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eGroundwater plays a vital role in sustaining agricultural productivity, urban development, and ecosystem services, particularly in arid and semi-arid regions such as central Iran. In these environments, where surface water resources are limited or highly variable, aquifers serve as the primary source of freshwater for both irrigation and domestic use. However, overexploitation of groundwater due to increasing population, agricultural expansion, and declining precipitation under climate change has led to alarming consequences, including groundwater depletion, land subsidence, reduced water quality, and diminished long-term water security \u003csup\u003e\u003cspan additionalcitationids=\"CR2\" citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e\u003c/sup\u003e. This escalating stress highlights the urgent need for robust groundwater monitoring and predictive tools to support sustainable resource management.\u003c/p\u003e\u003cp\u003eTraditional physically-based models such as MODFLOW have long been used for simulating groundwater flow and levels \u003csup\u003e\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e\u003c/sup\u003e. While these models offer valuable process-based insights, their application is often constrained by the requirement for extensive input data (e.g., aquifer geometry, hydraulic conductivity, recharge rates), which are frequently sparse or uncertain\u0026mdash;particularly in data-scarce regions like Iran \u003csup\u003e\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e\u003c/sup\u003e. Additionally, these models can be computationally intensive and difficult to calibrate under dynamically changing human-water interactions.\u003c/p\u003e\u003cp\u003eIn response to these challenges, data-driven models\u0026mdash;especially machine learning (ML) algorithms\u0026mdash;have emerged as powerful alternatives for groundwater prediction. Algorithms such as Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost) have shown promising results in capturing nonlinear relationships between groundwater levels and influencing factors such as precipitation, temperature, and human activities \u003csup\u003e\u003cspan additionalcitationids=\"CR7\" citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e\u003c/sup\u003e. These models offer greater flexibility, faster computation, and reduced dependency on physical assumptions, making them suitable for complex or data-limited environments \u003csup\u003e\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eSeveral studies have successfully applied ML methods to groundwater level forecasting in various regions of Iran and beyond \u003csup\u003e\u003cspan additionalcitationids=\"CR11 CR12\" citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e\u0026ndash;\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e\u003c/sup\u003e. For example, Vafadar, et al. \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e applied RF, SVM, and XGBoost for groundwater potential mapping in the Tehran and Karaj plains and found XGBoost to outperform the other models. Similarly, Zarafshan, et al. \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e evaluated various ML models for groundwater prediction in the Najafabad region and confirmed the potential of support vector-based approaches. Ibrahem Ahmed Osman, et al. \u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e, who applied XGBoost to predict groundwater depth in Selangor, Malaysia, and reported high predictive accuracy. However, many of these studies focused on either a single model or an aggregated spatial scale, often overlooking the spatial heterogeneity of hydrogeological conditions and the importance of local calibration.\u003c/p\u003e\u003cp\u003eThis study addresses this gap by conducting a comprehensive comparative analysis of RF, SVM, and XGBoost for predicting monthly groundwater levels in the Najafabad plain, Isfahan province. A novel contribution of this work is the spatial disaggregation of the study area into five distinct hydrogeological zones, enabling more detailed insight into the model performance under varying aquifer conditions. This zonal framework enhances the interpretability and applicability of the models for groundwater management in heterogeneous settings.\u003c/p\u003e\u003cp\u003eThe innovation of this research is threefold:\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eIt introduces a zone-based modeling approach to assess the spatial variability of ML model performance, rarely considered in previous Iranian studies.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIt compares three state-of-the-art ML models under consistent settings, evaluating their accuracy, robustness, and generalization ability.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eIt provides a feature importance analysis to identify the dominant climatic and anthropogenic drivers of groundwater fluctuations across zones.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThe outcomes of this research are expected to support targeted groundwater management strategies, inform adaptive policy-making, and contribute to the growing body of literature on machine learning applications in groundwater science. However, the study is not without limitations. The models rely solely on historical time series data, which may not fully capture future changes driven by unexpected anthropogenic or climatic shifts. Additionally, the spatial resolution of input data and the availability of high-quality monitoring records may influence the reliability of predictions, particularly in areas with sparse observational coverage. Finally, while machine learning models are powerful in capturing correlations, they do not inherently provide mechanistic insights into subsurface hydrological processes, and thus should be complemented by physical understanding for integrated groundwater management.\u003c/p\u003e"},{"header":"2. Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e\u003ch2\u003e2.1. Study area\u003c/h2\u003e\u003cp\u003eThe Najafabad Plain, located in Isfahan Province in central Iran, is a key agricultural region within the Zayandeh Rud Basin. It spans approximately 1,712 km\u0026sup2;, comprising a central alluvial aquifer of about 940.9 km\u0026sup2;, with the remaining 772.5 km\u0026sup2; consisting of surrounding highlands. Geographically, the plain lies between longitudes 50\u0026deg;57\u0026prime; and 51\u0026deg;44\u0026prime;26\u0026Prime; E and latitudes 32\u0026deg;20\u0026prime;13\u0026Prime; to 32\u0026deg;49\u0026prime;21\u0026Prime; N (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). The aquifer is hydrologically isolated from adjacent groundwater systems and functions as the principal reservoir for subsurface water in the area.\u003c/p\u003e\u003cp\u003eThe elevation of the plain ranges from roughly 2,906 meters above sea level in the northern highlands to approximately 1,538 meters in the southern lowlands, near the Zayandeh Rud River, which provides intermittent surface flow. The region experiences a semi-arid climate, with mean annual precipitation ranging from 158 to 172 mm, while the potential evapotranspiration exceeds 1,500 mm. The average annual temperature is approximately 15\u0026deg;C. Water usage in the region is dominated by agriculture, accounting for about 93.3% of the total water consumption. Domestic and industrial uses represent around 4.5% and 2.2%, respectively. Groundwater supplies nearly 72% of the total water demand, primarily through wells. In total, 15,933 groundwater extraction points have been identified, including 15,753 wells, 151 qanats, and 29 springs. Among them, 15,673 wells are located within the alluvial aquifer, collectively extracting approximately 881\u0026nbsp;million cubic meters of water annually \u003csup\u003e\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eIn recent years, declining precipitation and increasing water demand have exerted considerable pressure on groundwater resources. Although the development of new wells is legally restricted, unauthorized wells\u0026mdash;frequently reported by the Isfahan Regional Water Authority\u0026mdash;pose a significant challenge to sustainable groundwater management in the plain.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec4\" class=\"Section2\"\u003e\u003ch2\u003e2.2. Data Collection and Preprocessing\u003c/h2\u003e\u003cdiv id=\"Sec5\" class=\"Section3\"\u003e\u003ch2\u003e2.2.1. Meteorological, Surface Water, and Groundwater Data\u003c/h2\u003e\u003cp\u003eTo investigate groundwater level fluctuations in the Najafabad Plain, comprehensive hydrological and meteorological datasets were collected and preprocessed. Meteorological variables, including precipitation and temperature, were obtained from the Najafabad synoptic station, which served as the primary source for representing regional climatic conditions. Surface water data were acquired from three active hydrometric stations\u0026mdash;Lenj, Mousian, and Zefreh\u0026mdash;located along the Zayandeh Rud River, the principal surface water source in the region. These stations provide valuable insights into river flow dynamics and their interactions with the underlying groundwater system. The total annual surface inflow into the plain, including contributions from the Lenjanat, Mahyar-e-Shomali, and Karon sub-basins, was estimated at approximately 557.4\u0026nbsp;million cubic meters.\u003c/p\u003e\u003cp\u003eThese representative stations (see Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e) were selected based on their strategic locations and the availability of long-term, reliable data, ensuring adequate spatial and temporal coverage of the key hydrological and climatic variables influencing groundwater levels across the study area. During the study period (2011\u0026ndash;2021), the average annual precipitation was approximately 195 mm in the upland regions and 153.7 mm in the plain. March was identified as the wettest month, with average rainfall values of 31.6 mm in the uplands and 29.3 mm in the plain. Mean annual temperatures were estimated at around 13.8\u0026deg;C in the uplands and 15.4\u0026deg;C in the plain.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eSelected Stations for the Research\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"7\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eStation Name\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eStation Type\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eLongitude (\u0026deg;E)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eLatitude (\u0026deg;N)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eElevation (m)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eStart Year\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eEnd Year\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eNajafabad\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eSynoptic\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e51.389\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32.604\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1634\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1970\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2022\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZefreh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHydrometric\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e51.4986\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32.5019\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1623\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1965\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2022\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eMousian\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHydrometric\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e51.5261\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e32.5772\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1600\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1995\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2022\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLenej\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eHydrometric\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e51.5575\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e39.3236\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e\u003cp\u003e1646\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\"\u003e\u003cp\u003e1980\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c7\"\u003e\u003cp\u003e2022\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eGroundwater level data were collected from 53 observation wells distributed across the Najafabad Plain, corresponding to a monitoring density of approximately 1.4 wells per 25 km\u0026sup2; (see Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). These wells provided continuous records of groundwater level fluctuations throughout the aquifer system, which extends beneath approximately 87% of the study area and constitutes the primary source of groundwater extraction.\u003c/p\u003e\u003cp\u003eFor spatial analysis, the study area was divided into two major zones: the uplands and the plain. The uplands, primarily located in the northern and western parts of the region, cover an area of about 679.2 km\u0026sup2; and are predominantly composed of Jurassic and Cretaceous limestone formations. In contrast, the plain\u0026mdash;spanning approximately 932.2 km\u0026sup2;\u0026mdash;is characterized by Quaternary alluvial deposits with variable thicknesses and hydraulic conductivities. This zoning approach enabled differentiation of hydroclimatic and geological conditions across the region, thereby enhancing the accuracy of groundwater level modeling in response to meteorological and surface water drivers.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec6\" class=\"Section3\"\u003e\u003ch2\u003e2.2.2. Outlier Detection and Treatment\u003c/h2\u003e\u003cp\u003eTo enhance the quality and reliability of the dataset, outlier detection was performed using the boxplot method\u0026mdash;a widely adopted graphical technique in hydrological studies for identifying anomalous values. This approach effectively visualizes the distribution and dispersion of data, highlighting values that deviate significantly from the typical range. Given that the primary objective of this study is not the analysis of extreme events, the identified outliers\u0026mdash;presumed to be due to measurement or recording errors\u0026mdash;were excluded from further analysis. Subsequently, the resulting gaps in the time series were filled using interpolation techniques, ensuring data continuity and temporal consistency across all variables.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec7\" class=\"Section3\"\u003e\u003ch2\u003e2.2.3. Regional and Aquifer Zoning\u003c/h2\u003e\u003cp\u003eThe Najafabad aquifer, like many complex groundwater systems, exhibits significant spatial heterogeneity due to the combined influence of natural and anthropogenic factors such as groundwater abstraction, precipitation, riverbed infiltration, and temperature variability. Owing to the physiographic diversity of the plain, these factors impact different regions of the aquifer to varying degrees. For example, in the western parts of the aquifer\u0026mdash;located at a considerable distance from the Zayandeh-Rud River\u0026mdash;the influence of river discharge on groundwater levels is minimal to negligible. To improve the spatial resolution and accuracy of groundwater level modeling, the aquifer was subdivided into five distinct zones. This zoning was based on a conceptual understanding of the system, as well as the physical and geographical characteristics of the area, particularly in relation to the presence of the river and irrigation canals. The spatial layout and boundaries of these zones are depicted in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e, and their key characteristics are summarized in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eDescription of groundwater zones based on dominant influencing factors in the Najafabad aquifer.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"4\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZone\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eName\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMain Influential Factor\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eDescription\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRiver Corridor\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRiver discharge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eInfluenced by proximity to river\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eRight Side of Nekouabad Canal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigation canal (Nekouabad)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAffected by controlled irrigation system\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLeft Side of Nekouabad canal\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigation canal (Nekouabad)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eSimilar to Z2 but spatially distinct\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eKhamiran Irrigation Area\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eGroundwater abstraction\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eAffected by Khamiran irrigation scheme\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eWestern Highlands\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePrecipitation and natural recharge\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eMostly natural influence, limited access\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eZone 1 \u0026ndash; River Corridor Zone: This zone includes areas adjacent to the Zayandeh-Rud River, where groundwater levels are directly influenced by river discharge. Piezometric data from this region reflect the dynamic interaction between surface water and groundwater.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eZones 2 and 3 \u0026ndash; Nekouabad Irrigation Zones: These zones encompass agricultural areas affected by the Nekouabad irrigation network. Based on their geographical orientation, they are divided into: Zone 2: Right side of the Nekouabad irrigation system, and Zone 3: Left side of the Nekouabad irrigation system.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eZone 4 \u0026ndash; Khamiran Irrigation Zone: Located in the western part of the aquifer, this zone is mainly influenced by groundwater abstraction for agriculture under the Khamiran irrigation scheme.\u003c/p\u003e\u003c/li\u003e\u003cli\u003e\u003cp\u003eZone 5 \u0026ndash; Western Highland Zone: This zone includes piezometers situated in the elevated western highlands. Here, groundwater dynamics are primarily governed by natural recharge from precipitation, with limited anthropogenic impact.\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec8\" class=\"Section3\"\u003e\u003ch2\u003e2.2.4. Exploratory Data Analysis of Groundwater Level Fluctuations\u003c/h2\u003e\u003cp\u003eTo gain deeper insights into the temporal behavior of groundwater levels across the Najafabad Plain, an exploratory data analysis (EDA) was conducted using violin plots. These plots effectively combine boxplots and kernel density estimates to depict the distribution and variability of groundwater levels on both monthly and seasonal scales, across the five delineated observation zones (Z1 to Z5). The monthly and seasonal distributions are presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e and Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e, respectively.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eThe monthly violin plots (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) reveal distinct fluctuation patterns among the zones. Zone 1 (Z1), located near the river corridor, exhibits pronounced variability during both summer and winter months, as indicated by the wider spread of values. Zone 2 (Z2) shows a gradual rise in groundwater levels from April to December, along with considerable dispersion during the transitional months, potentially due to irrigation cycles. Zone 3 (Z3) maintains relatively stable levels, with the most notable fluctuations occurring between June and October\u0026mdash;likely reflecting seasonal agricultural demand. In Zone 4 (Z4), groundwater levels remain consistent throughout the year except for December, which displays a sharp increase in variability. Zone 5 (Z5), located in the western highlands, demonstrates a gradual rise in groundwater levels beginning in June and peaking around October, possibly due to delayed recharge from precipitation or irrigation return flows.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eSeasonal violin plots (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e) provide a broader view of temporal groundwater dynamics. Zones Z1, Z2, and Z3 show wider distributions during spring and summer, suggesting higher variability linked to increased evapotranspiration or agricultural withdrawals. In contrast, Z4 and Z5 exhibit narrower, more stable distributions across seasons, with a slight increase in variability during fall and winter. Median groundwater levels are generally higher during spring and summer, particularly in Z1 and Z2, which may reflect the influence of seasonal recharge or irrigation inputs.\u003c/p\u003e\u003cp\u003eThese visual analyses offer critical insights into the temporal variability of groundwater across different zones, highlighting areas that are more sensitive to environmental and anthropogenic influences. The outcomes of this analysis not only enhance conceptual understanding but also inform the subsequent stages of feature engineering, model training, and evaluation in the groundwater prediction framework.\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\u003ch2\u003e2.3. Groundwater Level Modeling Approach\u003c/h2\u003e\u003cdiv id=\"Sec10\" class=\"Section3\"\u003e\u003ch2\u003e2.3.1. Machine Learning Algorithms\u003c/h2\u003e\u003cp\u003eIn this study, three state-of-the-art supervised machine learning algorithms\u0026mdash;Random Forest (RF), Gradient Boosting (GB), and Support Vector Machine (SVM)\u0026mdash;were implemented to model and forecast groundwater level variations across the Najafabad aquifer. These algorithms are widely recognized for their robustness in handling nonlinear relationships, high-dimensional input spaces, and missing or noisy data, making them well-suited for hydrological applications \u003csup\u003e\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e,\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eRandom Forest (RF):\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eThis ensemble learning method is based on the bagging (bootstrap aggregating) technique. Multiple regression trees are constructed using bootstrapped samples of the training data. Each tree h\u003csub\u003ei\u003c/sub\u003e(x) makes a prediction, and the final RF output is the average of predictions from all \u0026#119872; trees:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$$\\:\\widehat{y}=\\frac{1}{M}\\sum\\:_{i=1}^{M}{h}_{i}\\left(x\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eRF minimizes the mean squared error (MSE) across trees and uses random subsets of features at each split to reduce correlation among trees. The objective function focuses on reducing variance while maintaining low bias. Feature importance is computed based on the mean decrease in impurity or prediction accuracy across all trees. In this study, the RF model was implemented using the scikit-learn library with hyperparameter tuning (e.g., number of trees, max depth) via grid search \u003csup\u003e\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e,\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eGradient Boosting (GB):\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eGradient Boosting builds models sequentially by fitting each new model to the residual errors of the ensemble prediction so far. The objective function combines a loss function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:L(y,\\widehat{y})\\)\u003c/span\u003e\u003c/span\u003e (e.g., squared error loss) and a regularization term \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\Omega\\:}\\left({f}_{t}\\right)\\)\u003c/span\u003e\u003c/span\u003e to penalize model complexity:\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\:{L}^{\\left(t\\right)}=\\sum\\:L\\left({y}_{i},{\\widehat{y}}_{i}^{(t-1)}+{f}_{t}\\left({x}_{i}\\right)\\right)+{\\Omega\\:}\\left({f}_{t}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eIn this study, the eXtreme Gradient Boosting (XGBoost) version was used, which improves efficiency and regularization. XGBoost uses second-order Taylor approximation of the loss function for optimization and supports shrinkage (learning rate) and column subsampling to reduce overfitting. The additive model is updated as:\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\:{\\widehat{y}}^{\\left(t\\right)}={\\widehat{y}}^{(t-1)}+\\eta\\:{f}_{t}\\left(x\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere \u0026#120578; is the learning rate. The model was implemented using the XGBoost library with hyperparameter tuning on parameters such as the number of estimators, learning rate, and maximum tree depth \u003csup\u003e\u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e,\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eSupport Vector Machine (SVM):\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003c/p\u003e\u003cp\u003eSupport Vector Regression (SVR) aims to find a function that approximates the underlying relationship between input features and target values while balancing model complexity and prediction accuracy. The SVR function is expressed as:\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$\\:f\\left(x\\right)=\u0026lang;w,\\varphi\\:\\left(x\\right)\u0026rang;+b$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere: \u0026#120601; (\u0026#119909;) maps the input vector \u0026#119909; into a high-dimensional feature space, and \u0026#119908; and \u0026#119887; are the model parameters.\u003c/p\u003e\u003cp\u003eThe SVR formulation attempts to minimize the model complexity (by keeping ∥\u0026#119908;∥ small) while ensuring that the predictions are within an ε-insensitive margin. The optimization problem is defined as:\u003cdiv id=\"Equ5\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ5\" name=\"EquationSource\"\u003e\n$$\\:{min}_{w,b,{\\xi\\:}_{i},{\\xi\\:}_{i}^{*}}=\\frac{1}{2}{‖w‖}^{2}+C\\sum\\:_{i=1}^{n}\\left({\\xi\\:}_{i}+{\\xi\\:}_{i}^{*}\\right)$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e5\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eSubject to:\u003cdiv id=\"Equ6\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ6\" name=\"EquationSource\"\u003e\n$$\\:\\left\\{\\begin{array}{c}{y}_{i}-\u0026lang;w,\\varphi\\:\\left({x}_{i}\\right)\u0026rang;-b\\le\\:\\epsilon\\:+{\\xi\\:}_{i}\\\\\\:\u0026lang;w,\\varphi\\:({x}_{i}\u0026rang;+b-{y}_{i}\\le\\:\\epsilon\\:+{\\xi\\:}_{i}^{*}\\\\\\:{\\xi\\:}_{i},{\\xi\\:}_{i}^{*}\\ge\\:0\\end{array}\\right.$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e6\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003ewhere: \u0026#119862; is the regularization parameter that controls the trade-off between the flatness of the function and the amount up to which deviations larger than ε are tolerated. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{\\xi\\:}_{i},{\\xi\\:}_{i}^{*}\\)\u003c/span\u003e\u003c/span\u003e are slack variables that allow violations of the ε margin.\u003c/p\u003e\u003cp\u003eTo model nonlinear relationships, a kernel function \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:K\\left({x}_{i},{x}_{j}\\right)=\u0026lang;\\varphi\\:\\left({x}_{i}\\right),\\varphi\\:\\left({x}_{j}\\right)\u0026rang;\\)\u003c/span\u003e\u003c/span\u003e is employed. In this study, the Radial Basis Function (RBF) kernel was used:\u003cdiv id=\"Equ7\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ7\" name=\"EquationSource\"\u003e\n$$\\:K\\left({x}_{i},{x}_{j}\\right)=\\text{e}\\text{x}\\text{p}(-\\gamma\\:{‖{x}_{i}-{x}_{j}‖}^{2})$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e7\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eThe RBF kernel enables the model to learn complex, nonlinear patterns in the data.\u003c/p\u003e\u003cp\u003eThe SVR model was implemented using the scikit-learn library in Python. The key hyperparameters\u0026mdash;\u0026#119862; (penalty), \u0026#120576; (margin width), and \u0026#120574; (kernel coefficient)\u0026mdash;were optimized through a grid search approach combined with cross-validation to ensure robust performance \u003csup\u003e\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e,\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eAll models were trained using the historical time series of meteorological variables (e.g., precipitation, temperature), surface water inflow, and groundwater levels across five hydrogeological zones (Section \u003cspan refid=\"Sec11\" class=\"InternalRef\"\u003e2.3.2\u003c/span\u003e). A combination of grid search and five-fold cross-validation was applied to fine-tune model hyperparameters and ensure robust generalization. Model performance was evaluated using standard statistical metrics, as detailed in Section \u003cspan refid=\"Sec13\" class=\"InternalRef\"\u003e2.3.4\u003c/span\u003e.\u003c/p\u003e\u003cp\u003eThe complete implementation codes, including data preprocessing, model training, and evaluation scripts, are provided in Supplementary Material (S2) for reference and reproducibility.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec11\" class=\"Section3\"\u003e\u003ch2\u003e2.3.2. Input Variables per Zone\u003c/h2\u003e\u003cp\u003eTo account for the spatial heterogeneity of hydrological and anthropogenic influences across the Najafabad aquifer, groundwater level modeling was conducted separately for five hydrogeological distinct zones (Z1\u0026ndash;Z5). This zonal approach allows for a more accurate representation of local dynamics by tailoring model inputs to the specific conditions of each area.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab3\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eInput Variables for Groundwater Level Change Modeling in the Defined Zones.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"8\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eZone\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"6\" nameend=\"c7\" namest=\"c2\"\u003e\u003cp\u003eInputs\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eOutput\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eMeteorological Station\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003ePrecipitation (P)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eTemperature (T)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eRiver Discharge (RD)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eIrrigation Volume (IV)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c7\"\u003e\u003cp\u003eGroundwater Abstraction (GwA)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c8\"\u003e\u003cp\u003eGroundwater Level (GwL)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eZefreh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eZefreh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eZefreh\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNajafabad\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eZ5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNajafabad\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e\u003cb\u003e-\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e\u003cb\u003e✓\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e✓: Variable included as model input, \u0026ndash;: Variable excluded due to unavailability or irrelevance.\u003c/p\u003e\u003cp\u003eThe selection of input variables in each zone was guided by both data availability and their relevance to groundwater level fluctuations. The monthly change in groundwater level (GwL) was considered the output variable across all zones, while the predictor variables included monthly precipitation (P), average temperature (T), river discharge (RD), irrigation volume (IV), and groundwater abstraction (GwA). Table\u0026nbsp;\u003cspan refid=\"Tab3\" class=\"InternalRef\"\u003e3\u003c/span\u003e provides an overview of the input variables used in each zone. The inclusion or exclusion of specific variables was influenced by each zone\u0026rsquo;s proximity to surface water resources, irrigation infrastructure, and elevation. For example, zones located downstream of the Nekouabad canal included irrigation volume as a model input, while zones without reliable abstraction data excluded that variable.\u003c/p\u003e\u003cp\u003eTo further substantiate the selection of input variables and enhance the understanding of interrelationships among the influencing factors, Pearson correlation heatmaps were generated separately for each of the five hydrogeological zones (Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). These heatmaps provide a clear visualization of the strength and direction of linear relationships between groundwater level change and its potential predictors. The results revealed spatial variability in correlation patterns across the zones, emphasizing the importance of localized analysis\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eIn Zone Z1, a moderate positive correlation was observed between groundwater level change and groundwater abstraction (r\u0026thinsp;=\u0026thinsp;0.44), indicating a significant influence of human withdrawals on the aquifer\u0026rsquo;s behavior. Zone Z2 exhibited a strong negative correlation between temperature and groundwater level change (r = \u0026minus;\u0026thinsp;0.51), suggesting that increased evapotranspiration during warmer periods may accelerate groundwater depletion. In Zone Z3, the groundwater system appeared to be influenced by a balanced combination of climatic variables (such as precipitation and temperature) and anthropogenic drivers (including irrigation and abstraction). In contrast, Zones Z4 and Z5 showed generally weaker correlations among the variables, potentially indicating more stable hydrological conditions or the presence of other unmeasured influencing factors.\u003c/p\u003e\u003cp\u003eThese statistical insights, derived from correlation analysis, reinforce the appropriateness of the zonal input configurations and highlight the necessity of incorporating both spatial and causal heterogeneity in the modeling of groundwater dynamics within the Najafabad aquifer.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec12\" class=\"Section3\"\u003e\u003ch2\u003e2.3.3. Model Training and Validation Strategy\u003c/h2\u003e\u003cp\u003eTo ensure robust and generalizable model performance, the available dataset for each hydrogeological zone was split into training and testing subsets. Seventy percent (70%) of the data was used for training the machine learning models, while the remaining 30% was reserved for testing and evaluation. The time series were divided chronologically to preserve the temporal structure and prevent data leakage. To fine-tune the model parameters and enhance prediction accuracy, a five-fold cross-validation (CV) strategy was employed during the training phase. In this method, the training data was partitioned into five equal subsets, and each subset was used once for validation while the remaining four subsets were used for training. The final model performance was averaged across the five folds, ensuring a balanced evaluation of bias and variance.\u003c/p\u003e\u003cp\u003eHyperparameter optimization was performed for each algorithm using a grid search approach, where multiple combinations of key parameters were systematically tested. The optimal hyperparameter set was selected based on minimizing the validation error, specifically the Root Mean Square Error (RMSE). To further prevent overfitting, each model incorporated algorithm-specific regularization mechanisms. For example, the Random Forest model utilized a maximum tree depth constraint and minimum sample split size, the Gradient Boosting model implemented learning rate and subsampling adjustments, and the Support Vector Machine model optimized the kernel function and regularization coefficient (C).\u003c/p\u003e\u003cp\u003eThis training-validation strategy ensured that the models were both accurate and resilient to noise and overfitting, and it enabled fair comparisons between algorithms across different zones of the Najafabad Plain.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec13\" class=\"Section3\"\u003e\u003ch2\u003e2.3.4. Model Evaluation Metrics\u003c/h2\u003e\u003cp\u003eTo evaluate the predictive accuracy of the developed machine learning models, three standard performance metrics were utilized: Coefficient of Determination (R\u0026sup2;), Root Mean Square Error (RMSE), and Mean Absolute Error (MAE). These metrics provide complementary insights into model performance in terms of accuracy, bias, and error distribution.\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eCoefficient of Determination (R\u0026sup2;) assesses how well the predicted values replicate the observed values. It is defined as:\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003cdiv id=\"Equ8\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ8\" name=\"EquationSource\"\u003e\n$$\\:{R}^{2}=1-\\frac{\\sum\\:_{i=1}^{n}{\\left({O}_{i}-{P}_{i}\\right)}^{2}}{\\sum\\:_{i=1}^{n}{\\left({O}_{i}-\\stackrel{-}{O}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e8\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eRoot Mean Square Error (RMSE) measures the square root of the average squared differences between predicted and observed values:\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003cdiv id=\"Equ9\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ9\" name=\"EquationSource\"\u003e\n$$\\:RMSE=\\sqrt{\\frac{1}{n}\\sum\\:_{i=1}^{n}{\\left({O}_{i}-{P}_{i}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e9\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cul\u003e\u003cli\u003e\u003cp\u003eMean Absolute Error (MAE) quantifies the average magnitude of absolute differences between predictions and actual observations:\u003c/p\u003e\u003c/li\u003e\u003c/ul\u003e\u003cdiv id=\"Equ10\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ10\" name=\"EquationSource\"\u003e\n$$\\:MAE=\\frac{1}{n}\\sum\\:_{i=1}^{n}\\left|{O}_{i}-{P}_{i}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e10\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eWhere O\u003csub\u003ei\u003c/sub\u003e and P\u003csub\u003ei\u003c/sub\u003e represent the observed and predicted values, respectively, and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\stackrel{-}{O}\\)\u003c/span\u003e\u003c/span\u003e is the mean of observed values. In summary, a higher R\u003csup\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/sup\u003e value closer to 1 indicates better model fit, while lower values of RMSE and MAE, approaching zero, reflect higher prediction accuracy and lower error magnitude.\u003c/p\u003e\u003cp\u003eThese metrics were computed for each model (RF, GB, SVM) and each zone to enable detailed comparison and identify the best-performing algorithm in each hydrogeological context. The selected metrics have been widely applied in hydrological and groundwater modeling studies \u003csup\u003e\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e,\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e\u003c/sup\u003e.\u003c/p\u003e\u003cp\u003eFor a detailed overview of the methodological framework including data preprocessing, zone delineation, model training, and evaluation steps, refer to the Methodology Flowchart provided in Supplementary Material (S1).\u003c/p\u003e\u003c/div\u003e\u003c/div\u003e"},{"header":"3. Results and Discussion","content":"\u003cp\u003eThis section presents the outcomes of the groundwater level modeling across the five defined zones of the Najafabad aquifer using three machine learning algorithms: Random Forest (RF), XGBoost (GB), and Support Vector Machine (SVM). The results are analyzed in terms of model accuracy, spatial variability, and the relative importance of input variables. Furthermore, the performance of the models is compared with findings from previous studies, and implications for regional groundwater management are discussed. The aim is to provide an in-depth understanding of how different factors affect groundwater level fluctuations and how data-driven models can support sustainable water resource planning in semi-arid regions.\u003c/p\u003e\u003cdiv id=\"Sec15\" class=\"Section2\"\u003e\u003ch2\u003e3.1. Model Performance per Zone\u003c/h2\u003e\u003cp\u003eThe performance of the three machine learning models\u0026mdash;XGBoost, Random Forest (RF), and Support Vector Machine (SVM)\u0026mdash;was evaluated across five hydrogeological zones using RMSE, MAE, and R\u0026sup2; during both training and testing phases (Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e). The scatter plots of predicted versus observed groundwater levels for each model and zone are illustrated in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, visually reinforcing the numerical results.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003ePerformance metrics (RMSE, MAE, R\u0026sup2;) of XGBoost, RF, and SVM models across hydrogeological zones Z1\u0026ndash;Z5 during training and testing.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"12\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c10\" colnum=\"10\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c11\" colnum=\"11\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c12\" colnum=\"12\"\u003e\u003c/div\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eIndex\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c7\" namest=\"c3\"\u003e\u003cp\u003eTraining\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"5\" nameend=\"c12\" namest=\"c8\"\u003e\u003cp\u003eTesting\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eZ1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003eZ2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003eZ3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eZ4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eZ5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eZ1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003eZ2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003eZ3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003eZ4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003eZ5\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003e\u003cb\u003eXGBoost\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eRMSE\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.82\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.81\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.64\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.09\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.35\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.85\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2.37\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.87\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e1.25\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.43\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eR\u003c/b\u003e\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.93\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.94\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.90\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e0.85\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e0.76\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.92\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eMAE\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.21\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.38\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.15\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.32\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e1.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e0.67\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.21\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eRMSE\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.74\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.43\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2.14\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.08\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.68\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.77\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2.77\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.93\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e1.23\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.83\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eR\u003c/b\u003e\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.82\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.88\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.86\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.70\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.9\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.75\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e0.77\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e0.59\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.85\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eMAE\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.42\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e1.7\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.41\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.72\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e1.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e2.8\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e0.66\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.48\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eRMSE\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.79\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.23\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2.34\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e1.13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.23\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.87\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2.05\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e4.16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e1.26\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eR\u003c/b\u003e\u003csup\u003e\u003cb\u003e\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e\u003c/b\u003e\u003c/sup\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.91\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.94\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.92\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.85\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0. 97\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.86\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e0.82\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e0.80\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e0.70\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.93\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003e\u003cb\u003eMAE\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.40\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e1.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.51\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e0.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e0.51\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c9\"\u003e\u003cp\u003e2.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c10\"\u003e\u003cp\u003e3.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c11\"\u003e\u003cp\u003e0.75\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c12\"\u003e\u003cp\u003e0.19\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eOverall, XGBoost outperformed the other models in most zones, demonstrating both accuracy and stability. It achieved the highest R\u0026sup2; values during testing in Z1 (0.91), Z3 (0.85), Z4 (0.76), and Z5 (0.92), with relatively low RMSEs. For instance, in Z5, XGBoost yielded a remarkably low RMSE of 0.43 m and MAE of 0.21 m, while maintaining a high R\u0026sup2; of 0.92, indicating excellent generalization and prediction capability in this zone. The SVM model, while slightly more variable, performed surprisingly well in Z5 (R\u0026sup2; = 0.93, RMSE\u0026thinsp;=\u0026thinsp;0.44 m, MAE\u0026thinsp;=\u0026thinsp;0.19 m), and also showed competitive accuracy in Z1 and Z2. However, its performance dropped significantly in Z3, where it recorded the highest RMSE (4.16 m) and MAE (3.2 m) among all models and zones during testing. This can also be observed in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e, where scatter points in Z3-SVM show a noticeable deviation from the 1:1 line, highlighting poor prediction consistency. These findings are consistent with those of Ibrahem Ahmed Osman, et al. \u003csup\u003e\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e\u003c/sup\u003e and Vafadar, et al. \u003csup\u003e\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e\u003c/sup\u003e, who also reported the superior performance of XGBoost in groundwater level prediction tasks, particularly in heterogeneous and data-limited environments.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eRandom Forest (RF) offered moderate results across the zones, with its best performance in Z5 (R\u0026sup2; = 0.85, RMSE\u0026thinsp;=\u0026thinsp;0.83 m, MAE\u0026thinsp;=\u0026thinsp;0.48 m). However, in Z2 and Z3, its testing RMSE exceeded 2.7 m, and R\u0026sup2; dropped to 0.75 and 0.77, respectively. These outcomes align with the findings of Abbas, et al. \u003csup\u003e\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e\u003c/sup\u003e, who recognized RF as a dependable model for groundwater-related predictions, while also highlighting its smoothing effect on sudden variations due to the ensemble averaging process. The scatter plots in these zones reflect the relatively dispersed nature of the predictions, especially when compared with XGBoost\u0026rsquo;s tighter clustering along the 1:1 line. A notable observation is that Zone 5 (Z5), despite its high elevation and potentially complex hydrogeological features, was modeled more successfully than Zones 2 and 3. All models achieved R\u0026sup2; values above 0.85 in Z5 during testing, with particularly low MAEs and RMSEs, suggesting more consistent and learnable patterns in the groundwater data of this zone. In contrast, Z3 posed a challenge for all models\u0026mdash;particularly for SVM\u0026mdash;likely due to higher variability or the presence of outliers. This is evident both in statistical metrics and in the wider spread of data points away from the 1:1 line in Fig.\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e5\u003c/span\u003e. These results are consistent with Dong, et al. \u003csup\u003e\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e\u003c/sup\u003e, who emphasized the importance of appropriate kernel selection and hyperparameter tuning when using SVM in regions with high variability. Furthermore, Zarafshan, et al. \u003csup\u003e\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e\u003c/sup\u003e found that although SVM and its variants like SVR can perform acceptably under various groundwater regimes, their performance is highly sensitive to data stability and spatial heterogeneity.\u003c/p\u003e\u003cp\u003eIn summary, XGBoost demonstrated the most balanced and robust performance across all zones, while SVM and RF showed strong localized performance, especially in zones with less variability. These findings emphasize the importance of considering local hydrogeological characteristics in selecting the optimal modeling approach for groundwater prediction.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec16\" class=\"Section2\"\u003e\u003ch2\u003e3.2. Importance of Input Variables\u003c/h2\u003e\u003cp\u003eTo assess the influence of various input variables on groundwater level prediction, the relative importance of each variable was evaluated using three machine learning algorithms: Support Vector Machine (SVM), Random Forest (RF), and Extreme Gradient Boosting (XGBoost). Figure\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e6\u003c/span\u003e presents the comparative importance values across the five hydrogeological zones (Z1 to Z5).\u003c/p\u003e\u003cp\u003eAcross all zones, precipitation (P) and temperature (T) emerged as the most influential variables, though their relative dominance varied by zone and algorithm. In Zone 1, temperature had a slightly higher influence than precipitation, especially under RF and XGBoost, while SVM showed nearly equal importance for recharge depth (RD) and T. In Zone 2, SVM gave the highest importance to temperature, exceeding 65%, while the other algorithms showed a more balanced influence between temperature and precipitation. Zone 3 showed similar trends, with temperature again dominating in all models, particularly in RF and XGBoost, while groundwater abstraction (GwA) had minimal contribution.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eIn Zone 4, XGBoost highlighted temperature as the most dominant variable (~\u0026thinsp;47%), followed by precipitation and recharge depth, whereas SVM and RF placed more emphasis on precipitation. Interestingly, Zone 5, which includes only two variables (P and T), demonstrated the clearest consensus among models: temperature was consistently the most influential, with values close to or above 65% in all models. Overall, temperature and precipitation consistently ranked as the top predictors, aligning with the hydro-climatic nature of groundwater dynamics in the region. In contrast, groundwater abstraction and irrigation volume showed lower relative importance in most zones, possibly due to weaker correlations or limited temporal resolution. These findings underline the spatial variability in predictor importance, highlighting the need to tailor groundwater modeling approaches to regional characteristics. Our results align closely with those of Nourani, et al. \u003csup\u003e\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e\u003c/sup\u003e and Costantini, et al. \u003csup\u003e\u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e\u003c/sup\u003e, who emphasized the predominant role of climatic drivers\u0026mdash;particularly precipitation and temperature\u0026mdash;over anthropogenic factors in arid and semi-arid aquifer systems. This consistency reinforces the conclusion that, in data-scarce regions like Najafabad, the predictive power of climatic variables outweighs that of human-induced parameters such as irrigation and pumping, whose impact may be masked by coarse temporal or spatial data resolution.\u003c/p\u003e\u003cp\u003eIn conclusion, this comparative analysis reaffirms the robustness of ensemble learning methods\u0026mdash;particularly XGBoost\u0026mdash;in predicting groundwater levels under complex hydrogeological conditions. It also underscores the importance of zone-based model calibration and input sensitivity analysis to account for local hydro-climatic variability and ensure the development of context-appropriate forecasting tools.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec17\" class=\"Section2\"\u003e\u003ch2\u003e3.3. Model Robustness and Generalization\u003c/h2\u003e\u003cp\u003eRobustness and generalization are critical attributes of any predictive model, particularly in groundwater studies where data quality, completeness, and variability can significantly influence model outcomes. In the present study, all three machine learning models\u0026mdash;Random Forest (RF), Support Vector Machine (SVM), and Extreme Gradient Boosting (XGBoost)\u0026mdash;were evaluated across five hydrogeological zones with varying characteristics to assess their stability and transferability (Table \u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eComparison of model performance metrics in training and testing phases across all zones.\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003ePhase\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eMean R\u0026sup2;\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eMean RMSE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eΔR\u0026sup2; (Train - Test)\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eΔRMSE (Test - Train)\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eXGBoost\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTraining\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.934\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e0.942\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e0.086\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e0.612\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTesting\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.848\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e1.554\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eRF\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTraining\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.832\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e1.214\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e0.084\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e0.492\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTesting\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.748\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e1.706\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e\u003cb\u003eSVM\u003c/b\u003e\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTraining\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.918\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e1.144\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c5\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e0.096\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c6\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003e0.612\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eTesting\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e\u003cp\u003e0.822\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e\u003cp\u003e1.756\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eAmong the models, XGBoost demonstrated the highest predictive performance during both training (mean R\u0026sup2; = 0.934, RMSE\u0026thinsp;=\u0026thinsp;0.942) and testing (mean R\u0026sup2; = 0.848, RMSE\u0026thinsp;=\u0026thinsp;1.554), with a moderate drop in performance (ΔR\u0026sup2; = 0.086, ΔRMSE\u0026thinsp;=\u0026thinsp;0.612). This relatively small gap highlights XGBoost\u0026rsquo;s robust generalization ability, especially considering its ability to capture nonlinear groundwater dynamics in both stable (e.g., Z1, Z5) and hydrologically complex regions (e.g., Z3). This robustness can be attributed to its embedded regularization mechanisms, capacity to manage multicollinearity, and iterative bias-variance reduction during training. The Random Forest (RF) model achieved lower overall predictive accuracy (R\u0026sup2; = 0.832 in training and 0.748 in testing) but exhibited the smallest RMSE difference (ΔRMSE\u0026thinsp;=\u0026thinsp;0.492), indicating good stability and a low tendency to overfit. While RF may underperform in zones with sharp hydrological fluctuations, it is relatively resilient to noisy or incomplete data, making it useful for preliminary analysis or areas with limited observations.\u003c/p\u003e\u003cp\u003eThe Support Vector Machine (SVM) model produced high accuracy during training (R\u0026sup2; = 0.918) but exhibited the largest generalization gap (ΔR\u0026sup2; = 0.096), indicating greater sensitivity to overfitting, especially in zones with high groundwater variability. Its testing performance dropped more sharply than the other models, which can be attributed to SVM's reliance on kernel function selection and sensitivity to feature distribution. These limitations necessitate careful preprocessing, normalization, and kernel tuning for improved robustness.\u003c/p\u003e\u003cp\u003eThese comparative results reinforce the importance of selecting models that balance accuracy with generalization, especially in regions such as the Najafabad plain where hydrogeological heterogeneity and climatic variability present significant modeling challenges. Among the three models, XGBoost emerged as the most robust and generalizable, offering both predictive strength and stability across diverse zones. Future studies should consider hybrid modeling strategies that integrate multiple machine learning models or external data sources (e.g., remote sensing, irrigation patterns, or socio-economic data) to enhance predictive robustness and transferability in similar complex groundwater systems.\u003c/p\u003e\u003c/div\u003e\u003cdiv id=\"Sec18\" class=\"Section2\"\u003e\u003ch2\u003e3.4. Implications for Groundwater Management and Future Directions\u003c/h2\u003e\u003cp\u003eThe findings of this study hold significant implications for groundwater management in data-scarce and hydrogeological diverse regions such as the Najafabad plain. By evaluating the performance of three machine learning algorithms\u0026mdash;XGBoost, Random Forest (RF), and Support Vector Machine (SVM)\u0026mdash;across five distinct zones, the research highlights the potential of data-driven approaches to enhance decision-making under conditions of uncertainty and limited observational data.\u003c/p\u003e\u003cp\u003eThe strong performance of XGBoost, particularly in terms of generalization and robustness, demonstrates its suitability for capturing complex and nonlinear groundwater dynamics. This makes it a valuable tool for supporting real-time groundwater level forecasting, resource allocation, and early warning systems. Moreover, the ability to quantify feature importance offers a transparent mechanism for identifying key drivers of groundwater fluctuations, such as precipitation and temperature, which can aid in the design of adaptive climate-resilient policies. From a management perspective, the zone-specific variability in model accuracy underscores the need for localized modeling frameworks. Zones like Z3, which exhibited more complex and unstable groundwater behavior, demand tailored calibration efforts and possibly the integration of external hydrological or anthropogenic indicators (e.g., land use, irrigation practices, pumping intensity) to improve predictive reliability.\u003c/p\u003e\u003cp\u003eHowever, several limitations of this study should be acknowledged. First, the accuracy of model predictions is inherently tied to the quality and resolution of input data. In this study, limited availability of high-resolution spatiotemporal data on irrigation and pumping rates may have influenced model performance, particularly in zones with high human activity. Second, while the machine learning models demonstrated strong predictive ability, they operate as black-box systems with limited capacity for causal interpretation, which may hinder their adoption in regulatory or policy contexts that require process transparency. Future research should explore hybrid modeling approaches that combine the strengths of data-driven methods with physically-based models to enhance both prediction and interpretability. Additionally, the integration of remotely sensed data, socio-economic indicators, and long-term groundwater abstraction records could help to further refine model inputs and improve spatial transferability. Developing dynamic models that incorporate feedback mechanisms from land use changes and climate scenarios could also strengthen the applicability of machine learning in strategic groundwater resource planning. In summary, this study affirms the promise of machine learning models\u0026mdash;particularly ensemble-based approaches like XGBoost\u0026mdash;for operational groundwater management. With continued refinement and integration of interdisciplinary data sources, these models can play a vital role in addressing the growing challenges of groundwater sustainability in semi-arid and arid environments.\u003c/p\u003e\u003c/div\u003e"},{"header":"4. Conclusion","content":"\u003cp\u003eThis study evaluated the predictive performance of three machine learning algorithms\u0026mdash;Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Support Vector Machine (SVM)\u0026mdash;in forecasting groundwater level fluctuations across five hydrogeological zones in the Najafabad plain, a semi-arid region in central Iran. Using climatic, hydrological, and anthropogenic input variables, the models were trained and tested to assess their accuracy, robustness, and generalization under varying groundwater conditions. Among the models, XGBoost consistently outperformed the others, achieving the highest average coefficient of determination during testing (mean R\u0026sup2; = 0.848) and the lowest root mean square error (RMSE\u0026thinsp;=\u0026thinsp;1.554). It also demonstrated a moderate performance gap between training and testing phases (ΔR\u0026sup2; = 0.086), indicating strong generalization ability and limited overfitting. This superior performance is attributed to its boosting mechanism, embedded regularization, and capacity to capture nonlinear groundwater dynamics even in complex zones (e.g., Z3).\u003c/p\u003e\u003cp\u003eRandom Forest (RF) showed moderate accuracy with mean R\u0026sup2; values of 0.832 in training and 0.748 in testing, and a relatively small RMSE difference (ΔRMSE\u0026thinsp;=\u0026thinsp;0.492), reflecting stable predictions with low overfitting risk. However, its performance declined notably in zones with highly variable groundwater levels. Support Vector Machine (SVM) achieved high training accuracy (mean R\u0026sup2; = 0.918) but exhibited the largest generalization gap (ΔR\u0026sup2; = 0.096) and higher testing RMSE (1.756), revealing sensitivity to data scaling and kernel selection, particularly in hydrologically volatile zones.\u003c/p\u003e\u003cp\u003eThese results highlight the trade-off between accuracy and robustness across models and emphasize the suitability of XGBoost for groundwater forecasting in semi-arid, hydrogeological heterogeneous regions like Najafabad. The ability of XGBoost to maintain high accuracy with moderate overfitting risk makes it an excellent tool for operational groundwater monitoring, resource allocation, and early warning systems. Nonetheless, limitations such as the lack of high-resolution anthropogenic data (e.g., pumping and irrigation rates) likely constrained model performance in certain zones. Furthermore, the \u0026ldquo;black-box\u0026rdquo; nature of these models limits direct causal interpretation, which can restrict their policy adoption.\u003c/p\u003e\u003cp\u003eFuture research should explore hybrid approaches combining data-driven models with physically-based groundwater models, and integrate diverse data sources such as remote sensing and socio-economic indicators. Additionally, dynamic modeling that incorporates land use and climate change scenarios could further improve forecasting reliability and sustainability in groundwater management. In conclusion, this study confirms that ensemble machine learning methods\u0026mdash;particularly XGBoost\u0026mdash;offer powerful, flexible, and scalable solutions for groundwater level prediction in complex semi-arid environments, supporting sustainable water resource management amid increasing climatic and anthropogenic pressures.\u003c/p\u003e"},{"header":"Declarations","content":"\u003cp\u003eAll authors have read, understood, and have complied as applicable with the statement on \u0026quot;Ethical responsibilities of Authors\u0026quot; as found in the Instructions for Authors.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eCompeting Interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFinancial Disclosure statement\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe author(s) received no specific funding for this work.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor contributions:\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAll authors contributed to the conceptualization and design of the study. Material preparation, data collection, and analysis were performed by Sahar Davari, Saeid Eslamian, and Mohammad Jamali. Code development was carried out by Mohammad Jamali and Sahar Davari. The first draft of the manuscript was written by Saeid Eslamian and Mohammad Jamali. Review and editing were conducted by Saeid Eslamian and Hamid Reza Safavi. All authors read and approved the final manuscript and provided comments on previous versions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eData availability\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe dataset generated and/or analyzed during the current study is available from the corresponding author upon reasonable request. The Python codes used for model development and evaluation are publicly available at the following GitHub repository: https://github.com/mohammadj74/Groundwater-ML-Najafabad\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEthics approval\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study is original research and has not been previously published or submitted elsewhere. The research does not involve any experiments on animals or human subjects.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eFamiglietti, J. S. The global groundwater crisis. \u003cem\u003eNature Climate Change\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 945-948, doi:10.1038/nclimate2425 (2014).\u003c/li\u003e\n\u003cli\u003eJamali, M., Yazdian, H., Bahman, G. \u0026amp; Eslamian, S. Water agriculture nexus a system dynamics approach for the next three decades. \u003cem\u003eScientific Reports\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 5946, doi:10.1038/s41598-025-90728-3 (2025).\u003c/li\u003e\n\u003cli\u003eWada, Y.\u003cem\u003e et al.\u003c/em\u003e Global depletion of groundwater resources. \u003cem\u003eGeophysical Research Letters\u003c/em\u003e \u003cstrong\u003e37\u003c/strong\u003e, doi:https://doi.org/10.1029/2010GL044571 (2010).\u003c/li\u003e\n\u003cli\u003eHarbaugh, A. W. \u003cem\u003eMODFLOW-2005, the US Geological Survey modular ground-water model: the ground-water flow process\u003c/em\u003e. Vol. 6 (US Department of the Interior, US Geological Survey Reston, VA, USA, 2005).\u003c/li\u003e\n\u003cli\u003eJahanshahi, A.\u003cem\u003e et al.\u003c/em\u003e Dependence of rainfall-runoff model transferability on climate conditions in Iran. \u003cem\u003eHydrological Sciences Journal\u003c/em\u003e \u003cstrong\u003e67\u003c/strong\u003e, 564-587, doi:10.1080/02626667.2022.2030867 (2022).\u003c/li\u003e\n\u003cli\u003eXu, T. \u0026amp; Liang, F. Machine learning for hydrologic sciences: An introductory overview. \u003cem\u003eWIREs Water\u003c/em\u003e \u003cstrong\u003e8\u003c/strong\u003e, e1533, doi:https://doi.org/10.1002/wat2.1533 (2021).\u003c/li\u003e\n\u003cli\u003eLange, H. \u0026amp; Sippel, S. in \u003cem\u003eForest-Water Interactions\u003c/em\u003e (eds Delphis F. Levia\u003cem\u003e et al.\u003c/em\u003e) 233-257 (Springer International Publishing, 2020).\u003c/li\u003e\n\u003cli\u003ePapacharalampous, G. \u0026amp; Tyralis, H. A review of machine learning concepts and methods for addressing challenges in probabilistic hydrological post-processing and forecasting. \u003cem\u003eFrontiers in Water\u003c/em\u003e \u003cstrong\u003eVolume 4 - 2022\u003c/strong\u003e, doi:10.3389/frwa.2022.961954 (2022).\u003c/li\u003e\n\u003cli\u003eEslamian, S. \u0026amp; Eslamian, F. \u003cem\u003eHandbook of hydroinformatics: volume i: classic soft-computing techniques\u003c/em\u003e. (Elsevier, 2022).\u003c/li\u003e\n\u003cli\u003eTao, H.\u003cem\u003e et al.\u003c/em\u003e Groundwater level prediction using machine learning models: A comprehensive review. \u003cem\u003eNeurocomputing\u003c/em\u003e \u003cstrong\u003e489\u003c/strong\u003e, 271-308, doi:https://doi.org/10.1016/j.neucom.2022.03.014 (2022).\u003c/li\u003e\n\u003cli\u003eSahoo, S., Russo, T. A., Elliott, J. \u0026amp; Foster, I. Machine learning algorithms for modeling groundwater level changes in agricultural regions of the U.S. \u003cem\u003eWater Resources Research\u003c/em\u003e \u003cstrong\u003e53\u003c/strong\u003e, 3878-3895, doi:https://doi.org/10.1002/2016WR019933 (2017).\u003c/li\u003e\n\u003cli\u003ePham, Q. B.\u003cem\u003e et al.\u003c/em\u003e Groundwater level prediction using machine learning algorithms in a drought-prone area. \u003cem\u003eNeural Computing and Applications\u003c/em\u003e \u003cstrong\u003e34\u003c/strong\u003e, 10751-10773, doi:10.1007/s00521-022-07009-7 (2022).\u003c/li\u003e\n\u003cli\u003eRaturi, M., Khare, D. \u0026amp; Patidar, N. in \u003cem\u003eApplications of Machine Learning in Hydroclimatology\u003c/em\u003e (eds Roshan Srivastav \u0026amp; Purna C. Nayak) 57-71 (Springer Nature Switzerland, 2025).\u003c/li\u003e\n\u003cli\u003eVafadar, S., Rahimzadegan, M. \u0026amp; Asadi, R. Evaluating the performance of machine learning methods and Geographic Information System (GIS) in identifying groundwater potential zones in Tehran-Karaj plain, Iran. \u003cem\u003eJournal of Hydrology\u003c/em\u003e \u003cstrong\u003e624\u003c/strong\u003e, 129952, doi:https://doi.org/10.1016/j.jhydrol.2023.129952 (2023).\u003c/li\u003e\n\u003cli\u003eZarafshan, P.\u003cem\u003e et al.\u003c/em\u003e Comparison of machine learning models for predicting groundwater level, case study: Najafabad region. \u003cem\u003eActa Geophysica\u003c/em\u003e \u003cstrong\u003e71\u003c/strong\u003e, 1817-1830, doi:10.1007/s11600-022-00948-8 (2023).\u003c/li\u003e\n\u003cli\u003eIbrahem Ahmed Osman, A., Najah Ahmed, A., Chow, M. F., Feng Huang, Y. \u0026amp; El-Shafie, A. Extreme gradient boosting (Xgboost) model to predict the groundwater levels in Selangor Malaysia. \u003cem\u003eAin Shams Engineering Journal\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 1545-1556, doi:https://doi.org/10.1016/j.asej.2020.11.011 (2021).\u003c/li\u003e\n\u003cli\u003eNaderi, M. \u0026amp; Hajiketabi, M. Quantification of normal and sustainable management practices for groundwater resources: example of the arid Najafabad alluvial aquifer in Isfahan Province, Iran. \u003cem\u003eHydrogeology Journal\u003c/em\u003e \u003cstrong\u003e31\u003c/strong\u003e, 195-218, doi:10.1007/s10040-023-02596-8 (2023).\u003c/li\u003e\n\u003cli\u003eKarimi, H.\u003cem\u003e et al.\u003c/em\u003e Enhancing groundwater quality prediction through ensemble machine learning techniques. \u003cem\u003eEnvironmental Monitoring and Assessment\u003c/em\u003e \u003cstrong\u003e197\u003c/strong\u003e, 21, doi:10.1007/s10661-024-13506-0 (2024).\u003c/li\u003e\n\u003cli\u003eBreiman, L. Random Forests. \u003cem\u003eMachine Learning\u003c/em\u003e \u003cstrong\u003e45\u003c/strong\u003e, 5-32, doi:10.1023/A:1010933404324 (2001).\u003c/li\u003e\n\u003cli\u003ePedregosa, F.\u003cem\u003e et al.\u003c/em\u003e Scikit-learn: Machine learning in Python. \u003cem\u003ethe Journal of machine Learning research\u003c/em\u003e \u003cstrong\u003e12\u003c/strong\u003e, 2825-2830 (2011).\u003c/li\u003e\n\u003cli\u003eChen, T. \u0026amp; Guestrin, C. in \u003cem\u003eProceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining\u003c/em\u003e 785\u0026ndash;794 (Association for Computing Machinery, San Francisco, California, USA, 2016).\u003c/li\u003e\n\u003cli\u003eKashani, A. \u0026amp; Safavi, H. R. Assessing groundwater drought in Iran using GRACE data and machine learning. \u003cem\u003eScientific Reports\u003c/em\u003e \u003cstrong\u003e15\u003c/strong\u003e, 14671, doi:10.1038/s41598-025-99342-9 (2025).\u003c/li\u003e\n\u003cli\u003eVapnik, V. \u003cem\u003eThe nature of statistical learning theory\u003c/em\u003e. (Springer science \u0026amp; business media, 2013).\u003c/li\u003e\n\u003cli\u003eSmola, A. J. \u0026amp; Sch\u0026ouml;lkopf, B. A tutorial on support vector regression. \u003cem\u003eStatistics and Computing\u003c/em\u003e \u003cstrong\u003e14\u003c/strong\u003e, 199-222, doi:10.1023/B:STCO.0000035301.49549.88 (2004).\u003c/li\u003e\n\u003cli\u003eYadav, A., Raj, A. \u0026amp; Yadav, B. Enhancing local-scale groundwater quality predictions using advanced machine learning approaches. \u003cem\u003eJournal of Environmental Management\u003c/em\u003e \u003cstrong\u003e370\u003c/strong\u003e, 122903, doi:https://doi.org/10.1016/j.jenvman.2024.122903 (2024).\u003c/li\u003e\n\u003cli\u003eThakur, S. \u0026amp; Karmakar, S. A Comparative Analysis of ANN, LSTM and Hybrid PSO-LSTM Algorithms for Groundwater Level Prediction. \u003cem\u003eTransactions of the Indian National Academy of Engineering\u003c/em\u003e \u003cstrong\u003e10\u003c/strong\u003e, 101-108, doi:10.1007/s41403-024-00505-3 (2025).\u003c/li\u003e\n\u003cli\u003eAbbas, F.\u003cem\u003e et al.\u003c/em\u003e (Water 2024, 16, 941., 2024).\u003c/li\u003e\n\u003cli\u003eDong, J., Tsai, G. \u0026amp; Olivares, C. I. Prediction of 35 Target Per- and Polyfluoroalkyl Substances (PFASs) in California Groundwater Using Multilabel Semisupervised Machine Learning. \u003cem\u003eACS ES\u0026amp;T Water\u003c/em\u003e \u003cstrong\u003e4\u003c/strong\u003e, 969-981, doi:10.1021/acsestwater.3c00134 (2024).\u003c/li\u003e\n\u003cli\u003eNourani, V., Ghareh Tapeh, A. H., Khodkar, K. \u0026amp; Huang, J. J. Assessing long-term climate change impact on spatiotemporal changes of groundwater level using autoregressive-based and ensemble machine learning models. \u003cem\u003eJournal of Environmental Management\u003c/em\u003e \u003cstrong\u003e336\u003c/strong\u003e, 117653, doi:https://doi.org/10.1016/j.jenvman.2023.117653 (2023).\u003c/li\u003e\n\u003cli\u003eCostantini, M., Colin, J. \u0026amp; Decharme, B. Projected Climate‐Driven Changes of Water Table Depth in the World\u0026apos;s Major Groundwater Basins. \u003cem\u003eEarth\u0026apos;s Future\u003c/em\u003e\u003cstrong\u003e11\u003c/strong\u003e, e2022EF003068 (2023).\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true},"keywords":"Groundwater level prediction, Machine Learning, Najafabad Plain, Random Forest, Extreme Gradient Boosting, Support Vector Machine","lastPublishedDoi":"10.21203/rs.3.rs-6992628/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-6992628/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eAccurate groundwater level prediction is vital for sustainable water management, particularly in arid and semi-arid regions under climatic and human-induced stress. This study investigates the performance of three machine learning algorithms\u0026mdash;Extreme Gradient Boosting (XGBoost), Random Forest (RF), and Support Vector Machine (SVM)\u0026mdash;to forecast groundwater levels in five hydrogeological zones of the Najafabad plain, Iran. Models were trained using climatic (precipitation, temperature), hydrological (previous groundwater level), and anthropogenic (irrigation, pumping) inputs. Performance was evaluated using R\u0026sup2;, RMSE, and MAE during training and testing phases. XGBoost outperformed others with a mean testing R\u0026sup2; of 0.848 and RMSE as low as 1.554, showing high robustness, especially in zones with stable to moderately varying water tables. RF achieved moderate accuracy (mean R\u0026sup2; = 0.748) with consistent error margins. SVM showed high training performance (R\u0026sup2; = 0.918) but reduced generalization (testing R\u0026sup2; = 0.822), indicating overfitting tendencies. Results highlight the effectiveness of ensemble models like XGBoost for heterogeneous aquifer systems. However, limited resolution of anthropogenic data remains a challenge. This study reinforces the value of machine learning\u0026mdash;particularly gradient boosting\u0026mdash;as a reliable, scalable approach for groundwater monitoring in data-scarce, complex environments.\u003c/p\u003e","manuscriptTitle":"Application of Machine Learning Algorithms for Groundwater Level Prediction in the Najafabad Plain","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-07-18 12:12:18","doi":"10.21203/rs.3.rs-6992628/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Revision requested","date":"2025-08-04T09:07:22+00:00","index":"","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-22T12:13:29+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"322517559862690340124330675356420607719","date":"2025-07-21T14:42:26+00:00","index":"hide","fulltext":""},{"type":"editorInvitedReview","content":"","date":"2025-07-20T17:29:09+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"174078478021936164441312152622609746501","date":"2025-07-16T16:16:29+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"164481256497012667199349524603727444128","date":"2025-07-16T10:02:12+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"325967037286570084091792341432430296387","date":"2025-07-16T09:52:12+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-07-16T09:25:30+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-07-09T09:00:24+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"","date":"2025-07-08T18:28:01+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2025-07-04T01:28:39+00:00","index":"","fulltext":""},{"type":"submitted","content":"Scientific Reports","date":"2025-07-03T18:21:59+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"[email protected]","identity":"scientific-reports","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"scirep","sideBox":"Learn more about [Scientific Reports](http://www.nature.com/srep/)","snPcode":"","submissionUrl":"","title":"Scientific Reports","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Scientific Reports","inReviewEnabled":true,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"0c0083d7-be24-4262-a87f-821e44898492","owner":[],"postedDate":"July 18th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[{"id":51704240,"name":"Earth and environmental sciences/Environmental sciences"},{"id":51704241,"name":"Earth and environmental sciences/Hydrology"}],"tags":[],"updatedAt":"2025-12-15T16:11:02+00:00","versionOfRecord":{"articleIdentity":"rs-6992628","link":"https://doi.org/10.1038/s41598-025-32376-1","journal":{"identity":"scientific-reports","isVorOnly":false,"title":"Scientific Reports"},"publishedOn":"2025-12-10 15:57:17","publishedOnDateReadable":"December 10th, 2025"},"versionCreatedAt":"2025-07-18 12:12:18","video":"","vorDoi":"10.1038/s41598-025-32376-1","vorDoiUrl":"https://doi.org/10.1038/s41598-025-32376-1","workflowStages":[]},"version":"v1","identity":"rs-6992628","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-6992628","identity":"rs-6992628","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}

Text is read by the "Ask this paper" AI Q&A widget below. Extraction quality varies by source — PMC NXML preserves structure cleanly, OA-HTML may include some navigation residue, and OA-PDF can have broken hyphenation. The publisher copy (via DOI) is the canonical version.

My notes (saved in your browser only)

Ask this paper AI returns verbatim quotes from the full text · source: preprint-html

Answers must be backed by verbatim quotes from this paper's full text. Hallucinated quotes are dropped automatically; if no verbatim passage answers the question, we say so. How this works

Citation neighborhood (no data yet)

We don't have any in-corpus citations linked to this paper yet. This is a recent paper (2025) — citers typically take a year or two to land, and the OpenAlex reference graph may still be filling in.

Source provenance

europepmc
last seen: 2026-05-20T01:45:00.602351+00:00