A Novel Deep Learning Framework for Field Scale Wheat Yield Prediction | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article A Novel Deep Learning Framework for Field Scale Wheat Yield Prediction M Lokeshwari, GIRISH KUMAR JHA, A Praveenkumar, Jyoti Kumari, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-7061170/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 25 Feb, 2026 Read the published version in Theoretical and Applied Genetics → Version 1 posted 5 You are reading this latest preprint version Abstract Hand-held or vehicle-mounted active proximal sensing technologies offer a rapid, non-destructive method for real-time crop monitoring through spectral vegetation indices. This study integrates such proximal sensing data into a deep learning framework for field-scale wheat yield prediction. Specifically, wheat yield is predicted using normalized difference vegetation indices (NDVIs), canopy temperatures (CTs), and plant height (PH) through a deep neural network (DNN) optimized using a genetic algorithm (GA). The model is trained on data from 3,350 diverse wheat germplasm grown under irrigated and rainfed conditions at two locations during the 2020–21 winter season. Comparative analysis demonstrates that the GA-optimized DNN outperforms traditional machine learning models such as Random Forest Regression (RFR), Least Absolute Shrinkage and Selection Operator (LASSO), and Support Vector Regression (SVR). Among individual feature groups, NDVIs measured at five wheat growth stages showing strong predictive capability, with R² values ≥60% under irrigated and ≥50% under rainfed conditions. Additionally, RFR is employed to identify the most influential features within each group. This pioneering study introduces the first-ever application of a GA-optimized deep neural network, leveraging handheld or vehicle-mounted proximal sensing data for predicting crop yield, in the context of Indian agriculture. The proposed approach offers a robust and scalable solution for pre-harvest yield estimation, supporting breeders and researchers in efficient genotype selection and contributing to the achievement of sustainable development goals. agricultural prediction deep neural network genetic algorithm spectral vegetation indices Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Key Messages A genetic algorithm-optimized deep neural network was developed using proximal sensing data to accurately predict wheat yield at field scale, outperforming traditional machine learning models under diverse conditions. Introduction Accurate crop yield prediction plays a pivotal role in ensuring global food security and advancing sustainable agricultural systems (Pantazi et al. 2016; Wang et al. 2014). Timely field-scale yield prediction significantly boosts crop production and profitability while reducing environmental degradation, resource waste and production risks (McBratney et al. 2005; Panda, Ames, and Panigrahi 2010). In addition, precise crop yield prediction enables the rapid and efficient identification of genotypes from a vast pool of potentially promising genotypes under breeding and research programs (Elsayed et al. 2017; Bendig et al. 2015). In order to address aforementioned challenges, the developments of new innovative and efficient crop yield prediction models are highly important. Traditional crop growth models, such as the Decision Support System for Agrotechnology Transfer (DSSAT) has been developed to improve crop yield prediction (Basso and Liu 2019). However, their high predictive power is limited by the need for climatic and management data (Wang et al. 2020; Zhang et al. 2020). Traditional statistical models, like linear regression, predict yield by building simple regression equations between dependent and independent variables (Basso and Liu 2019). These models face challenges in predicting non-linear interactions like weather, soil, fertilizer, and pests, and may cause over-fitting when multiple input variables are present (Carew et al. 2009). Machine learning, leveraged by its data-driven capabilities, has the capacity to construct robust non-linear regression models (Elavarasan et al. 2018). With the exponential increase in data from diverse sources, machine learning-based regression techniques such as Least Absolute Shrinkage and Selection Operator (LASSO), Artificial Neural Networks (ANN) (Fieuzal et al. 2017), Partial Least Squares Regression (PLSR) ((Rischbeck et al. 2016), Random Forest Regression (RFR) (Aghighi et al. 2018), and Support Vector Regression (SVR) (Kuwata and Shibasaki 2016) have showcased significant potential in predicting crop yields. By leveraging multi-layered neural architectures, deep learning models can extract hierarchical features from large datasets and achieve high predictive accuracy in yield estimations (You et al., 2017). Recent studies have explored large-scale wheat yield prediction using deep learning frameworks that integrate remote sensing, soil, and climate data (Cao et al., 2022; Han et al., 2020; Wolanin et al., 2020). However, these models often rely on satellite-based data, which may lack the spatial and temporal resolution required for site-specific precision agriculture (Evans & Shen, 2021; Sagan et al., 2021). Moreover, satellite imagery poses challenges such as cloud interference, low revisit frequency, and extensive pre-processing requirements (Cammarano et al., 2020; Huang & Kuo, 2018). Proximal sensing technologies, particularly hand-held and vehicle-mounted devices, offer a compelling alternative(Hamazaki and Iwata 2022; Messina et al. 2025). These sensors such as the GreenSeeker collect real-time normalized difference vegetation index (NDVI) and canopy temperature (CT) data, which serve as proxies for plant vigor, biomass, and water status (Padilla et al., 2018; Jaradat, 2019; Amin et al., 2023). Unlike satellite sensors, proximal devices are less affected by weather conditions and provide an plot-level resolution suitable for breeding trials(M. Lokeshwari et al. 2025). Consequently, they are well-suited for integrating into field-scale crop yield prediction pipelines to inform genotype selection and resource allocation. Accurate crop yield prediction and precision cultivation relies heavily on the application of deep learning frameworks. To obtain the best performance results of this model, choosing optimal hyper-parameters like number of layers, nodes, learning rate and loss function is crucial. The current study proposes a new hybrid approach, named GA optimized DNN, which combines genetic algorithms and deep neural networks to make precise predictions on wheat yield at a field level. Genetic algorithms simulate natural selection and evolution processes, ultimately determining the most suitable hyper-parameters for the deep neural network model. Utilizing multiple interconnected layers of neurons, deep neural networks provide the capability to recognize complex patterns, enabling better predictions. Overall, this hybrid GA optimized DNN model can offer an effective solution for accurately predicting crop yield, specifically for wheat cultivation. In recent years, the application of deep learning models in crop yield prediction has yielded positive results. For instance, Aula et al. (2022) enhanced the accuracy of mid-season wheat yield prediction by incorporating proximal sensing measurements and weather features into their deep neural network (DNN) model. Similarly, Abbas et al. (2020) utilized machine learning and deep learning algorithms such as linear regression, elastic net regression, k-nearest neighbor, support vector regression, and DNN to predict potato tuber yield using data on soil and crop characteristics obtained through proximal sensing. The main contributions of this study are as follows: Development of a novel GA-optimized deep neural network (GA-DNN) model that integrates proximal sensing data (NDVI, CT, and PH) for accurate wheat yield prediction at field scale under both irrigated and rainfed conditions. Evaluation of feature importance and model performance, highlighting the predictive power of NDVI at different wheat growth stages, and comparing GA-DNN with conventional machine learning models such as Random Forest, LASSO, and SVR. Materials and methods Data description This study uses real-world data to predict the performance of a diverse set of wheat germplasm under different locations as well as environments. Approximately 3350 wheat accession were grown in augmented block design during the winter season 2020–2021 fewer than two environmental conditions i.e. , irrigated and rainfed at ICAR-National Bureau of Plant Genetic Resources (NBPGR), Issapur Farm, New Delhi and at Agharkar Research Institute, Pune under ICAR-DBT network project. The traits measured included grain yield (GY) in gm − 2 , plant height (PH) in (cm), canopy temperature (CT) in (°C), and normalized difference vegetation indices (NDVI). Plant height was measured as the length from ground level to the apex of the spike excluding awns. CT and NDVI data were collected during the growing seasons at different growth stages from tillering to senescence (ground cover, heading, anthesis, grain filling, and maturity) (Zadoks et al. 1974 ) using handheld Infrared Thermometer (IRT) and Handheld GreenSeeker respectively (Table 1 ). Table 1 Description of the dataset Category Item Description Yield Irrigated wheat yield (gm − 2 ) Rainfed wheat yield (gm − 2 ) Harvested yield at two different locations under two growing environments NDVI Normalized Difference Vegetation Indices (NDVI) at five growth stages NDVI data were collected during ground cover, heading, anthesis, grain filling and maturity stages of two environments under two growing environments CT Canopy Temperature (CT,°C) at four growth stages CT data were collected during heading, anthesis, grain filling and maturity stage at two locations under two growing environments PH Plant Height (cm) Measured at two locations under two growing environments Data preprocessing In this dataset, each attribute has its own range of measurements. To ensure precise predictions, the dataset was rescaled using Eq. (1): $$\:{X}^{{\prime\:}}=\frac{X-min\left(X\right)}{max\left(X\right)-min\left(X\right)}$$ 1 Where, X′ is the rescaled value, X is the attributes value, min(X) is minimum of the attributes value and max(X) is the maximum of attributes value. The models are constructed during the training phase, and their predictive abilities are assessed in the testing phase. The flow diagram of the models used to predict wheat yield is shown in Fig. 1 . Least Absolute Shrinkage and Selection Operator (LASSO) LASSO is a linear regression technique that incorporates regularization and feature selection. It adds a penalty equal to the absolute value of the magnitude of coefficients to the regression model, which helps in reducing over fitting. This penalty can shrink some coefficients to zero, effectively selecting a simpler model that involves only a subset of the available features. LASSO is particularly useful for datasets with many features, as it simplifies models and enhances interpretability. The regularization strength is controlled by a parameter, typically chosen through cross-validation. While LASSO helps in feature reduction and preventing over fitting, it may introduce bias, especially with large coefficients or highly correlated variables. Random Forest Regression (RFR) RFR is an ensemble machine learning algorithm used for regression problems. It builds multiple decision trees during training and outputs the mean prediction of the individual trees for more accurate and stable results. RFR is effective in handling large datasets with a mix of categorical and numerical features, and it's robust against over fitting. It also provides insights into feature importance, doesn't require feature scaling, and can handle non-linear relationships. However, key hyper parameters like the number of trees and tree depth need careful tuning for optimal performance. RFR is widely used due to its versatility and generally strong performance across a range of regression problems. Support Vector Regression (SVR) Support Vector Regression (SVR) is a machine learning algorithm used for regression tasks. It extends the concept of Support Vector Machines (SVM) to regression, focusing on fitting a hyper plane within a specified tolerance level to minimize errors. SVR is known for its ability to handle both linear and non-linear relationships through the use of different kernel functions. It's robust against outliers and works well with high-dimensional data but requires careful tuning of hyper parameters like the regularization parameter, kernel type and epsilon value. SVR is effective in scenarios where precision is crucial and is particularly favored for its generalization capabilities. Deep Neural Network (DNN) A Deep Neural Network (DNN) is an advanced type of neural network that consists of multiple layers. These networks are designed to recognize intricate patterns and relationships within data, making them particularly effective for tasks such as image classification, speech recognition, natural language processing, and more. The basic outline of what constitutes a DNN and how it works: Multiple Layers The term "deep" refers to the presence of multiple layers in the network. These typically include an input layer, several hidden layers, and an output layer. The addition of more layers enables the network to learn increasingly complex data representations. Neurons (or Nodes) Each layer is composed of units known as neurons or nodes. These nodes perform calculations by processing input data through weighted sums and applying activation functions, before passing the output to the next layer. Activation Functions To enable the network to model non-linear relationships, activation functions are applied after the weighted input is computed. Common examples include ReLU (Rectified Linear Unit), sigmoid , and tanh functions. Weights and Biases Every connection between neurons carries a weight, and each neuron has an associated bias. These parameters are crucial in determining the output and are fine-tuned during the training process to optimize performance. Forward Propagation In forward propagation, data moves through the network from the input layer to the output layer. At each layer, the data is transformed by weights, biases, and activation functions. Loss Function The network's performance is evaluated using a loss function. This function measures the difference between the network's predictions and the actual target values. Common loss functions include mean squared error for regression tasks and cross-entropy for classification tasks. Backpropagation and Optimization After computing the loss, backpropagation is employed to calculate how much each weight and bias contributed to the error. These gradients are then used by optimization algorithms (like Stochastic Gradient Descent or Adam) to update the weights and minimize the loss over time. Training and Testing During training, the network learns by processing input data, evaluating loss, and adjusting parameters. Once trained, the model is tested on new, unseen data to determine its effectiveness. Regularization and Dropout To ensure that the model generalizes well to new data, techniques like L1/L2 regularization and dropout (which temporarily disables certain neurons during training) are applied. These methods help avoid overfitting and improve the model’s robustness. The general architecture of DNN model is visualized in Fig. 2 . Genetic algorithm for hyper-parameter tuning In the context of deep learning, two categories of parameters are involved. The first includes those learned directly from the data during training, namely weights and biases. The second comprises the learning algorithm parameters, commonly referred to as hyper-parameters. These include the number of neurons in each layer, the number of hidden layers, learning rate, dropout rate, and the regularization function. In this study, genetic algorithms were employed to tune the hyper-parameters of the DNN model. The range of values considered for optimization is detailed in Table 2 . Genetic algorithms are based on the principles of natural selection and evolution, and are well-suited for solving complex problems, including hyper-parameter tuning in deep neural networks. The optimization process starts with generating a population of potential solutions, each represented as a chromosome encoding a specific set of hyper-parameters. These candidate solutions are assessed using a fitness function, which evaluates their performance on the given task. The most successful individuals, identified by their fitness scores, are selected for the next generation. New candidate solutions are then created using genetic operations such as crossover (which combines parts of two solutions) and mutation (which introduces random changes). This process continues over multiple generations, progressively enhancing the quality of solutions. Through repeated application of selection, crossover, and mutation, the genetic algorithm effectively explores the solution space and moves toward globally optimal solutions. The algorithm terminates when predefined stopping criteria are met, such as a set number of generations or a target performance level. The best solution obtained through this iterative process represents the optimized hyper-parameter configuration for the DNN, contributing to improved model performance. The overall workflow of the genetic algorithm is depicted in Fig. 3. Table 2 Parameter space for hyper-parameter optimization Models Hyper-parameters Parameter space Encoding type DNN Number of layers [1–20] Integer Number of neurons [8–32] Integer Dropout rate [0–1] Continuous Learning rate [0–1] Continuous L2 penalty [0–1] Continuous LASSO λ [0–10] Continuous RFR n_estimators [100–500] Integer SVR C [0–10] Continuous kernel ['linear', 'rbf', 'poly'] Categorical The study utilizes a genetic algorithm and employs 10-fold cross-validation on the training data to identify the most suitable hyper-parameters. The performance of the chosen model is assessed using the test set. The selected GA parameters utilized for hyper-parameter tuning are displayed in Table 3 . Evaluation criteria To evaluate the performance of the proposed wheat yield prediction models, two widely used metrics were employed: the coefficient of determination (R²) and the root mean squared error (RMSE), as defined in Equations ( 2 ) and ( 3 ), respectively. R² reflects the proportion of variability in the actual yield that is accounted for by the model’s predictions, serving as an indicator of the model’s goodness-of-fit. A value of R² closer to 1.0 suggests a stronger predictive performance and a higher degree of alignment between predicted and observed yields. In contrast, RMSE quantifies the average prediction error by taking the square root of the mean squared differences between the predicted and actual values. A lower RMSE indicates higher accuracy, with predictions closely matching the true yield values. Combined, these metrics offer a robust assessment of the model’s effectiveness in capturing the complex, non-linear interactions between phenotypic traits and wheat grain yield. $$\:{R}^{2}=\frac{{\sum\:}_{i=1}^{N}{\left[\left({E}_{i}-\overline{E}\right)\left({O}_{i}-\overline{O}\right)\right]}^{2}}{{\sum\:}_{i=1}^{N}{\left({E}_{i}-\overline{E}\right)}^{2}{\sum\:}_{i=1}^{N}{\left({O}_{i}-\overline{O}\right)}^{2}}$$ 2 $$\:RMSE=\sqrt{\frac{{\sum\:}_{i=1}^{N}{\left({E}_{i}-{O}_{i}\right)}^{2}}{N}}$$ 3 Where, \(\:{E}_{i}\) represents the estimated values and \(\:{O}_{i}\) denotes observed values. \(\:\overline{O}\) denotes the mean value among the observed values, \(\:\overline{E}\) signifies the mean value among the estimated values, and N represents total number of observations. Table 3 Parameters utilized by Genetic Algorithm (GA) for each datasets Parameters Value Population size 100 Generations 25 Mutation rate 0.1 Crossover rate 0.5 Fitness function RMSE Results The study was implemented using Python 3.9 in an environment tailored for machine learning and time series forecasting tasks. The primary libraries utilized include NumPy for numerical operations, Pandas for data handling and manipulation, Scikit-learn for data preprocessing and model evaluation, and TensorFlow/Keras for constructing and training deep learning models. For visual representation of data and results, Matplotlib and Seaborn were employed. Model development and training were carried out on a system equipped with an Intel Core i7 processor, 16 GB RAM, and an NVIDIA GeForce GTX GPU to facilitate faster computation. For performance comparison, three machine learning models are LASSO, Support Vector Regression (SVR), and Random Forest Regression (RFR) which were developed alongside the proposed GA-optimized Deep Neural Network (DNN). The hyper-parameters for all models were fine-tuned using a genetic algorithm (GA), and the implementations were carried out uniformly across the same computing environment to ensure fairness. During the training process, the following hyper-parameters were employed and the values optimized by genetic algorithm for crop yield prediction were shown in Table 4 . Stochastic Gradient Descent (SGD) with a mini-batch size of 64. Utilization of Adam optimizer with an optimized learning rate. Batch normalization is implemented before activation in all hidden layers, excluding the initial hidden layer. Training of the model extended up to a maximum of 200 iterations. Activation functions using Rectified Linear Unit (ReLU) for all neurons in the networks, except for the output layer, which remained devoid of any activation function. Implementation of L2 regularization to prevent over fitting across all hidden layers. Table 4 Optimized hyper-parameter values for crop yield prediction model Models Hyperparameters Wheat Delhi Pune Irrigated Rainfed Irrigated Rainfed DNN Number of hidden layers 2 3 2 3 Number of neurons in each layer 12 6 13 5 Dropout rate 0.01 0.001 0.01 0.001 Learning rate 0,03 0.003 0.03 0.003 L2 penalty 0.001 0.01 0.001 0.01 LASSO λ 0.4 0.5 0.2 0.5 RFR n_estimators 200 310 230 360 SVR C 0.001 0.01 0.001 0.01 Kernel rbf poly rbf rbf Table 5 compares the performance of GA optimized different machine learning models Lasso, Random Forest Regressor (RFR), Support Vector Regressor (SVR), and GA-Deep Neural Network (DNN)) across two locations (Delhi and Pune) and two environmental conditions (Irrigated and Rainfed). The model’s performances are evaluated using two metrics: Root Mean Square Error (RMSE) and R-squared (R 2 ). Delhi under irrigated conditions, the DNN has the lowest RMSE (78.26 for training and 81.16 for validation), indicating that it predicts the target variable with the least error among the models. The DNN in Pune under Irrigated conditions has an R 2 of 84.89% for training and 81.87% for validation, suggesting it explains a significant portion of variance in the data. The performance of models varies across locations and environments. For example, models tend to perform better (lower RMSE and higher R 2 ) under irrigated conditions compared to rainfed conditions in Delhi. This variation could be due to differences in the underlying data patterns and complexities in each scenario. Among all models, the DNN generally shows superior performance in both RMSE and R 2 across different conditions and locations, indicating its high predictive accuracy and ability to explain a large portion of the variance in the data. Lasso regression, while simpler, tends to have higher RMSE and lower R 2 values, suggesting it may not capture complex patterns in the data as effectively as the other models. On the other hand, the RFR model, being a non-parametric model, performed better than Lasso by capturing nonlinear effects. SVR demonstrated comparable performance due to its capability to utilize different kernel functions (linear, polynomial, radial basis function, etc.) to transform data into higher dimensions, enabling it to capture complex relationships between variables. The Training RMSE and R 2 indicate how well the model fits the training data, while the Validation RMSE and R 2 show how well the model generalizes to new, unseen data. A model with good performance on training data but poor on validation data might be over fitting. However, in our results, the proposed model shows consistent performance in both training and validation, which is a good sign. The DNN's ability to leverage deep learning architecture and optimize hyper-parameters through genetic algorithm optimization resulted in significantly higher accuracy and robustness in predicting wheat yield under various conditions. The DNN's capacity to capture intricate nonlinear relationships within the data contributed to its superior performance. These findings highlight the importance of incorporating advanced optimization techniques, such as genetic algorithms, in fine-tuning deep learning architectures. This hybrid approach can greatly enhance predictive accuracy, especially when dealing with complex agricultural datasets. Further, the probability density functions of the grain yield (observed yield) and the predicted yield were plotted to assess whether the proposed GA optimized DNN model can accurately represent the distributional properties of the observed grain yield. As shown in Fig. 4 , our model successfully approximates the distributional properties of the observed yield. However, it should be noted that the variance of the predicted yield is considerably less than that of the observed data. This implies that the GA optimized DNN model's predictions tend to be more centralized around the mean value. Table 5 Prediction performance of different models Models Location Environment Training Validation RMSE R 2 (%) RMSE R 2 (%) Lasso Delhi Irrigated 123.15 32.13 134.39 28.87 Rainfed 106.13 36.56 113.12 24.98 Pune Irrigated 159.39 34.43 182.82 24.11 Rainfed 147.78 35.12 152.38 25.09 RFR Delhi Irrigated 96.36 54.45 101.12 42.04 Rainfed 78.76 50.18 85.17 49.34 Pune Irrigated 116.63 51.65 123.89 54.88 Rainfed 54.88 57.99 58.77 51.12 SVR Delhi Irrigated 93.12 61.19 94.37 58.09 Rainfed 76.71 57.87 73.71 54.12 Pune Irrigated 79.4 63.34 72.58 68.44 Rainfed 57.76 51.65 62.36 48.87 DNN Delhi Irrigated 78.26 83.43 81.16 79.34 Rainfed 69.68 81.56 71.06 76.56 Pune Irrigated 58.78 84.89 61.89 81.87 Rainfed 48.57 79.07 50.76 77.98 Table 6 Wheat yield prediction performances of Genetic Algorithm optimized Deep Neural Network (GA-DNN) on each feature group Location Environment Model RMSE R 2 Delhi Irrigated DNN (NDVI) 77.26 61.23 DNN (CT) 129.44 17.14 DNN (PH) 132.46 15.87 Rainfed DNN (NDVI) 78.15 60.87 DNN (CT) 123.34 12.78 DNN (PH) 99.59 27.67 Pune Irrigated DNN (NDVI) 76.33 51.23 DNN (CT) 82.1 32.76 DNN (PH) 89.78 29.67 Rainfed DNN (NDVI) 56.21 50.67 DNN (CT) 87.95 19.98 DNN (PH) 92.13 13.87 Further, to assess the individual contributions of feature groups such as NDVIs, CTs, and PH in predicting crop yield, we employed GA optimized DNN model to capture linear and nonlinear effects of individual feature groups. Table 6 presents the yield prediction performance of the GA optimized DNN model using these three input variables for both irrigated and rainfed environments. By considering our results, we can conclude that NDVIs (R 2 = 77%, RMSE = 75.67 g/m 2 ) is a crucial factor in accurately predicting crop yield, surpassing the predictive power of CTs (R 2 = 56%, RMSE = 98.06 g/m 2 ) when used in isolation. This emphasizes the importance of utilizing NDVI as a significant input variable when employing predictive models for crop yield estimation. Because NDVI helps track the density and health of the developing plants, ensuring they progress uniformly and detecting stress factors that might hinder growth or yield potential (Nuarsa et al. 2011 ). NDVIs (Normalized Difference Vegetation Index) and CTs (Canopy Temperatures) are typically composed of numerous variables, but not all variables contribute equally to yield prediction. Hence, it's crucial to identify the significant variables and exclude redundant ones to maintain the accuracy of predictive models. In this study, random forest evaluates features by measuring how much decrease in prediction errors due to features when making decisions across its trees (Speiser et al. 2019 ). Features contributing more to reducing prediction error are assigned higher importance scores. The “feature_importances” attribute provides these scores, helping select the most influential features for prediction. As shown in Fig. 5 , NDVI is an important factor for grain yield prediction, which was earlier reported by Hassan et al. 2019 . Interestingly, the model selected NDVI during ground cover (early growth), anthesis (flowering) and maturity (harvest readiness) for yield prediction. High NDVI values typically correspond to robust ground cover with high vegetative growth indicating superior genotypes with high vigor’s, favorable growing conditions and often correlating with higher potential yield. Moving into the grain filling stage, NDVI becomes a reflection of the plant's photosynthetic activity and physiological health, directly impacting its ability to convert light into biomass, thereby influencing the potential grain yield. As the crop progresses toward maturity, NDVI captures changes in vegetation vigor, signaling the onset of senescence and declining photosynthetic processes. Therefore, early season prediction of grain yield may be achieved based on NDVI value using DNN model. Among abiotic stresses, drought significantly impedes wheat production, leading to substantial yield losses (Upadhyay et al. 2023 ) and this model may also be used for rainfed condition yield prediction. In wheat yield prediction, canopy temperature serves as an indicator of stress, water status, and potential productivity (Raun et al. 2001 ). Elevated temperatures during grain filling may indicate stress, impacting the duration and efficiency of grain-filling processes, potentially reducing yield, and affecting grain quality attributes like grain size and weight. Plant height indicates overall crop health, growth, and developmental stage (Rutkoski et al. 2016 ) and assists in assessing vegetative vigor, potential lodging risks, and resource allocation shifts between vegetative and reproductive growth. Thus rapid and easily scorable traits such as vegetation indices and canopy temperature may be used for early and in season identification of high yielding genotypes under optimum and stressed environments. Discussion The results clearly demonstrate that the Deep Neural Network (DNN) model optimized using a Genetic Algorithm (GA) outperforms conventional machine learning methods in predicting wheat yield. The DNN showed excellent generalization ability, delivering stable and reliable performance on both training and validation datasets. Its ability to model complex nonlinear relationships is crucial in capturing intricate interactions among vegetative and environmental factors. The variation in model performance between irrigated and rainfed conditions highlight the sensitivity of yield prediction models to environmental variability. Irrigated conditions, being more stable, generally yielded higher prediction accuracy. Nevertheless, the GA-DNN also maintained strong predictive ability under rainfed conditions, suggesting robustness and adaptability to stress-prone environments. The comparative underperformance of LASSO can be attributed to its linear structure, which is insufficient for capturing the nonlinear dynamics of crop yield data. RFR and SVR performed moderately well, with SVR benefiting from the flexibility of kernel functions to map data into higher dimensions. The use of genetic algorithms for hyper-parameter optimization significantly improved model accuracy. Traditional manual tuning often fails to explore the full hyper-parameter space efficiently, whereas GA ensured convergence towards near-optimal configurations. The superior contribution of NDVI as an input feature is biologically plausible. NDVI serves as a proxy for plant vigor, canopy density, and photosynthetic activity. High NDVI values at key growth stages i.e. ground cover, anthesis, and maturity were found to correlate with high grain yield, consistent with previous literature (e.g., Hassan et al., 2019 ). The limited predictive value of canopy temperature (CT) and plant height (PH) when used alone, suggests these features are more supportive than primary drivers of yield prediction. CT, although indicative of stress and water status, requires contextual coupling with NDVI or growth stage information for effective interpretation. Overall, the proposed model offers a scalable and practical approach for early-season and in-season yield prediction, particularly useful for breeders and researchers to identify promising genotypes under both optimal and stressed conditions. It also shows potential in addressing drought stress scenarios, given the inclusion of CT and PH alongside NDVI in predictive modeling. Conclusion Wheat improvement necessitates the evaluation of numerous genotypes across diverse environmental conditions. Physiological traits like vegetation indices (VIs) are closely linked to grain yield, and their rapid measurement can significantly aid in field-scale selection of high-yielding genotypes within a growing season. This study presents an innovative approach for field-scale wheat yield prediction by utilizing a Deep Neural Network (DNN) model optimized with a genetic algorithm, using inputs such as normalized difference vegetation indices, canopy temperature, and plant height. The proposed method demonstrated superior performance compared to other machine learning techniques, including LASSO, Random Forest Regression (RFR), and Support Vector Regression (SVR). Despite the challenges associated with selecting optimal hyper-parameters for DNNs, this research is among the first to successfully apply genetic algorithms for this purpose, marking a significant methodological advancement. The findings highlight the substantial potential of integrating proximal sensing data with genetically optimized DNN models for accurate crop yield prediction. Yet, further assessments across various crop types, seasons, developmental stages, and environmental contexts are essential to affirm its robustness. Future endeavors will explore advanced deep learning techniques like convolution and recurrent neural networks for yield prediction. Additionally, incorporating supplementary indices and agronomic traits stands as a promising avenue to augment the precision of yield prediction models. Declarations Author Contribution Statement LM and GKJ developed the proposed method of yield estimation, LM and PA conducted all statistical analyses, LM, GKJ and PA drafted the manuscript. JK and SN for data curation and compilation, GPS and PVP review and editing. All authors have read and approved the final manuscript. Funding : The first author acknowledges The Graduate School, ICAR-Indian Agricultural Research Institute, for providing all the facility throughout her doctoral studies. The field data was collected under ICAR-DBT funded Network Project “Germplasm Characterization and Trait Discovery in Wheat using Genomics Approaches and its Integration for Improving Climate Resilience, Productivity and Nutritional quality Sub Project-3: Evaluation of wheat germplasm for abiotic stresses” with the scheme code 40003 and project number 1012267. Conflict of interest: The authors declare that they have no conflict of interest References Abbas F, Afzaal H, Farooque AA, Tang S (2020) Crop yield prediction through proximal sensing and machine learning algorithms. J. Agron. 10(7). https://doi,org/10,3390/AGRONOMY10071046 Aghighi H, Azadbakht M, Ashourloo D, Shahrabi H, S, Radiom S (2018). Machine Learning Regression Techniques for the Silage Maize Yield Prediction Using Time-Series Images of Landsat 8 OLI. IEEE J Sel Top Appl Earth Obs Remote Sens 11(12): 4563–4577. https://doi,org/10,1109/JSTARS,2018,2823361 Amin Md N, Islam Md M, Coyne C, J, Carpenter-Boggs L, McGee RJ (2023) Spectral indices for characterizing lentil accessions in the dryland of Pacific Northwest. Genet. Resour. Crop Evol. 71: 167-179. https://doi,org/10,1007/s10722-023-01614-8 Aula L, Omara P, Nambi E, Oyebiyi FB, Dhillon J, Eickhoff E, Carpenter J, Raun WR (2022). Erratum to: Active optical sensor measurements and weather variables for predicting winter wheat yield. J. Agron. 114(5): 3052. https://doi,org/10,1002/agj2,21177 Basso B, Liu L (2019). Seasonal crop yield forecast: Methods applications and accuracies, Adv. Agron. 154: 201–255. https://doi,org/10,1016/bs,agron,2018,11,002 Bendig J, Yu K, Aasen H, Bolten A, Bennertz S, Broscheit J, Gnyp ML, Bareth G (2015). Combining UAV- based plant height from crop surface models visible and near infrared vegetation indices for biomass monitoring in barley. Int J Appl Earth Obs Geoinf. 39: 79–87. https://doi,org/10,1016/j,jag,2015,02,012 Breiman L (2001). Random Forests.Mach. Learn. 45: 5-32.http://dx.doi.org/10.1023/A:1010933404324 Cai Y, Guan K, Peng J, Wang S, Seifert C, Wardlow B, Li Z (2018). A high-performance and in-season classification system of field-level crop types using time-series Landsat data and a machine learning approach. Remote Sens. Environ. 210: 35–47. https://doi,org/10,1016/j,rse,2018,02,045 Cammarano D, Zha H, Wilson L, Li Y, Batchelor W, D, Miao Y (2020). A remote sensing-based approach to management zone delineation in small scale farming systems. Agronomy 10(11): 1–14. https://doi,org/10,3390/agronomy10111767 Cao J, Wang H, Li J, Wu W, Li W, Cao J, Wang H, Li J, Tian Q, Niyogi D (2022). Citation: Improving the Forecasting of Winter Wheat Yields in Northern China with Machine Learning-Dynamical Hybrid Subseasonal-to-Seasonal Ensemble Prediction. Remote Sens 14(7): 1707. https://doi,org/10,3390/rs14071707 Carew R, Smith E, G, Grant C (2009). Factors Influencing Wheat Yield and Variability: Evidence from Manitoba Canada. J. Agric. Appl. Econ. 41(3): 625–639, https://doi,org/10,1017/s1074070800003114 Elavarasan D, Vincent DR, Sharma V, Zomaya AY, Srinivasan K (2018). Forecasting yield by integrating agrarian factors and machine learning models: A survey. Comput Electron Agric. 155: 257–282. https://doi,org/10,1016/j,compag,2018,10,024 Elsayed S, Elhoweity M, Ibrahim HH, Dewir YH, Migdadi HM, Schmidhalter U (2017). Thermal imaging and passive reflectance sensing to estimate the water status and grain yield of wheat under different irrigation regimes. Agric. Water Manag. 189: 98–110. https://doi,org/10,1016/j,agwat,2017,05,001 Evans FH, Shen J (2021). Long-term hindcasts of wheat yield in fields using remotely sensed phenology climate data and machine learning. Remote Sens. 13(13): 2435. https://doi,org/10,3390/rs13132435 Fieuzal R, Marais Sicre C, Baup F (2017). Estimation of corn yield using multi-temporal optical and radar satellite data and artificial neural networks. Int J Appl Earth Obs Geoinf. 57: 14–23. https://doi,org/10,1016/j,jag,2016,12,011 Filippi P, Jones EJ, Wimalathunge NS, Somarathna SN, Pozza LE, Ugbaje SU, Jephcott TG, Paterson SE, Whelan BM, Bishop TFA (2019). An approach to forecast grain crop yield using multi-layered multi- farm data sets and machine learning. Precis. Agric. 20(5): 1015–1029. https://doi,org/10,1007/s11119- 018-09628-4 Han J, Zhang Z, Cao J, Luo Y, Zhang L, Li Z, Zhang J (2020). Prediction of winter wheat yield based on multi- source data and machine learning in China. Remote Sens. 12(2). https://doi,org/10,3390/rs12020236 Hamazaki, K., & Iwata, H. (2022). Bayesian optimization of multivariate genomic prediction models based on secondary traits for improved accuracy gains and phenotyping costs. Theor Appl Gen, 1-16. Hassan MA, Yang M, Rasheed A, Yang G, Reynolds M, Xia X, Xiao Y, He Z (2019). A rapid monitoring of NDVI across the wheat growth cycle for grain yield prediction using a multi-spectral UAV platform. Plant Sci. 282: 95–103. https://doi,org/10,1016/j,plantsci,2018,10,022 Huang CJ, Kuo PH (2018). A Deep CNN-LSTM Model for Particulate Matter (PM2.5) Forecasting in Smart Cities. Sensors 18(7): 2220. https://doi,org/10,3390/s18072220 Jaradat AA (2019). Comparative assessment of einkorn and emmer wheat phenomes: III. Phenology. Genet. Resour. Crop Evol. 66(8): 1727–1760. https://doi,org/10,1007/s10722-019-00816-3 Kuwata K, Shibasaki R (2016). Estimating Corn Yield in the United States with Modis Evi and Machine Learning Methods. ISPRS Ann. 131–136. https://doi,org/10,5194/isprsannals-iii-8-131-2016 Lokeshwari M, Jha, G. . K., Achal Lama, A. Praveenkumar, Jyoti Kumari, K. J. Yashavantha Kumar, & Rajender Parsad. (2025). Wheat yield prediction through artificial bee colony-enhanced convolutional neural network. Indian J. Genet. Plant Breed, 85(01), 78–86. https://doi.org/10.31742/ISGPB.85.1.9 McBratney A, Whelan B, Ancev T, Bouma J (2005). Future directions of precision agriculture. Precis. Agric. 6(1): 7–23. https://doi,org/10,1007/S11119-005-0681-8 Messina, C., Garcia-Abadillo, J., Powell, O., Tomura, S., Zare, A., Ganapathysubramanian, B., & Cooper, M. (2025). Toward a general framework for AI-enabled prediction in crop improvement. Theor Appl Gen, 1-15. Nuarsa IW, Nishio F, Nishio F, Hongo C (2011). Relationship between Rice Spectral and Rice Yield Using Modis Data. J. Agric. Sci. 3(2). https://doi,org/10,5539/jas,v3n2p80 Padilla FM, Gallardo M, Peña-Fleitas MT, De Souza R, Thompson RB (2018). Proximal optical sensors for nitrogen management of vegetable crops: A review. Sensors 18(7): 1–23. https://doi,org/10,3390/s18072083 Panda SS, Ames DP, Panigrahi S (2010). Application of vegetation indices for agricultural crop yield prediction using neural network techniques. Remote Sens. 2(3): 673–696. https://doi,org/10,3390/RS2030673 Pantazi XE, Moshou D, Alexandridis T, Whetton RL, Mouazen AM (2016). Wheat yield prediction using machine learning and advanced sensing techniques. Comput Electron Agric 121: 57–65. https://doi,org/10,1016/J,COMPAG,2015,11,018 Pask AJ, Pietragalla J, Mullan DM, Reynolds MP (2012). Physiological breeding II: a field guide to wheat phenotyping. CIMMYT. http://hdl.handle.net/10883/1288 Raun WR, Solie JB, Johnson GV, Stone ML, Lukina EV, Thomason WE, Schepers JS (2001). In‐Season Prediction of Potential Grain Yield in Winter Wheat Using Canopy Reflectance. J. Agron. 93(1): 131– 138. https://doi,org/10,2134/agronj2001,931131x Rischbeck P, Elsayed S, Mistele B, Barmeier G, Heil K, Schmidhalter U (2016). Data fusion of spectral thermal and canopy height parameters for improved yield prediction of drought stressed spring barley. Eur J Agron 78: 44–59. https://doi,org/10,1016/j,eja,2016,04,013 Rutkoski J, Poland J, Mondal S, Autrique E, Pérez LG, Crossa J, Reynolds M, Singh R (2016). Canopy Temperature and Vegetation Indices from High-Throughput Phenotyping Improve Accuracy of Pedigree and Genomic Selection for Grain Yield in Wheat. G3 6(9): 2799–2808. https://doi,org/10,1534/g3,116,032888 Sagan V, Maimaitijiang M, Bhadra S, Maimaitiyiming M, Brown DR, Sidike P, Fritschi FB (2021). Field-scale crop yield prediction using multi-temporal WorldView-3 and PlanetScope satellite data and deep learning. ISPRS J. Photogramm. Remote Sens. 265–281. https://doi,org/10,1016/j,isprsjprs,2021,02,008 Speiser JL, Miller ME, Tooze J, Ip E (2019). A comparison of random forest variable selection methods for classification prediction modeling. Expert Syst. Appl. 134: 93–101. https://doi,org/10,1016/j,eswa,2019,05,028 Upadhyay D, Budhlakoti N, Mishra DC, Kumari J, Gahlaut V, Chaudhary N, Padaria JC, Sareen S, Kumar S (2023). Characterization of stress-induced changes in morphological physiological and biochemical properties of Indian bread wheat ( Triticum aestivum L,) under deficit irrigation. Genet. Resour. Crop Evol. 70(8): 2353–2366. https://doi,org/10,1007/s10722-023-01693-7 Wang L, Tian Y, Yao X, Zhu Y, Cao W (2014). Predicting grain yield and protein content in wheat by fusing multi-sensor and multi-temporal remote-sensing images. Field Crops Res. 164(1): 178–188. https://doi,org/10,1016/j,fcr,2014,05,001 Wang X, Huang J, Feng Q, Yin D (2020). Winter wheat yield prediction at county level and uncertainty analysis in main wheat-producing regions of China with deep learning approaches. Remote Sens. 12(11). https://doi,org/10,3390/rs12111744 Wolanin A, Mateo-Garciá G, Camps-Valls G, Gómez-Chova L, Meroni M, Duveiller G, Liangzhi Y, Guanter L (2020). Estimating and understanding crop yields with explainable deep learning in the Indian Wheat Belt. Environ. Res. Lett. 15(2). https://doi,org/10,1088/1748-9326/ab68ac You J, Li X, Low M, Lobell D, Ermon S (2017). Deep Gaussian Process for Crop Yield Prediction Based on Remote Sensing Data. AAAI Conference on Artificial Intelligence. 31(1). https://doi.org/10.1609/aaai.v31i1.11172 Zadoks JC, Chang TT, Konzak CF (1974). A decimal code for the growth stages of cereals. Weed Res. 14(6): 415–421. https://doi,org/10,1111/j,1365-3180,1974,tb01084,x Zhang L, Zhang Z, Luo Y, Cao J, Tao F (2020). Combining optical fluorescence thermal satellite and environmental data to predict county-level maize yield in China using machine learning approaches. Remote Sens. 12(1). https://doi,org/10,3390/RS12010021 Cite Share Download PDF Status: Published Journal Publication published 25 Feb, 2026 Read the published version in Theoretical and Applied Genetics → Version 1 posted Editorial decision: Major revisions 25 Sep, 2025 Reviewers agreed at journal 28 Aug, 2025 Reviewers invited by journal 07 Aug, 2025 Editor assigned by journal 07 Jul, 2025 First submitted to journal 06 Jul, 2025 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-7061170","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":497260712,"identity":"7963541a-96e2-4acd-9867-ec917ef07576","order_by":0,"name":"M Lokeshwari","email":"","orcid":"","institution":"ICAR Indian Agricultural Statistics Research Institute","correspondingAuthor":false,"prefix":"","firstName":"M","middleName":"","lastName":"Lokeshwari","suffix":""},{"id":497260713,"identity":"ab5edaab-e77c-4b6d-b90e-eb8f715a4e19","order_by":1,"name":"GIRISH KUMAR JHA","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7klEQVRIie2RsYrCQBCGZxmYNBttI57kFSJpxS19DUNau6vEQkHYKuDbLJaRBW0OrhVszOUJUhxculsXC6uspeB+xewOzMc/ywJ4PK8IApRsY6/sagrvPa1wQExuCj2VdFcoujVOJZ6F1aHdgxDB93H5u5h+EGD1c+5QxjpIdPgFWcFzuoxUbhajNF10KVsCzSTMORhloNAonIYu5dBKELxf0+dArd1KjARlKIEVUU6sUdqtJEbRoYyy4lynQ6ZOnNDxlnh3xKaVExHssqpp1Ur0g21Vd6aU9rA/Asht7Ri3KZuHhv05pj0ej+c9+QcTrj341jmksgAAAABJRU5ErkJggg==","orcid":"https://orcid.org/0000-0002-3872-4843","institution":"Indian Agricultural Statistics Research Institute","correspondingAuthor":true,"prefix":"","firstName":"GIRISH","middleName":"KUMAR","lastName":"JHA","suffix":""},{"id":497260714,"identity":"1ec030de-3f50-4d60-bc15-49906459f9c6","order_by":2,"name":"A Praveenkumar","email":"","orcid":"","institution":"ICAR Indian Agricultural Statistics Research Institute","correspondingAuthor":false,"prefix":"","firstName":"A","middleName":"","lastName":"Praveenkumar","suffix":""},{"id":497260715,"identity":"1129a2fc-339a-4ae1-b858-8c8d0f574de1","order_by":3,"name":"Jyoti Kumari","email":"","orcid":"","institution":"NBPGR: National Bureau of Plant Genetic Resources","correspondingAuthor":false,"prefix":"","firstName":"Jyoti","middleName":"","lastName":"Kumari","suffix":""},{"id":497260716,"identity":"3008f2e7-00d5-415d-a83a-62dc1bcda48c","order_by":4,"name":"Sudhir Navathe","email":"","orcid":"","institution":"Agharkar Research Institute","correspondingAuthor":false,"prefix":"","firstName":"Sudhir","middleName":"","lastName":"Navathe","suffix":""},{"id":497260717,"identity":"8c741806-aa61-419c-9bf4-3bfa99069590","order_by":5,"name":"Gyanendra Pratap Singh","email":"","orcid":"","institution":"NBPGR: National Bureau of Plant Genetic Resources","correspondingAuthor":false,"prefix":"","firstName":"Gyanendra","middleName":"Pratap","lastName":"Singh","suffix":""},{"id":497260718,"identity":"2288c10a-5bf8-416e-9179-4a6d418b56bd","order_by":6,"name":"P V Vara Prasad","email":"","orcid":"","institution":"Kansas State University","correspondingAuthor":false,"prefix":"","firstName":"P","middleName":"V Vara","lastName":"Prasad","suffix":""}],"badges":[],"createdAt":"2025-07-07 04:34:22","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-7061170/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-7061170/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s00122-026-05166-0","type":"published","date":"2026-02-25T15:57:37+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":88974351,"identity":"8b2bbd2c-9d21-42d4-8ff8-ab1e63ef8331","added_by":"auto","created_at":"2025-08-13 10:07:58","extension":"jpeg","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":445708,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow of GA-DNN model for wheat yield prediction\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage1.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7061170/v1/0b2723a7f94a08a0ce523260.jpeg"},{"id":88975990,"identity":"5aaf3c6e-b268-4c53-a688-55a908088c8e","added_by":"auto","created_at":"2025-08-13 10:23:59","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":173928,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe general architecture of DNN model\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-7061170/v1/844727cc88cdbc2074d379fa.jpeg"},{"id":88974352,"identity":"1a72a438-2f90-4135-99b1-65d1d5718eb3","added_by":"auto","created_at":"2025-08-13 10:07:58","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":52166,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eWorkflow of genetic algorithm\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-7061170/v1/27e3ea84b5d9de485b74aed6.png"},{"id":88975163,"identity":"ce1c8366-3b1c-4d4a-b427-f30694f29fa2","added_by":"auto","created_at":"2025-08-13 10:15:58","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":140033,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eThe probability density functions of the grain yield (observed yield) and the predicted yield by the proposed model\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-7061170/v1/9c757f191f65d857ac33365a.png"},{"id":88975166,"identity":"3323e084-9a47-4bd6-8e64-7bd835b694f7","added_by":"auto","created_at":"2025-08-13 10:15:59","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":90382,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eRandom Forest for feature importance\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-7061170/v1/778d9dd50db83e512c175294.png"},{"id":103765452,"identity":"e00eca3f-ea40-4be7-8827-69f1fabe8ade","added_by":"auto","created_at":"2026-03-02 16:02:22","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1810717,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-7061170/v1/68452e18-0eef-4dc8-bc11-894cfa637400.pdf"}],"financialInterests":"","formattedTitle":"A Novel Deep Learning Framework for Field Scale Wheat Yield Prediction","fulltext":[{"header":"Key Messages","content":"\u003cp\u003eA genetic algorithm-optimized deep neural network was developed using proximal sensing data to accurately predict wheat yield at field scale, outperforming traditional machine learning models under diverse conditions.\u003c/p\u003e"},{"header":"Introduction","content":"\u003cp\u003eAccurate crop yield prediction plays a pivotal role in ensuring global food security and advancing sustainable agricultural systems\u0026nbsp;(Pantazi et al. 2016; Wang et al. 2014). Timely field-scale yield prediction significantly boosts crop production and profitability while reducing environmental degradation, resource waste and production risks (McBratney et al. 2005; Panda, Ames, and Panigrahi 2010). In addition, precise crop yield prediction enables the rapid and efficient identification of genotypes from a vast pool of potentially promising genotypes under breeding and research programs (Elsayed et al. 2017; Bendig et al. 2015). In order to address aforementioned challenges, the developments of new innovative and efficient crop yield prediction models are highly important.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eTraditional crop growth models, such as the Decision Support System for Agrotechnology Transfer (DSSAT) has been developed to improve crop yield prediction (Basso and Liu 2019). However, their high predictive power is limited by the need for climatic and management data (Wang et al. 2020; Zhang et al. 2020). Traditional statistical models, like linear regression, predict yield by building simple regression equations between dependent and independent variables (Basso and Liu 2019). These models face challenges in predicting non-linear interactions like weather, soil, fertilizer, and pests, and may cause over-fitting when multiple input variables are present (Carew et al. 2009).\u003c/p\u003e\n\u003cp\u003eMachine learning, leveraged by its data-driven capabilities, has the capacity to construct robust non-linear regression models\u0026nbsp;(Elavarasan et al. 2018). With the exponential increase in data from diverse sources, machine learning-based regression techniques such as Least Absolute Shrinkage and Selection Operator (LASSO), Artificial Neural Networks (ANN) (Fieuzal et al. 2017), Partial Least Squares Regression (PLSR) ((Rischbeck et al. 2016), Random Forest Regression (RFR) (Aghighi et al. 2018), and Support Vector Regression (SVR) (Kuwata and Shibasaki 2016) have showcased significant potential in predicting crop yields. By leveraging multi-layered neural architectures, deep learning models can extract hierarchical features from large datasets and achieve high predictive accuracy in yield estimations (You et al., 2017).\u003c/p\u003e\n\u003cp\u003eRecent studies have explored large-scale wheat yield prediction using deep learning frameworks that integrate remote sensing, soil, and climate data (Cao et al., 2022; Han et al., 2020; Wolanin et al., 2020). However, these models often rely on satellite-based data, which may lack the spatial and temporal resolution required for site-specific precision agriculture (Evans \u0026amp; Shen, 2021; Sagan et al., 2021). Moreover, satellite imagery poses challenges such as cloud interference, low revisit frequency, and extensive pre-processing requirements (Cammarano et al., 2020; Huang \u0026amp; Kuo, 2018). Proximal sensing technologies, particularly hand-held and vehicle-mounted devices, offer a compelling alternative(Hamazaki and Iwata 2022; Messina et al. 2025). These sensors such as the GreenSeeker collect real-time normalized difference vegetation index (NDVI) and canopy temperature (CT) data, which serve as proxies for plant vigor, biomass, and water status (Padilla et al., 2018; Jaradat, 2019; Amin et al., 2023). Unlike satellite sensors, proximal devices are less affected by weather conditions and provide an plot-level resolution suitable for breeding trials(M. Lokeshwari et al. 2025). Consequently, they are well-suited for integrating into field-scale crop yield prediction pipelines to inform genotype selection and resource allocation.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAccurate crop yield prediction and precision cultivation relies heavily on the application of deep learning frameworks. To obtain the best performance results of this model, choosing optimal hyper-parameters like number of layers, nodes, learning rate and loss function is crucial. The current study proposes a new hybrid approach, named GA optimized DNN, which combines genetic algorithms and deep neural networks to make precise predictions on wheat yield at a field level. Genetic algorithms simulate natural selection and evolution processes, ultimately determining the most suitable hyper-parameters for the deep neural network model. Utilizing multiple interconnected layers of neurons, deep neural networks provide the capability to recognize complex patterns, enabling better predictions. Overall, this hybrid GA optimized DNN model can offer an effective solution for accurately predicting crop yield, specifically for wheat cultivation.\u003c/p\u003e\n\u003cp\u003eIn recent years, the application of deep learning models in crop yield prediction has yielded positive results. For instance, Aula et al. (2022) enhanced the accuracy of mid-season wheat yield prediction by incorporating proximal sensing measurements and weather features into their deep neural network (DNN) model. Similarly, Abbas et al. (2020) utilized machine learning and deep learning algorithms such as linear regression, elastic net regression, k-nearest neighbor, support vector regression, and DNN to predict potato tuber yield using data on soil and crop characteristics obtained through proximal sensing. The main contributions of this study are as follows:\u003c/p\u003e\n\u003col style=\"list-style-type: lower-roman;\"\u003e\n \u003cli\u003eDevelopment of a novel GA-optimized deep neural network (GA-DNN) model that integrates proximal sensing data (NDVI, CT, and PH) for accurate wheat yield prediction at field scale under both irrigated and rainfed conditions.\u003c/li\u003e\n \u003cli\u003eEvaluation of feature importance and model performance, highlighting the predictive power of NDVI at different wheat growth stages, and comparing GA-DNN with conventional machine learning models such as Random Forest, LASSO, and SVR.\u003c/li\u003e\n\u003c/ol\u003e"},{"header":"Materials and methods","content":"\u003cp\u003e\u003cstrong\u003eData description\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis study uses real-world data to predict the performance of a diverse set of wheat germplasm under different locations as well as environments. Approximately 3350 wheat accession were grown in augmented block design during the winter season 2020\u0026ndash;2021 fewer than two environmental conditions \u003cem\u003ei.e.\u003c/em\u003e, irrigated and rainfed at ICAR-National Bureau of Plant Genetic Resources (NBPGR), Issapur Farm, New Delhi and at Agharkar Research Institute, Pune under ICAR-DBT network project. The traits measured included grain yield (GY) in gm\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e, plant height (PH) in (cm), canopy temperature (CT) in (\u0026deg;C), and normalized difference vegetation indices (NDVI). Plant height was measured as the length from ground level to the apex of the spike excluding awns. CT and NDVI data were collected during the growing seasons at different growth stages from tillering to senescence (ground cover, heading, anthesis, grain filling, and maturity) (Zadoks et al. \u003cspan class=\"CitationRef\"\u003e1974\u003c/span\u003e) using handheld Infrared Thermometer (IRT) and Handheld GreenSeeker respectively (Table \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e).\u003c/p\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003ctable id=\"Tab1\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eDescription of the dataset\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCategory\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eItem\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eDescription\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eYield\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eIrrigated wheat yield (gm\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e)\u003c/p\u003e\n \u003cp\u003eRainfed wheat yield (gm\u003csup\u003e\u0026minus;\u0026thinsp;2\u003c/sup\u003e)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eHarvested yield at two different locations under two growing environments\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNDVI\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNormalized Difference Vegetation Indices (NDVI) at five growth stages\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNDVI data were collected during ground cover, heading, anthesis, grain filling and maturity stages of two environments under two growing environments\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCT\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCanopy Temperature (CT,\u0026deg;C) at four growth stages\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCT data were collected during heading, anthesis, grain filling and maturity stage at two locations under two growing environments\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ePH\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ePlant Height (cm)\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMeasured at two locations under two growing environments\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cstrong\u003eData preprocessing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn this dataset, each attribute has its own range of measurements. To ensure precise predictions, the dataset was rescaled using Eq. (1):\u003c/p\u003e\n\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e$$\\:{X}^{{\\prime\\:}}=\\frac{X-min\\left(X\\right)}{max\\left(X\\right)-min\\left(X\\right)}$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003eWhere, X\u0026prime; is the rescaled value, X is the attributes value, min(X) is minimum of the attributes value and max(X) is the maximum of attributes value. The models are constructed during the training phase, and their predictive abilities are assessed in the testing phase. The flow diagram of the models used to predict wheat yield is shown in Fig. \u003cspan class=\"InternalRef\"\u003e1\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLeast Absolute Shrinkage and Selection Operator (LASSO)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eLASSO is a linear regression technique that incorporates regularization and feature selection. It adds a penalty equal to the absolute value of the magnitude of coefficients to the regression model, which helps in reducing over fitting. This penalty can shrink some coefficients to zero, effectively selecting a simpler model that involves only a subset of the available features. LASSO is particularly useful for datasets with many features, as it simplifies models and enhances interpretability. The regularization strength is controlled by a parameter, typically chosen through cross-validation. While LASSO helps in feature reduction and preventing over fitting, it may introduce bias, especially with large coefficients or highly correlated variables.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRandom Forest Regression (RFR)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRFR is an ensemble machine learning algorithm used for regression problems. It builds multiple decision trees during training and outputs the mean prediction of the individual trees for more accurate and stable results. RFR is effective in handling large datasets with a mix of categorical and numerical features, and it\u0026apos;s robust against over fitting. It also provides insights into feature importance, doesn\u0026apos;t require feature scaling, and can handle non-linear relationships. However, key hyper parameters like the number of trees and tree depth need careful tuning for optimal performance. RFR is widely used due to its versatility and generally strong performance across a range of regression problems.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSupport Vector Regression (SVR)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eSupport Vector Regression (SVR) is a machine learning algorithm used for regression tasks. It extends the concept of Support Vector Machines (SVM) to regression, focusing on fitting a hyper plane within a specified tolerance level to minimize errors. SVR is known for its ability to handle both linear and non-linear relationships through the use of different kernel functions. It\u0026apos;s robust against outliers and works well with high-dimensional data but requires careful tuning of hyper parameters like the regularization parameter, kernel type and epsilon value. SVR is effective in scenarios where precision is crucial and is particularly favored for its generalization capabilities.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDeep Neural Network (DNN)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA Deep Neural Network (DNN) is an advanced type of neural network that consists of multiple layers. These networks are designed to recognize intricate patterns and relationships within data, making them particularly effective for tasks such as image classification, speech recognition, natural language processing, and more. The basic outline of what constitutes a DNN and how it works:\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eMultiple Layers\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe term \u003cem\u003e\u0026quot;deep\u0026quot;\u003c/em\u003e refers to the presence of multiple layers in the network. These typically include an input layer, several hidden layers, and an output layer. The addition of more layers enables the network to learn increasingly complex data representations.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNeurons (or Nodes)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEach layer is composed of units known as neurons or nodes. These nodes perform calculations by processing input data through weighted sums and applying activation functions, before passing the output to the next layer.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eActivation Functions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo enable the network to model non-linear relationships, activation functions are applied after the weighted input is computed. Common examples include \u003cem\u003eReLU\u003c/em\u003e (Rectified Linear Unit), \u003cem\u003esigmoid\u003c/em\u003e, and \u003cem\u003etanh\u003c/em\u003e functions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eWeights and Biases\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eEvery connection between neurons carries a weight, and each neuron has an associated bias. These parameters are crucial in determining the output and are fine-tuned during the training process to optimize performance.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eForward Propagation\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn forward propagation, data moves through the network from the input layer to the output layer. At each layer, the data is transformed by weights, biases, and activation functions.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eLoss Function\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe network\u0026apos;s performance is evaluated using a loss function. This function measures the difference between the network\u0026apos;s predictions and the actual target values. Common loss functions include mean squared error for regression tasks and cross-entropy for classification tasks.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eBackpropagation and Optimization\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAfter computing the loss, backpropagation is employed to calculate how much each weight and bias contributed to the error. These gradients are then used by optimization algorithms (like Stochastic Gradient Descent or Adam) to update the weights and minimize the loss over time.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eTraining and Testing\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eDuring training, the network learns by processing input data, evaluating loss, and adjusting parameters. Once trained, the model is tested on new, unseen data to determine its effectiveness.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eRegularization and Dropout\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo ensure that the model generalizes well to new data, techniques like L1/L2 regularization and dropout (which temporarily disables certain neurons during training) are applied. These methods help avoid overfitting and improve the model\u0026rsquo;s robustness. The general architecture of DNN model is visualized in Fig. \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eGenetic algorithm for hyper-parameter tuning\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn the context of deep learning, two categories of parameters are involved. The first includes those learned directly from the data during training, namely weights and biases. The second comprises the learning algorithm parameters, commonly referred to as hyper-parameters. These include the number of neurons in each layer, the number of hidden layers, learning rate, dropout rate, and the regularization function. In this study, genetic algorithms were employed to tune the hyper-parameters of the DNN model. The range of values considered for optimization is detailed in Table \u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e. Genetic algorithms are based on the principles of natural selection and evolution, and are well-suited for solving complex problems, including hyper-parameter tuning in deep neural networks. The optimization process starts with generating a population of potential solutions, each represented as a chromosome encoding a specific set of hyper-parameters. These candidate solutions are assessed using a fitness function, which evaluates their performance on the given task. The most successful individuals, identified by their fitness scores, are selected for the next generation. New candidate solutions are then created using genetic operations such as crossover (which combines parts of two solutions) and mutation (which introduces random changes). This process continues over multiple generations, progressively enhancing the quality of solutions. Through repeated application of selection, crossover, and mutation, the genetic algorithm effectively explores the solution space and moves toward globally optimal solutions. The algorithm terminates when predefined stopping criteria are met, such as a set number of generations or a target performance level. The best solution obtained through this iterative process represents the optimized hyper-parameter configuration for the DNN, contributing to improved model performance. The overall workflow of the genetic algorithm is depicted in Fig.\u0026nbsp;3.\u003c/p\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003ctable id=\"Tab2\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eParameter space for hyper-parameter optimization\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eModels\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eHyper-parameters\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eParameter space\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eEncoding type\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" rowspan=\"5\"\u003e\n \u003cp\u003eDNN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNumber of layers\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[1\u0026ndash;20]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eInteger\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eNumber of neurons\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[8\u0026ndash;32]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eInteger\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eDropout rate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[0\u0026ndash;1]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eContinuous\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLearning rate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[0\u0026ndash;1]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eContinuous\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eL2 penalty\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[0\u0026ndash;1]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eContinuous\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eLASSO\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u0026lambda;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[0\u0026ndash;10]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eContinuous\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRFR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003en_estimators\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[100\u0026ndash;500]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eInteger\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\" rowspan=\"2\"\u003e\n \u003cp\u003eSVR\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eC\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[0\u0026ndash;10]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eContinuous\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ekernel\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e[\u0026apos;linear\u0026apos;, \u0026apos;rbf\u0026apos;, \u0026apos;poly\u0026apos;]\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCategorical\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003cp\u003eThe study utilizes a genetic algorithm and employs 10-fold cross-validation on the training data to identify the most suitable hyper-parameters. The performance of the chosen model is assessed using the test set. The selected GA parameters utilized for hyper-parameter tuning are displayed in Table \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eEvaluation criteria\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eTo evaluate the performance of the proposed wheat yield prediction models, two widely used metrics were employed: the coefficient of determination (R\u0026sup2;) and the root mean squared error (RMSE), as defined in Equations (\u003cspan class=\"InternalRef\"\u003e2\u003c/span\u003e) and (\u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e), respectively. R\u0026sup2; reflects the proportion of variability in the actual yield that is accounted for by the model\u0026rsquo;s predictions, serving as an indicator of the model\u0026rsquo;s goodness-of-fit. A value of R\u0026sup2; closer to 1.0 suggests a stronger predictive performance and a higher degree of alignment between predicted and observed yields. In contrast, RMSE quantifies the average prediction error by taking the square root of the mean squared differences between the predicted and actual values. A lower RMSE indicates higher accuracy, with predictions closely matching the true yield values. Combined, these metrics offer a robust assessment of the model\u0026rsquo;s effectiveness in capturing the complex, non-linear interactions between phenotypic traits and wheat grain yield.\u003c/p\u003e\n\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e$$\\:{R}^{2}=\\frac{{\\sum\\:}_{i=1}^{N}{\\left[\\left({E}_{i}-\\overline{E}\\right)\\left({O}_{i}-\\overline{O}\\right)\\right]}^{2}}{{\\sum\\:}_{i=1}^{N}{\\left({E}_{i}-\\overline{E}\\right)}^{2}{\\sum\\:}_{i=1}^{N}{\\left({O}_{i}-\\overline{O}\\right)}^{2}}$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\n \u003cdiv class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e$$\\:RMSE=\\sqrt{\\frac{{\\sum\\:}_{i=1}^{N}{\\left({E}_{i}-{O}_{i}\\right)}^{2}}{N}}$$\u003c/div\u003e\n \u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eWhere, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{E}_{i}\\)\u003c/span\u003e\u003c/span\u003e represents the estimated values and \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:{O}_{i}\\)\u003c/span\u003e\u003c/span\u003e denotes observed values. \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\overline{O}\\)\u003c/span\u003e\u003c/span\u003e denotes the mean value among the observed values, \u003cspan class=\"InlineEquation\"\u003e\u003cspan class=\"mathinline\"\u003e\\(\\:\\overline{E}\\)\u003c/span\u003e\u003c/span\u003e signifies the mean value among the estimated values, and N represents total number of observations.\u003c/p\u003e\n\u003cp\u003e\u003c/p\u003e\n\u003ctable id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eParameters utilized by Genetic Algorithm (GA) for each datasets\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eParameters\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eValue\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003ePopulation size\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e100\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eGenerations\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e25\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eMutation rate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.1\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eCrossover rate\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e0.5\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eFitness function\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003eRMSE\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003c/p\u003e"},{"header":"Results","content":"\u003cp\u003eThe study was implemented using Python 3.9 in an environment tailored for machine learning and time series forecasting tasks. The primary libraries utilized include \u003cem\u003eNumPy\u003c/em\u003e for numerical operations, \u003cem\u003ePandas\u003c/em\u003e for data handling and manipulation, \u003cem\u003eScikit-learn\u003c/em\u003e for data preprocessing and model evaluation, and \u003cem\u003eTensorFlow/Keras\u003c/em\u003e for constructing and training deep learning models. For visual representation of data and results, \u003cem\u003eMatplotlib\u003c/em\u003e and \u003cem\u003eSeaborn\u003c/em\u003e were employed. Model development and training were carried out on a system equipped with an Intel Core i7 processor, 16 GB RAM, and an NVIDIA GeForce GTX GPU to facilitate faster computation. For performance comparison, three machine learning models are LASSO, Support Vector Regression (SVR), and Random Forest Regression (RFR) which were developed alongside the proposed GA-optimized Deep Neural Network (DNN). The hyper-parameters for all models were fine-tuned using a genetic algorithm (GA), and the implementations were carried out uniformly across the same computing environment to ensure fairness.\u003c/p\u003e\u003cp\u003eDuring the training process, the following hyper-parameters were employed and the values optimized by genetic algorithm for crop yield prediction were shown in Table\u0026nbsp;\u003cspan refid=\"Tab4\" class=\"InternalRef\"\u003e4\u003c/span\u003e.\u003c/p\u003e\u003cp\u003e\u003col\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eStochastic Gradient Descent (SGD) with a mini-batch size of 64.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eUtilization of Adam optimizer with an optimized learning rate.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eBatch normalization is implemented before activation in all hidden layers, excluding the initial hidden layer.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eTraining of the model extended up to a maximum of 200 iterations.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eActivation functions using Rectified Linear Unit (ReLU) for all neurons in the networks, except for the output layer, which remained devoid of any activation function.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003cspan\u003e\u003cli\u003e\u003cp\u003eImplementation of L2 regularization to prevent over fitting across all hidden layers.\u003c/p\u003e\u003c/li\u003e\u003c/span\u003e\u003c/ol\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab4\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eOptimized hyper-parameter values for crop yield prediction model\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"6\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eModels\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eHyperparameters\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"4\" nameend=\"c6\" namest=\"c3\"\u003e\u003cp\u003eWheat\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c4\" namest=\"c3\"\u003e\u003cp\u003eDelhi\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003ePune\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c6\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"4\" rowspan=\"5\"\u003e\u003cp\u003eDNN\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNumber of hidden layers\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e3\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eNumber of neurons in each layer\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e6\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e5\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eDropout rate\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eLearning rate\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0,03\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.003\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.03\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.003\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eL2 penalty\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLASSO\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eλ\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.5\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.2\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.5\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\"\u003e\u003cp\u003eRFR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003en_estimators\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e200\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e310\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e230\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e360\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eSVR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eC\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e0.001\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003e0.01\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\"\u003e\u003cp\u003eKernel\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003erbf\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003epoly\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003erbf\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003erbf\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eTable\u0026nbsp;\u003cspan refid=\"Tab5\" class=\"InternalRef\"\u003e5\u003c/span\u003e compares the performance of GA optimized different machine learning models Lasso, Random Forest Regressor (RFR), Support Vector Regressor (SVR), and GA-Deep Neural Network (DNN)) across two locations (Delhi and Pune) and two environmental conditions (Irrigated and Rainfed). The model\u0026rsquo;s performances are evaluated using two metrics: Root Mean Square Error (RMSE) and R-squared (R\u003csup\u003e2\u003c/sup\u003e). Delhi under irrigated conditions, the DNN has the lowest RMSE (78.26 for training and 81.16 for validation), indicating that it predicts the target variable with the least error among the models. The DNN in Pune under Irrigated conditions has an R\u003csup\u003e2\u003c/sup\u003e of 84.89% for training and 81.87% for validation, suggesting it explains a significant portion of variance in the data. The performance of models varies across locations and environments. For example, models tend to perform better (lower RMSE and higher R\u003csup\u003e2\u003c/sup\u003e) under irrigated conditions compared to rainfed conditions in Delhi. This variation could be due to differences in the underlying data patterns and complexities in each scenario. Among all models, the DNN generally shows superior performance in both RMSE and R\u003csup\u003e2\u003c/sup\u003e across different conditions and locations, indicating its high predictive accuracy and ability to explain a large portion of the variance in the data. Lasso regression, while simpler, tends to have higher RMSE and lower R\u003csup\u003e2\u003c/sup\u003e values, suggesting it may not capture complex patterns in the data as effectively as the other models.\u003c/p\u003e\u003cp\u003eOn the other hand, the RFR model, being a non-parametric model, performed better than Lasso by capturing nonlinear effects. SVR demonstrated comparable performance due to its capability to utilize different kernel functions (linear, polynomial, radial basis function, etc.) to transform data into higher dimensions, enabling it to capture complex relationships between variables. The Training RMSE and R\u003csup\u003e2\u003c/sup\u003e indicate how well the model fits the training data, while the Validation RMSE and R\u003csup\u003e2\u003c/sup\u003e show how well the model generalizes to new, unseen data. A model with good performance on training data but poor on validation data might be over fitting. However, in our results, the proposed model shows consistent performance in both training and validation, which is a good sign. The DNN's ability to leverage deep learning architecture and optimize hyper-parameters through genetic algorithm optimization resulted in significantly higher accuracy and robustness in predicting wheat yield under various conditions. The DNN's capacity to capture intricate nonlinear relationships within the data contributed to its superior performance. These findings highlight the importance of incorporating advanced optimization techniques, such as genetic algorithms, in fine-tuning deep learning architectures. This hybrid approach can greatly enhance predictive accuracy, especially when dealing with complex agricultural datasets.\u003c/p\u003e\u003cp\u003eFurther, the probability density functions of the grain yield (observed yield) and the predicted yield were plotted to assess whether the proposed GA optimized DNN model can accurately represent the distributional properties of the observed grain yield. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e4\u003c/span\u003e, our model successfully approximates the distributional properties of the observed yield. However, it should be noted that the variance of the predicted yield is considerably less than that of the observed data. This implies that the GA optimized DNN model's predictions tend to be more centralized around the mean value.\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab5\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003ePrediction performance of different models\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"8\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eModels\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eLocation\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eEnvironment\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"3\" nameend=\"c6\" namest=\"c4\"\u003e\u003cp\u003eTraining\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c8\" namest=\"c7\"\u003e\u003cp\u003eValidation\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c5\" namest=\"c4\"\u003e\u003cp\u003eRMSE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c6\"\u003e\u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e(%)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003eRMSE\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e(%)\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eLasso\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eDelhi\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e123.15\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e32.13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e134.39\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e28.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e106.13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e36.56\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e113.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e24.98\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003ePune\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e159.39\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e34.43\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e182.82\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e24.11\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e147.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e35.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e152.38\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e25.09\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eRFR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eDelhi\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e96.36\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e54.45\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e101.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e42.04\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e78.76\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e50.18\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e85.17\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e49.34\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003ePune\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e116.63\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e51.65\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e123.89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e54.88\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e54.88\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e57.99\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e58.77\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e51.12\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eSVR\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eDelhi\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e93.12\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e61.19\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e94.37\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e58.09\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e76.71\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e57.87\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e73.71\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e54.12\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003ePune\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e79.4\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e63.34\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e72.58\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e68.44\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e57.76\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e51.65\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e62.36\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e48.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"3\" rowspan=\"4\"\u003e\u003cp\u003eDNN\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003eDelhi\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e78.26\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e83.43\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e81.16\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e79.34\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e69.68\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e81.56\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e71.06\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e76.56\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"1\" rowspan=\"2\"\u003e\u003cp\u003ePune\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e58.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e84.89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e61.89\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e81.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e48.57\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colspan=\"2\" nameend=\"c6\" namest=\"c5\"\u003e\u003cp\u003e79.07\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c7\"\u003e\u003cp\u003e50.76\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c8\"\u003e\u003cp\u003e77.98\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003e\u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab6\" border=\"1\"\u003e\u003ccaption language=\"En\"\u003e\u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\u003cdiv class=\"CaptionContent\"\u003e\u003cp\u003eWheat yield prediction performances of Genetic Algorithm optimized Deep Neural Network (GA-DNN) on each feature group\u003c/p\u003e\u003c/div\u003e\u003c/caption\u003e\u003ccolgroup cols=\"5\"\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e\u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e\u003cthead\u003e\u003ctr\u003e\u003cth align=\"left\" colname=\"c1\"\u003e\u003cp\u003eLocation\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c2\"\u003e\u003cp\u003eEnvironment\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c3\"\u003e\u003cp\u003eModel\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c4\"\u003e\u003cp\u003eRMSE\u003c/p\u003e\u003c/th\u003e\u003cth align=\"left\" colname=\"c5\"\u003e\u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e\u003c/th\u003e\u003c/tr\u003e\u003c/thead\u003e\u003ctbody\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e\u003cp\u003eDelhi\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (NDVI)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e77.26\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e61.23\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (CT)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e129.44\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e17.14\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (PH)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e132.46\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e15.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (NDVI)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e78.15\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e60.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (CT)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e123.34\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e12.78\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (PH)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e99.59\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e27.67\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c1\" morerows=\"5\" rowspan=\"6\"\u003e\u003cp\u003ePune\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eIrrigated\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (NDVI)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e76.33\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e51.23\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (CT)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e82.1\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e32.76\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (PH)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e89.78\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e29.67\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c2\" morerows=\"2\" rowspan=\"3\"\u003e\u003cp\u003eRainfed\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (NDVI)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e56.21\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e50.67\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (CT)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e87.95\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e19.98\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003ctr\u003e\u003ctd align=\"left\" colname=\"c3\"\u003e\u003cp\u003eDNN (PH)\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c4\"\u003e\u003cp\u003e92.13\u003c/p\u003e\u003c/td\u003e\u003ctd align=\"left\" colname=\"c5\"\u003e\u003cp\u003e13.87\u003c/p\u003e\u003c/td\u003e\u003c/tr\u003e\u003c/tbody\u003e\u003c/colgroup\u003e\u003c/table\u003e\u003c/div\u003e\u003c/p\u003e\u003cp\u003eFurther, to assess the individual contributions of feature groups such as NDVIs, CTs, and PH in predicting crop yield, we employed GA optimized DNN model to capture linear and nonlinear effects of individual feature groups. Table\u0026nbsp;\u003cspan refid=\"Tab6\" class=\"InternalRef\"\u003e6\u003c/span\u003e presents the yield prediction performance of the GA optimized DNN model using these three input variables for both irrigated and rainfed environments. By considering our results, we can conclude that NDVIs (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;77%, RMSE\u0026thinsp;=\u0026thinsp;75.67 g/m\u003csup\u003e2\u003c/sup\u003e) is a crucial factor in accurately predicting crop yield, surpassing the predictive power of CTs (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;56%, RMSE\u0026thinsp;=\u0026thinsp;98.06 g/m\u003csup\u003e2\u003c/sup\u003e) when used in isolation. This emphasizes the importance of utilizing NDVI as a significant input variable when employing predictive models for crop yield estimation. Because NDVI helps track the density and health of the developing plants, ensuring they progress uniformly and detecting stress factors that might hinder growth or yield potential (Nuarsa et al. \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2011\u003c/span\u003e).\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eNDVIs (Normalized Difference Vegetation Index) and CTs (Canopy Temperatures) are typically composed of numerous variables, but not all variables contribute equally to yield prediction. Hence, it's crucial to identify the significant variables and exclude redundant ones to maintain the accuracy of predictive models. In this study, random forest evaluates features by measuring how much decrease in prediction errors due to features when making decisions across its trees (Speiser et al. \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). Features contributing more to reducing prediction error are assigned higher importance scores. The \u0026ldquo;feature_importances\u0026rdquo; attribute provides these scores, helping select the most influential features for prediction. As shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e5\u003c/span\u003e, NDVI is an important factor for grain yield prediction, which was earlier reported by Hassan et al. \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2019\u003c/span\u003e. Interestingly, the model selected NDVI during ground cover (early growth), anthesis (flowering) and maturity (harvest readiness) for yield prediction. High NDVI values typically correspond to robust ground cover with high vegetative growth indicating superior genotypes with high vigor\u0026rsquo;s, favorable growing conditions and often correlating with higher potential yield. Moving into the grain filling stage, NDVI becomes a reflection of the plant's photosynthetic activity and physiological health, directly impacting its ability to convert light into biomass, thereby influencing the potential grain yield. As the crop progresses toward maturity, NDVI captures changes in vegetation vigor, signaling the onset of senescence and declining photosynthetic processes.\u003c/p\u003e\u003cp\u003e\u003c/p\u003e\u003cp\u003eTherefore, early season prediction of grain yield may be achieved based on NDVI value using DNN model. Among abiotic stresses, drought significantly impedes wheat production, leading to substantial yield losses (Upadhyay et al. \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2023\u003c/span\u003e) and this model may also be used for rainfed condition yield prediction. In wheat yield prediction, canopy temperature serves as an indicator of stress, water status, and potential productivity (Raun et al. \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2001\u003c/span\u003e). Elevated temperatures during grain filling may indicate stress, impacting the duration and efficiency of grain-filling processes, potentially reducing yield, and affecting grain quality attributes like grain size and weight. Plant height indicates overall crop health, growth, and developmental stage (Rutkoski et al. \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2016\u003c/span\u003e) and assists in assessing vegetative vigor, potential lodging risks, and resource allocation shifts between vegetative and reproductive growth. Thus rapid and easily scorable traits such as vegetation indices and canopy temperature may be used for early and in season identification of high yielding genotypes under optimum and stressed environments.\u003c/p\u003e"},{"header":"Discussion","content":"\u003cp\u003eThe results clearly demonstrate that the Deep Neural Network (DNN) model optimized using a Genetic Algorithm (GA) outperforms conventional machine learning methods in predicting wheat yield. The DNN showed excellent generalization ability, delivering stable and reliable performance on both training and validation datasets. Its ability to model complex nonlinear relationships is crucial in capturing intricate interactions among vegetative and environmental factors. The variation in model performance between irrigated and rainfed conditions highlight the sensitivity of yield prediction models to environmental variability. Irrigated conditions, being more stable, generally yielded higher prediction accuracy. Nevertheless, the GA-DNN also maintained strong predictive ability under rainfed conditions, suggesting robustness and adaptability to stress-prone environments.\u003c/p\u003e\u003cp\u003eThe comparative underperformance of LASSO can be attributed to its linear structure, which is insufficient for capturing the nonlinear dynamics of crop yield data. RFR and SVR performed moderately well, with SVR benefiting from the flexibility of kernel functions to map data into higher dimensions. The use of genetic algorithms for hyper-parameter optimization significantly improved model accuracy. Traditional manual tuning often fails to explore the full hyper-parameter space efficiently, whereas GA ensured convergence towards near-optimal configurations.\u003c/p\u003e\u003cp\u003eThe superior contribution of NDVI as an input feature is biologically plausible. NDVI serves as a proxy for plant vigor, canopy density, and photosynthetic activity. High NDVI values at key growth stages i.e. ground cover, anthesis, and maturity were found to correlate with high grain yield, consistent with previous literature (e.g., Hassan et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The limited predictive value of canopy temperature (CT) and plant height (PH) when used alone, suggests these features are more supportive than primary drivers of yield prediction. CT, although indicative of stress and water status, requires contextual coupling with NDVI or growth stage information for effective interpretation.\u003c/p\u003e\u003cp\u003eOverall, the proposed model offers a scalable and practical approach for early-season and in-season yield prediction, particularly useful for breeders and researchers to identify promising genotypes under both optimal and stressed conditions. It also shows potential in addressing drought stress scenarios, given the inclusion of CT and PH alongside NDVI in predictive modeling.\u003c/p\u003e"},{"header":"Conclusion","content":"\u003cp\u003eWheat improvement necessitates the evaluation of numerous genotypes across diverse environmental conditions. Physiological traits like vegetation indices (VIs) are closely linked to grain yield, and their rapid measurement can significantly aid in field-scale selection of high-yielding genotypes within a growing season. This study presents an innovative approach for field-scale wheat yield prediction by utilizing a Deep Neural Network (DNN) model optimized with a genetic algorithm, using inputs such as normalized difference vegetation indices, canopy temperature, and plant height. The proposed method demonstrated superior performance compared to other machine learning techniques, including LASSO, Random Forest Regression (RFR), and Support Vector Regression (SVR). Despite the challenges associated with selecting optimal hyper-parameters for DNNs, this research is among the first to successfully apply genetic algorithms for this purpose, marking a significant methodological advancement. The findings highlight the substantial potential of integrating proximal sensing data with genetically optimized DNN models for accurate crop yield prediction.\u003c/p\u003e\u003cp\u003eYet, further assessments across various crop types, seasons, developmental stages, and environmental contexts are essential to affirm its robustness. Future endeavors will explore advanced deep learning techniques like convolution and recurrent neural networks for yield prediction. Additionally, incorporating supplementary indices and agronomic traits stands as a promising avenue to augment the precision of yield prediction models.\u003c/p\u003e\u003cp\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthor Contribution Statement\u0026nbsp;\u003c/strong\u003eLM and GKJ developed the proposed method of yield estimation, LM and PA conducted all statistical analyses, LM, GKJ and PA drafted the manuscript. JK and SN for data curation and compilation, GPS and PVP review and editing. All authors have read and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e: The first author acknowledges The Graduate School, ICAR-Indian Agricultural Research Institute, for providing all the facility throughout her doctoral studies. The field data was collected under ICAR-DBT funded Network Project \u0026ldquo;Germplasm Characterization and Trait Discovery in Wheat using Genomics Approaches and its Integration for Improving Climate Resilience, Productivity and Nutritional quality Sub Project-3: Evaluation of wheat germplasm for abiotic stresses\u0026rdquo; with the scheme code 40003 and project number 1012267.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eConflict of interest:\u0026nbsp;\u003c/strong\u003e The authors declare that they have no conflict of interest\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n \u003cli\u003eAbbas F, Afzaal H, Farooque AA, Tang S (2020) Crop yield prediction through proximal sensing and machine learning algorithms. J. Agron. 10(7). https://doi,org/10,3390/AGRONOMY10071046\u003c/li\u003e\n \u003cli\u003eAghighi H, Azadbakht M, Ashourloo D, Shahrabi H, S, Radiom S (2018). Machine Learning Regression \u0026nbsp; \u0026nbsp;\u0026nbsp;Techniques for the Silage Maize Yield Prediction Using Time-Series Images of Landsat 8 OLI. IEEE J \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Sel Top Appl Earth Obs Remote Sens 11(12): 4563\u0026ndash;4577. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;https://doi,org/10,1109/JSTARS,2018,2823361\u003c/li\u003e\n \u003cli\u003eAmin Md N, Islam Md M, Coyne C, J, Carpenter-Boggs L, McGee RJ (2023) Spectral indices for characterizing \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;lentil accessions in the dryland of Pacific Northwest. Genet. Resour. Crop Evol. 71: 167-179. \u0026nbsp; \u0026nbsp;https://doi,org/10,1007/s10722-023-01614-8\u003c/li\u003e\n \u003cli\u003eAula L, Omara P, Nambi E, Oyebiyi FB, Dhillon J, Eickhoff E, Carpenter J, Raun WR (2022). Erratum to: \u0026nbsp;Active optical sensor measurements and weather variables for predicting winter wheat yield. J. Agron. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;114(5): 3052. https://doi,org/10,1002/agj2,21177\u003c/li\u003e\n \u003cli\u003eBasso B, Liu L (2019). Seasonal crop yield forecast: Methods applications and accuracies, Adv. Agron. 154: \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;201\u0026ndash;255. https://doi,org/10,1016/bs,agron,2018,11,002\u003c/li\u003e\n \u003cli\u003eBendig J, Yu K, Aasen H, Bolten A, Bennertz S, Broscheit J, Gnyp ML, Bareth G (2015). Combining UAV-\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;based plant height from crop surface models visible and near infrared vegetation indices for biomass \u0026nbsp;monitoring in barley. Int J Appl Earth Obs Geoinf. 39: 79\u0026ndash;87. https://doi,org/10,1016/j,jag,2015,02,012\u003c/li\u003e\n \u003cli\u003eBreiman L (2001). Random Forests.Mach. Learn. 45:\u0026nbsp;5-32.http://dx.doi.org/10.1023/A:1010933404324\u003c/li\u003e\n \u003cli\u003eCai Y, Guan K, Peng J, Wang S, Seifert C, Wardlow B, Li Z (2018). A high-performance and in-season \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;classification system of field-level crop types using time-series Landsat data and a machine learning \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;approach. Remote Sens. Environ. 210: 35\u0026ndash;47. https://doi,org/10,1016/j,rse,2018,02,045\u003c/li\u003e\n \u003cli\u003eCammarano D, Zha H, Wilson L, Li Y, Batchelor W, D, Miao Y (2020). A remote sensing-based approach to \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;management zone delineation in small scale farming systems. Agronomy 10(11): 1\u0026ndash;14. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;https://doi,org/10,3390/agronomy10111767\u003c/li\u003e\n \u003cli\u003eCao J, Wang H, Li J, Wu W, Li W, Cao J, Wang H, Li J, Tian Q, Niyogi D (2022). Citation: Improving the \u0026nbsp;Forecasting of Winter Wheat Yields in Northern China with Machine Learning-Dynamical Hybrid \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Subseasonal-to-Seasonal Ensemble Prediction. Remote Sens 14(7): 1707. \u0026nbsp;https://doi,org/10,3390/rs14071707\u003c/li\u003e\n \u003cli\u003eCarew R, Smith E, G, Grant C (2009). Factors Influencing Wheat Yield and Variability: Evidence from \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Manitoba Canada. J. Agric. Appl. Econ. 41(3): 625\u0026ndash;639, https://doi,org/10,1017/s1074070800003114\u003c/li\u003e\n \u003cli\u003eElavarasan D, Vincent DR, Sharma V, Zomaya AY, Srinivasan K (2018). Forecasting yield by integrating \u0026nbsp;\u0026nbsp;agrarian factors and machine learning models: A survey. Comput Electron Agric. 155: 257\u0026ndash;282. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;https://doi,org/10,1016/j,compag,2018,10,024\u003c/li\u003e\n \u003cli\u003eElsayed S, Elhoweity M, Ibrahim HH, Dewir YH, Migdadi HM, Schmidhalter U (2017). Thermal imaging and \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;passive reflectance sensing to estimate the water status and grain yield of wheat under different \u0026nbsp; \u0026nbsp;irrigation regimes. Agric. Water Manag. 189: 98\u0026ndash;110. https://doi,org/10,1016/j,agwat,2017,05,001\u003c/li\u003e\n \u003cli\u003eEvans FH, Shen J (2021). Long-term hindcasts of wheat yield in fields using remotely sensed phenology climate \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;data and machine learning. Remote Sens. 13(13): 2435. \u0026nbsp; https://doi,org/10,3390/rs13132435\u003c/li\u003e\n \u003cli\u003eFieuzal R, Marais Sicre C, Baup F (2017). Estimation of corn yield using multi-temporal optical and radar \u0026nbsp;satellite data and artificial neural networks. Int J Appl Earth Obs Geoinf. 57: 14\u0026ndash;23. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;https://doi,org/10,1016/j,jag,2016,12,011\u003c/li\u003e\n \u003cli\u003eFilippi P, Jones EJ, Wimalathunge NS, Somarathna SN, Pozza LE, Ugbaje SU, Jephcott TG, Paterson SE, Whelan BM, Bishop TFA (2019). An approach to forecast grain crop yield using multi-layered multi- farm data sets and machine learning. Precis. Agric. 20(5): 1015\u0026ndash;1029. https://doi,org/10,1007/s11119- 018-09628-4\u003c/li\u003e\n \u003cli\u003eHan J, Zhang Z, Cao J, Luo Y, Zhang L, Li Z, Zhang J (2020). Prediction of winter wheat yield based on multi- source data and machine learning in China. Remote Sens. 12(2). https://doi,org/10,3390/rs12020236\u003c/li\u003e\n \u003cli\u003eHamazaki, K., \u0026amp; Iwata, H. (2022). Bayesian optimization of multivariate genomic prediction models based on \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;secondary traits for improved accuracy gains and phenotyping costs. Theor Appl Gen, \u0026nbsp; \u0026nbsp;1-16.\u003c/li\u003e\n \u003cli\u003eHassan MA, Yang M, Rasheed A, Yang G, Reynolds M, Xia X, Xiao Y, He Z (2019). A rapid monitoring of \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;NDVI across the wheat growth cycle for grain yield prediction using a multi-spectral UAV platform. \u0026nbsp; \u0026nbsp;\u0026nbsp;Plant Sci. 282: 95\u0026ndash;103. https://doi,org/10,1016/j,plantsci,2018,10,022\u003c/li\u003e\n \u003cli\u003eHuang CJ, Kuo PH (2018). A Deep CNN-LSTM Model for Particulate Matter (PM2.5) Forecasting in Smart \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Cities. Sensors 18(7): 2220. https://doi,org/10,3390/s18072220\u003c/li\u003e\n \u003cli\u003eJaradat AA (2019). Comparative assessment of einkorn and emmer wheat phenomes: III. \u0026nbsp;Phenology. Genet. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Resour. Crop Evol. 66(8): 1727\u0026ndash;1760. https://doi,org/10,1007/s10722-019-00816-3\u003c/li\u003e\n \u003cli\u003eKuwata K, Shibasaki R (2016). Estimating Corn Yield in the United States with Modis Evi and Machine Learning Methods. ISPRS Ann. 131\u0026ndash;136. https://doi,org/10,5194/isprsannals-iii-8-131-2016\u003c/li\u003e\n \u003cli\u003eLokeshwari M, Jha, G. . K., Achal Lama, A. Praveenkumar, Jyoti Kumari, K. J. Yashavantha Kumar, \u0026amp;\u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Rajender Parsad. (2025). Wheat yield prediction through artificial bee colony-enhanced convolutional \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;neural network. Indian J. Genet.\u0026nbsp;Plant Breed, 85(01), 78\u0026ndash;86. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;https://doi.org/10.31742/ISGPB.85.1.9\u003c/li\u003e\n \u003cli\u003eMcBratney A, Whelan B, Ancev T, Bouma J (2005). Future directions of precision agriculture. Precis. Agric. 6(1): 7\u0026ndash;23. https://doi,org/10,1007/S11119-005-0681-8\u003c/li\u003e\n \u003cli\u003eMessina, C., Garcia-Abadillo, J., Powell, O., Tomura, S., Zare, A., Ganapathysubramanian, B., \u0026amp; Cooper, M. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;(2025). Toward a general framework for AI-enabled prediction in crop improvement. Theor Appl Gen, \u0026nbsp; \u0026nbsp;\u0026nbsp;1-15.\u003c/li\u003e\n \u003cli\u003eNuarsa IW, Nishio F, Nishio F, Hongo C (2011). Relationship between Rice Spectral and Rice Yield Using \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Modis Data. J. Agric. Sci. 3(2). https://doi,org/10,5539/jas,v3n2p80\u003c/li\u003e\n \u003cli\u003ePadilla FM, Gallardo M, Pe\u0026ntilde;a-Fleitas MT, De Souza R, Thompson RB (2018). Proximal optical sensors for \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;nitrogen management of vegetable crops: A review. Sensors 18(7): 1\u0026ndash;23. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;https://doi,org/10,3390/s18072083\u003c/li\u003e\n \u003cli\u003ePanda SS, Ames DP, Panigrahi S (2010). Application of vegetation indices for agricultural crop yield prediction \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;using neural network techniques. Remote Sens. 2(3): 673\u0026ndash;696. https://doi,org/10,3390/RS2030673\u003c/li\u003e\n \u003cli\u003ePantazi XE, Moshou D, Alexandridis T, Whetton RL, Mouazen AM (2016). Wheat yield prediction using \u0026nbsp;\u0026nbsp;machine learning and advanced sensing techniques. Comput Electron Agric 121: 57\u0026ndash;65. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;https://doi,org/10,1016/J,COMPAG,2015,11,018\u003c/li\u003e\n \u003cli\u003ePask AJ, Pietragalla J, Mullan DM, Reynolds MP (2012). Physiological breeding II: a field guide to wheat \u0026nbsp;phenotyping. CIMMYT. http://hdl.handle.net/10883/1288\u003c/li\u003e\n \u003cli\u003eRaun WR, Solie JB, Johnson GV, Stone ML, Lukina EV, Thomason WE, Schepers JS (2001). In‐Season \u0026nbsp; \u0026nbsp;\u0026nbsp;Prediction of Potential Grain Yield in Winter Wheat Using Canopy Reflectance. J. Agron. 93(1): 131\u0026ndash;\u0026nbsp;\u0026nbsp;138. https://doi,org/10,2134/agronj2001,931131x\u003c/li\u003e\n \u003cli\u003eRischbeck P, Elsayed S, Mistele B, Barmeier G, Heil K, Schmidhalter U (2016). Data fusion of spectral thermal \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;and canopy height parameters for improved yield prediction of drought stressed spring barley. Eur J \u0026nbsp;\u0026nbsp;Agron 78: 44\u0026ndash;59. https://doi,org/10,1016/j,eja,2016,04,013\u003c/li\u003e\n \u003cli\u003eRutkoski J, Poland J, Mondal S, Autrique E, P\u0026eacute;rez LG, Crossa J, Reynolds M, Singh R (2016). Canopy \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Temperature and Vegetation Indices from High-Throughput Phenotyping Improve Accuracy of \u0026nbsp; \u0026nbsp;\u0026nbsp;Pedigree and Genomic Selection for Grain Yield in Wheat. G3 6(9): 2799\u0026ndash;2808. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;https://doi,org/10,1534/g3,116,032888\u003c/li\u003e\n \u003cli\u003eSagan V, Maimaitijiang M, Bhadra S, Maimaitiyiming M, Brown DR, Sidike P, Fritschi FB (2021). Field-scale \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;crop yield prediction using multi-temporal WorldView-3 and PlanetScope satellite data and deep \u0026nbsp; \u0026nbsp; \u0026nbsp;learning. ISPRS J. Photogramm. Remote Sens. 265\u0026ndash;281. \u0026nbsp; \u0026nbsp;https://doi,org/10,1016/j,isprsjprs,2021,02,008\u003c/li\u003e\n \u003cli\u003eSpeiser JL, Miller ME, Tooze J, Ip E (2019). A comparison of random forest variable selection methods for \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;classification prediction modeling. Expert Syst. Appl. 134: 93\u0026ndash;101. \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;https://doi,org/10,1016/j,eswa,2019,05,028\u003c/li\u003e\n \u003cli\u003eUpadhyay D, Budhlakoti N, Mishra DC, Kumari J, Gahlaut V, Chaudhary N, Padaria JC, Sareen S, Kumar S \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;(2023). Characterization of stress-induced changes in morphological physiological and biochemical \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;properties of Indian bread wheat (\u003cem\u003eTriticum aestivum\u003c/em\u003e L,) under deficit irrigation. Genet. Resour. Crop \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;Evol. 70(8): 2353\u0026ndash;2366. https://doi,org/10,1007/s10722-023-01693-7\u003c/li\u003e\n \u003cli\u003eWang L, Tian Y, Yao X, Zhu Y, Cao W (2014). Predicting grain yield and protein content in wheat by fusing \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;multi-sensor and multi-temporal remote-sensing images. Field Crops Res. 164(1): 178\u0026ndash;188. \u0026nbsp;\u0026nbsp;https://doi,org/10,1016/j,fcr,2014,05,001\u003c/li\u003e\n \u003cli\u003eWang X, Huang J, Feng Q, Yin D (2020). Winter wheat yield prediction at county level and uncertainty analysis \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;in main wheat-producing regions of China with deep learning approaches. Remote Sens. 12(11). \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;https://doi,org/10,3390/rs12111744\u003c/li\u003e\n \u003cli\u003eWolanin A, Mateo-Garci\u0026aacute; G, Camps-Valls G, G\u0026oacute;mez-Chova L, Meroni M, Duveiller G, Liangzhi Y, Guanter L \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;(2020). Estimating and understanding crop yields with explainable deep learning in the Indian Wheat Belt. Environ. Res. Lett. 15(2). https://doi,org/10,1088/1748-9326/ab68ac\u003c/li\u003e\n \u003cli\u003eYou J, Li X, Low M, Lobell D, Ermon S (2017). Deep Gaussian Process for Crop Yield Prediction Based on \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;Remote Sensing Data. AAAI Conference on Artificial Intelligence. 31(1). \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;https://doi.org/10.1609/aaai.v31i1.11172\u003c/li\u003e\n \u003cli\u003eZadoks JC, Chang TT, Konzak CF (1974). A decimal code for the growth stages of cereals. Weed Res. 14(6): \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp;\u0026nbsp;415\u0026ndash;421. https://doi,org/10,1111/j,1365-3180,1974,tb01084,x\u003c/li\u003e\n \u003cli\u003eZhang L, Zhang Z, Luo Y, Cao J, Tao F (2020). Combining optical fluorescence thermal satellite and \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; \u0026nbsp; environmental data to predict county-level maize yield in China using machine learning approaches. \u0026nbsp; \u0026nbsp; Remote Sens. 12(1). https://doi,org/10,3390/RS12010021\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"theoretical-and-applied-genetics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"taag","sideBox":"Learn more about [Theoretical and Applied Genetics](https://www.springer.com/journal/122)","snPcode":"122","submissionUrl":"https://submission.nature.com/new-submission/122/3","title":"Theoretical and Applied Genetics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"agricultural prediction, deep neural network, genetic algorithm, spectral vegetation indices","lastPublishedDoi":"10.21203/rs.3.rs-7061170/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-7061170/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eHand-held or vehicle-mounted active proximal sensing technologies offer a rapid, non-destructive method for real-time crop monitoring through spectral vegetation indices. This study integrates such proximal sensing data into a deep learning framework for field-scale wheat yield prediction. Specifically, wheat yield is predicted using normalized difference vegetation indices (NDVIs), canopy temperatures (CTs), and plant height (PH) through a deep neural network (DNN) optimized using a genetic algorithm (GA). The model is trained on data from 3,350 diverse wheat germplasm grown under irrigated and rainfed conditions at two locations during the 2020–21 winter season. Comparative analysis demonstrates that the GA-optimized DNN outperforms traditional machine learning models such as Random Forest Regression (RFR), Least Absolute Shrinkage and Selection Operator (LASSO), and Support Vector Regression (SVR). Among individual feature groups, NDVIs measured at five wheat growth stages showing strong predictive capability, with R² values ≥60% under irrigated and ≥50% under rainfed conditions. Additionally, RFR is employed to identify the most influential features within each group. This pioneering study introduces the first-ever application of a GA-optimized deep neural network, leveraging handheld or vehicle-mounted proximal sensing data for predicting crop yield, in the context of Indian agriculture. The proposed approach offers a robust and scalable solution for pre-harvest yield estimation, supporting breeders and researchers in efficient genotype selection and contributing to the achievement of sustainable development goals.\u003c/p\u003e","manuscriptTitle":"A Novel Deep Learning Framework for Field Scale Wheat Yield Prediction","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2025-08-13 10:07:54","doi":"10.21203/rs.3.rs-7061170/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"decision","content":"Major revisions","date":"2025-09-25T09:45:27+00:00","index":"","fulltext":""},{"type":"reviewerAgreed","content":"","date":"2025-08-29T02:39:02+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2025-08-07T15:35:09+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2025-07-07T08:14:20+00:00","index":"","fulltext":""},{"type":"submitted","content":"Theoretical and Applied Genetics","date":"2025-07-07T00:34:15+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"theoretical-and-applied-genetics","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"taag","sideBox":"Learn more about [Theoretical and Applied Genetics](https://www.springer.com/journal/122)","snPcode":"122","submissionUrl":"https://submission.nature.com/new-submission/122/3","title":"Theoretical and Applied Genetics","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"8987ff5c-db52-407a-9595-dc544ab29c17","owner":[],"postedDate":"August 13th, 2025","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2026-03-02T16:00:48+00:00","versionOfRecord":{"articleIdentity":"rs-7061170","link":"https://doi.org/10.1007/s00122-026-05166-0","journal":{"identity":"theoretical-and-applied-genetics","isVorOnly":false,"title":"Theoretical and Applied Genetics"},"publishedOn":"2026-02-25 15:57:37","publishedOnDateReadable":"February 25th, 2026"},"versionCreatedAt":"2025-08-13 10:07:54","video":"","vorDoi":"10.1007/s00122-026-05166-0","vorDoiUrl":"https://doi.org/10.1007/s00122-026-05166-0","workflowStages":[]},"version":"v1","identity":"rs-7061170","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-7061170","identity":"rs-7061170","version":["v1"]},"buildId":"8U1c8b4HqxoKbykW_rLl7","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.