Parsimonious machine learning models to estimate environmental footprints of crop production | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Parsimonious machine learning models to estimate environmental footprints of crop production Farhang Raymand, Koen J.J. Kuipers, Sarah Sim, P. James Joyce, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-9214313/v1 This work is licensed under a CC BY 4.0 License Status: Under Review Version 1 posted 7 You are reading this latest preprint version Abstract Crop production is a major driver of anthropogenic impacts on the environment. Quantifying these impacts requires primary farm-level life cycle inventory data, which are sparse and difficult to collect. Therefore, parsimonious models are needed to predict environmental footprints of crop production. Here, we test the ability of five machine learning modelling techniques (Random Forest (RF), K-Nearest Neighbours (KNN), Artificial Neural Network (ANN), Generalised Boosting Method (GBM), and Linear Modelling (LM)) to predict biodiversity, water and climate footprints of crop production based on limited aggregated farming information, using life cycle data from 121 crops across 57 countries. We found RF to be the most parsimonious model with four predictors for biodiversity (R 2 = 0.88), GBM using five predictors for water (R 2 = 0.89) and ANN for climate footprint (R 2 = 0.58) with four predictors. Uncertainty in model predictions is +/- a factor 2.1, 5.1, and 4.0 (95% confidence interval) for the three footprints, respectively. Key farming information to predict biodiversity and climate footprints are yield, fertiliser, electricity and climatic region. Irrigation, fertiliser and pesticide are important predictors of water footprint. Our study offers predictive models highlighting key predictors of environmental footprints of crop production, to be prioritised for data gap filling. crop production machine learning life cycle assessment biodiversity footprint climate footprint predictor importance Figures Figure 1 Figure 2 Figure 3 1. Introduction The agricultural sector covers about 37% of land on earth (FAO, 2022 ). It is a major driver of biodiversity loss (Norris, 2008 ; Willet et al., 2019 ), and it is responsible for about 30% of all global greenhouse gas (GHG) emissions (Laborde et al., 2021 ; Poore & Nemecek, 2018 ; Vermeulen et al., 2012 ). Moreover, agriculture is responsible for more than 70% of global water consumption, mostly due to crop irrigation (FAO, 2022 ; Hoogeveen et al., 2015 ). With a global population expected to exceed 9.7 billion people by 2050 (UN, 2022), the adverse environmental impacts of the agricultural sector are expected to further increase. Environmental impacts of agricultural crop production can be quantified using Life Cycle Assessment (LCA). LCA can be used to identify agricultural practices and inputs that contribute most to environmental impacts, allowing decision makers to identify improvement options while considering potential burden shifts (Borghino et al., 2021 ; Algunaibet & Guillén-Gosálbez, 2019). However, LCA studies require considerable input data, collection of which is time-consuming (Bellon-Maurel et al., 2014 ; Zhao et al., 2025 ). Where data are lacking, LCA practitioners can either collect new information, extrapolate information from similar crops (e.g. obtain average fuel input of wheat products to estimate fuel consumption of soft wheat) or use proxies (e.g. use electricity input of tomatoes cultivated in greenhouses to represent those of peppers cultivated in greenhouse) (Zargar et al., 2022 ; Canals et al., 2011 ). More recently, advances in data-driven tools have enabled the utilisation of machine learning algorithms to support LCA applications such as data collection and estimating environmental footprints, as these algorithms can enable automation, prediction through discovery and learning from hidden patterns (Romeiko et al., 2023 , Donati et al., 2022 ). These studies, however, have predominantly focused on single crops, such as rice, wheat, sugarcane, or corn, and employed a relatively narrow range of modelling techniques, without a systematic comparison of alternative modelling techniques (Khanali et al., 2017; Khoshnevisan et al., 2013 ; Lam et al., 2021 ; Pishgar-Komleh et al., 2017 ; Pishgar-Komleh et al., 2020 ; Nabavi-Pelesaraei et al., 2018 ; Lee et al., 2020 ). Moreover, these studies have not evaluated whether a set of key predictors is sufficient to predict environmental footprints accurately (Khanali et al., 2017; Khoshnevisan et al., 2013 ; Lam et al., 2021 ; Pishgar-Komleh et al., 2017 ; Pishgar-Komleh et al., 2020 ; Nabavi-Pelesaraei et al., 2018 ; Lee et al., 2020 ). An analysis to find the most parsimonious model (i.e. modelling footprints well while using the fewest on-farm, foreground farming information) capable of reliably estimating impacts is yet missing. Here, we aim to develop a predictive model capable of estimating life cycle crop production environmental footprints (cradle-to-gate) based on on-farm information only. To this end, we evaluate (i) the extent to which different machine learning modelling techniques can be reliably used to predict life cycle environmental footprints of crop production in terms of biodiversity loss, water use and climate change, based on data related to a limited set of on-farm information concerning climate region and crop type as well as key inputs of material, land area, energy and water; and (ii) identify the most parsimonious models and the farming information that contribute most to their predictive power. We focus on biodiversity, water and climate, as these footprints are particularly important in food production and agricultural systems (Laborde et al., 2021 ; Hoogeveen et al., 2015 ; Willet et al., 2019 ; Jacob-Lopes et al., 2021 ; Dudley & Alexander, 2017 ). We first obtained biodiversity, water, and climate footprints of 991 crops available in common life cycle inventory (LCI) databases. We then applied five machine learning modelling techniques; Random Forest (RF), K-Nearest Neighbours (KNN), Artificial Neural Network (ANN), Generalised Boosting Method (GBM) and a regularised Linear Model (LM), to predict these three footprints and identify the most parsimonious model per footprint based on an initial set of data for 16 farming information types. In addition, we identify the most important farming information driving model predictions. 2. Materials and Methods 2.1 General Approach Our general approach for modelling environmental footprints and determining the relative importance of farming information in terms of predictive power consists of three steps (Fig. 1 ). First, we quantified life cycle biodiversity, water and climate footprints (cradle-to-gate) of food crops available in common (agricultural) life cycle inventory (LCI) databases (see section 2.2 ). Second, we grouped on-farm information from the LCIs into a set of 16 predictors based on their purpose and unit of measurement (see section 2.3 ) and applied five machine learning modelling techniques to predict the three environmental footprints based on the selected predictors (see section 2.4 ). We compared the performance of the machine learning modelling techniques via cross-validation. Third, we identified which predictors are most relevant to include, and we assessed the trade-off between utilizing fewer predictors and change in predictive performance (see section 2.5 ). 2.2 Environmental footprints We used the ecoinvent database (version 3.9.1, cut-off allocation system model), agribalyse database (version 3.11) and World Food LCA Database (WFLDB) (version 3.10) to obtain LCIs for available agricultural crop production activities (Wernet et al., 2016 ; Asselin- Balençon et al., 2020 ; Nemecek et al., 2019). We retrieved LCIs for 991 crop production activities representing 121 crop species cultivated in 57 regions (see Table S1 and Figures S2 and S3 of the supplementary information for further details). We normalised all inputs per 1 kilogram of produced crop, i.e., the functional unit. We considered impacts from the extraction of raw materials up to the farm gate. We used the ReCiPe 2016 methodology to quantify the biodiversity, water, and climate footprints (Huijbregts et al., 2017 ). We chose the endpoint indicator damage to ecosystems (referred to as biodiversity footprint in this work) [species-yr] which includes terrestrial, marine and freshwater biodiversity loss. We also used two midpoint indicators, namely water use [m 3 ] and climate change [kg CO 2 eq], referred to as respectively water footprint and climate footprint in this work. We used the Brightway 2.5 Python package to calculate footprints for each ReCiPe 2016 perspective (Hierarchist, Individualist, Egalitarian) and to retrieve the crop LCIs (Mutel, 2017 ). Here, we show the results based on the Hierarchist perspective (Section 3 ), as within the ReCiPe documentation and literature, the Hierarchist perspective is the “default” for LCA calculations (Huijbregts et al., 2017 ). The results based on the Individualist and Egalitarian perspectives can be found in sections S8 and S9 of the supplementary information. 2.3 Predictor selection and preprocessing We selected on-farm information concerning climate region and crop type as well as key inputs of material, land area, energy and water as predictors and grouped them based on their function (i.e., used towards the same goal) and form (i.e., unit). For instance, calcium ammonium and urea both serve as nitrogen fertilisers and are measured in kilograms. We therefore aggregated them in a single predictor by summing their nitrogen content. Furthermore, we included categorical predictors representing the climate region of crop production activity and whether the crop is perennial or annual (for an overview of the implementation of these predictors see S4). This resulted in 16 predictors representing six overarching predictor themes (Table 1 ). For a detailed overview of predictors and farming information, see Table S5. Table 1 – Overview of predictors used to quantify environmental footprints of crop production. Predictor theme Predictor Unit Description Energy Electricity kWh Electricity used to run greenhouses and supporting activities in other productions Fuel MJ Diesel fuel burned in agricultural machinery Heat MJ Heat used in greenhouses Fertilisers & soil amendment Potassium fertilisers Kg active ingredient (K 2 O) Potassium rich fertilisers such as potassium chloride, potassium sulfite, vinasse Nitrogen fertilisers kg active ingredient (N) Nitrogen rich fertilisers such as urea, ammonia, ammonium nitrate, calcium nitrate Phosphorus fertilisers kg active ingredient (P 2 O 5 ) Phosphorus rich fertilisers such as superphosphates, P 2 O 5 , filter cake Manure kg manure Manure used for fertilisation Micronutrients kg micronutrient Materials that supply micronutrients (e.g., Fe, Mo, Ni, Zn, Cu) to the crop such as calcium borates, zinc, manganese sulfate Soil improvement kg soil conditioner materials used to correct soil pH and condition soil, such as lime, gypsum, dolomite Stimulants kg stimulant Materials that enhance growth, such as biostimulants, growth regulators Plant protection Plant protection material kg active ingredient Pesticide, herbicide, fungicides and other materials used for protecting the crops and plants Water use Irrigation water m 3 Water used for irrigation Tap water m 3 Water used for dilution of pesticide Yield Yield kg/ha Yield of each crop calculated based on land occupation Categorical information Climate region - Whether crop is located in a tropical, arid, temperate, continental or polar climate (according to S4) Crop type - Whether crop is perennial or annual To check for multicollinearity, we calculated Variance Inflation Factors (VIFs) for the predictor variables and observed that all VIF scores were smaller than 5, which is indicative of relatively low multicollinearity among the predictors (Ferre, 2009 ). We then scaled all predictor values to the range from 0 to 1 because some machine learning algorithms, particularly ANNs, are sensitive to data scaling (Pedregosa et al., 2011 ). We rescaled the predictor values using the MinMaxScaler function of the Scikit-learn package in Python (Pedregosa et al., 2011 ). 2.4 Machine learning modelling techniques We used five machine learning modelling techniques to model footprints: Random Forest (RF), K-Nearest Neighbours (KNN), Artificial Neural Network (ANN), Generalised Boosting Method (GBM) and a regularised Linear Model (LM). All five machine learning modelling techniques have been shown to achieve high predictive performance in estimating footprints in agricultural, construction and chemical LCA studies (Meng et al., 2019 ; Romeiko et al., 2020 ; Thilakarathna et al., 2020 ; Xikai et al., 2019 ). RF is a regression and classification algorithm comprised of an ensemble of decision trees. Each tree is constructed based on a random subset of data and predictors and operates by iteratively splitting the data at each node, based on the predictor that minimises variance, thus achieving more uniform subsets of data on either branch of the node (Luan et al., 2020 ). The output of the RF is an average of outputs of individual trees, making it resistant to scaling and data availability issues (Breiman, 2001 ). KNN regression is a predictive algorithm that estimates a data point's target value by averaging the target values of a certain number (k) of its nearest neighbours, with k being specified by the user. KNN is a non-parametric algorithm, with no assumptions on the underlying data. It makes predictions solely based on the distances between data points, also making the algorithm sensitive to the number of neighbours and distance metric chosen (Peterson, 2009 ). ANNs emulate the learning process of the human brain. They consist of input, output and hidden layers, within which, neurons or nodes exist with activation functions that project data to nonlinear spaces. ANNs are therefore capable of identifying non-linear relationships (Das et al., 2016 ; Song et al., 2017 ). GBMs are ensemble learning techniques that combine multiple weak learners, such as decision trees, to create a strong predictive model through an additive training strategy, where each new learner is trained to correct the errors of the previous ones (Qui et al., 2022). LMs are algorithms that model the relationship between dependent and independent variables using a linear equation (Pedregosa et al., 2011 ). The LM implementation used here includes a regularization term to prevent overfitting and is commonly known as Ridge Regression (McDonald, 2009 ). 2.5 Model fitting and evaluation We used a five-fold block cross-validation to evaluate the models’ performance to predict footprints of crops that are not part of the training data and to avoid overfitting. In each run, 80% of the data (i.e. four folds) is used for training, and the remaining 20% of the data (i.e. one fold) is used for validation of the model, resulting in five runs. We split the data based on crop species, ensuring the same crops would not be represented in both training and validation sets to prevent data leakage. To quantify the performance of each model, we calculated for each validation set the explained variance (R 2 ) and the Root Mean Squared Error (RMSE) (Chai & Draxler, 2014 ). We then averaged the R 2 and RMSE over the five folds to obtain the overall performance for each model. We applied each model 16 times (once using 1 predictor, once using two, and so on until 16 predictors). We optimised each model by finding the predictor subset that minimises the cross-validated RMSE. We then tuned the hyperparameters of each machine learning modelling technique (e.g. number of neighbours in KNN, number of trees in RF, learning rate in GBM; Table S6) using Bayesian optimization, with the goal of maximizing predictive performance. We performed the optimizations using the Optuna package in Python (Akiba et al., 2019 ). The models with optimal hyperparameters are specified in section S7 of the supplementary information. We built all modelling pipelines using the Scikit-learn Python package (Pedregosa et al., 2011 ). To determine the most parsimonious model per footprint, we used the knee (also known as “elbow”) method to identify the number of predictors beyond which the inclusion of more predictors yields a negligible increase in predictive performance (i.e. lowering RMSE) (Satopaa et al., 2011 ). The knee point is defined as the point of maximum curvature in any function, with curvature being a mathematical measure of how much a function differs from a straight line (Satopaa et al., 2011 ). As a result, maximum curvature captures where performance improvements start to level off (Satopaa et al., 2011 ). We used the Python package kneed (Arvai, 2020 ) to find the knee point on the line connecting the best performing machine learning modelling technique for each predictor subset size. To quantify the importance of each predictor to the environmental footprint predictions, we used Shapley values. Shapley values are based on game theory and explain machine learning model outputs by calculating the average marginal contribution of each predictor across all possible subsets of predictors (Sundararajan & Najmi, 2020 ). We used the SHAP python package to calculate the Shapley values (Lundberg & Lee, 2017). 3. Results and Discussion 3.1 Predictive power and parsimony Based on all predictors considered (Table 1 ), we found block cross-validated R 2 values for the water and biodiversity footprint models of 0.90 (Fig. 2 ) using GBM (best performing model). For the climate footprint model, we found a block cross-validated R 2 of 0.64. This performance is lower than typically reported for predictive models of climate footprints, as R² scores between 0.64 and 0.92 are reported in previous work focused on single crop species ( Khanali et al., 2017 ; Khoshnevisan et al., 2013 ; Lee et al., 2020 ; Romeiko et al., 2020 ). The comparatively lower R 2 of our model may reflect that we i) considered only on-farm information as predictors, excluding information on infrastructure, crop residue treatment, packaging, labour, machinery and seed production which were included in previous work and are known to be additional drivers of the climate footprint in some crop-country combinations (Khanali et al., 2017; Lam et al., 2021 ; Nabavi-Pelesaraei et al., 2018 ; Verge et al., 2007 ; Liu et al., 2016 ; Yan et al., 2014 ; Aguilera et al., 2015 ), ii) included a large variety of crops and countries, and iii) performed block cross-validation to test predictions for crop species that are not included in the training set. On the one hand, our approach to train the models on a heterogeneous database of crops enhances its generalizability to other crops, locations and practices. On the other hand, our models will yield lower predictive performance for the footprints of a specific crop species compared to models specifically trained on that particular crop or closely related data, especially considering that multiple crops only have one entry in our dataset. We found that when using a small subset of predictors, a prediction performance could be attained similar to that of using a larger set of predictors (Fig. 2 ). We observed that using subsets of more than five predictors (identified via the knee method) hardly increases or even decreases performance as measured by both RMSE (in natural log-space) and R 2 (e.g., for climate footprints), underscoring the importance of parsimony in model selection. This decrease in performance is due to predictors adding more noise than information to the models’ learning, especially predictors which are important only to a few specific crops. KNN is most sensitive to this, leading to a significant decrease in performance past the knee point. In contrast, previous work has relied on all available input data for modelling footprints (Romeiko et al., 2023 ; Kahanali et al., 2017; Khoshnevisan et al., 2013 ; Lam et al., 2021 ; Pishgar-Komleh et al., 2017 ; Pishgar-Komleh et al., 2020 ; Nabavi-Pelesaraei et al., 2018 ; Lee et al., 2020 ). Overall, GBM, ANN and RF outperform the other two machine learning modelling techniques (KNN and LM) over the majority of predictor subsets for the estimation of all three footprints. We found RF to be the most parsimonious model with four predictors for the biodiversity footprint (R 2 = 0.88, RMSE = 0.37). We found GBM using five predictors as the most parsimonious model for water footprint (R 2 = 0.89, RMSE = 0.83) and ANN for climate footprint (R 2 = 0.58, RMSE = 0.71) with four predictors. These RMSE values correspond to an uncertainty of +/- a factor 2.1, 5.1, and 4.0 (95% confidence interval) for biodiversity, water, and climate footprints, respectively. RMSE and R 2 reveal consistent patterns (i.e. RMSE is lowest when R 2 is highest, both indicating the best model; Fig. 2 ). 3.2 Relative contribution of different predictors to model prediction Yield is the most important predictor of the biodiversity footprint (Fig. 3 ), as land use is the primary driver of biodiversity loss and increasing yield directly reduces the land use area required for crop production (Hald-Mortensen, 2023 ; Jaureguiberry et al., 2022 ). It should be noted that whilst our predictors (on-farm information) likely represent heterogenous farming practices (e.g. monoculture vs intercropping, or organic vs conventional agriculture) which influence the biodiversity footprint, our models have not been trained with farm management as an input parameter. In addition, land use intensity effects on local biodiversity are not captured in the ReCiPe 2016 methodology (Huijbregts et al., 2017 ). Phosphorus fertiliser use, irrigation water, and heat are the next most important predictors for the biodiversity footprint. Phosphorus impacts biodiversity via eutrophication of freshwater (Huijbregts et al., 2017 ; Sud, 2020). For the water footprint, we found that irrigation, nitrogen fertiliser, electricity use and plant protection inputs were the most important predictors (Fig. 3 ). The latter may reflect the water use in the upstream production and application of fertilisers and pesticides (Sud, 2020). While we find that the biodiversity and water footprints are largely driven by a single predictor (yield and irrigation, respectively), we find that contributions to climate footprints are more evenly spread across several predictors (yield with 42.1%, nitrogen fertiliser input with 25.9%, climate region with 21% and electricity use with 10.8%). The importance of yield in climate footprint estimation is in line with other studies finding associations between decreasing footprints and increased yield (Lam et al., 2021 ; Pishgar-Komleh et al., 2017 ). Nitrogen fertiliser is also known to be a driver of climate footprints through N 2 O and NOx emissions from the soil subsequent to its application (Rosegrant et al., 2002). Lastly, land use change based emissions vary across countries and climate zones even for the same activity (Reinhard et al., 2017 ). Therefore climate region information helps the model distinguish between these types of emissions. While using country-level data would be more precise, it would greatly increase dimensionality, reducing the robustness of the model. 3.3 Implications, limitations and future research directions The block cross-validation revealed that, based on groups of on-farm information, machine learning models can predict the biodiversity and water footprints of the production of ‘new’ crops (i.e., crop species on which the models were not trained) with relatively high accuracy (R 2 > 0.89). This could be partially explained by biodiversity and water footprints being strongly related to single pressures (namely land and water use, respectively) (Reinhard et al., 2017 ; Fróna et al., 2019). Climate footprints are more challenging to predict, partially due to the many different sources and drivers of GHG emissions (Poore & Nemecek, 2018 ; Khoshnevisan et al., 2013 ; Lam et al., 2021 ). We recommend the use of GBM, RF and ANN (in order of performance) over KNN and LM when predicting footprints of crop production based on on-farm information, and we recommend caution when applying predictive models to estimate footprints that are related to many different processes such as climate footprints. Our analysis showed that a small set of on-farm predictors can achieve nearly the same prediction accuracy as using many predictors. To our knowledge, no previous work has examined the trade-off between the number of predictors used and prediction accuracy in the context of crop production footprints. We recommend prioritizing data collection or data gap-filling of agricultural yield, (nitrogen and phosphorus) fertiliser use, pesticide use, irrigation, and electricity use when aiming to assess biodiversity, water, or climate footprints. While most of the existing literature examines the contribution of farming information to footprints in a mechanistic manner, the only similar machine learning-based study we found, also identified fertiliser as a key driver of climate footprints, despite using different predictors (Romeiko et al., 2020 ). Our methodology can be used to provide impact estimates for screening purposes when detailed data are unavailable or a quick identification of high-impact crop production activities is desired. Our methodology can also be used to reveal major drivers of other footprints and other products to guide LCI data collection and data gap-filling efforts. Lastly, we found that categorical predictors such as climatic region of the farm location (i.e. tropical, arid, temperate, continental and polar) and crop type (i.e. perennial or annual) provide additional, valuable and relatively easy information to the models, although they lack the detail to provide good performance on their own (see S12 for performance metrics of modelling techniques using only categorical predictors as well as categorical predictors and most important numerical predictor; see S13 for an overview of categorical predictors). We encountered several challenges that exposed opportunities for future research. First, the databases we used are mostly comprised of national average crop production LCIs, which do not capture the full variety of farming management practices (e.g. various land use intensities). This means that we have been unable to test the accuracy of our predictions for the same crops in the same region but with different management characteristics. Compiling a higher-resolution database covering a wider range of farming practices and temporal variation would improve the validity and representativeness of model predictions. Second, the aggregation of farming information into predictors, such as the aggregation of specific fertilisers (e.g., ammonium nitrate) into fertiliser groups (e.g., nitrogen fertiliser) based on their active ingredients (e.g., kg nitrogen), or aggregating electricity inputs into a single predictor regardless of generation method, masks potential differences in impacts within these groups. Although we distinguished manure from nitrogen, phosphorus, and potassium fertilisers, we did not distinguish other organic fertilisers from inorganic fertilisers, which may have different impacts depending on their production processes and associated on-field emissions (García Castellanos et al., 2023 ). Thus alternative predictor grouping strategies could affect model outcomes and insights. Third, here we focused on biodiversity, water and climate footprints as they are important to agricultural systems. Future research could evaluate the applicability of ML models to other impact indicators. Finally, we omitted several information including packaging material, on-farm preprocessing, seed and seedling input, and land use change, due to the boundaries of our study (i.e. considering only on-farm or foreground information). While this limits the explicit inclusion of upstream and indirect processes, expanding the system boundary in future work could allow the evaluation of identified relationships in a more comprehensive setting. Given the scarcity of LCI data for many crop-country-practice combinations, and the difficulty in collecting them, here we evaluated the potential for predicting impacts using machine learning techniques and a more limited set of data than normally captured in LCIs (i.e. on-farm information). Overall, this study shows that machine learning is a promising tool for predicting some environmental footprints of crop production based on on-farm information, although prediction accuracy varies by footprint. We found that climate region of the farm, yield, irrigation, fertiliser and pesticide inputs, and the use of electricity and heat, are sufficient to achieve predictive accuracy close to models using more inputs. The parsimonious models developed in this work provide a systematic approach to identifying priorities for data gap-filling, thereby facilitating a broadening of the range of crops for which environmental footprints can be estimated. Declarations Competing interests: The authors declare no competing interests relevant to the content of this article. Author Contribution All authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by F.R.. All authors discussed the results. F.R. wrote the original draft with inputs from all authors. All authors worked on the revisions to the paper and approved this manuscript. Acknowledgement We would like to thank Katy Armstrong and Florence Bohnes, internal reviewers from Unilever’s Safety, Environmental and Regulatory Sciences group, for their insightful comments. Moreover, we would like to thank Selwyn Hoeks from Radboud University and Leonardo Contreas from Unilever for their help with programming issues and code review. Data Availability All scripts that support the findings of this study are available in Github at: https://github.com/FRaymand/parsimonious-crop-footprint-estimator References Aguilera, E., Guzmán, G., & Alonso, A. (2015). Greenhouse gas emissions from conventional and organic cropping systems in Spain. I. Herbaceous crops. Agronomy for Sustainable Development , 35(2), 713–724. https://doi.org/10.1007/s13593-014-0267-9 Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. (2019). Optuna: A Next-generation Hyperparameter Optimization Framework. Proceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining , 2623–2631. https://doi.org/10.1145/3292500.3330701 Algunaibet, I. M., & Guillén-Gosálbez, G. (2019). Life cycle burden-shifting in energy systems designed to minimize greenhouse gas emissions: Novel analytical method and application to the United States. Journal of Cleaner Production , 229, 886–901. Arvai, K. (2020). kneed (Version 0.7.0) [Computer software]. https://doi.org/10.5281/zenodo.6496267 Asselin-Balençon, A., Broekema, R., Teulon, H., Gastaldi, G., Houssier, J., Moutia, A., … & Colomb, V. (2020). AGRIBALYSE v3. 0: the French agricultural and food LCI database. Methodology for the food products . Ed. ADEME . Bellon-Maurel, V., Short, M. D., Roux, P., Schulz, M., & Peters, G. M. (2014). Streamlining life cycle inventory data generation in agriculture using traceability data and information and communication technologies–part I: concepts and technical basis. Journal of Cleaner Production , 69, 60–66. https://doi.org/10.1016/j.jclepro.2014.01.079 Borghino, N., Corson, M., Nitschelm, L., Wilfart, A., Fleuet, J., Moraine, M., … Godinot, O. (2021). Contribution of LCA to decision making: A scenario analysis in territorial agricultural production systems. Journal of Environmental Management , 287, 112288. Breiman, L. (2001). Random Forests. Machine Learning 45, 5–32. https://doi.org/10.1023/A:1010933404324 Canals, L. M. I., Azapagic, A., Doka, G., Jefferies, D., King, H., Mutel, C., Nemecek, T., Roches, A., Sim, S., Stichnothe, H., Thoma, G., & Williams, A. (2011). Approaches for addressing life cycle assessment data gaps for bio-based products. Journal of Industrial Ecology , 15(5), 707–725. https://doi.org/10.1111/j.1530-9290.2011.00369.x Chai, T., & Draxler, R. R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? -Arguments against avoiding RMSE in the literature. Geoscientific Model Development , 7(3), 1247–1250. https://doi.org/10.5194/gmd-7-1247-2014 Das, P. J., Preuss, C., & Mazumder, B. (2016). Artificial Neural Network as Helping Tool for Drug Formulation and Drug Administration Strategies. Artificial Neural Network for Drug Design, Delivery and Disposition , 263–276. https://doi.org/10.1016/B978-0-12-801559-9.00013-2 Donati, F., Dente, S. M., Li, C., Vilaysouk, X., Froemelt, A., Nishant, R., … Hashimoto, S. (2022). The future of artificial intelligence in the context of industrial ecology. Journal of Industrial Ecology , 26 (4), 1175–1181. Dudley, N., & Alexander, S. (2017). Agriculture and biodiversity: a review. Biodiversity , 18(2–3), 45–49. https://doi.org/10.1080/14888386.2017.1351892 FAO. 2022. The State of the World’s Land and Water Resources for Food and Agriculture – Systems at breaking point. Main report. Rome. https://doi.org/10.4060/cb9910en Fróna, D., Szenderák, J., & Harangi-Rákos, M. (2021). The challenge of feeding the world sustainably. Chall. Feed. World Sustain , 11, 1–36. https://doi.org/10.17226/26007 Ferre, J. (2009). Variance inflation factor: An overview. Science Direct , 3, 1–24. García Castellanos, B., García García, B., & García García, J. (2023). Economic and environmental effects of replacing inorganic fertilizers with organic fertilizers in three rainfed crops in a semi-arid area. Sustainability , 15 (24), 16897. https://doi.org/10.3390/su152416897 Hald-Mortensen, C. (2023). The main drivers of biodiversity loss: a brief overview. Journal of Ecology and Natural Resources , 7(3), 000346. https://doi.org/10.23880/jenr-16000346 Hoogeveen, J., Faurès, J. M., Peiser, L., Burke, J., & Van De Giesen, N. (2015). GlobWat - A global water balance model to assess water use in irrigated agriculture. Hydrology and Earth System Sciences , 19(9), 3829–3844. https://doi.org/10.5194/hess-19-3829-2015 Huijbregts, M. A., Steinmann, Z. J., Elshout, P. M., Stam, G., Verones, F., Vieira, M., Zijp, M., Hollander, A. & Van Zelm, R. (2017). ReCiPe2016: a harmonised life cycle impact assessment method at midpoint and endpoint level. The International Journal of Life Cycle Assessment , 22, 138–147. Jacob-Lopes, E., Zepka, L. Q., & Deprá, M. C. (2021). Sustainability metrics and indicators of environmental impact: industrial and agricultural life cycle assessment . Elsevier. https://doi.org/10.1016/C2020-0 -00268-9Khanali, M., Mobli, H., & Hosseinzadeh-Bandbafha, H. (2017). Modeling of yield and environmental impact categories in tea processing units based on artificial neural networks. Environmental Science and Pollution Research , 24(34), 26324–26340. https://doi.org/10.1007/s11356-017-0234-5 Jaureguiberry, P., Titeux, N., Wiemers, M., Bowler, D. E., Coscieme, L., Golden, A. S., … Purvis, A. (2022). The direct drivers of recent global anthropogenic biodiversity loss. Science advances , 8(45), eabm9982. https://doi.org/10.1126/sciadv.abm9982 Khoshnevisan, B., Rafiee, S., Omid, M., Mousazadeh, H., & Sefeedpari, P. (2013). Prognostication of environmental indices in potato production using artificial neural networks. Journal of Cleaner Production , 52, 402–409. https://doi.org/10.1016/j.jclepro.2013.03.028 Laborde, D., Mamun, A., Martin, W., Piñeiro, V., & Vos, R. (2021). Agricultural subsidies and global greenhouse gas emissions. Nature Communications , 12(1), 2601. https://doi.org/10.1038/s41467-021-22703-1 Lam, W. Y., Sim, S., Kulak, M., van Zelm, R., Schipper, A. M., & Huijbregts, M. A. J. (2021). Drivers of variability in greenhouse gas footprints of crop production. Journal of Cleaner Production , 315, 128121. https://doi.org/10.1016/j.jclepro.2021.128121 Lee, E. K., Zhang, W. J., Zhang, X., Adler, P. R., Lin, S., Feingold, B. J., … Romeiko, X. X. (2020). Projecting life-cycle environmental impacts of corn production in the US Midwest under future climate scenarios using a machine learning approach. Science of The Total Environment , 714, 136697. https://doi.org/10.1016/j.scitotenv.2020.136697 Liu, C., Cutforth, H., Chai, Q., & Gan, Y. (2016). Farming tactics to reduce the carbon footprint of crop cultivation in semiarid areas. A review. Agronomy for Sustainable Development , 36(4), 69. https://doi.org/10.1007/s13593-016-0404-8 Luan, J., Zhang, C., Xu, B., Xue, Y., & Ren, Y. (2020). The predictive performances of random forest models with limited sample size and different species traits. Fisheries Research , 227, 105534. https://doi.org/10.1016/j.fishres.2020.105534Lundberg , S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. Advances in neural information processing systems , 30. McDonald, G. C. (2009). Ridge regression. Wiley Interdisciplinary Reviews: Computational Statistics , 1(1), 93–100. Meng, F., LaFleur, C., Wijesinghe, A., & Colvin, J. (2019). Data-driven approach to fill in data gaps for life cycle inventory of dual fuel technology. Fuel , 246, 187–195. https://doi.org/10.1016/j.fuel.2019.02.124 Mutel, C. (2017). Brightway: An open source framework for Life Cycle Assessment. Journal of Open Source Software , 2(12), 236. https://doi.org/10.21105/joss.00236 Nabavi-Pelesaraei, A., Rafiee, S., Mohtasebi, S. S., Hosseinzadeh-Bandbafha, H., & Chau, K. W. (2018). Integration of artificial intelligence methods and life cycle assessment to predict energy output and environmental impacts of paddy production. Science of the Total Environment , 631, 1279–1294. https://doi.org/10.1016/j.scitotenv.2018.03.088 Nemecek, T., Bengoa, X., Lansche, J., Roesch, A., Faist-Emmenegger, M., Rossi, V., … Riedener, E. (2019). World food LCA database. Methodological Guidelines for the Life Cycle Inventory of Agricultural Products , 3. Norris, K. (2008). Agriculture and biodiversity conservation: opportunity knocks. Conservation Letters , 1(1), 2–11. https://doi.org/10.1111/j.1755-263x.2008.00007.x Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., … Duchesnay, É. (2011). Scikit-learn: Machine learning in Python. Journal of Machine Learning Research , 12, 2825–2830. https://doi.org/10.48550/arXiv.1201.0490 Peterson, L. (2009). K-nearest neighbor. Scholarpedia , 4(2), 1883. https://doi.org/10.4249/SCHOLARPEDIA.1883 Pishgar-Komleh, S. H., Akram, A., Keyhani, A., Raei, M., Elshout, P. M. F., Huijbregts, M. A. J., & Van Zelm, R. (2017). Variability in the carbon footprint of open-field tomato production in Iran-A case study of Alborz and East-Azerbaijan provinces. Journal of Cleaner Production , 142, 1510–1517. http://dx.doi.org/10.1016/j.jclepro.2016.11.154 Pishgar-Komleh, S. H., Akram, A., Keyhani, A., Sefeedpari, P., Shine, P., & Brandao, M. (2020). Integration of life cycle assessment, artificial neural networks, and metaheuristic optimization algorithms for optimization of tomato-based cropping systems in Iran. International Journal of Life Cycle Assessment , 25(3), 620–632. https://doi.org/10.1007/s11367-019-01707-6 Poore, J., & Nemecek, T. (2018). Reducing food’s environmental impacts through producers and consumers. Science , 360(6392), 987–992. https://doi.org/10.1126/science.aaq0216 Qiu, R., Liu, C., Cui, N., Gao, Y., Li, L., Wu, Z., Jiang, S., & Hu, M. (2022). Generalized Extreme Gradient Boosting model for predicting daily global solar radiation for locations without historical data. Energy Conversion and Management , 258. https://doi.org/10.1016/j.enconman.2022.115488 Reinhard J., Moreno-Ruiz E., & Gmünder S. (2017) Consideration of land use change in ecoinvent version 3.3: Method, Implementation and Illustration. White paper, ecoinvent Association, Zürich, Switzerland Romeiko, X. X., Guo, Z., Pang, Y., Lee, E. K., & Zhang, X. (2020). Comparing machine learning approaches for predicting spatially explicit life cycle global warming and eutrophication impacts from corn production. Sustainability , 12(4), 1481. https://doi.org/10.3390/su12041481 Romeiko, X. X., Zhang, X., Pang, Y., Gao, F., Xu, M., Lin, S., & Babbitt, C. (2023). A review of machine learning applications in life cycle assessment studies. Science of the Total Environment , 168969. Rosegrant, M. W., Cai, X., & Cline, S. A. ( https://hdl.handle.net/10568/1567952002 ). World water and food to 2025: dealing with scarcity. Intl Food Policy Res Inst . https://hdl.handle.net/10568/156795 Satopaa, V., Albrecht, J., Irwin, D., & Raghavan, B. (2011, June). Finding a" kneedle" in a haystack: Detecting knee points in system behavior. In 2011 31st international conference on distributed computing systems workshops (pp. 166–171). IEEE . https://doi.org/10.1109/ICDCSW.2011.20 Song, R., Keller, A. A., & Suh, S. (2017). Rapid Life-Cycle Impact Screening Using Artificial Neural Networks. Environmental Science and Technology , 51(18), 10777–10785. https://doi.org/10.1021/acs.est.7b02862 Sud, M. (2020). Managing the biodiversity impacts of fertiliser and pesticide use: Overview and insights from trends and policies across selected OECD countries. OECD Environment Working Papers , 155. https://doi.org/10.1787/63942249-enCabot , M. I., Lado, J., & Sanjuan, N. (2025). Sundararajan, M., & Najmi, A. (2020, November). The many Shapley values for model explanation. In International conference on machine learning (pp. 9269–9278). PMLR . https://proceedings.mlr.press/v119/sundararajan20b.html Thilakarathna, P. S. M., Seo, S., Baduge, K. S. K., Lee, H., Mendis, P., & Foliente, G. (2020). Embodied carbon analysis and benchmarking emissions of high and ultra-high strength concrete using machine learning algorithms. Journal of Cleaner Production , 262, 121281. https://doi.org/10.1016/j.jclepro.2020.121281 UN (2022). United Nations Department of Economic and Social Affairs, Population Division (2022). World Population Prospects 2022: Summary of Results. UN DESA/POP/2022/TR/NO. 3. Verge, X. P. C., De Kimpe, C., & Desjardins, R. L. (2007). Agricultural production, greenhouse gas emissions and mitigation potential. Agricultural and forest meteorology , 142(2–4), 255–269. https://doi.org/10.1016/j.agrformet.2006.06.011 Vermeulen, S. J., Campbell, B. M., & Ingram, J. S. (2012). Climate change and food systems. Annual Review of Environment and Resources , 37(1), 195–222. https://doi.org/10.1146/annurev-environ-020411-130608 Wernet, G., Bauer, C., Steubing, B., Reinhard, J., Moreno-Ruiz, E., & Weidema, B. (2016). The ecoinvent database version 3 (part I): overview and methodology. The International Journal of Life Cycle Assessment , 21, 1218–1230. https://doi.org/10.1007/s11367-016-1087-8 Willett, W., Rockström, J., Loken, B., Springmann, M., Lang, T., Vermeulen, S., … Murray, C. J. L. (2019). Food in the Anthropocene: the EAT–Lancet Commission on healthy diets from sustainable food systems. The Lancet , 393(10170), 447–492. https://doi.org/10.1016/S0140-6736(18)31788-4 Willet, J., Wetser, K., Vreeburg, J., & Rijnaarts, H. H. (2019). Review of methods to assess sustainability of industrial water use. Water Resources and Industry , 21, 100110. Xikai, M., Lixiong, W., Jiwei, L., Xiaoli, Q., & Tongyao, W. (2019). Comparison of regression models for estimation of carbon emissions during building’s lifecycle using designing factors: a case study of residential buildings in Tianjin, China. Energy and Buildings , 204, 109519. https://doi.org/10.1016/j.enbuild.2019.109519 Yan, M., Cheng, K., Luo, T., & Pan, G. (2014). Carbon footprint of crop production and the significance for greenhouse gas reduction in the agriculture sector of China. In Assessment of Carbon Footprint in Different Industrial Sectors, Volume 1 (pp. 247–264). Singapore: Springer Singapore. https://doi.org/10.1007/978-981-4560-41-2_10 Zargar, S., Yao, Y., & Tu, Q. (2022). A review of inventory modeling methods for missing data in life cycle assessment. Journal of Industrial Ecology , 26(5), 1676–1689. https://doi.org/10.1111/jiec.13305 Zhao, B., Jiang, J., Xu, M., & Tu, Q. (2025). A data-centric investigation on the challenges of machine learning methods for bridging life cycle inventory data gaps. Journal of Industrial Ecology , 29 (3), 955–966. Additional Declarations No competing interests reported. Supplementary Files FarhangRaymandSupplementaryInformationParsimoniousmachinelearningmodelstoestimateenvironmentalfootprintsofcropproduction.docx Cite Share Download PDF Status: Under Review Version 1 posted Reviews received at journal 18 May, 2026 Reviewers agreed at journal 21 Apr, 2026 Reviewers agreed at journal 08 Apr, 2026 Reviewers invited by journal 06 Apr, 2026 Editor assigned by journal 27 Mar, 2026 Submission checks completed at journal 25 Mar, 2026 First submitted to journal 24 Mar, 2026 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-9214313","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":620986552,"identity":"d07d1e56-129d-4234-991c-7df7229b030b","order_by":0,"name":"Farhang Raymand","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA7UlEQVRIiWNgGAWjYBAC9gYGNmQ+sxyYeoBHC88BNC3GYCqBFC2JDQS1sJ999uDnjloG8/Yzpht+7rBO3z4jge0BXi086eaGvWeOM8icyTG72XsmPXfOjQR2A3xa7BnS2CR4244xSDDkmN3gbTucO0MigU0Cry38z9gk/4K08L8xu/m37XC6BEEtEmls0rxtNQwSEjlmt4G2JBCh5RmbtGzbAR4JiWdlt2Xb0g1n8DxsI+CwNDbJt211chL8ydtuvm2zlpdgTz4m8QGPFig4zMPAwGEA5TA2ENbAwFAHxOwPiFE5CkbBKBgFIxAAAG2CRpvtu5lYAAAAAElFTkSuQmCC","orcid":"","institution":"Radboud University","correspondingAuthor":true,"prefix":"","firstName":"Farhang","middleName":"","lastName":"Raymand","suffix":""},{"id":620986553,"identity":"6967116b-b6d1-4efb-b825-2f349a408bad","order_by":1,"name":"Koen J.J. Kuipers","email":"","orcid":"","institution":"Radboud University","correspondingAuthor":false,"prefix":"","firstName":"Koen","middleName":"J.J.","lastName":"Kuipers","suffix":""},{"id":620986554,"identity":"57e681d2-290e-4462-b104-d73de93c6852","order_by":2,"name":"Sarah Sim","email":"","orcid":"","institution":"Unilever","correspondingAuthor":false,"prefix":"","firstName":"Sarah","middleName":"","lastName":"Sim","suffix":""},{"id":620986555,"identity":"f25a0642-0efb-4d09-b2ed-b9874395bc31","order_by":3,"name":"P. James Joyce","email":"","orcid":"","institution":"Unilever","correspondingAuthor":false,"prefix":"","firstName":"P.","middleName":"James","lastName":"Joyce","suffix":""},{"id":620986556,"identity":"833cb184-4909-4b71-9330-fc273f552290","order_by":4,"name":"Aafke M. Schipper","email":"","orcid":"","institution":"Radboud University","correspondingAuthor":false,"prefix":"","firstName":"Aafke","middleName":"M.","lastName":"Schipper","suffix":""},{"id":620986557,"identity":"138f2601-bb82-4f78-b0fe-f47af9378fe6","order_by":5,"name":"Mark A.J. Huijbregts","email":"","orcid":"","institution":"Radboud University","correspondingAuthor":false,"prefix":"","firstName":"Mark","middleName":"A.J.","lastName":"Huijbregts","suffix":""}],"badges":[],"createdAt":"2026-03-24 16:12:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-9214313/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-9214313/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":106792274,"identity":"b74dadd8-0894-4ede-8515-eb6e5c10beb5","added_by":"auto","created_at":"2026-04-13 13:31:47","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":158018,"visible":true,"origin":"","legend":"\u003cp\u003eGeneral approach used to develop models to predict crop production footprints based on farming information, as well as relative importance of information\u003c/p\u003e","description":"","filename":"floatimage1.png","url":"https://assets-eu.researchsquare.com/files/rs-9214313/v1/fe8e5ef8b39931efed3be608.png"},{"id":106792275,"identity":"55a5bbb3-7cfa-4b40-851f-2ce61c0f310d","added_by":"auto","created_at":"2026-04-13 13:31:47","extension":"jpeg","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":563446,"visible":true,"origin":"","legend":"\u003cp\u003eModel predictive performance in relation to the number of predictors used for the biodiversity, water and climate footprints of crop production quantified based on the Hierarchist perspective. Predictive performance is expressed as R\u003csup\u003e2\u003c/sup\u003e (a) and RMSE (b; in natural log-space) obtained as mean values based on a 5-fold block cross-validation per machine learning modelling technique (KNN in orange, GBM in pink, LM in light green, RF in dark green, and ANN in blue). The vertical lines represent the knee points corresponding with the most parsimonious model per footprint. Results for the ReCiPe2016 Egalitarian and Individualist perspectives show similar patterns (see Section S8 and S9).\u003c/p\u003e","description":"","filename":"floatimage2.jpeg","url":"https://assets-eu.researchsquare.com/files/rs-9214313/v1/5d625a9bd50f4a9644156fa1.jpeg"},{"id":106792273,"identity":"aa2ab7d5-0ae2-4d33-b0d9-b982829cd4c6","added_by":"auto","created_at":"2026-04-13 13:31:47","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":50027,"visible":true,"origin":"","legend":"\u003cp\u003eRelative importance of predictors for the prediction of biodiversity footprint, water footprint and climate footprint based on the most parsimonious model (biodiversity footprint: RF with 4 predictors; water footprint: GBM with 5 predictors; climate footprint: ANN with 4 predictors) and the Hierarchist perspective. Results for the Individualistic and Egalitarian perspectives were similar (see section S10 and S11).\u003c/p\u003e","description":"","filename":"floatimage3.png","url":"https://assets-eu.researchsquare.com/files/rs-9214313/v1/0bac8a3f60bd7ffb07b5708e.png"},{"id":106792294,"identity":"1ba46399-6969-440e-bb7f-4d724adff651","added_by":"auto","created_at":"2026-04-13 13:31:59","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1431667,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-9214313/v1/4b4ea7d1-5106-474a-ba07-07e2c8c1f237.pdf"},{"id":106792272,"identity":"8cc1685b-5e2c-4fc3-8235-03b34ac68aee","added_by":"auto","created_at":"2026-04-13 13:31:47","extension":"docx","order_by":0,"title":"","display":"","copyAsset":false,"role":"supplement","size":855875,"visible":true,"origin":"","legend":"","description":"","filename":"FarhangRaymandSupplementaryInformationParsimoniousmachinelearningmodelstoestimateenvironmentalfootprintsofcropproduction.docx","url":"https://assets-eu.researchsquare.com/files/rs-9214313/v1/a10895aa2c2e5b430c91cdd6.docx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Parsimonious machine learning models to estimate environmental footprints of crop production","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe agricultural sector covers about 37% of land on earth (FAO, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). It is a major driver of biodiversity loss (Norris, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2008\u003c/span\u003e; Willet et al., \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2019\u003c/span\u003e), and it is responsible for about 30% of all global greenhouse gas (GHG) emissions (Laborde et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Poore \u0026amp; Nemecek, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Vermeulen et al., \u003cspan citationid=\"CR51\" class=\"CitationRef\"\u003e2012\u003c/span\u003e). Moreover, agriculture is responsible for more than 70% of global water consumption, mostly due to crop irrigation (FAO, \u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Hoogeveen et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2015\u003c/span\u003e). With a global population expected to exceed 9.7\u0026nbsp;billion people by 2050 (UN, 2022), the adverse environmental impacts of the agricultural sector are expected to further increase. Environmental impacts of agricultural crop production can be quantified using Life Cycle Assessment (LCA). LCA can be used to identify agricultural practices and inputs that contribute most to environmental impacts, allowing decision makers to identify improvement options while considering potential burden shifts (Borghino et al., \u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Algunaibet \u0026amp; Guill\u0026eacute;n-Gos\u0026aacute;lbez, 2019). However, LCA studies require considerable input data, collection of which is time-consuming (Bellon-Maurel et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2014\u003c/span\u003e; Zhao et al., \u003cspan citationid=\"CR58\" class=\"CitationRef\"\u003e2025\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eWhere data are lacking, LCA practitioners can either collect new information, extrapolate information from similar crops (e.g. obtain average fuel input of wheat products to estimate fuel consumption of soft wheat) or use proxies (e.g. use electricity input of tomatoes cultivated in greenhouses to represent those of peppers cultivated in greenhouse) (Zargar et al., \u003cspan citationid=\"CR57\" class=\"CitationRef\"\u003e2022\u003c/span\u003e; Canals et al., \u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). More recently, advances in data-driven tools have enabled the utilisation of machine learning algorithms to support LCA applications such as data collection and estimating environmental footprints, as these algorithms can enable automation, prediction through discovery and learning from hidden patterns (Romeiko et al., \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2023\u003c/span\u003e, Donati et al., \u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). These studies, however, have predominantly focused on single crops, such as rice, wheat, sugarcane, or corn, and employed a relatively narrow range of modelling techniques, without a systematic comparison of alternative modelling techniques (Khanali et al., 2017; Khoshnevisan et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Lam et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Nabavi-Pelesaraei et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Lee et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Moreover, these studies have not evaluated whether a set of key predictors is sufficient to predict environmental footprints accurately (Khanali et al., 2017; Khoshnevisan et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Lam et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Nabavi-Pelesaraei et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Lee et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). An analysis to find the most parsimonious model (i.e. modelling footprints well while using the fewest on-farm, foreground farming information) capable of reliably estimating impacts is yet missing.\u003c/p\u003e \u003cp\u003eHere, we aim to develop a predictive model capable of estimating life cycle crop production environmental footprints (cradle-to-gate) based on on-farm information only. To this end, we evaluate (i) the extent to which different machine learning modelling techniques can be reliably used to predict life cycle environmental footprints of crop production in terms of biodiversity loss, water use and climate change, based on data related to a limited set of on-farm information concerning climate region and crop type as well as key inputs of material, land area, energy and water; and (ii) identify the most parsimonious models and the farming information that contribute most to their predictive power. We focus on biodiversity, water and climate, as these footprints are particularly important in food production and agricultural systems (Laborde et al., \u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Hoogeveen et al., \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2015\u003c/span\u003e; Willet et al., \u003cspan citationid=\"CR54\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Jacob-Lopes et al., \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Dudley \u0026amp; Alexander, \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). We first obtained biodiversity, water, and climate footprints of 991 crops available in common life cycle inventory (LCI) databases. We then applied five machine learning modelling techniques; Random Forest (RF), K-Nearest Neighbours (KNN), Artificial Neural Network (ANN), Generalised Boosting Method (GBM) and a regularised Linear Model (LM), to predict these three footprints and identify the most parsimonious model per footprint based on an initial set of data for 16 farming information types. In addition, we identify the most important farming information driving model predictions.\u003c/p\u003e"},{"header":"2. Materials and Methods","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 General Approach\u003c/h2\u003e \u003cp\u003eOur general approach for modelling environmental footprints and determining the relative importance of farming information in terms of predictive power consists of three steps (Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). First, we quantified life cycle biodiversity, water and climate footprints (cradle-to-gate) of food crops available in common (agricultural) life cycle inventory (LCI) databases (see section \u003cspan refid=\"Sec4\" class=\"InternalRef\"\u003e2.2\u003c/span\u003e). Second, we grouped on-farm information from the LCIs into a set of 16 predictors based on their purpose and unit of measurement (see section \u003cspan refid=\"Sec5\" class=\"InternalRef\"\u003e2.3\u003c/span\u003e) and applied five machine learning modelling techniques to predict the three environmental footprints based on the selected predictors (see section \u003cspan refid=\"Sec6\" class=\"InternalRef\"\u003e2.4\u003c/span\u003e). We compared the performance of the machine learning modelling techniques via cross-validation. Third, we identified which predictors are most relevant to include, and we assessed the trade-off between utilizing fewer predictors and change in predictive performance (see section \u003cspan refid=\"Sec7\" class=\"InternalRef\"\u003e2.5\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Environmental footprints\u003c/h2\u003e \u003cp\u003eWe used the ecoinvent database (version 3.9.1, cut-off allocation system model), agribalyse database (version 3.11) and World Food LCA Database (WFLDB) (version 3.10) to obtain LCIs for available agricultural crop production activities (Wernet et al., \u003cspan citationid=\"CR52\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Asselin- Balen\u0026ccedil;on et al., 2020 ; Nemecek et al., 2019). We retrieved LCIs for 991 crop production activities representing 121 crop species cultivated in 57 regions (see Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e and Figures S2 and S3 of the supplementary information for further details). We normalised all inputs per 1 kilogram of produced crop, i.e., the functional unit. We considered impacts from the extraction of raw materials up to the farm gate.\u003c/p\u003e \u003cp\u003eWe used the ReCiPe 2016 methodology to quantify the biodiversity, water, and climate footprints (Huijbregts et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). We chose the endpoint indicator damage to ecosystems (referred to as biodiversity footprint in this work) [species-yr] which includes terrestrial, marine and freshwater biodiversity loss. We also used two midpoint indicators, namely water use [m\u003csup\u003e3\u003c/sup\u003e] and climate change [kg CO\u003csub\u003e2\u003c/sub\u003e eq], referred to as respectively water footprint and climate footprint in this work.\u003c/p\u003e \u003cp\u003eWe used the Brightway 2.5 Python package to calculate footprints for each ReCiPe 2016 perspective (Hierarchist, Individualist, Egalitarian) and to retrieve the crop LCIs (Mutel, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Here, we show the results based on the Hierarchist perspective (Section \u003cspan refid=\"Sec8\" class=\"InternalRef\"\u003e3\u003c/span\u003e), as within the ReCiPe documentation and literature, the Hierarchist perspective is the \u0026ldquo;default\u0026rdquo; for LCA calculations (Huijbregts et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). The results based on the Individualist and Egalitarian perspectives can be found in sections S8 and S9 of the supplementary information.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Predictor selection and preprocessing\u003c/h2\u003e \u003cp\u003eWe selected on-farm information concerning climate region and crop type as well as key inputs of material, land area, energy and water as predictors and grouped them based on their function (i.e., used towards the same goal) and form (i.e., unit). For instance, calcium ammonium and urea both serve as nitrogen fertilisers and are measured in kilograms. We therefore aggregated them in a single predictor by summing their nitrogen content. Furthermore, we included categorical predictors representing the climate region of crop production activity and whether the crop is perennial or annual (for an overview of the implementation of these predictors see S4). This resulted in 16 predictors representing six overarching predictor themes (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e). For a detailed overview of predictors and farming information, see Table S5.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003e\u0026ndash; Overview of predictors used to quantify environmental footprints of crop production.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"4\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePredictor \u003c/p\u003e \u003cp\u003etheme\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePredictor\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eUnit\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDescription\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"2\" rowspan=\"3\"\u003e \u003cp\u003eEnergy\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eElectricity\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekWh\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eElectricity used to run greenhouses and supporting activities in other productions\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eFuel\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMJ\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eDiesel fuel burned in agricultural machinery\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eHeat\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMJ\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eHeat used in greenhouses\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"6\" rowspan=\"7\"\u003e \u003cp\u003eFertilisers \u0026amp; \u003c/p\u003e \u003cp\u003esoil amendment\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePotassium fertilisers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eKg active ingredient (K\u003csub\u003e2\u003c/sub\u003eO)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePotassium rich fertilisers such as potassium chloride, potassium sulfite, vinasse\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNitrogen fertilisers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg active ingredient (N)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eNitrogen rich fertilisers such as urea, ammonia, ammonium nitrate, calcium nitrate\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePhosphorus fertilisers\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg active ingredient (P\u003csub\u003e2\u003c/sub\u003eO\u003csub\u003e5\u003c/sub\u003e)\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePhosphorus rich fertilisers such as superphosphates, P\u003csub\u003e2\u003c/sub\u003eO\u003csub\u003e5\u003c/sub\u003e, filter cake\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eManure\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg manure\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eManure used for fertilisation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eMicronutrients\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg micronutrient\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMaterials that supply micronutrients (e.g., Fe, Mo, Ni, Zn, Cu) to the crop such as\u003c/p\u003e \u003cp\u003e calcium borates, zinc, manganese sulfate\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eSoil improvement\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg soil conditioner\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ematerials used to correct soil pH and condition soil, such as lime, gypsum, dolomite\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eStimulants\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg stimulant\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMaterials that enhance growth, such as biostimulants, growth regulators\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003ePlant protection\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ePlant protection material\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg active ingredient\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003ePesticide, herbicide, fungicides and other materials used for protecting the crops and plants\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eWater use\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eIrrigation water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003em\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWater used for irrigation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eTap water\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003em\u003csup\u003e3\u003c/sup\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWater used for dilution of pesticide\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eYield\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eYield\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003ekg/ha\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eYield of each crop calculated based on land occupation\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\" morerows=\"1\" rowspan=\"2\"\u003e \u003cp\u003eCategorical information\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eClimate region\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWhether crop is located in a tropical, arid, temperate, continental or polar climate (according to S4)\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCrop type\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e-\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c4\"\u003e \u003cp\u003eWhether crop is perennial or annual\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTo check for multicollinearity, we calculated Variance Inflation Factors (VIFs) for the predictor variables and observed that all VIF scores were smaller than 5, which is indicative of relatively low multicollinearity among the predictors (Ferre, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2009\u003c/span\u003e). We then scaled all predictor values to the range from 0 to 1 because some machine learning algorithms, particularly ANNs, are sensitive to data scaling (Pedregosa et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). We rescaled the predictor values using the MinMaxScaler function of the Scikit-learn package in Python (Pedregosa et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2011\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Machine learning modelling techniques\u003c/h2\u003e \u003cp\u003eWe used five machine learning modelling techniques to model footprints: Random Forest (RF), K-Nearest Neighbours (KNN), Artificial Neural Network (ANN), Generalised Boosting Method (GBM) and a regularised Linear Model (LM). All five machine learning modelling techniques have been shown to achieve high predictive performance in estimating footprints in agricultural, construction and chemical LCA studies (Meng et al., \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e2019\u003c/span\u003e; Romeiko et al., \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Thilakarathna et al., \u003cspan citationid=\"CR48\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Xikai et al., \u003cspan citationid=\"CR55\" class=\"CitationRef\"\u003e2019\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eRF is a regression and classification algorithm comprised of an ensemble of decision trees. Each tree is constructed based on a random subset of data and predictors and operates by iteratively splitting the data at each node, based on the predictor that minimises variance, thus achieving more uniform subsets of data on either branch of the node (Luan et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The output of the RF is an average of outputs of individual trees, making it resistant to scaling and data availability issues (Breiman, \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2001\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eKNN regression is a predictive algorithm that estimates a data point's target value by averaging the target values of a certain number (k) of its nearest neighbours, with k being specified by the user. KNN is a non-parametric algorithm, with no assumptions on the underlying data. It makes predictions solely based on the distances between data points, also making the algorithm sensitive to the number of neighbours and distance metric chosen (Peterson, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2009\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eANNs emulate the learning process of the human brain. They consist of input, output and hidden layers, within which, neurons or nodes exist with activation functions that project data to nonlinear spaces. ANNs are therefore capable of identifying non-linear relationships (Das et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2016\u003c/span\u003e\u0026nbsp;; Song et al., \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e2017\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eGBMs are ensemble learning techniques that combine multiple weak learners, such as decision trees, to create a strong predictive model through an additive training strategy, where each new learner is trained to correct the errors of the previous ones (Qui et al., 2022).\u003c/p\u003e \u003cp\u003eLMs are algorithms that model the relationship between dependent and independent variables using a linear equation (Pedregosa et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). The LM implementation used here includes a regularization term to prevent overfitting and is commonly known as Ridge Regression (McDonald, \u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e2009\u003c/span\u003e).\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e\u003cem\u003e2.5 Model fitting and evaluation\u003c/em\u003e\u003c/h2\u003e \u003cp\u003eWe used a five-fold block cross-validation to evaluate the models\u0026rsquo; performance to predict footprints of crops that are not part of the training data and to avoid overfitting. In each run, 80% of the data (i.e. four folds) is used for training, and the remaining 20% of the data (i.e. one fold) is used for validation of the model, resulting in five runs. We split the data based on crop species, ensuring the same crops would not be represented in both training and validation sets to prevent data leakage. To quantify the performance of each model, we calculated for each validation set the explained variance (R\u003csup\u003e2\u003c/sup\u003e) and the Root Mean Squared Error (RMSE) (Chai \u0026amp; Draxler, \u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e2014\u003c/span\u003e). We then averaged the R\u003csup\u003e2\u003c/sup\u003e and RMSE over the five folds to obtain the overall performance for each model.\u003c/p\u003e \u003cp\u003eWe applied each model 16 times (once using 1 predictor, once using two, and so on until 16 predictors). We optimised each model by finding the predictor subset that minimises the cross-validated RMSE. We then tuned the hyperparameters of each machine learning modelling technique (e.g. number of neighbours in KNN, number of trees in RF, learning rate in GBM; Table S6) using Bayesian optimization, with the goal of maximizing predictive performance. We performed the optimizations using the Optuna package in Python (Akiba et al., \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2019\u003c/span\u003e). The models with optimal hyperparameters are specified in section S7 of the supplementary information. We built all modelling pipelines using the Scikit-learn Python package (Pedregosa et al., \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2011\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eTo determine the most parsimonious model per footprint, we used the knee (also known as \u0026ldquo;elbow\u0026rdquo;) method to identify the number of predictors beyond which the inclusion of more predictors yields a negligible increase in predictive performance (i.e. lowering RMSE) (Satopaa et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). The knee point is defined as the point of maximum curvature in any function, with curvature being a mathematical measure of how much a function differs from a straight line (Satopaa et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). As a result, maximum curvature captures where performance improvements start to level off (Satopaa et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2011\u003c/span\u003e). We used the Python package kneed (Arvai, \u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e2020\u003c/span\u003e) to find the knee point on the line connecting the best performing machine learning modelling technique for each predictor subset size.\u003c/p\u003e \u003cp\u003eTo quantify the importance of each predictor to the environmental footprint predictions, we used Shapley values. Shapley values are based on game theory and explain machine learning model outputs by calculating the average marginal contribution of each predictor across all possible subsets of predictors (Sundararajan \u0026amp; Najmi, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). We used the SHAP python package to calculate the Shapley values (Lundberg \u0026amp; Lee, 2017).\u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results and Discussion","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.1 Predictive power and parsimony\u003c/h2\u003e \u003cp\u003eBased on all predictors considered (Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e), we found block cross-validated R\u003csup\u003e2\u003c/sup\u003e values for the water and biodiversity footprint models of 0.90 (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e) using GBM (best performing model). For the climate footprint model, we found a block cross-validated R\u003csup\u003e2\u003c/sup\u003e of 0.64. This performance is lower than typically reported for predictive models of climate footprints, as R\u0026sup2; scores between 0.64 and 0.92 are reported in previous work focused on single crop species ( Khanali et al., 2017 ; Khoshnevisan et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2013\u003c/span\u003e ; Lee et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2020\u003c/span\u003e ; Romeiko et al., \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). The comparatively lower R\u003csup\u003e2\u003c/sup\u003e of our model may reflect that we i) considered only on-farm information as predictors, excluding information on infrastructure, crop residue treatment, packaging, labour, machinery and seed production which were included in previous work and are known to be additional drivers of the climate footprint in some crop-country combinations (Khanali et al., 2017; Lam et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Nabavi-Pelesaraei et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Verge et al., \u003cspan citationid=\"CR50\" class=\"CitationRef\"\u003e2007\u003c/span\u003e; Liu et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2016\u003c/span\u003e; Yan et al., \u003cspan citationid=\"CR56\" class=\"CitationRef\"\u003e2014\u003c/span\u003e; Aguilera et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2015\u003c/span\u003e), ii) included a large variety of crops and countries, and iii) performed block cross-validation to test predictions for crop species that are not included in the training set. On the one hand, our approach to train the models on a heterogeneous database of crops enhances its generalizability to other crops, locations and practices. On the other hand, our models will yield lower predictive performance for the footprints of a specific crop species compared to models specifically trained on that particular crop or closely related data, especially considering that multiple crops only have one entry in our dataset.\u003c/p\u003e \u003cp\u003eWe found that when using a small subset of predictors, a prediction performance could be attained similar to that of using a larger set of predictors (Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e). We observed that using subsets of more than five predictors (identified via the knee method) hardly increases or even decreases performance as measured by both RMSE (in natural log-space) and R\u003csup\u003e2\u003c/sup\u003e (e.g., for climate footprints), underscoring the importance of parsimony in model selection. This decrease in performance is due to predictors adding more noise than information to the models\u0026rsquo; learning, especially predictors which are important only to a few specific crops. KNN is most sensitive to this, leading to a significant decrease in performance past the knee point. In contrast, previous work has relied on all available input data for modelling footprints (Romeiko et al., \u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Kahanali et al., 2017; Khoshnevisan et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Lam et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e2020\u003c/span\u003e; Nabavi-Pelesaraei et al., \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Lee et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eOverall, GBM, ANN and RF outperform the other two machine learning modelling techniques (KNN and LM) over the majority of predictor subsets for the estimation of all three footprints. We found RF to be the most parsimonious model with four predictors for the biodiversity footprint (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.88, RMSE\u0026thinsp;=\u0026thinsp;0.37). We found GBM using five predictors as the most parsimonious model for water footprint (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.89, RMSE\u0026thinsp;=\u0026thinsp;0.83) and ANN for climate footprint (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.58, RMSE\u0026thinsp;=\u0026thinsp;0.71) with four predictors. These RMSE values correspond to an uncertainty of +/- a factor 2.1, 5.1, and 4.0 (95% confidence interval) for biodiversity, water, and climate footprints, respectively. RMSE and R\u003csup\u003e2\u003c/sup\u003e reveal consistent patterns (i.e. RMSE is lowest when R\u003csup\u003e2\u003c/sup\u003e is highest, both indicating the best model; Fig.\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Relative contribution of different predictors to model prediction\u003c/h2\u003e \u003cp\u003eYield is the most important predictor of the biodiversity footprint (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e), as land use is the primary driver of biodiversity loss and increasing yield directly reduces the land use area required for crop production (Hald-Mortensen, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2023\u003c/span\u003e; Jaureguiberry et al., \u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). It should be noted that whilst our predictors (on-farm information) likely represent heterogenous farming practices (e.g. monoculture vs intercropping, or organic vs conventional agriculture) which influence the biodiversity footprint, our models have not been trained with farm management as an input parameter. In addition, land use intensity effects on local biodiversity are not captured in the ReCiPe 2016 methodology (Huijbregts et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Phosphorus fertiliser use, irrigation water, and heat are the next most important predictors for the biodiversity footprint. Phosphorus impacts biodiversity via eutrophication of freshwater (Huijbregts et al., \u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Sud, 2020).\u003c/p\u003e \u003cp\u003eFor the water footprint, we found that irrigation, nitrogen fertiliser, electricity use and plant protection inputs were the most important predictors (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The latter may reflect the water use in the upstream production and application of fertilisers and pesticides (Sud, 2020).\u003c/p\u003e \u003cp\u003eWhile we find that the biodiversity and water footprints are largely driven by a single predictor (yield and irrigation, respectively), we find that contributions to climate footprints are more evenly spread across several predictors (yield with 42.1%, nitrogen fertiliser input with 25.9%, climate region with 21% and electricity use with 10.8%). The importance of yield in climate footprint estimation is in line with other studies finding associations between decreasing footprints and increased yield (Lam et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e; Pishgar-Komleh et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Nitrogen fertiliser is also known to be a driver of climate footprints through N\u003csub\u003e2\u003c/sub\u003eO and NOx emissions from the soil subsequent to its application (Rosegrant et al., 2002). Lastly, land use change based emissions vary across countries and climate zones even for the same activity (Reinhard et al., \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2017\u003c/span\u003e). Therefore climate region information helps the model distinguish between these types of emissions. While using country-level data would be more precise, it would greatly increase dimensionality, reducing the robustness of the model.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Implications, limitations and future research directions\u003c/h2\u003e \u003cp\u003eThe block cross-validation revealed that, based on groups of on-farm information, machine learning models can predict the biodiversity and water footprints of the production of \u0026lsquo;new\u0026rsquo; crops (i.e., crop species on which the models were not trained) with relatively high accuracy (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;\u0026gt;\u0026thinsp;0.89). This could be partially explained by biodiversity and water footprints being strongly related to single pressures (namely land and water use, respectively) (Reinhard et al., \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2017\u003c/span\u003e; Fr\u0026oacute;na et al., 2019). Climate footprints are more challenging to predict, partially due to the many different sources and drivers of GHG emissions (Poore \u0026amp; Nemecek, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e2018\u003c/span\u003e; Khoshnevisan et al., \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2013\u003c/span\u003e; Lam et al., \u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). We recommend the use of GBM, RF and ANN (in order of performance) over KNN and LM when predicting footprints of crop production based on on-farm information, and we recommend caution when applying predictive models to estimate footprints that are related to many different processes such as climate footprints.\u003c/p\u003e \u003cp\u003eOur analysis showed that a small set of on-farm predictors can achieve nearly the same prediction accuracy as using many predictors. To our knowledge, no previous work has examined the trade-off between the number of predictors used and prediction accuracy in the context of crop production footprints. We recommend prioritizing data collection or data gap-filling of agricultural yield, (nitrogen and phosphorus) fertiliser use, pesticide use, irrigation, and electricity use when aiming to assess biodiversity, water, or climate footprints.\u003c/p\u003e \u003cp\u003eWhile most of the existing literature examines the contribution of farming information to footprints in a mechanistic manner, the only similar machine learning-based study we found, also identified fertiliser as a key driver of climate footprints, despite using different predictors (Romeiko et al., \u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). Our methodology can be used to provide impact estimates for screening purposes when detailed data are unavailable or a quick identification of high-impact crop production activities is desired.\u003c/p\u003e \u003cp\u003eOur methodology can also be used to reveal major drivers of other footprints and other products to guide LCI data collection and data gap-filling efforts. Lastly, we found that categorical predictors such as climatic region of the farm location (i.e. tropical, arid, temperate, continental and polar) and crop type (i.e. perennial or annual) provide additional, valuable and relatively easy information to the models, although they lack the detail to provide good performance on their own (see S12 for performance metrics of modelling techniques using only categorical predictors as well as categorical predictors and most important numerical predictor; see S13 for an overview of categorical predictors).\u003c/p\u003e \u003cp\u003eWe encountered several challenges that exposed opportunities for future research. First, the databases we used are mostly comprised of national average crop production LCIs, which do not capture the full variety of farming management practices (e.g. various land use intensities). This means that we have been unable to test the accuracy of our predictions for the same crops in the same region but with different management characteristics. Compiling a higher-resolution database covering a wider range of farming practices and temporal variation would improve the validity and representativeness of model predictions. Second, the aggregation of farming information into predictors, such as the aggregation of specific fertilisers (e.g., ammonium nitrate) into fertiliser groups (e.g., nitrogen fertiliser) based on their active ingredients (e.g., kg nitrogen), or aggregating electricity inputs into a single predictor regardless of generation method, masks potential differences in impacts within these groups. Although we distinguished manure from nitrogen, phosphorus, and potassium fertilisers, we did not distinguish other organic fertilisers from inorganic fertilisers, which may have different impacts depending on their production processes and associated on-field emissions (Garc\u0026iacute;a Castellanos et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Thus alternative predictor grouping strategies could affect model outcomes and insights. Third, here we focused on biodiversity, water and climate footprints as they are important to agricultural systems. Future research could evaluate the applicability of ML models to other impact indicators. Finally, we omitted several information including packaging material, on-farm preprocessing, seed and seedling input, and land use change, due to the boundaries of our study (i.e. considering only on-farm or foreground information). While this limits the explicit inclusion of upstream and indirect processes, expanding the system boundary in future work could allow the evaluation of identified relationships in a more comprehensive setting.\u003c/p\u003e \u003cp\u003eGiven the scarcity of LCI data for many crop-country-practice combinations, and the difficulty in collecting them, here we evaluated the potential for predicting impacts using machine learning techniques and a more limited set of data than normally captured in LCIs (i.e. on-farm information). Overall, this study shows that machine learning is a promising tool for predicting some environmental footprints of crop production based on on-farm information, although prediction accuracy varies by footprint. We found that climate region of the farm, yield, irrigation, fertiliser and pesticide inputs, and the use of electricity and heat, are sufficient to achieve predictive accuracy close to models using more inputs. The parsimonious models developed in this work provide a systematic approach to identifying priorities for data gap-filling, thereby facilitating a broadening of the range of crops for which environmental footprints can be estimated.\u003c/p\u003e \u003c/div\u003e"},{"header":"Declarations","content":" \u003cp\u003e \u003cstrong\u003eCompeting interests:\u003c/strong\u003e \u003cp\u003eThe authors declare no competing interests relevant to the content of this article.\u003c/p\u003e \u003c/p\u003e\u003ch2\u003eAuthor Contribution\u003c/h2\u003e\u003cp\u003eAll authors contributed to the study conception and design. Material preparation, data collection and analysis were performed by F.R.. All authors discussed the results. F.R. wrote the original draft with inputs from all authors. All authors worked on the revisions to the paper and approved this manuscript.\u003c/p\u003e\u003ch2\u003eAcknowledgement\u003c/h2\u003e\u003cp\u003eWe would like to thank Katy Armstrong and Florence Bohnes, internal reviewers from Unilever\u0026rsquo;s Safety, Environmental and Regulatory Sciences group, for their insightful comments. Moreover, we would like to thank Selwyn Hoeks from Radboud University and Leonardo Contreas from Unilever for their help with programming issues and code review.\u003c/p\u003e\u003ch2\u003eData Availability\u003c/h2\u003e\u003cp\u003eAll scripts that support the findings of this study are available in Github at: https://github.com/FRaymand/parsimonious-crop-footprint-estimator\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\u003cli\u003e\u003cspan\u003eAguilera, E., Guzm\u0026aacute;n, G., \u0026amp; Alonso, A. (2015). Greenhouse gas emissions from conventional and organic cropping systems in Spain. I. Herbaceous crops. \u003cem\u003eAgronomy for Sustainable Development\u003c/em\u003e, 35(2), 713\u0026ndash;724. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s13593-014-0267-9\u003c/span\u003e\u003cspan address=\"10.1007/s13593-014-0267-9\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAkiba, T., Sano, S., Yanase, T., Ohta, T., \u0026amp; Koyama, M. (2019). Optuna: A Next-generation Hyperparameter Optimization Framework. \u003cem\u003eProceedings of the ACM SIGKDD International Conference on Knowledge Discovery and Data Mining\u003c/em\u003e, 2623\u0026ndash;2631. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1145/3292500.3330701\u003c/span\u003e\u003cspan address=\"10.1145/3292500.3330701\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAlgunaibet, I. M., \u0026amp; Guill\u0026eacute;n-Gos\u0026aacute;lbez, G. (2019). Life cycle burden-shifting in energy systems designed to minimize greenhouse gas emissions: Novel analytical method and application to the United States. \u003cem\u003eJournal of Cleaner Production\u003c/em\u003e, 229, 886\u0026ndash;901.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eArvai, K. (2020). kneed (Version 0.7.0) [Computer software]. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5281/zenodo.6496267\u003c/span\u003e\u003cspan address=\"10.5281/zenodo.6496267\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eAsselin-Balen\u0026ccedil;on, A., Broekema, R., Teulon, H., Gastaldi, G., Houssier, J., Moutia, A., \u0026hellip; \u0026amp; Colomb, V. (2020). AGRIBALYSE v3. 0: the French agricultural and food LCI database. \u003cem\u003eMethodology for the food products\u003c/em\u003e. \u003cem\u003eEd. ADEME\u003c/em\u003e.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBellon-Maurel, V., Short, M. D., Roux, P., Schulz, M., \u0026amp; Peters, G. M. (2014). Streamlining life cycle inventory data generation in agriculture using traceability data and information and communication technologies\u0026ndash;part I: concepts and technical basis. \u003cem\u003eJournal of Cleaner Production\u003c/em\u003e, 69, 60\u0026ndash;66. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jclepro.2014.01.079\u003c/span\u003e\u003cspan address=\"10.1016/j.jclepro.2014.01.079\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBorghino, N., Corson, M., Nitschelm, L., Wilfart, A., Fleuet, J., Moraine, M., \u0026hellip; Godinot, O. (2021). Contribution of LCA to decision making: A scenario analysis in territorial agricultural production systems. \u003cem\u003eJournal of Environmental Management\u003c/em\u003e, 287, 112288.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eBreiman, L. (2001). Random Forests. \u003cem\u003eMachine Learning\u003c/em\u003e 45, 5\u0026ndash;32. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1023/A:1010933404324\u003c/span\u003e\u003cspan address=\"10.1023/A:1010933404324\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eCanals, L. M. I., Azapagic, A., Doka, G., Jefferies, D., King, H., Mutel, C., Nemecek, T., Roches, A., Sim, S., Stichnothe, H., Thoma, G., \u0026amp; Williams, A. (2011). Approaches for addressing life cycle assessment data gaps for bio-based products. \u003cem\u003eJournal of Industrial Ecology\u003c/em\u003e, 15(5), 707\u0026ndash;725. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/j.1530-9290.2011.00369.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1530-9290.2011.00369.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eChai, T., \u0026amp; Draxler, R. R. (2014). Root mean square error (RMSE) or mean absolute error (MAE)? -Arguments against avoiding RMSE in the literature. \u003cem\u003eGeoscientific Model Development\u003c/em\u003e, 7(3), 1247\u0026ndash;1250. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5194/gmd-7-1247-2014\u003c/span\u003e\u003cspan address=\"10.5194/gmd-7-1247-2014\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDas, P. J., Preuss, C., \u0026amp; Mazumder, B. (2016). Artificial Neural Network as Helping Tool for Drug Formulation and Drug Administration Strategies. \u003cem\u003eArtificial Neural Network for Drug Design, Delivery and Disposition\u003c/em\u003e, 263\u0026ndash;276. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/B978-0-12-801559-9.00013-2\u003c/span\u003e\u003cspan address=\"10.1016/B978-0-12-801559-9.00013-2\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDonati, F., Dente, S. M., Li, C., Vilaysouk, X., Froemelt, A., Nishant, R., \u0026hellip; Hashimoto, S. (2022). The future of artificial intelligence in the context of industrial ecology. \u003cem\u003eJournal of Industrial Ecology\u003c/em\u003e, \u003cem\u003e26\u003c/em\u003e(4), 1175\u0026ndash;1181.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eDudley, N., \u0026amp; Alexander, S. (2017). Agriculture and biodiversity: a review. \u003cem\u003eBiodiversity\u003c/em\u003e, 18(2\u0026ndash;3), 45\u0026ndash;49. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1080/14888386.2017.1351892\u003c/span\u003e\u003cspan address=\"10.1080/14888386.2017.1351892\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFAO. 2022. The State of the World\u0026rsquo;s Land and Water Resources for Food and Agriculture \u0026ndash; Systems at breaking point. Main report. Rome. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.4060/cb9910en\u003c/span\u003e\u003cspan address=\"10.4060/cb9910en\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFr\u0026oacute;na, D., Szender\u0026aacute;k, J., \u0026amp; Harangi-R\u0026aacute;kos, M. (2021). The challenge of feeding the world sustainably. \u003cem\u003eChall. Feed. World Sustain\u003c/em\u003e, 11, 1\u0026ndash;36. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.17226/26007\u003c/span\u003e\u003cspan address=\"10.17226/26007\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eFerre, J. (2009). Variance inflation factor: An overview. \u003cem\u003eScience Direct\u003c/em\u003e, 3, 1\u0026ndash;24.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eGarc\u0026iacute;a Castellanos, B., Garc\u0026iacute;a Garc\u0026iacute;a, B., \u0026amp; Garc\u0026iacute;a Garc\u0026iacute;a, J. (2023). Economic and environmental effects of replacing inorganic fertilizers with organic fertilizers in three rainfed crops in a semi-arid area. \u003cem\u003eSustainability\u003c/em\u003e, \u003cem\u003e15\u003c/em\u003e(24), 16897. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/su152416897\u003c/span\u003e\u003cspan address=\"10.3390/su152416897\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHald-Mortensen, C. (2023). The main drivers of biodiversity loss: a brief overview. \u003cem\u003eJournal of Ecology and Natural Resources\u003c/em\u003e, 7(3), 000346. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.23880/jenr-16000346\u003c/span\u003e\u003cspan address=\"10.23880/jenr-16000346\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHoogeveen, J., Faur\u0026egrave;s, J. M., Peiser, L., Burke, J., \u0026amp; Van De Giesen, N. (2015). GlobWat - A global water balance model to assess water use in irrigated agriculture. \u003cem\u003eHydrology and Earth System Sciences\u003c/em\u003e, 19(9), 3829\u0026ndash;3844. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.5194/hess-19-3829-2015\u003c/span\u003e\u003cspan address=\"10.5194/hess-19-3829-2015\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eHuijbregts, M. A., Steinmann, Z. J., Elshout, P. M., Stam, G., Verones, F., Vieira, M., Zijp, M., Hollander, A. \u0026amp; Van Zelm, R. (2017). ReCiPe2016: a harmonised life cycle impact assessment method at midpoint and endpoint level. \u003cem\u003eThe International Journal of Life Cycle Assessment\u003c/em\u003e, 22, 138\u0026ndash;147.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJacob-Lopes, E., Zepka, L. Q., \u0026amp; Depr\u0026aacute;, M. C. (2021). \u003cem\u003eSustainability metrics and indicators of environmental impact: industrial and agricultural life cycle assessment\u003c/em\u003e. Elsevier. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/C2020-0\u003c/span\u003e\u003cspan address=\"10.1016/C2020-0\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e-00268-9Khanali, M., Mobli, H., \u0026amp; Hosseinzadeh-Bandbafha, H. (2017). Modeling of yield and environmental impact categories in tea processing units based on artificial neural networks. \u003cem\u003eEnvironmental Science and Pollution Research\u003c/em\u003e, 24(34), 26324\u0026ndash;26340. https://doi.org/10.1007/s11356-017-0234-5\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eJaureguiberry, P., Titeux, N., Wiemers, M., Bowler, D. E., Coscieme, L., Golden, A. S., \u0026hellip; Purvis, A. (2022). The direct drivers of recent global anthropogenic biodiversity loss. \u003cem\u003eScience advances\u003c/em\u003e, 8(45), eabm9982. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/sciadv.abm9982\u003c/span\u003e\u003cspan address=\"10.1126/sciadv.abm9982\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eKhoshnevisan, B., Rafiee, S., Omid, M., Mousazadeh, H., \u0026amp; Sefeedpari, P. (2013). Prognostication of environmental indices in potato production using artificial neural networks. \u003cem\u003eJournal of Cleaner Production\u003c/em\u003e, 52, 402\u0026ndash;409. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jclepro.2013.03.028\u003c/span\u003e\u003cspan address=\"10.1016/j.jclepro.2013.03.028\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLaborde, D., Mamun, A., Martin, W., Pi\u0026ntilde;eiro, V., \u0026amp; Vos, R. (2021). Agricultural subsidies and global greenhouse gas emissions. \u003cem\u003eNature Communications\u003c/em\u003e, 12(1), 2601. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1038/s41467-021-22703-1\u003c/span\u003e\u003cspan address=\"10.1038/s41467-021-22703-1\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLam, W. Y., Sim, S., Kulak, M., van Zelm, R., Schipper, A. M., \u0026amp; Huijbregts, M. A. J. (2021). Drivers of variability in greenhouse gas footprints of crop production. \u003cem\u003eJournal of Cleaner Production\u003c/em\u003e, 315, 128121. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jclepro.2021.128121\u003c/span\u003e\u003cspan address=\"10.1016/j.jclepro.2021.128121\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLee, E. K., Zhang, W. J., Zhang, X., Adler, P. R., Lin, S., Feingold, B. J., \u0026hellip; Romeiko, X. X. (2020). Projecting life-cycle environmental impacts of corn production in the US Midwest under future climate scenarios using a machine learning approach. \u003cem\u003eScience of The Total Environment\u003c/em\u003e, 714, 136697. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.scitotenv.2020.136697\u003c/span\u003e\u003cspan address=\"10.1016/j.scitotenv.2020.136697\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLiu, C., Cutforth, H., Chai, Q., \u0026amp; Gan, Y. (2016). Farming tactics to reduce the carbon footprint of crop cultivation in semiarid areas. A review. \u003cem\u003eAgronomy for Sustainable Development\u003c/em\u003e, 36(4), 69. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s13593-016-0404-8\u003c/span\u003e\u003cspan address=\"10.1007/s13593-016-0404-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eLuan, J., Zhang, C., Xu, B., Xue, Y., \u0026amp; Ren, Y. (2020). The predictive performances of random forest models with limited sample size and different species traits. \u003cem\u003eFisheries Research\u003c/em\u003e, 227, 105534. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.fishres.2020.105534Lundberg\u003c/span\u003e\u003cspan address=\"https://doi.org/10.1016/j.fishres.2020.105534Lundberg\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, S. M., \u0026amp; Lee, S. I. (2017). A unified approach to interpreting model predictions. \u003cem\u003eAdvances in neural information processing systems\u003c/em\u003e, 30.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMcDonald, G. C. (2009). Ridge regression. Wiley Interdisciplinary Reviews: \u003cem\u003eComputational Statistics\u003c/em\u003e, 1(1), 93\u0026ndash;100.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMeng, F., LaFleur, C., Wijesinghe, A., \u0026amp; Colvin, J. (2019). Data-driven approach to fill in data gaps for life cycle inventory of dual fuel technology. \u003cem\u003eFuel\u003c/em\u003e, 246, 187\u0026ndash;195. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.fuel.2019.02.124\u003c/span\u003e\u003cspan address=\"10.1016/j.fuel.2019.02.124\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eMutel, C. (2017). Brightway: An open source framework for Life Cycle Assessment. \u003cem\u003eJournal of Open Source Software\u003c/em\u003e, 2(12), 236. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.21105/joss.00236\u003c/span\u003e\u003cspan address=\"10.21105/joss.00236\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNabavi-Pelesaraei, A., Rafiee, S., Mohtasebi, S. S., Hosseinzadeh-Bandbafha, H., \u0026amp; Chau, K. W. (2018). Integration of artificial intelligence methods and life cycle assessment to predict energy output and environmental impacts of paddy production. \u003cem\u003eScience of the Total Environment\u003c/em\u003e, 631, 1279\u0026ndash;1294. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.scitotenv.2018.03.088\u003c/span\u003e\u003cspan address=\"https://doi.org/10.1016/j.scitotenv.2018.03.088\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e Nemecek, T., Bengoa, X., Lansche, J., Roesch, A., Faist-Emmenegger, M., Rossi, V., \u0026hellip; Riedener, E. (2019). World food LCA database. \u003cem\u003eMethodological Guidelines for the Life Cycle Inventory of Agricultural Products\u003c/em\u003e, 3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eNorris, K. (2008). Agriculture and biodiversity conservation: opportunity knocks. \u003cem\u003eConservation Letters\u003c/em\u003e, 1(1), 2\u0026ndash;11. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/j.1755-263x.2008.00007.x\u003c/span\u003e\u003cspan address=\"10.1111/j.1755-263x.2008.00007.x\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., \u0026hellip; Duchesnay, \u0026Eacute;. (2011). Scikit-learn: Machine learning in Python. \u003cem\u003eJournal of Machine Learning Research\u003c/em\u003e, 12, 2825\u0026ndash;2830. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.48550/arXiv.1201.0490\u003c/span\u003e\u003cspan address=\"10.48550/arXiv.1201.0490\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePeterson, L. (2009). K-nearest neighbor. \u003cem\u003eScholarpedia\u003c/em\u003e, 4(2), 1883. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.4249/SCHOLARPEDIA.1883\u003c/span\u003e\u003cspan address=\"10.4249/SCHOLARPEDIA.1883\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePishgar-Komleh, S. H., Akram, A., Keyhani, A., Raei, M., Elshout, P. M. F., Huijbregts, M. A. J., \u0026amp; Van Zelm, R. (2017). Variability in the carbon footprint of open-field tomato production in Iran-A case study of Alborz and East-Azerbaijan provinces. \u003cem\u003eJournal of Cleaner Production\u003c/em\u003e, 142, 1510\u0026ndash;1517. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttp://dx.doi.org/10.1016/j.jclepro.2016.11.154\u003c/span\u003e\u003cspan address=\"10.1016/j.jclepro.2016.11.154\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePishgar-Komleh, S. H., Akram, A., Keyhani, A., Sefeedpari, P., Shine, P., \u0026amp; Brandao, M. (2020). Integration of life cycle assessment, artificial neural networks, and metaheuristic optimization algorithms for optimization of tomato-based cropping systems in Iran. \u003cem\u003eInternational Journal of Life Cycle Assessment\u003c/em\u003e, 25(3), 620\u0026ndash;632. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11367-019-01707-6\u003c/span\u003e\u003cspan address=\"10.1007/s11367-019-01707-6\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003ePoore, J., \u0026amp; Nemecek, T. (2018). Reducing food\u0026rsquo;s environmental impacts through producers and consumers. \u003cem\u003eScience\u003c/em\u003e, 360(6392), 987\u0026ndash;992. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1126/science.aaq0216\u003c/span\u003e\u003cspan address=\"10.1126/science.aaq0216\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eQiu, R., Liu, C., Cui, N., Gao, Y., Li, L., Wu, Z., Jiang, S., \u0026amp; Hu, M. (2022). Generalized Extreme Gradient Boosting model for predicting daily global solar radiation for locations without historical data. \u003cem\u003eEnergy Conversion and Management\u003c/em\u003e, 258. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.enconman.2022.115488\u003c/span\u003e\u003cspan address=\"10.1016/j.enconman.2022.115488\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eReinhard J., Moreno-Ruiz E., \u0026amp; Gm\u0026uuml;nder S. (2017) Consideration of land use change in ecoinvent version 3.3: Method, Implementation and Illustration. White paper, ecoinvent Association, Z\u0026uuml;rich, Switzerland\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRomeiko, X. X., Guo, Z., Pang, Y., Lee, E. K., \u0026amp; Zhang, X. (2020). Comparing machine learning approaches for predicting spatially explicit life cycle global warming and eutrophication impacts from corn production. \u003cem\u003eSustainability\u003c/em\u003e, 12(4), 1481. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.3390/su12041481\u003c/span\u003e\u003cspan address=\"10.3390/su12041481\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRomeiko, X. X., Zhang, X., Pang, Y., Gao, F., Xu, M., Lin, S., \u0026amp; Babbitt, C. (2023). A review of machine learning applications in life cycle assessment studies. \u003cem\u003eScience of the Total Environment\u003c/em\u003e, 168969.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eRosegrant, M. W., Cai, X., \u0026amp; Cline, S. A. (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://hdl.handle.net/10568/1567952002\u003c/span\u003e\u003cspan address=\"https://hdl.handle.net/10568/1567952002\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e). World water and food to 2025: dealing with scarcity. \u003cem\u003eIntl Food Policy Res Inst\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://hdl.handle.net/10568/156795\u003c/span\u003e\u003cspan address=\"https://hdl.handle.net/10568/156795\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSatopaa, V., Albrecht, J., Irwin, D., \u0026amp; Raghavan, B. (2011, June). Finding a\" kneedle\" in a haystack: Detecting knee points in system behavior. In 2011 31st international conference on distributed computing systems workshops (pp. 166\u0026ndash;171). \u003cem\u003eIEEE\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1109/ICDCSW.2011.20\u003c/span\u003e\u003cspan address=\"10.1109/ICDCSW.2011.20\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSong, R., Keller, A. A., \u0026amp; Suh, S. (2017). Rapid Life-Cycle Impact Screening Using Artificial Neural Networks. \u003cem\u003eEnvironmental Science and Technology\u003c/em\u003e, 51(18), 10777\u0026ndash;10785. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1021/acs.est.7b02862\u003c/span\u003e\u003cspan address=\"10.1021/acs.est.7b02862\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSud, M. (2020). Managing the biodiversity impacts of fertiliser and pesticide use: Overview and insights from trends and policies across selected OECD countries. \u003cem\u003eOECD Environment Working Papers\u003c/em\u003e, 155. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1787/63942249-enCabot\u003c/span\u003e\u003cspan address=\"10.1787/63942249-enCabot\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e, M. I., Lado, J., \u0026amp; Sanjuan, N. (2025).\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eSundararajan, M., \u0026amp; Najmi, A. (2020, November). The many Shapley values for model explanation. In International conference on machine learning (pp. 9269\u0026ndash;9278). \u003cem\u003ePMLR\u003c/em\u003e. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://proceedings.mlr.press/v119/sundararajan20b.html\u003c/span\u003e\u003cspan address=\"https://proceedings.mlr.press/v119/sundararajan20b.html\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eThilakarathna, P. S. M., Seo, S., Baduge, K. S. K., Lee, H., Mendis, P., \u0026amp; Foliente, G. (2020). Embodied carbon analysis and benchmarking emissions of high and ultra-high strength concrete using machine learning algorithms. \u003cem\u003eJournal of Cleaner Production\u003c/em\u003e, 262, 121281. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.jclepro.2020.121281\u003c/span\u003e\u003cspan address=\"10.1016/j.jclepro.2020.121281\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eUN (2022). United Nations Department of Economic and Social Affairs, Population Division (2022). World Population Prospects 2022: Summary of Results. UN DESA/POP/2022/TR/NO. 3.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVerge, X. P. C., De Kimpe, C., \u0026amp; Desjardins, R. L. (2007). Agricultural production, greenhouse gas emissions and mitigation potential. \u003cem\u003eAgricultural and forest meteorology\u003c/em\u003e, 142(2\u0026ndash;4), 255\u0026ndash;269. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.agrformet.2006.06.011\u003c/span\u003e\u003cspan address=\"10.1016/j.agrformet.2006.06.011\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eVermeulen, S. J., Campbell, B. M., \u0026amp; Ingram, J. S. (2012). Climate change and food systems. \u003cem\u003eAnnual Review of Environment and Resources\u003c/em\u003e, 37(1), 195\u0026ndash;222. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1146/annurev-environ-020411-130608\u003c/span\u003e\u003cspan address=\"10.1146/annurev-environ-020411-130608\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWernet, G., Bauer, C., Steubing, B., Reinhard, J., Moreno-Ruiz, E., \u0026amp; Weidema, B. (2016). The ecoinvent database version 3 (part I): overview and methodology. \u003cem\u003eThe International Journal of Life Cycle Assessment\u003c/em\u003e, 21, 1218\u0026ndash;1230. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/s11367-016-1087-8\u003c/span\u003e\u003cspan address=\"10.1007/s11367-016-1087-8\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWillett, W., Rockstr\u0026ouml;m, J., Loken, B., Springmann, M., Lang, T., Vermeulen, S., \u0026hellip; Murray, C. J. L. (2019). Food in the Anthropocene: the EAT\u0026ndash;Lancet Commission on healthy diets from sustainable food systems. \u003cem\u003eThe Lancet\u003c/em\u003e, 393(10170), 447\u0026ndash;492. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/S0140-6736(18)31788-4\u003c/span\u003e\u003cspan address=\"10.1016/S0140-6736(18)31788-4\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eWillet, J., Wetser, K., Vreeburg, J., \u0026amp; Rijnaarts, H. H. (2019). Review of methods to assess sustainability of industrial water use. \u003cem\u003eWater Resources and Industry\u003c/em\u003e, 21, 100110.\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eXikai, M., Lixiong, W., Jiwei, L., Xiaoli, Q., \u0026amp; Tongyao, W. (2019). Comparison of regression models for estimation of carbon emissions during building\u0026rsquo;s lifecycle using designing factors: a case study of residential buildings in Tianjin, China. \u003cem\u003eEnergy and Buildings\u003c/em\u003e, 204, 109519. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1016/j.enbuild.2019.109519\u003c/span\u003e\u003cspan address=\"10.1016/j.enbuild.2019.109519\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eYan, M., Cheng, K., Luo, T., \u0026amp; Pan, G. (2014). Carbon footprint of crop production and the significance for greenhouse gas reduction in the agriculture sector of China. \u003cem\u003eIn Assessment of Carbon Footprint in Different Industrial Sectors, Volume 1\u003c/em\u003e (pp. 247\u0026ndash;264). Singapore: Springer Singapore. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1007/978-981-4560-41-2_10\u003c/span\u003e\u003cspan address=\"10.1007/978-981-4560-41-2_10\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZargar, S., Yao, Y., \u0026amp; Tu, Q. (2022). A review of inventory modeling methods for missing data in life cycle assessment. \u003cem\u003eJournal of Industrial Ecology\u003c/em\u003e, 26(5), 1676\u0026ndash;1689. \u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://doi.org/10.1111/jiec.13305\u003c/span\u003e\u003cspan address=\"10.1111/jiec.13305\" targettype=\"DOI\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e\u003c/span\u003e\u003c/li\u003e \u003cli\u003e\u003cspan\u003eZhao, B., Jiang, J., Xu, M., \u0026amp; Tu, Q. (2025). A data-centric investigation on the challenges of machine learning methods for bridging life cycle inventory data gaps. \u003cem\u003eJournal of Industrial Ecology\u003c/em\u003e, \u003cem\u003e29\u003c/em\u003e(3), 955\u0026ndash;966.\u003c/span\u003e\u003c/li\u003e\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"journal-of-industrial-ecology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"44498","submissionUrl":"https://submission.springernature.com/new-submission/44498/3","title":"Journal of Industrial Ecology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"crop production, machine learning, life cycle assessment, biodiversity footprint, climate footprint, predictor importance","lastPublishedDoi":"10.21203/rs.3.rs-9214313/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-9214313/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eCrop production is a major driver of anthropogenic impacts on the environment. Quantifying these impacts requires primary farm-level life cycle inventory data, which are sparse and difficult to collect. Therefore, parsimonious models are needed to predict environmental footprints of crop production. Here, we test the ability of five machine learning modelling techniques (Random Forest (RF), K-Nearest Neighbours (KNN), Artificial Neural Network (ANN), Generalised Boosting Method (GBM), and Linear Modelling (LM)) to predict biodiversity, water and climate footprints of crop production based on limited aggregated farming information, using life cycle data from 121 crops across 57 countries. We found RF to be the most parsimonious model with four predictors for biodiversity (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.88), GBM using five predictors for water (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.89) and ANN for climate footprint (R\u003csup\u003e2\u003c/sup\u003e\u0026thinsp;=\u0026thinsp;0.58) with four predictors. Uncertainty in model predictions is +/- a factor 2.1, 5.1, and 4.0 (95% confidence interval) for the three footprints, respectively. Key farming information to predict biodiversity and climate footprints are yield, fertiliser, electricity and climatic region. Irrigation, fertiliser and pesticide are important predictors of water footprint. Our study offers predictive models highlighting key predictors of environmental footprints of crop production, to be prioritised for data gap filling.\u003c/p\u003e","manuscriptTitle":"Parsimonious machine learning models to estimate environmental footprints of crop production","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2026-04-13 13:31:42","doi":"10.21203/rs.3.rs-9214313/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"editorInvitedReview","content":"","date":"2026-05-18T08:17:27+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"324136794562225580561690254793656493377","date":"2026-04-21T06:27:31+00:00","index":"hide","fulltext":""},{"type":"reviewerAgreed","content":"210686740693189844055864636953302712173","date":"2026-04-08T22:33:47+00:00","index":"hide","fulltext":""},{"type":"reviewersInvited","content":"","date":"2026-04-06T21:11:01+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2026-03-27T16:34:13+00:00","index":"","fulltext":""},{"type":"checksComplete","content":"","date":"2026-03-25T09:21:06+00:00","index":"","fulltext":""},{"type":"submitted","content":"Journal of Industrial Ecology","date":"2026-03-24T15:58:15+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"journal-of-industrial-ecology","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"","sideBox":"","snPcode":"44498","submissionUrl":"https://submission.springernature.com/new-submission/44498/3","title":"Journal of Industrial Ecology","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"stoa","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"08ea34a5-d958-4340-8d07-40894b9f9344","owner":[],"postedDate":"April 13th, 2026","published":true,"recentEditorialEvents":[{"type":"editorInvitedReview","content":"","date":"2026-05-18T08:17:27+00:00","index":55,"fulltext":""}],"rejectedJournal":[],"revision":"","amendment":"","status":"under-review","subjectAreas":[],"tags":[],"updatedAt":"2026-04-13T13:31:42+00:00","versionOfRecord":[],"versionCreatedAt":"2026-04-13 13:31:42","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-9214313","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-9214313","identity":"rs-9214313","version":["v1"]},"buildId":"XKTyCvWXoU3ODBz1xrDgd","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.