Optimizing Sustainability: A Deep Learning Approach on Data Augmentation of Indonesia Palm Oil Products Emission | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Optimizing Sustainability: A Deep Learning Approach on Data Augmentation of Indonesia Palm Oil Products Emission Imam Tahyudin, Ades Tikaningsih, Yaya Suryana, Hanung Adi Nugroho, and 5 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-3675682/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Life Cycle Assessment (LCA) is a widely used methodology for quantifying the environmental impacts of products, including the carbon footprint. However, conducting LCA studies for complex systems, such as the palm oil industry in Indonesia, can be challenging due to limited data availability. This study proposes a novel approach called the Anonymization Through Data Synthesis (ADS-GAN) based on a deep learning approach to augment carbon footprint data for LCA assessments of palm oil products in Indonesia. This approach addresses the data size limitation and enhances the comprehensiveness of carbon footprint assessments. An original dataset comprising information on various palm oil life cycle stages, including plantation operations, milling, refining, transportation, and waste management. The number of original data is 195 obtained from the Sustainable Production Systems and Life Assessment Research Centre of Indonesia's National Innovation Research Agency (BRIN). To measure the performance of prediction accuracy, this study used regression models: Random Forest Regressor (RFR), Gradient Boosting Regressor (GBR), and Adaptive Boosting Regressor (ABR). The best-augmented data size is 1000 data. In addition, the best algorithm is the Random Forest Regressor, resulting in the MAE, MSE, and MSLE values are 0.0031, 6.127072889081567e-05, and 5.838479552074619e-05 respectively. The proposed ADS-GAN offers a valuable tool for LCA practitioners and decision-makers in the palm oil industry to conduct more accurate and comprehensive carbon footprint assessments. By augmenting the dataset, this technique enables a better understanding of the environmental impacts of palm oil products, facilitating informed decision-making and the development of sustainable practices. ADS-GAN Carbon Footprint Data Augmentation Deep Learning Environmental Impacts LCA Regression Palm Oil Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 I. Introduction Indonesia is the world's largest producer of palm oil and accounts for most of the global supply. The industry contributes significantly, ranging from 1.5–2.5% of Indonesia's gross domestic product (GDP) (Nations & Programme, 2023 ). According to data from the Central Statistical Agency (BPS), the total area of palm coconut plantations in Indonesia is about 11.9 million hectares, a figure about three times higher than in 2000, when about four million acres of Indonesian land were used for palm coke plantations. The Indonesian palm coconut industry can also be seen from documented data from the Food and Agriculture Organization (FAO), which provides information that by 2021, the most productive sector in Indonesia will be coca palm, with a total production of about 25.130.562 tons. This figure is much higher when compared to the production of rice products, the primary raw material in Indonesia, with an average production of 21.280.164 tons (Sari et al., 2021 ). The industry has experienced rapid growth and expansion, driven by favorable climate conditions, abundant land resources, and government support. The country's tropical climate provides an ideal environment for palm coconut cultivation, enabling high yields and efficient production. Indonesian palm coconut farming is spread across regions, with significant production areas in Sumatera Island, Kalimantan (Kalimantan), and Papua (Yan, 2022 ). However, the palm coconut industry in Indonesia has challenges. Environmental problems such as deforestation and habitat loss have been linked to the expansion of palm and coconut plantations, raising concerns about the conservation of biodiversity and the destruction of vital ecosystems, including tropical rainforests (Union et al., 2018 ). However, it is essential to note that not all palm and coconut plantations are created equally. Several farms have been established on degraded land or agricultural debris, which can help reduce the pressure on natural forests (Denis J Murphy, 2019). Life Cycle Assessment (LCA) for assessing the environmental impact of palm oil products is essential because of the unique challenges and controversies associated with the oil industry (Siregar et al., 2020 ). LCA considers the entire life cycle, including cultivation, processing, transportation, and disposal of products, thus providing a holistic understanding of environmental consequences. This enables it to thoroughly assess the environmental impacts of palm oil production, such as deforestation, soil degradation, carbon emissions, water pollution, and biodiversity loss. The LCA enables the comparison of carbon footprints in various palm oil production systems, enabling the identification of emission points and the development of strategies to reduce carbon intensity. Palm oil production requires substantial water resources, and improper management can lead to water shortages and pollution (Amalia Yunia Rahmawati, 2020 ). Data processing will support LCA analysis, as LCA is important in the palm coconut industry. Machine learning (ML) is a popular and robust method for analyzing data sets. Unfortunately, the data supporting this case in Indonesia is so limited that the ML performance could be more satisfactory. In this context, additional data sets are needed to improve ML accuracy. Added data integrates primary data, collected through surveys and field measurements, with secondary data from various sources, such as databases, literature, expert knowledge, or data sets created based on theoretical experiments. LCA relies on data availability, and the palm coconut industry often faces data gaps, especially regarding specific environmental impacts. Additional data helps fill this gap by combining primary data collected through field measurements and surveys with secondary data from relevant sources. Added data contributes to the improved quality and reliability of LCA results. Combining primary and secondary data can reduce the potential bias and uncertainty associated with a single data source (Lohr, 2018 ). Primary data collection enables location-specific information, while secondary source data provides a broader context and generalizable information. Expanded data enables incorporating of local knowledge and stakeholders, which can significantly improve the accuracy and relevance of LCA assessments (Nations & Programme, 2023 ). Therefore, expanded data contributes to developing evidence-based policies and strategies that promote sustainable palm oil production and minimize environmental impact. There are several techniques for data augmentation, such as noise injection (Gutowski, 2018 ) and Anonymization Through Data Synthesis (ADS-GAN). ADS-GAN is a deep learning technique used to generate data based on its population. The ADS-GAN network consists of two main parts: a generator and a discriminator. This technology overcomes data constraints and can improve prediction accuracy (Chen et al., 2021 ),(A. Khan et al., 2021 ). Therefore, this study aims to solve the problem of limiting the analysis of palm coconut products with the ADS-GAN technique. In addition, to determine the accuracy of the prediction, it best uses some regression algorithms: Random Forest Regression (RFR), Gradient Boosting Regressions (GBR), and Adaptive Boosting Regression (ABR). II. literature review Previous studies have successfully investigated carbon footprints in palm oil products using machine learning approaches. However, research that applies the latest machine learning techniques, such as the Generative Adversarial Network (GAN), to analyze the carbon footprint of palm petroleum products still needs to be completed. In the health sector, the implementation of synthetic data has successfully overcome the constraints of medical data that often need to be updated. Synthetic data generated through generative hostile tissue (GAN) contributes positively to the performance of artificial intelligence algorithms, as seen in kidney cell carcinoma research. Continuing this, a related study added synthetic data using virtual sensors to diagnose defects on rotating machines with 42 different classes. The results recorded an improvement in training accuracy of about 6–15% and validation precision of about 44–49%. Accordingly, studies on the design of fractional factorial experiments are used to evaluate the impact of basic data augmentation methods on the detection of COVID-19 (Chen et al., 2021 ),(A. Khan et al., 2021 ),(Davila et al., 2023 ) In the area of time series data analysis, research on data augmentation has been the focus of some researchers. In his experiment, the researchers found that the implementation of data augmentation significantly improves the performance of the model. At the same time, other studies have a similar focus on applying data amplification techniques to timelines and integrating them into the context of timeline classification using simulated neural networks (Fawaz et al., 2018 ),(Iwana & Uchida, 2021 ). Later, in the context of computer vision, the introduction of RLR (random local rotation) geometric augmentation techniques was done to manipulate local information on images without adding non-original pixel values. Other studies of data augmentation that can be applied in the computer vision domain include thoroughly studied augmentation strategies, feature-based amplification techniques, and meta-learning-based amplification. However, to determine the transformation or combination of the best transformation, it is often necessary to go through trial-and-error processes (Alomar et al., 2023 ),(Mumuni & Mumuni, 2022 ). Similar research focused on developing a Bisindo translator model that translates sign language signals into text using a machine learning approach based on convolutional neural networks (CNN). As a result, this effort increased the model's performance to an accuracy of 94.38%. Other research highlighted the Convolutional Neural Network Model (CNN) towards data augmentation through the use of technology that can be applied to the classification of respiratory signals, thus enabling focus on relevant data (Fadillah et al., 2021 ),(Rozo et al., 2022 ). Some studies have also applied the data augmentation concept in the context of life cycle analysis (LCA), in which research explores the application of machine learning in life cycle analysis supported by dynamic data research to predict the environmental performance of a product or service using artificial intelligence techniques. Using natural language processing (NLP) and random forest algorithms to train models to provide quick predictions for LCA practitioners and testers in implementing LCA, Similar things have been analyzed to predict the impact of LCA on electricity consumption. Comparing feed-forward (NN) neural networks and repetitive neural (RNN) networks, although limited to one data set (Ghoroghi et al., 2022 ),(Koyamparambath et al., 2022 ),(Portolani et al., 2022 ). The study examined the environmental impact of two palm oil production systems using the ISO standard for life cycle analysis. Although this study improves transparency and reproductivity, it does not consider land-use changes, which can be an essential factor in the reality of palm oil production in some areas. A study evaluates the potential for using artificial intelligence (AI) to improve life cycle impact assessments (LCAs) accuracy and efficiency. It uses a qualitative approach to analyze the potential use of AI in LCAs. Similarly, further research aims to identify challenges and opportunities in conducting life-cycle analysis (LACs) on palm oil production systems. By highlighting the need for spatially explicit methods and the impact of land-use change, this study provides in-depth insights (Heinz Stichnothe, 2011 ),(Oduque de Jesus et al., 2021 ),(Stichnothe & Bessou, 2017 ). In the context of augmented data prediction of CO 2 emission tracks and estimates of coconut oil and palm oil yields, Several researchers have undertaken exploration for life cycle analysis (LCA) of crude palm oil production in the Philippines. This study provides insight into the potential environmental impact of palm oil production. Another study aims to develop a machine-learning model to predict total global CO 2 emissions and understand the impact of the COVID-19 pandemic on future carbon emissions. The palm oil results are also being analyzed, considering weather and soil humidity changes. The evaluation of the performance of various machine learning algorithms achieved high accuracy (95%) compared to traditional methods (Espino et al., 2019 ),(Meng & Noman, 2022 ),(N. Khan et al., 2022 ). In order to improve the performance of systems or learning models, focusing on using synthetic data or data augmentation, a researcher concentrated on utilizing two GAN architectures, GAN and CGAN. The investigation results showed that the synthesized CGAN-based data had reached 63%. Compared to the study, the augmentation of data to improve active learning performance (AL) in the classification showed a significant improvement in AL performance (Aziira et al., 2020 ),(Fonseca & Bacao, 2023 ). III. MATERIAL AND METHOD The research method will be explained through a flowchart to provide an understanding of the research implementation, as depicted in Figure 1. A. Data Description The dataset comes from the Research Centre for Sustainable Production Systems and Life Assessment of the National Innovation Research Agency (BRIN). The amount of original data collected was 195 samples with 19 variables. Table I includes a complete description of the variables included in this process. TABLE I. DESCRIPTION OF THE DATA USED Variable Description Land (Ha) The amount of land in hectares used for oil palm cultivation. Water (m3) The amount of water in cubic meters used for oil palm cultivation and processing. Dolomite A mineral fertilizer used to improve soil pH and provide calcium and magnesium. Urea (Ton) A nitrogen fertilizer used to increase crop yields. Bunch Ash (Ton) A byproduct of oil palm processing that can be used as a fertilizer. Rock Phosphate (Ton) A phosphorus fertilizer used to improve crop yields. Borate (Ton) A boron fertilizer used to improve crop yields and quality. Paraquat (Ton) A herbicide used to control weeds in oil palm plantations. Glyphosate (Ton) A herbicide used to control weeds in oil palm plantations. FFB (Ton) The amount of fresh fruit bunches (FFB) produced in tons. Water for boiler (m3) The amount of water in cubic meters used to generate steam for oil palm processing. Transport (kWh) The cost of transporting FFB to the processing mill and CPO to the market. Genset (kWh) The cost of generating electricity using diesel generators. Electricity (kWh) The amount of electricity in kilowatt-hours used for oil palm processing. CPO (Ton) The amount of crude palm oil (CPO) produced in tons. Productivity (area - FFB) The amount of FFB produced per hectare of land. Productivity (area - CPO) The amount of CPO produced per hectare of land. Gas Emission (CO 2 -eq/ton) The amount of greenhouse gas emissions in carbon dioxide equivalents per ton of CPO produced. Biodiesel (Litres) The amount of biodiesel produced in litres. This research attempts to overcome the limitations of original data with data synthesis techniques. Using synthetic data in this context can add value to an effort to understand and reduce the carbon impact of palm oil products in Indonesia. B. Synthetic Data Augmentation Large and representative datasets play a crucial role in guaranteeing the accuracy and precision of model predictions. To overcome data constraints using synthesis methods. Synthetic data can address data shortages, improve model performance, and improve data collection efficiency (Lamberti, 2023). One of the latest methods that got severe attention is generative adversarial networks (GANs). Generative Adversarial Networks (GANs) are a type of deep learning model that can generate synthetic data that mimics accurate data. The learning process in the GAN involves one generator and one discriminatory nerve network that plays the minimax zero-sum game. Generators create synthetic data, while discriminators evaluate whether the data is natural or synthetic. The generator improves its ability to create realistic synthetic data through this hostile process. In the end, if everything goes well, the generator can generate synthetic data that appears natural but is difficult for discriminators (or humans) to distinguish as accurate or synthesized (Huang, 2021). The research investigates the potential and limitations of three methods, namely Anonymization Through Data Synthesis (ADS-GAN), Private Aggregation of Teacher Ensembles (PATE-GAN) and Differentially Private (DP-GAN). Involving 195 samples of original data, the synthesis process is performed randomly to produce additional amounts of data that vary: 1000, 5000, and 10000. All variables in the data set are considered sensitive, except for the primary target variable, 'Gas Emissions (CO 2 -eq/tonne)'. The choice of these three methods is made with consideration of the superiority of each. ADSGAN was chosen because of its ability to generate synthetic data that matches the characteristics of the original data. In contrast, PATEGAN and DPGAN were selected for their focus on differential privacy concepts to protect the confidentiality of individual data, placing particular emphasis on privacy security (Rastogi et al., 2023),(Fan, n.d.). The resulting synthetic data will be analyzed using three regression models: Random Forest Regressor (RFR), Gradient Boosting Regresor (GBR), and Adaptive Boosting regressor (ABR), to evaluate the synthesis quality of the three methods. C. Data Synthesis Process Across Multiple GANs The synthesis data process uses Generative Adversarial Networks (GANs) using synthcity libraries. Synthcity serves as a library that captures the entire workflow in producing and evaluating synthetic data. This library provides synthesis data plugins such as ADS-GAN, PATE-GAN, and DP-GAN (Bedorf, 2023). Using synthcity, the data synthesis phase can be run more efficiently. This library provides the ability to manage workflows comprehensively, including implementing various types of GANs such as Madigan and pagan. Each plugin has a unique role in generating synthetic data that matches the research needs. The synthcity library is the foundation that records the entire process of producing and evaluating synthetic data. A typical workflow using Synthcity involves a set of detailed steps. Here are the general steps found in the synth city workflow (Qian, n.d.): 1) Loading the dataset using a DataLoader A class called DataLoader offers a standardized interface for loading and storing many kinds of input data, including survival, time series, and tabular data. It was employed in this work to import the original data, which made up 195 data points from the sample and served as the foundation for the synthesis process. This data collection includes several characteristics about the production of palm oil and its effects on the environment. Still, it focuses on the critical variable "gas emissions (CO 2 -eq/tonne)." 2) Training the Generator using a Plugin Through the Plugin class in Synthcity, users can design, train, and use a range of data generators. Every plugin uses a different algorithm to produce data. The fit() method offered by the plugin is used to train the generator. The plugins employed in the context of this study are "pategan", "dpgan", and "adsgan". 3) Generating Synthetic Data Users can create fake data by using the generate() method once the plugin has been trained. Conditional generation is also possible with some plugins. 4) Evaluating Synthetic Data A wide range of measures are available from Synthcity to assess the integrity, use, and privacy of synthetic data. Users can do assessments using the Metrics class. D. Regression Algorithm 1) Adaptive Boosting regressor Adaboost is a supervised machine-learning algorithm that can solve classification and regression problems. The advantage of this algorithm is that it handles complex data and feature interactions. It can handle complex data and feature interactions, prevent overfitting, and is relatively easy to use (Jyotsna Vadakkanmarveettil, 2021). The drawback is that it is sensitive to specific parameters because the Adaboost algorithm is sensitive to data distribution and may not work well if the data has specific characteristics (Hornyák & Iantovics, 2023). 2) Random Forest Regressor RF (Random Forest) is an ML algorithm that uses several decision trees to make predictions. The advantage of this algorithm is its ability to handle complex or unstructured data and provide accurate results in making predictions. However, the weakness of this algorithm is the possibility of overfitting, namely the prediction model that is too fit with the training data so that it cannot be used for test data (Davila et al., 2023). 3) Gradient Boosting Regressor Gradient Boosting is an ML algorithm that uses several weak predictive models (weak learners) to produce solid predictive models. The advantage of this algorithm is its ability to handle complex or unstructured data and provide accurate results when making predictions. However, the weakness of this algorithm is the computation time, which is quite long and sensitive to the parameters used (Fadillah et al., 2021). E. Performance Analysis Analysis of the performance of machine learning algorithms in regression models involves matrix calculations as an evaluation method. This process assesses how much the model can predict target values (dependent variables) based on input variables (independent variables). There are several standard evaluation metrics used for regression models, and most of them can be calculated using a payoff matrix using the following metric values (Silitonga et al., 2020)(Chicco et al., 2021)(Jadon et al., 2022): 1) Mean Absolut Error (MAE) MAE is used to calculate the error value between the predicted value obtained from the model results and the actual value, and then absolute it. 2) Mean Square Error (MSE) MSE measures the average squared difference between the model's predicted values and the actual values of the observed data 3) Mean Squared Logarithmic Error (MSLE) MSLE is a variant of Mean Squared Error (MSE), which uses the logarithm of the predicted value and the actual value. MSLE measures the average of the squared logarithmic differences between predicted values and actual values. The best model is the model that obtains the lesser value of MAE, MSE, and MSLE. IV. RESULT and DISCUSSION The analysis of the original palm and coconut data provides an initial understanding of the data collected. The primary dataset briefly describes the target variable, which can be seen in Figure 2. Through Figure 2, a statistical histogram visualization of descriptive emission variables provides in-depth insight into data distribution on data sets. The amount of data, or count, can be estimated through the number of bars on the histogram. The histogram's average gas emission indicated by the central distribution position is about 3.79 CO 2 -eq/ton. At the same time, the standard deviation of about 2.39 shows a significant spread of the average value. Min and Max, which represent the minimum and maximum values of the distribution, are seen on the x-axis of the histogram, with the smallest values reaching 0.09 CO 2 -eq/ton and the most outstanding value reaching 25.48 CO 2 /ton. In addition, a median value of around 2.39 CO 2 per ton can be found in the middle of such a distribution. Histograms, as a form of visualization, provide a more intuitive understanding of data distribution and reinforce descriptive statistical findings related to emission variables in the palm coconut industry. A. Downstream Task Performances With Synthetic Datasets In this work, data synthesis was carried out using the Generative Adversarial Networks (GAN) model. The experiments were conducted with variations of data amounts, i.e., 1000, 5000, and 10000, to evaluate the model's performance. The models assessed included ADS-GAN, PATE-GAN, and DP-GAN. The evaluation used three main metrics: Jensen-Shannon Divergence (JSD), Wasserstein distance, and Kolmogorov-Smirnov Test (KS Test). The use of these metrics was aimed at measuring to what extent the distribution of synthetic data generated by the GAN model was similar to that of the original data distribution. Jensen-Shannon Divergence provides information about the differences in probability distribution between two datasets, while Wasserstein distance estimates how much change is needed to transform one distribution into another. The Kolmogorov-Smirnov Test (KS Test), on the other hand, is used to test the extent to which two probability distributions are simultaneous (Yoon et al., 2020),(Schiappa, 2019). TABLE II COMPARISON OF SYNTHETIC DATA GENERATION IN TERMS OF JSD, WASSERSTEIN DISTANCE AND KS TEST ADS-GAN PATE-GAN DP-GAN 1000 Wasserstein distance 0.082203 1.568033 0.203662 JSD 0.018408 0.018411 0.042283 KS Test 0.628880 0.175439 0.641026 5000 Wasserstein distance 0.504504 0.723084 1.936032 JSD 0.022357 0.036514 0.058279 KS Test 0.336032 0.402159 0.191633 10000 Wasserstein distance 0.469520 0.736796 1.975019 JSD 0.020051 0.036710 0.059375 KS Test 0.372470 0.388664 0.174089 Based on Table II, we compared the performance of three synthetic data generation approaches, namely ADS-GAN, PATE-GAN, and DP-GAN, using Wasserstein Distance, Jensen-Shannon Divergence (JSD), and KS Test evaluation metrics in three different scenarios: 1000, 5000, and 10000 synthetical data. Then, in the case of 1000 synthetic data, ADS-GAN showed superior performance with a low Wasserstein Distance (0.082203) compared to PATE-GAN (1.568033) and DP-GAN (0.203662). Similarly, in JSD, the ADS -GAN (0.018408) had a meager value, showing high similarities with the original data distribution. KS Test on ADS –GAN (0.628880) also showed promising results in measuring distribution equality. Meanwhile, ADS-GAN performed well in the 5000 synthetic data scenario with a stable low Wasserstein Distance (0.504504). PATE-GAN also showed promising results, while DP-GAN had higher values on the metric. Nevertheless, ADS-GAN remained superior with low JSD (0.022357) and adequate KS Test (0.336032). Finally, on 10000 synthetic data, ADs-GAN and PATE-GAN again showed good performance with relatively low Wasserstein Distance. ADS -GAN (0.469520) had low J SD (0.020051) and satisfactory KS Test (0.372470). PATEGAN also showed a good result, while DP-GAN showed higher values on both metrics. Overall, ADS-GAN shows consistent performance with the ability to generate synthetic data close to the original data distribution, compared to PATE-GAN and DP-GAN. In each test scenario (1000, 5000, and 10,000 synthetical data), ADS -GAN consistently stands out with superior performance in evaluation metrics such as Wasserstein Distance, Jensen-Shannon Divergence (JSD), and KS Test. Specifically, in scenarios with 1000 synthetic data, the ADS-GAN implementation performed better than 5000 and 10000 systems. This is mainly seen from the Wasserstein and JSD Distance metric evaluations. In this context, the low Wasserstein Distance value and the low JSD value in 1000 scenarios indicate that the distribution of synthetical data generated by the ADS-GAN is closer to the original distribution than the scenario with more significant amounts of synthesized data. Thus, this conclusion affirms that ADS-GAN can deliver optimal performance, especially in scenarios with more limited quantities of synthetic data. This is an important consideration, especially when computational efficiency or limitation of synthesized data is a significant factor. B. Performance of Regression Models on Synthetic Data Regression is a supervised machine learning model directed at predicting continuous values. A regression model describes the relationship between a dependent variable and one or more independent variables. This research uses regression analysis to predict gas emissions from palm oil products. Building this model involves comparing two sets of data, namely 1000 samples of synthetic palm oil data, as a first step before comparing model performance using accurate data. The selection of 1000 synthetic data samples as a basis for comparison was based on the finding that this number provided optimal model performance, as seen from the evaluation using the Wasserstein Distance, Jensen-Shannon Divergence (JSD), and KS Test metrics documented in Table II. Then, the algorithms used include Random Forest Regression (RFR), Gradient Boosting Regression (GBR), and Adaptive Boosting Regression (ABR). They were chosen because of their advantages in handling data complexity and their ability to capture non-linear relationships between variables. Meanwhile, the dependent variable in the context of this regression model is "Gas Emissions (CO 2 -eq/ton)," which refers to the amount of greenhouse gases emitted per metric ton of coconut palm produced or processed. Meanwhile, the independent variable involves 19 features that can affect emission levels. In building a regression model, dividing data into two subgroups is crucial to training and testing the model's performance. Training sets are used to train models, which means models are "learned" from patterns and relationships in the data. Instead, a testing set is used to test the performance of models on data that have never been seen before (Kumar, 2020). This study applies the separation of data with an 80/20 ratio. The ratio has a good chance of understanding and adapting to data variations while still providing a good test on data that has never been seen before. In the following explanation, we will present an analysis of the results of the Gradient Boosting Regression (GBR), Random Forest (RF), and AdaBoost models using 1000 synthetic data samples. This analysis will provide a deeper understanding of how well these three models can provide predictions of carbon emissions on palm oil data, illustrating the strengths and weaknesses of each model in a more specific applicable context. 1) Gradient Boosting Regression (GBR) The Gradient Boosting Regressor (GBR) model has shown excellent performance in predicting palm coconut carbon emissions. The GBR has an MAE rate of around 0.00341. In addition, the mean squared error (MSE) of around 6.671114421152337e-05 shows the ability of the model to reduce predictive errors, and the MSE is more sensitive to significant errors because it squares the difference between the prediction and the actual value. A low MSE value shows the model's ability to reduce prediction errors effectively. The GBR model also has an MSLE of around 6.323034859557045e-05, which is useful when measuring errors on a logarithmic scale. Low MSLE values indicate that the model has a minimal error rate on the logarithm scale. Figure 3 illustrates a variable that plays a vital role in predicting emissions of palm oil products using the Gradient Boosting Regressor algorithm. The top three variables are 'Genset,' 'Transport,' and 'Urea.' As is well known, the Genset is the primary source of electricity in many palm and coconut plantations (Yusniati et al., 2018), its use involves fossil fuels such as solar or diesel, which contribute to greenhouse gas emissions. The more the generator is used, the higher the greenhouse gas emissions generated (Evers, 2020). Transportation also plays a vital role in the supply chain of the palm coconut industry (Cheah et al., 2023), including the transportation of crops, fertilizers, pesticides, and labor. If transportation uses conventional vehicles with fossil fuels, this will impact CO 2 emissions. Furthermore, using urea as fertilizer requires energy and can produce greenhouse gas emissions, mainly if used excessively, which produces nitrogen oxide (NOx) (Dong et al., 2022). 2) Random Forest Regression In analyzing the performance of predicting emissions of palm oil products using the Random Forest Regression algorithm, the model shows excellent results. The mean absolute error MAE has a value of around 0.003132978845438705, indicating high prediction accuracy, while the MSE of around 6.127072889081567e-05 shows the ability of the model to reduce prediction errors. In addition, the MSLE, which is only about 5.838479552074619e-05, describes a low error rate in the logarithm scale. With these results, the Random Forest Regression model proved to have excellent performance in predicting emissions of palm oil products, with high accuracy, a solid ability to reduce predictive errors, and minimal errors in the logarithmic scale. Predicting palm oil emissions using Random Forest Regressor algorithms yielded essential values, as shown in Figure 4. The feature importance score in the Random Forest Regression algorithm measures the extent to which each variable contributes to the accuracy of the model prediction. As shown in Figure 4, Random Forest Regression has the same result as gradient boosting regression. Features such as 'Genset,' 'Transport,' and 'Urea' found in Random Forest Regression can be regarded as variables that significantly influence the outcome of emission predictions. This means these variables have a decisive role in influencing the emission rate of palm oil products. This understanding can help further decision-making and modeling to reduce or control emissions. 3) Adaptive Boosting Regression This model shows quite good results in the performance analysis of the palm oil product emission prediction using the Adaptive Boosting Regressor algorithm. The MAE has a value of around 0.012349103068078711. A low MAE indicates a reasonable degree of accuracy in the prediction. MSE around 0.00019333446314396462 describes the ability of a model to reduce predictive error, and a low MSE indicates that the model has a minimum error rate. Meanwhile, The MSLE has a value of about 0.00018800274586211268, which indicates a low logarithm scale error rate. Based on these evaluation metrics, the adaptive boosting regression model produces accurate predictions, has a solid ability to reduce predictive errors, and has minimal errors in the logarithmic scale. Thus, the adaptive boosting regression algorithm also shows good performance in predicting emissions of palm oil products, which can be used as a basis for better managing these emissions. Predicting palm oil emissions using the adaptive boosting regression algorithm yielded essential values, as shown in Figure 5. As shown in Figure 5, the top three critical features when using the adaptive boosting regression algorithm are the variables waterfor_boiler(m3), FFB (ton), and Urea. Using water for boilers (m3) in the palm coconut industry can impact carbon emissions. The palm coconut plant has a fuel consumption of 0.37 liters of diesel per ton of fresh fruit (TBS), equivalent to 5.0 kilograms of CO (Hong, 2023). In addition, FFB (ton) in the palm coke industry can also contribute to carbon emissions, with each ton of TBS requiring an allocation of 2 kilometers of transportation (Saswattecha et al., 2015). Urea use as a fertilizer requires energy and can produce greenhouse gas emissions, mainly if used excessively, which produces nitrogen oxide (NOx) (Dong et al., 2022). C. Comparison of Regression Model Performance on Original Data and Augmented Data This section presents the evaluation results of three regression models: gradient boosting, random forest, and adaptive boosting. The results of this evaluation are presented for two different situations: first, when the models are applied to a dataset of 195 samples (original data); second, after the data synthesis process using the ADS-GAN method, where the resulting dataset consists of 1000 samples. Table III and Table IV display the model evaluation based on the calculation of the primary metrics, namely Mean Absolute Error (MAE), Mean Squared Error (MSE), and Mean Squared Logarithmic Error (MSLE) for both data conditions. TABLE III. ORIGINAL DATA EVALUATION Model MAE MSE MSLE GBR 0.00413079449910069 9.725835958226041e-05 3.0886335234002594e-05 RF 0.004699656505053924 0.0002242838981326051 6.50071636956904e-05 ABR 0.013249024047795621 0.00038966505062171193 0.00018540864271605393 TABLE IV. ADS-GAN AUGMENTED DATA EVALUATION Model MAE MSE MSLE GBR 0.0034182220980073807 6.671114421152337e-05 6.323034859557045e-05 RF 0.003132978845438705 6.127072889081567e-05 5.838479552074619e-05 ABR 0.012349103068078711 0.00019333446314396462 0.00018800274586211268 Here is a comparative explanation of the performance of three algorithm models, namely the Gradient Boosting Regressor (GBR), the Random Forest Regressor (RFR), and the Adaptive Boosting Regressor (ABR), in predicting palm coconut emissions. The comparison was done using two datasets: original data with 195 samples and data synthesized with 1000 samples, documented in Tables III and IV. Performance evaluation was done considering four main metrics: mean absolute error (MAE), mean square error, mean logarithmic error and squared logarithmic error. Models with original palm coconut data from Table III show that the Gradient Boosting Regressor (GBR) algorithm has the lowest MAE value of around 0.0041, compared to Random Forest Regressors (RFR) and Adaptive Boosting Regressors (ABR). This value indicates that the GBR prediction has a low and accurate error rate. In addition, GBR also has the smallest of around 0,0002. In the MSLE, adaptive boosting regressions (ABR) have the lowest value of about 0,0001. These results show that GBR is the best choice for predicting palm coconut emissions compared to RFR and ABR. Models with the synthesized data from Table IV show that Random Forest (RF) has the lowest MAE value, around 0.0031, compared to Random Forest Regressor (RFR) and Adaptive Boosting Regressor (ABR). RF also has the smallest MSE value of around 0.00006127072 and the lower value MSLE of about 0.000005838479552074619. The results confirmed that the synthetic palm data delivered the best performance on the RF algorithm model in predicting coconut emissions compared to the other two, the Random Forest Regressor and the AdaBoost Regressor. Comparisons between the performance of models using original and synthesized data, as documented in Tables III and IV, reveal significant differences. The results of the analysis showed that the Random Forest algorithm on the synthesis data showed performance improvements, with mean absolute error (MAE) values of 0.003132978845438705, mean square error (MSE) around 6.127072889081567e-05, and mean squared logarithmic error (MSLE) around 5.838479552074619e-05. It should be emphasized that the exponents in the MSE and MSLE values indicate that the prediction error on the synthetic data is deficient, close to zero. Further, looking at the performance of the Gradient Boosting Regressor (GBR) model from the synthesis data, it was seen that the model succeeded in reducing the MAE and MSE values to 0.0034182220980073807 and 6.671114421152337e-05, respectively. This comparison describes an increase in the accuracy of GBR predictions on synthesis data compared to models using original data. The exponents on MSE values again show that the error rate generated by the GBR model of the synthesized data is minimal. At the same time, the adaptive boosting regression (ABR) model performance of the synthesis data also showed a significant decrease in MAE and MSE values. MAE values were around 0.012349103068078711, while MSE reached 0.00019333446314396462. Based on a comparison of model performance results on accurate and synthetic data, it can be concluded that the best algorithm for synthetical data is Random Forest. This is evident from the significant increase in prediction accuracy, with deficient mean absolute error (MAE) values close to zero, as well as mean square error (MSE) and mean squared logarithmic error (MSLE) that indicate minimal levels of prediction. Error on synthetic data. The Random Forest Regression (RFR) model used to predict carbon emissions in synthetic datasets can be seen through the distribution of residues in the residual plot Fig. 6. Residues are the difference between actual and predictive values generated by the regression model. The average residual distribution, i.e., distributed evenly around zero values, is an essential indication in evaluating the model's conformity to actual data. The setting of model parameters is a critical factor in achieving optimal performance. By setting n_estimators=700, the RFR model is made with 700 decision trees, providing predictive stability and preventing potential overfitting. The parameter max_depth=50 is used to control the maximum depth of each tree, while min_samples_split=5 and min_ samples_leaf=5 determine the minimum number of samples required to divide the internal node and form the leaf node in the decision tree. Random_state=42 controls randomness in model development, ensuring consistency of results every time a model is run. The residual distribution analysis in Figure 6 gives a positive picture of the performance of the Random Forest Regressor model. The general pattern indicates that the overall residues are evenly spread around the zero line, showing the consistency of the model in predicting the value. Although some points are away from zero, most of the residue is close to 0, indicating a model's ability to predict accurately. Furthermore, separating residues into positive and negative categories opens up additional insights. Damaging residue tends to have a smaller magnitude compared to positive residue. This suggests that models are more accurate in predicting smaller values while may encounter difficulties forecasting larger ones. Giving background colors to the positive ("lightgreen") and negative ("lightcoral") residues helps visualize where models tend to perform over-predictions and under-prediction. (residue negative). D. Prediction Model The CO 2 emission model prediction results are presented as a scatter plot using Microsoft Excel to illustrate the predictive, descriptive analysis visually. The scatter plot, seen in Figure 7, shows the distribution of points between the actual value and the predicted value. In descriptive analysis, these scatter plot patterns carry valuable information regarding model performance. The gathering of many points around the point (0,0) indicates that the model can predict CO 2 emissions well in cases with low values. However, scattered dots around the plot indicate variations in model performance, where some predictions approximate the actual values well, while others may have more significant errors. Detection of points far from the diagonal line (0,0) can provide insight into outliers or cases where the model has significant prediction errors. By understanding these visual characteristics, the analysis provides A solid basis for further decision-making regarding model performance, Identifying specific patterns or trends in its predictions, and Highlighting points that require further investigation. Figure 8 is a column diagram that illustrates the comparison of average CO 2 emissions between actual and predicted values throughout the Indonesia. In this context, the average actual CO 2 emissions are stated to be 0.05, while the average estimated CO 2 emissions from the model is 0.04. This comparison shows that the average predicted value of CO 2 emissions is slightly lower than that of actual CO 2 emissions. In other words, the model predicts CO 2 emissions at slightly lower values than they do. This phenomenon can be interpreted as a tendency for underprediction, where the model tends to estimate values that are less than the actual value of CO 2 emissions. By understanding these differences, further interpretation can be made to investigate specific factors or features that may influence model performance so that the model can be improved to provide more accurate predictions. Overall, the analysis results show that the use of the synthesized data can significantly improve the model's performance, especially in the context of palm coconut emission predictions. This success is reflected in smaller critical evaluation values such as MAE, MSE, and MSLE on models with synthetic data compared to original ones. These smaller error values indicate that models with synthesis data can better estimate actual values with a lower error rate. This performance improvement can be explained by the ability of the model to recognize patterns and relationships in more complex data, as is the case in data synthesis. Thus, models tend to be more precise and responsive to variations in synthetic datasets, resulting in more accurate predictions related to palm coconut emissions. E. Conclusion and Future Work The research addressed the constraints of limited data sources by applying data augmentation techniques. The researchers increased the number of samples from 195 to 1,000. Its primary purpose is to support predictive analysis of palm oil carbon footprints. To ensure the success of the data augmentation process, a comparison of the performance of the three machine learning approaches, namely the Random Forest Regressor, the Gradient Enhanced Regresor, and the Adaptive Enhanced Regresior, was carried out. This study used the ADS-GAN, PATE-GAN, and DP-GAN methods to produce synthetic data sets. The study results showed that using synthesized data with ADs-GAN yielded the best performance, significantly affecting the performance of regression models with excellent improvements in synthetical data. This analysis confirms that the synthesis data significantly improves model performance in predicting palm coconut emissions. Models using Random Forest synthetic data (RF) showed lower error values than the original models. On the synthesis data, RF produced MAE, MSE, and MSLE values of 0.003132978845438705, 6.127072889081567e-05, and 5.838479552074619e-05. These results showed that using synthesized data can significantly improve the accuracy of palm oil emission predictions. Furthermore, future tasks involve collecting more original data sets and exploring new techniques to improve palm oil carbon footprint predictive analysis quality continuously. Abbreviations LCA Life Cycle Assessment GANs Generative Adversial Networks ADS-GAN Anonymization Through Data Synthesis PATE-GAN Private Aggregation of Teacher Ensembles DP-GAN Differentially Private GBR Gradient Boosting Regression ABR Adaptive Boosting Regression RFR Random Forest Regression MAE Mean Absolute Error MSE Mean Squared Error MSLE Mean Squared Log Error ML Machine Learning JSD Jensen-Shannon Divergence KS Test Kolmogorov-Smirnov Test. Declarations Authors’ information Imam Tahyudin: His Doctoral program obtained since 2018 from Kanazawa University, Japan. Since 2023, he has been an associate professor (Computer) at Universitas Amikom Purwokerto. He is the author of a number of journal articles and conference papers. His research interests include machine learning, big data, data science, knowledge discovery, and decision support systems. Hanung Adi Nugroho: His research interest includes signal and image processing and analysis, computer vision, medical imaging, medical instrumentation, and statistical pattern recognition. His Doctoral program obtained from The University of Queensland, Australia (2005). Agus Bejo: is a computer scientist specializing in VLSI, Design of processor and compiler, embedded system, biometric authentication, smartcard application based, IoT. He obtained the doctoral program from Chulalongkorn University, Thailand and the Department of Communications and Integrated Systems, Tokyo Institute of Technology, Japan in 2009 and 2014 respectively. Ade Tikaningsih: She is a student from Informatic program in Universitas Amikom Purwokerto. Her research passion in machine learning, data mining, and decision support system. Ade Nurhopipan: Her Master program graduated from Gadjah Mada University. Her research interest includes machine learning, image processing, NLP, and decision support system. Puji Lestari: She is a student from Information System program in Universitas Amikom Purwokerto. Her research passion in machine learning, data mining, and decision support system, web programming. Yaya Suryana: His research interest includes data mining, machine learning, AI, and decision support system. His doctoral program from Tsukuba University, Japan. He is a research group chairman of Applied Instrumentation Systems, Acquisition & Big Data and Continuous Measurement of BRIN. Nugroho Adi Sasongko: His research interest includes sustainable development, renewable energy technologies, life-cycle assessment, and biofuel production. He is a Director of Center for Sustainable Production Systems Research and Life Cycle Assessment, National Research and Innovation Agency (BRIN). He obtained the doctoral program from Tsukuba University, Japan. Ahmad Ismed Yanuar: He is a member of research group of Carbon Cycle Management and Standardization of Life Cycle Assessment of BRIN. His research focus for Life Cycle Assessment (LCA) for palm oil product. Availability of data and materials Not applicable. Computing interests The authors declare that they have no competing interests. Funding This work was not funded. Author’s contributions IT, AT, YS conceived of the presented idea, developed the theory, implemented the research, and wrote the manuscript. IT, AN, AT, PL contributed to the design, and to the analysis of the results. All authors reviewed and approved the final manuscript. Acknowledgements This research is supported by Research and Innovation Agency (BRIN) of Indonesia Republic from Visting Research Program. Furthermore, it is supported by Study Program of Engineering profession Program (PSPPI), Universitas Gadjah Mada. The authors would like to thank LPPM Universitas Amikom Purwokerto for supporting this research. References Alomar, K., Aysel, H. I., & Cai, X. (2023). Data Augmentation in Classification and Segmentation: A Survey and New Strategies. Journal of Imaging , 9 (2). https://doi.org/10.3390/jimaging9020046 Amalia Yunia Rahmawati. (2020). Life Cycle Assessment of Palm Oil at United Plantations Berhad 2022 . July , 1–23. Aziira, A. H., Setiawan, N. A., & Soesanti, I. (2020). Generation of Synthetic Continuous Numerical Data Using Generative Adversarial Networks. Journal of Physics: Conference Series , 1577 (1). https://doi.org/10.1088/1742-6596/1577/1/012027 Bedorf,A.(2023). Synthcity .Ccaim.Cam.Ac.Uk.https://ccaim.cam.ac.uk/synthcity/ Cheah, W. Y., Siti-Dina, R. P., Leng, S. T. K., Er, A. C., & Show, P. L. (2023). Circular bioeconomy in palm oil industry: Current practices and future perspectives. Environmental Technology and Innovation , 30 . https://doi.org/10.1016/j.eti.2023.103050 Chen, R. J., Lu, M. Y., Chen, T. Y., Williamson, D. F. K., & Mahmood, F. (2021). Synthetic data in machine learning for medicine and healthcare. Nature Biomedical Engineering , 5 (6), 493–497. https://doi.org/10.1038/s41551-021-00751-8 Chicco, D., Warrens, M. J., & Jurman, G. (2021). The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. PeerJ Computer Science , 7 , 1–24. https://doi.org/10.7717/PEERJ-CS.623 Davila, M. H., Baldeon-Calisto, M., Murillo, J. J., Puente-Mejia, B., Navarrete, D., Riofrío, D., Peréz, N., Benítez, D. S., & Moyano, R. F. (2023). Analyzing the Effect of Basic Data Augmentation for COVID-19 Detection through a Fractional Factorial Experimental Design. Emerging Science Journal , 7 (Special Issue), 1–16. https://doi.org/10.28991/ESJ-2023-SPER-01 Denis J Murphy. (2019). Growing palm oil on former farmland cuts deforestation, CO₂ and biodiversity loss . Theconversation.Com. https://theconversation.com/growing-palm-oil-on-former-farmland-cuts-deforestation-co-and-biodiversity-loss-127312 Dong, D., Yang, W., Sun, H., Kong, S., & Xu, H. (2022). Effects of Split Application of Urea on Greenhouse Gas and Ammonia Emissions From a Rainfed Maize Field in Northeast China. Frontiers in Environmental Science , 9 (January), 1–11. https://doi.org/10.3389/fenvs.2021.798383 Espino, M. T. M., De Ramos, R. M. Q., & Bellotindos, L. M. (2019). Life cycle assessment of the oil palm production in the Philippines: A cradle to gate approach. Nature Environment and Pollution Technology , 18 (3), 709–718. Evers, S. (2020). Palm oil: research shows that new plantations produce double the emissions of mature ones . Theconversation.Com. https://theconversation.com/palm-oil-research-shows-that-new-plantations-produce-double-the-emissions-of-mature-ones-130330 Fadillah, R. Z., Irawan, A., Susanty, M., & Artikel, I. (2021). Data Augmentasi Untuk Mengatasi Keterbatasan Data Pada Model Penerjemah Bahasa Isyarat Indonesia (BISINDO). Jurnal Informatika , 8 (2), 208–214. https://ejournal.bsi.ac.id/ejurnal/index.php/ji/article/view/10768 Fan, L. (n.d.). A Survey of Differentially Private Generative Adversarial Networks . Fawaz, H. I., Forestier, G., Weber, J., Idoumghar, L., & Muller, P.-A. (2018). Data augmentation using synthetic data for time series classification with deep residual networks . http://arxiv.org/abs/1808.02455 Fonseca, J., & Bacao, F. (2023). Improving Active Learning Performance through the Use of Data Augmentation. International Journal of Intelligent Systems , 2023 , 1–17. https://doi.org/10.1155/2023/7941878 Ghoroghi, A., Rezgui, Y., Petri, I., & Beach, T. (2022). Advances in application of machine learning to life cycle assessment: a literature review. International Journal of Life Cycle Assessment , 27 (3), 433–456. https://doi.org/10.1007/s11367-022-02030-3 Gutowski, T. G. (2018). A Critique of Life Cycle Assessment ; Where Are the People ? Procedia CIRP , 69 (May), 11–15. https://doi.org/10.1016/j.procir.2018.01.002 Heinz Stichnothe, F. S. (2011). LCA of two palm oil production systems. Biomass and Bioenergy , 35 (9). https://doi.org/https://doi.org/10.1016/j.biombioe.2011.06.001 Hong, W. O. (2023). Review on Carbon Footprint of the Palm Oil Industry: Insights into Recent Developments. International Journal of Sustainable Development and Planning , 18 (2), 447–455. https://doi.org/10.18280/ijsdp.180213 Hornyák, O., & Iantovics, L. B. (2023). AdaBoost Algorithm Could Lead to Weak Results for Data with Certain Characteristics. Mathematics , 11 (8). https://doi.org/10.3390/math11081801 Huang, D. (2021). Synthetic data generation using Generative Adversarial Networks (GANs) . Medium.Com. https://medium.com/data-science-at-microsoft/synthetic-data-generation-using-generative-adversarial-networks-gans-part-1-47ecbf46b575 Iwana, B. K., & Uchida, S. (2021). An empirical survey of data augmentation for time series classification with neural networks. In PLoS ONE (Vol. 16, Issue 7 July). https://doi.org/10.1371/journal.pone.0254841 Jadon, A., Patil, A., & Jadon, S. (2022). A Comprehensive Survey of Regression Based Loss Functions for Time Series Forecasting . http://arxiv.org/abs/2211.02989 Jyotsna Vadakkanmarveettil. (2021). AdaBoost – An Easy Guide (2021) . U-next.Com. https://u-next.com/blogs/data-science/adaboost/ Khan, A., Hwang, H., & Kim, H. S. (2021). Synthetic data augmentation and deep learning for the fault diagnosis of rotating machines. Mathematics , 9 (18). https://doi.org/10.3390/math9182336 Khan, N., Kamaruddin, M. A., Ullah Sheikh, U., Zawawi, M. H., Yusup, Y., Bakht, M. P., & Mohamed Noor, N. (2022). Prediction of Oil Palm Yield Using Machine Learning in the Perspective of Fluctuating Weather and Soil Moisture Conditions: Evaluation of a Generic Workflow. Plants , 11 (13). https://doi.org/10.3390/plants11131697 Koyamparambath, A., Adibi, N., Szablewski, C., Adibi, S. A., & Sonnemann, G. (2022). Implementing Artificial Intelligence Techniques to Predict Environmental Impacts: Case of Construction Products. Sustainability (Switzerland) , 14 (6), 1–12. https://doi.org/10.3390/su14063699 Kumar, S. (2020). Data splitting technique to fit any Machine Learning Model . Towardsdatascience.Com. https://towardsdatascience.com/data-splitting-technique-to-fit-any-machine-learning-model-c0d7f3f1c790 Lamberti, A. (2023). The benefits and limitations of generating synthetic data . Syntheticus.Ai. https://syntheticus.ai/blog/the-benefits-and-limitations-of-generating-synthetic-data Lohr, S. L. (2018). Measuring Uncertainty with Multiple Sources of Data . Meng, Y., & Noman, H. (2022). Predicting CO2 Emission Footprint Using AI through Machine Learning. Atmosphere , 13 (11), 1–15. https://doi.org/10.3390/atmos13111871 Mumuni, A., & Mumuni, F. (2022). Data augmentation: A comprehensive survey of modern approaches. Array , 16 (November), 100258. https://doi.org/10.1016/j.array.2022.100258 Nations, U., & Programme, D. (2023). Indonesia: Sustainable Palm Oil . Www.Undp.Org. https://www.undp.org/facs/indonesia-sustainable-palm-oil Oduque de Jesus, J., Oliveira-Esquerre, K., & Lima Medeiros, D. (2021). Integration of Artificial Intelligence and Life Cycle Assessment Methods. IOP Conference Series: Materials Science and Engineering , 1196 (1), 012028. https://doi.org/10.1088/1757-899x/1196/1/012028 Portolani, P., Vitali, A., Cornago, S., Rovelli, D., Brondi, C., Low, J. S. C., Ramakrishna, S., & Ballarino, A. (2022). Machine learning to forecast electricity hourly LCA impacts due to a dynamic electricity technology mix. Frontiers in Sustainability , 3 (Lci). https://doi.org/10.3389/frsus.2022.1037497 Qian, Z. (n.d.). Synthcity : facilitating innovative use cases of synthetic data in different data modalities . 1–14. Rastogi, A., Garamendi, J. F., Fern, A., & Guitart, A. (2023). Synthetic Data Generator For Adaptive Interventions In Global Health . 1–9. Rozo, A., Moeyersons, J., Morales, J., Garcia van der Westen, R., Lijnen, L., Smeets, C., Jantzen, S., Monpellier, V., Ruttens, D., Van Hoof, C., Van Huffel, S., Groenendaal, W., & Varon, C. (2022). Data Augmentation and Transfer Learning for Data Quality Assessment in Respiratory Monitoring. Frontiers in Bioengineering and Biotechnology , 10 (February), 1–14. https://doi.org/10.3389/fbioe.2022.806761 Sari, D. W., Hidayat, F. N., & Abdul, I. (2021). Efficiency of land use in smallholder palm oil plantations in indonesia: A stochastic frontier approach. Forest and Society , 5 (1), 75–89. https://doi.org/10.24259/fs.v5i1.10912 Saswattecha, K., Cuevas Romero, M., Hein, L., Jawjit, W., & Kroeze, C. (2015). Non-CO2 greenhouse gas emissions from palm oil production in Thailand. Journal of Integrative Environmental Sciences , 12 , 67–85. https://doi.org/10.1080/1943815X.2015.1110184 Schiappa, M. (2019). Performance Metrics in Machine Learning . Towardsdatascience.Com. https://towardsdatascience.com/metrics-ml-2563f9e47faa Silitonga, P. D. P., Himawan, H., & Damanik, R. (2020). Forecasting acceptance of new students using double exponential smoothing method. Journal of Critical Reviews , 7 (1), 300–305. https://doi.org/10.31838/jcr.07.01.57 Siregar, K., Ichwana, I., & Nasution, I. S. (2020). Implementation of Life Cycle Assessment ( LCA ) for oil palm industry in Aceh Implementation of Life Cycle Assessment ( LCA ) for oil palm industry in Aceh Province , Indonesia . October . https://doi.org/10.1088/1755-1315/542/1/012046 Stichnothe, H., & Bessou, C. (2017). Challenges for Life Cycle Assessment Of Palm Oil Production System. Indonesian Journal of Life Cycle Assessment and Sustainability , 1 (2), 1–9. https://doi.org/10.52394/ijolcas.v1i2.28 Union, I., Conservation, F. O. R., & Nature, O. F. (2018). Oil palm and biodiversity: a situation analysis by the IUCN Oil Palm Task Force. In Oil palm and biodiversity: a situation analysis by the IUCN Oil Palm Task Force . https://doi.org/10.2305/iucn.ch.2018.11.en Yan. (2022). 10 largest oil palm plantations in Indonesia . Indonesiabusinesspost.Com. https://indonesiabusinesspost.com/insider/10-largest-oil-palm-plantations-in-indonesia/ Yoon, J., Drumright, L. N., & Van Der Schaar, M. (2020). Anonymization through data synthesis using generative adversarial networks (ADS-GAN). IEEE Journal of Biomedical and Health Informatics , 24 (8), 2378–2388. https://doi.org/10.1109/JBHI.2020.2980262 Yusniati, Parinduri, L., & Sulaiman, O. K. (2018). Biomass analysis at palm oil factory as an electric power plant. Journal of Physics: Conference Series , 1007 (1). https://doi.org/10.1088/1742-6596/1007/1/012053 Additional Declarations No competing interests reported. Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-3675682","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":255388499,"identity":"59fd1edb-ed66-4353-bea5-97731f7006fe","order_by":0,"name":"Imam Tahyudin","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAABT0lEQVRIie3RPUvDQBgH8CcEnK7NJoUO9SNECsWh1g/icuEgXRppKZRAhUaE6xLtmqDoV4gUOkcO0yXqehAQRHQuBMTRBFK9Nrg73B+Oe+N391wCICPzD7NT9Dhregh2toLWez8DUJzaNsFrEm+TsExAJKBQ8fBfAiCQ6i57D5SvzonOyR0b3HSOq4g9zAf0xZghFqUrGxqao55xobC62eKAyVDnZp/5C2LRCjUTnw4N/4ISL4xh3wuV8wOR4JyoRsB7OqssVItqqJVUKLaCZ9SEewpKkBUsPr/e/czIpCDXE5FoaU6OSqSX38IK4rCsMLcgj66aE6NMRhybS8OPP/oMRUuLoshMvCc88d2oCXFcIx7beEvjqrvgq/bYuFySeYpOx9atS6KkP8JNDZE3sO324Ww6fRW+WPFbAPZCYUEVxrXNqZCG8weRkZGRkfkGMNSIVkFP43IAAAAASUVORK5CYII=","orcid":"","institution":"Universitas Amikom Purwokerto","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Imam","middleName":"","lastName":"Tahyudin","suffix":""},{"id":255388501,"identity":"ec2441ad-9c85-44b6-9533-577c87fc3e94","order_by":1,"name":"Ades Tikaningsih","email":"","orcid":"","institution":"Universitas Amikom Purwokerto","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ades","middleName":"","lastName":"Tikaningsih","suffix":""},{"id":255388504,"identity":"2bb4f1f4-49d2-442f-b495-dfedaeb36efa","order_by":2,"name":"Yaya Suryana","email":"","orcid":"","institution":"National Research and Innovation Agency (BRIN)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yaya","middleName":"","lastName":"Suryana","suffix":""},{"id":255388506,"identity":"6fd1463f-27b4-4ecd-9697-03228e15916c","order_by":3,"name":"Hanung Adi Nugroho","email":"","orcid":"","institution":"Universitas Gadjah Mada Yogyakarta","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Hanung","middleName":"Adi","lastName":"Nugroho","suffix":""},{"id":255388507,"identity":"ab438b44-3075-47d0-abc0-4f0bdbc6fcd6","order_by":4,"name":"Ade Nurhopipah","email":"","orcid":"","institution":"Universitas Amikom Purwokerto","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ade","middleName":"","lastName":"Nurhopipah","suffix":""},{"id":255388508,"identity":"842cc01a-bea8-44dc-b11c-0ee80fca0c79","order_by":5,"name":"Nugroho Adi Sasongko","email":"","orcid":"","institution":"National Research and Innovation Agency (BRIN)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Nugroho","middleName":"Adi","lastName":"Sasongko","suffix":""},{"id":255388509,"identity":"0c8cfb02-82e2-4ead-b8e0-f84bcb0e507d","order_by":6,"name":"Agus Bejo","email":"","orcid":"","institution":"Universitas Gadjah Mada Yogyakarta","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Agus","middleName":"","lastName":"Bejo","suffix":""},{"id":255388510,"identity":"c768cc86-23fa-4293-9808-24de29e3cf29","order_by":7,"name":"Puji Lestari","email":"","orcid":"","institution":"Universitas Amikom Purwokerto","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Puji","middleName":"","lastName":"Lestari","suffix":""},{"id":255388511,"identity":"02f7bc09-19f0-43cf-9b03-2fd36308bfcc","order_by":8,"name":"Ahmad Ismed Yanuar","email":"","orcid":"","institution":"National Research and Innovation Agency (BRIN)","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ahmad","middleName":"Ismed","lastName":"Yanuar","suffix":""}],"badges":[],"createdAt":"2023-11-28 09:29:27","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-3675682/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-3675682/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":47675029,"identity":"d24c0db6-2479-4e0a-932f-214e2f67f24c","added_by":"auto","created_at":"2023-12-06 04:14:44","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":45503,"visible":true,"origin":"","legend":"\u003cp\u003eResearch workflow\u003c/p\u003e","description":"","filename":"Fig1.Researchworkflow.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/196bb372edef20ae0bed539f.png"},{"id":47675718,"identity":"e010f98e-0fab-4808-bbc3-a203aa5465d7","added_by":"auto","created_at":"2023-12-06 04:22:44","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":31710,"visible":true,"origin":"","legend":"\u003cp\u003eDescriptive statistics of target variables\u003c/p\u003e","description":"","filename":"Fig2.Descriptivestatisticsoftargetvariables.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/e79fe86a6304f8d7bcb80e4a.png"},{"id":47675027,"identity":"bf3c16bc-29b6-4b47-92d1-f2cf369096df","added_by":"auto","created_at":"2023-12-06 04:14:44","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":64063,"visible":true,"origin":"","legend":"\u003cp\u003eFeature importance of GBR\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/1453e13f8e9a487bf75f4e44.png"},{"id":47675031,"identity":"e07c5d17-c03c-4296-b2aa-83c60a803cae","added_by":"auto","created_at":"2023-12-06 04:14:44","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":55831,"visible":true,"origin":"","legend":"\u003cp\u003eFeature Importance RF\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/94d90fa15987bdd75068d10b.png"},{"id":47675032,"identity":"eae822bc-0e16-4595-8b49-ceaeb4e48b36","added_by":"auto","created_at":"2023-12-06 04:14:44","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":260966,"visible":true,"origin":"","legend":"\u003cp\u003eFeature Importance Adaboost\u003c/p\u003e","description":"","filename":"Fig5.FeatureImportanceAdaboost.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/1827170b939afe50025ef062.png"},{"id":47675033,"identity":"50dbe094-4f7e-4a18-a7fa-dabef7251c51","added_by":"auto","created_at":"2023-12-06 04:14:45","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":68988,"visible":true,"origin":"","legend":"\u003cp\u003eRandom Forest Regression\u003c/p\u003e","description":"","filename":"Fig6.RandomForestRegression.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/aff5c01ae3ce010150fd6dff.png"},{"id":47675034,"identity":"d0ff1689-3287-4d13-b8b8-c5debde5e18f","added_by":"auto","created_at":"2023-12-06 04:14:45","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":160612,"visible":true,"origin":"","legend":"\u003cp\u003eActual vs predicted CO\u003csub\u003e2 \u003c/sub\u003eemissions\u003c/p\u003e","description":"","filename":"Fig7.ActualvspredictedCO2emissions.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/fd4c1b84421b9d6d9d5467ca.png"},{"id":47675719,"identity":"1dc2aa50-a12f-46f9-8227-bad52d847d7b","added_by":"auto","created_at":"2023-12-06 04:22:44","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":18366,"visible":true,"origin":"","legend":"\u003cp\u003eAverage actual and predicted\u003c/p\u003e","description":"","filename":"Fig8.Averageactualandpredicted.png","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/de86fea8d1a53949a70edd4b.png"},{"id":48568605,"identity":"0cc6dd0f-c6f2-43ad-9f18-be00a0f6b8d8","added_by":"auto","created_at":"2023-12-20 23:22:21","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":1240990,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-3675682/v1/4c98fd2f-1364-4065-9482-3330f9310bff.pdf"}],"financialInterests":"No competing interests reported.","formattedTitle":"Optimizing Sustainability: A Deep Learning Approach on Data Augmentation of Indonesia Palm Oil Products Emission","fulltext":[{"header":"I. Introduction","content":"\u003cp\u003eIndonesia is the world's largest producer of palm oil and accounts for most of the global supply. The industry contributes significantly, ranging from 1.5\u0026ndash;2.5% of Indonesia's gross domestic product (GDP) (Nations \u0026amp; Programme, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). According to data from the Central Statistical Agency (BPS), the total area of palm coconut plantations in Indonesia is about 11.9\u0026nbsp;million hectares, a figure about three times higher than in 2000, when about four million acres of Indonesian land were used for palm coke plantations. The Indonesian palm coconut industry can also be seen from documented data from the Food and Agriculture Organization (FAO), which provides information that by 2021, the most productive sector in Indonesia will be coca palm, with a total production of about 25.130.562 tons. This figure is much higher when compared to the production of rice products, the primary raw material in Indonesia, with an average production of 21.280.164 tons (Sari et al., \u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe industry has experienced rapid growth and expansion, driven by favorable climate conditions, abundant land resources, and government support. The country's tropical climate provides an ideal environment for palm coconut cultivation, enabling high yields and efficient production. Indonesian palm coconut farming is spread across regions, with significant production areas in Sumatera Island, Kalimantan (Kalimantan), and Papua (Yan, \u003cspan citationid=\"CR47\" class=\"CitationRef\"\u003e2022\u003c/span\u003e). However, the palm coconut industry in Indonesia has challenges. Environmental problems such as deforestation and habitat loss have been linked to the expansion of palm and coconut plantations, raising concerns about the conservation of biodiversity and the destruction of vital ecosystems, including tropical rainforests (Union et al., \u003cspan citationid=\"CR46\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). However, it is essential to note that not all palm and coconut plantations are created equally. Several farms have been established on degraded land or agricultural debris, which can help reduce the pressure on natural forests (Denis J Murphy, 2019).\u003c/p\u003e \u003cp\u003eLife Cycle Assessment (LCA) for assessing the environmental impact of palm oil products is essential because of the unique challenges and controversies associated with the oil industry (Siregar et al., \u003cspan citationid=\"CR44\" class=\"CitationRef\"\u003e2020\u003c/span\u003e). LCA considers the entire life cycle, including cultivation, processing, transportation, and disposal of products, thus providing a holistic understanding of environmental consequences. This enables it to thoroughly assess the environmental impacts of palm oil production, such as deforestation, soil degradation, carbon emissions, water pollution, and biodiversity loss. The LCA enables the comparison of carbon footprints in various palm oil production systems, enabling the identification of emission points and the development of strategies to reduce carbon intensity. Palm oil production requires substantial water resources, and improper management can lead to water shortages and pollution (Amalia Yunia Rahmawati, \u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2020\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eData processing will support LCA analysis, as LCA is important in the palm coconut industry. Machine learning (ML) is a popular and robust method for analyzing data sets. Unfortunately, the data supporting this case in Indonesia is so limited that the ML performance could be more satisfactory. In this context, additional data sets are needed to improve ML accuracy. Added data integrates primary data, collected through surveys and field measurements, with secondary data from various sources, such as databases, literature, expert knowledge, or data sets created based on theoretical experiments.\u003c/p\u003e \u003cp\u003eLCA relies on data availability, and the palm coconut industry often faces data gaps, especially regarding specific environmental impacts. Additional data helps fill this gap by combining primary data collected through field measurements and surveys with secondary data from relevant sources. Added data contributes to the improved quality and reliability of LCA results. Combining primary and secondary data can reduce the potential bias and uncertainty associated with a single data source (Lohr, \u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e2018\u003c/span\u003e). Primary data collection enables location-specific information, while secondary source data provides a broader context and generalizable information. Expanded data enables incorporating of local knowledge and stakeholders, which can significantly improve the accuracy and relevance of LCA assessments (Nations \u0026amp; Programme, \u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e2023\u003c/span\u003e). Therefore, expanded data contributes to developing evidence-based policies and strategies that promote sustainable palm oil production and minimize environmental impact.\u003c/p\u003e \u003cp\u003eThere are several techniques for data augmentation, such as noise injection (Gutowski, \u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e2018\u003c/span\u003e) and Anonymization Through Data Synthesis (ADS-GAN). ADS-GAN is a deep learning technique used to generate data based on its population. The ADS-GAN network consists of two main parts: a generator and a discriminator. This technology overcomes data constraints and can improve prediction accuracy (Chen et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2021\u003c/span\u003e),(A. Khan et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2021\u003c/span\u003e). Therefore, this study aims to solve the problem of limiting the analysis of palm coconut products with the ADS-GAN technique. In addition, to determine the accuracy of the prediction, it best uses some regression algorithms: Random Forest Regression (RFR), Gradient Boosting Regressions (GBR), and Adaptive Boosting Regression (ABR).\u003c/p\u003e"},{"header":"II. literature review","content":"\u003cp\u003ePrevious studies have successfully investigated carbon footprints in palm oil products using machine learning approaches. However, research that applies the latest machine learning techniques, such as the Generative Adversarial Network (GAN), to analyze the carbon footprint of palm petroleum products still needs to be completed.\u003c/p\u003e \u003cp\u003eIn the health sector, the implementation of synthetic data has successfully overcome the constraints of medical data that often need to be updated. Synthetic data generated through generative hostile tissue (GAN) contributes positively to the performance of artificial intelligence algorithms, as seen in kidney cell carcinoma research. Continuing this, a related study added synthetic data using virtual sensors to diagnose defects on rotating machines with 42 different classes. The results recorded an improvement in training accuracy of about 6\u0026ndash;15% and validation precision of about 44\u0026ndash;49%. Accordingly, studies on the design of fractional factorial experiments are used to evaluate the impact of basic data augmentation methods on the detection of COVID-19 (Chen et al., \u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e2021\u003c/span\u003e),(A. Khan et al., \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e2021\u003c/span\u003e),(Davila et al., \u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e2023\u003c/span\u003e)\u003c/p\u003e \u003cp\u003eIn the area of time series data analysis, research on data augmentation has been the focus of some researchers. In his experiment, the researchers found that the implementation of data augmentation significantly improves the performance of the model. At the same time, other studies have a similar focus on applying data amplification techniques to timelines and integrating them into the context of timeline classification using simulated neural networks (Fawaz et al., \u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e2018\u003c/span\u003e),(Iwana \u0026amp; Uchida, \u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e2021\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eLater, in the context of computer vision, the introduction of RLR (random local rotation) geometric augmentation techniques was done to manipulate local information on images without adding non-original pixel values. Other studies of data augmentation that can be applied in the computer vision domain include thoroughly studied augmentation strategies, feature-based amplification techniques, and meta-learning-based amplification. However, to determine the transformation or combination of the best transformation, it is often necessary to go through trial-and-error processes (Alomar et al., \u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e2023\u003c/span\u003e),(Mumuni \u0026amp; Mumuni, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eSimilar research focused on developing a Bisindo translator model that translates sign language signals into text using a machine learning approach based on convolutional neural networks (CNN). As a result, this effort increased the model's performance to an accuracy of 94.38%. Other research highlighted the Convolutional Neural Network Model (CNN) towards data augmentation through the use of technology that can be applied to the classification of respiratory signals, thus enabling focus on relevant data (Fadillah et al., \u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e2021\u003c/span\u003e),(Rozo et al., \u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eSome studies have also applied the data augmentation concept in the context of life cycle analysis (LCA), in which research explores the application of machine learning in life cycle analysis supported by dynamic data research to predict the environmental performance of a product or service using artificial intelligence techniques. Using natural language processing (NLP) and random forest algorithms to train models to provide quick predictions for LCA practitioners and testers in implementing LCA, Similar things have been analyzed to predict the impact of LCA on electricity consumption. Comparing feed-forward (NN) neural networks and repetitive neural (RNN) networks, although limited to one data set (Ghoroghi et al., \u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e2022\u003c/span\u003e),(Koyamparambath et al., \u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e2022\u003c/span\u003e),(Portolani et al., \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eThe study examined the environmental impact of two palm oil production systems using the ISO standard for life cycle analysis. Although this study improves transparency and reproductivity, it does not consider land-use changes, which can be an essential factor in the reality of palm oil production in some areas. A study evaluates the potential for using artificial intelligence (AI) to improve life cycle impact assessments (LCAs) accuracy and efficiency. It uses a qualitative approach to analyze the potential use of AI in LCAs. Similarly, further research aims to identify challenges and opportunities in conducting life-cycle analysis (LACs) on palm oil production systems. By highlighting the need for spatially explicit methods and the impact of land-use change, this study provides in-depth insights (Heinz Stichnothe, \u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e2011\u003c/span\u003e),(Oduque de Jesus et al., \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e2021\u003c/span\u003e),(Stichnothe \u0026amp; Bessou, \u003cspan citationid=\"CR45\" class=\"CitationRef\"\u003e2017\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn the context of augmented data prediction of CO\u003csub\u003e2\u003c/sub\u003e emission tracks and estimates of coconut oil and palm oil yields, Several researchers have undertaken exploration for life cycle analysis (LCA) of crude palm oil production in the Philippines. This study provides insight into the potential environmental impact of palm oil production. Another study aims to develop a machine-learning model to predict total global CO\u003csub\u003e2\u003c/sub\u003e emissions and understand the impact of the COVID-19 pandemic on future carbon emissions. The palm oil results are also being analyzed, considering weather and soil humidity changes. The evaluation of the performance of various machine learning algorithms achieved high accuracy (95%) compared to traditional methods (Espino et al., \u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e2019\u003c/span\u003e),(Meng \u0026amp; Noman, \u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e2022\u003c/span\u003e),(N. Khan et al., \u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e2022\u003c/span\u003e).\u003c/p\u003e \u003cp\u003eIn order to improve the performance of systems or learning models, focusing on using synthetic data or data augmentation, a researcher concentrated on utilizing two GAN architectures, GAN and CGAN. The investigation results showed that the synthesized CGAN-based data had reached 63%. Compared to the study, the augmentation of data to improve active learning performance (AL) in the classification showed a significant improvement in AL performance (Aziira et al., \u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e2020\u003c/span\u003e),(Fonseca \u0026amp; Bacao, \u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e2023\u003c/span\u003e).\u003c/p\u003e"},{"header":"III. MATERIAL AND METHOD","content":"\u003cp\u003eThe research method will be explained through a flowchart to provide an understanding of the research implementation, as depicted in Figure 1.\u003c/p\u003e\n\u003ch2\u003eA. Data Description\u003c/h2\u003e\n\u003cp\u003eThe dataset comes from the Research Centre for Sustainable Production Systems and Life Assessment of the National Innovation Research Agency (BRIN). The amount of original data collected was 195 samples with 19 variables. Table I includes a complete description of the variables included in this process.\u003c/p\u003e\n\u003cp\u003eTABLE\u0026nbsp;I. DESCRIPTION OF THE DATA USED\u003c/p\u003e\n\u003cdiv\u003e\n \u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"317\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eVariable\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eDescription\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eLand (Ha)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of land in hectares used for oil palm cultivation.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eWater (m3)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of water in cubic meters used for oil palm cultivation and processing.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eDolomite\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA mineral fertilizer used to improve soil pH and provide calcium and magnesium.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eUrea (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA nitrogen fertilizer used to increase crop yields.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eBunch Ash (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA byproduct of oil palm processing that can be used as a fertilizer.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eRock Phosphate (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA phosphorus fertilizer used to improve crop yields.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eBorate (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA boron fertilizer used to improve crop yields and quality.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eParaquat (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA herbicide used to control weeds in oil palm plantations.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eGlyphosate (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eA herbicide used to control weeds in oil palm plantations.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eFFB (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of fresh fruit bunches (FFB) produced in tons.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eWater for boiler (m3)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of water in cubic meters used to generate steam for oil palm processing.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eTransport (kWh)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe cost of transporting FFB to the processing mill and CPO to the market.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eGenset (kWh)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe cost of generating electricity using diesel generators.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eElectricity (kWh)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of electricity in kilowatt-hours used for oil palm processing.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eCPO (Ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of crude palm oil (CPO) produced in tons.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eProductivity (area - FFB)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of FFB produced per hectare of land.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eProductivity (area - CPO)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of CPO produced per hectare of land.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eGas Emission (CO\u003csub\u003e2\u003c/sub\u003e-eq/ton)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of greenhouse gas emissions in carbon dioxide equivalents per ton of CPO produced.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.444794952681388%\"\u003e\n \u003cp\u003e\u003cem\u003eBiodiesel (Litres)\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"72.55520504731861%\" valign=\"top\"\u003e\n \u003cp\u003eThe amount of biodiesel produced in litres.\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n\u003c/div\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eThis research attempts to overcome the limitations of original data with data synthesis techniques. Using synthetic data in this context can add value to an effort to understand and reduce the carbon impact of palm oil products in Indonesia.\u003c/p\u003e\n\u003ch2\u003eB. Synthetic Data Augmentation\u003c/h2\u003e\n\u003cp\u003eLarge and representative datasets play a crucial role in guaranteeing the accuracy and precision of model predictions. To overcome data constraints using synthesis methods. Synthetic data can address data shortages, improve model performance, and improve data collection efficiency \u0026nbsp;(Lamberti, 2023). One of the latest methods that got severe attention is generative adversarial networks (GANs).\u003c/p\u003e\n\u003cp\u003eGenerative Adversarial Networks (GANs) are a type of deep learning model that can generate synthetic data that mimics accurate data. The learning process in the GAN involves one generator and one discriminatory nerve network that plays the minimax zero-sum game. Generators create synthetic data, while discriminators evaluate whether the data is natural or synthetic. The generator improves its ability to create realistic synthetic data through this hostile process. In the end, if everything goes well, the generator can generate synthetic data that appears natural but is difficult for discriminators (or humans) to distinguish as accurate or synthesized (Huang, 2021).\u003c/p\u003e\n\u003cp\u003eThe research investigates the potential and limitations of three methods, namely Anonymization Through Data Synthesis (ADS-GAN), Private Aggregation of Teacher Ensembles (PATE-GAN) and Differentially Private (DP-GAN). Involving 195 samples of original data, the synthesis process is performed randomly to produce additional amounts of data that vary: 1000, 5000, and 10000. All variables in the data set are considered sensitive, except for the primary target variable, \u0026apos;Gas Emissions (CO\u003csub\u003e2\u003c/sub\u003e-eq/tonne)\u0026apos;. The choice of these three methods is made with consideration of the superiority of each. ADSGAN was chosen because of its ability to generate synthetic data that matches the characteristics of the original data. In contrast, PATEGAN and DPGAN were selected for their focus on differential privacy concepts to protect the confidentiality of individual data, placing particular emphasis on privacy security (Rastogi et al., 2023),(Fan, n.d.).\u003c/p\u003e\n\u003cp\u003eThe resulting synthetic data will be analyzed using three regression models: Random Forest Regressor (RFR), Gradient Boosting Regresor (GBR), and Adaptive Boosting regressor (ABR), to evaluate the synthesis quality of the three methods.\u003c/p\u003e\n\u003ch2\u003eC. Data Synthesis Process Across Multiple GANs\u003c/h2\u003e\n\u003cp\u003eThe synthesis data process uses Generative Adversarial Networks (GANs) using synthcity libraries. Synthcity serves as a library that captures the entire workflow in producing and evaluating synthetic data. This library provides synthesis data plugins such as ADS-GAN, PATE-GAN, and DP-GAN (Bedorf, 2023).\u003c/p\u003e\n\u003cp\u003eUsing synthcity, the data synthesis phase can be run more efficiently. This library provides the ability to manage workflows comprehensively, including implementing various types of GANs such as Madigan and pagan. Each plugin has a unique role in generating synthetic data that matches the research needs. The synthcity library is the foundation that records the entire process of producing and evaluating synthetic data. A typical workflow using Synthcity involves a set of detailed steps. Here are the general steps found in the synth city workflow (Qian, n.d.):\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e1) Loading the dataset using a DataLoader\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA class called DataLoader offers a standardized interface for loading and storing many kinds of input data, including survival, time series, and tabular data. It was employed in this work to import the original data, which made up 195 data points from the sample and served as the foundation for the synthesis process. This data collection includes several characteristics about the production of palm oil and its effects on the environment. Still, it focuses on the critical variable \u0026quot;gas emissions (CO\u003csub\u003e2\u003c/sub\u003e-eq/tonne).\u0026quot;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2) Training the Generator using a Plugin\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThrough the Plugin class in Synthcity, users can design, train, and use a range of data generators. Every plugin uses a different algorithm to produce data. The fit() method offered by the plugin is used to train the generator. The plugins employed in the context of this study are \u0026quot;pategan\u0026quot;, \u0026quot;dpgan\u0026quot;, and \u0026quot;adsgan\u0026quot;.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3) Generating Synthetic Data\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eUsers can create fake data by using the generate() method once the plugin has been trained. Conditional generation is also possible with some plugins.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e4) Evaluating Synthetic Data\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eA wide range of measures are available from Synthcity to assess the integrity, use, and privacy of synthetic data. Users can do assessments using the Metrics class.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eD. Regression Algorithm\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e1) Adaptive Boosting regressor\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eAdaboost is a supervised machine-learning algorithm that can solve classification and regression problems. The advantage of this algorithm is that it handles complex data and feature interactions. It can handle complex data and feature interactions, prevent overfitting, and is relatively easy to use (Jyotsna Vadakkanmarveettil, 2021). The drawback is that it is sensitive to specific parameters because the Adaboost algorithm is sensitive to data distribution and may not work well if the data has specific characteristics (Horny\u0026aacute;k \u0026amp; Iantovics, 2023).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2) Random Forest Regressor\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eRF (Random Forest) is an ML algorithm that uses several decision trees to make predictions. The advantage of this algorithm is its ability to handle complex or unstructured data and provide accurate results in making predictions. However, the weakness of this algorithm is the possibility of overfitting, namely the prediction model that is too fit with the training data so that it cannot be used for test data\u0026nbsp;(Davila et al., 2023).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3) Gradient Boosting Regressor\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eGradient Boosting is an ML algorithm that uses several weak predictive models (weak learners) to produce solid predictive models. The advantage of this algorithm is its ability to handle complex or unstructured data and provide accurate results when making predictions. However, the weakness of this algorithm is the computation time, which is quite long and sensitive to the parameters used (Fadillah et al., 2021).\u003c/p\u003e\n\u003ch2\u003eE. Performance Analysis\u003c/h2\u003e\n\u003cp\u003eAnalysis of the performance of machine learning algorithms in regression models involves matrix calculations as an evaluation method. This process assesses how much the model can predict target values (dependent variables) based on input variables (independent variables). There are several standard evaluation metrics used for regression models, and most of them can be calculated using a payoff matrix using the following metric values (Silitonga et al., 2020)(Chicco et al., 2021)(Jadon et al., 2022):\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e1) \u0026nbsp; \u0026nbsp;Mean Absolut Error (MAE)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMAE is used to calculate the error value between the predicted value obtained from the model results and the actual value, and then absolute it.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2) \u0026nbsp; \u0026nbsp;Mean Square Error (MSE)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMSE measures the average squared difference between the model\u0026apos;s predicted values and the actual values of the observed data\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e3) \u0026nbsp; \u0026nbsp;Mean Squared Logarithmic Error (MSLE) \u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eMSLE is a variant of Mean Squared Error (MSE), which uses the logarithm of the predicted value and the actual value. MSLE measures the average of the squared logarithmic differences between predicted values and actual values. The best model is the model that obtains the lesser value of MAE, MSE, and MSLE.\u003c/p\u003e"},{"header":"IV. RESULT and DISCUSSION","content":"\u003cp\u003eThe analysis of the original palm and coconut data provides an initial understanding of the data collected. The primary dataset briefly describes the target variable, which can be seen in Figure 2.\u003c/p\u003e\n\u003cp\u003eThrough Figure 2, a statistical histogram visualization of descriptive emission variables provides in-depth insight into data distribution on data sets. The amount of data, or count, can be estimated through the number of bars on the histogram. The histogram\u0026apos;s average gas emission indicated by the central distribution position is about 3.79 CO\u003csub\u003e2\u003c/sub\u003e-eq/ton. At the same time, the standard deviation of about 2.39 shows a significant spread of the average value. Min and Max, which represent the minimum and maximum values of the distribution, are seen on the x-axis of the histogram, with the smallest values reaching 0.09 CO\u003csub\u003e2\u003c/sub\u003e-eq/ton and the most outstanding value reaching 25.48 CO\u003csub\u003e2\u003c/sub\u003e/ton. In addition, a median value of around 2.39 CO\u003csub\u003e2\u003c/sub\u003e per ton can be found in the middle of such a distribution. Histograms, as a form of visualization, provide a more intuitive understanding of data distribution and reinforce descriptive statistical findings related to emission variables in the palm coconut industry.\u003c/p\u003e\n\u003ch2\u003eA. Downstream Task Performances With Synthetic Datasets\u003c/h2\u003e\n\u003cp\u003eIn this work, data synthesis was carried out using the Generative Adversarial Networks (GAN) model. The experiments were conducted with variations of data amounts, i.e., 1000, 5000, and 10000, to evaluate the model\u0026apos;s performance. The models assessed included ADS-GAN, PATE-GAN, and DP-GAN. The evaluation used three main metrics: Jensen-Shannon Divergence (JSD), Wasserstein distance, and Kolmogorov-Smirnov Test (KS Test). The use of these metrics was aimed at measuring to what extent the distribution of synthetic data generated by the GAN model was similar to that of the original data distribution. Jensen-Shannon Divergence provides information about the differences in probability distribution between two datasets, while Wasserstein distance estimates how much change is needed to transform one distribution into another. The Kolmogorov-Smirnov Test (KS Test), on the other hand, is used to test the extent to which two probability distributions are simultaneous (Yoon et al., 2020),(Schiappa, 2019).\u003c/p\u003e\n\u003cp\u003eTABLE\u0026nbsp;II\u0026nbsp;COMPARISON OF SYNTHETIC DATA GENERATION IN TERMS OF JSD, WASSERSTEIN DISTANCE AND KS TEST\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"324\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"15.432098765432098%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"23.14814814814815%\" valign=\"top\"\u003e\n \u003cp\u003e\u0026nbsp;\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003eADS-GAN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003ePATE-GAN\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.59259259259259%\" valign=\"top\"\u003e\n \u003cp\u003eDP-GAN\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"15.432098765432098%\" rowspan=\"3\" valign=\"top\"\u003e\n \u003cp\u003e1000\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"23.14814814814815%\" valign=\"top\"\u003e\n \u003cp\u003eWasserstein distance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003e0.082203\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003e1.568033\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.59259259259259%\" valign=\"top\"\u003e\n \u003cp\u003e0.203662\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.37226277372263%\" valign=\"top\"\u003e\n \u003cp\u003eJSD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.018408\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.018411\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.802919708029197%\" valign=\"top\"\u003e\n \u003cp\u003e0.042283\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.37226277372263%\" valign=\"top\"\u003e\n \u003cp\u003eKS Test\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.628880\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.175439\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.802919708029197%\" valign=\"top\"\u003e\n \u003cp\u003e0.641026\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"15.432098765432098%\" rowspan=\"3\" valign=\"top\"\u003e\n \u003cp\u003e5000\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"23.14814814814815%\" valign=\"top\"\u003e\n \u003cp\u003eWasserstein distance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003e0.504504\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003e0.723084\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.59259259259259%\" valign=\"top\"\u003e\n \u003cp\u003e1.936032\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.37226277372263%\" valign=\"top\"\u003e\n \u003cp\u003eJSD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.022357\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.036514\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.802919708029197%\" valign=\"top\"\u003e\n \u003cp\u003e0.058279\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.37226277372263%\" valign=\"top\"\u003e\n \u003cp\u003eKS Test\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.336032\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.402159\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.802919708029197%\" valign=\"top\"\u003e\n \u003cp\u003e0.191633\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"15.432098765432098%\" rowspan=\"3\" valign=\"top\"\u003e\n \u003cp\u003e10000\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"23.14814814814815%\" valign=\"top\"\u003e\n \u003cp\u003eWasserstein distance\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003e0.469520\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"21.91358024691358%\" valign=\"top\"\u003e\n \u003cp\u003e0.736796\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"17.59259259259259%\" valign=\"top\"\u003e\n \u003cp\u003e1.975019\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.37226277372263%\" valign=\"top\"\u003e\n \u003cp\u003eJSD\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.020051\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.036710\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.802919708029197%\" valign=\"top\"\u003e\n \u003cp\u003e0.059375\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"27.37226277372263%\" valign=\"top\"\u003e\n \u003cp\u003eKS Test\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.372470\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.912408759124087%\" valign=\"top\"\u003e\n \u003cp\u003e0.388664\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"20.802919708029197%\" valign=\"top\"\u003e\n \u003cp\u003e0.174089\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eBased on Table II, we compared the performance of three synthetic data generation approaches, namely ADS-GAN, PATE-GAN, and DP-GAN, using Wasserstein Distance, Jensen-Shannon Divergence (JSD), and KS Test evaluation metrics in three different scenarios: 1000, 5000, and 10000 synthetical data. Then, in the case of 1000 synthetic data, ADS-GAN showed superior performance with a low Wasserstein Distance (0.082203) compared to PATE-GAN (1.568033) and DP-GAN (0.203662). Similarly, in JSD, the ADS -GAN (0.018408) had a meager value, showing high similarities with the original data distribution. KS Test on ADS \u0026ndash;GAN (0.628880) also showed promising results in measuring distribution equality.\u003c/p\u003e\n\u003cp\u003eMeanwhile, ADS-GAN performed well in the 5000 synthetic data scenario with a stable low Wasserstein Distance (0.504504). PATE-GAN also showed promising results, while DP-GAN had higher values on the metric. Nevertheless, ADS-GAN remained superior with low JSD (0.022357) and adequate KS Test (0.336032). Finally, on 10000 synthetic data, ADs-GAN and PATE-GAN again showed good performance with relatively low Wasserstein Distance. ADS -GAN (0.469520) had low J SD (0.020051) and satisfactory KS Test (0.372470). PATEGAN also showed a good result, while DP-GAN showed higher values on both metrics.\u003c/p\u003e\n\u003cp\u003eOverall, ADS-GAN shows consistent performance with the ability to generate synthetic data close to the original data distribution, compared to PATE-GAN and DP-GAN. In each test scenario (1000, 5000, and 10,000 synthetical data), ADS -GAN consistently stands out with superior performance in evaluation metrics such as Wasserstein Distance, Jensen-Shannon Divergence (JSD), and KS Test. Specifically, in scenarios with 1000 synthetic data, the ADS-GAN implementation performed better than 5000 and 10000 systems. This is mainly seen from the Wasserstein and JSD Distance metric evaluations. In this context, the low Wasserstein Distance value and the low JSD value in 1000 scenarios indicate that the distribution of synthetical data generated by the ADS-GAN is closer to the original distribution than the scenario with more significant amounts of synthesized data.\u003c/p\u003e\n\u003cp\u003eThus, this conclusion affirms that ADS-GAN can deliver optimal performance, especially in scenarios with more limited quantities of synthetic data. This is an important consideration, especially when computational efficiency or limitation of synthesized data is a significant factor.\u003c/p\u003e\n\u003ch2\u003eB. Performance of Regression Models on Synthetic Data\u003c/h2\u003e\n\u003cp\u003eRegression is a supervised machine learning model directed at predicting continuous values. A regression model describes the relationship between a dependent variable and one or more independent variables. This research uses regression analysis to predict gas emissions from palm oil products. Building this model involves comparing two sets of data, namely 1000 samples of synthetic palm oil data, as a first step before comparing model performance using accurate data. The selection of 1000 synthetic data samples as a basis for comparison was based on the finding that this number provided optimal model performance, as seen from the evaluation using the Wasserstein Distance, Jensen-Shannon Divergence (JSD), and KS Test metrics documented in Table II.\u003c/p\u003e\n\u003cp\u003eThen, the algorithms used include Random Forest Regression (RFR), Gradient Boosting Regression (GBR), and Adaptive Boosting Regression (ABR). They were chosen because of their advantages in handling data complexity and their ability to capture non-linear relationships between variables. Meanwhile, the dependent variable in the context of this regression model is \u0026quot;Gas Emissions (CO\u003csub\u003e2\u003c/sub\u003e-eq/ton),\u0026quot; which refers to the amount of greenhouse gases emitted per metric ton of coconut palm produced or processed. Meanwhile, the independent variable involves 19 features that can affect emission levels. In building a regression model, dividing data into two subgroups is crucial to training and testing the model\u0026apos;s performance. Training sets are used to train models, which means models are \u0026quot;learned\u0026quot; from patterns and relationships in the data. Instead, a testing set is used to test the performance of models on data that have never been seen before (Kumar, 2020). This study applies the separation of data with an 80/20 ratio. The ratio has a good chance of understanding and adapting to data variations while still providing a good test on data that has never been seen before.\u003c/p\u003e\n\u003cp\u003eIn the following explanation, we will present an analysis of the results of the Gradient Boosting Regression (GBR), Random Forest (RF), and AdaBoost models using 1000 synthetic data samples. This analysis will provide a deeper understanding of how well these three models can provide predictions of carbon emissions on palm oil data, illustrating the strengths and weaknesses of each model in a more specific applicable context.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e1) \u0026nbsp;Gradient Boosting Regression (GBR)\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe Gradient Boosting Regressor (GBR) model has shown excellent performance in predicting palm coconut carbon emissions. The GBR has an MAE rate of around 0.00341. In addition, the mean squared error (MSE) of around 6.671114421152337e-05 shows the ability of the model to reduce predictive errors, and the MSE is more sensitive to significant errors because it squares the difference between the prediction and the actual value. A low MSE value shows the model\u0026apos;s ability to reduce prediction errors effectively. The GBR model also has an MSLE of around 6.323034859557045e-05, which is useful when measuring errors on a logarithmic scale. Low MSLE values indicate that the model has a minimal error rate on the logarithm scale.\u003c/p\u003e\n\u003cp\u003eFigure 3 illustrates a variable that plays a vital role in predicting emissions of palm oil products using the Gradient Boosting Regressor algorithm. The top three variables are \u0026apos;Genset,\u0026apos; \u0026apos;Transport,\u0026apos; and \u0026apos;Urea.\u0026apos; As is well known, the Genset is the primary source of electricity in many palm and coconut plantations (Yusniati et al., 2018), its use involves fossil fuels such as solar or diesel, which contribute to greenhouse gas emissions. The more the generator is used, the higher the greenhouse gas emissions generated (Evers, 2020). Transportation also plays a vital role in the supply chain of the palm coconut industry (Cheah et al., 2023), including the transportation of crops, fertilizers, pesticides, and labor. If transportation uses conventional vehicles with fossil fuels, this will impact CO\u003csub\u003e2\u003c/sub\u003e emissions. Furthermore, using urea as fertilizer requires energy and can produce greenhouse gas emissions, mainly if used excessively, which produces nitrogen oxide (NOx) (Dong et al., 2022).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003e2) Random Forest Regression\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIn analyzing the performance of predicting emissions of palm oil products using the Random Forest Regression algorithm, the model shows excellent results. The mean absolute error MAE has a value of around 0.003132978845438705, indicating high prediction accuracy, while the MSE of around 6.127072889081567e-05 shows the ability of the model to reduce prediction errors. In addition, the MSLE, which is only about 5.838479552074619e-05, describes a low error rate in the logarithm scale. With these results, the Random Forest Regression model proved to have excellent performance in predicting emissions of palm oil products, with high accuracy, a solid ability to reduce predictive errors, and minimal errors in the logarithmic scale.\u003c/p\u003e\n\u003cp\u003ePredicting palm oil emissions using Random Forest Regressor algorithms yielded essential values, as shown in Figure 4.\u003c/p\u003e\n\u003cp\u003eThe feature importance score in the Random Forest Regression algorithm measures the extent to which each variable contributes to the accuracy of the model prediction. As shown in Figure 4, Random Forest Regression has the same result as gradient boosting regression. Features such as \u0026apos;Genset,\u0026apos; \u0026apos;Transport,\u0026apos; and \u0026apos;Urea\u0026apos; found in Random Forest Regression can be regarded as variables that significantly influence the outcome of emission predictions. This means these variables have a decisive role in influencing the emission rate of palm oil products. This understanding can help further decision-making and modeling to reduce or control emissions.\u003c/p\u003e\n\u003ch3\u003e3) \u0026nbsp;Adaptive Boosting Regression\u003c/h3\u003e\n\u003cp\u003eThis model shows quite good results in the performance analysis of the palm oil product emission prediction using the Adaptive Boosting Regressor algorithm. The MAE has a value of around 0.012349103068078711. A low MAE indicates a reasonable degree of accuracy in the prediction. MSE around 0.00019333446314396462 describes the ability of a model to reduce predictive error, and a low MSE indicates that the model has a minimum error rate. Meanwhile, The MSLE has a value of about 0.00018800274586211268, which indicates a low logarithm scale error rate. Based on these evaluation metrics, the adaptive boosting regression model produces accurate predictions, has a solid ability to reduce predictive errors, and has minimal errors in the logarithmic scale. Thus, the adaptive boosting regression algorithm also shows good performance in predicting emissions of palm oil products, which can be used as a basis for better managing these emissions.\u003c/p\u003e\n\u003cp\u003ePredicting palm oil emissions using the adaptive boosting regression algorithm yielded essential values, as shown in Figure 5.\u003c/p\u003e\n\u003cp\u003eAs shown in Figure 5, the top three critical features when using the adaptive boosting regression algorithm are the variables waterfor_boiler(m3), FFB (ton), and Urea. Using water for boilers (m3) in the palm coconut industry can impact carbon emissions. The palm coconut plant has a fuel consumption of 0.37 liters of diesel per ton of fresh fruit (TBS), equivalent to 5.0 kilograms of CO (Hong, 2023). In addition, FFB (ton) in the palm coke industry can also contribute to carbon emissions, with each ton of TBS requiring an allocation of 2 kilometers of transportation (Saswattecha et al., 2015). Urea use as a fertilizer requires energy and can produce greenhouse gas emissions, mainly if used excessively, which produces nitrogen oxide (NOx) (Dong et al., 2022).\u003c/p\u003e\n\u003ch2\u003eC. Comparison of Regression Model Performance on Original Data and Augmented Data\u003c/h2\u003e\n\u003cp\u003eThis section presents the evaluation results of three regression models: gradient boosting, random forest, and adaptive boosting. The results of this evaluation are presented for two different situations: first, when the models are applied to a dataset of 195 samples (original data); second, after the data synthesis process using the ADS-GAN method, where the resulting dataset consists of 1000 samples. Table III and Table IV display the model evaluation based on the calculation of the primary metrics, namely Mean Absolute Error (MAE), Mean Squared Error (MSE), and Mean Squared Logarithmic Error (MSLE) for both data conditions.\u003c/p\u003e\n\u003cp\u003eTABLE\u0026nbsp;III. ORIGINAL DATA EVALUATION\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"324\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.312693498452012%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.102167182662537%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.721362229102166%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.86377708978328%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMSLE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.312693498452012%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eGBR\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.102167182662537%\" valign=\"top\"\u003e\n \u003cp\u003e0.00413079449910069\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.721362229102166%\" valign=\"top\"\u003e\n \u003cp\u003e9.725835958226041e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.86377708978328%\" valign=\"top\"\u003e\n \u003cp\u003e3.0886335234002594e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.312693498452012%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eRF\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.102167182662537%\" valign=\"top\"\u003e\n \u003cp\u003e0.004699656505053924\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.721362229102166%\" valign=\"top\"\u003e\n \u003cp\u003e0.0002242838981326051\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.86377708978328%\" valign=\"top\"\u003e\n \u003cp\u003e6.50071636956904e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"13.312693498452012%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eABR\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.102167182662537%\" valign=\"top\"\u003e\n \u003cp\u003e0.013249024047795621\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.721362229102166%\" valign=\"top\"\u003e\n \u003cp\u003e0.00038966505062171193\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.86377708978328%\" valign=\"top\"\u003e\n \u003cp\u003e0.00018540864271605393\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eTABLE\u0026nbsp;IV. ADS-GAN AUGMENTED DATA EVALUATION\u003c/p\u003e\n\u003ctable border=\"1\" cellspacing=\"0\" cellpadding=\"0\" width=\"324\"\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd width=\"16.666666666666668%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eModel\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.77777777777778%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.617283950617285%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.938271604938272%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cstrong\u003eMSLE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"16.666666666666668%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eGBR\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.77777777777778%\" valign=\"top\"\u003e\n \u003cp\u003e0.0034182220980073807\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.617283950617285%\" valign=\"top\"\u003e\n \u003cp\u003e6.671114421152337e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.938271604938272%\" valign=\"top\"\u003e\n \u003cp\u003e6.323034859557045e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"16.666666666666668%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eRF\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.77777777777778%\" valign=\"top\"\u003e\n \u003cp\u003e0.003132978845438705\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.617283950617285%\" valign=\"top\"\u003e\n \u003cp\u003e6.127072889081567e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.938271604938272%\" valign=\"top\"\u003e\n \u003cp\u003e5.838479552074619e-05\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd width=\"16.666666666666668%\" valign=\"top\"\u003e\n \u003cp\u003e\u003cem\u003eABR\u003c/em\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"27.77777777777778%\" valign=\"top\"\u003e\n \u003cp\u003e0.012349103068078711\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"25.617283950617285%\" valign=\"top\"\u003e\n \u003cp\u003e0.00019333446314396462\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd width=\"29.938271604938272%\" valign=\"top\"\u003e\n \u003cp\u003e0.00018800274586211268\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n\u003c/table\u003e\n\u003cp\u003e\u003cbr\u003e\u003c/p\u003e\n\u003cp\u003eHere is a comparative explanation of the performance of three algorithm models, namely the Gradient Boosting Regressor (GBR), the Random Forest Regressor (RFR), and the Adaptive Boosting Regressor (ABR), in predicting palm coconut emissions. The comparison was done using two datasets: original data with 195 samples and data synthesized with 1000 samples, documented in Tables III and IV. Performance evaluation was done considering four main metrics: mean absolute error (MAE), mean square error, mean logarithmic error and squared logarithmic error.\u003c/p\u003e\n\u003cp\u003eModels with original palm coconut data from Table III show that the Gradient Boosting Regressor (GBR) algorithm has the lowest MAE value of around 0.0041, compared to Random Forest Regressors (RFR) and Adaptive Boosting Regressors (ABR). This value indicates that the GBR prediction has a low and accurate error rate. In addition, GBR also has the smallest of around 0,0002. In the MSLE, adaptive boosting regressions (ABR) have the lowest value of about 0,0001. These results show that GBR is the best choice for predicting palm coconut emissions compared to RFR and ABR.\u003c/p\u003e\n\u003cp\u003eModels with the synthesized data from Table IV show that Random Forest (RF) has the lowest MAE value, around 0.0031, compared to Random Forest Regressor (RFR) and Adaptive Boosting Regressor (ABR). RF also has the smallest MSE value of around 0.00006127072 and the lower value MSLE of about 0.000005838479552074619. The results confirmed that the synthetic palm data delivered the best performance on the RF algorithm model in predicting coconut emissions compared to the other two, the Random Forest Regressor and the AdaBoost Regressor.\u003c/p\u003e\n\u003cp\u003eComparisons between the performance of models using original and synthesized data, as documented in Tables III and IV, reveal significant differences. The results of the analysis showed that the Random Forest algorithm on the synthesis data showed performance improvements, with mean absolute error (MAE) values of 0.003132978845438705, mean square error (MSE) around 6.127072889081567e-05, and mean squared logarithmic error (MSLE) around 5.838479552074619e-05. It should be emphasized that the exponents in the MSE and MSLE values indicate that the prediction error on the synthetic data is deficient, close to zero.\u003c/p\u003e\n\u003cp\u003eFurther, looking at the performance of the Gradient Boosting Regressor (GBR) model from the synthesis data, it was seen that the model succeeded in reducing the MAE and MSE values to 0.0034182220980073807 and 6.671114421152337e-05, respectively. This comparison describes an increase in the accuracy of GBR predictions on synthesis data compared to models using original data. The exponents on MSE values again show that the error rate generated by the GBR model of the synthesized data is minimal. At the same time, the adaptive boosting regression (ABR) model performance of the synthesis data also showed a significant decrease in MAE and MSE values. MAE values were around 0.012349103068078711, while MSE reached 0.00019333446314396462.\u003c/p\u003e\n\u003cp\u003eBased on a comparison of model performance results on accurate and synthetic data, it can be concluded that the best algorithm for synthetical data is Random Forest. This is evident from the significant increase in prediction accuracy, with deficient mean absolute error (MAE) values close to zero, as well as mean square error (MSE) and mean squared logarithmic error (MSLE) that indicate minimal levels of prediction. Error on synthetic data. The Random Forest Regression (RFR) model used to predict carbon emissions in synthetic datasets can be seen through the distribution of residues in the residual plot Fig. 6.\u003c/p\u003e\n\u003cp\u003eResidues are the difference between actual and predictive values generated by the regression model. The average residual distribution, i.e., distributed evenly around zero values, is an essential indication in evaluating the model\u0026apos;s conformity to actual data. The setting of model parameters is a critical factor in achieving optimal performance. By setting n_estimators=700, the RFR model is made with 700 decision trees, providing predictive stability and preventing potential overfitting. The parameter max_depth=50 is used to control the maximum depth of each tree, while min_samples_split=5 and min_ samples_leaf=5 determine the minimum number of samples required to divide the internal node and form the leaf node in the decision tree. Random_state=42 controls randomness in model development, ensuring consistency of results every time a model is run.\u003c/p\u003e\n\u003cp\u003eThe residual distribution analysis in Figure 6 gives a positive picture of the performance of the Random Forest Regressor model. The general pattern indicates that the overall residues are evenly spread around the zero line, showing the consistency of the model in predicting the value. Although some points are away from zero, most of the residue is close to 0, indicating a model\u0026apos;s ability to predict accurately. Furthermore, separating residues into positive and negative categories opens up additional insights. Damaging residue tends to have a smaller magnitude compared to positive residue. This suggests that models are more accurate in predicting smaller values while may encounter difficulties forecasting larger ones. Giving background colors to the positive (\u0026quot;lightgreen\u0026quot;) and negative (\u0026quot;lightcoral\u0026quot;) residues helps visualize where models tend to perform over-predictions and under-prediction. (residue negative).\u003c/p\u003e\n\u003ch2\u003eD. Prediction Model\u003c/h2\u003e\n\u003cp\u003eThe CO\u003csub\u003e2\u003c/sub\u003e emission model prediction results are presented as a scatter plot using Microsoft Excel to illustrate the predictive, descriptive analysis visually. The scatter plot, seen in Figure 7, shows the distribution of points between the actual value and the predicted value. In descriptive analysis, these scatter plot patterns carry valuable information regarding model performance.\u003c/p\u003e\n\u003cp\u003eThe gathering of many points around the point (0,0) indicates that the model can predict CO\u003csub\u003e2\u003c/sub\u003e emissions well in cases with low values. However, scattered dots around the plot indicate variations in model performance, where some predictions approximate the actual values well, while others may have more significant errors. Detection of points far from the diagonal line (0,0) can provide insight into outliers or cases where the model has significant prediction errors.\u003c/p\u003e\n\u003cp\u003eBy understanding these visual characteristics, the analysis provides A solid basis for further decision-making regarding model performance, Identifying specific patterns or trends in its predictions, and Highlighting points that require further investigation.\u003c/p\u003e\n\u003cp\u003eFigure 8 is a column diagram that illustrates the comparison of average CO\u003csub\u003e2\u003c/sub\u003e emissions between actual and predicted values throughout the Indonesia. In this context, the average actual CO\u003csub\u003e2\u003c/sub\u003e emissions are stated to be 0.05, while the average estimated CO\u003csub\u003e2\u003c/sub\u003e emissions from the model is 0.04. This comparison shows that the average predicted value of CO\u003csub\u003e2\u003c/sub\u003e emissions is slightly lower than that of actual CO\u003csub\u003e2\u003c/sub\u003e emissions. In other words, the model predicts CO\u003csub\u003e2\u003c/sub\u003e emissions at slightly lower values than they do. This phenomenon can be interpreted as a tendency for underprediction, where the model tends to estimate values that are less than the actual value of CO\u003csub\u003e2\u003c/sub\u003e emissions. By understanding these differences, further interpretation can be made to investigate specific factors or features that may influence model performance so that the model can be improved to provide more accurate predictions.\u003c/p\u003e\n\u003cp\u003eOverall, the analysis results show that the use of the synthesized data can significantly improve the model\u0026apos;s performance, especially in the context of palm coconut emission predictions. This success is reflected in smaller critical evaluation values such as MAE, MSE, and MSLE on models with synthetic data compared to original ones. These smaller error values indicate that models with synthesis data can better estimate actual values with a lower error rate. This performance improvement can be explained by the ability of the model to recognize patterns and relationships in more complex data, as is the case in data synthesis. Thus, models tend to be more precise and responsive to variations in synthetic datasets, resulting in more accurate predictions related to palm coconut emissions.\u003c/p\u003e\n\u003ch2\u003eE. Conclusion and Future Work\u003c/h2\u003e\n\u003cp\u003eThe research addressed the constraints of limited data sources by applying data augmentation techniques. The researchers increased the number of samples from 195 to 1,000. Its primary purpose is to support predictive analysis of palm oil carbon footprints. To ensure the success of the data augmentation process, a comparison of the performance of the three machine learning approaches, namely the Random Forest Regressor, the Gradient Enhanced Regresor, and the Adaptive Enhanced Regresior, was carried out. This study used the ADS-GAN, PATE-GAN, and DP-GAN methods to produce synthetic data sets. The study results showed that using synthesized data with ADs-GAN yielded the best performance, significantly affecting the performance of regression models with excellent improvements in synthetical data.\u003c/p\u003e\n\u003cp\u003eThis analysis confirms that the synthesis data significantly improves model performance in predicting palm coconut emissions. Models using Random Forest synthetic data (RF) showed lower error values than the original models. On the synthesis data, RF produced MAE, MSE, and MSLE values of 0.003132978845438705, 6.127072889081567e-05, and 5.838479552074619e-05. These results showed that using synthesized data can significantly improve the accuracy of palm oil emission predictions. Furthermore, future tasks involve collecting more original data sets and exploring new techniques to improve palm oil carbon footprint predictive analysis quality continuously.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cdiv class=\"DefinitionList\"\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eLCA\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eLife Cycle Assessment\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGANs\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGenerative Adversial Networks\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eADS-GAN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAnonymization Through Data Synthesis\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003ePATE-GAN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003ePrivate Aggregation of Teacher Ensembles\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eDP-GAN\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eDifferentially Private\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eGBR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eGradient Boosting Regression\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eABR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eAdaptive Boosting Regression\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eRFR\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eRandom Forest Regression\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMAE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Absolute Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMSE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Squared Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eMSLE\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMean Squared Log Error\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eML\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eMachine Learning\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eJSD\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eJensen-Shannon Divergence\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv class=\"DefinitionListEntry\"\u003e \u003cdiv class=\"Term\"\u003eKS Test\u003c/div\u003e \u003cdiv class=\"Description\"\u003e \u003cp\u003eKolmogorov-Smirnov Test.\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003c/div\u003e"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eAuthors\u0026rsquo; information\u0026nbsp;\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eImam Tahyudin: His Doctoral program obtained since 2018 from Kanazawa University, Japan. Since 2023, he has been an associate professor (Computer) at Universitas Amikom Purwokerto. He is the author of a number of journal articles and conference papers. His research interests include machine learning, big data, data science, knowledge discovery, and decision support systems.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eHanung Adi Nugroho: His research interest includes signal and image processing and analysis, computer vision, medical imaging, medical instrumentation, and statistical pattern recognition. His Doctoral program obtained from The University of Queensland, Australia (2005).\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAgus Bejo: is a computer scientist specializing in VLSI, Design of processor and compiler, embedded system, biometric authentication, smartcard application based, IoT. He obtained the doctoral program from Chulalongkorn University, Thailand and the Department of Communications and Integrated Systems, Tokyo Institute of Technology, Japan in 2009 and 2014 respectively.\u003c/p\u003e\n\u003cp\u003eAde Tikaningsih: She is a student from Informatic program in Universitas Amikom Purwokerto. Her research passion in machine learning, data mining, and decision support system.\u003c/p\u003e\n\u003cp\u003eAde Nurhopipan: Her Master program graduated from Gadjah Mada University. Her research interest includes machine learning, image processing, NLP, and decision support system.\u003c/p\u003e\n\u003cp\u003ePuji Lestari: She is a student from Information System program in Universitas Amikom Purwokerto. Her research passion in machine learning, data mining, and decision support system, web programming.\u003c/p\u003e\n\u003cp\u003eYaya Suryana: His research interest includes data mining, machine learning, AI, and decision support system. His doctoral program from Tsukuba University, Japan. He is a research group chairman of Applied Instrumentation Systems, Acquisition \u0026amp; Big Data and Continuous Measurement of BRIN.\u003c/p\u003e\n\u003cp\u003eNugroho Adi Sasongko: His research interest includes sustainable development, renewable energy technologies, life-cycle assessment, and biofuel production. He is a Director of Center for Sustainable Production Systems Research and Life Cycle Assessment, National Research and Innovation Agency (BRIN). He obtained the doctoral program from Tsukuba University, Japan.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003eAhmad Ismed Yanuar: He is a member of research group of Carbon Cycle Management and Standardization of Life Cycle Assessment of BRIN. His research focus for Life Cycle Assessment (LCA) for palm oil product.\u0026nbsp;\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAvailability of data and materials\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eNot applicable.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eComputing interests\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare that they have no competing interests.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFunding\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was not funded.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAuthor\u0026rsquo;s contributions\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eIT, AT, YS conceived of the presented idea, developed the theory, implemented the research, and wrote the manuscript. IT, AN, AT, PL contributed to the design, and to the analysis of the results. All authors reviewed and approved the final manuscript.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAcknowledgements\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis research is supported by Research and Innovation Agency (BRIN) of Indonesia Republic from Visting Research Program. Furthermore, it is supported by Study Program of Engineering profession Program (PSPPI), Universitas Gadjah Mada. The authors would like to thank LPPM Universitas Amikom Purwokerto for supporting this research.\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eAlomar, K., Aysel, H. I., \u0026amp; Cai, X. (2023). Data Augmentation in Classification and Segmentation: A Survey and New Strategies. \u003cem\u003eJournal of Imaging\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(2). https://doi.org/10.3390/jimaging9020046\u003c/li\u003e\n\u003cli\u003eAmalia Yunia Rahmawati. (2020). \u003cem\u003eLife Cycle Assessment of Palm Oil at United Plantations Berhad 2022\u003c/em\u003e. \u003cem\u003eJuly\u003c/em\u003e, 1\u0026ndash;23.\u003c/li\u003e\n\u003cli\u003eAziira, A. H., Setiawan, N. A., \u0026amp; Soesanti, I. (2020). Generation of Synthetic Continuous Numerical Data Using Generative Adversarial Networks. \u003cem\u003eJournal of Physics: Conference Series\u003c/em\u003e, \u003cem\u003e1577\u003c/em\u003e(1). https://doi.org/10.1088/1742-6596/1577/1/012027\u003c/li\u003e\n\u003cli\u003eBedorf,A.(2023).\u003cem\u003eSynthcity\u003c/em\u003e.Ccaim.Cam.Ac.Uk.https://ccaim.cam.ac.uk/synthcity/\u003c/li\u003e\n\u003cli\u003eCheah, W. Y., Siti-Dina, R. P., Leng, S. T. K., Er, A. C., \u0026amp; Show, P. L. (2023). Circular bioeconomy in palm oil industry: Current practices and future perspectives. \u003cem\u003eEnvironmental Technology and Innovation\u003c/em\u003e, \u003cem\u003e30\u003c/em\u003e. https://doi.org/10.1016/j.eti.2023.103050\u003c/li\u003e\n\u003cli\u003eChen, R. J., Lu, M. Y., Chen, T. Y., Williamson, D. F. K., \u0026amp; Mahmood, F. (2021). Synthetic data in machine learning for medicine and healthcare. \u003cem\u003eNature Biomedical Engineering\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e(6), 493\u0026ndash;497. https://doi.org/10.1038/s41551-021-00751-8\u003c/li\u003e\n\u003cli\u003eChicco, D., Warrens, M. J., \u0026amp; Jurman, G. (2021). The coefficient of determination R-squared is more informative than SMAPE, MAE, MAPE, MSE and RMSE in regression analysis evaluation. \u003cem\u003ePeerJ Computer Science\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e, 1\u0026ndash;24. https://doi.org/10.7717/PEERJ-CS.623\u003c/li\u003e\n\u003cli\u003eDavila, M. H., Baldeon-Calisto, M., Murillo, J. J., Puente-Mejia, B., Navarrete, D., Riofr\u0026iacute;o, D., Per\u0026eacute;z, N., Ben\u0026iacute;tez, D. S., \u0026amp; Moyano, R. F. (2023). Analyzing the Effect of Basic Data Augmentation for COVID-19 Detection through a Fractional Factorial Experimental Design. \u003cem\u003eEmerging Science Journal\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(Special Issue), 1\u0026ndash;16. https://doi.org/10.28991/ESJ-2023-SPER-01\u003c/li\u003e\n\u003cli\u003eDenis J Murphy. (2019). \u003cem\u003eGrowing palm oil on former farmland cuts deforestation, CO₂ and biodiversity loss\u003c/em\u003e. Theconversation.Com. https://theconversation.com/growing-palm-oil-on-former-farmland-cuts-deforestation-co-and-biodiversity-loss-127312\u003c/li\u003e\n\u003cli\u003eDong, D., Yang, W., Sun, H., Kong, S., \u0026amp; Xu, H. (2022). Effects of Split Application of Urea on Greenhouse Gas and Ammonia Emissions From a Rainfed Maize Field in Northeast China. \u003cem\u003eFrontiers in Environmental Science\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(January), 1\u0026ndash;11. https://doi.org/10.3389/fenvs.2021.798383\u003c/li\u003e\n\u003cli\u003eEspino, M. T. M., De Ramos, R. M. Q., \u0026amp; Bellotindos, L. M. (2019). Life cycle assessment of the oil palm production in the Philippines: A cradle to gate approach. \u003cem\u003eNature Environment and Pollution Technology\u003c/em\u003e, \u003cem\u003e18\u003c/em\u003e(3), 709\u0026ndash;718.\u003c/li\u003e\n\u003cli\u003eEvers, S. (2020). \u003cem\u003ePalm oil: research shows that new plantations produce double the emissions of mature ones\u003c/em\u003e. Theconversation.Com. https://theconversation.com/palm-oil-research-shows-that-new-plantations-produce-double-the-emissions-of-mature-ones-130330\u003c/li\u003e\n\u003cli\u003eFadillah, R. Z., Irawan, A., Susanty, M., \u0026amp; Artikel, I. (2021). Data Augmentasi Untuk Mengatasi Keterbatasan Data Pada Model Penerjemah Bahasa Isyarat Indonesia (BISINDO). \u003cem\u003eJurnal Informatika\u003c/em\u003e, \u003cem\u003e8\u003c/em\u003e(2), 208\u0026ndash;214. https://ejournal.bsi.ac.id/ejurnal/index.php/ji/article/view/10768\u003c/li\u003e\n\u003cli\u003eFan, L. (n.d.). \u003cem\u003eA Survey of Differentially Private Generative Adversarial Networks\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eFawaz, H. I., Forestier, G., Weber, J., Idoumghar, L., \u0026amp; Muller, P.-A. (2018). \u003cem\u003eData augmentation using synthetic data for time series classification with deep residual networks\u003c/em\u003e. http://arxiv.org/abs/1808.02455\u003c/li\u003e\n\u003cli\u003eFonseca, J., \u0026amp; Bacao, F. (2023). Improving Active Learning Performance through the Use of Data Augmentation. \u003cem\u003eInternational Journal of Intelligent Systems\u003c/em\u003e, \u003cem\u003e2023\u003c/em\u003e, 1\u0026ndash;17. https://doi.org/10.1155/2023/7941878\u003c/li\u003e\n\u003cli\u003eGhoroghi, A., Rezgui, Y., Petri, I., \u0026amp; Beach, T. (2022). Advances in application of machine learning to life cycle assessment: a literature review. \u003cem\u003eInternational Journal of Life Cycle Assessment\u003c/em\u003e, \u003cem\u003e27\u003c/em\u003e(3), 433\u0026ndash;456. https://doi.org/10.1007/s11367-022-02030-3\u003c/li\u003e\n\u003cli\u003eGutowski, T. G. (2018). A Critique of Life Cycle Assessment ; Where Are the People ? \u003cem\u003eProcedia CIRP\u003c/em\u003e, \u003cem\u003e69\u003c/em\u003e(May), 11\u0026ndash;15. https://doi.org/10.1016/j.procir.2018.01.002\u003c/li\u003e\n\u003cli\u003eHeinz Stichnothe, F. S. (2011). LCA of two palm oil production systems. \u003cem\u003eBiomass and Bioenergy\u003c/em\u003e, \u003cem\u003e35\u003c/em\u003e(9). https://doi.org/https://doi.org/10.1016/j.biombioe.2011.06.001\u003c/li\u003e\n\u003cli\u003eHong, W. O. (2023). Review on Carbon Footprint of the Palm Oil Industry: Insights into Recent Developments. \u003cem\u003eInternational Journal of Sustainable Development and Planning\u003c/em\u003e, \u003cem\u003e18\u003c/em\u003e(2), 447\u0026ndash;455. https://doi.org/10.18280/ijsdp.180213\u003c/li\u003e\n\u003cli\u003eHorny\u0026aacute;k, O., \u0026amp; Iantovics, L. B. (2023). AdaBoost Algorithm Could Lead to Weak Results for Data with Certain Characteristics. \u003cem\u003eMathematics\u003c/em\u003e, \u003cem\u003e11\u003c/em\u003e(8). https://doi.org/10.3390/math11081801\u003c/li\u003e\n\u003cli\u003eHuang, D. (2021). \u003cem\u003eSynthetic data generation using Generative Adversarial Networks (GANs)\u003c/em\u003e. Medium.Com. https://medium.com/data-science-at-microsoft/synthetic-data-generation-using-generative-adversarial-networks-gans-part-1-47ecbf46b575\u003c/li\u003e\n\u003cli\u003eIwana, B. K., \u0026amp; Uchida, S. (2021). An empirical survey of data augmentation for time series classification with neural networks. In \u003cem\u003ePLoS ONE\u003c/em\u003e (Vol. 16, Issue 7 July). https://doi.org/10.1371/journal.pone.0254841\u003c/li\u003e\n\u003cli\u003eJadon, A., Patil, A., \u0026amp; Jadon, S. (2022). \u003cem\u003eA Comprehensive Survey of Regression Based Loss Functions for Time Series Forecasting\u003c/em\u003e. http://arxiv.org/abs/2211.02989\u003c/li\u003e\n\u003cli\u003eJyotsna Vadakkanmarveettil. (2021). \u003cem\u003eAdaBoost \u0026ndash; An Easy Guide (2021)\u003c/em\u003e. U-next.Com. https://u-next.com/blogs/data-science/adaboost/\u003c/li\u003e\n\u003cli\u003eKhan, A., Hwang, H., \u0026amp; Kim, H. S. (2021). Synthetic data augmentation and deep learning for the fault diagnosis of rotating machines. \u003cem\u003eMathematics\u003c/em\u003e, \u003cem\u003e9\u003c/em\u003e(18). https://doi.org/10.3390/math9182336\u003c/li\u003e\n\u003cli\u003eKhan, N., Kamaruddin, M. A., Ullah Sheikh, U., Zawawi, M. H., Yusup, Y., Bakht, M. P., \u0026amp; Mohamed Noor, N. (2022). Prediction of Oil Palm Yield Using Machine Learning in the Perspective of Fluctuating Weather and Soil Moisture Conditions: Evaluation of a Generic Workflow. \u003cem\u003ePlants\u003c/em\u003e, \u003cem\u003e11\u003c/em\u003e(13). https://doi.org/10.3390/plants11131697\u003c/li\u003e\n\u003cli\u003eKoyamparambath, A., Adibi, N., Szablewski, C., Adibi, S. A., \u0026amp; Sonnemann, G. (2022). Implementing Artificial Intelligence Techniques to Predict Environmental Impacts: Case of Construction Products. \u003cem\u003eSustainability (Switzerland)\u003c/em\u003e, \u003cem\u003e14\u003c/em\u003e(6), 1\u0026ndash;12. https://doi.org/10.3390/su14063699\u003c/li\u003e\n\u003cli\u003eKumar, S. (2020). \u003cem\u003eData splitting technique to fit any Machine Learning Model\u003c/em\u003e. Towardsdatascience.Com. https://towardsdatascience.com/data-splitting-technique-to-fit-any-machine-learning-model-c0d7f3f1c790\u003c/li\u003e\n\u003cli\u003eLamberti, A. (2023). \u003cem\u003eThe benefits and limitations of generating synthetic data\u003c/em\u003e. Syntheticus.Ai. https://syntheticus.ai/blog/the-benefits-and-limitations-of-generating-synthetic-data\u003c/li\u003e\n\u003cli\u003eLohr, S. L. (2018). \u003cem\u003eMeasuring Uncertainty with Multiple Sources of Data\u003c/em\u003e.\u003c/li\u003e\n\u003cli\u003eMeng, Y., \u0026amp; Noman, H. (2022). Predicting CO2 Emission Footprint Using AI through Machine Learning. \u003cem\u003eAtmosphere\u003c/em\u003e, \u003cem\u003e13\u003c/em\u003e(11), 1\u0026ndash;15. https://doi.org/10.3390/atmos13111871\u003c/li\u003e\n\u003cli\u003eMumuni, A., \u0026amp; Mumuni, F. (2022). Data augmentation: A comprehensive survey of modern approaches. \u003cem\u003eArray\u003c/em\u003e, \u003cem\u003e16\u003c/em\u003e(November), 100258. https://doi.org/10.1016/j.array.2022.100258\u003c/li\u003e\n\u003cli\u003eNations, U., \u0026amp; Programme, D. (2023). \u003cem\u003eIndonesia: Sustainable Palm Oil\u003c/em\u003e. Www.Undp.Org. https://www.undp.org/facs/indonesia-sustainable-palm-oil\u003c/li\u003e\n\u003cli\u003eOduque de Jesus, J., Oliveira-Esquerre, K., \u0026amp; Lima Medeiros, D. (2021). Integration of Artificial Intelligence and Life Cycle Assessment Methods. \u003cem\u003eIOP Conference Series: Materials Science and Engineering\u003c/em\u003e, \u003cem\u003e1196\u003c/em\u003e(1), 012028. https://doi.org/10.1088/1757-899x/1196/1/012028\u003c/li\u003e\n\u003cli\u003ePortolani, P., Vitali, A., Cornago, S., Rovelli, D., Brondi, C., Low, J. S. C., Ramakrishna, S., \u0026amp; Ballarino, A. (2022). Machine learning to forecast electricity hourly LCA impacts due to a dynamic electricity technology mix. \u003cem\u003eFrontiers in Sustainability\u003c/em\u003e, \u003cem\u003e3\u003c/em\u003e(Lci). https://doi.org/10.3389/frsus.2022.1037497\u003c/li\u003e\n\u003cli\u003eQian, Z. (n.d.). \u003cem\u003eSynthcity : facilitating innovative use cases of synthetic data in different data modalities\u003c/em\u003e. 1\u0026ndash;14.\u003c/li\u003e\n\u003cli\u003eRastogi, A., Garamendi, J. F., Fern, A., \u0026amp; Guitart, A. (2023). \u003cem\u003eSynthetic Data Generator For Adaptive Interventions In Global Health\u003c/em\u003e. 1\u0026ndash;9.\u003c/li\u003e\n\u003cli\u003eRozo, A., Moeyersons, J., Morales, J., Garcia van der Westen, R., Lijnen, L., Smeets, C., Jantzen, S., Monpellier, V., Ruttens, D., Van Hoof, C., Van Huffel, S., Groenendaal, W., \u0026amp; Varon, C. (2022). Data Augmentation and Transfer Learning for Data Quality Assessment in Respiratory Monitoring. \u003cem\u003eFrontiers in Bioengineering and Biotechnology\u003c/em\u003e, \u003cem\u003e10\u003c/em\u003e(February), 1\u0026ndash;14. https://doi.org/10.3389/fbioe.2022.806761\u003c/li\u003e\n\u003cli\u003eSari, D. W., Hidayat, F. N., \u0026amp; Abdul, I. (2021). Efficiency of land use in smallholder palm oil plantations in indonesia: A stochastic frontier approach. \u003cem\u003eForest and Society\u003c/em\u003e, \u003cem\u003e5\u003c/em\u003e(1), 75\u0026ndash;89. https://doi.org/10.24259/fs.v5i1.10912\u003c/li\u003e\n\u003cli\u003eSaswattecha, K., Cuevas Romero, M., Hein, L., Jawjit, W., \u0026amp; Kroeze, C. (2015). Non-CO2 greenhouse gas emissions from palm oil production in Thailand. \u003cem\u003eJournal of Integrative Environmental Sciences\u003c/em\u003e, \u003cem\u003e12\u003c/em\u003e, 67\u0026ndash;85. https://doi.org/10.1080/1943815X.2015.1110184\u003c/li\u003e\n\u003cli\u003eSchiappa, M. (2019). \u003cem\u003ePerformance Metrics in Machine Learning\u003c/em\u003e. Towardsdatascience.Com. https://towardsdatascience.com/metrics-ml-2563f9e47faa\u003c/li\u003e\n\u003cli\u003eSilitonga, P. D. P., Himawan, H., \u0026amp; Damanik, R. (2020). Forecasting acceptance of new students using double exponential smoothing method. \u003cem\u003eJournal of Critical Reviews\u003c/em\u003e, \u003cem\u003e7\u003c/em\u003e(1), 300\u0026ndash;305. https://doi.org/10.31838/jcr.07.01.57\u003c/li\u003e\n\u003cli\u003eSiregar, K., Ichwana, I., \u0026amp; Nasution, I. S. (2020). \u003cem\u003eImplementation of Life Cycle Assessment ( LCA ) for oil palm industry in Aceh Implementation of Life Cycle Assessment ( LCA ) for oil palm industry in Aceh Province , Indonesia\u003c/em\u003e. \u003cem\u003eOctober\u003c/em\u003e. https://doi.org/10.1088/1755-1315/542/1/012046\u003c/li\u003e\n\u003cli\u003eStichnothe, H., \u0026amp; Bessou, C. (2017). Challenges for Life Cycle Assessment Of Palm Oil Production System. \u003cem\u003eIndonesian Journal of Life Cycle Assessment and Sustainability\u003c/em\u003e, \u003cem\u003e1\u003c/em\u003e(2), 1\u0026ndash;9. https://doi.org/10.52394/ijolcas.v1i2.28\u003c/li\u003e\n\u003cli\u003eUnion, I., Conservation, F. O. R., \u0026amp; Nature, O. F. (2018). Oil palm and biodiversity: a situation analysis by the IUCN Oil Palm Task Force. In \u003cem\u003eOil palm and biodiversity: a situation analysis by the IUCN Oil Palm Task Force\u003c/em\u003e. https://doi.org/10.2305/iucn.ch.2018.11.en\u003c/li\u003e\n\u003cli\u003eYan. (2022). \u003cem\u003e10 largest oil palm plantations in Indonesia\u003c/em\u003e. Indonesiabusinesspost.Com. https://indonesiabusinesspost.com/insider/10-largest-oil-palm-plantations-in-indonesia/\u003c/li\u003e\n\u003cli\u003eYoon, J., Drumright, L. N., \u0026amp; Van Der Schaar, M. (2020). Anonymization through data synthesis using generative adversarial networks (ADS-GAN). \u003cem\u003eIEEE Journal of Biomedical and Health Informatics\u003c/em\u003e, \u003cem\u003e24\u003c/em\u003e(8), 2378\u0026ndash;2388. https://doi.org/10.1109/JBHI.2020.2980262\u003c/li\u003e\n\u003cli\u003eYusniati, Parinduri, L., \u0026amp; Sulaiman, O. K. (2018). Biomass analysis at palm oil factory as an electric power plant. \u003cem\u003eJournal of Physics: Conference Series\u003c/em\u003e, \u003cem\u003e1007\u003c/em\u003e(1). https://doi.org/10.1088/1742-6596/1007/1/012053\u003cstrong\u003e\u003cbr\u003e \u003c/strong\u003e\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"ADS-GAN, Carbon Footprint, Data Augmentation, Deep Learning, Environmental Impacts, LCA, Regression, Palm Oil","lastPublishedDoi":"10.21203/rs.3.rs-3675682/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-3675682/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eLife Cycle Assessment (LCA) is a widely used methodology for quantifying the environmental impacts of products, including the carbon footprint. However, conducting LCA studies for complex systems, such as the palm oil industry in Indonesia, can be challenging due to limited data availability. This study proposes a novel approach called the Anonymization Through Data Synthesis (ADS-GAN) based on a deep learning approach to augment carbon footprint data for LCA assessments of palm oil products in Indonesia. This approach addresses the data size limitation and enhances the comprehensiveness of carbon footprint assessments. An original dataset comprising information on various palm oil life cycle stages, including plantation operations, milling, refining, transportation, and waste management. The number of original data is 195 obtained from the Sustainable Production Systems and Life Assessment Research Centre of Indonesia's National Innovation Research Agency (BRIN). To measure the performance of prediction accuracy, this study used regression models: Random Forest Regressor (RFR), Gradient Boosting Regressor (GBR), and Adaptive Boosting Regressor (ABR). The best-augmented data size is 1000 data. In addition, the best algorithm is the Random Forest Regressor, resulting in the MAE, MSE, and MSLE values are 0.0031, 6.127072889081567e-05, and 5.838479552074619e-05 respectively. The proposed ADS-GAN offers a valuable tool for LCA practitioners and decision-makers in the palm oil industry to conduct more accurate and comprehensive carbon footprint assessments. By augmenting the dataset, this technique enables a better understanding of the environmental impacts of palm oil products, facilitating informed decision-making and the development of sustainable practices.\u003c/p\u003e","manuscriptTitle":"Optimizing Sustainability: A Deep Learning Approach on Data Augmentation of Indonesia Palm Oil Products Emission","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2023-12-06 04:14:39","doi":"10.21203/rs.3.rs-3675682/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"87385306-ea46-47e2-8690-73628e5c9151","owner":[],"postedDate":"December 6th, 2023","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-03-15T06:21:49+00:00","versionOfRecord":[],"versionCreatedAt":"2023-12-06 04:14:39","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-3675682","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-3675682","identity":"rs-3675682","version":["v1"]},"buildId":"rHA-KDH7Qsr4HCuvH75dn","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.