Data-driven framework for the techno-economic assessment of sustainable aviation fuel from pyrolysis. | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Data-driven framework for the techno-economic assessment of sustainable aviation fuel from pyrolysis. Jude Okolie, Keon Moradi, Brooke Rogachuk, Bala Nagaraju Narra, and 3 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4595354/v1 This work is licensed under a CC BY 4.0 License Status: Published Journal Publication published 02 Dec, 2024 Read the published version in BioEnergy Research → Version 1 posted 5 You are reading this latest preprint version Abstract The aviation sector plays a crucial role in quickly moving people and goods around the world. It also greatly helps in the economic growth and social integration of countries. As the industry continues to experience rapid growth, there is a tendency for an increase in emissions associated with the industry. Sustainable aviation fuel (SAF) presents a way to reduce the environmental effects of the aviation industry by providing a clean-burning, renewable substitute for conventional jet fuel. SAF can be produced from diverse processes and feedstocks. Fast pyrolysis (FP) is a promising thermochemical process for SAF production due to its advantages including low-cost feedstocks, faster reaction times, and simpler technology, making it more cost-effective and scalable compared to other thermochemical processes. However, the preliminary estimation of the economic viability of FP for SAF production is complex and tedious requiring detailed process models and several assumptions. Moreover, the relationship between the feedstock properties and the minimum selling price of fuel (MSP) is often challenging to estimate. To address these challenges, the present study developed a data-driven framework for preliminary estimation of the MSP of SAF from FP. The target output feature is MSP. To enhance model accuracy and predictions, synthetic data was created using Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), and hyperparameter optimization was conducted using Grid Search. Five surrogate models were evaluated: linear regression, gradient boost regression (GBR), random forest (RF), extreme boost regression (XGBoost), and Elastic net. GBR and RF showed the most promise based on metrics like R², RMSE, and MAE for both original and synthetic datasets. Specifically, GBR achieved a Train R² of 0.9999 and a Test R² of 0.9277, while RF had Train and Test R² scores of 0.9789 and 0.9255, respectively. The use of data from the VAE notably enhanced model accuracy. Additionally, a publicly available GUI has been developed for researchers to estimate the MSP of Sustainable Aviation Fuel (SAF) based on biomass properties, plant capacity, and location. Pyrolysis Sustainable Aviation fuel Machine learning Techno-economic Biofuels Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 Figure 7 Figure 8 Figure 9 Figure 10 Figure 11 Figure 12 Figure 13 1. Introduction The aviation industry has contributed immensely to the immediate transportation of people and goods across the continents. It also provides a significant contribution towards the economic growth and social integration of a nation. According to the International Air Transport Association (IATA), the aviation sector is on the trajectory of exponential growth of global passenger counts expected to double over two decades, reaching 8.2 billion by 2030 [1]. This projection underscores the burdening demand for air travel, stimulated by population growth, affordability of airfare, and improved standards of living. It should be mentioned that the growth in the aviation industry is also accompanied by severe environmental impacts. The aviation sector contributes significantly to global greenhouse gas emissions. These emissions include carbon dioxide and other gases that have worsened the greenhouse effect [2]. Since 2009, fuel use by commercial airlines has been climbing annually, hitting an all-time high of 359 billion liters in 2019 [3]. However, even with the travel limits imposed in 2020 to curb the spread of COVID-19, which reduced air travel significantly, commercial airlines still used around 101 billion liters of aviation fuel that year [3]. Moreover, the aviation industry accounts for about 2% of total carbon dioxide emissions [3]. Therefore, there is an urgent need to decarbonize the aviation industry to ensure mitigation of greenhouse gas emissions while keeping up with its growth. Sustainable aviation fuel (SAF) presents a way to reduce the environmental effects of the aviation industry by providing a clean-burning, renewable substitute for conventional jet fuel. SAF can be produced from diverse processes and feedstocks. These processes are categorized into thermochemical, biological, and integrated processes [1]. The advantages and limitations of each SAF production pathway have been documented in a previous study [4]. Among the production methods, thermochemical processes are advantageous due to their high efficiency and the ability to convert a wide range of biomass feedstocks including waste cooking oil, municipal waste, algae, and lignocellulosic biomass [5]. Pyrolysis is a thermochemical process that offers a low-carbon pathway for SAF production. During fast pyrolysis (FP), biogenic waste is heated in an air-free environment facilitating its breakdown into bio-oil, gases, and solid char [6]. FP operates at high temperatures between 400–600°C and short residence times, efficiently maximizing the production of liquid bio-oil, which can then be refined into biofuels including SAF via hydro-cracking and isomerization [1]. It should be mentioned that using bio-oil directly as fuel is not feasible due to its high oxygen content and characteristics such as thermal instability, corrosiveness, and low energy density [7]. To meet Sustainable Aviation Fuel (SAF) standards and ensure compatibility with current aircraft systems, further refinement of the bio-oil is necessary. The production of SAF from FP and subsequent hydroprocessing is advantageous for several reasons including low-cost feedstocks, faster reaction times, and simpler technology, making it more cost-effective and scalable compared to other thermochemical processes [8], [9]. A recent study showed that while FP-hydroprocessing and gasification-Fischer Tropsch (GFT) processes are two economically feasible thermochemical processes for producing SAF, the latter produced lower CO 2 emissions compared to GFT and fossil-based aviation fuel [1]. Although promising, FP pathways for SAF face several challenges related to high oxygen content and instability of the bio-oil produced. These require significant refining to meet fuel standards, catalysts deactivation during hydroprocessing as well as detailed understanding of the relationship between the physicochemical properties of the biogenic waste and bio-oil properties. To scale up FP- hydroprocessing for commercial SAF production while addressing the challenges mentioned earlier, techno-economic analysis (TEA) studies are required. TEA is used to assess the economic viability of a new process or a full-scale plant, especially in conceptual designs where there is no information about a similar industrial facility [10]. In the case of FP-hydroprocessing, TEA is used to assess its economic viability and how several factors influence the profitability of the process at varying capacities. However, performing TEA analysis requires the development of rigorous process models, large data collection from literature and several assumptions related to equipment purchase cost (EPC), biomass properties and plant capacity. Moreover, the relationship between the feedstock properties and the minimum selling price of fuel (MSP) is often challenging to estimate. Several studies have reported the TEA of SAF in literature. For instance, Rogachuk and Okolie [1], performed a detailed economic evaluation of GFT and FP for SAF production, however, they focused on waste tires, and it is not clear if their economic model could be applicable to other feedstocks. In another study, Michailos and Bridgewater[11] compared the economic feasibility of producing SAF from three different bio-oil upgrading routes. Their results do not explain how plant capacity and varying location as well as biomass properties influence the MSP. Recently, Saeed et al.[12] also explored the economic feasibility of SAF production via chemical looping technologies. Their study showed that an economically feasible SAF with negative emission can be produced by integrated chemical looping technology. Despite the ample amount of studies on the TEA of SAF, there are several literature gaps. Integrating feedstock properties including proximate and ultimate analysis has seldom been reported. Additionally, studies that developed an economic model while considering different production capacities and different regions are limited. Furthermore, the complexity and assumptions inherent in TEA studies have not been addressed. Therein lies the motivation of the present study. The proposed study integrated process simulation results with experimental data to develop a data-driven machine learning (ML) framework for easy prediction of the MSP of SAF from FP. Data-driven ML methods have been used in biofuel technology to study the relationship between process conditions and yield by several researchers [13], [14], [15], [16]. In addition, ML methods are employed for catalyst optimization, biofuel yield prediction and char predictions [13] [17]. However, its application as a complementary tool to TEA is scarcely reported. Therefore, the key objectives of the present study are as follows: combine process models with experimental data from the literature to develop a robust economic analysis database. Following which two deep learning methods would be used to generate additional synthetic data for ML models. Develop relationships between biomass properties, production capacity and regions with the MSP of SAF. Finally, a publicly available graphical user interface (GUI) will be developed to help researchers and industry practices in the preliminary estimation of the MSP of SAF from FP. 2. Methodology The overall methodology adopted in this study is presented in Fig. 1 . The steps include the development of a robust dataset comprising process model development in Aspen Plus, experiment data collection from literature and deep learning methods for synthetic data generation. Details of each method are meticulously presented in subsequent sections. 2.1 Data collection A biomass database was created by collecting the proximate and ultimate analysis of 31 different biomass materials. These biomasses represent several categories of biogenic wastes including energy crops, animal manure, agricultural wastes, woody biomass, sewage sludge and food wastes. The proximate, ultimate, and compositional analysis of the biogenic waste materials were compiled from the literature and Phyllis database [18]. The dataset is comprised of 13 input features including Carbon %, Hydrogen %, Nitrogen %, Oxygen %, Sulfur%, Volatile Matter %, Fixed Carbon Content %, Ash content %, Cellulose %, Hemicellulose %, Lignin content %, Location, and Plant Capacity (kg/hr), and one target output feature MSP. It should be noted that all proximate and ultimate analysis data are either received on a dry basis or converted to a dry basis and normalized to sum to 100%. The study explores three different locations: the US, China, and the UK. Additionally, two different plant capacities, 25,000 kg/hr and 50,000 kg/hr, are examined. Location factors of 0.66, 1.0, and 1.04 were used to represent China, the US, and the UK, respectively [17]. Based on two different plant capacities and three locations, a total of 186 data points were collected, representing biomass with varying proximate, ultimate, and compositional analyses, as well as differences in plant capacity and location. The entire dataset is presented in Table S1 of the supplementary materials while Table 1 shows the summary statistics of the dataset. Figure 2 presents a histogram plot that shows the distribution of various features in the dataset. In each plot, the x-axis represents a continuous variable, while the y-axis corresponds to the frequency of observations within the bins defined by the x-axis. The shapes of the distributions suggest the variability of each biomass property within the dataset. The histogram shows the distribution of the MSP for different biomass materials indicating the range of prices from the use of different feedstocks, location and production capacities. Table 1 Summary statistics of the original dataset. Note all data are presented on a dry basis while the proximate, ultimate and compositional analyses are expressed in percentages. The MSP is the only output feature in the Table. Location is not presented in the table. Features Count Mean Std Minimum value 25th percentile 50th percentile 75th percentile Maximum value C (wt.%) 186 49.51 4.00 40.02 47.01 48.87 52.85 60.46 H (wt.%) 186 6.42 1.04 5.32 5.98 6.18 6.50 10.76 N (wt.%) 186 1.79 1.95 0.10 0.49 1.17 2.19 8.21 O (wt.%) 186 42.06 5.06 27.36 39.37 43.28 45.56 52.86 S (wt.%) 186 0.22 0.39 0 0 0 0.31 1.55 Volatile Matter (VM) (wt.%) 186 75.47 9.17 49.36 71.70 76.56 81.80 94.16 Ash content (Ash) (wt.%) 186 8.70 9.88 0.40 1.70 4.13 14.80 42.02 Fixed carbon (FC) (wt.%) 186 15.84 6.38 3.37 11.83 16.61 20.61 26.60 Cellulose content (Cel) (wt.%) 186 33.69 9.96 6.92 25.19 35.00 42.60 49.88 Hemicellulose content (Hem) (wt.%) 186 26.99 9.74 12.80 22.00 25.00 30.00 55.42 Lignin content (Lig) (wt.%) 186 21.84 10.99 1.60 14.50 22.40 26.80 53.40 Plant capacity. (kg/hr) 186 37500.00 12533.74 25000.00 25000.00 37500.00 50000.00 50000.00 Minimum selling price (MSP) US$/L 186 0.79 0.24 0.40 0.64 0.73 0.99 1.59 A total number of 186 data points were used as shown in Table 1 . Furthermore, the range of proximate, ultimate, and compositional analyses was carefully selected to capture different varieties of biomass feedstock. That way the range of biomass materials properties can develop data-driven models that can predict the MSP of SAF produced from feedstocks within and outside the data range. 2.2 Process model development and economic evaluation A model for fast pyrolysis was developed using the advanced features of Aspen Plus Version 11, a comprehensive process modeling software licensed through the University of Oklahoma. The simulation was meticulously conducted under specific conditions, namely steady-state and isothermal, to ensure the utmost accuracy and consistency in the results. Details of the FP-hydroprocessing model including the underlying assumptions have been meticulously defined in our previous study [1]. The Aspen Plus model was refined to accommodate different feedstock based on their proximate and ultimate analysis. A total of 31 different simulations were developed for various feedstocks, each with unique compositions, to achieve the most accurate model possible. During FP, biomass was defined as non-conventional solids by using a calculator block incorporated into the RYield block, and its function is executed using a programmed FORTRAN subroutine statement. Details of the non-conventional reactor modelling in Aspen Plus can be found elsewhere [19]. An economic assessment was performed to determine the MSP of SAF following a generalized approach reported in our previous study [1][20]. This approach involves the appraisal of the Capital Expenditure (CapEx) and the Operational Expenditure (OpEx). The CaPEx calculations are based on the EPC and the installation costs. In contrast, the OpEx was estimated to include both Variable Operating Cost (VOC) and Fixed Operating Cost (FOC), considering components of labor costs and EPC. Details of CapEx and OpEx calculations including the assumptions have been documented in our recent study [1]. Since the proposed plant location and capacity varies, the economic model was adjusted to account for the plant location and salary of laborers. For instance, the labor cost for the U.S. is different from China and these costs were based on the minimum wages in the respective countries. Furthermore, the EPC was estimated using the Aspen Plus Economic Analyzer while the cost was adjusted to the year 2023 using the chemical engineering plant cost index (CEPCI). The MSP of SAF which is defined as the lowest price at which the fuel can be sold to cover all costs of production and ensure economic viability. The MSP is estimated based on the sum of capital and operating expenses over the project's lifetime. 2.3 Data Preprocessing and Synthetic Data Generation The dataset includes a categorical variable, 'Location,' which was transformed via one-hot encoding prior to training the Generative Adversarial Network (GAN) model. This transformation resulted in high-dimensional data with binary values assigned to the location feature. The dataset was partitioned into training and testing sets (X_train, X_test, Y_train, Y_test) using the train_test_split() function, configured with a test size of 0.2 and a random state set to 42. Numerical data within the dataset was normalized using the MinMaxScaler() function before model training. In order to improve the performance of surrogate models for the prediction of MSP of SAF, synthetic data generation methods based on deep learning approaches were applied. Two synthetic data generation methods including the Generative Adversarial Network (GAN) and Variational Autoencoder (VAE) were used and compared. GAN and VAE models were utilized to generate 5,000 records. The generated data underwent a cleaning process to ensure it conformed to the characteristics of the real data, which involved removing any data points that fell outside the predefined range. The z-score was calculated for the features in the synthetic data to identify and remove outliers. Additionally, bounds were set based on the minimum and maximum values observed in the original data, and these bounds were used to further refine the synthetic data. The GAN is a neural network algorithm for generative modeling. GANs consist of two neural networks, the generator G and the discriminator D, which are trained simultaneously through a competitive process [21]. The generator “G” creates samples that are intended to come from the same distribution as the training data, while the discriminator D evaluates them. Details of the GANs' operating mechanism are shown in Fig. 3 . The main objective of the generator is to produce data that is indistinguishable from real data. It takes a random noise vector z (from a latent space) as input and generates samples that resemble the training data. The discriminator's job is to distinguish between actual data samples and fake samples produced by the generator. It outputs a scalar representing the probability that a given sample is real. GAN is trained through a min-max game where D and G have competing objectives. The GAN model comprises a generator and a discriminator. The generator features a dense network with 32 units and utilizes LeakyReLU as the activation function in its hidden layers to introduce nonlinearity, enabling the model to tackle complex patterns effectively. The output layer of the generator is bifurcated: the first segment employs a tanh activation function to scale outputs to a range of [-1, 1], while the second segment consists of three units using a softmax function, which facilitates the generation of one-hot encoded data. The dimensionality of the latent space ( latent_dim ) and the number of units in the hidden layers are adjustable hyperparameters, allowing customization according to the complexity of the dataset. Compared to GANs, the VAEs are a type of generative model that uses neural networks to encode data into a latent space and then decode it back to the original space. Unlike traditional autoencoders, VAEs are probabilistic models that produce a distribution over the latent space (Fig. 3 ). The encoder part of a VAE maps input data x to a latent representation z. However, instead of encoding an input as a single point, it encodes it as a distribution over the latent space. The decoder part then samples from this distribution to generate a sample ′x′ that approximates the original input x. The loss function for VAEs has two terms: the reconstruction loss and the KL divergence. The reconstruction loss ensures the decoded samples match the original inputs, while the KL divergence regularizes the encoder by comparing the distribution of latent variables with a prior (often a Gaussian distribution). 2.4 Surrogate models development Five different ML models including linear regression, gradient boost regression (GBR), random forest (RF), extreme boost regression (XGBoost) and Elastic net were selected for the development of surrogate models. These models were selected due to their promising results in biofuel yield optimization, catalyst design, product yield prediction and new materials development [13], [15], [16], [22]. GBR is an ensemble machine learning model that trains using a boosting strategy [23]. A gradient-boosting leaf is obtained by taking the average of the output values and using them as an initial value to calculate the initial errors. Then, the GBR model begins constructing a decision tree, which is a binary tree used by ML models to split input data into smaller subgroups, and this new tree is assigned a learning rate which it uses to calculate new prediction values from the previous prediction value and error. This process occurs recursively until the maximum depth limit has been passed or until a satisfactory error value has been obtained. Various hyperparameters of the GBR model need to be tuned to properly optimize the predictive algorithm. The RF Regressor is also an ensemble learning algorithm that utilizes random feature bagging and selection strategies to output an average of the base model predictions. It simultaneously trains several base decision trees on random subsets of the original data, and for qualitative data, it averages the results of those decision trees for a final prediction output [24]. Various hyperparameters must be passed to RF Regressors to optimize their performance, including the number of decision trees that can be running in parallel, the maximum number of variables in each tree, and the total number of decision trees to be constructed [25]. The XGBoost represents a regularized regression method that constructs multiple decision trees sequentially, with each new tree aiming to correct the errors made by the previous ones. This process helps in minimizing the overall prediction error. XGBoost is well-regarded for its ability to handle large datasets efficiently, its versatility in accommodating various types of predictive modeling problems, and its robustness in dealing with a wide range of data types. ElasticNet is a regression technique that combines the properties of both ridge and lasso regression methods. It is particularly useful when dealing with highly correlated data or when the number of predictors exceeds the number of observations. ElasticNet incorporates penalties from both L1 and L2 regularization, allowing it to shrink some coefficients to zero (like lasso) for variable selection, while also evenly distributing the penalty among the remaining coefficients (like ridge). 2.5 Model evaluation and hyperparameter optimization In order to evaluate the performance of each ML model the Regression coefficient (R 2 ), root-mean-square error (RMSE) and mean absolute error (MAE) were used as quantifiable metrics. Details of these criteria as well as the underlying equations can be found elsewhere [26]. It should be mentioned that these metrics were selected because they provide a comprehensive view of the surrogate model performance from different viewpoints. R² indicates how much of the variance in the dependent variable is explained by the predictors, offering insight into the overall effectiveness of the model. RMSE provides a scale-dependent measure of the average magnitude of the prediction errors, penalizing larger errors more severely due to its squaring of residuals. MAE, on the other hand, offers a straightforward, scale-dependent average error magnitude that is easy to interpret and less sensitive to outliers than RMSE. The combination of these metrics provides and holistic and balanced understanding of model accuracy, consistency, and the nature of errors. A K-fold cross-validation method with 5 folds was employed to validate the performance of each model which offers robustness and computational efficiency. The cross-validation method partitions the data into complementary subsets conducting the analyses on one subset called the training set and validating the analyses on the other subset. In our methodology data was shuffled to ensure randomness in the cross-validation splits, controlled by a random state. After developing the ML models, hyperparameter optimization was performed to help identify the optimized hyperparameter combinations that produce the most promising accuracy. The Grid search approach was used during hyperparameter optimization. This approach has been proven advantageous in our previous study [21]. In Grid Search, all possible combinations of the specified hyperparameters are systematically explored. A set of hyperparameters was initially defined during the training process as shown in Table 2 . Three distinct values were initially selected for each hyperparameter to ensure a wide range of options for the tuning process to explore. This approach helps to assess how sensitive the model's performance is to different settings of each hyperparameter, thus improving the ability to find the optimal configuration that yields the best results. For Linear Regression, no hyperparameters were selected to tune. For RF, n_estimators (100, 200, 300), max_depth (None, 5, 10), and min_samples_split (2, 5, 10) were selected. For GBR, n_estimators (100, 200, 300), learning_rate (0.01, 0.1, 0.5), and max_depth (3, 5, 7) were selected. For Elastic Net, hyperparameters such as alpha (0.1, 0.5, 1.0), and l1_ratio (0.1, 0.5, 0.9) were chosen. For XGBoost, hyperparameters such as n_estimators (100, 200, 300), learning_rate (0.01, 0.1, 0.5), and max_depth (3, 5, 7) were selected. Table 2 Hyperparameters initially selected during the ML model development. Model Hyperparameters Selected Linear Regression None Random Forest n_estimators 100, 200, 300 max_depth None, 5, 10 min_samples_split 2, 5, 10 Gradient Boosting n_estimators 100, 200, 300 learning_rate 0.01, 0.1, 0.5 max_depth 3, 5, 7 Elastic Net alpha 0.1, 0.5, 1.0 l1_ratio 0.1, 0.5, 0.9 XGBoost n_estimators 100, 200, 300 learning_rate 0.01, 0.1, 0.5 max_depth 3,5,7 3. Results and Discussions 3.1 Parametric study The influence of proximate, ultimate, and compositional analysis of biomass on the MSP of SAF for different locations is presented in Figs. 4 and 5 . The results show that geographical locations influence the MSP of SAF. Prices in the US and UK are higher compared to prices in China. Furthermore, the figures show that a change in the biomass properties does not have a pronounced impact on the MSP compared to the geographical location which had a considerable impact. The higher MSP of SAF in the UK and US compared to China can be attributed to several factors. Firstly, the regulatory environment in the UK and US often imposes stricter sustainability and environmental standards, which can increase production costs. These countries may also have higher labor and operational costs. Additionally, the UK and US might have more developed markets for SAF, with policies supporting its use, such as tax incentives or mandatory blending requirements, which can drive up prices due to increased demand. Conversely, China might benefit from lower production costs, government subsidies, and a rapidly scaling biofuels industry, which can help keep prices lower. These factors combined explain the regional differences in the minimum selling price of SAF. The influence of location on the MSP of biofuels has been reported by some researchers [17], [27]. Rogers et al.[17] reported a considerable difference in the Levelized cost of hydrogen (LCOH) produced from hydrothermal gasification at different locations. The authors attributed this observation to the varying capital and operating costs in different countries. The production capacity has a significant impact on the MSP of SAF as shown in Fig. 6 . A larger production capacity led to a decline in the MSP. The trend follows the economies of scale, where the average cost of producing SAF decreases as the quantity of output increases. This reduction in unit cost is primarily due to the more efficient use of resources, better procurement terms for bulk materials, and the spreading of fixed costs (like capital investments and administrative expenses) over a larger volume of production. Consequently, producers with higher capacities are often able to offer SAF at more competitive prices. The MSP of SAF obtained from biomass pyrolysis is compared to other SAF production technologies including gasification- Fischer Tropsch process (GFT), alcohol to jet fuel (ATJ), hydroprocessing of esters and fatty acids (HEFA) and direct sugar to hydrocarbons (DSHC). The results presented in Fig. 7 show varying MSPs for different technologies. It should be noted that the results were extracted from various published studies that use different feedstocks, locations and the year of study differs. Additionally, the prices were also computed at different plant capacities which could significantly influence the MSP. Regardless the data presented in Fig. 7 provides a basis for comparing the economic viability of various SAF production pathways. As seen in Fig. 7 , GFT has the lowest range, starting close to 0.40 US $ /L up to about 0.91 US $ /L. On the contrary, ATJ shows a price range between 0.75–1.38 US $ /L. DSHC shows the most significant variation, with prices starting near 2.21 US $ /L and reaching above 2.5US $ /L. The price ranges for each technology can be attributed to factors such as feedstock costs, production efficiency, technological maturity, and scale of production. For instance, DSHC relies on sugars as the primary feedstock, which can be expensive, especially if sourced from food-grade crops. The cost can be influenced by agricultural market prices, which are subject to volatility. Moreover, converting sugars to hydrocarbons can involve complex chemical processes characterized by low product yield [28], [30]. The low yield could significantly influence the process economics as documented in a previous study [28]. The lower cost of HEFA compared to other SAF production pathways could be attributed to several reasons. HEFA technology commonly uses waste oils, animal fats, and non-food vegetable oils, which can be less expensive and more readily available than the food-grade sugars required for DSHC. Furthermore, the hydroprocessing step in HEFA technology is well understood and can be highly efficient, leading to higher yields of SAF from the feedstock and potentially lower production costs. HEFA processes can produce valuable co-products such as propane and naphtha, which can be sold to offset the cost of fuel production. 3.2 Synthetic Dataset Generation About 5000 synthetic values based on the original dataset using GAN and VAE were generated. These datasets mimic the distribution of the original data. However, after applying data cleaning techniques on the synthetic data including the removal of out-of-range or unreasonable values, the dataset was reduced to 4862 for VAE and 346 records from GAN. A box plot comparing the original dataset with GAN and VAE datasets is presented in Fig. 8 while Tables 3 and 4 show the summary statistics of the dataset. The GAN synthetic data has generally higher variability as shown by larger standard deviations for many features compared to the original data. This might reflect greater diversity but also indicates that the synthetic data might be capturing additional noise or generating unrealistic values such as negative values for Cel, Hem, Lig, and MSP. The VAE-generated data exhibits mean values that closely approximate the original data across most features but show significantly less variability. While this might suggest that the VAE is better at capturing the central tendency of the data, the reduced variability could imply that it is not capturing the full diversity present in the original dataset. The analysis of various components in the datasets reveals distinct differences between the original data and the synthetic data generated from both GAN and VAE. For Carbon Content (C), while the original dataset shows a broader spread potentially inclusive of outliers, the GAN synthetic dataset demonstrates an even wider distribution, contrasting with the VAE's much narrower interquartile range (IQR). This trend of narrower variability with VAE-generated data is also evident in the distribution of Hydrogen Content (H), where the GAN distribution is wider than the original, and the VAEs are significantly narrower. For elements such as Nitrogen (N), Oxygen (O), Sulfur (S), and Volatile Matter (VM), the VAE synthetic data maintains a noticeably narrower spread compared to both the original and GAN datasets. In terms of Ash and Fixed Carbon (FC), the original dataset displays a widespread, with the GAN data showing a similar distribution for Ash but a wider range for FC. Conversely, the VAE data for these features appears more condensed, indicating less variability. Considering the compositional analysis of biomass like Cellulose (Cel), Hemicellulose (Hem), and Lignin (Lig), the GAN data reveals extreme outliers with a broader distribution for Hem and Lig, while VAE-generated data avoids such extremes and aligns more closely with the original dataset's distribution. Plant Capacity analysis shows the original data concentrated around a middle capacity range, whereas the GAN data is more dispersed, and VAE data is markedly more compact. For the MSP, the original data presents a slightly skewed distribution with higher variability, which is more pronounced in the GAN data. The VAE, on the other hand, shows a tight distribution, indicating a significant reduction in variation in the synthetic selling prices. Regarding the location features (China, UK, US), which are categorical and represented through one-hot encoding, both GAN and VAE maintain the original distribution without introducing outliers, as expected for binary categorical data. This analysis underscores the varying effectiveness and characteristics of GAN and VAE in replicating and altering the statistical properties of the original dataset. Table 3 Summary statistics of the GAN synthetic dataset. Features Count Mean Std Minimum value 25th percentile 50th percentile 75th percentile Maximum value C (wt.%) 346.00 52.85 4.93 40.34 49.19 53.69 56.80 60.23 H (wt.%) 346.00 7.00 1.19 5.33 5.96 6.86 7.84 10.19 N (wt.%) 346.00 3.93 1.86 0.14 2.49 3.97 5.35 7.95 O (wt.%) 346.00 47.33 4.56 30.06 44.82 48.89 50.86 52.77 S (wt.%) 346.00 0.41 0.28 0.00 0.18 0.36 0.60 1.24 Volatile Matter (VM) (wt.%) 346.00 85.68 6.02 62.78 81.74 87.00 90.52 93.89 Ash content (Ash) (wt.%) 346.00 9.71 6.80 0.40 4.44 8.03 13.91 29.28 Fixed carbon (FC) (wt.%) 346.00 22.83 2.96 10.98 21.37 23.56 25.13 26.50 Cellulose content (Cel) (wt.%) 346.00 31.99 8.89 7.48 26.56 32.12 39.51 48.53 Hemicellulose content (Hem) (wt.%) 346.00 40.27 12.21 12.87 30.53 43.36 50.99 55.28 Lignin content (Lig) (wt.%) 346.00 26.33 14.22 1.74 15.15 24.42 38.22 53.03 Plant capacity (kg/hr) 346.00 41222.49 5739.17 25495.57 37100.14 41958.52 46114.67 49672.01 MSP (US$/L) 346.00 0.92 0.24 0.40 0.75 0.93 1.10 1.53 Table 4 Summary statistics of the VAE synthetic dataset. Features Count Mean Std Minimum value 25th percentile 50th percentile 75th percentile Maximum value C (wt.%) 4862.00 49.61 0.32 48.48 49.41 49.65 49.81 50.31 H (wt.%) 4862.00 6.47 0.29 5.59 6.26 6.50 6.72 6.95 N (wt.%) 4862.00 1.99 0.40 0.74 1.72 2.04 2.31 2.58 O (wt.%) 4862.00 41.94 0.38 41.37 41.63 41.86 42.18 43.22 S (wt.%) 4862.00 0.31 0.08 0.09 0.25 0.32 0.38 0.45 Volatile Matter (VM) (wt.%) 4862.00 75.34 0.75 74.18 74.74 75.24 75.81 77.86 Ash content (Ash) (wt.%) 4862.00 9.34 2.52 2.17 7.52 9.59 11.39 13.41 Fixed carbon (FC) (wt.%) 4862.00 15.76 0.42 15.03 15.45 15.67 15.99 17.21 Cellulose content (Cel) (wt.%) 4862.00 33.66 1.46 31.46 32.51 33.35 34.54 38.49 Hemicellulose content (Hem) (wt.%) 4862.00 27.15 1.52 22.38 26.12 27.36 28.37 29.56 Lignin content (Lig) (wt.%) 4862.00 22.29 1.01 18.34 21.69 22.43 23.07 23.87 Plant capacity (kg/hr) 4862.00 37464.86 395.08 35993.07 37185.51 37571.23 37773.73 38231.48 MSP (US$/L) 4862.00 0.81 0.05 0.65 0.78 0.81 0.84 0.88 3.3 Evaluation of Surrogate Model Performance Using the original and synthetic dataset five different surrogate models were developed for the prediction of the MSP of SAF from biomass pyrolysis. A K-fold cross-validation method with 5 folds was employed to validate the performance of each model which offers robustness and computational efficiency. The cross-validation method partitions the data into complementary subsets conducting the analyses on one subset called the training set and validating the analyses on the other subset. In our methodology data was shuffled to ensure randomness in the cross-validation splits, controlled by a random state. The model evaluation results using different datasets were presented in Tables 5 – 7 , while Figs. 8 – 10 show ranges of R 2 , RMSE and MAE from cross-validation. Note that these are model evaluation results after hyperparameter optimization has been performed. Model evaluation results before hyperparameter tunning are shown in Tables S2- S4 of the supplementary materials. After model evaluation to select the most promising model, hyperparameter tuning was performed to successfully identify the optimal hyperparameters for each ML model applied to the original data as well as data generated by GAN and VAE. For the original dataset, the best hyperparameters yielded an RF regressor with a max_depth of 10, min_samples_split of 2, and n_estimators set to 300, while the GBR performed optimally with a learning_rate of 0.5, max_depth of 3, and n_estimators of 200. The Elastic Net model achieved its peak performance with an alpha value of 0.1 and an l1_ratio of 0.1. Similarly, the XGBoost regressor showed superior results with a learning rate of 0.5, max_depth of 3, and n_estimators of 100. When applied to GAN-generated data, the RF regressor reached its optimum with max_depth set to None, min_samples_split of 2, and n_estimators of 200, while the GBR showed improved performance with a learning_rate of 0.1, max_depth of 3, and n_estimators of 300. The Elastic Net model exhibited similar hyperparameters with an alpha of 0.1 and an l1_ratio of 0.1. For VAE-generated data, the RF regressor maintained its peak performance with max_depth set to None, min_samples_split of 2, and n_estimators of 200, while the GBR demonstrated enhanced results with a learning_rate of 0.1, max_depth of 7, and n_estimators of 300. The Elastic Net model retained its optimal hyperparameters with an alpha of 0.1 and an l1_ratio of 0.1, while the XGBoost regressor showed improved performance with a learning_rate of 0.1, max_depth of 7, and n_estimators of 100. Based on the values of R2, RMSE and MAE it was observed that hyperparameter tuning helped improve the model performance for both the original and synthetic datasets. From the results, it was observed that the GBR exhibited the highest performance with a Train and Test R² of 0.9999 and 0.9277, respectively, suggesting a strong predictive ability on the original dataset. It also displayed the lowest RMSE and MAE scores, indicating high precision of predictions, and suggesting good generalizability. Linear Regression achieved Train and Test R² scores of 0.9412 and 0.9384, respectively, suggesting a strong linear relationship. This is also reflected by a low RMSE of 0.0564 for training data and 0.0676 for testing data, MAE value of 0.0413 for training data and 0.0509 for testing data. RF showed slightly lower train and test R² scores of 0.9789 and 0.9255 but excelled in capturing non-linear complexities inherent in the data. Elastic Net scored an R² of 0.8768, demonstrating reasonable predictive power with regularization benefits and the lowest R 2 suggests slight underfitting and high RMSE signifying large errors in predictions compared to other models. XGBoost delivered an R² of 0.9992 for training data and 0.9677 for testing data, showcasing its robustness in handling various data patterns and suggesting a need to watch for overfitting. The results from the GAN-generated data indicate that the XGBoost Model stands out as the top performer achieving R 2 of 0.9965 for the training data and 0.5966 for the test data indicating its ability of predictive power of the generated synthetic data. But the model also displayed the lowest RMSE and MAE scores underscoring its precision in making predictions with minimal errors. The GBR Model demonstrated strong performance with R 2 scores of 0.9982 for train data and 0.5469 for the test data. It is slightly lower than XGBoost but showcases robust predictive capability, especially considering the synthetic nature of the data. The Linear regression Model exhibited limited capability with training and test R 2 values of 0.6071 and 0.4155 for test data indicating that it struggles to capture the complexities within synthetic data. Both RMSE and MAE were higher on the test data suggesting model predictions were less accurate for new data. RF model does not perform well compared to the GBR and XGBoost, though exhibits competitive performance with R 2 of 0.9348 for training data and 0.5180 for testing data. Though it captured a strong linear relationship, its predictive accuracy was lower compared to other ensemble methods. Elastic Net also displayed reasonable predictive power with an R 2 of 0.5836, leveraging its regularization benefits to provide stable performance on the GAN data. The results from the VAE-Generated data show the exceptional performance of several ML models with GBR leading. This is also reflected by the low range of R2, RMSE and MAE in Fig. 10 , compared to Figs. 8 and 9 . GBR achieved outstanding R 2 scores of 0.999 for the training data and 0.998 for testing data showing its predictive capability on synthetic data. Furthermore, it demonstrated low RMSE and MAE scores, underscoring its precision in predictions with minimal errors. XGBoost exhibited strong performance with R 2 scores of 0.999 for training data and 0.998 for testing and low RMSE and MAE scores reinforcing its effectiveness in capturing the underlying patterns in VAE-generated data. RF demonstrated competitive performance with R 2 scores of training data with 0.999 and 0.998 for testing data. The model showcased its ability to capture the non-linear relationships with VAE-generated data. Linear regression scores are not better than ensemble methods, but the performance is good with R 2 scores of 0.985 with training and 0.984 with testing and showed the capability of capturing the linear relationship between the independent and dependent variables. Based on the comparison results between the original and synthetic datasets, it was observed that the use of VAE data significantly improved the model accuracy. Table 5 Model performance evaluation for the original dataset. Linear Regression Random Forest Gradient Boosting Elastic Net XGBoost Train R² 0.9412 0.9789 0.9999 0.8768 0.9992 Test R² 0.9384 0.9256 0.9277 0.8478 0.9677 Train RMSE 0.0564 0.0337 0.0018 0.0816 0.0064 Test RMSE 0.0676 0.0742 0.0732 0.1061 0.0489 Train MAE 0.0413 0.0182 0.0011 0.0620 0.0044 Test MAE 0.0509 0.0455 0.0299 0.0746 0.0216 Table 6 Model performance evaluation for the GAN dataset. Linear Regression Random Forest Gradient Boosting Elastic Net XGBoost Train R² 0.6071 0.9348 0.9982 0.5836 0.9965 Test R² 0.4155 0.5180 0.5469 0.4325 0.5956 Train RMSE 0.1554 0.0633 0.0106 0.1600 0.0147 Test RMSE 0.1690 0.1535 0.1488 0.1665 0.1406 Train MAE 0.1251 0.0504 0.0084 0.1277 0.0111 Test MAE 0.1329 0.1257 0.1185 0.1302 0.1124 Table 7 Model performance evaluation for the VAE dataset. Linear Regression Random Forest Gradient Boosting Elastic Net XGBoost Train R² 0.9852 0.9998 0.9999 0.8900 0.9999 Test R² 0.9847 0.9975 0.9976 0.8935 0.9979 Train RMSE 0.0055 0.0006 0.0002 0.0151 0.0006 Test RMSE 0.0056 0.0023 0.0022 0.0148 0.0021 Train MAE 0.0044 0.0003 0.0001 0.0105 0.0004 Test MAE 0.0045 0.0009 0.0011 0.0103 0.0011 Augmented data comprising of the original and synthetic dataset were also used to develop and evaluate the ML models. The results are presented in Tables 8 and 9 . The performance of the models using augmented datasets does not compare to that of the synthetic datasets in terms of accuracy. Table 8 Model performance with GAN – original dataset (Augmented) Linear Regression Random Forest Gradient Boosting Elastic Net XGBoost Train R² 0.4380 0.9464 0.9882 0.4062 0.9850 Test R² 0.3326 0.9211 0.9761 0.1912 0.9775 Train RMSE 0.1869 0.0577 0.0271 0.1921 0.0305 Test RMSE 0.1972 0.0678 0.0373 0.2171 0.0362 Train MAE 0.1487 0.0442 0.0218 0.1521 0.0237 Test MAE 0.1527 0.0410 0.0283 0.1658 0.0267 Table 9 Model performance with VAE – original dataset (Augmented) Linear Regression Random Forest Gradient Boosting Elastic Net XGBoost Train R² 0.7715 0.9877 0.9974 0.6064 0.9949 Test R² 0.5514 0.9333 0.9877 0.3758 0.9846 Train RMSE 0.0288 0.0067 0.0031 0.0377 0.0043 Test RMSE 0.1616 0.0623 0.0268 0.1907 0.0299 Train MAE 0.0124 0.0011 0.0020 0.0216 0.0023 Test MAE 0.1363 0.0381 0.0138 0.1562 0.0192 Feature analysis was performed using the VAE dataset based on the GBR model since it is one of the most promising. The results of the feature analysis are presented in Fig. 11 . It should be mentioned that larger values of the feature indicate a greater impact on the MSP of SAF which is the model output. It shows that plant capacity, measured in kilograms per hour and location are the most crucial factors, suggesting that larger capacities might affect costs and efficiencies, thus influencing MSP significantly. Location results indicates that that where the plant is situated affects the MSP due to variables like resource availability, transportation costs, and local economic conditions. The nitrogen content of the biomass, represented as N (%), also plays a moderate role, possibly because it impacts the production process and costs. Similarly, the content of hemicellulose (Hem %) and cellulose (Cel %) in the biomass has a moderate influence, likely due to their effects on the efficiency and quality of the biofuel production process. Other elements like ash content, carbon content (C %), and lignin content (Lig %) show lower importance but are still relevant, affecting the combustion properties and processing of biomass. The least influential features include oxygen, hydrogen, fixed carbon, volatile matter, and sulfur contents, which, while impacting biomass characteristics, have a minimal direct effect on pricing compared to other factors. This analysis underscores that the plant's operational scale and its geographic location are pivotal in shaping the economic aspects of sustainable aviation fuel production. A publicly available graphical user interface has been developed and can be accessed through the barcode in Fig. 12 . The GUI enables easy evaluation of the MSP of SAF based on several input features presented in Fig. 12 . 3.5 Study limitations and future recommendations A framework for preliminary economic evaluation of the MSP of SAF from pyrolysis was presented using surrogate models. Although there are several assumptions used during the process modelling and TEA assessment, it is important to understand how the model results deviate from experimental findings for different biomass materials. These are not covered in this study. Additionally, during the TEA appraisal, the gate fees were assumed to be constant for all locations and biomass classes. The gate fees play a crucial role in determining the economic viability and market pricing of SAF. It refers to the cost incurred for the delivery and processing of feedstock at a production facility. These fees can vary depending on the type of biomass, its source, and the specific handling requirements. Feedstocks such as agricultural residues and food waste often incur gate fees. Since gate fees differ across different countries, it is also influenced by the specific type of waste management and recycling treatment employed, with fees generally rising due to increased costs and changing market conditions. While a flat gate fee was considered for all biomass in this study, future studies should focus on a holistic evaluation of the impact of gate fees on the economic model. The TEA model should also be extended to other regions or develop economic models for various countries. The project dataset is very small, and this could limit our model's ability to learn complex patterns and potentially affect the generalizability of the results to larger and more varied datasets. Although data generation techniques like GAN and VAE are used, these methods are strongly reliant on data generation but may introduce subtle biases or fail to capture real-world data distribution, possibly affecting model performance. Future studies would employ a more comprehensive biomass database with a robust dataset comprising validated experimental results from the literature. Additionally, AI could be integrated with LLM for meticulously screening academic articles to facilitate easy and fast data generation. Such an approach has been utilized by Zheng et al.[31] for text mining and the prediction of MOF synthesis. While this study only focused on TEA, lifecycle assessment (LCA) should also be formed to make an informed decision on the economic and environmental assessment of a technology. Additionally, only the MSP of SAF was used as an output variable in this study. Future studies should implement metrics such as payback period, net present value, internal rate of return, and global warming potential in making technology decisions and these metrics should be implemented in the data-driven framework. 4. Conclusions The present study evaluates the use of data-driven surrogate models combined with process simulation for the prediction of MSP of SAF from biomass pyrolysis. A biomass database was developed by collecting the proximate and ultimate analysis of 31 different biomass materials ranging from agricultural wastes, woody biomass, sewage sludge and food wastes. The dataset is comprised of 13 input features including Carbon %, Hydrogen %, Nitrogen %, Oxygen %, Sulfur%, Volatile Matter %, Fixed Carbon Content %, Ash content %, Cellulose %, Hemicellulose %, Lignin content %, Location, and Plant Capacity (kg/hr), and one target output feature MSP. To improve the model accuracy and prediction, GAN and VAE were used to generate synthetic data while hyperparameter optimization based on Grid Search was also formed. Among the five surrogate models used i.e. the linear regression, gradient boost regression (GBR), random forest (RF), extreme boost regression (XGBoost) and Elastic net, the GBR and RF appear to be most promising in terms of R 2 , RMSE and MAE for the original and synthetic datasets. GBR exhibited the highest performance with a Train and Test R² of 0.9999 and 0.9277, and RF showed slightly lower train and test R² scores of 0.9789 and 0.9255. Based on the comparison results between the original and synthetic datasets, it was observed that the use of VAE data significantly improved the model accuracy. A publicly available GUI was developed and made available for researchers to perform a preliminary estimation of the MSP of SAF as a function of biomass properties, plant capacity and location. References B. E. Rogachuk and J. A. Okolie, “Comparative assessment of pyrolysis and Gasification-Fischer Tropsch for sustainable aviation fuel production from waste tires,” Energy Convers Manag , vol. 302, p. 118110, Feb. 2024, doi: 10.1016/J.ENCONMAN.2024.118110. J. Heyne, B. Rauch, P. Le Clercq, and M. Colket, “Sustainable aviation fuel prescreening tools and procedures,” Fuel , vol. 290, p. 120004, Apr. 2021, doi: 10.1016/J.FUEL.2020.120004. N. Montoya Sánchez et al. , “Conversion of waste to sustainable aviation fuel via Fischer–Tropsch synthesis: Front-end design decisions,” Energy Sci Eng , vol. 10, no. 5, pp. 1763–1789, May 2022, doi: 10.1002/ESE3.1072. J. A. Okolie et al. , “Multi-criteria decision analysis for the evaluation and screening of sustainable aviation fuel production pathways,” iScience , vol. 26, no. 6, Jun. 2023, doi: 10.1016/J.ISCI.2023.106944. M. von Kurnatowski and M. Bortz, “Modeling and multi-criteria optimization of a process for h2o2 electrosynthesis,” Processes , vol. 9, no. 2, pp. 1–24, Feb. 2021, doi: 10.3390/pr9020399. Y. Elkasabi, C. A. Mullen, A. L. M. T. Pighinelli, and A. A. Boateng, “Hydrodeoxygenation of fast-pyrolysis bio-oils from various feedstocks using carbon-supported catalysts,” Fuel Processing Technology , vol. 123, pp. 11–18, Jul. 2014, doi: 10.1016/J.FUPROC.2014.01.039. Y. C. Liu and W. C. Wang, “Process design and evaluations for producing pyrolytic jet fuel,” Biofuels, Bioproducts and Biorefining , vol. 14, no. 2, pp. 249–264, Mar. 2020, doi: 10.1002/BBB.2061. T. Dickerson and J. Soria, “Catalytic fast pyrolysis: A review,” Energies (Basel) , vol. 6, no. 1, pp. 514–538, 2013, doi: 10.3390/en6010514. S. H. Chang, “Bio-oil derived from palm empty fruit bunches: Fast pyrolysis, liquefaction and future prospects,” Biomass Bioenergy , vol. 119, pp. 263–276, Dec. 2018, doi: 10.1016/j.biombioe.2018.09.033. F. J. Gutiérrez Ortiz, “Techno-economic assessment of supercritical processes for biofuel production,” Journal of Supercritical Fluids , vol. 160, p. 104788, Jun. 2020, doi: 10.1016/j.supflu.2020.104788. S. Michailos and A. Bridgwater, “A comparative techno-economic assessment of three bio-oil upgrading routes for aviation biofuel production,” Int J Energy Res , vol. 43, no. 13, pp. 7206–7228, Oct. 2019, doi: 10.1002/ER.4745. M. N. Saeed, M. Shahrivar, G. D. Surywanshi, T. R. Kumar, T. Mattisson, and A. H. Soleimanisalim, “Production of aviation fuel with negative emissions via chemical looping gasification of biogenic residues: Full chain process modelling and techno-economic analysis,” Fuel Processing Technology , vol. 241, p. 107585, Mar. 2023, doi: 10.1016/J.FUPROC.2022.107585. H. Li et al. , “Machine-learning-aided thermochemical treatment of biomass: a review,” Biofuel Research Journal , vol. 10, no. 1, pp. 1786–1809, Mar. 2023, doi: 10.18331/BRJ2023.10.1.4. D. Chen, C. Shang, and Z. P. Liu, “Machine-learning atomic simulation for heterogeneous catalysis,” npj Computational Materials 2023 9:1 , vol. 9, no. 1, pp. 1–9, Jan. 2023, doi: 10.1038/s41524-022-00959-5. “Machine learning-based optimization of catalytic hydrodeoxygenation of biomass pyrolysis oil,” J Clean Prod , p. 140738, Jan. 2024, doi: 10.1016/J.JCLEPRO.2024.140738. F. Elmaz, Ö. Yücel, and A. Y. Mutlu, “Predictive modeling of biomass gasification with machine learning-based regression methods,” Energy , vol. 191, p. 116541, Jan. 2020, doi: 10.1016/J.ENERGY.2019.116541. S. Rodgers et al. , “A surrogate model for the economic evaluation of renewable hydrogen production from biomass feedstocks via supercritical water gasification,” Int J Hydrogen Energy , vol. 49, pp. 277–294, Jan. 2024, doi: 10.1016/J.IJHYDENE.2023.08.016. “Phyllis2 - Database for the physico-chemical composition of (treated) lignocellulosic biomass, micro- and macroalgae, various feedstocks for biogas production and biochar.” Accessed: Apr. 22, 2024. [Online]. Available: https://phyllis.nl/ J. A. Okolie, S. Nanda, A. K. Dalai, and J. A. Kozinski, “Hydrothermal gasification of soybean straw and flax straw for hydrogen-rich syngas production: Experimental and thermodynamic modeling,” Energy Convers Manag , vol. 208, p. 112545, Mar. 2020, doi: 10.1016/J.ENCONMAN.2020.112545. J. A. Okolie, F. O. Omoarukhe, E. I. Epelle, C. C. Ogbaga, A. A. Adeleke, and P. U. Okoye, “Biomethane and propylene glycol synthesis via a novel integrated catalytic transfer hydrogenolysis, carbon capture and biomethanation process,” Chemical Engineering Journal Advances , vol. 16, p. 100523, Nov. 2023, doi: 10.1016/J.CEJA.2023.100523. J. A. Okolie, “Can biomass structural composition be predicted from a small dataset using a hybrid deep learning approach?,” Ind Crops Prod , vol. 203, p. 117191, Nov. 2023, doi: 10.1016/J.INDCROP.2023.117191. F. Elmaz, Ö. Yücel, and A. Y. Mutlu, “Predictive modeling of biomass gasification with machine learning-based regression methods,” Energy , vol. 191, p. 116541, Jan. 2020, doi: 10.1016/j.energy.2019.116541. R. Noori, M. A. Abdoli, A. Ameri Ghasrodashti, and M. Jalili Ghazizade, “Prediction of municipal solid waste generation with combination of support vector machine and principal component analysis: A case study of Mashhad,” Environ Prog Sustain Energy , vol. 28, no. 2, pp. 249–258, Jul. 2009, doi: 10.1002/EP.10317. E. Scornet, G. Biau, and J. P. Vert, “Consistency of random forests,” Ann Stat , vol. 43, no. 4, pp. 1716–1741, Aug. 2015, doi: 10.1214/15-AOS1321. E. Vigneau, P. Courcoux, R. Symoneaux, L. Guérin, and A. Villière, “Random forests: A machine learning methodology to highlight the volatile organic compounds involved in olfactory perception,” Food Qual Prefer , vol. 68, pp. 135–145, Sep. 2018, doi: 10.1016/j.foodqual.2018.02.008. G. C. Umenweke, I. C. Afolabi, E. I. Epelle, and J. A. Okolie, “Machine learning methods for modeling conventional and hydrothermal gasification of waste biomass: A review,” Bioresour Technol Rep , vol. 17, p. 100976, Feb. 2022, doi: 10.1016/J.BITEB.2022.100976. D. W. Stewart, Y. R. Cortés-Peña, Y. Li, A. S. Stillwell, M. Khanna, and J. S. Guest, “Implications of Biorefinery Policy Incentives and Location-Specific Economic Parameters for the Financial Viability of Biofuels,” Environ Sci Technol , vol. 57, no. 6, pp. 2262–2271, Feb. 2023, doi: 10.1021/ACS.EST.2C07936/ASSET/IMAGES/LARGE/ES2C07936_0004.JPEG. S. Michailos, “Process design, economic evaluation and life cycle assessment of jet fuel production from sugar cane residue,” Environ Prog Sustain Energy , vol. 37, no. 3, pp. 1227–1235, May 2018, doi: 10.1002/EP.12840. G. Yao, M. D. Staples, R. Malina, and W. E. Tyner, “Stochastic techno-economic analysis of alcohol-to-jet fuel production,” Biotechnol Biofuels , vol. 10, no. 1, pp. 1–13, Jan. 2017, doi: 10.1186/S13068-017-0702-7/FIGURES/5. S. Niekamp, U. R. Bharadwaj, J. Sadhukhan, and M. K. Chryssanthopoulos, “A multi-criteria decision support framework for sustainable asset management and challenges in its application,” Journal of Industrial and Production Engineering , vol. 32, no. 1, pp. 23–36, 2015, doi: 10.1080/21681015.2014.1000401. Z. Zheng, O. Zhang, C. Borgs, J. T. Chayes, and O. M. Yaghi, “ChatGPT Chemistry Assistant for Text Mining and the Prediction of MOF Synthesis,” J Am Chem Soc , vol. 145, no. 32, pp. 18048–18062, Aug. 2023, doi: 10.1021/JACS.3C05819/ASSET/IMAGES/LARGE/JA3C05819_0007.JPEG. Supplementary Files Supplementarymaterials.docx Cite Share Download PDF Status: Published Journal Publication published 02 Dec, 2024 Read the published version in BioEnergy Research → Version 1 posted Reviewers agreed at journal 22 Jun, 2024 Reviewers invited by journal 21 Jun, 2024 Editor invited by journal 18 Jun, 2024 Editor assigned by journal 18 Jun, 2024 First submitted to journal 17 Jun, 2024 You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4595354","acceptedTermsAndConditions":true,"allowDirectSubmit":false,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":317492361,"identity":"4876c8f0-e1d0-4bce-a101-404b33bb4f68","order_by":0,"name":"Jude Okolie","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAA5klEQVRIiWNgGAWjYBAC9gYGAyCVwMDHzMD4ACQiQUgLI0wLGzMDswGJWhgY2CSI0zIjeeMDhoo0OTZ25mPVPH/s7CUbmB9+wK8lrdiA4UyOMRszW9pt3rbkxNkMbMZ4bWKckWMmwdhWkdjGzGN2m7eBOUGOgQe/44BazH9AtPB/K+b5U28P1ML8A58WQaAtDIxtOSBb2IDoMONsBh42vLZI8zwrlkg4kwbyi7Hk3LbjiTOb2cws8GnhY0/e+OFDRbIcP//hhx/e/Km2lzje/PgGPi1gkIDCYyaofhSMglEwCkYBIQAAKRI7jOOTo2oAAAAASUVORK5CYII=","orcid":"https://orcid.org/0000-0001-9769-7307","institution":"University of Oklahoma Norman Campus: The University of Oklahoma","correspondingAuthor":true,"prefix":"","firstName":"Jude","middleName":"","lastName":"Okolie","suffix":""},{"id":317492362,"identity":"42aa8a29-2e53-4877-9870-4c0e4213d67c","order_by":1,"name":"Keon Moradi","email":"","orcid":"","institution":"University of Oklahama Norman Campus: The University of Oklahoma","correspondingAuthor":false,"prefix":"","firstName":"Keon","middleName":"","lastName":"Moradi","suffix":""},{"id":317492363,"identity":"b1a9a212-5e7f-48ab-8819-73a3e63b3145","order_by":2,"name":"Brooke Rogachuk","email":"","orcid":"","institution":"University of Oklahama Norman Campus: The University of Oklahoma","correspondingAuthor":false,"prefix":"","firstName":"Brooke","middleName":"","lastName":"Rogachuk","suffix":""},{"id":317492364,"identity":"200df18f-b6d9-4038-83dd-5acf37bf201e","order_by":3,"name":"Bala Nagaraju Narra","email":"","orcid":"","institution":"University of Oklahama Norman Campus: The University of Oklahoma","correspondingAuthor":false,"prefix":"","firstName":"Bala","middleName":"Nagaraju","lastName":"Narra","suffix":""},{"id":317492365,"identity":"efc6e5fd-a1c4-4f06-ba30-cc0679880a91","order_by":4,"name":"Chukwuma C. Ogbaga","email":"","orcid":"","institution":"Philomath University","correspondingAuthor":false,"prefix":"","firstName":"Chukwuma","middleName":"C.","lastName":"Ogbaga","suffix":""},{"id":317492366,"identity":"281c4df6-35fb-4037-8b0b-324e46af9c42","order_by":5,"name":"Patrick Okoye","email":"","orcid":"","institution":"National Autonomous University of Mexico: Universidad Nacional Autonoma de Mexico","correspondingAuthor":false,"prefix":"","firstName":"Patrick","middleName":"","lastName":"Okoye","suffix":""},{"id":317492367,"identity":"c80f4805-5ddb-4fcd-8baa-049fe2cfdb7d","order_by":6,"name":"Adekunle Adeleke","email":"","orcid":"","institution":"Nile University of Nigeria","correspondingAuthor":false,"prefix":"","firstName":"Adekunle","middleName":"","lastName":"Adeleke","suffix":""}],"badges":[],"createdAt":"2024-06-17 16:26:18","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4595354/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4595354/v1","draftVersion":[],"editorialEvents":[{"content":"https://doi.org/10.1007/s12155-024-10803-x","type":"published","date":"2024-12-02T15:57:34+00:00"}],"editorialNote":"","failedWorkflow":false,"files":[{"id":60135964,"identity":"c6ddd1a5-5b49-436b-9da5-089080101442","added_by":"auto","created_at":"2024-07-12 08:03:43","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":936536,"visible":true,"origin":"","legend":"\u003cp\u003eSchematics of the data-driven framework for TEA assessment.\u003c/p\u003e","description":"","filename":"1.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/ae6600edea8858624b29913f.png"},{"id":60135962,"identity":"5ab35b2e-d4c5-4f18-9b09-f75e4b51f62c","added_by":"auto","created_at":"2024-07-12 08:03:43","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":138266,"visible":true,"origin":"","legend":"\u003cp\u003eHistogram showing the distribution of the original dataset.\u003c/p\u003e","description":"","filename":"2.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/0847d2f08302a3913b4da655.png"},{"id":60135970,"identity":"320a8ddf-46ef-4be8-833d-90e95522e0fe","added_by":"auto","created_at":"2024-07-12 08:03:44","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":133273,"visible":true,"origin":"","legend":"\u003cp\u003e(a) Schematics of the GABN architecture (b) Schematics of the VAE architecture.\u003c/p\u003e","description":"","filename":"3.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/67979e825744faaba2f9d240.png"},{"id":60135967,"identity":"b861a63e-4d4a-4608-a34b-b048f0a64d73","added_by":"auto","created_at":"2024-07-12 08:03:44","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":131654,"visible":true,"origin":"","legend":"\u003cp\u003eImpact of ultimate analysis on the MSP of SAF.\u003c/p\u003e","description":"","filename":"4.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/6ac1c4e3c5d3a0f8f1d93be0.png"},{"id":60135969,"identity":"71fb5546-be08-46c9-a05b-39a7d32e51f8","added_by":"auto","created_at":"2024-07-12 08:03:44","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":260546,"visible":true,"origin":"","legend":"\u003cp\u003eImpact of proximate and compositional analysis on the MSP of SAF.\u003c/p\u003e","description":"","filename":"5.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/fb262c0555f8230d30817c5d.png"},{"id":60136874,"identity":"e167c2fc-0670-425c-9404-526b4e19af11","added_by":"auto","created_at":"2024-07-12 08:11:44","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":21280,"visible":true,"origin":"","legend":"\u003cp\u003eImpact of plant capacity on the MSP\u003c/p\u003e","description":"","filename":"6.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/4eca4f65aaa9aeed19eef3c2.png"},{"id":60135976,"identity":"9a0686c5-d5a3-4c13-85fd-e929cb1d190a","added_by":"auto","created_at":"2024-07-12 08:03:44","extension":"png","order_by":7,"title":"Figure 7","display":"","copyAsset":false,"role":"figure","size":30601,"visible":true,"origin":"","legend":"\u003cp\u003eComparison of the MSP of different SAF production technologies. GFT represents gasification- Fischer Tropsch process; ATJ represents alcohol to jet fuel; HEFA represents Hydroprocessing of esters and fatty acids; DSHC represents Direct sugar to hydrocarbons; FP represents fast pyrolysis. Data obtained from [4], [28], [29].\u003c/p\u003e","description":"","filename":"7.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/d1ae5574f30a554c332da607.png"},{"id":60136876,"identity":"48f1b9f9-d4b3-4e50-a05d-36ad47913258","added_by":"auto","created_at":"2024-07-12 08:11:44","extension":"png","order_by":8,"title":"Figure 8","display":"","copyAsset":false,"role":"figure","size":205086,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eBox plot comparing data range between the original and synthetic dataset.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"8.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/ca9d2adb1430ab869daa021b.png"},{"id":60135963,"identity":"775d966b-3d1e-4646-a655-86ca0687fe08","added_by":"auto","created_at":"2024-07-12 08:03:43","extension":"png","order_by":9,"title":"Figure 9","display":"","copyAsset":false,"role":"figure","size":60843,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 8: Cross-validated model performance metrics for the original dataset.\u003c/p\u003e","description":"","filename":"08.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/036b61f5b3e96d7dd1ef568a.png"},{"id":60136877,"identity":"8b9aa66b-af93-45d4-a9b8-9e0bd8769687","added_by":"auto","created_at":"2024-07-12 08:11:44","extension":"png","order_by":10,"title":"Figure 10","display":"","copyAsset":false,"role":"figure","size":73938,"visible":true,"origin":"","legend":"\u003cp\u003eFigure \u0026nbsp;9: Cross-validated model performance for GAN dataset\u003c/p\u003e","description":"","filename":"9.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/48988d5bea1bebdfd8c8d0af.png"},{"id":60135973,"identity":"8c9316e3-ac5c-4b43-af3a-adaf6f5cb1b1","added_by":"auto","created_at":"2024-07-12 08:03:44","extension":"png","order_by":11,"title":"Figure 11","display":"","copyAsset":false,"role":"figure","size":47500,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 10: Cross-validated model performance for VAE dataset.\u003c/p\u003e","description":"","filename":"10.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/ced3245c8f0f71f596f1f79b.png"},{"id":60135975,"identity":"fcb88352-6fa3-4e3b-bd29-f4edfee8c2df","added_by":"auto","created_at":"2024-07-12 08:03:44","extension":"png","order_by":12,"title":"Figure 12","display":"","copyAsset":false,"role":"figure","size":20917,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 11: Feature importance of the model inputs on the MSP Predictions using SHAP values\u003c/p\u003e","description":"","filename":"11.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/ce73ce0d8a30de556dc3ae08.png"},{"id":60136875,"identity":"c3911841-23f6-4fab-873e-b58090c1d4e1","added_by":"auto","created_at":"2024-07-12 08:11:44","extension":"png","order_by":13,"title":"Figure 13","display":"","copyAsset":false,"role":"figure","size":73374,"visible":true,"origin":"","legend":"\u003cp\u003eFigure 12: \u0026nbsp;Barcode to the develop GUI\u003c/p\u003e","description":"","filename":"12.png","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/d9802668cadd5e4f4c3718b4.png"},{"id":70965252,"identity":"5bd3b513-09ad-4d67-87fc-c59c4985c2bd","added_by":"auto","created_at":"2024-12-09 16:18:09","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":3480475,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/7b4be76e-6281-4548-a63c-672fe043d58f.pdf"},{"id":60135965,"identity":"4e9b2200-bad8-4a23-bdac-36a98d391cc5","added_by":"auto","created_at":"2024-07-12 08:03:43","extension":"docx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":940510,"visible":true,"origin":"","legend":"","description":"","filename":"Supplementarymaterials.docx","url":"https://assets-eu.researchsquare.com/files/rs-4595354/v1/4bb7716d281d862b4bba4bed.docx"}],"financialInterests":"","formattedTitle":"Data-driven framework for the techno-economic assessment of sustainable aviation fuel from pyrolysis.","fulltext":[{"header":"1. Introduction","content":"\u003cp\u003eThe aviation industry has contributed immensely to the immediate transportation of people and goods across the continents. It also provides a significant contribution towards the economic growth and social integration of a nation. According to the International Air Transport Association (IATA), the aviation sector is on the trajectory of exponential growth of global passenger counts expected to double over two decades, reaching 8.2\u0026nbsp;billion by 2030 [1]. This projection underscores the burdening demand for air travel, stimulated by population growth, affordability of airfare, and improved standards of living. It should be mentioned that the growth in the aviation industry is also accompanied by severe environmental impacts.\u003c/p\u003e \u003cp\u003eThe aviation sector contributes significantly to global greenhouse gas emissions. These emissions include carbon dioxide and other gases that have worsened the greenhouse effect [2]. Since 2009, fuel use by commercial airlines has been climbing annually, hitting an all-time high of 359\u0026nbsp;billion liters in 2019 [3]. However, even with the travel limits imposed in 2020 to curb the spread of COVID-19, which reduced air travel significantly, commercial airlines still used around 101\u0026nbsp;billion liters of aviation fuel that year [3]. Moreover, the aviation industry accounts for about 2% of total carbon dioxide emissions [3]. Therefore, there is an urgent need to decarbonize the aviation industry to ensure mitigation of greenhouse gas emissions while keeping up with its growth.\u003c/p\u003e \u003cp\u003eSustainable aviation fuel (SAF) presents a way to reduce the environmental effects of the aviation industry by providing a clean-burning, renewable substitute for conventional jet fuel. SAF can be produced from diverse processes and feedstocks. These processes are categorized into thermochemical, biological, and integrated processes [1]. The advantages and limitations of each SAF production pathway have been documented in a previous study [4]. Among the production methods, thermochemical processes are advantageous due to their high efficiency and the ability to convert a wide range of biomass feedstocks including waste cooking oil, municipal waste, algae, and lignocellulosic biomass [5].\u003c/p\u003e \u003cp\u003ePyrolysis is a thermochemical process that offers a low-carbon pathway for SAF production. During fast pyrolysis (FP), biogenic waste is heated in an air-free environment facilitating its breakdown into bio-oil, gases, and solid char [6]. FP operates at high temperatures between 400\u0026ndash;600\u0026deg;C and short residence times, efficiently maximizing the production of liquid bio-oil, which can then be refined into biofuels including SAF via hydro-cracking and isomerization [1]. It should be mentioned that using bio-oil directly as fuel is not feasible due to its high oxygen content and characteristics such as thermal instability, corrosiveness, and low energy density [7]. To meet Sustainable Aviation Fuel (SAF) standards and ensure compatibility with current aircraft systems, further refinement of the bio-oil is necessary.\u003c/p\u003e \u003cp\u003eThe production of SAF from FP and subsequent hydroprocessing is advantageous for several reasons including low-cost feedstocks, faster reaction times, and simpler technology, making it more cost-effective and scalable compared to other thermochemical processes [8], [9]. A recent study showed that while FP-hydroprocessing and gasification-Fischer Tropsch (GFT) processes are two economically feasible thermochemical processes for producing SAF, the latter produced lower CO\u003csub\u003e2\u003c/sub\u003e emissions compared to GFT and fossil-based aviation fuel [1]. Although promising, FP pathways for SAF face several challenges related to high oxygen content and instability of the bio-oil produced. These require significant refining to meet fuel standards, catalysts deactivation during hydroprocessing as well as detailed understanding of the relationship between the physicochemical properties of the biogenic waste and bio-oil properties.\u003c/p\u003e \u003cp\u003eTo scale up FP- hydroprocessing for commercial SAF production while addressing the challenges mentioned earlier, techno-economic analysis (TEA) studies are required. TEA is used to assess the economic viability of a new process or a full-scale plant, especially in conceptual designs where there is no information about a similar industrial facility [10]. In the case of FP-hydroprocessing, TEA is used to assess its economic viability and how several factors influence the profitability of the process at varying capacities. However, performing TEA analysis requires the development of rigorous process models, large data collection from literature and several assumptions related to equipment purchase cost (EPC), biomass properties and plant capacity. Moreover, the relationship between the feedstock properties and the minimum selling price of fuel (MSP) is often challenging to estimate.\u003c/p\u003e \u003cp\u003eSeveral studies have reported the TEA of SAF in literature. For instance, Rogachuk and Okolie [1], performed a detailed economic evaluation of GFT and FP for SAF production, however, they focused on waste tires, and it is not clear if their economic model could be applicable to other feedstocks. In another study, Michailos and Bridgewater[11] compared the economic feasibility of producing SAF from three different bio-oil upgrading routes. Their results do not explain how plant capacity and varying location as well as biomass properties influence the MSP. Recently, Saeed et al.[12] also explored the economic feasibility of SAF production via chemical looping technologies. Their study showed that an economically feasible SAF with negative emission can be produced by integrated chemical looping technology. Despite the ample amount of studies on the TEA of SAF, there are several literature gaps. Integrating feedstock properties including proximate and ultimate analysis has seldom been reported. Additionally, studies that developed an economic model while considering different production capacities and different regions are limited. Furthermore, the complexity and assumptions inherent in TEA studies have not been addressed. Therein lies the motivation of the present study. The proposed study integrated process simulation results with experimental data to develop a data-driven machine learning (ML) framework for easy prediction of the MSP of SAF from FP.\u003c/p\u003e \u003cp\u003eData-driven ML methods have been used in biofuel technology to study the relationship between process conditions and yield by several researchers [13], [14], [15], [16]. In addition, ML methods are employed for catalyst optimization, biofuel yield prediction and char predictions [13] [17]. However, its application as a complementary tool to TEA is scarcely reported. Therefore, the key objectives of the present study are as follows: combine process models with experimental data from the literature to develop a robust economic analysis database. Following which two deep learning methods would be used to generate additional synthetic data for ML models. Develop relationships between biomass properties, production capacity and regions with the MSP of SAF. Finally, a publicly available graphical user interface (GUI) will be developed to help researchers and industry practices in the preliminary estimation of the MSP of SAF from FP.\u003c/p\u003e"},{"header":"2. Methodology","content":"\u003cp\u003eThe overall methodology adopted in this study is presented in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The steps include the development of a robust dataset comprising process model development in Aspen Plus, experiment data collection from literature and deep learning methods for synthetic data generation. Details of each method are meticulously presented in subsequent sections.\u003c/p\u003e \u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data collection\u003c/h2\u003e \u003cp\u003eA biomass database was created by collecting the proximate and ultimate analysis of 31 different biomass materials. These biomasses represent several categories of biogenic wastes including energy crops, animal manure, agricultural wastes, woody biomass, sewage sludge and food wastes. The proximate, ultimate, and compositional analysis of the biogenic waste materials were compiled from the literature and Phyllis database [18]. The dataset is comprised of 13 input features including Carbon %, Hydrogen %, Nitrogen %, Oxygen %, Sulfur%, Volatile Matter %, Fixed Carbon Content %, Ash content %, Cellulose %, Hemicellulose %, Lignin content %, Location, and Plant Capacity (kg/hr), and one target output feature MSP. It should be noted that all proximate and ultimate analysis data are either received on a dry basis or converted to a dry basis and normalized to sum to 100%. The study explores three different locations: the US, China, and the UK. Additionally, two different plant capacities, 25,000 kg/hr and 50,000 kg/hr, are examined. Location factors of 0.66, 1.0, and 1.04 were used to represent China, the US, and the UK, respectively [17].\u003c/p\u003e \u003cp\u003eBased on two different plant capacities and three locations, a total of 186 data points were collected, representing biomass with varying proximate, ultimate, and compositional analyses, as well as differences in plant capacity and location. The entire dataset is presented in \u003cb\u003eTable \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e\u003c/b\u003e of the supplementary materials while Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e shows the summary statistics of the dataset. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e2\u003c/span\u003e presents a histogram plot that shows the distribution of various features in the dataset. In each plot, the x-axis represents a continuous variable, while the y-axis corresponds to the frequency of observations within the bins defined by the x-axis. The shapes of the distributions suggest the variability of each biomass property within the dataset. The histogram shows the distribution of the MSP for different biomass materials indicating the range of prices from the use of different feedstocks, location and production capacities.\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eSummary statistics of the original dataset. Note all data are presented on a dry basis while the proximate, ultimate and compositional analyses are expressed in percentages. The MSP is the only output feature in the Table. Location is not presented in the table.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"9\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c6\" colnum=\"6\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c7\" colnum=\"7\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c8\" colnum=\"8\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c9\" colnum=\"9\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eFeatures\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eCount\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMean\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eStd\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eMinimum value\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c6\"\u003e \u003cp\u003e25th percentile\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c7\"\u003e \u003cp\u003e50th percentile\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c8\"\u003e \u003cp\u003e75th percentile\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c9\"\u003e \u003cp\u003eMaximum value\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eC (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e49.51\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e4.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e40.02\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e47.01\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e48.87\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e52.85\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e60.46\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eH (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e6.42\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1.04\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e5.32\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e5.98\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e6.18\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e6.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e10.76\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eN (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e1.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e1.95\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.10\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.49\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e1.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e2.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e8.21\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eO (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e42.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e5.06\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e27.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e39.37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e43.28\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e45.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e52.86\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eS (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.22\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.39\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.31\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e1.55\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eVolatile Matter (VM) (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e75.47\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9.17\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e49.36\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e71.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e76.56\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e81.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e94.16\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eAsh content (Ash) (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e8.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9.88\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e1.70\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e4.13\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e14.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e42.02\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eFixed carbon (FC) (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e15.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e6.38\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e3.37\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e11.83\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e16.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e20.61\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e26.60\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eCellulose content (Cel) (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e33.69\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9.96\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e6.92\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e25.19\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e35.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e42.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e49.88\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eHemicellulose content (Hem) (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e26.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e9.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e12.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e22.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e25.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e30.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e55.42\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eLignin content (Lig) (wt.%)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e21.84\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e10.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e1.60\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e14.50\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e22.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e26.80\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e53.40\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003ePlant capacity.\u003c/b\u003e\u003c/p\u003e \u003cp\u003e\u003cb\u003e(kg/hr)\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e37500.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e12533.74\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e25000.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e25000.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e37500.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e50000.00\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e50000.00\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003e\u003cb\u003eMinimum selling price (MSP) US$/L\u003c/b\u003e\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e186\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.79\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.24\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c5\"\u003e \u003cp\u003e0.40\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c6\"\u003e \u003cp\u003e0.64\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c7\"\u003e \u003cp\u003e0.73\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c8\"\u003e \u003cp\u003e0.99\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c9\"\u003e \u003cp\u003e1.59\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003cp\u003eA total number of 186 data points were used as shown in Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. Furthermore, the range of proximate, ultimate, and compositional analyses was carefully selected to capture different varieties of biomass feedstock. That way the range of biomass materials properties can develop data-driven models that can predict the MSP of SAF produced from feedstocks within and outside the data range.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Process model development and economic evaluation\u003c/h2\u003e \u003cp\u003eA model for fast pyrolysis was developed using the advanced features of Aspen Plus Version 11, a comprehensive process modeling software licensed through the University of Oklahoma. The simulation was meticulously conducted under specific conditions, namely steady-state and isothermal, to ensure the utmost accuracy and consistency in the results. Details of the FP-hydroprocessing model including the underlying assumptions have been meticulously defined in our previous study [1]. The Aspen Plus model was refined to accommodate different feedstock based on their proximate and ultimate analysis. A total of 31 different simulations were developed for various feedstocks, each with unique compositions, to achieve the most accurate model possible. During FP, biomass was defined as non-conventional solids by using a calculator block incorporated into the RYield block, and its function is executed using a programmed FORTRAN subroutine statement. Details of the non-conventional reactor modelling in Aspen Plus can be found elsewhere [19].\u003c/p\u003e \u003cp\u003eAn economic assessment was performed to determine the MSP of SAF following a generalized approach reported in our previous study [1][20]. This approach involves the appraisal of the Capital Expenditure (CapEx) and the Operational Expenditure (OpEx). The CaPEx calculations are based on the EPC and the installation costs. In contrast, the OpEx was estimated to include both Variable Operating Cost (VOC) and Fixed Operating Cost (FOC), considering components of labor costs and EPC. Details of CapEx and OpEx calculations including the assumptions have been documented in our recent study [1]. Since the proposed plant location and capacity varies, the economic model was adjusted to account for the plant location and salary of laborers. For instance, the labor cost for the U.S. is different from China and these costs were based on the minimum wages in the respective countries. Furthermore, the EPC was estimated using the Aspen Plus Economic Analyzer while the cost was adjusted to the year 2023 using the chemical engineering plant cost index (CEPCI).\u003c/p\u003e \u003cp\u003eThe MSP of SAF which is defined as the lowest price at which the fuel can be sold to cover all costs of production and ensure economic viability. The MSP is estimated based on the sum of capital and operating expenses over the project's lifetime.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec5\" class=\"Section2\"\u003e \u003ch2\u003e2.3 Data Preprocessing and Synthetic Data Generation\u003c/h2\u003e \u003cp\u003eThe dataset includes a categorical variable, 'Location,' which was transformed via one-hot encoding prior to training the Generative Adversarial Network (GAN) model. This transformation resulted in high-dimensional data with binary values assigned to the location feature. The dataset was partitioned into training and testing sets (X_train, X_test, Y_train, Y_test) using the train_test_split() function, configured with a test size of 0.2 and a random state set to 42. Numerical data within the dataset was normalized using the MinMaxScaler() function before model training.\u003c/p\u003e \u003cp\u003eIn order to improve the performance of surrogate models for the prediction of MSP of SAF, synthetic data generation methods based on deep learning approaches were applied. Two synthetic data generation methods including the Generative Adversarial Network (GAN) and Variational Autoencoder (VAE) were used and compared. GAN and VAE models were utilized to generate 5,000 records. The generated data underwent a cleaning process to ensure it conformed to the characteristics of the real data, which involved removing any data points that fell outside the predefined range. The z-score was calculated for the features in the synthetic data to identify and remove outliers. Additionally, bounds were set based on the minimum and maximum values observed in the original data, and these bounds were used to further refine the synthetic data.\u003c/p\u003e \u003cp\u003eThe GAN is a neural network algorithm for generative modeling. GANs consist of two neural networks, the generator G and the discriminator D, which are trained simultaneously through a competitive process [21]. The generator \u0026ldquo;G\u0026rdquo; creates samples that are intended to come from the same distribution as the training data, while the discriminator D evaluates them. Details of the GANs' operating mechanism are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The main objective of the generator is to produce data that is indistinguishable from real data. It takes a random noise vector z (from a latent space) as input and generates samples that resemble the training data. The discriminator's job is to distinguish between actual data samples and fake samples produced by the generator. It outputs a scalar representing the probability that a given sample is real. GAN is trained through a min-max game where D and G have competing objectives.\u003c/p\u003e \u003cp\u003eThe GAN model comprises a generator and a discriminator. The generator features a dense network with 32 units and utilizes LeakyReLU as the activation function in its hidden layers to introduce nonlinearity, enabling the model to tackle complex patterns effectively. The output layer of the generator is bifurcated: the first segment employs a tanh activation function to scale outputs to a range of [-1, 1], while the second segment consists of three units using a softmax function, which facilitates the generation of one-hot encoded data. The dimensionality of the latent space (\u003cb\u003elatent_dim\u003c/b\u003e) and the number of units in the hidden layers are adjustable hyperparameters, allowing customization according to the complexity of the dataset.\u003c/p\u003e \u003cp\u003eCompared to GANs, the VAEs are a type of generative model that uses neural networks to encode data into a latent space and then decode it back to the original space. Unlike traditional autoencoders, VAEs are probabilistic models that produce a distribution over the latent space (Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e3\u003c/span\u003e). The encoder part of a VAE maps input data x to a latent representation z. However, instead of encoding an input as a single point, it encodes it as a distribution over the latent space.\u003c/p\u003e \u003cp\u003eThe decoder part then samples from this distribution to generate a sample \u0026prime;x\u0026prime; that approximates the original input x. The loss function for VAEs has two terms: the reconstruction loss and the KL divergence. The reconstruction loss ensures the decoded samples match the original inputs, while the KL divergence regularizes the encoder by comparing the distribution of latent variables with a prior (often a Gaussian distribution).\u003c/p\u003e \u003cp\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.4 Surrogate models development\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eFive different ML models including linear regression, gradient boost regression (GBR), random forest (RF), extreme boost regression (XGBoost) and Elastic net were selected for the development of surrogate models. These models were selected due to their promising results in biofuel yield optimization, catalyst design, product yield prediction and new materials development [13], [15], [16], [22].\u003c/p\u003e \u003cp\u003eGBR is an ensemble machine learning model that trains using a boosting strategy [23]. A gradient-boosting leaf is obtained by taking the average of the output values and using them as an initial value to calculate the initial errors. Then, the GBR model begins constructing a decision tree, which is a binary tree used by ML models to split input data into smaller subgroups, and this new tree is assigned a learning rate which it uses to calculate new prediction values from the previous prediction value and error. This process occurs recursively until the maximum depth limit has been passed or until a satisfactory error value has been obtained. Various hyperparameters of the GBR model need to be tuned to properly optimize the predictive algorithm.\u003c/p\u003e \u003cp\u003eThe RF Regressor is also an ensemble learning algorithm that utilizes random feature bagging and selection strategies to output an average of the base model predictions. It simultaneously trains several base decision trees on random subsets of the original data, and for qualitative data, it averages the results of those decision trees for a final prediction output [24]. Various hyperparameters must be passed to RF Regressors to optimize their performance, including the number of decision trees that can be running in parallel, the maximum number of variables in each tree, and the total number of decision trees to be constructed [25].\u003c/p\u003e \u003cp\u003eThe XGBoost represents a regularized regression method that constructs multiple decision trees sequentially, with each new tree aiming to correct the errors made by the previous ones. This process helps in minimizing the overall prediction error. XGBoost is well-regarded for its ability to handle large datasets efficiently, its versatility in accommodating various types of predictive modeling problems, and its robustness in dealing with a wide range of data types.\u003c/p\u003e \u003cp\u003eElasticNet is a regression technique that combines the properties of both ridge and lasso regression methods. It is particularly useful when dealing with highly correlated data or when the number of predictors exceeds the number of observations. ElasticNet incorporates penalties from both L1 and L2 regularization, allowing it to shrink some coefficients to zero (like lasso) for variable selection, while also evenly distributing the penalty among the remaining coefficients (like ridge).\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Model evaluation and hyperparameter optimization\u003c/h2\u003e \u003cp\u003e \u003cdiv class=\"BlockQuote\"\u003e \u003cp\u003eIn order to evaluate the performance of each ML model the Regression coefficient (R\u003csup\u003e2\u003c/sup\u003e), root-mean-square error (RMSE) and mean absolute error (MAE) were used as quantifiable metrics. Details of these criteria as well as the underlying equations can be found elsewhere [26]. It should be mentioned that these metrics were selected because they provide a comprehensive view of the surrogate model performance from different viewpoints. R\u0026sup2; indicates how much of the variance in the dependent variable is explained by the predictors, offering insight into the overall effectiveness of the model. RMSE provides a scale-dependent measure of the average magnitude of the prediction errors, penalizing larger errors more severely due to its squaring of residuals. MAE, on the other hand, offers a straightforward, scale-dependent average error magnitude that is easy to interpret and less sensitive to outliers than RMSE. The combination of these metrics provides and holistic and balanced understanding of model accuracy, consistency, and the nature of errors. A K-fold cross-validation method with 5 folds was employed to validate the performance of each model which offers robustness and computational efficiency. The cross-validation method partitions the data into complementary subsets conducting the analyses on one subset called the training set and validating the analyses on the other subset. In our methodology data was shuffled to ensure randomness in the cross-validation splits, controlled by a random state.\u003c/p\u003e \u003cp\u003eAfter developing the ML models, hyperparameter optimization was performed to help identify the optimized hyperparameter combinations that produce the most promising accuracy. The Grid search approach was used during hyperparameter optimization. This approach has been proven advantageous in our previous study [21]. In Grid Search, all possible combinations of the specified hyperparameters are systematically explored. A set of hyperparameters was initially defined during the training process as shown in Table\u0026nbsp;\u003cspan refid=\"Tab2\" class=\"InternalRef\"\u003e2\u003c/span\u003e. Three distinct values were initially selected for each hyperparameter to ensure a wide range of options for the tuning process to explore. This approach helps to assess how sensitive the model's performance is to different settings of each hyperparameter, thus improving the ability to find the optimal configuration that yields the best results.\u003c/p\u003e \u003cp\u003eFor Linear Regression, no hyperparameters were selected to tune. For RF, n_estimators (100, 200, 300), max_depth (None, 5, 10), and min_samples_split (2, 5, 10) were selected. For GBR, n_estimators (100, 200, 300), learning_rate (0.01, 0.1, 0.5), and max_depth (3, 5, 7) were selected. For Elastic Net, hyperparameters such as alpha (0.1, 0.5, 1.0), and l1_ratio (0.1, 0.5, 0.9) were chosen. For XGBoost, hyperparameters such as n_estimators (100, 200, 300), learning_rate (0.01, 0.1, 0.5), and max_depth (3, 5, 7) were selected.\u003c/p\u003e \u003c/div\u003e \u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab2\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 2\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eHyperparameters initially selected during the ML model development.\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"3\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eModel\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colspan=\"2\" nameend=\"c3\" namest=\"c2\"\u003e \u003cp\u003eHyperparameters Selected\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eLinear Regression\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003eNone\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e\u0026nbsp;\u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRandom Forest\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en_estimators\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100, 200, 300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emax_depth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003eNone, 5, 10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emin_samples_split\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e2, 5, 10\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGradient Boosting\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en_estimators\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100, 200, 300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003elearning_rate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.01, 0.1, 0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emax_depth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3, 5, 7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eElastic Net\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003ealpha\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1, 0.5, 1.0\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003el1_ratio\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.1, 0.5, 0.9\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003en_estimators\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e100, 200, 300\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003elearning_rate\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e0.01, 0.1, 0.5\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e\u0026nbsp;\u003c/td\u003e \u003ctd align=\"left\" colname=\"c2\"\u003e \u003cp\u003emax_depth\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"left\" colname=\"c3\"\u003e \u003cp\u003e3,5,7\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003c/div\u003e"},{"header":"3. Results and Discussions","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e\n \u003ch2\u003e3.1 Parametric study\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eThe influence of proximate, ultimate, and compositional analysis of biomass on the MSP of SAF for different locations is presented in Figs. \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e. The results show that geographical locations influence the MSP of SAF. Prices in the US and UK are higher compared to prices in China. Furthermore, the figures show that a change in the biomass properties does not have a pronounced impact on the MSP compared to the geographical location which had a considerable impact. The higher MSP of SAF in the UK and US compared to China can be attributed to several factors. Firstly, the regulatory environment in the UK and US often imposes stricter sustainability and environmental standards, which can increase production costs. These countries may also have higher labor and operational costs. Additionally, the UK and US might have more developed markets for SAF, with policies supporting its use, such as tax incentives or mandatory blending requirements, which can drive up prices due to increased demand. Conversely, China might benefit from lower production costs, government subsidies, and a rapidly scaling biofuels industry, which can help keep prices lower. These factors combined explain the regional differences in the minimum selling price of SAF. The influence of location on the MSP of biofuels has been reported by some researchers [17], [27]. Rogers et al.[17] reported a considerable difference in the Levelized cost of hydrogen (LCOH) produced from hydrothermal gasification at different locations. The authors attributed this observation to the varying capital and operating costs in different countries.\u003c/p\u003e\n \u003cp\u003eThe production capacity has a significant impact on the MSP of SAF as shown in Fig. \u003cspan class=\"InternalRef\"\u003e6\u003c/span\u003e. A larger production capacity led to a decline in the MSP. The trend follows the economies of scale, where the average cost of producing SAF decreases as the quantity of output increases.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eThis reduction in unit cost is primarily due to the more efficient use of resources, better procurement terms for bulk materials, and the spreading of fixed costs (like capital investments and administrative expenses) over a larger volume of production. Consequently, producers with higher capacities are often able to offer SAF at more competitive prices.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eThe MSP of SAF obtained from biomass pyrolysis is compared to other SAF production technologies including gasification- Fischer Tropsch process (GFT), alcohol to jet fuel (ATJ), hydroprocessing of esters and fatty acids (HEFA) and direct sugar to hydrocarbons (DSHC). The results presented in Fig. \u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e show varying MSPs for different technologies. It should be noted that the results were extracted from various published studies that use different feedstocks, locations and the year of study differs. Additionally, the prices were also computed at different plant capacities which could significantly influence the MSP. Regardless the data presented in Fig. \u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e provides a basis for comparing the economic viability of various SAF production pathways.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eAs seen in Fig. \u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e, GFT has the lowest range, starting close to 0.40 US\u003cspan\u003e$\u003c/span\u003e/L up to about 0.91 US\u003cspan\u003e$\u003c/span\u003e/L. On the contrary, ATJ shows a price range between 0.75\u0026ndash;1.38 US\u003cspan\u003e$\u003c/span\u003e/L. DSHC shows the most significant variation, with prices starting near 2.21 US\u003cspan\u003e$\u003c/span\u003e/L and reaching above 2.5US\u003cspan\u003e$\u003c/span\u003e/L. The price ranges for each technology can be attributed to factors such as feedstock costs, production efficiency, technological maturity, and scale of production. For instance, DSHC relies on sugars as the primary feedstock, which can be expensive, especially if sourced from food-grade crops. The cost can be influenced by agricultural market prices, which are subject to volatility. Moreover, converting sugars to hydrocarbons can involve complex chemical processes characterized by low product yield [28], [30]. The low yield could significantly influence the process economics as documented in a previous study [28]. The lower cost of HEFA compared to other SAF production pathways could be attributed to several reasons. HEFA technology commonly uses waste oils, animal fats, and non-food vegetable oils, which can be less expensive and more readily available than the food-grade sugars required for DSHC. Furthermore, the hydroprocessing step in HEFA technology is well understood and can be highly efficient, leading to higher yields of SAF from the feedstock and potentially lower production costs. HEFA processes can produce valuable co-products such as propane and naphtha, which can be sold to offset the cost of fuel production.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec10\" class=\"Section2\"\u003e\n \u003ch2\u003e3.2 Synthetic Dataset Generation\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eAbout 5000 synthetic values based on the original dataset using GAN and VAE were generated. These datasets mimic the distribution of the original data. However, after applying data cleaning techniques on the synthetic data including the removal of out-of-range or unreasonable values, the dataset was reduced to 4862 for VAE and 346 records from GAN. A box plot comparing the original dataset with GAN and VAE datasets is presented in Fig. \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e while Tables \u003cspan class=\"InternalRef\"\u003e3\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e4\u003c/span\u003e show the summary statistics of the dataset.\u003c/p\u003e\n \u003cp\u003eThe GAN synthetic data has generally higher variability as shown by larger standard deviations for many features compared to the original data. This might reflect greater diversity but also indicates that the synthetic data might be capturing additional noise or generating unrealistic values such as negative values for Cel, Hem, Lig, and MSP. The VAE-generated data exhibits mean values that closely approximate the original data across most features but show significantly less variability. While this might suggest that the VAE is better at capturing the central tendency of the data, the reduced variability could imply that it is not capturing the full diversity present in the original dataset.\u003c/p\u003e\n \u003cp\u003eThe analysis of various components in the datasets reveals distinct differences between the original data and the synthetic data generated from both GAN and VAE. For Carbon Content (C), while the original dataset shows a broader spread potentially inclusive of outliers, the GAN synthetic dataset demonstrates an even wider distribution, contrasting with the VAE\u0026apos;s much narrower interquartile range (IQR). This trend of narrower variability with VAE-generated data is also evident in the distribution of Hydrogen Content (H), where the GAN distribution is wider than the original, and the VAEs are significantly narrower.\u003c/p\u003e\n \u003cp\u003eFor elements such as Nitrogen (N), Oxygen (O), Sulfur (S), and Volatile Matter (VM), the VAE synthetic data maintains a noticeably narrower spread compared to both the original and GAN datasets. In terms of Ash and Fixed Carbon (FC), the original dataset displays a widespread, with the GAN data showing a similar distribution for Ash but a wider range for FC. Conversely, the VAE data for these features appears more condensed, indicating less variability.\u003c/p\u003e\n \u003cp\u003eConsidering the compositional analysis of biomass like Cellulose (Cel), Hemicellulose (Hem), and Lignin (Lig), the GAN data reveals extreme outliers with a broader distribution for Hem and Lig, while VAE-generated data avoids such extremes and aligns more closely with the original dataset\u0026apos;s distribution. Plant Capacity analysis shows the original data concentrated around a middle capacity range, whereas the GAN data is more dispersed, and VAE data is markedly more compact.\u003c/p\u003e\n \u003cp\u003eFor the MSP, the original data presents a slightly skewed distribution with higher variability, which is more pronounced in the GAN data. The VAE, on the other hand, shows a tight distribution, indicating a significant reduction in variation in the synthetic selling prices. Regarding the location features (China, UK, US), which are categorical and represented through one-hot encoding, both GAN and VAE maintain the original distribution without introducing outliers, as expected for binary categorical data. This analysis underscores the varying effectiveness and characteristics of GAN and VAE in replicating and altering the statistical properties of the original dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable id=\"Tab3\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 3\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eSummary statistics of the GAN synthetic dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"9\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eFeatures\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCount\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMean\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eStd\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMinimum value\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e25th percentile\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e50th percentile\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e75th percentile\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMaximum value\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eC (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e52.85\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e40.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49.19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e53.69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e56.80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e60.23\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eH (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.19\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5.33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5.96\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7.84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10.19\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eN (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.49\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e3.97\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5.35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7.95\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eO (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e47.33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4.56\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e30.06\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e44.82\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e48.89\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e50.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e52.77\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eS (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.41\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.28\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.36\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.60\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.24\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eVolatile Matter (VM) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e85.68\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.02\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e62.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e81.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e87.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e90.52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e93.89\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eAsh content (Ash) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e9.71\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.80\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4.44\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e8.03\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e13.91\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e29.28\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eFixed carbon (FC) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22.83\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.96\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e10.98\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e21.37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e23.56\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e25.13\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e26.50\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eCellulose content (Cel) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e31.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e8.89\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7.48\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e26.56\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e32.12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e39.51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e48.53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eHemicellulose content (Hem) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e40.27\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e12.21\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e12.87\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e30.53\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e43.36\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e50.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e55.28\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eLignin content (Lig) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e26.33\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e14.22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15.15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e24.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e38.22\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e53.03\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003ePlant capacity (kg/hr)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e41222.49\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5739.17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e25495.57\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e37100.14\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e41958.52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e46114.67\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49672.01\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMSP (US$/L)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e346.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.92\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.75\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.93\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.10\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.53\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\u003cbr\u003e\u003c/div\u003e\u0026nbsp;\u003ctable id=\"Tab4\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 4\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eSummary statistics of the VAE synthetic dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"9\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eFeatures\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eCount\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMean\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eStd\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMinimum value\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e25th percentile\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e50th percentile\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003e75th percentile\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eMaximum value\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eC (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49.61\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.32\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e48.48\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49.41\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49.65\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e49.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e50.31\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eH (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.47\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e5.59\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.26\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.50\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.72\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e6.95\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eN (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.40\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.72\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.04\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.31\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.58\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eO (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e41.94\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e41.37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e41.63\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e41.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e42.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e43.22\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eS (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.31\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.09\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.25\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.32\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.45\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eVolatile Matter (VM) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e75.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.75\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e74.18\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e74.74\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e75.24\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e75.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e77.86\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eAsh content (Ash) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e9.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e2.17\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e7.52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e9.59\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e11.39\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e13.41\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eFixed carbon (FC) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15.76\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.42\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15.03\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15.45\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15.67\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e15.99\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e17.21\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eCellulose content (Cel) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e33.66\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.46\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e31.46\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e32.51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e33.35\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e34.54\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e38.49\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eHemicellulose content (Hem) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e27.15\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.52\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22.38\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e26.12\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e27.36\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e28.37\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e29.56\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eLignin content (Lig) (wt.%)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22.29\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e1.01\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e18.34\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e21.69\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e22.43\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e23.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e23.87\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003ePlant capacity (kg/hr)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e37464.86\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e395.08\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e35993.07\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e37185.51\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e37571.23\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e37773.73\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e38231.48\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eMSP (US$/L)\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e4862.00\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.05\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.65\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.78\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.81\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.84\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.88\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec11\" class=\"Section2\"\u003e\n \u003ch2\u003e3.3 Evaluation of Surrogate Model Performance\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eUsing the original and synthetic dataset five different surrogate models were developed for the prediction of the MSP of SAF from biomass pyrolysis. A K-fold cross-validation method with 5 folds was employed to validate the performance of each model which offers robustness and computational efficiency. The cross-validation method partitions the data into complementary subsets conducting the analyses on one subset called the training set and validating the analyses on the other subset. In our methodology data was shuffled to ensure randomness in the cross-validation splits, controlled by a random state. The model evaluation results using different datasets were presented in Tables \u003cspan class=\"InternalRef\"\u003e5\u003c/span\u003e\u0026ndash;\u003cspan class=\"InternalRef\"\u003e7\u003c/span\u003e, while Figs. \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e\u0026ndash;\u003cspan class=\"InternalRef\"\u003e10\u003c/span\u003e show ranges of R\u003csup\u003e2\u003c/sup\u003e, RMSE and MAE from cross-validation. Note that these are model evaluation results after hyperparameter optimization has been performed. Model evaluation results before hyperparameter tunning are shown in Tables S2- S4 of the supplementary materials.\u003c/p\u003e\n \u003cp\u003eAfter model evaluation to select the most promising model, hyperparameter tuning was performed to successfully identify the optimal hyperparameters for each ML model applied to the original data as well as data generated by GAN and VAE. For the original dataset, the best hyperparameters yielded an RF regressor with a max_depth of 10, min_samples_split of 2, and n_estimators set to 300, while the GBR performed optimally with a learning_rate of 0.5, max_depth of 3, and n_estimators of 200. The Elastic Net model achieved its peak performance with an alpha value of 0.1 and an l1_ratio of 0.1.\u003c/p\u003e\n \u003cp\u003eSimilarly, the XGBoost regressor showed superior results with a learning rate of 0.5, max_depth of 3, and n_estimators of 100. When applied to GAN-generated data, the RF regressor reached its optimum with max_depth set to None, min_samples_split of 2, and n_estimators of 200, while the GBR showed improved performance with a learning_rate of 0.1, max_depth of 3, and n_estimators of 300. The Elastic Net model exhibited similar hyperparameters with an alpha of 0.1 and an l1_ratio of 0.1. For VAE-generated data, the RF regressor maintained its peak performance with max_depth set to None, min_samples_split of 2, and n_estimators of 200, while the GBR demonstrated enhanced results with a learning_rate of 0.1, max_depth of 7, and n_estimators of 300. The Elastic Net model retained its optimal hyperparameters with an alpha of 0.1 and an l1_ratio of 0.1, while the XGBoost regressor showed improved performance with a learning_rate of 0.1, max_depth of 7, and n_estimators of 100. Based on the values of R2, RMSE and MAE it was observed that hyperparameter tuning helped improve the model performance for both the original and synthetic datasets.\u003c/p\u003e\n \u003cp\u003eFrom the results, it was observed that the GBR exhibited the highest performance with a Train and Test R\u0026sup2; of 0.9999 and 0.9277, respectively, suggesting a strong predictive ability on the original dataset. It also displayed the lowest RMSE and MAE scores, indicating high precision of predictions, and suggesting good generalizability.\u003c/p\u003e\n \u003cp\u003eLinear Regression achieved Train and Test R\u0026sup2; scores of 0.9412 and 0.9384, respectively, suggesting a strong linear relationship. This is also reflected by a low RMSE of 0.0564 for training data and 0.0676 for testing data, MAE value of 0.0413 for training data and 0.0509 for testing data. RF showed slightly lower train and test R\u0026sup2; scores of 0.9789 and 0.9255 but excelled in capturing non-linear complexities inherent in the data. Elastic Net scored an R\u0026sup2; of 0.8768, demonstrating reasonable predictive power with regularization benefits and the lowest R\u003csup\u003e2\u003c/sup\u003e suggests slight underfitting and high RMSE signifying large errors in predictions compared to other models. XGBoost delivered an R\u0026sup2; of 0.9992 for training data and 0.9677 for testing data, showcasing its robustness in handling various data patterns and suggesting a need to watch for overfitting.\u003c/p\u003e\n \u003cp\u003eThe results from the GAN-generated data indicate that the XGBoost Model stands out as the top performer achieving R\u003csup\u003e2\u003c/sup\u003e of 0.9965 for the training data and 0.5966 for the test data indicating its ability of predictive power of the generated synthetic data. But the model also displayed the lowest RMSE and MAE scores underscoring its precision in making predictions with minimal errors. The GBR Model demonstrated strong performance with R\u003csup\u003e2\u003c/sup\u003e scores of 0.9982 for train data and 0.5469 for the test data. It is slightly lower than XGBoost but showcases robust predictive capability, especially considering the synthetic nature of the data. The Linear regression Model exhibited limited capability with training and test R\u003csup\u003e2\u003c/sup\u003e values of 0.6071 and 0.4155 for test data indicating that it struggles to capture the complexities within synthetic data. Both RMSE and MAE were higher on the test data suggesting model predictions were less accurate for new data.\u003c/p\u003e\n \u003cp\u003eRF model does not perform well compared to the GBR and XGBoost, though exhibits competitive performance with R\u003csup\u003e2\u003c/sup\u003e of 0.9348 for training data and 0.5180 for testing data. Though it captured a strong linear relationship, its predictive accuracy was lower compared to other ensemble methods. Elastic Net also displayed reasonable predictive power with an R\u003csup\u003e2\u003c/sup\u003e of 0.5836, leveraging its regularization benefits to provide stable performance on the GAN data.\u003c/p\u003e\n \u003cp\u003eThe results from the VAE-Generated data show the exceptional performance of several ML models with GBR leading. This is also reflected by the low range of R2, RMSE and MAE in Fig. \u003cspan class=\"InternalRef\"\u003e10\u003c/span\u003e, compared to Figs. \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e9\u003c/span\u003e. GBR achieved outstanding R\u003csup\u003e2\u003c/sup\u003e scores of 0.999 for the training data and 0.998 for testing data showing its predictive capability on synthetic data. Furthermore, it demonstrated low RMSE and MAE scores, underscoring its precision in predictions with minimal errors.\u003c/p\u003e\n \u003cp\u003eXGBoost exhibited strong performance with R\u003csup\u003e2\u003c/sup\u003e scores of 0.999 for training data and 0.998 for testing and low RMSE and MAE scores reinforcing its effectiveness in capturing the underlying patterns in VAE-generated data.\u003c/p\u003e\n \u003cp\u003eRF demonstrated competitive performance with R\u003csup\u003e2\u003c/sup\u003e scores of training data with 0.999 and 0.998 for testing data. The model showcased its ability to capture the non-linear relationships with VAE-generated data. Linear regression scores are not better than ensemble methods, but the performance is good with R\u003csup\u003e2\u003c/sup\u003e scores of 0.985 with training and 0.984 with testing and showed the capability of capturing the linear relationship between the independent and dependent variables. Based on the comparison results between the original and synthetic datasets, it was observed that the use of VAE data significantly improved the model accuracy.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable id=\"Tab5\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 5\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eModel performance evaluation for the original dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"6\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLinear Regression\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eForest\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eElastic Net\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9412\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9789\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9999\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.8768\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9992\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9384\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9256\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9277\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.8478\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9677\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0564\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0337\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0018\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0816\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0064\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0676\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0742\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0732\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1061\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0489\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0413\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0182\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0011\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0620\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0044\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0509\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0455\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0299\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0746\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0216\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\u003cbr\u003e\u003c/div\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\u003cbr\u003e\u003c/div\u003e\u0026nbsp;\u003ctable id=\"Tab6\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 6\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eModel performance evaluation for the GAN dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"6\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLinear Regression\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eForest\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eElastic Net\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.6071\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9348\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9982\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.5836\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9965\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.4155\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.5180\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.5469\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.4325\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.5956\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain\u003c/strong\u003e\u003c/p\u003e\n \u003cp\u003e\u003cstrong\u003eRMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1554\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0633\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0106\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1600\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0147\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1690\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1535\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1488\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1665\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1406\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1251\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0504\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0084\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1277\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0111\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1329\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1257\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1185\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1302\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1124\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\u003cbr\u003e\u003c/div\u003e\u0026nbsp;\u003ctable id=\"Tab7\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 7\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eModel performance evaluation for the VAE dataset.\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"6\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLinear Regression\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRandom\u003c/p\u003e\n \u003cp\u003eForest\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eElastic Net\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9852\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9998\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9999\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.8900\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9999\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9847\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9975\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9976\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.8935\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9979\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0055\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0006\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0002\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0151\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0006\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0056\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0023\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0022\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0148\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0021\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0044\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0003\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0001\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0105\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0004\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0045\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0009\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0011\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0103\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0011\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eAugmented data comprising of the original and synthetic dataset were also used to develop and evaluate the ML models. The results are presented in Tables \u003cspan class=\"InternalRef\"\u003e8\u003c/span\u003e and \u003cspan class=\"InternalRef\"\u003e9\u003c/span\u003e. The performance of the models using augmented datasets does not compare to that of the synthetic datasets in terms of accuracy.\u003c/p\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\u0026nbsp;\u003ctable id=\"Tab8\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 8\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eModel performance with GAN \u0026ndash; original dataset (Augmented)\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"6\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLinear Regression\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eElastic Net\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.4380\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9464\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9882\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.4062\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9850\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.3326\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9211\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9761\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1912\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9775\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1869\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0577\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0271\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1921\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0305\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1972\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0678\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0373\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.2171\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0362\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1487\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0442\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0218\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1521\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0237\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1527\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0410\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0283\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1658\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0267\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"gridtable\"\u003e\n \u003cdiv align=\"char\" class=\"colspec\"\u003e\u003cbr\u003e\u003c/div\u003e\u0026nbsp;\u003ctable id=\"Tab9\" border=\"1\"\u003e\n \u003ccaption language=\"En\"\u003e\n \u003cdiv class=\"CaptionNumber\"\u003eTable 9\u003c/div\u003e\n \u003cdiv class=\"CaptionContent\"\u003e\n \u003cp\u003eModel performance with VAE \u0026ndash; original dataset (Augmented)\u003c/p\u003e\n \u003c/div\u003e\n \u003c/caption\u003e\n \u003ccolgroup cols=\"6\"\u003e\u003c/colgroup\u003e\n \u003cthead\u003e\n \u003ctr\u003e\n \u003cth align=\"left\"\u003e\u0026nbsp;\u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eLinear Regression\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eRandom Forest\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eGradient Boosting\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eElastic Net\u003c/p\u003e\n \u003c/th\u003e\n \u003cth align=\"left\"\u003e\n \u003cp\u003eXGBoost\u003c/p\u003e\n \u003c/th\u003e\n \u003c/tr\u003e\n \u003c/thead\u003e\n \u003ctbody\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.7715\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9877\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9974\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.6064\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9949\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest R\u0026sup2;\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.5514\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9333\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9877\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.3758\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.9846\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0288\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0067\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0031\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0377\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0043\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest RMSE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1616\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0623\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0268\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1907\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0299\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTrain MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0124\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0011\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0020\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0216\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0023\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003ctr\u003e\n \u003ctd align=\"left\"\u003e\n \u003cp\u003e\u003cstrong\u003eTest MAE\u003c/strong\u003e\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1363\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0381\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0138\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.1562\u003c/p\u003e\n \u003c/td\u003e\n \u003ctd align=\"char\"\u003e\n \u003cp\u003e0.0192\u003c/p\u003e\n \u003c/td\u003e\n \u003c/tr\u003e\n \u003c/tbody\u003e\n \u003c/table\u003e\n \u003c/div\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eFeature analysis was performed using the VAE dataset based on the GBR model since it is one of the most promising. The results of the feature analysis are presented in Fig. \u003cspan class=\"InternalRef\"\u003e11\u003c/span\u003e. It should be mentioned that larger values of the feature indicate a greater impact on the MSP of SAF which is the model output. It shows that plant capacity, measured in kilograms per hour and location are the most crucial factors, suggesting that larger capacities might affect costs and efficiencies, thus influencing MSP significantly. Location results indicates that that where the plant is situated affects the MSP due to variables like resource availability, transportation costs, and local economic conditions. The nitrogen content of the biomass, represented as N (%), also plays a moderate role, possibly because it impacts the production process and costs. Similarly, the content of hemicellulose (Hem %) and cellulose (Cel %) in the biomass has a moderate influence, likely due to their effects on the efficiency and quality of the biofuel production process. Other elements like ash content, carbon content (C %), and lignin content (Lig %) show lower importance but are still relevant, affecting the combustion properties and processing of biomass. The least influential features include oxygen, hydrogen, fixed carbon, volatile matter, and sulfur contents, which, while impacting biomass characteristics, have a minimal direct effect on pricing compared to other factors. This analysis underscores that the plant\u0026apos;s operational scale and its geographic location are pivotal in shaping the economic aspects of sustainable aviation fuel production.\u003c/p\u003e\n \u003cp\u003eA publicly available graphical user interface has been developed and can be accessed through the barcode in Fig. \u003cspan class=\"InternalRef\"\u003e12\u003c/span\u003e. The GUI enables easy evaluation of the MSP of SAF based on several input features presented in Fig. \u003cspan class=\"InternalRef\"\u003e12\u003c/span\u003e.\u003c/p\u003e\n \u003c/div\u003e\n\u003c/div\u003e\n\u003cdiv id=\"Sec12\" class=\"Section2\"\u003e\n \u003ch2\u003e3.5 Study limitations and future recommendations\u003c/h2\u003e\n \u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eA framework for preliminary economic evaluation of the MSP of SAF from pyrolysis was presented using surrogate models. Although there are several assumptions used during the process modelling and TEA assessment, it is important to understand how the model results deviate from experimental findings for different biomass materials. These are not covered in this study. Additionally, during the TEA appraisal, the gate fees were assumed to be constant for all locations and biomass classes. The gate fees play a crucial role in determining the economic viability and market pricing of SAF. It refers to the cost incurred for the delivery and processing of feedstock at a production facility. These fees can vary depending on the type of biomass, its source, and the specific handling requirements. Feedstocks such as agricultural residues and food waste often incur gate fees. Since gate fees differ across different countries, it is also influenced by the specific type of waste management and recycling treatment employed, with fees generally rising due to increased costs and changing market conditions. While a flat gate fee was considered for all biomass in this study, future studies should focus on a holistic evaluation of the impact of gate fees on the economic model. The TEA model should also be extended to other regions or develop economic models for various countries. The project dataset is very small, and this could limit our model\u0026apos;s ability to learn complex patterns and potentially affect the generalizability of the results to larger and more varied datasets. Although data generation techniques like GAN and VAE are used, these methods are strongly reliant on data generation but may introduce subtle biases or fail to capture real-world data distribution, possibly affecting model performance. Future studies would employ a more comprehensive biomass database with a robust dataset comprising validated experimental results from the literature. Additionally, AI could be integrated with LLM for meticulously screening academic articles to facilitate easy and fast data generation. Such an approach has been utilized by Zheng et al.[31] for text mining and the prediction of MOF synthesis.\u003c/p\u003e\n \u003c/div\u003e\n \u003cp\u003eWhile this study only focused on TEA, lifecycle assessment (LCA) should also be formed to make an informed decision on the economic and environmental assessment of a technology. Additionally, only the MSP of SAF was used as an output variable in this study. Future studies should implement metrics such as payback period, net present value, internal rate of return, and global warming potential in making technology decisions and these metrics should be implemented in the data-driven framework.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"4. Conclusions","content":"\u003cdiv class=\"BlockQuote\"\u003e\n \u003cp\u003eThe present study evaluates the use of data-driven surrogate models combined with process simulation for the prediction of MSP of SAF from biomass pyrolysis. A biomass database was developed by collecting the proximate and ultimate analysis of 31 different biomass materials ranging from agricultural wastes, woody biomass, sewage sludge and food wastes. The dataset is comprised of 13 input features including Carbon %, Hydrogen %, Nitrogen %, Oxygen %, Sulfur%, Volatile Matter %, Fixed Carbon Content %, Ash content %, Cellulose %, Hemicellulose %, Lignin content %, Location, and Plant Capacity (kg/hr), and one target output feature MSP. To improve the model accuracy and prediction, GAN and VAE were used to generate synthetic data while hyperparameter optimization based on Grid Search was also formed. Among the five surrogate models used i.e. the linear regression, gradient boost regression (GBR), random forest (RF), extreme boost regression (XGBoost) and Elastic net, the GBR and RF appear to be most promising in terms of R\u003csup\u003e2\u003c/sup\u003e, RMSE and MAE for the original and synthetic datasets. GBR exhibited the highest performance with a Train and Test R\u0026sup2; of 0.9999 and 0.9277, and RF showed slightly lower train and test R\u0026sup2; scores of 0.9789 and 0.9255. Based on the comparison results between the original and synthetic datasets, it was observed that the use of VAE data significantly improved the model accuracy. A publicly available GUI was developed and made available for researchers to perform a preliminary estimation of the MSP of SAF as a function of biomass properties, plant capacity and location.\u003c/p\u003e\n\u003c/div\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eB. E. Rogachuk and J. A. Okolie, \u0026ldquo;Comparative assessment of pyrolysis and Gasification-Fischer Tropsch for sustainable aviation fuel production from waste tires,\u0026rdquo; \u003cem\u003eEnergy Convers Manag\u003c/em\u003e, vol. 302, p. 118110, Feb. 2024, doi: 10.1016/J.ENCONMAN.2024.118110.\u003c/li\u003e\n\u003cli\u003eJ. Heyne, B. Rauch, P. Le Clercq, and M. Colket, \u0026ldquo;Sustainable aviation fuel prescreening tools and procedures,\u0026rdquo; \u003cem\u003eFuel\u003c/em\u003e, vol. 290, p. 120004, Apr. 2021, doi: 10.1016/J.FUEL.2020.120004.\u003c/li\u003e\n\u003cli\u003eN. Montoya S\u0026aacute;nchez \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Conversion of waste to sustainable aviation fuel via Fischer\u0026ndash;Tropsch synthesis: Front-end design decisions,\u0026rdquo; \u003cem\u003eEnergy Sci Eng\u003c/em\u003e, vol. 10, no. 5, pp. 1763\u0026ndash;1789, May 2022, doi: 10.1002/ESE3.1072.\u003c/li\u003e\n\u003cli\u003eJ. A. Okolie \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Multi-criteria decision analysis for the evaluation and screening of sustainable aviation fuel production pathways,\u0026rdquo; \u003cem\u003eiScience\u003c/em\u003e, vol. 26, no. 6, Jun. 2023, doi: 10.1016/J.ISCI.2023.106944.\u003c/li\u003e\n\u003cli\u003eM. von Kurnatowski and M. Bortz, \u0026ldquo;Modeling and multi-criteria optimization of a process for h2o2 electrosynthesis,\u0026rdquo; \u003cem\u003eProcesses\u003c/em\u003e, vol. 9, no. 2, pp. 1\u0026ndash;24, Feb. 2021, doi: 10.3390/pr9020399.\u003c/li\u003e\n\u003cli\u003eY. Elkasabi, C. A. Mullen, A. L. M. T. Pighinelli, and A. A. Boateng, \u0026ldquo;Hydrodeoxygenation of fast-pyrolysis bio-oils from various feedstocks using carbon-supported catalysts,\u0026rdquo; \u003cem\u003eFuel Processing Technology\u003c/em\u003e, vol. 123, pp. 11\u0026ndash;18, Jul. 2014, doi: 10.1016/J.FUPROC.2014.01.039.\u003c/li\u003e\n\u003cli\u003eY. C. Liu and W. C. Wang, \u0026ldquo;Process design and evaluations for producing pyrolytic jet fuel,\u0026rdquo; \u003cem\u003eBiofuels, Bioproducts and Biorefining\u003c/em\u003e, vol. 14, no. 2, pp. 249\u0026ndash;264, Mar. 2020, doi: 10.1002/BBB.2061.\u003c/li\u003e\n\u003cli\u003eT. Dickerson and J. Soria, \u0026ldquo;Catalytic fast pyrolysis: A review,\u0026rdquo; \u003cem\u003eEnergies (Basel)\u003c/em\u003e, vol. 6, no. 1, pp. 514\u0026ndash;538, 2013, doi: 10.3390/en6010514.\u003c/li\u003e\n\u003cli\u003eS. H. Chang, \u0026ldquo;Bio-oil derived from palm empty fruit bunches: Fast pyrolysis, liquefaction and future prospects,\u0026rdquo; \u003cem\u003eBiomass Bioenergy\u003c/em\u003e, vol. 119, pp. 263\u0026ndash;276, Dec. 2018, doi: 10.1016/j.biombioe.2018.09.033.\u003c/li\u003e\n\u003cli\u003eF. J. Guti\u0026eacute;rrez Ortiz, \u0026ldquo;Techno-economic assessment of supercritical processes for biofuel production,\u0026rdquo; \u003cem\u003eJournal of Supercritical Fluids\u003c/em\u003e, vol. 160, p. 104788, Jun. 2020, doi: 10.1016/j.supflu.2020.104788.\u003c/li\u003e\n\u003cli\u003eS. Michailos and A. Bridgwater, \u0026ldquo;A comparative techno-economic assessment of three bio-oil upgrading routes for aviation biofuel production,\u0026rdquo; \u003cem\u003eInt J Energy Res\u003c/em\u003e, vol. 43, no. 13, pp. 7206\u0026ndash;7228, Oct. 2019, doi: 10.1002/ER.4745.\u003c/li\u003e\n\u003cli\u003eM. N. Saeed, M. Shahrivar, G. D. Surywanshi, T. R. Kumar, T. Mattisson, and A. H. Soleimanisalim, \u0026ldquo;Production of aviation fuel with negative emissions via chemical looping gasification of biogenic residues: Full chain process modelling and techno-economic analysis,\u0026rdquo; \u003cem\u003eFuel Processing Technology\u003c/em\u003e, vol. 241, p. 107585, Mar. 2023, doi: 10.1016/J.FUPROC.2022.107585.\u003c/li\u003e\n\u003cli\u003eH. Li \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;Machine-learning-aided thermochemical treatment of biomass: a review,\u0026rdquo; \u003cem\u003eBiofuel Research Journal\u003c/em\u003e, vol. 10, no. 1, pp. 1786\u0026ndash;1809, Mar. 2023, doi: 10.18331/BRJ2023.10.1.4.\u003c/li\u003e\n\u003cli\u003eD. Chen, C. Shang, and Z. P. Liu, \u0026ldquo;Machine-learning atomic simulation for heterogeneous catalysis,\u0026rdquo; \u003cem\u003enpj Computational Materials 2023 9:1\u003c/em\u003e, vol. 9, no. 1, pp. 1\u0026ndash;9, Jan. 2023, doi: 10.1038/s41524-022-00959-5.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Machine learning-based optimization of catalytic hydrodeoxygenation of biomass pyrolysis oil,\u0026rdquo; \u003cem\u003eJ Clean Prod\u003c/em\u003e, p. 140738, Jan. 2024, doi: 10.1016/J.JCLEPRO.2024.140738.\u003c/li\u003e\n\u003cli\u003eF. Elmaz, \u0026Ouml;. Y\u0026uuml;cel, and A. Y. Mutlu, \u0026ldquo;Predictive modeling of biomass gasification with machine learning-based regression methods,\u0026rdquo; \u003cem\u003eEnergy\u003c/em\u003e, vol. 191, p. 116541, Jan. 2020, doi: 10.1016/J.ENERGY.2019.116541.\u003c/li\u003e\n\u003cli\u003eS. Rodgers \u003cem\u003eet al.\u003c/em\u003e, \u0026ldquo;A surrogate model for the economic evaluation of renewable hydrogen production from biomass feedstocks via supercritical water gasification,\u0026rdquo; \u003cem\u003eInt J Hydrogen Energy\u003c/em\u003e, vol. 49, pp. 277\u0026ndash;294, Jan. 2024, doi: 10.1016/J.IJHYDENE.2023.08.016.\u003c/li\u003e\n\u003cli\u003e\u0026ldquo;Phyllis2 - Database for the physico-chemical composition of (treated) lignocellulosic biomass, micro- and macroalgae, various feedstocks for biogas production and biochar.\u0026rdquo; Accessed: Apr. 22, 2024. [Online]. Available: https://phyllis.nl/\u003c/li\u003e\n\u003cli\u003eJ. A. Okolie, S. Nanda, A. K. Dalai, and J. A. Kozinski, \u0026ldquo;Hydrothermal gasification of soybean straw and flax straw for hydrogen-rich syngas production: Experimental and thermodynamic modeling,\u0026rdquo; \u003cem\u003eEnergy Convers Manag\u003c/em\u003e, vol. 208, p. 112545, Mar. 2020, doi: 10.1016/J.ENCONMAN.2020.112545.\u003c/li\u003e\n\u003cli\u003eJ. A. Okolie, F. O. Omoarukhe, E. I. Epelle, C. C. Ogbaga, A. A. Adeleke, and P. U. Okoye, \u0026ldquo;Biomethane and propylene glycol synthesis via a novel integrated catalytic transfer hydrogenolysis, carbon capture and biomethanation process,\u0026rdquo; \u003cem\u003eChemical Engineering Journal Advances\u003c/em\u003e, vol. 16, p. 100523, Nov. 2023, doi: 10.1016/J.CEJA.2023.100523.\u003c/li\u003e\n\u003cli\u003eJ. A. Okolie, \u0026ldquo;Can biomass structural composition be predicted from a small dataset using a hybrid deep learning approach?,\u0026rdquo; \u003cem\u003eInd Crops Prod\u003c/em\u003e, vol. 203, p. 117191, Nov. 2023, doi: 10.1016/J.INDCROP.2023.117191.\u003c/li\u003e\n\u003cli\u003eF. Elmaz, \u0026Ouml;. Y\u0026uuml;cel, and A. Y. Mutlu, \u0026ldquo;Predictive modeling of biomass gasification with machine learning-based regression methods,\u0026rdquo; \u003cem\u003eEnergy\u003c/em\u003e, vol. 191, p. 116541, Jan. 2020, doi: 10.1016/j.energy.2019.116541.\u003c/li\u003e\n\u003cli\u003eR. Noori, M. A. Abdoli, A. Ameri Ghasrodashti, and M. Jalili Ghazizade, \u0026ldquo;Prediction of municipal solid waste generation with combination of support vector machine and principal component analysis: A case study of Mashhad,\u0026rdquo; \u003cem\u003eEnviron Prog Sustain Energy\u003c/em\u003e, vol. 28, no. 2, pp. 249\u0026ndash;258, Jul. 2009, doi: 10.1002/EP.10317.\u003c/li\u003e\n\u003cli\u003eE. Scornet, G. Biau, and J. P. Vert, \u0026ldquo;Consistency of random forests,\u0026rdquo; \u003cem\u003eAnn Stat\u003c/em\u003e, vol. 43, no. 4, pp. 1716\u0026ndash;1741, Aug. 2015, doi: 10.1214/15-AOS1321.\u003c/li\u003e\n\u003cli\u003eE. Vigneau, P. Courcoux, R. Symoneaux, L. Gu\u0026eacute;rin, and A. Villi\u0026egrave;re, \u0026ldquo;Random forests: A machine learning methodology to highlight the volatile organic compounds involved in olfactory perception,\u0026rdquo; \u003cem\u003eFood Qual Prefer\u003c/em\u003e, vol. 68, pp. 135\u0026ndash;145, Sep. 2018, doi: 10.1016/j.foodqual.2018.02.008.\u003c/li\u003e\n\u003cli\u003eG. C. Umenweke, I. C. Afolabi, E. I. Epelle, and J. A. Okolie, \u0026ldquo;Machine learning methods for modeling conventional and hydrothermal gasification of waste biomass: A review,\u0026rdquo; \u003cem\u003eBioresour Technol Rep\u003c/em\u003e, vol. 17, p. 100976, Feb. 2022, doi: 10.1016/J.BITEB.2022.100976.\u003c/li\u003e\n\u003cli\u003eD. W. Stewart, Y. R. Cort\u0026eacute;s-Pe\u0026ntilde;a, Y. Li, A. S. Stillwell, M. Khanna, and J. S. Guest, \u0026ldquo;Implications of Biorefinery Policy Incentives and Location-Specific Economic Parameters for the Financial Viability of Biofuels,\u0026rdquo; \u003cem\u003eEnviron Sci Technol\u003c/em\u003e, vol. 57, no. 6, pp. 2262\u0026ndash;2271, Feb. 2023, doi: 10.1021/ACS.EST.2C07936/ASSET/IMAGES/LARGE/ES2C07936_0004.JPEG.\u003c/li\u003e\n\u003cli\u003eS. Michailos, \u0026ldquo;Process design, economic evaluation and life cycle assessment of jet fuel production from sugar cane residue,\u0026rdquo; \u003cem\u003eEnviron Prog Sustain Energy\u003c/em\u003e, vol. 37, no. 3, pp. 1227\u0026ndash;1235, May 2018, doi: 10.1002/EP.12840.\u003c/li\u003e\n\u003cli\u003eG. Yao, M. D. Staples, R. Malina, and W. E. Tyner, \u0026ldquo;Stochastic techno-economic analysis of alcohol-to-jet fuel production,\u0026rdquo; \u003cem\u003eBiotechnol Biofuels\u003c/em\u003e, vol. 10, no. 1, pp. 1\u0026ndash;13, Jan. 2017, doi: 10.1186/S13068-017-0702-7/FIGURES/5.\u003c/li\u003e\n\u003cli\u003eS. Niekamp, U. R. Bharadwaj, J. Sadhukhan, and M. K. Chryssanthopoulos, \u0026ldquo;A multi-criteria decision support framework for sustainable asset management and challenges in its application,\u0026rdquo; \u003cem\u003eJournal of Industrial and Production Engineering\u003c/em\u003e, vol. 32, no. 1, pp. 23\u0026ndash;36, 2015, doi: 10.1080/21681015.2014.1000401.\u003c/li\u003e\n\u003cli\u003eZ. Zheng, O. Zhang, C. Borgs, J. T. Chayes, and O. M. Yaghi, \u0026ldquo;ChatGPT Chemistry Assistant for Text Mining and the Prediction of MOF Synthesis,\u0026rdquo; \u003cem\u003eJ Am Chem Soc\u003c/em\u003e, vol. 145, no. 32, pp. 18048\u0026ndash;18062, Aug. 2023, doi: 10.1021/JACS.3C05819/ASSET/IMAGES/LARGE/JA3C05819_0007.JPEG.\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":false,"highlight":"","institution":"","isAcceptedByJournal":true,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":true,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"bioenergy-research","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bere","sideBox":"Learn more about [BioEnergy Research](https://www.springer.com/journal/12155)","snPcode":"12155","submissionUrl":"https://submission.nature.com/new-submission/12155/3","title":"BioEnergy Research","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false},"keywords":"Pyrolysis, Sustainable Aviation fuel, Machine learning, Techno-economic, Biofuels","lastPublishedDoi":"10.21203/rs.3.rs-4595354/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4595354/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eThe aviation sector plays a crucial role in quickly moving people and goods around the world. It also greatly helps in the economic growth and social integration of countries. As the industry continues to experience rapid growth, there is a tendency for an increase in emissions associated with the industry. Sustainable aviation fuel (SAF) presents a way to reduce the environmental effects of the aviation industry by providing a clean-burning, renewable substitute for conventional jet fuel. SAF can be produced from diverse processes and feedstocks. Fast pyrolysis (FP) is a promising thermochemical process for SAF production due to its advantages including low-cost feedstocks, faster reaction times, and simpler technology, making it more cost-effective and scalable compared to other thermochemical processes. However, the preliminary estimation of the economic viability of FP for SAF production is complex and tedious requiring detailed process models and several assumptions. Moreover, the relationship between the feedstock properties and the minimum selling price of fuel (MSP) is often challenging to estimate. To address these challenges, the present study developed a data-driven framework for preliminary estimation of the MSP of SAF from FP. The target output feature is MSP. To enhance model accuracy and predictions, synthetic data was created using Generative Adversarial Networks (GAN) and Variational Autoencoders (VAE), and hyperparameter optimization was conducted using Grid Search. Five surrogate models were evaluated: linear regression, gradient boost regression (GBR), random forest (RF), extreme boost regression (XGBoost), and Elastic net. GBR and RF showed the most promise based on metrics like R\u0026sup2;, RMSE, and MAE for both original and synthetic datasets. Specifically, GBR achieved a Train R\u0026sup2; of 0.9999 and a Test R\u0026sup2; of 0.9277, while RF had Train and Test R\u0026sup2; scores of 0.9789 and 0.9255, respectively. The use of data from the VAE notably enhanced model accuracy. Additionally, a publicly available GUI has been developed for researchers to estimate the MSP of Sustainable Aviation Fuel (SAF) based on biomass properties, plant capacity, and location.\u003c/p\u003e","manuscriptTitle":"Data-driven framework for the techno-economic assessment of sustainable aviation fuel from pyrolysis.","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-07-12 08:03:39","doi":"10.21203/rs.3.rs-4595354/v1","editorialEvents":[{"type":"communityComments","content":0},{"type":"reviewerAgreed","content":"","date":"2024-06-22T23:04:31+00:00","index":0,"fulltext":""},{"type":"reviewersInvited","content":"","date":"2024-06-21T20:57:12+00:00","index":"","fulltext":""},{"type":"editorInvited","content":"BioEnergy Research","date":"2024-06-18T07:22:09+00:00","index":"","fulltext":""},{"type":"editorAssigned","content":"","date":"2024-06-18T06:39:28+00:00","index":"","fulltext":""},{"type":"submitted","content":"BioEnergy Research","date":"2024-06-17T12:26:08+00:00","index":"","fulltext":""}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"bioenergy-research","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":false,"externalIdentity":"bere","sideBox":"Learn more about [BioEnergy Research](https://www.springer.com/journal/12155)","snPcode":"12155","submissionUrl":"https://submission.nature.com/new-submission/12155/3","title":"BioEnergy Research","twitterHandle":"","acdcEnabled":true,"dfaEnabled":true,"editorialSystem":"em","reportingPortfolio":"Springer Hybrid","inReviewEnabled":true,"inReviewRevisionsEnabled":false}}],"origin":"","ownerIdentity":"bb94d4ae-d7ed-4a70-b0c2-d2aac20db71c","owner":[],"postedDate":"July 12th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"published-in-journal","subjectAreas":[],"tags":[],"updatedAt":"2024-12-09T16:10:04+00:00","versionOfRecord":{"articleIdentity":"rs-4595354","link":"https://doi.org/10.1007/s12155-024-10803-x","journal":{"identity":"bioenergy-research","isVorOnly":false,"title":"BioEnergy Research"},"publishedOn":"2024-12-02 15:57:34","publishedOnDateReadable":"December 2nd, 2024"},"versionCreatedAt":"2024-07-12 08:03:39","video":"","vorDoi":"10.1007/s12155-024-10803-x","vorDoiUrl":"https://doi.org/10.1007/s12155-024-10803-x","workflowStages":[]},"version":"v1","identity":"rs-4595354","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4595354","identity":"rs-4595354","version":["v1"]},"buildId":"qtupq5eGEP_6zYnWcrvyt","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.