Machine-learning-based bitter taste threshold prediction model for bitter substances: fusing molecular docking binding energy with molecular descriptor features | Research Square window.SnipcartSettings = { analytics: { enabled: false } }; (function() { var accessVector = localStorage.getItem('access_vector') || ''; window.dataLayer = window.dataLayer || []; if (accessVector) { window.dataLayer.push({ user: { profile: { profileInfo: { snid: accessVector } } } }); } })(); (function(w,d,s,l,i){w[l]=w[l]||[];w[l].push({'gtm.start':new Date().getTime(),event:'gtm.js'});var f=d.getElementsByTagName(s)[0],j=d.createElement(s),dl=l!='dataLayer'?'&l='+l:'';j.async=true;j.src='https://www.googletagmanager.com/gtm.js?id='+i+dl;f.parentNode.insertBefore(j,f);})(window,document,'script','dataLayer','GTM-K279D39R'); Browse Preprints In Review Journals COVID-19 Preprints AJE Video Bytes Research Tools Research Promotion AJE Professional Editing AJE Rubriq About Preprint Platform In Review Editorial Policies Our Team Advisory Board Help Center Sign In Submit a Preprint Cite Share Download PDF Research Article Machine-learning-based bitter taste threshold prediction model for bitter substances: fusing molecular docking binding energy with molecular descriptor features Can Chen, Haichao Deng, Huijie Wei, Yaqing Wang, Ning Xia, Jianwen Teng, and 2 more This is a preprint; it has not been peer reviewed by a journal. https://doi.org/ 10.21203/rs.3.rs-4439031/v1 This work is licensed under a CC BY 4.0 License Status: Posted Version 1 posted You are reading this latest preprint version Abstract Establishing the bitterness threshold of molecules is vital for their application in healthy foods. Although numerous studies have utilized Mathematical algorithms to identify bitter chemicals, few models can accurately forecast the bitterness threshold. This study investigates the binding mode of bitter substances to the TAS2R14 receptor, establishing the relationship between the threshold and binding energy. Subsequently, a structure-taste relationship model was constructed using random forest (RF), extreme gradient boosting (XGBoost), categorical boosting (CatBoost), and gradient boosting decision tree (GBDT) algorithms. Results showed R-squared values of 0.906, 0.889, 0.936, and 0.877, respectively, suggesting a relatively good predictive capability for the bitterness threshold. Among these models, CatBoost performed optimally. The CatBoost model was then employed to predict the bitter thresholds of 223 compounds. The model provides a precise reference for detecting the bitterness thresholds of a wide range of chemicals and dangerous substances. Binding Energy Bitterness threshold Machine learning Molecular docking TAS2R14 Figures Figure 1 Figure 2 Figure 3 Figure 4 Figure 5 Figure 6 INTRODUCTION In the study of taste perception, bitter taste has always been regarded as a fundamental taste[ 1 ]. Many foods and drinks with bitter taste are widely accepted and consumed, and they offer positive health effects. For instance, three of the most popular beverages – tea, coffee, and chocolate – are known for their distinctive bitter flavors and aromas[ 1 ]. Additionally, the overall flavor of beer, wine, cider, and citrus juice is inextricably linked to bitterness. However, excessive bitterness can be unpleasant. The amount of bitterness directly impacts target demographics and customer acceptance. Therefore, identifying bitter substances in foods and managing them as a way to process foods with appropriate bitter flavors can have a positive economic impact. Currently, sensory assessment by humans is the method most commonly used method to describe the bitter flavor. The threshold test and scale approach are generally utilized to measure bitterness intensity. The lowest sample concentration, which each evaluator perceives as bitter after two or three trials, is considered the bitterness threshold[ 2 ]. The Dose-over-threshold (DoT) factor, known as the bitterness active value[ 3 ], is frequently employed with the bitterness threshold and computed as the ratio of the individual concentration to the corresponding tastant's recognition threshold[ 4 ]. A DoT factor ≥ 1 indicates that the bitter tastant contributes to the overall bitterness intensity. The higher the DoT factor, the more bitter the tastant exhibits, and thus the more the tastant contributes to the overall bitterness intensity[ 1 ]. At present, commonly employed techniques for determining the key bitter substances in food include the following: the DoT factor method that determines whether the substance is the key bitter substance in food by directly obtaining the bitter threshold and DoT factor of the substance[ 4 ]; Taste dilution analysis (TDA) method that is utilized to grade food solutions, determine the TD factor between different fractions by continuously diluting the fractions, and finally determine the substances in the most bitter fraction, to determine the key bitter substances in the food[ 5 ]; Spectrum descriptive analysis method (SDAM) is also used to remove the dilution process. It directly uses different concentrations of a bitter substance as a strength reference for strength. It evaluates the bitterness strength of various fractions in the food solution to determine its bitter substances[ 6 ]. These techniques can be employed alongside a variety of high-precision experimental tools, such as nuclear magnetic resonance (NMR) and liquid chromatography mass spectrometry (LC-MS). The advantage of the DoT method over TDA and SDAM is that it eliminates the laborious fractionation stage and combines the bitterness threshold, allowing for quick quantification of known bitter chemicals and a significant reduction in experimental components. However, the bitterness threshold approach has its drawbacks. Firstly, obtaining a single bitter substance can be challenging, often requiring separation and purification through column chromatography or artificial synthesis. Secondly, it is impossible to establish the threshold of toxic bitter chemicals using this method. Thirdly, many sensory evaluators are necessary to evaluate bitter compounds. Lastly, determining the bitter taste threshold involves a cumbersome experimental process and calculation method[ 1 ]. In recent years, various mathematical models have been developed to determine the bitterness of a single molecule based on its chemical structure. Bitter taste classification models employing machine learning techniques (e.g., SVM, KNN, GBM, RF, and DNN) include BitterX[ 7 ], e-bitter[ 8 ], Bittersweet[ 9 ], BitterPredict[ 10 ], BitterSweetForest[ 11 ], and Premexotac[ 12 ]. Bittermatch[ 13 ] is employed to Predict the binding of bitter substances to bitter receptors based on the GBDT algorithm and chemical descriptors. On the other hand, models utilizing different algorithms, such as the iBitter-SCM[ 14 ] and the BERT4Bitter[ 15 ] bitter peptide prediction models, have been employed. This is based on a bidirectional encoder representation from transformers (BERT) based model and the scoring card method (SCM) model, which is a unique algorithm that is different from machine learning, and it provides a different method of classifying bitter peptides by determining the amino acid sequence of the peptide to predict the bitter peptide. The appellate models have made significant contributions to identifying bitter substances and investigating bitter taste mechanisms; however, they cannot predict the intensity of bitter substances or the bitter taste threshold. It is important to note that identifying bitter compounds is only the first step in identifying key bitter substances in food. In further investigation, it is necessary to screen all bitter compounds and identify those that have the greatest impact on food bitterness. Therefore, it is crucial to establish a prediction model for the bitterness threshold. The BitterIntense model classifies bitter substances into two categories according to their thresholds: below 0.1 mM as very bitter and above 0.1 mM as mild bitter, and then it performs bitter threshold classification modeling[ 16 ]. Although this approach roughly predict the bitter threshold of the substance, it could not provide exact bitter threshold data for bitter substances. Molecular docking technology is extensively employed for investigating ligandprotein receptor intera-ctions and for virtual screening purposes[ 17 ]. Presently, 19 activated amino acid residues within TAS2R14 have been rigorously examined in experimental studies[ 18 ]. Based on this empirical evidence, it can be inferred that substances which dock with these specific amino acid residues can be effectively screened using molecular docking technology. Additionally, the binding energy of the bitter substance bound to TAS2R14 can be determined through the application of the LibDock algorithm and LibDockScore. Consequently, the bitter taste threshold can be derived from the binding energy associated with this particular conformation, when combined with machine learning techniques. The utilization of computer simulations holds the potential to substantially reduce the number of required experiments, equipment costs, and labor involved in these investigations. Based on the aforementioned information, this study involved machine learning modeling of 96 bitter compounds with known thresholds, using binding energy in molecular docking and chemical descriptors as features. Various algorithms were employed to develop the models, including GBDT, XGBoost, CatBoost, and RF algorithms. The main objectives of this study are as follows: Firstly, to establish an accurate model for predicting the bitterness threshold of specific bitter compounds. Secondly, to explain the binding mechanism between bitter substances and bitter receptor TAS2R14 using molecular docking technology. Thirdly, to identify the features related to the bitterness threshold and select the most accurate features for more modeling methods. Lastly, To predict the bitterness threshold of 223 bitter compounds with unidentified bitter thresholds. MATERIALS AND METHODS 2.1 Data collection 262 substances were collected from the BitterDB website ( https://bitterdb.agri.huji.ac.il/dbbitter.php ) and some literature sources. The 3D structures of bitter substances were obtained according to the Pubchem ID from the Pubchem website ( https://pubchem.ncbi.nlm.nih.gov/ ), and the structures were downloaded in sdf format. The thresholds are transformed according to the following equation PT=-lg(T×10 − 6). where T represents the bitterness threshold in mg/kg and PT is the bitterness threshold after processing, as shown in Table S1 . Due to the absence of X-ray crystallographic structure for TAS2R14, several models representing the 3D structure of the 25 human TAS2Rs have been developed, leveraging the high homology among GPCR receptors (Paredes Ramos & M López Vilariño, 2023)[ 19 ]. The protein model of TAS2R14 was obtained from the AlphaFoldDB website, numbered Q9NYV8[ 20 , 21 ]. After performing molecular docking, the small molecule conformation with the highest LibDockScore was selected from multiple results of each molecule and exported to sdf format for subsequent molecular descriptor calculation. Finally, the 96 substances were successfully docked to obtain the binding energy. 2.2 Molecular docking of the bitter substances with TAS2R14 The water molecule present in the TAS2R14 structure was removed and replaced with a hydrogen atom. Discovery Studio software (DS) of the Receptor-Ligand Interactions module was used to define seven active sites, and the first one was selected. This site corresponds to the binding site proposed in current research for TAS2R14[ 18 ]. The Prepare Ligands tools of the Small Molecules module were used to optimize the energetic conformations. DS generated all possible optimal conformations of the small molecules. In the LibDock protocol of the Discovery Studio software, the first site was located at coordinates: x = 12.209, y = 3.273, z = -7.255, with a radius of 17.4 Å. For each optimal conformation set, up to 10 docking results were generated. The molecular docking results were evaluated based on the LibDockScore between the peptides and T2R14. After performing molecular docking, the small molecule conformation with the highest LibDockScore was selected from multiple results of each molecule and exported to sdf format for subsequent molecular descriptor calculation. Then, binding energy of the ligand–protein complex was calculated by using “Calculate Binding Energy” protocol which uses CHARMm implicit solvent model. Discovery Studio software generated a 2D results map of the molecular docking process. Finally, the 96 substances were successfully docked to obtain the binding energy. 2.3 Molecular descriptor calculation and feature processing. The 208 descriptors of structural features of the 96 bitter compounds were computed using the MolecularDescriptorCalculator module of the RDKit package. Due to a high number of zero values for some descriptors, they were removed as unsuitable features. The final descriptors used are displayed in Table S2 . The Scikit-learn library was used for machine learning modeling. The Python machine learning package Scikit-learn offers a variety of commonly used machine learning tools and techniques, such as model selection, dimensionality reduction, clustering, classification, regression, and preprocessing modules. In this study, four algorithms were employed, namely GBDT, XGBoost, CatBoost, and RF. The 96 data sets were randomly divided into a ratio of 8:2 for the training and test sets. The training set and testing set were then normalized using the formula Z=(X-u)/s, where X represents the input property descriptor data, u is the mean of the training set and s is the standard deviation of the training set. This normalization process ensures that the variance of the data in each dimension is 1 and the mean is 0, preventing features with excessively large dimensions from influencing the prediction outcomes. Since there are a large number of remaining features, feature selection is performed on the training set using the embedding method in sklearn in order to prevent dimensionality loss. A 5-fold cross-validation was applied exclusively on the training data to identify the optimal hyperparameters for each method from empirical candidates based on the cross-validated R2, which were used to estimate the performance of the different models. 2.4 Regression Modeling and Model Evaluation R-Square (R2), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE) were utilized to evaluate the model. The statistical metric R2 is used to assess how well the regression model fits the data. The average of the absolute differences between the model predictions and the actual observations is known as the Mean Absolute Error (MAE). It reflects the magnitude of the true error. MSE measures the mean squared difference between the model predictions and the actual observations. RMSE is the square root of the mean squared difference between the model predictions and the actual observations. It measures the average magnitude of the prediction errors, taking into account the squared differences between the predicted and actual values. The formulas of these four values are expressed as follows: $${\text{R}}^{2}=1-\frac{{\text{S}\text{S}}_{\text{r}\text{e}\text{s}}}{{\text{s}\text{s}}_{\text{t}\text{o}\text{t}}}=1-\frac{{\sum }_{\text{i}}{({\text{y}}_{\text{i}}-{\text{f}}_{\text{i}})}^{2}}{{\sum }_{\text{i}}{({\text{y}}_{\text{i}}-\stackrel{-}{\text{y}})}^{2}}$$ 1 $$\text{M}\text{A}\text{E}=\frac{1}{\text{n}}\sum _{\text{i}=1}^{\text{n}}\left|{\left({\text{y}}_{\text{i}}-{\text{f}}_{\text{i}}\right)}^{2}\right|$$ 2 $$\text{M}\text{A}\text{S}\text{E}=\frac{1}{\text{n}}\sum _{\text{i}=1}^{\text{n}}{\left({\text{y}}_{\text{i}}-{\text{f}}_{\text{i}}\right)}^{2}$$ 3 $$\text{R}\text{M}\text{S}\text{E}=\sqrt{\frac{1}{\text{n}}\sum _{\text{i}=1}^{\text{n}}{\left({\text{y}}_{\text{i}}-{\text{f}}_{\text{i}}\right)}^{2}}$$ 4 where yi, fi, and y̅ represent the true value, predicted value, and average value, respectively. SSres is the sum of squares of residuals, known as the residual sum of squares. SStot is the total sum of squares. n is the number of samples. A model interpretation package called Shapley additive explanation (SHAP) ( https://github.com/slundberg/shap ) was created in Python and put into action by Python scripts. It was employed to assess the feature importance of the four models (GBDT, RF, CatBoost, and XGBoost). 2.5 Prediction of unknown bitterness threshold Obtained 223 substances from the BitterDB website that are only identified to have bitterness but with undetermined bitterness thresholds. The binding energies for 223 of these substances were then determined by molecular docking, following the method introduced in section 2.2 . The molecular descriptor and feature scaling were then determined in accordance with the method introduced in section 2.3. The prediction of PT was conducted using the CatBoost model and transformed into a bitter threshold. 2.6 Statistical Analysis All analyses were implemented on the Anaconda3 platform, including scikit-learn joblib, pandas, numpy, and matplotlib for data processing and analysis. RESULTS AND DISCUSSION 3.1 grouping of 262 compounds A dataset of 262 samples was obtained from the BitterDB website and some literature, which compiles information on the bitterness of various foods from several resources, including Gouda cheese, fennel, beer, whole wheat bread crumbs, oats, red wine, asparagus, carrots, and roasted cocoa nibs. These substances mainly include polypeptides, alkaloids, phenols, terpenoids, and saponins (Table S1 ). In recent years, many studies have been conducted on Quantitative Structure-Activity Relationship modeling using bitter peptides based on molecular descriptors and the efficacy of some molecular peptides. Bitter peptides are mainly produced by proteins during the food fermentation process and offer significant benefits for cardiovascular, neurological, immune, and nutritional systems[ 22 ]. Phenolic compounds, which are organic compounds found abundantly in plants, have gained attention as an emerging field of nutrition in recent decades. A growing body of evidence suggests that the intake of polyphenols may play a vital role in regulating metabolism, body weight, chronic diseases, and cell proliferation[ 23 ]. Phenolic compounds contribute to the bitterness and astringency of various beverages and fruits, such as oranges, coffee, and tea[ 1 ]. Alkaloids, a class of nitrogenous organic compounds, are present in natural constituents of vegetables and drinks, some food processing additives, and pollutants[ 24 ]. Most alkaloids exhibit a bitter taste and possess a potent biological activity, some of which may be toxic to humans. Terpenoids (terpenes and their oxygenated derivatives) are important secondary metabolites found in medicinal plants and some well-known foods and beverages, contributing to their pronounced bitterness. Saponins, on the other hand, are secondary plant metabolites composed of steroids, triterpenes, or alkaloids with one or more sugar chains[ 25 ]. They are bitter and astringent substances found in certain foods. Additionally, fatty acids, fatty alcohols and other categories are also important bitter substances. 3.2 Binding mode of bitter protein TAS2R14 with bitter substances The binding residues of the TAS2R14 protein that have been experimentally identified include Trp66, Leu85, Thr86, Asn87, Trp89, Thr90, Asn93, His94, Thr182, Ser183, Phe186, Ile187, Tyr240, Ala241, Phe243, Phe247, Ile263, Gln266, and Gly269 [ 18 ]. Molecular docking showed similar results to the experimental outcome, as shown in Fig. 1 . The binding pocket of the TAS2R14 receptor is quite spacious and can tolerate agonist structures that largely exceed the size of unmodified agonists [ 26 ]. No residues were found in the TAS2R14 binding pocket where mutations resulted in a complete loss of function of all agonists, whereas studies of other receptors have identified such restricted targets. The Asn24 and Ile27 residues of TAS2R1[ 27 ]. Gly281, and Ser285 residues of TAS2R4[ 28 ], Trp261 residues of TAS2R16[ 29 ] have been found to be important binding sites for activating proteins. Moreover, TAS2R14 provides a large number of agonist-selective contacts that may outweigh all other promiscuous TAS2Rs[ 18 ]. After screening substances that respond to the appeal TAS2R14 binding site, the docking site map is shown in Fig. 1 and Table S3 . And the binding energies of the 96 successfully docked substances are shown in Table S1 . Regarding substance categorization, almost all classes bound a portion of their substances to TAS2R14, further demonstrating the broad coordination ability of TAS2R14 [ 18 , 26 , 30 ]. Nearly all of the 24 fatty acids, fatty alcohols, and glycerolipids were successfully docked, and all three belong to long-chain lipids with a very long carbon chain that allows easier access to the centers of the TM1 and TM2 chains and less site blockage. These three substances demonstrate high LibDockScore (> 100). LibDockScore is a scoring algorithm that rates the degree of ligand-protein stabilization[ 31 ]. Therefore, it indicates that long-chain analogs can form more stable complexes with TAS2R14 proteins, and some of them have low thresholds (< 1mg/Kg). In addition, most of Cinnamic acids and derivatives and flavonoids bind to TAS2R14. The mode of interaction between the TAS2R14 receptor and the ligand has received extensive attention. The results of the binding site reflect two patterns, as shown in Fig. 1 and Table S3 . while alkaloids, amino acids and their derivatives, and benzenes do not interact with Lys8 and Trp89 and mainly interact hydrophobically with Ile187, Cinnamic acids and their derivatives, Prenol lipids, and flavonoids interact hydrophobically mainly with Leu85 and Trp89, and they partially bind to Ile187. Interestingly, the first group usually had lower LibDockScore and higher thresholds, suggesting that unstable binding affects the bitterness thresholds. It can be found that substances with binding energies below − 100 have low thresholds (< 50mg/Kg), while some substances with high thresholds, such as amino acids, have high binding energies. The magnitude of the binding energy indicates the strength of the interaction, and substances with low binding energies indicate that they interact more strongly with TAS2R14, consequently exhibiting stronger bitterness intensity and lower bitterness threshold[ 32 ]. However, there are some exceptions. Long-chain substances have the lowest threshold, yet their binding energies vary, indicating that the threshold is not solely influenced by binding strength. Long-chain substances might more readily bind with TAS2R14 to activate it, but the binding interactions are not necessarily strong. Figure 2 shows the binding of the five known TAS2R14 activators. In agreement with some findings, Leu85 and Trp89 are amongst the most important binding residues[ 33 ]. Four substances, Lupulone, Adlupulone, cis-isohumulone, and trans-Isoadhumulone, were found to interact hydrophobically with Leu85 and Trp89, with the main interactions being alkyl and pi-alkyl. Residue Trp89 is deep in the binding pocket[ 34 ]; thus, the ability to bind to residue Trp89 may be more stable. Residue Thr86 also undergoes hydrogen bonding interactions with trans-Isoadhumulone cis-isohumulone, with Thr86 and Asn93 being the most common hydrogen bonding interactions in the study[ 26 , 33 ]. In this study, both residues also show binding to a great extent, with residue Thr86 having more binding ability than Asn93, especially since the flavonoid glycosides almost all interact with THR86 by hydrogen bonding. In addition, some studies have concluded that flufenamic acid, mainly pi-pi, interacts with Trp89, Phe186, and Phe247[ 35 ]. The mechanism of action of the benzoic acid analog TAS2R14 includes p-p interactions between the ligand and Phe186, Phe243, and His94 as well as hydrogen bonding with Asn93 and Thr86[ 26 ]. These binding residues were also found in our study. With further research, we will explore more binding sites and activators, which will deepen our study of the binding mode of TAS2R14 and the relationship between binding energy and threshold. 3.3 Feature selection The interaction between bitter compounds and bitter receptors is the key factor underlying bitterness, and the strength of this interaction determines the intensity of bitterness[ 32 , 36 ]. Binding energy as a quantitative measure for evaluating the strength of receptor-ligand binding is closely related to the bitter taste threshold. Since only one receptor's binding energy can be used as a feature for modeling, using receptors that are able to bind a wide range of substances will improve our model's ability to predict substance types. A total of 25 bitter receptors have been identified to bind various types of bitter substances. However, TAS2R10, TAS2R14, and TAS2R46 are broadly tuned receptors known to respond to a variety of natural and artificial bitter compounds, while TAS2R3, TAS2R5, TAS2R8, TAS2R13, TAS2R41, TAS2R49, and TAS2R50, are activated by very few bitter compounds, and TAS2R42, TAS2R45, TAS2R48, and TAS2R60 are currently unidentified substances[ 1 ]. TAS2R10 and TAS2R14 showed less preference among agonists, while TAS2R14 could bind more chemical scaffolds and more agonists[ 26 , 30 ]. Most receptors embody a high degree of selectivity for structure; for example, TAS2R16 and TAS2R38 show a strong tendency to recognize β-dglucopyranosides and N–C = S group in molecules[ 1 ]. In addition, TAS2R7 is suited for receptors for metal cations such as zinc, magnesium, calcium, aluminum, copper and manganese[ 37 ]. TAS2R46, on the other hand, tends to detect sesquiterpene lactones and related compounds[ 18 ]. Molecular descriptors are numerical features used to represent and describe the structure of a molecule, and the MolecularDescriptorCalculato module calculates 208 molecular descriptors, with 93 descriptors remaining after removing 0 and most of the descriptors with 0 (and adding binding energies) as shown in Table S2 . Feature selection was performed on the training set by embedding method decision tree. The top 15 features selected according to gain are shown in Fig. 3 . Due to the stochastic nature of decision tree feature selection, to prevent some features from being missed, some of the features were selected from the remaining 78 features that were considered to be important at the time the other bitterness models were built. Weichen Bo et al[ 38 ]. used an artificial neural network algorithm to build a bitter-non-bitter classification model, and based on the important properties they derived to classify bitter substances, we chose heavy atom molecular weight (HeavyAtomMolWt), the hydrophobicity feature (BCUT2D_LOGPHI), backbone atom electronegativity (EState_VSA1 and EState_VSA4), molecular shape (Kappa2). MinEStateIndex was added based on the conserved site modeling[ 33 ], in addition, BitterIntense, as a model for classification thresholding, also provides two important features, heavyatom and molar refractivity. HeavyAtom is replaced by HeavyAtomMolWt, molar refractivity is replaced by SMR_VSA5. Finally, we added molecular descriptors, as shown in Fig. 3 , and them a correlation analysis was performed, removing substances with too much correlation to reduce multicollinearity, and the final results are shown in Fig. 3 . The predictive model with higher R-squared bitter taste thresholds was developed through multiple selections, and the feasibility of the binding energy as a feature was demonstrated by the fact that the binding energy was ranked 12 and successfully retained from multiple selections. Since the randomness of the embedding method does not fully represent the importance of the features, they will be further evaluated by the SHAP algorithm in the modeling below. 3.4 Machine learning model analysis Four regression models, namely RF, XGBoost, CatBoost and GBDT, were employed for the prediction of bitterness thresholds, and the prediction results are shown in Table S4 . Figure 4 displays the optimal model results for each algorithm, which demonstrates a good correlation in both the training set and the test set. The R2 values of PT in the training set were 0.916, 0.893, 0.948 and 0.893, respectively. Similarly, the R2 values of PT in the test set were 0.906, 0.889, 0.936, and 0.877, respectively. The slight difference between the R2 of the training set and the test set indicates a low risk of overfitting. The findings demonstrate that all four machine learning models could successfully estimate the PT of a substance, with the CatBoost algorithm showing the best results in both the training and test sets. To further evaluate the models, three metrics, namely MAE, MSE, and RMSE, were utilized. These three indicators assess the degree of error between the true and predicted values, and the closer the three indicators are to 0, the better the model works. Table 1 demonstrates that the MAE, MSE, and RMSE exhibit opposite tendencies to R2 for these four models. Among these models, the three indicators of CatBoost were the smallest (i.e., 0.936, 0.317, 0.174, 0.417, respectively). Hence, the CatBoost model demonstrated the highest effectiveness and can be used to forecast the bitterness threshold in future studies. In addition, models based on RF and GBDT algorithms also displayed good results. Table 1 Model Evaluation Algorithm R 2 MAE MSE RMSE GBDT 0.877 0.477 0.333 0.577 XGBoost 0.889 0.423 0.299 0.546 RF 0.906 0.370 0.253 0.503 CatBoost 0.936 0.317 0.174 0.417 Table S4 lists the training and true values for all test sets, and due to the high R-squared of the model itself, most of the substance predictions are almost identical to the true values. However, there are some flaws in the model; the four categories of Chalcones, Flavanols, and Organooxygen compounds appeared less in the training set, the Flavanols category was not trained, and the remaining 2 types had 1–3 substances trained in the training set. As can be observed, the prediction value of Flavanols is too high. At the same time, the other two substances are predicted accurately, indicating that the model is adequately trained for the training set, and it can effectively predict the various structures of the substances appearing in the training set. However, it may have poorer prediction ability of unknown types. This needs to be improved by collecting more data. Secondly, Amino acids, Peptides, Benzoic acids, and derivatives also showed inaccurate predictions, which is a category of substances that the model needs to focus on subsequently. In contrast, the rest of Glycerolipids, Flavonoid glycosides, and Fatty acids were more accurate. Among the four algorithms used, GBDT, RF, CatBoost, and XGBoost are all integrated algorithms[ 39 ]. Ensemble algorithms combine multiple simple models to create a more powerful model, thereby improving prediction performance. An ensemble technique built on a gradient boosting decision tree is called XGBoost or CatBoost, while GBDT and RF are ensemble approaches based on decision trees[ 39 ]. The performance of the model heavily relies on the selection of features. Since our feature selection method uses decision trees, our four algorithms are based on decision trees. RF and GBDT are two distinct Decision Tree-based methods. RF constructs numerous randomized decision trees to achieve integration, whereas GBDT iteratively optimizes the model[ 40 ]. RF is more suitable for high-dimensional problems. Both CatBoost and XGBoost are GBDT-based algorithms. XGBoost is fast in terms of execution speed and memory efficiency, and it can fit the data better and find the global optimal solution faster, so it is more suitable for large data volumes. CatBoost, on the other hand, employs ordered boosting strategies to prevent prediction bias caused by target leakage [ 41 ]. Due to the small sample size, CatBoost has a built-in regularization mechanism that reduces the risk of overfitting and thus performs better when dealing with small samples or noisy data. 3.5 Feature analysis of 4 models In Fig. 5 A, the feature importance of the CatBoost, GBDT, XGBoost and RF model is ranked by SHAP and measured based on the absolute value of the average Shapley value. Among these models, BCUT2D_MRLOW, Chi2v, PEOE_VSA7 are the three highest ranked feature. Figure 5 B presents the summary_polt, which illustrates the effect of each feature on the CatBoost model. In the plot, the PEOE_VSA7, Chi2v and SMR_VSA5 and Binding Energy points red and blue areas at the same time on value < 0, Explain that these features have a positive effect on the predicted value of PT (There is some degree of inverse relationship between the eigenvalues and the results predicted by the model). the BCUT2D_MRLOW points are clustered into two regions: the red region with positive SHAP values (> 0) and the blue region with negative SHAP values (< 0). This indicates that higher BCUT2D_MRLOW values have a positive effect on PT, while lower BCUT2D_MRLOW values have a negative effect. Figure 5 C shows the force distribution of the CatBoost model (arranged by similarity). Based on the distribution, it can be observed that the predictive pattern of the features has two ways for all the samples of the model, and the main difference is that the BCUT2D_MRLOW feature contributes positively and negatively to the model. Chi2v is a molecular descriptor based on atomic charge distribution and molecular topology used to encode the electrical, chemical reactivity, and structural characteristics of molecules. The SMR_VSA descriptor explains the molar refractive index of a molecule and its effect within the receptor or on the way to the receptor. SMR_VSA5 is defined as the sum of the vi of all atoms i. pi denotes the contribution to the molar refractive index of atom i calculated in the SMR descriptor, computed in the range of 0.440 to 0.485[ 42 ]. The BCUT2D descriptor is based on the Burden matrix, which encodes the bond strengths between atoms in a molecule. It changes the diagonal elements of the matrix to include atomic properties and then performs eigenvalue decomposition to obtain the highest and lowest eigenvalues. PEOE_VSA8 (a molecular surface area descriptor). Mollogp, SlogP_VSA6 is related to the partition coefficient (logP) and is a characterization of the hydrophobicity of the molecule. Binding Energy, as mentioned earlier, is a descriptor for assessing the strength of ligand-receptor binding and is related to hydrophobic interactions, charge interactions, hydrogen bonding interactions, etc. These molecular descriptors and binding energies are important features for constructing bitter taste threshold models. Molecular descriptors are used to quantitatively describe the physical and chemical information of molecules[ 33 ]. Thus, the physical properties they respond to influence the magnitude of the bitterness threshold. Currently, there are several studies on the bittering mechanism, among which the recognized properties are hydrophobicity, which is the main property found in various types of studies, where researchers have suggested that substances with a strong bitter taste are accompanied by a high degree of hydrophobicity [ 32 , 35 , 38 ]. This relationship can be attributed to the influence of hydrophobic molecules on various factors, such as the distribution of molecules in the cell membrane, the ability to bind to bitter taste receptors, and the interaction force between molecules, thereby affecting the perceived intensity of bitter taste. The molar refractivity is a measure of the extent to which the electron density distribution within a molecule can be distorted[ 16 ]. Bitter taste prediction models Premexotac [ 12 ] and a recent artificial neural network prediction model[ 38 ] considered it as an important feature for bitterness classification, while BitterIntense[ 16 ] model is a prediction of the bitterness threshold and this model gained a high importance score in molar refractivity, SMR_VSA5 is a molar refractive index descriptor, and it was ranked 6th in the importance of the model, so we got the same conclusion that the distribution of the molecule's electron cloud affects the intensity of bitterness. The binding energy contributed to the construction of CatBoost model and RF model, but its importance score in GBDT and XGBoost model is almost 0. The comparison of the features shows that the importance of the binding energy is the 11th feature for constructing the bitterness threshold after hydrophobicity and Molar refractivity. Although the contribution of binding energy is 0 in GBDT and XGBoost models, the R-square of these two models is low (< 0.9), which suggests that the inability to find the relationship between the binding energy and the bitter taste threshold may be a reason for the low prediction effect of these two algorithms. Finally, this paper concludes that factors such as the electronic environment of the molecule, hydrophobicity, chemical reactivity, and the strength of binding to the TAS2R14 protein (binding energy) influence the magnitude of the bitterness threshold. 3.6 Threshold prediction for bitter substances Table S5 presents the predictions for 223 compounds, including their names, structure, binding energy, PT, threshold (mg/kg), Chi2v, PEOE_VSA7 and BCUT2D_MRLOW values. Figure 6 A shows the distribution of thresholds for 223 substances, and the result shows that the thresholds for most substances are concentrated in the range of 100–300 mg/kg. To further investigate these substances, a three-dimensional scatter plot of three features was drawn based on these substances, as shown in Fig. 6 B. From this plot, it can be observed that sample points with close bitterness thresholds are also closer together in the plot and show a gradual change in colour, indicating that the importance of these three features in regression modelling. Secondly, most of the substances with thresholds lower than 1000 are in the region where Chi2v is less than 2.5 and BCUT2D_MRLOW is greater than − 0.5, and the lower the Chi2v value, the higher the BCUT2D_MRLOW value and the lower the threshold. The substances with high thresholds (> 1000mg/Kg) are more likely to be affected by the BCUT2D_MRLOW feature because these substances are all at the bottom of the 3D map space; however, they are also modulated by PEOE_VSA7, which slightly decreases the thresholds of these substances as PEOE_VSA7 increases. CONCLUSION In this study, we initially analyzed the binding mode of bitter substances with the TAS2R14 receptor protein and identified the amino acid residues Lys85, Trp89, and Thr86 crucial for bitter substances to interact with T2R14. They interact with bitter compounds through hydrophobic and hydrogen bond interactions, respectively. Subsequently, a bitterness threshold prediction model was established based on four algorithms. CatBoost is an ensemble algorithm based on the GBDT algorithm. It’s highest r-square model was constructed and used to accurately predict the bitterness threshold of bitterness. Finally, the important factors that impact the prediction of bitterness threshold by the model are found to be molar refractivity, hydrophobicity, the electronic environment of the molecule, reactivity, and the strength of binding to the TAS2R14 protein. The main disadvantage of the current model is the small amount of data, and the determination of the bitterness threshold may vary across different studies of the same substance due to subjective factors. Therefore, collecting more bitter taste threshold data in the future will increase the generalization ability of the model. Secondly, we used the classical LibDock algorithm for molecular docking, while newer scoring functions based on machine learning are gradually being investigated[ 43 ], which may be able to be applied in the future for the assessment of the binding status of bitter substances to bitter receptors and bitter taste threshold modeling. However, it was undeniable that the bitterness threshold prediction models proposed in this study hold potential applications and it can eliminate complex sensory experiments to quickly determine the bitterness threshold of a large number of substances through initial screening of food bitter substances. Additionally, these models can serve as alternatives to sensory experiments in determining the bitterness threshold of some toxic substances and offer valuable guidance for the subsequent exploration and masking of key bitter compounds by determining the bitterness threshold. Abbreviations RF, Random forest; XGBoost, Extreme Gradient Boosting; CatBoost, Categorical Boosting; GBDT, Gradient Boosting Decision Tree; DOT, Dose-over-threshold; TDA, Taste Dilution Analysis; SDAM, Spectrum Descriptive Analysis Method; NMR, Nuclear Magnetic Resonance; LC-MS, Liquid Chromatography Mass Spectrometry; BERT, Bidirectional Encoder Representation from Transformers; SCM, Scoring Card Method; DS, Discovery Studio; R2, R-Square; MAE, Mean Absolute Error; MSE, Mean Squared Error; RMSE, Root Mean Squard Error; SHAP, Shapley Additive Pxplanation; SDF, Structure-Data File; PT, threshold after logarithmic processing; kNN, k-Nearest Neighbor; SVM, Support Vector Machine; GBM, Gradient Boosting Machine; DNN, Deep Neural Network. Declarations ACKNOWLEDGMENTS The authors would like to express their gratitude to EditSprings (https://www.editsprings.cn ) for the expert linguistic services provided. The authors would like to thank the high-performance computing platform of Guangxi University, Nanning, China. FUNDING SOURCES This work was supported by the National Natural Science Foundation of China (32160571), and Guangxi Key Research and Development Program (Guike AB21220068). SUPPORTING INFORMATION DESCRIPTION Bitterness threshold and classification of all substances (Table S1) (XLSX) All characteristics of a substance (Table S2) (XLSX) Docking results for 96 substances (Table S3) (XLSX) Predicted values of bitterness thresholds for four models (Table S4) (XLSX) Predicted bitter threshold for 223 bitter substances (Table S5) (XLSX) AUTHOR CONTRIBUTIONS Can Chen: Investigation, Data curation, Writing - original draft, Writing - review & editing. Haichao Deng: Data Curation, Formal Analysis. Huijie Wei: Investigation, Data curation. Yaqing Wang: Conceptualization, Supervision. Ning Xia: Conceptualization. Jianwen Teng: Formal analysis, Methodology. Qisong Zhang: Conceptualization. Li Huang: Conceptualization, Supervision, Resources, Formal analysis, Methodology, writing-review & editing NOTE The authors declare no competing financial interest. DATA AVAILABILITY The script and the parameter files are available in the GitHub repository: https://github.com/pandaness/Bitter-model.git References Yan J., Tong H. (2023) An overview of bitter compounds in foodstuffs: Classifications, evaluation methods for sensory contribution, separation and identification techniques, and mechanism of bitter taste transduction COMPREHENSIVE REVIEWS IN FOOD SCIENCE AND FOOD SAFETY 22:187-232 https://doi.org/10.1111/1541-4337.13067 Li H., Li L.F., Zhang Z.J., Wu C.J., Yu S.J. (2021) Sensory evaluation, chemical structures, and threshold concentrations of bitter-tasting compounds in common foodstuffs derived from plants and maillard reaction: A review CRITICAL REVIEWS IN FOOD SCIENCE AND NUTRITION 1-41 https://doi.org/10.1080/10408398.2021.1973956 Seo M.W., Yang D.S., Kays S.J., Lee G.P., Park K.W. (2009) Sesquiterpene Lactones and Bitterness in Korean Leaf Lettuce Cultivars HORTSCIENCE 44:246-249 https://doi.org/10.21273/HORTSCI.44.2.246 Scharbert S., Hofmann T. (2005) Molecular Definition of Black Tea Taste by Means of Quantitative Studies, Taste Reconstitution, and Omission Experiments JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 53:5377-5384 https://doi.org/10.1021/jf050294d Frank O., Ottinger H., Hofmann T. (2001) Characterization of an intense bitter-tasting 1H,4H-quinolizinium-7-olate by application of the taste dilution analysis, a novel bioassay for the screening and identification of taste-active compounds in foods JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 49:231-238 https://doi.org/10.1021/jf0010073 Liu X., Jiang D., Peterson D.G. (2014) Identification of bitter peptides in whey protein hydrolysate JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 62:5719-5725 https://doi.org/10.1021/jf4019728 Huang W., Shen Q., Su X., Ji M., Liu X., Chen Y., Lu S., Zhuang H., Zhang J. (2016) BitterX: a tool for understanding bitter taste in humans Scientific Reports 6 https://doi.org/10.1038/srep23450 Zheng S., Jiang M., Zhao C., Zhu R., Hu Z., Xu Y., Lin F. (2018) e-Bitter: Bitterant Prediction by the Consensus Voting From the Machine-Learning Methods Frontiers in Chemistry 6 https://doi.org/10.3389/fchem.2018.00082 Tuwani R., Wadhwa S., Bagler G. (2019) BitterSweet: Building machine learning models for predicting the bitter and sweet taste of small molecules Scientific Reports 9 https://doi.org/10.1038/s41598-019-43664-y Dagan-Wiener A., Nissim I., Ben Abu N., Borgonovo G., Bassoli A., Niv M.Y. (2017) Bitter or not? BitterPredict, a tool for predicting taste from chemical structure Scientific Reports 7 https://doi.org/10.1038/s41598-017-12359-7 Banerjee P., Preissner R. (2018) BitterSweetForest: A Random Forest Based Binary Classifier to Predict Bitterness and Sweetness of Chemical Compounds Frontiers in Chemistry 6 https://doi.org/10.3389/fchem.2018.00093 De León G., Fröhlich E., Fink E., Di Pizio A., Salar-Behzadi S. (2022) Premexotac: Machine learning bitterants predictor for advancing pharmaceutical development INTERNATIONAL JOURNAL OF PHARMACEUTICS 628:122263 https://doi.org/10.1016/j.ijpharm.2022.122263 Margulis E., Slavutsky Y., Lang T., Behrens M., Benjamini Y., Niv M.Y. (2022) BitterMatch: recommendation systems for matching molecules with bitter taste receptors Journal of Cheminformatics 14 https://doi.org/10.1186/s13321-022-00612-9 Charoenkwan P., Yana J., Schaduangrat N., Nantasenamat C., Hasan M.M., Shoombuatong W. (2020) iBitter-SCM: Identification and characterization of bitter peptides using a scoring card method with propensity scores of dipeptides GENOMICS 112:2813-2822 https://doi.org/10.1016/j.ygeno.2020.03.019 Charoenkwan P., Nantasenamat C., Hasan M.M., Manavalan B., Shoombuatong W. (2021) BERT4Bitter: a bidirectional encoder representations from transformers (BERT)-based model for improving the prediction of bitter peptides BIOINFORMATICS 37:2556-2562 https://doi.org/10.1093/bioinformatics/btab133 Margulis E., Dagan-Wiener A., Ives R.S., Jaffari S., Siems K., Niv M.Y. (2021) Intense bitterness of molecules: Machine learning for expediting drug discovery Computational and Structural Biotechnology Journal 19:568-576 https://doi.org/10.1016/j.csbj.2020.12.030 Zhao W., Su L., Huo S., Yu Z., Li J., Liu J. (2023) Virtual screening, molecular docking and identification of umami peptides derived from Oncorhynchus mykiss Food Science and Human Wellness 12:89-93 https://doi.org/10.1016/j.fshw.2022.07.026 Nowak S., Di Pizio A., Levit A., Niv M.Y., Meyerhof W., Behrens M. (2018) Reengineering the ligand sensitivity of the broadly tuned human bitter taste receptor TAS2R14 Biochimica et Biophysica Acta (BBA) - General Subjects 1862:2162-2173 https://doi.org/10.1016/j.bbagen.2018.07.009 Paredes Ramos M., M López Vilariño J. (2023) Hop bitterness in beer evaluated by computational analysis JOURNAL OF THE INSTITUTE OF BREWING 129 https://doi.org/10.58430/jib.v129i2.20 Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Žídek A., Potapenko A. (2021) Highly accurate protein structure prediction with AlphaFold NATURE 596:583-589 https://doi.org/10.1038/s41586-021-03819-2 Varadi M., Anyango S., Deshpande M., Nair S., Natassia C., Yordanova G., Yuan D., Stroe O., Wood G., Laydon A. (2022) AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models NUCLEIC ACIDS RESEARCH 50:D439-D444https://doi.org/10.1093/nar/gkab1061 Agyei D., Tsopmo A., Udenigwe C.C. (2018) Bioinformatics and peptidomics approaches to the discovery and analysis of food-derived bioactive peptides ANALYTICAL AND BIOANALYTICAL CHEMISTRY 410:3463-3472 https://doi.org/10.1007/s00216-018-0974-1 Cory H., Passarelli S., Szeto J., Tamez M., Mattei J. (2018) The Role of Polyphenols in Human Health and Food Systems: A Mini-Review Frontiers in nutrition (Lausanne) 5:87 https://doi.org/10.3389/fnut.2018.00087 Chen C., Lin L. Alkaloids in Diet, in: J. Xiao, S.D. Sarker, Y. Asakawa(2019) Handbook of Dietary Phytochemicals, Springer Singapore, Singapore:1-35 Abdelrahman M., Jogaiah S. Isolation and Characterization of Triterpenoid and Steroidal Saponins, in: M. Abdelrahman, S. Jogaiah(2020) Bioactive Molecules in Plant Defense: Saponins, Springer International Publishing, Cham:59-78 Karaman R., Nowak S., Di Pizio A., Kitaneh H., Abu-Jaish A., Meyerhof W., Niv M.Y., Behrens M. (2016) Probing the Binding Pocket of the Broadly Tuned Human Bitter Taste Receptor TAS2R14 by Chemical Modification of Cognate Agonists Chemical Biology & Drug Design 88:66-75 https://doi.org/10.1111/cbdd.12734 Singh N., Pydi S.P., Upadhyaya J., Chelikani P. (2011) Structural Basis of Activation of Bitter Taste Receptor T2R1 and Comparison with Class A G-protein-coupled Receptors (GPCRs) JOURNAL OF BIOLOGICAL CHEMISTRY 286:36032-36041 https://doi.org/10.1074/jbc.M111.246983 Pydi S.P., Bhullar R.P., Chelikani P. (2012) Constitutively active mutant gives novel insights into the mechanism of bitter taste receptor activation JOURNAL OF NEUROCHEMISTRY 122:537-544 https://doi.org/10.1111/j.1471-4159.2012.07808.x Thomas A., Sulli C., Davidson E., Berdougo E., Phillips M., Puffer B.A., Paes C., Doranz B.J., Rucker J.B. (2017) The Bitter Taste Receptor TAS2R16 Achieves High Specificity and Accommodates Diverse Glycoside Ligands by using a Two-faced Binding Pocket Scientific Reports 7 https://doi.org/10.1038/s41598-017-07256-y Bayer S., Mayer A.I., Borgonovo G., Morini G., Di Pizio A., Bassoli A. (2021) Chemoinformatics View on Bitter Taste Receptor Agonists in Food JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 69:13916-13924 https://doi.org/10.1021/acs.jafc.1c05057 https://doi.org/10.1016/j.chroma.2019.460474 Fang Y., Chen S., Lin D., Yao S. (2019) A new tetrapeptide biomimetic chromatographic resin for antibody separation with high adsorption capacity and selectivity JOURNAL OF CHROMATOGRAPHY A 1604:460474 https://doi.org/10.1016/j.chroma.2019.460474 Acevedo W., González-Nilo F., Agosin E. (2016) Docking and Molecular Dynamics of Steviol Glycoside–Human Bitter Receptor Interactions JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 64:7585-7596 https://doi.org/10.1021/acs.jafc.6b02840 Cui Z., Zhang N., Zhou T., Zhou X., Meng H., Yu Y., Zhang Z., Zhang Y., Wang W., Liu Y. (2023) Conserved Sites and Recognition Mechanisms of T1R1 and T2R14 Receptors Revealed by Ensemble Docking and Molecular Descriptors and Fingerprints Combined with Machine Learning JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 71:5630-5645 https://doi.org/10.1021/acs.jafc.3c00591 Woo J.A., Castaño M., Goss A., Kim D., Lewandowski E.M., Chen Y., Liggett S.B. (2019) Differential long‐term regulation of TAS2R14 by structurally distinct agonists The FASEB Journal 33:12213-12225 https://doi.org/10.1096/fj.201802627RR Levit A., Nowak S., Peters M., Wiener A., Meyerhof W., Behrens M., Niv M.Y. (2013) The bitter pill: clinical drugs that activate the human bitter taste receptor TAS2R14 The FASEB Journal 28:1181-1197 https://doi.org/10.1096/fj.13-242594.v De León G., Fröhlich E., Salar-Behzadi S. (2021) Bitter taste in silico: A review on virtual ligand screening and characterization methods for TAS2R-bitterant interactions INTERNATIONAL JOURNAL OF PHARMACEUTICS 600:120486 https://doi.org/10.1016/j.ijpharm.2021.120486 Behrens M., Redel U., Blank K., Meyerhof W. (2019) The human bitter taste receptor TAS2R7 facilitates the detection of bitter salts BIOCHEMICAL AND BIOPHYSICAL RESEARCH COMMUNICATIONS 512:877-881 https://doi.org/10.1016/j.bbrc.2019.03.139 Bo W., Qin D., Zheng X., Wang Y., Ding B., Li Y., Liang G. (2022) Prediction of bitterant and sweetener using structure-taste relationship models based on an artificial neural network FOOD RESEARCH INTERNATIONAL 153:110974 https://doi.org/10.1016/j.foodres.2022.110974 Jabeur S.B., Gharib C., Mefteh-Wali S., Arfi W.B. (2021) CatBoost model and artificial intelligence techniques for corporate failure prediction TECHNOLOGICAL FORECASTING AND SOCIAL CHANGE 166:120658 https://doi.org/10.1016/j.techfore.2021.120658 Jun M. (2021) A comparison of a gradient boosting decision tree, random forests, and artificial neural networks to model urban land use changes: the case of the Seoul metropolitan area International journal of geographical information science : IJGIS 35:2149-2167 https://doi.org/10.1080/13658816.2021.1887490 Zhou B., Bartholmai B.J., Kalra S., Osborn T., Zhang X. (2021) Lung mass density prediction using machine learning based on ultrasound surface wave elastography and pulmonary function testing JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA 149:1318 https://doi.org/10.1121/10.0003575 Moorthy N.S.H.N., Cerqueira N.M.F.S., Ramos M.J., Fernandes P.A. (2012) QSAR and pharmacophore analysis of thiosemicarbazone derivatives as ribonucleotide reductase inhibitors MEDICINAL CHEMISTRY RESEARCH 21:739-746 https://doi.org/10.1007/s00044-011-9580-x Ballester P.J. (2019) Selecting machine-learning scoring functions for structure-based virtual screening Drug Discovery Today: Technologies 32-33:81-87 https://doi.org/10.1016/j.ddtec.2020.09.001 Additional Declarations No competing interests reported. Supplementary Files TableS1.xlsx TableS2.xlsx TableS3.xlsx TableS4.xlsx TableS5.xlsx Cite Share Download PDF Status: Posted Version 1 posted You are reading this latest preprint version Research Square lets you share your work early, gain feedback from the community, and start making changes to your manuscript prior to peer review in a journal. As a division of Research Square Company, we’re committed to making research communication faster, fairer, and more useful. We do this by developing innovative software and high quality services for the global research community. Our growing team is made up of researchers and industry professionals working together to solve the most critical problems facing scientific publishing. Also discoverable on Platform About Our Team In Review Editorial Policies Advisory Board Help Center Resources Author Services Accessibility API Access RSS feed Manage Cookie Preferences © Research Square 2026 | ISSN 2693-5015 (online) Privacy Policy Terms of Service Do Not Sell My Personal Information {"props":{"pageProps":{"initialData":{"identity":"rs-4439031","acceptedTermsAndConditions":true,"allowDirectSubmit":true,"archivedVersions":[],"articleType":"Research Article","associatedPublications":[],"authors":[{"id":307082518,"identity":"24510baf-dd48-4d0a-acb7-75180da88283","order_by":0,"name":"Can Chen","email":"","orcid":"","institution":"Guangxi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Can","middleName":"","lastName":"Chen","suffix":""},{"id":307082519,"identity":"e3aaf045-d1e3-41ed-924f-52e659af43df","order_by":1,"name":"Haichao Deng","email":"","orcid":"","institution":"Baihui Pharmaceutical Group co, LTD","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Haichao","middleName":"","lastName":"Deng","suffix":""},{"id":307082520,"identity":"20821b8d-e2c1-4fd0-9e72-3205bf4864cd","order_by":2,"name":"Huijie Wei","email":"","orcid":"","institution":"Guangxi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Huijie","middleName":"","lastName":"Wei","suffix":""},{"id":307082521,"identity":"9d722691-1674-492e-8ebd-d7a51c7e6ae7","order_by":3,"name":"Yaqing Wang","email":"","orcid":"","institution":"Guangxi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Yaqing","middleName":"","lastName":"Wang","suffix":""},{"id":307082522,"identity":"db899bad-9e2a-430a-b44b-e8c45a6fbbaf","order_by":4,"name":"Ning Xia","email":"","orcid":"","institution":"Guangxi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Ning","middleName":"","lastName":"Xia","suffix":""},{"id":307082523,"identity":"bc5e16f1-856b-4bda-9328-1662132c2aa0","order_by":5,"name":"Jianwen Teng","email":"","orcid":"","institution":"Guangxi University","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Jianwen","middleName":"","lastName":"Teng","suffix":""},{"id":307082524,"identity":"9bc788f4-61cd-4892-961a-54343912d20b","order_by":6,"name":"Qisong Zhang","email":"","orcid":"","institution":"Medical College, Guangxi University, Nanning","correspondingAuthor":false,"submittingAuthor":false,"prefix":"","firstName":"Qisong","middleName":"","lastName":"Zhang","suffix":""},{"id":307082525,"identity":"3f855f30-4574-4c26-ab0e-c5a132787a62","order_by":7,"name":"Li Huang","email":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAAZAAAAAyAQMAAABI0h/eAAAABlBMVEX///8AAABVwtN+AAAACXBIWXMAAA7EAAAOxAGVKw4bAAAAvklEQVRIiWNgGAWjYLCCDwb/5djYmw8Qr4NxRgWzMR/PsQTitTBznGFOnCeRo0Ccct32s8+kGdvY0tsYchgYflRsI6zF7Ey6mXRhG09uG8PZA4w9Z24ToeVAGpv0zDaJ3DbGvgRmxjZitJx/xibN22aQzsbMY0CklhtAW3jOJCSwsRGv5Rmz5YyKA4ZtPGwJB4nzy/k0xhsfDA7Iy89/fPDBjwoitAABiwSMdYAo9UDA/IFYlaNgFIyCUTBCAQAclzoCjEZg/wAAAABJRU5ErkJggg==","orcid":"","institution":"Guangxi University","correspondingAuthor":true,"submittingAuthor":false,"prefix":"","firstName":"Li","middleName":"","lastName":"Huang","suffix":""}],"badges":[],"createdAt":"2024-05-18 01:53:29","currentVersionCode":1,"declarations":"","doi":"10.21203/rs.3.rs-4439031/v1","doiUrl":"https://doi.org/10.21203/rs.3.rs-4439031/v1","draftVersion":[],"editorialEvents":[],"editorialNote":"","failedWorkflow":false,"files":[{"id":57438398,"identity":"47708d3e-211e-4fb6-ac23-59971a1253ea","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"png","order_by":1,"title":"Figure 1","display":"","copyAsset":false,"role":"figure","size":1437605,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eDocking results for 96 substances. note: some substances are replaced by Pubchem ID in the figure.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure1.png","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/7d503f4f62e171eaaa6e2414.png"},{"id":57438593,"identity":"3496a014-4de9-4526-a703-6ca79db5cf35","added_by":"auto","created_at":"2024-05-30 17:10:11","extension":"png","order_by":2,"title":"Figure 2","display":"","copyAsset":false,"role":"figure","size":2301868,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003e2D interaction diagram of (a) Lupulone, (b) trans-Isoadhumulone, (c) Caffeine, (d) cis-isohumulone, (e) Adlupulone, The hydrogen bonding interactions in the figure are Conventional Hydrogen Bonds and van der Waals and carbon Hydrogen Bond; Hydrophobic interactions are Pi-Pi Stacked, Alkyl, Pi-Alkyl; Electrostatic interactions are Pi-Anion.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure2.png","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/ccd821ad4f9f7f05c36c944e.png"},{"id":57438400,"identity":"254a8076-dd48-4db6-8b2f-d9d3bf8481d5","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"png","order_by":3,"title":"Figure 3","display":"","copyAsset":false,"role":"figure","size":2033469,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eCorrelation analysis and feature selection.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure3.png","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/8a1cd16e351600c364c3b198.png"},{"id":57438399,"identity":"dd7b818e-8ff1-44f8-95b9-03b508476158","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"png","order_by":4,"title":"Figure 4","display":"","copyAsset":false,"role":"figure","size":1466774,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eR2 of the test set and training set of (A) XGBoost, (B) CatBoost, (C)GBDT, (D) RF, (E) SVR algorithms.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure4.png","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/70abcd2b8a15028a22213a1b.png"},{"id":57438405,"identity":"331cca1b-a807-4ef4-9d67-563d0b06fff5","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"png","order_by":5,"title":"Figure 5","display":"","copyAsset":false,"role":"figure","size":1997605,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003eSHAP analysis of the features influence. (A) The importance by the (a) CatBoost, (b) GBDT, (c) XGBoost, (d) RF model. (B) SHAP value of each feature and (C) the SHAP explanation of all instance by the CatBoost mode.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure5.png","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/0caec056cabbdbc71490a51f.png"},{"id":57438406,"identity":"cb9d50c1-dcef-464e-b49c-2f0e10bbde2b","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"png","order_by":6,"title":"Figure 6","display":"","copyAsset":false,"role":"figure","size":962899,"visible":true,"origin":"","legend":"\u003cp\u003e\u003cstrong\u003ebitterness threshold prediction result of 223 compounds. (A) distribution plot and (B) scatter plot of bitterness threshold.\u003c/strong\u003e\u003c/p\u003e","description":"","filename":"Figure6.png","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/a231f92fba34c98f4ad841c0.png"},{"id":57529531,"identity":"7807d981-e2ed-4a8f-9880-8bd29c2fc12d","added_by":"auto","created_at":"2024-06-01 04:31:45","extension":"pdf","order_by":0,"title":"","display":"","copyAsset":false,"role":"manuscript-pdf","size":12034587,"visible":true,"origin":"","legend":"","description":"","filename":"manuscript.pdf","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/e0aed7cf-ae51-4cb8-9eb3-f490342b3b8e.pdf"},{"id":57438396,"identity":"c4b9f141-6ee4-439f-b147-1ce08503f685","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"xlsx","order_by":1,"title":"","display":"","copyAsset":false,"role":"supplement","size":38120,"visible":true,"origin":"","legend":"","description":"","filename":"TableS1.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/77960321f93597dcfea9d5b8.xlsx"},{"id":57438401,"identity":"3d873c72-56d4-4d8e-aaa9-7bb4a72b3657","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"xlsx","order_by":2,"title":"","display":"","copyAsset":false,"role":"supplement","size":94021,"visible":true,"origin":"","legend":"","description":"","filename":"TableS2.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/1d02de9a0c477167e025a380.xlsx"},{"id":57438594,"identity":"3ce759ed-c4d8-419a-8cd3-dc2fcdf896e6","added_by":"auto","created_at":"2024-05-30 17:10:11","extension":"xlsx","order_by":3,"title":"","display":"","copyAsset":false,"role":"supplement","size":23908,"visible":true,"origin":"","legend":"","description":"","filename":"TableS3.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/9fec56a8c90789a393c5d3e3.xlsx"},{"id":57438403,"identity":"7323de5a-67bc-4a20-a7e3-f8f8b253ac70","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"xlsx","order_by":4,"title":"","display":"","copyAsset":false,"role":"supplement","size":18711,"visible":true,"origin":"","legend":"","description":"","filename":"TableS4.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/aeac4ef08382e89e4b3d3b48.xlsx"},{"id":57438407,"identity":"448b0236-f065-444a-950a-fa5c56d81640","added_by":"auto","created_at":"2024-05-30 17:02:11","extension":"xlsx","order_by":5,"title":"","display":"","copyAsset":false,"role":"supplement","size":699263,"visible":true,"origin":"","legend":"","description":"","filename":"TableS5.xlsx","url":"https://assets-eu.researchsquare.com/files/rs-4439031/v1/71da421c823b19595dbed11d.xlsx"}],"financialInterests":"No competing interests reported.","formattedTitle":"Machine-learning-based bitter taste threshold prediction model for bitter substances: fusing molecular docking binding energy with molecular descriptor features","fulltext":[{"header":"INTRODUCTION","content":"\u003cp\u003eIn the study of taste perception, bitter taste has always been regarded as a fundamental taste[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Many foods and drinks with bitter taste are widely accepted and consumed, and they offer positive health effects. For instance, three of the most popular beverages \u0026ndash; tea, coffee, and chocolate \u0026ndash; are known for their distinctive bitter flavors and aromas[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Additionally, the overall flavor of beer, wine, cider, and citrus juice is inextricably linked to bitterness. However, excessive bitterness can be unpleasant. The amount of bitterness directly impacts target demographics and customer acceptance. Therefore, identifying bitter substances in foods and managing them as a way to process foods with appropriate bitter flavors can have a positive economic impact.\u003c/p\u003e \u003cp\u003eCurrently, sensory assessment by humans is the method most commonly used method to describe the bitter flavor. The threshold test and scale approach are generally utilized to measure bitterness intensity. The lowest sample concentration, which each evaluator perceives as bitter after two or three trials, is considered the bitterness threshold[\u003cspan citationid=\"CR2\" class=\"CitationRef\"\u003e2\u003c/span\u003e]. The Dose-over-threshold (DoT) factor, known as the bitterness active value[\u003cspan citationid=\"CR3\" class=\"CitationRef\"\u003e3\u003c/span\u003e], is frequently employed with the bitterness threshold and computed as the ratio of the individual concentration to the corresponding tastant's recognition threshold[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]. A DoT factor\u0026thinsp;\u0026ge;\u0026thinsp;1 indicates that the bitter tastant contributes to the overall bitterness intensity. The higher the DoT factor, the more bitter the tastant exhibits, and thus the more the tastant contributes to the overall bitterness intensity[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAt present, commonly employed techniques for determining the key bitter substances in food include the following: the DoT factor method that determines whether the substance is the key bitter substance in food by directly obtaining the bitter threshold and DoT factor of the substance[\u003cspan citationid=\"CR4\" class=\"CitationRef\"\u003e4\u003c/span\u003e]; Taste dilution analysis (TDA) method that is utilized to grade food solutions, determine the TD factor between different fractions by continuously diluting the fractions, and finally determine the substances in the most bitter fraction, to determine the key bitter substances in the food[\u003cspan citationid=\"CR5\" class=\"CitationRef\"\u003e5\u003c/span\u003e]; Spectrum descriptive analysis method (SDAM) is also used to remove the dilution process. It directly uses different concentrations of a bitter substance as a strength reference for strength. It evaluates the bitterness strength of various fractions in the food solution to determine its bitter substances[\u003cspan citationid=\"CR6\" class=\"CitationRef\"\u003e6\u003c/span\u003e]. These techniques can be employed alongside a variety of high-precision experimental tools, such as nuclear magnetic resonance (NMR) and liquid chromatography mass spectrometry (LC-MS).\u003c/p\u003e \u003cp\u003eThe advantage of the DoT method over TDA and SDAM is that it eliminates the laborious fractionation stage and combines the bitterness threshold, allowing for quick quantification of known bitter chemicals and a significant reduction in experimental components. However, the bitterness threshold approach has its drawbacks. Firstly, obtaining a single bitter substance can be challenging, often requiring separation and purification through column chromatography or artificial synthesis. Secondly, it is impossible to establish the threshold of toxic bitter chemicals using this method. Thirdly, many sensory evaluators are necessary to evaluate bitter compounds. Lastly, determining the bitter taste threshold involves a cumbersome experimental process and calculation method[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eIn recent years, various mathematical models have been developed to determine the bitterness of a single molecule based on its chemical structure. Bitter taste classification models employing machine learning techniques (e.g., SVM, KNN, GBM, RF, and DNN) include BitterX[\u003cspan citationid=\"CR7\" class=\"CitationRef\"\u003e7\u003c/span\u003e], e-bitter[\u003cspan citationid=\"CR8\" class=\"CitationRef\"\u003e8\u003c/span\u003e], Bittersweet[\u003cspan citationid=\"CR9\" class=\"CitationRef\"\u003e9\u003c/span\u003e], BitterPredict[\u003cspan citationid=\"CR10\" class=\"CitationRef\"\u003e10\u003c/span\u003e], BitterSweetForest[\u003cspan citationid=\"CR11\" class=\"CitationRef\"\u003e11\u003c/span\u003e], and Premexotac[\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e]. Bittermatch[\u003cspan citationid=\"CR13\" class=\"CitationRef\"\u003e13\u003c/span\u003e] is employed to Predict the binding of bitter substances to bitter receptors based on the GBDT algorithm and chemical descriptors. On the other hand, models utilizing different algorithms, such as the iBitter-SCM[\u003cspan citationid=\"CR14\" class=\"CitationRef\"\u003e14\u003c/span\u003e] and the BERT4Bitter[\u003cspan citationid=\"CR15\" class=\"CitationRef\"\u003e15\u003c/span\u003e] bitter peptide prediction models, have been employed. This is based on a bidirectional encoder representation from transformers (BERT) based model and the scoring card method (SCM) model, which is a unique algorithm that is different from machine learning, and it provides a different method of classifying bitter peptides by determining the amino acid sequence of the peptide to predict the bitter peptide. The appellate models have made significant contributions to identifying bitter substances and investigating bitter taste mechanisms; however, they cannot predict the intensity of bitter substances or the bitter taste threshold. It is important to note that identifying bitter compounds is only the first step in identifying key bitter substances in food. In further investigation, it is necessary to screen all bitter compounds and identify those that have the greatest impact on food bitterness. Therefore, it is crucial to establish a prediction model for the bitterness threshold. The BitterIntense model classifies bitter substances into two categories according to their thresholds: below 0.1 mM as very bitter and above 0.1 mM as mild bitter, and then it performs bitter threshold classification modeling[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Although this approach roughly predict the bitter threshold of the substance, it could not provide exact bitter threshold data for bitter substances.\u003c/p\u003e \u003cp\u003eMolecular docking technology is extensively employed for investigating ligandprotein receptor intera-ctions and for virtual screening purposes[\u003cspan citationid=\"CR17\" class=\"CitationRef\"\u003e17\u003c/span\u003e]. Presently, 19 activated amino acid residues within TAS2R14 have been rigorously examined in experimental studies[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Based on this empirical evidence, it can be inferred that substances which dock with these specific amino acid residues can be effectively screened using molecular docking technology. Additionally, the binding energy of the bitter substance bound to TAS2R14 can be determined through the application of the LibDock algorithm and LibDockScore. Consequently, the bitter taste threshold can be derived from the binding energy associated with this particular conformation, when combined with machine learning techniques. The utilization of computer simulations holds the potential to substantially reduce the number of required experiments, equipment costs, and labor involved in these investigations.\u003c/p\u003e \u003cp\u003eBased on the aforementioned information, this study involved machine learning modeling of 96 bitter compounds with known thresholds, using binding energy in molecular docking and chemical descriptors as features. Various algorithms were employed to develop the models, including GBDT, XGBoost, CatBoost, and RF algorithms. The main objectives of this study are as follows: Firstly, to establish an accurate model for predicting the bitterness threshold of specific bitter compounds. Secondly, to explain the binding mechanism between bitter substances and bitter receptor TAS2R14 using molecular docking technology. Thirdly, to identify the features related to the bitterness threshold and select the most accurate features for more modeling methods. Lastly, To predict the bitterness threshold of 223 bitter compounds with unidentified bitter thresholds.\u003c/p\u003e"},{"header":"MATERIALS AND METHODS","content":"\u003cdiv id=\"Sec3\" class=\"Section2\"\u003e \u003ch2\u003e2.1 Data collection\u003c/h2\u003e \u003cp\u003e262 substances were collected from the BitterDB website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://bitterdb.agri.huji.ac.il/dbbitter.php\u003c/span\u003e\u003cspan address=\"https://bitterdb.agri.huji.ac.il/dbbitter.php\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) and some literature sources. The 3D structures of bitter substances were obtained according to the Pubchem ID from the Pubchem website (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://pubchem.ncbi.nlm.nih.gov/\u003c/span\u003e\u003cspan address=\"https://pubchem.ncbi.nlm.nih.gov/\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e), and the structures were downloaded in sdf format. The thresholds are transformed according to the following equation\u003c/p\u003e \u003cp\u003ePT=-lg(T\u0026times;10\u0026thinsp;\u0026minus;\u0026thinsp;6).\u003c/p\u003e \u003cp\u003ewhere T represents the bitterness threshold in mg/kg and PT is the bitterness threshold after processing, as shown in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. Due to the absence of X-ray crystallographic structure for TAS2R14, several models representing the 3D structure of the 25 human TAS2Rs have been developed, leveraging the high homology among GPCR receptors (Paredes Ramos \u0026amp; M L\u0026oacute;pez Vilari\u0026ntilde;o, 2023)[\u003cspan citationid=\"CR19\" class=\"CitationRef\"\u003e19\u003c/span\u003e]. The protein model of TAS2R14 was obtained from the AlphaFoldDB website, numbered Q9NYV8[\u003cspan citationid=\"CR20\" class=\"CitationRef\"\u003e20\u003c/span\u003e, \u003cspan citationid=\"CR21\" class=\"CitationRef\"\u003e21\u003c/span\u003e]. After performing molecular docking, the small molecule conformation with the highest LibDockScore was selected from multiple results of each molecule and exported to sdf format for subsequent molecular descriptor calculation. Finally, the 96 substances were successfully docked to obtain the binding energy.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec4\" class=\"Section2\"\u003e \u003ch2\u003e2.2 Molecular docking of the bitter substances with TAS2R14\u003c/h2\u003e \u003cp\u003eThe water molecule present in the TAS2R14 structure was removed and replaced with a hydrogen atom. Discovery Studio software (DS) of the Receptor-Ligand Interactions module was used to define seven active sites, and the first one was selected. This site corresponds to the binding site proposed in current research for TAS2R14[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. The Prepare Ligands tools of the Small Molecules module were used to optimize the energetic conformations. DS generated all possible optimal conformations of the small molecules. In the LibDock protocol of the Discovery Studio software, the first site was located at coordinates: x\u0026thinsp;=\u0026thinsp;12.209, y\u0026thinsp;=\u0026thinsp;3.273, z = -7.255, with a radius of 17.4 \u0026Aring;. For each optimal conformation set, up to 10 docking results were generated. The molecular docking results were evaluated based on the LibDockScore between the peptides and T2R14. After performing molecular docking, the small molecule conformation with the highest LibDockScore was selected from multiple results of each molecule and exported to sdf format for subsequent molecular descriptor calculation. Then, binding energy of the ligand\u0026ndash;protein complex was calculated by using \u0026ldquo;Calculate Binding Energy\u0026rdquo; protocol which uses CHARMm implicit solvent model. Discovery Studio software generated a 2D results map of the molecular docking process. Finally, the 96 substances were successfully docked to obtain the binding energy.\u003c/p\u003e \u003cp\u003e \u003cb\u003e2.3 Molecular descriptor calculation and feature processing.\u003c/b\u003e \u003c/p\u003e \u003cp\u003eThe 208 descriptors of structural features of the 96 bitter compounds were computed using the MolecularDescriptorCalculator module of the RDKit package. Due to a high number of zero values for some descriptors, they were removed as unsuitable features. The final descriptors used are displayed in Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e. The Scikit-learn library was used for machine learning modeling. The Python machine learning package Scikit-learn offers a variety of commonly used machine learning tools and techniques, such as model selection, dimensionality reduction, clustering, classification, regression, and preprocessing modules. In this study, four algorithms were employed, namely GBDT, XGBoost, CatBoost, and RF. The 96 data sets were randomly divided into a ratio of 8:2 for the training and test sets. The training set and testing set were then normalized using the formula Z=(X-u)/s, where X represents the input property descriptor data, u is the mean of the training set and s is the standard deviation of the training set. This normalization process ensures that the variance of the data in each dimension is 1 and the mean is 0, preventing features with excessively large dimensions from influencing the prediction outcomes. Since there are a large number of remaining features, feature selection is performed on the training set using the embedding method in sklearn in order to prevent dimensionality loss. A 5-fold cross-validation was applied exclusively on the training data to identify the optimal hyperparameters for each method from empirical candidates based on the cross-validated R2, which were used to estimate the performance of the different models.\u003c/p\u003e \u003cdiv id=\"Sec5\" class=\"Section3\"\u003e \u003ch2\u003e2.4 Regression Modeling and Model Evaluation\u003c/h2\u003e \u003cp\u003eR-Square (R2), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE) were utilized to evaluate the model. The statistical metric R2 is used to assess how well the regression model fits the data. The average of the absolute differences between the model predictions and the actual observations is known as the Mean Absolute Error (MAE). It reflects the magnitude of the true error. MSE measures the mean squared difference between the model predictions and the actual observations. RMSE is the square root of the mean squared difference between the model predictions and the actual observations. It measures the average magnitude of the prediction errors, taking into account the squared differences between the predicted and actual values. The formulas of these four values are expressed as follows:\u003cdiv id=\"Equ1\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ1\" name=\"EquationSource\"\u003e\n$${\\text{R}}^{2}=1-\\frac{{\\text{S}\\text{S}}_{\\text{r}\\text{e}\\text{s}}}{{\\text{s}\\text{s}}_{\\text{t}\\text{o}\\text{t}}}=1-\\frac{{\\sum }_{\\text{i}}{({\\text{y}}_{\\text{i}}-{\\text{f}}_{\\text{i}})}^{2}}{{\\sum }_{\\text{i}}{({\\text{y}}_{\\text{i}}-\\stackrel{-}{\\text{y}})}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e1\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ2\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ2\" name=\"EquationSource\"\u003e\n$$\\text{M}\\text{A}\\text{E}=\\frac{1}{\\text{n}}\\sum _{\\text{i}=1}^{\\text{n}}\\left|{\\left({\\text{y}}_{\\text{i}}-{\\text{f}}_{\\text{i}}\\right)}^{2}\\right|$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e2\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ3\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ3\" name=\"EquationSource\"\u003e\n$$\\text{M}\\text{A}\\text{S}\\text{E}=\\frac{1}{\\text{n}}\\sum _{\\text{i}=1}^{\\text{n}}{\\left({\\text{y}}_{\\text{i}}-{\\text{f}}_{\\text{i}}\\right)}^{2}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e3\u003c/div\u003e\u003c/div\u003e\u003cdiv id=\"Equ4\" class=\"Equation\"\u003e\u003cdiv format=\"TEX\" class=\"mathdisplay\" id=\"FileID_Equ4\" name=\"EquationSource\"\u003e\n$$\\text{R}\\text{M}\\text{S}\\text{E}=\\sqrt{\\frac{1}{\\text{n}}\\sum _{\\text{i}=1}^{\\text{n}}{\\left({\\text{y}}_{\\text{i}}-{\\text{f}}_{\\text{i}}\\right)}^{2}}$$\u003c/div\u003e\u003cdiv class=\"EquationNumber\"\u003e4\u003c/div\u003e\u003c/div\u003e\u003c/p\u003e \u003cp\u003ewhere yi, fi, and y̅ represent the true value, predicted value, and average value, respectively. SSres is the sum of squares of residuals, known as the residual sum of squares. SStot is the total sum of squares. n is the number of samples.\u003c/p\u003e \u003cp\u003eA model interpretation package called Shapley additive explanation (SHAP) (\u003cspan class=\"ExternalRef\"\u003e\u003cspan class=\"RefSource\"\u003ehttps://github.com/slundberg/shap\u003c/span\u003e\u003cspan address=\"https://github.com/slundberg/shap\" targettype=\"URL\" class=\"RefTarget\"\u003e\u003c/span\u003e\u003c/span\u003e) was created in Python and put into action by Python scripts. It was employed to assess the feature importance of the four models (GBDT, RF, CatBoost, and XGBoost).\u003c/p\u003e \u003c/div\u003e \u003c/div\u003e \u003cdiv id=\"Sec6\" class=\"Section2\"\u003e \u003ch2\u003e2.5 Prediction of unknown bitterness threshold\u003c/h2\u003e \u003cp\u003eObtained 223 substances from the BitterDB website that are only identified to have bitterness but with undetermined bitterness thresholds. The binding energies for 223 of these substances were then determined by molecular docking, following the method introduced in section \u003cspan refid=\"Sec4\" class=\"InternalRef\"\u003e2.2\u003c/span\u003e. The molecular descriptor and feature scaling were then determined in accordance with the method introduced in section 2.3. The prediction of PT was conducted using the CatBoost model and transformed into a bitter threshold.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec7\" class=\"Section2\"\u003e \u003ch2\u003e2.6 Statistical Analysis\u003c/h2\u003e \u003cp\u003eAll analyses were implemented on the Anaconda3 platform, including scikit-learn joblib, pandas, numpy, and matplotlib for data processing and analysis.\u003c/p\u003e \u003c/div\u003e"},{"header":"RESULTS AND DISCUSSION","content":"\u003cdiv id=\"Sec9\" class=\"Section2\"\u003e \u003ch2\u003e3.1 grouping of 262 compounds\u003c/h2\u003e \u003cp\u003eA dataset of 262 samples was obtained from the BitterDB website and some literature, which compiles information on the bitterness of various foods from several resources, including Gouda cheese, fennel, beer, whole wheat bread crumbs, oats, red wine, asparagus, carrots, and roasted cocoa nibs. These substances mainly include polypeptides, alkaloids, phenols, terpenoids, and saponins (Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e). In recent years, many studies have been conducted on Quantitative Structure-Activity Relationship modeling using bitter peptides based on molecular descriptors and the efficacy of some molecular peptides. Bitter peptides are mainly produced by proteins during the food fermentation process and offer significant benefits for cardiovascular, neurological, immune, and nutritional systems[\u003cspan citationid=\"CR22\" class=\"CitationRef\"\u003e22\u003c/span\u003e]. Phenolic compounds, which are organic compounds found abundantly in plants, have gained attention as an emerging field of nutrition in recent decades. A growing body of evidence suggests that the intake of polyphenols may play a vital role in regulating metabolism, body weight, chronic diseases, and cell proliferation[\u003cspan citationid=\"CR23\" class=\"CitationRef\"\u003e23\u003c/span\u003e]. Phenolic compounds contribute to the bitterness and astringency of various beverages and fruits, such as oranges, coffee, and tea[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. Alkaloids, a class of nitrogenous organic compounds, are present in natural constituents of vegetables and drinks, some food processing additives, and pollutants[\u003cspan citationid=\"CR24\" class=\"CitationRef\"\u003e24\u003c/span\u003e]. Most alkaloids exhibit a bitter taste and possess a potent biological activity, some of which may be toxic to humans. Terpenoids (terpenes and their oxygenated derivatives) are important secondary metabolites found in medicinal plants and some well-known foods and beverages, contributing to their pronounced bitterness. Saponins, on the other hand, are secondary plant metabolites composed of steroids, triterpenes, or alkaloids with one or more sugar chains[\u003cspan citationid=\"CR25\" class=\"CitationRef\"\u003e25\u003c/span\u003e]. They are bitter and astringent substances found in certain foods. Additionally, fatty acids, fatty alcohols and other categories are also important bitter substances.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec10\" class=\"Section2\"\u003e \u003ch2\u003e3.2 Binding mode of bitter protein TAS2R14 with bitter substances\u003c/h2\u003e \u003cp\u003eThe binding residues of the TAS2R14 protein that have been experimentally identified include Trp66, Leu85, Thr86, Asn87, Trp89, Thr90, Asn93, His94, Thr182, Ser183, Phe186, Ile187, Tyr240, Ala241, Phe243, Phe247, Ile263, Gln266, and Gly269 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e]. Molecular docking showed similar results to the experimental outcome, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e. The binding pocket of the TAS2R14 receptor is quite spacious and can tolerate agonist structures that largely exceed the size of unmodified agonists [\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. No residues were found in the TAS2R14 binding pocket where mutations resulted in a complete loss of function of all agonists, whereas studies of other receptors have identified such restricted targets. The Asn24 and Ile27 residues of TAS2R1[\u003cspan citationid=\"CR27\" class=\"CitationRef\"\u003e27\u003c/span\u003e]. Gly281, and Ser285 residues of TAS2R4[\u003cspan citationid=\"CR28\" class=\"CitationRef\"\u003e28\u003c/span\u003e], Trp261 residues of TAS2R16[\u003cspan citationid=\"CR29\" class=\"CitationRef\"\u003e29\u003c/span\u003e] have been found to be important binding sites for activating proteins. Moreover, TAS2R14 provides a large number of agonist-selective contacts that may outweigh all other promiscuous TAS2Rs[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eAfter screening substances that respond to the appeal TAS2R14 binding site, the docking site map is shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Table \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e. And the binding energies of the 96 successfully docked substances are shown in Table \u003cspan refid=\"MOESM1\" class=\"InternalRef\"\u003eS1\u003c/span\u003e. Regarding substance categorization, almost all classes bound a portion of their substances to TAS2R14, further demonstrating the broad coordination ability of TAS2R14 [\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e, \u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Nearly all of the 24 fatty acids, fatty alcohols, and glycerolipids were successfully docked, and all three belong to long-chain lipids with a very long carbon chain that allows easier access to the centers of the TM1 and TM2 chains and less site blockage. These three substances demonstrate high LibDockScore (\u0026gt;\u0026thinsp;100). LibDockScore is a scoring algorithm that rates the degree of ligand-protein stabilization[\u003cspan citationid=\"CR31\" class=\"CitationRef\"\u003e31\u003c/span\u003e]. Therefore, it indicates that long-chain analogs can form more stable complexes with TAS2R14 proteins, and some of them have low thresholds (\u0026lt;\u0026thinsp;1mg/Kg). In addition, most of Cinnamic acids and derivatives and flavonoids bind to TAS2R14.\u003c/p\u003e \u003cp\u003eThe mode of interaction between the TAS2R14 receptor and the ligand has received extensive attention. The results of the binding site reflect two patterns, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig1\" class=\"InternalRef\"\u003e1\u003c/span\u003e and Table \u003cspan refid=\"MOESM3\" class=\"InternalRef\"\u003eS3\u003c/span\u003e. while alkaloids, amino acids and their derivatives, and benzenes do not interact with Lys8 and Trp89 and mainly interact hydrophobically with Ile187, Cinnamic acids and their derivatives, Prenol lipids, and flavonoids interact hydrophobically mainly with Leu85 and Trp89, and they partially bind to Ile187. Interestingly, the first group usually had lower LibDockScore and higher thresholds, suggesting that unstable binding affects the bitterness thresholds. It can be found that substances with binding energies below \u0026minus;\u0026thinsp;100 have low thresholds (\u0026lt;\u0026thinsp;50mg/Kg), while some substances with high thresholds, such as amino acids, have high binding energies. The magnitude of the binding energy indicates the strength of the interaction, and substances with low binding energies indicate that they interact more strongly with TAS2R14, consequently exhibiting stronger bitterness intensity and lower bitterness threshold[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e]. However, there are some exceptions. Long-chain substances have the lowest threshold, yet their binding energies vary, indicating that the threshold is not solely influenced by binding strength. Long-chain substances might more readily bind with TAS2R14 to activate it, but the binding interactions are not necessarily strong.\u003c/p\u003e \u003cp\u003eFigure\u0026nbsp;\u003cspan refid=\"Fig5\" class=\"InternalRef\"\u003e2\u003c/span\u003e shows the binding of the five known TAS2R14 activators. In agreement with some findings, Leu85 and Trp89 are amongst the most important binding residues[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Four substances, Lupulone, Adlupulone, cis-isohumulone, and trans-Isoadhumulone, were found to interact hydrophobically with Leu85 and Trp89, with the main interactions being alkyl and pi-alkyl. Residue Trp89 is deep in the binding pocket[\u003cspan citationid=\"CR34\" class=\"CitationRef\"\u003e34\u003c/span\u003e]; thus, the ability to bind to residue Trp89 may be more stable. Residue Thr86 also undergoes hydrogen bonding interactions with trans-Isoadhumulone cis-isohumulone, with Thr86 and Asn93 being the most common hydrogen bonding interactions in the study[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. In this study, both residues also show binding to a great extent, with residue Thr86 having more binding ability than Asn93, especially since the flavonoid glycosides almost all interact with THR86 by hydrogen bonding. In addition, some studies have concluded that flufenamic acid, mainly pi-pi, interacts with Trp89, Phe186, and Phe247[\u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e]. The mechanism of action of the benzoic acid analog TAS2R14 includes p-p interactions between the ligand and Phe186, Phe243, and His94 as well as hydrogen bonding with Asn93 and Thr86[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e]. These binding residues were also found in our study. With further research, we will explore more binding sites and activators, which will deepen our study of the binding mode of TAS2R14 and the relationship between binding energy and threshold.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec11\" class=\"Section2\"\u003e \u003ch2\u003e3.3 Feature selection\u003c/h2\u003e \u003cp\u003eThe interaction between bitter compounds and bitter receptors is the key factor underlying bitterness, and the strength of this interaction determines the intensity of bitterness[\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR36\" class=\"CitationRef\"\u003e36\u003c/span\u003e]. Binding energy as a quantitative measure for evaluating the strength of receptor-ligand binding is closely related to the bitter taste threshold. Since only one receptor's binding energy can be used as a feature for modeling, using receptors that are able to bind a wide range of substances will improve our model's ability to predict substance types.\u003c/p\u003e \u003cp\u003eA total of 25 bitter receptors have been identified to bind various types of bitter substances. However, TAS2R10, TAS2R14, and TAS2R46 are broadly tuned receptors known to respond to a variety of natural and artificial bitter compounds, while TAS2R3, TAS2R5, TAS2R8, TAS2R13, TAS2R41, TAS2R49, and TAS2R50, are activated by very few bitter compounds, and TAS2R42, TAS2R45, TAS2R48, and TAS2R60 are currently unidentified substances[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. TAS2R10 and TAS2R14 showed less preference among agonists, while TAS2R14 could bind more chemical scaffolds and more agonists[\u003cspan citationid=\"CR26\" class=\"CitationRef\"\u003e26\u003c/span\u003e, \u003cspan citationid=\"CR30\" class=\"CitationRef\"\u003e30\u003c/span\u003e]. Most receptors embody a high degree of selectivity for structure; for example, TAS2R16 and TAS2R38 show a strong tendency to recognize β-dglucopyranosides and N\u0026ndash;C\u0026thinsp;=\u0026thinsp;S group in molecules[\u003cspan citationid=\"CR1\" class=\"CitationRef\"\u003e1\u003c/span\u003e]. In addition, TAS2R7 is suited for receptors for metal cations such as zinc, magnesium, calcium, aluminum, copper and manganese[\u003cspan citationid=\"CR37\" class=\"CitationRef\"\u003e37\u003c/span\u003e]. TAS2R46, on the other hand, tends to detect sesquiterpene lactones and related compounds[\u003cspan citationid=\"CR18\" class=\"CitationRef\"\u003e18\u003c/span\u003e].\u003c/p\u003e \u003cp\u003eMolecular descriptors are numerical features used to represent and describe the structure of a molecule, and the MolecularDescriptorCalculato module calculates 208 molecular descriptors, with 93 descriptors remaining after removing 0 and most of the descriptors with 0 (and adding binding energies) as shown in Table \u003cspan refid=\"MOESM2\" class=\"InternalRef\"\u003eS2\u003c/span\u003e. Feature selection was performed on the training set by embedding method decision tree. The top 15 features selected according to gain are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e3\u003c/span\u003e. Due to the stochastic nature of decision tree feature selection, to prevent some features from being missed, some of the features were selected from the remaining 78 features that were considered to be important at the time the other bitterness models were built. Weichen Bo et al[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. used an artificial neural network algorithm to build a bitter-non-bitter classification model, and based on the important properties they derived to classify bitter substances, we chose heavy atom molecular weight (HeavyAtomMolWt), the hydrophobicity feature (BCUT2D_LOGPHI), backbone atom electronegativity (EState_VSA1 and EState_VSA4), molecular shape (Kappa2). MinEStateIndex was added based on the conserved site modeling[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e], in addition, BitterIntense, as a model for classification thresholding, also provides two important features, heavyatom and molar refractivity. HeavyAtom is replaced by HeavyAtomMolWt, molar refractivity is replaced by SMR_VSA5. Finally, we added molecular descriptors, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e3\u003c/span\u003e, and them a correlation analysis was performed, removing substances with too much correlation to reduce multicollinearity, and the final results are shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig6\" class=\"InternalRef\"\u003e3\u003c/span\u003e. The predictive model with higher R-squared bitter taste thresholds was developed through multiple selections, and the feasibility of the binding energy as a feature was demonstrated by the fact that the binding energy was ranked 12 and successfully retained from multiple selections. Since the randomness of the embedding method does not fully represent the importance of the features, they will be further evaluated by the SHAP algorithm in the modeling below.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec12\" class=\"Section2\"\u003e \u003ch2\u003e3.4 Machine learning model analysis\u003c/h2\u003e \u003cp\u003eFour regression models, namely RF, XGBoost, CatBoost and GBDT, were employed for the prediction of bitterness thresholds, and the prediction results are shown in Table \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e. Figure\u0026nbsp;\u003cspan refid=\"Fig2\" class=\"InternalRef\"\u003e4\u003c/span\u003e displays the optimal model results for each algorithm, which demonstrates a good correlation in both the training set and the test set. The R2 values of PT in the training set were 0.916, 0.893, 0.948 and 0.893, respectively. Similarly, the R2 values of PT in the test set were 0.906, 0.889, 0.936, and 0.877, respectively. The slight difference between the R2 of the training set and the test set indicates a low risk of overfitting. The findings demonstrate that all four machine learning models could successfully estimate the PT of a substance, with the CatBoost algorithm showing the best results in both the training and test sets. To further evaluate the models, three metrics, namely MAE, MSE, and RMSE, were utilized. These three indicators assess the degree of error between the true and predicted values, and the closer the three indicators are to 0, the better the model works. Table\u0026nbsp;\u003cspan refid=\"Tab1\" class=\"InternalRef\"\u003e1\u003c/span\u003e demonstrates that the MAE, MSE, and RMSE exhibit opposite tendencies to R2 for these four models. Among these models, the three indicators of CatBoost were the smallest (i.e., 0.936, 0.317, 0.174, 0.417, respectively). Hence, the CatBoost model demonstrated the highest effectiveness and can be used to forecast the bitterness threshold in future studies. In addition, models based on RF and GBDT algorithms also displayed good results.\u003c/p\u003e \u003cp\u003e \u003cdiv class=\"gridtable\"\u003e\u003ctable float=\"Yes\" id=\"Tab1\" border=\"1\"\u003e \u003ccaption language=\"En\"\u003e \u003cdiv class=\"CaptionNumber\"\u003eTable 1\u003c/div\u003e \u003cdiv class=\"CaptionContent\"\u003e \u003cp\u003eModel Evaluation\u003c/p\u003e \u003c/div\u003e \u003c/caption\u003e \u003ccolgroup cols=\"5\"\u003e \u003cdiv align=\"left\" class=\"colspec\" colname=\"c1\" colnum=\"1\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c2\" colnum=\"2\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c3\" colnum=\"3\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c4\" colnum=\"4\"\u003e\u003c/div\u003e \u003cdiv align=\"char\" char=\".\" class=\"colspec\" colname=\"c5\" colnum=\"5\"\u003e\u003c/div\u003e \u003cthead\u003e \u003ctr\u003e \u003cth align=\"left\" colname=\"c1\"\u003e \u003cp\u003eAlgorithm\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c2\"\u003e \u003cp\u003eR\u003csup\u003e2\u003c/sup\u003e\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c3\"\u003e \u003cp\u003eMAE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c4\"\u003e \u003cp\u003eMSE\u003c/p\u003e \u003c/th\u003e \u003cth align=\"left\" colname=\"c5\"\u003e \u003cp\u003eRMSE\u003c/p\u003e \u003c/th\u003e \u003c/tr\u003e \u003c/thead\u003e \u003ctbody\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eGBDT\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.877\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.477\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.333\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.577\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eXGBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.889\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.423\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.299\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.546\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eRF\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.906\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.370\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.253\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.503\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003ctr\u003e \u003ctd align=\"left\" colname=\"c1\"\u003e \u003cp\u003eCatBoost\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c2\"\u003e \u003cp\u003e0.936\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c3\"\u003e \u003cp\u003e0.317\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c4\"\u003e \u003cp\u003e0.174\u003c/p\u003e \u003c/td\u003e \u003ctd align=\"char\" char=\".\" colname=\"c5\"\u003e \u003cp\u003e0.417\u003c/p\u003e \u003c/td\u003e \u003c/tr\u003e \u003c/tbody\u003e \u003c/colgroup\u003e \u003c/table\u003e\u003c/div\u003e \u003c/p\u003e \u003cp\u003eTable \u003cspan refid=\"MOESM4\" class=\"InternalRef\"\u003eS4\u003c/span\u003e lists the training and true values for all test sets, and due to the high R-squared of the model itself, most of the substance predictions are almost identical to the true values. However, there are some flaws in the model; the four categories of Chalcones, Flavanols, and Organooxygen compounds appeared less in the training set, the Flavanols category was not trained, and the remaining 2 types had 1\u0026ndash;3 substances trained in the training set. As can be observed, the prediction value of Flavanols is too high. At the same time, the other two substances are predicted accurately, indicating that the model is adequately trained for the training set, and it can effectively predict the various structures of the substances appearing in the training set. However, it may have poorer prediction ability of unknown types. This needs to be improved by collecting more data. Secondly, Amino acids, Peptides, Benzoic acids, and derivatives also showed inaccurate predictions, which is a category of substances that the model needs to focus on subsequently. In contrast, the rest of Glycerolipids, Flavonoid glycosides, and Fatty acids were more accurate.\u003c/p\u003e \u003cp\u003eAmong the four algorithms used, GBDT, RF, CatBoost, and XGBoost are all integrated algorithms[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. Ensemble algorithms combine multiple simple models to create a more powerful model, thereby improving prediction performance. An ensemble technique built on a gradient boosting decision tree is called XGBoost or CatBoost, while GBDT and RF are ensemble approaches based on decision trees[\u003cspan citationid=\"CR39\" class=\"CitationRef\"\u003e39\u003c/span\u003e]. The performance of the model heavily relies on the selection of features. Since our feature selection method uses decision trees, our four algorithms are based on decision trees. RF and GBDT are two distinct Decision Tree-based methods. RF constructs numerous randomized decision trees to achieve integration, whereas GBDT iteratively optimizes the model[\u003cspan citationid=\"CR40\" class=\"CitationRef\"\u003e40\u003c/span\u003e]. RF is more suitable for high-dimensional problems. Both CatBoost and XGBoost are GBDT-based algorithms. XGBoost is fast in terms of execution speed and memory efficiency, and it can fit the data better and find the global optimal solution faster, so it is more suitable for large data volumes. CatBoost, on the other hand, employs ordered boosting strategies to prevent prediction bias caused by target leakage [\u003cspan citationid=\"CR41\" class=\"CitationRef\"\u003e41\u003c/span\u003e]. Due to the small sample size, CatBoost has a built-in regularization mechanism that reduces the risk of overfitting and thus performs better when dealing with small samples or noisy data.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec13\" class=\"Section2\"\u003e \u003ch2\u003e3.5 Feature analysis of 4 models\u003c/h2\u003e \u003cp\u003eIn Fig.\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e5\u003c/span\u003eA, the feature importance of the CatBoost, GBDT, XGBoost and RF model is ranked by SHAP and measured based on the absolute value of the average Shapley value. Among these models, BCUT2D_MRLOW, Chi2v, PEOE_VSA7 are the three highest ranked feature. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e5\u003c/span\u003eB presents the summary_polt, which illustrates the effect of each feature on the CatBoost model. In the plot, the PEOE_VSA7, Chi2v and SMR_VSA5 and Binding Energy points red and blue areas at the same time on value\u0026thinsp;\u0026lt;\u0026thinsp;0, Explain that these features have a positive effect on the predicted value of PT (There is some degree of inverse relationship between the eigenvalues and the results predicted by the model). the BCUT2D_MRLOW points are clustered into two regions: the red region with positive SHAP values (\u0026gt;\u0026thinsp;0) and the blue region with negative SHAP values (\u0026lt;\u0026thinsp;0). This indicates that higher BCUT2D_MRLOW values have a positive effect on PT, while lower BCUT2D_MRLOW values have a negative effect. Figure\u0026nbsp;\u003cspan refid=\"Fig3\" class=\"InternalRef\"\u003e5\u003c/span\u003eC shows the force distribution of the CatBoost model (arranged by similarity). Based on the distribution, it can be observed that the predictive pattern of the features has two ways for all the samples of the model, and the main difference is that the BCUT2D_MRLOW feature contributes positively and negatively to the model.\u003c/p\u003e \u003cp\u003eChi2v is a molecular descriptor based on atomic charge distribution and molecular topology used to encode the electrical, chemical reactivity, and structural characteristics of molecules. The SMR_VSA descriptor explains the molar refractive index of a molecule and its effect within the receptor or on the way to the receptor. SMR_VSA5 is defined as the sum of the vi of all atoms i. pi denotes the contribution to the molar refractive index of atom i calculated in the SMR descriptor, computed in the range of 0.440 to 0.485[\u003cspan citationid=\"CR42\" class=\"CitationRef\"\u003e42\u003c/span\u003e]. The BCUT2D descriptor is based on the Burden matrix, which encodes the bond strengths between atoms in a molecule. It changes the diagonal elements of the matrix to include atomic properties and then performs eigenvalue decomposition to obtain the highest and lowest eigenvalues. PEOE_VSA8 (a molecular surface area descriptor). Mollogp, SlogP_VSA6 is related to the partition coefficient (logP) and is a characterization of the hydrophobicity of the molecule. Binding Energy, as mentioned earlier, is a descriptor for assessing the strength of ligand-receptor binding and is related to hydrophobic interactions, charge interactions, hydrogen bonding interactions, etc. These molecular descriptors and binding energies are important features for constructing bitter taste threshold models. Molecular descriptors are used to quantitatively describe the physical and chemical information of molecules[\u003cspan citationid=\"CR33\" class=\"CitationRef\"\u003e33\u003c/span\u003e]. Thus, the physical properties they respond to influence the magnitude of the bitterness threshold.\u003c/p\u003e \u003cp\u003eCurrently, there are several studies on the bittering mechanism, among which the recognized properties are hydrophobicity, which is the main property found in various types of studies, where researchers have suggested that substances with a strong bitter taste are accompanied by a high degree of hydrophobicity [\u003cspan citationid=\"CR32\" class=\"CitationRef\"\u003e32\u003c/span\u003e, \u003cspan citationid=\"CR35\" class=\"CitationRef\"\u003e35\u003c/span\u003e, \u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e]. This relationship can be attributed to the influence of hydrophobic molecules on various factors, such as the distribution of molecules in the cell membrane, the ability to bind to bitter taste receptors, and the interaction force between molecules, thereby affecting the perceived intensity of bitter taste. The molar refractivity is a measure of the extent to which the electron density distribution within a molecule can be distorted[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e]. Bitter taste prediction models Premexotac [\u003cspan citationid=\"CR12\" class=\"CitationRef\"\u003e12\u003c/span\u003e] and a recent artificial neural network prediction model[\u003cspan citationid=\"CR38\" class=\"CitationRef\"\u003e38\u003c/span\u003e] considered it as an important feature for bitterness classification, while BitterIntense[\u003cspan citationid=\"CR16\" class=\"CitationRef\"\u003e16\u003c/span\u003e] model is a prediction of the bitterness threshold and this model gained a high importance score in molar refractivity, SMR_VSA5 is a molar refractive index descriptor, and it was ranked 6th in the importance of the model, so we got the same conclusion that the distribution of the molecule's electron cloud affects the intensity of bitterness. The binding energy contributed to the construction of CatBoost model and RF model, but its importance score in GBDT and XGBoost model is almost 0. The comparison of the features shows that the importance of the binding energy is the 11th feature for constructing the bitterness threshold after hydrophobicity and Molar refractivity. Although the contribution of binding energy is 0 in GBDT and XGBoost models, the R-square of these two models is low (\u0026lt;\u0026thinsp;0.9), which suggests that the inability to find the relationship between the binding energy and the bitter taste threshold may be a reason for the low prediction effect of these two algorithms. Finally, this paper concludes that factors such as the electronic environment of the molecule, hydrophobicity, chemical reactivity, and the strength of binding to the TAS2R14 protein (binding energy) influence the magnitude of the bitterness threshold.\u003c/p\u003e \u003c/div\u003e \u003cdiv id=\"Sec14\" class=\"Section2\"\u003e \u003ch2\u003e3.6 Threshold prediction for bitter substances\u003c/h2\u003e \u003cp\u003eTable \u003cspan refid=\"MOESM5\" class=\"InternalRef\"\u003eS5\u003c/span\u003e presents the predictions for 223 compounds, including their names, structure, binding energy, PT, threshold (mg/kg), Chi2v, PEOE_VSA7 and BCUT2D_MRLOW values. Figure\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e6\u003c/span\u003eA shows the distribution of thresholds for 223 substances, and the result shows that the thresholds for most substances are concentrated in the range of 100\u0026ndash;300 mg/kg. To further investigate these substances, a three-dimensional scatter plot of three features was drawn based on these substances, as shown in Fig.\u0026nbsp;\u003cspan refid=\"Fig4\" class=\"InternalRef\"\u003e6\u003c/span\u003eB. From this plot, it can be observed that sample points with close bitterness thresholds are also closer together in the plot and show a gradual change in colour, indicating that the importance of these three features in regression modelling. Secondly, most of the substances with thresholds lower than 1000 are in the region where Chi2v is less than 2.5 and BCUT2D_MRLOW is greater than \u0026minus;\u0026thinsp;0.5, and the lower the Chi2v value, the higher the BCUT2D_MRLOW value and the lower the threshold. The substances with high thresholds (\u0026gt;\u0026thinsp;1000mg/Kg) are more likely to be affected by the BCUT2D_MRLOW feature because these substances are all at the bottom of the 3D map space; however, they are also modulated by PEOE_VSA7, which slightly decreases the thresholds of these substances as PEOE_VSA7 increases.\u003c/p\u003e \u003c/div\u003e"},{"header":"CONCLUSION","content":"\u003cp\u003eIn this study, we initially analyzed the binding mode of bitter substances with the TAS2R14 receptor protein and identified the amino acid residues Lys85, Trp89, and Thr86 crucial for bitter substances to interact with T2R14. They interact with bitter compounds through hydrophobic and hydrogen bond interactions, respectively. Subsequently, a bitterness threshold prediction model was established based on four algorithms. CatBoost is an ensemble algorithm based on the GBDT algorithm. It\u0026rsquo;s highest r-square model was constructed and used to accurately predict the bitterness threshold of bitterness. Finally, the important factors that impact the prediction of bitterness threshold by the model are found to be molar refractivity, hydrophobicity, the electronic environment of the molecule, reactivity, and the strength of binding to the TAS2R14 protein. The main disadvantage of the current model is the small amount of data, and the determination of the bitterness threshold may vary across different studies of the same substance due to subjective factors. Therefore, collecting more bitter taste threshold data in the future will increase the generalization ability of the model. Secondly, we used the classical LibDock algorithm for molecular docking, while newer scoring functions based on machine learning are gradually being investigated[\u003cspan citationid=\"CR43\" class=\"CitationRef\"\u003e43\u003c/span\u003e], which may be able to be applied in the future for the assessment of the binding status of bitter substances to bitter receptors and bitter taste threshold modeling. However, it was undeniable that the bitterness threshold prediction models proposed in this study hold potential applications and it can eliminate complex sensory experiments to quickly determine the bitterness threshold of a large number of substances through initial screening of food bitter substances. Additionally, these models can serve as alternatives to sensory experiments in determining the bitterness threshold of some toxic substances and offer valuable guidance for the subsequent exploration and masking of key bitter compounds by determining the bitterness threshold.\u003c/p\u003e"},{"header":"Abbreviations","content":"\u003cp\u003eRF, Random forest; XGBoost, Extreme Gradient Boosting; CatBoost, Categorical Boosting; GBDT, Gradient Boosting Decision Tree; DOT, Dose-over-threshold; TDA, Taste Dilution Analysis; SDAM, Spectrum Descriptive Analysis Method; NMR, Nuclear Magnetic Resonance; LC-MS, Liquid Chromatography Mass Spectrometry; BERT, Bidirectional Encoder Representation from Transformers; SCM, Scoring Card Method; DS, Discovery Studio; R2, R-Square; MAE, Mean Absolute Error; MSE, Mean Squared Error; RMSE, Root Mean Squard Error; SHAP, Shapley Additive Pxplanation; SDF, Structure-Data File; PT, threshold after logarithmic processing; kNN, k-Nearest Neighbor; SVM, Support Vector Machine; GBM, Gradient Boosting Machine; DNN, Deep Neural Network.\u003c/p\u003e\n"},{"header":"Declarations","content":"\u003cp\u003e\u003cstrong\u003eACKNOWLEDGMENTS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors would like to express their gratitude to EditSprings (https://www.editsprings.cn ) for the expert linguistic services provided. The authors would like to thank the high-performance computing platform of Guangxi University, Nanning, China.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eFUNDING SOURCES\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThis work was supported by the National Natural Science Foundation of China (32160571), and Guangxi Key Research and Development Program (Guike AB21220068).\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eSUPPORTING INFORMATION DESCRIPTION\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eBitterness threshold and classification of all substances (Table S1) (XLSX)\u003c/p\u003e\n\u003cp\u003eAll characteristics of a substance (Table S2) (XLSX)\u003c/p\u003e\n\u003cp\u003eDocking results for 96 substances (Table S3) (XLSX)\u003c/p\u003e\n\u003cp\u003ePredicted values of bitterness thresholds for four models (Table S4) (XLSX)\u003c/p\u003e\n\u003cp\u003ePredicted bitter threshold for 223 bitter substances (Table S5) (XLSX)\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eAUTHOR CONTRIBUTIONS\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eCan Chen: Investigation, Data curation, Writing - original draft, Writing - review \u0026amp; editing.\u003c/p\u003e\n\u003cp\u003eHaichao Deng: Data Curation, Formal Analysis.\u003c/p\u003e\n\u003cp\u003eHuijie Wei: Investigation, Data curation.\u003c/p\u003e\n\u003cp\u003eYaqing Wang: Conceptualization, Supervision.\u003c/p\u003e\n\u003cp\u003eNing Xia: Conceptualization.\u003c/p\u003e\n\u003cp\u003eJianwen Teng: Formal analysis, Methodology. \u0026nbsp;\u003c/p\u003e\n\u003cp\u003eQisong Zhang:\u0026nbsp;Conceptualization.\u003c/p\u003e\n\u003cp\u003eLi Huang: Conceptualization, Supervision, Resources, Formal analysis, Methodology, writing-review \u0026amp; editing\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eNOTE\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe authors declare no competing financial interest.\u003c/p\u003e\n\u003cp\u003e\u003cstrong\u003eDATA AVAILABILITY\u003c/strong\u003e\u003c/p\u003e\n\u003cp\u003eThe script and the parameter files are available in the GitHub repository: https://github.com/pandaness/Bitter-model.git\u003c/p\u003e"},{"header":"References","content":"\u003col\u003e\n\u003cli\u003eYan J., Tong H. (2023) An overview of bitter compounds in foodstuffs: Classifications, evaluation methods for sensory contribution, separation and identification techniques, and mechanism of bitter taste transduction COMPREHENSIVE REVIEWS IN FOOD SCIENCE AND FOOD SAFETY 22:187-232 https://doi.org/10.1111/1541-4337.13067\u003c/li\u003e\n\u003cli\u003eLi H., Li L.F., Zhang Z.J., Wu C.J., Yu S.J. (2021) Sensory evaluation, chemical structures, and threshold concentrations of bitter-tasting compounds in common foodstuffs derived from plants and maillard reaction: A review CRITICAL REVIEWS IN FOOD SCIENCE AND NUTRITION 1-41 https://doi.org/10.1080/10408398.2021.1973956\u003c/li\u003e\n\u003cli\u003eSeo M.W., Yang D.S., Kays S.J., Lee G.P., Park K.W. (2009) Sesquiterpene Lactones and Bitterness in Korean Leaf Lettuce Cultivars HORTSCIENCE 44:246-249 https://doi.org/10.21273/HORTSCI.44.2.246\u003c/li\u003e\n\u003cli\u003eScharbert S., Hofmann T. (2005) Molecular Definition of Black Tea Taste by Means of Quantitative Studies, Taste Reconstitution, and Omission Experiments JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 53:5377-5384 https://doi.org/10.1021/jf050294d\u003c/li\u003e\n\u003cli\u003eFrank O., Ottinger H., Hofmann T. (2001) Characterization of an intense bitter-tasting 1H,4H-quinolizinium-7-olate by application of the taste dilution analysis, a novel bioassay for the screening and identification of taste-active compounds in foods JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 49:231-238 https://doi.org/10.1021/jf0010073\u003c/li\u003e\n\u003cli\u003eLiu X., Jiang D., Peterson D.G. (2014) Identification of bitter peptides in whey protein hydrolysate JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 62:5719-5725 https://doi.org/10.1021/jf4019728\u003c/li\u003e\n\u003cli\u003eHuang W., Shen Q., Su X., Ji M., Liu X., Chen Y., Lu S., Zhuang H., Zhang J. (2016) BitterX: a tool for understanding bitter taste in humans Scientific Reports 6 https://doi.org/10.1038/srep23450\u003c/li\u003e\n\u003cli\u003eZheng S., Jiang M., Zhao C., Zhu R., Hu Z., Xu Y., Lin F. (2018) e-Bitter: Bitterant Prediction by the Consensus Voting From the Machine-Learning Methods Frontiers in Chemistry 6 https://doi.org/10.3389/fchem.2018.00082\u003c/li\u003e\n\u003cli\u003eTuwani R., Wadhwa S., Bagler G. (2019) BitterSweet: Building machine learning models for predicting the bitter and sweet taste of small molecules Scientific Reports 9 https://doi.org/10.1038/s41598-019-43664-y\u003c/li\u003e\n\u003cli\u003eDagan-Wiener A., Nissim I., Ben Abu N., Borgonovo G., Bassoli A., Niv M.Y. (2017) Bitter or not? BitterPredict, a tool for predicting taste from chemical structure Scientific Reports 7 https://doi.org/10.1038/s41598-017-12359-7\u003c/li\u003e\n\u003cli\u003eBanerjee P., Preissner R. (2018) BitterSweetForest: A Random Forest Based Binary Classifier to Predict Bitterness and Sweetness of Chemical Compounds Frontiers in Chemistry 6 https://doi.org/10.3389/fchem.2018.00093\u003c/li\u003e\n\u003cli\u003eDe Le\u0026oacute;n G., Fr\u0026ouml;hlich E., Fink E., Di Pizio A., Salar-Behzadi S. (2022) Premexotac: Machine learning bitterants predictor for advancing pharmaceutical development INTERNATIONAL JOURNAL OF PHARMACEUTICS 628:122263 https://doi.org/10.1016/j.ijpharm.2022.122263\u003c/li\u003e\n\u003cli\u003eMargulis E., Slavutsky Y., Lang T., Behrens M., Benjamini Y., Niv M.Y. (2022) BitterMatch: recommendation systems for matching molecules with bitter taste receptors Journal of Cheminformatics 14 https://doi.org/10.1186/s13321-022-00612-9\u003c/li\u003e\n\u003cli\u003eCharoenkwan P., Yana J., Schaduangrat N., Nantasenamat C., Hasan M.M., Shoombuatong W. (2020) iBitter-SCM: Identification and characterization of bitter peptides using a scoring card method with propensity scores of dipeptides GENOMICS 112:2813-2822 https://doi.org/10.1016/j.ygeno.2020.03.019\u003c/li\u003e\n\u003cli\u003eCharoenkwan P., Nantasenamat C., Hasan M.M., Manavalan B., Shoombuatong W. (2021) BERT4Bitter: a bidirectional encoder representations from transformers (BERT)-based model for improving the prediction of bitter peptides BIOINFORMATICS 37:2556-2562 https://doi.org/10.1093/bioinformatics/btab133\u003c/li\u003e\n\u003cli\u003eMargulis E., Dagan-Wiener A., Ives R.S., Jaffari S., Siems K., Niv M.Y. (2021) Intense bitterness of molecules: Machine learning for expediting drug discovery Computational and Structural Biotechnology Journal 19:568-576 https://doi.org/10.1016/j.csbj.2020.12.030\u003c/li\u003e\n\u003cli\u003eZhao W., Su L., Huo S., Yu Z., Li J., Liu J. (2023) Virtual screening, molecular docking and identification of umami peptides derived from Oncorhynchus mykiss Food Science and Human Wellness 12:89-93 https://doi.org/10.1016/j.fshw.2022.07.026\u003c/li\u003e\n\u003cli\u003eNowak S., Di Pizio A., Levit A., Niv M.Y., Meyerhof W., Behrens M. (2018) Reengineering the ligand sensitivity of the broadly tuned human bitter taste receptor TAS2R14 Biochimica et Biophysica Acta (BBA) - General Subjects 1862:2162-2173 https://doi.org/10.1016/j.bbagen.2018.07.009\u003c/li\u003e\n\u003cli\u003eParedes Ramos M., M L\u0026oacute;pez Vilari\u0026ntilde;o J. (2023) Hop bitterness in beer evaluated by computational analysis JOURNAL OF THE INSTITUTE OF BREWING 129 https://doi.org/10.58430/jib.v129i2.20\u003c/li\u003e\n\u003cli\u003eJumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O., Tunyasuvunakool K., Bates R., Ž\u0026iacute;dek A., Potapenko A. (2021) Highly accurate protein structure prediction with AlphaFold NATURE 596:583-589 https://doi.org/10.1038/s41586-021-03819-2\u003c/li\u003e\n\u003cli\u003eVaradi M., Anyango S., Deshpande M., Nair S., Natassia C., Yordanova G., Yuan D., Stroe O., Wood G., Laydon A. (2022) AlphaFold Protein Structure Database: massively expanding the structural coverage of protein-sequence space with high-accuracy models NUCLEIC ACIDS RESEARCH 50:D439-D444https://doi.org/10.1093/nar/gkab1061\u003c/li\u003e\n\u003cli\u003eAgyei D., Tsopmo A., Udenigwe C.C. (2018) Bioinformatics and peptidomics approaches to the discovery and analysis of food-derived bioactive peptides ANALYTICAL AND BIOANALYTICAL CHEMISTRY 410:3463-3472 https://doi.org/10.1007/s00216-018-0974-1\u003c/li\u003e\n\u003cli\u003eCory H., Passarelli S., Szeto J., Tamez M., Mattei J. (2018) The Role of Polyphenols in Human Health and Food Systems: A Mini-Review Frontiers in nutrition (Lausanne) 5:87 https://doi.org/10.3389/fnut.2018.00087\u003c/li\u003e\n\u003cli\u003eChen C., Lin L. Alkaloids in Diet, in: J. Xiao, S.D. Sarker, Y. Asakawa(2019) Handbook of Dietary Phytochemicals, Springer Singapore, Singapore:1-35\u003c/li\u003e\n\u003cli\u003eAbdelrahman M., Jogaiah S. Isolation and Characterization of Triterpenoid and Steroidal Saponins, in: M. Abdelrahman, S. Jogaiah(2020) Bioactive Molecules in Plant Defense: Saponins, Springer International Publishing, Cham:59-78\u003c/li\u003e\n\u003cli\u003eKaraman R., Nowak S., Di Pizio A., Kitaneh H., Abu-Jaish A., Meyerhof W., Niv M.Y., Behrens M. (2016) Probing the Binding Pocket of the Broadly Tuned Human Bitter Taste Receptor TAS2R14 by Chemical Modification of Cognate Agonists Chemical Biology \u0026amp; Drug Design 88:66-75 https://doi.org/10.1111/cbdd.12734\u003c/li\u003e\n\u003cli\u003eSingh N., Pydi S.P., Upadhyaya J., Chelikani P. (2011) Structural Basis of Activation of Bitter Taste Receptor T2R1 and Comparison with Class A G-protein-coupled Receptors (GPCRs) JOURNAL OF BIOLOGICAL CHEMISTRY 286:36032-36041 https://doi.org/10.1074/jbc.M111.246983\u003c/li\u003e\n\u003cli\u003ePydi S.P., Bhullar R.P., Chelikani P. (2012) Constitutively active mutant gives novel insights into the mechanism of bitter taste receptor activation JOURNAL OF NEUROCHEMISTRY 122:537-544 https://doi.org/10.1111/j.1471-4159.2012.07808.x\u003c/li\u003e\n\u003cli\u003eThomas A., Sulli C., Davidson E., Berdougo E., Phillips M., Puffer B.A., Paes C., Doranz B.J., Rucker J.B. (2017) The Bitter Taste Receptor TAS2R16 Achieves High Specificity and Accommodates Diverse Glycoside Ligands by using a Two-faced Binding Pocket Scientific Reports 7 https://doi.org/10.1038/s41598-017-07256-y\u003c/li\u003e\n\u003cli\u003eBayer S., Mayer A.I., Borgonovo G., Morini G., Di Pizio A., Bassoli A. (2021) Chemoinformatics View on Bitter Taste Receptor Agonists in Food JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 69:13916-13924 https://doi.org/10.1021/acs.jafc.1c05057 https://doi.org/10.1016/j.chroma.2019.460474\u003c/li\u003e\n\u003cli\u003eFang Y., Chen S., Lin D., Yao S. (2019) A new tetrapeptide biomimetic chromatographic resin for antibody separation with high adsorption capacity and selectivity JOURNAL OF CHROMATOGRAPHY A 1604:460474 https://doi.org/10.1016/j.chroma.2019.460474\u003c/li\u003e\n\u003cli\u003eAcevedo W., Gonz\u0026aacute;lez-Nilo F., Agosin E. (2016) Docking and Molecular Dynamics of Steviol Glycoside\u0026ndash;Human Bitter Receptor Interactions JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 64:7585-7596 https://doi.org/10.1021/acs.jafc.6b02840\u003c/li\u003e\n\u003cli\u003eCui Z., Zhang N., Zhou T., Zhou X., Meng H., Yu Y., Zhang Z., Zhang Y., Wang W., Liu Y. (2023) Conserved Sites and Recognition Mechanisms of T1R1 and T2R14 Receptors Revealed by Ensemble Docking and Molecular Descriptors and Fingerprints Combined with Machine Learning JOURNAL OF AGRICULTURAL AND FOOD CHEMISTRY 71:5630-5645 https://doi.org/10.1021/acs.jafc.3c00591\u003c/li\u003e\n\u003cli\u003eWoo J.A., Casta\u0026ntilde;o M., Goss A., Kim D., Lewandowski E.M., Chen Y., Liggett S.B. (2019) Differential long‐term regulation of TAS2R14 by structurally distinct agonists The FASEB Journal 33:12213-12225 https://doi.org/10.1096/fj.201802627RR\u003c/li\u003e\n\u003cli\u003eLevit A., Nowak S., Peters M., Wiener A., Meyerhof W., Behrens M., Niv M.Y. (2013) The bitter pill: clinical drugs that activate the human bitter taste receptor TAS2R14 The FASEB Journal 28:1181-1197 https://doi.org/10.1096/fj.13-242594.v\u003c/li\u003e\n\u003cli\u003eDe Le\u0026oacute;n G., Fr\u0026ouml;hlich E., Salar-Behzadi S. (2021) Bitter taste in silico: A review on virtual ligand screening and characterization methods for TAS2R-bitterant interactions INTERNATIONAL JOURNAL OF PHARMACEUTICS 600:120486 https://doi.org/10.1016/j.ijpharm.2021.120486\u003c/li\u003e\n\u003cli\u003eBehrens M., Redel U., Blank K., Meyerhof W. (2019) The human bitter taste receptor TAS2R7 facilitates the detection of bitter salts BIOCHEMICAL AND BIOPHYSICAL RESEARCH COMMUNICATIONS 512:877-881 https://doi.org/10.1016/j.bbrc.2019.03.139\u003c/li\u003e\n\u003cli\u003eBo W., Qin D., Zheng X., Wang Y., Ding B., Li Y., Liang G. (2022) Prediction of bitterant and sweetener using structure-taste relationship models based on an artificial neural network FOOD RESEARCH INTERNATIONAL 153:110974 https://doi.org/10.1016/j.foodres.2022.110974\u003c/li\u003e\n\u003cli\u003eJabeur S.B., Gharib C., Mefteh-Wali S., Arfi W.B. (2021) CatBoost model and artificial intelligence techniques for corporate failure prediction TECHNOLOGICAL FORECASTING AND SOCIAL CHANGE 166:120658 https://doi.org/10.1016/j.techfore.2021.120658\u003c/li\u003e\n\u003cli\u003eJun M. (2021) A comparison of a gradient boosting decision tree, random forests, and artificial neural networks to model urban land use changes: the case of the Seoul metropolitan area International journal of geographical information science : IJGIS 35:2149-2167 https://doi.org/10.1080/13658816.2021.1887490\u003c/li\u003e\n\u003cli\u003eZhou B., Bartholmai B.J., Kalra S., Osborn T., Zhang X. (2021) Lung mass density prediction using machine learning based on ultrasound surface wave elastography and pulmonary function testing JOURNAL OF THE ACOUSTICAL SOCIETY OF AMERICA 149:1318 https://doi.org/10.1121/10.0003575\u003c/li\u003e\n\u003cli\u003eMoorthy N.S.H.N., Cerqueira N.M.F.S., Ramos M.J., Fernandes P.A. (2012) QSAR and pharmacophore analysis of thiosemicarbazone derivatives as ribonucleotide reductase inhibitors MEDICINAL CHEMISTRY RESEARCH 21:739-746 https://doi.org/10.1007/s00044-011-9580-x\u003c/li\u003e\n\u003cli\u003eBallester P.J. (2019) Selecting machine-learning scoring functions for structure-based virtual screening Drug Discovery Today: Technologies 32-33:81-87 https://doi.org/10.1016/j.ddtec.2020.09.001\u003c/li\u003e\n\u003c/ol\u003e"}],"fulltextSource":"","fullText":"","funders":[],"hasAdminPriorityOnWorkflow":false,"hasManuscriptDocX":true,"hasOptedInToPreprint":true,"hasPassedJournalQc":"","hasAnyPriority":false,"hideJournal":true,"highlight":"","institution":"","isAcceptedByJournal":false,"isAuthorSuppliedPdf":false,"isDeskRejected":"","isHiddenFromSearch":false,"isInQc":false,"isInWorkflow":false,"isPdf":false,"isPdfUpToDate":true,"isWithdrawnOrRetracted":false,"journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true},"keywords":"Binding Energy, Bitterness threshold, Machine learning, Molecular docking, TAS2R14","lastPublishedDoi":"10.21203/rs.3.rs-4439031/v1","lastPublishedDoiUrl":"https://doi.org/10.21203/rs.3.rs-4439031/v1","license":{"name":"CC BY 4.0","url":"https://creativecommons.org/licenses/by/4.0/"},"manuscriptAbstract":"\u003cp\u003eEstablishing the bitterness threshold of molecules is vital for their application in healthy foods. Although numerous studies have utilized Mathematical algorithms to identify bitter chemicals, few models can accurately forecast the bitterness threshold. This study investigates the binding mode of bitter substances to the TAS2R14 receptor, establishing the relationship between the threshold and binding energy. Subsequently, a structure-taste relationship model was constructed using random forest (RF), extreme gradient boosting (XGBoost), categorical boosting (CatBoost), and gradient boosting decision tree (GBDT) algorithms. Results showed R-squared values of 0.906, 0.889, 0.936, and 0.877, respectively, suggesting a relatively good predictive capability for the bitterness threshold. Among these models, CatBoost performed optimally. The CatBoost model was then employed to predict the bitter thresholds of 223 compounds.\u003c/p\u003e \u003cp\u003eThe model provides a precise reference for detecting the bitterness thresholds of a wide range of chemicals and dangerous substances.\u003c/p\u003e","manuscriptTitle":"Machine-learning-based bitter taste threshold prediction model for bitter substances: fusing molecular docking binding energy with molecular descriptor features","msid":"","msnumber":"","nonDraftVersions":[{"code":1,"date":"2024-05-30 17:02:06","doi":"10.21203/rs.3.rs-4439031/v1","editorialEvents":[{"type":"communityComments","content":0}],"status":"published","journal":{"display":true,"email":"
[email protected]","identity":"researchsquare","isNatureJournal":false,"hasQc":true,"allowDirectSubmit":true,"externalIdentity":"","sideBox":"","snPcode":"","submissionUrl":"/submission","title":"Research Square","twitterHandle":"researchsquare","acdcEnabled":true,"dfaEnabled":false,"editorialSystem":"","reportingPortfolio":"","inReviewEnabled":false,"inReviewRevisionsEnabled":true}}],"origin":"","ownerIdentity":"5c5c10d9-9b46-49a9-8ac9-5b47fab5b6b4","owner":[],"postedDate":"May 30th, 2024","published":true,"recentEditorialEvents":[],"rejectedJournal":[],"revision":"","amendment":"","status":"posted","subjectAreas":[],"tags":[],"updatedAt":"2024-06-01T04:23:31+00:00","versionOfRecord":[],"versionCreatedAt":"2024-05-30 17:02:06","video":"","vorDoi":"","vorDoiUrl":"","workflowStages":[]},"version":"v1","identity":"rs-4439031","journalConfig":"researchsquare"},"__N_SSP":true},"page":"/article/[identity]/[[...version]]","query":{"redirect":"/article/rs-4439031","identity":"rs-4439031","version":["v1"]},"buildId":"WrCJVZZCHTDjtuVLN7oU0","isFallback":false,"isExperimentalCompile":false,"dynamicIds":[84888],"gssp":true,"scriptLoader":[]}
Text is read by the "Ask this paper" AI Q&A widget below.
Extraction quality varies by source — PMC NXML preserves structure
cleanly, OA-HTML may include some navigation residue, and OA-PDF can
have broken hyphenation. The publisher copy
(via DOI)
is the canonical version.